@twilio/video-node-sdk - v1.0.0-rc.3
    Preparing search index...

    @twilio/video-node-sdk - v1.0.0-rc.3

    @twilio/video-node-sdk

    CI

    Server-side Node.js SDK for Twilio Video Group Rooms with raw media frame access. Built on a native C++ addon over WebRTC, it lets you push and receive decoded video and audio frames from Node.js on realtime.

    API reference

    Note: this is a beta release of the Twilio Media SDK for Node.js. It is provided for evaluation purposes only and should not be used with production traffic. During the beta period this SDK is not HIPAA eligible.

    npm install @twilio/video-node-sdk
    

    The native binary is prebuilt and bundled — no build step required.

    Requirements:

    • Node.js >= 24.0.0
    • Linux x86-64, glibc >= 2.34 (Ubuntu 22.04+, Debian 12+), or
    • macOS 26+ on x86-64 (Intel, or Apple Silicon with Node running under Rosetta)

    Alpine/musl, native arm64 and Windows are not supported. See Platform Support for the full requirements.

    The connect() function takes a standard Twilio Video Access Token with a VideoGrant, the same token format used by the JavaScript SDK. Generate one using the twilio helper library. See User Identity and Access Tokens for details.

    const { connect, createLocalVideoTrack } = require('@twilio/video-node-sdk');

    async function main() {
    const videoTrack = createLocalVideoTrack('my-camera');

    const room = await connect(token, {
    name: 'my-room',
    videoTracks: [videoTrack],
    });

    console.log('Connected:', room.name, room.sid);

    // Push I420 video frames. connect() resolves only once the Room is connected;
    // frames written before that are dropped and counted in getWriteStats().
    videoTrack.write({
    format: 'I420',
    width: 1280,
    height: 720,
    y: { data: yPlane, stride: 1280, width: 1280, height: 720 },
    u: { data: uPlane, stride: 640, width: 640, height: 360 },
    v: { data: vPlane, stride: 640, width: 640, height: 360 },
    });

    async function trackSubscribed(track) {
    if (track.kind !== 'video') return;
    // Awaiting each frame is the backpressure. The loop ends by itself when the
    // track is unsubscribed or the Room disconnects.
    for await (const frame of track.frames()) {
    console.log(`${frame.width}x${frame.height} @ ${frame.timestamp}us`);
    frame.close?.();
    }
    }

    function participantConnected(participant) {
    participant.on('trackSubscribed', trackSubscribed);

    participant.tracks.forEach(publication => {
    if (publication.isSubscribed) {
    trackSubscribed(publication.track);
    }
    });
    }

    // participantConnected does not fire for participants already in the Room, and a
    // track can finish subscribing before this listener is attached. Seed from
    // room.participants and check isSubscribed on the publications found there.
    room.participants.forEach(participantConnected);
    room.on('participantConnected', participantConnected);

    room.on('disconnected', () => {
    room.dispose();
    });
    }

    main().catch(err => {
    console.error('Error:', err);
    process.exit(1);
    });

    Call room.dispose() when you are done with a Room. Until you do, the process does not exit on its own. disconnect() leaves the session but does not release the native resources behind it.

    The disconnected event is the clearest place to dispose, as in the quick start.

    For what each object owns, what teardown releases, and the ordering the SDK guarantees during teardown, see LIFECYCLE.md.

    This SDK shares the same Room/Participant/Track model and event names as the Twilio Video JavaScript SDK, but is designed for server-side media processing rather than browser-based conferencing. Key differences:

    • No device capture. createLocalVideoTrack() and createLocalAudioTrack() return pushable tracks with no media constraints. You supply raw frames via track.write() instead of capturing from a camera or microphone.

    • No rendering. There is no track.attach(element). Remote media arrives as raw decoded frames (I420 video, PCM audio) through for await (const frame of track.frames()).

    • Fixed audio input format. LocalAudioTrack.write() accepts only 48 kHz mono S16LE PCM. Received audio may vary in sample rate and channel count.

    • No adaptive simulcast or track priority. Published video tracks always use standard priority with simulcast disabled; the deprecated TrackPriority API is not exposed. Bandwidth and remote render-size hints are configurable instead via bandwidthProfile, encodingParameters, and RemoteVideoTrack.setContentPreferences().

    • Synchronous track creation. createLocalVideoTrack() returns a track immediately (no async device permissions). Tracks can be passed to connect() or published later via localParticipant.publishTrack().

    Function Description
    connect(token, options?) Connect to a room. Returns Promise<Room>.
    createLocalVideoTrack(name?) Create a pushable local video track.
    createLocalAudioTrack(name?) Create a pushable local audio track.
    createLocalDataTrack(name | options?) Create a local data track. Options are name, ordered, and one of maxPacketLifeTime/maxRetransmits.
    createLocalTracks(options?) Create local audio and/or video tracks. With no options, returns both. If either audio or video is specified, the other defaults to false. Each key accepts true/false or a per-track options object (e.g. { name }). Returns Promise<(LocalAudioTrack | LocalVideoTrack)[]>.
    twilioErrorFromCode(code, message?) Build a TwilioError (or matching subclass) from a numeric error code.
    setLogLevel(level) Set native log level. Accepts a name ('off' | 'fatal' | 'error' | 'warning' | 'info' | 'debug' | 'trace' | 'all') or the equivalent number 0 (off) through 7 (all).
    getVersion() Returns the native SDK version string.
    MAX_QUEUE_CEILING Upper bound (1024) accepted for any maxQueue, on frames() and on source.maxQueue.
    SDK_LOCAL_CODE The code (0) carried by errors the SDK raises locally, which Twilio never assigns a code to. Match on the error class instead.
    Export Description
    Room A connected video room. Emits events, exposes participants.
    LocalParticipant The local participant. Publish/unpublish tracks.
    RemoteParticipant A remote participant. Emits trackSubscribed/trackUnsubscribed.
    LocalVideoTrack Pushable video track (write(frame)).
    LocalAudioTrack Pushable audio track (write(frame), clearBuffer()).
    LocalDataTrack Send arbitrary data (send). Create via createLocalDataTrack(name | options?).
    RemoteVideoTrack Receive video frames (frames()), getStats(), frameDropped event.
    RemoteAudioTrack Receive audio frames (frames()), getStats(), frameDropped event.
    RemoteDataTrack Receive data messages (message event).
    TrackPublication Base class for published tracks (trackSid, trackName, kind, isTrackEnabled).
    LocalTrackPublication Local publication. Exposes track and unpublish(). Subclassed per kind (LocalVideoTrackPublication, …).
    RemoteTrackPublication Remote publication. Exposes track and isSubscribed. Subclassed per kind (RemoteVideoTrackPublication, …).
    TwilioError Base error class, carrying a numeric code. One subclass per known Twilio code, plus SDK-local errors (NativeBindingLoadError, RoomConnectTimeoutError, …).
    ErrorCode Enum of Twilio Video error codes.

    participantConnected is not emitted for participants who were already in the Room when connect() resolved. They are part of the Room's starting state: read them from room.participants.

    A participant who was already publishing emits trackSubscribed after connect() resolves. Subscriptions that completed before the listener was attached are not replayed, and appear in participant.tracks with isSubscribed set to true.

    Event Handler Signature
    disconnected (room: Room, error?: TwilioError) => void
    connectFailure (error: TwilioError) => void
    reconnecting (error?: TwilioError) => void
    reconnected () => void
    participantConnected (participant: RemoteParticipant) => void
    participantDisconnected (participant: RemoteParticipant) => void
    participantReconnecting (participant: RemoteParticipant) => void
    participantReconnected (participant: RemoteParticipant) => void
    recordingStarted () => void
    recordingStopped () => void
    dominantSpeakerChanged (participant: RemoteParticipant | null) => void
    transcription (transcriptionJson: string) => void

    The Room re-emits every track event in RemoteParticipant Events, appending the RemoteParticipant that emitted it as the last argument. Handle every participant's tracks from one place instead of attaching a listener to each participant.

    Event Handler Signature
    trackSubscribed (track: RemoteVideoTrack | RemoteAudioTrack | RemoteDataTrack, publication: RemoteTrackPublication) => void
    trackUnsubscribed (track: RemoteVideoTrack | RemoteAudioTrack | RemoteDataTrack, publication: RemoteTrackPublication) => void
    trackSubscriptionFailed (error: TwilioError, publication: RemoteTrackPublication) => void
    trackPublished (publication: RemoteTrackPublication) => void
    trackUnpublished (publication: RemoteTrackPublication) => void
    trackEnabled (publication: RemoteTrackPublication) => void
    trackDisabled (publication: RemoteTrackPublication) => void
    videoTrackSwitchedOff (track: RemoteVideoTrack) => void
    videoTrackSwitchedOn (track: RemoteVideoTrack) => void
    networkQualityLevelChanged (level: number) => void
    Event Handler Signature
    trackPublished (publication: LocalTrackPublication) => void
    trackPublicationFailed (error: TwilioError, localTrack?: LocalTrack) => void
    networkQualityLevelChanged (level: number) => void

    LocalTrackPublication exposes track (the local track instance) and an unpublish() method:

    const pub = room.localParticipant.tracks.get(trackSid); // LocalTrackPublication
    pub.unpublish(); // unpublishes the underlying track

    RemoteTrackPublication exposes track (the subscribed remote track, if any) and isSubscribed.

    Returns a snapshot of WebRTC stats per peer connection. Rejects if the room is disconnected.

    const reports = await room.getStats();
    // reports[i]: {
    // peerConnectionId, localAudioTrackStats, localVideoTrackStats,
    // remoteAudioTrackStats, remoteVideoTrackStats
    // }

    Push raw I420 video frames into a room. Frames written before connect() resolves are dropped, and counted in getWriteStats().

    const track = createLocalVideoTrack('camera');
    track.write({
    format: 'I420', // optional; only 'I420' is accepted
    width,
    height, // both must be positive and even
    y: { data: yBuffer, stride: yStride, width, height },
    u: { data: uBuffer, stride: uStride, width: width / 2, height: height / 2 },
    v: { data: vBuffer, stride: vStride, width: width / 2, height: height / 2 },
    timestamp, // optional microseconds; defaults to monotonic now
    rotation, // optional 0 | 90 | 180 | 270
    });
    track.enabled = false; // mute

    Buffers are copied synchronously, so they can be reused as soon as write() returns.

    write() returns false when the frame was dropped rather than encoded - most often because the encoder sink has not attached yet, but also when libwebrtc's adapter rate-limits or rejects the resolution. It throws TypeError/RangeError on invalid input.

    Optionally pin the frame size at creation, so a mismatched frame is rejected instead of silently rescaled:

    const track = createLocalVideoTrack({
    name: 'camera',
    source: { type: 'raw', format: 'I420', width: 1280, height: 720, fps: 30 },
    });

    Publish-side counters:

    const { framesWritten, framesDropped, sendQueueDepth, maxQueue, lastTimestamp } =
    track.getWriteStats();

    Video publish is synchronous - write() hands the frame straight to the encoder - so there is no SDK-side send queue and sendQueueDepth/maxQueue are always 0. A framesDropped here means the frame was rejected, not shed from a queue.

    Push raw PCM audio samples into a room. Format is fixed to 48kHz mono S16LE.

    const track = createLocalAudioTrack('mic');
    const accepted = track.write({
    pcm, // Buffer of interleaved int16 samples
    frames, // samples per channel, e.g. 480 for a 10ms chunk
    timestamp, // optional microseconds
    });

    Unlike video, audio publish has a real send queue, drained one 10 ms chunk at a time. It is bounded (~500 ms by default) so a producer running faster than real time cannot accumulate latency. A write() is accepted only if it fits whole in the remaining space; otherwise nothing is buffered, write() returns false, and the rejection is counted. A single write() larger than the bound can never fit, so size maxQueue to the largest burst you intend to publish:

    const track = createLocalAudioTrack({
    name: 'mic',
    // maxQueue is in 10ms chunks: 200 => ~2s of smoothing. It binds the
    // process-wide audio device, so it applies to every local audio track.
    source: { type: 'raw', format: 'PCM_S16LE', sampleRate: 48000, channels: 1, maxQueue: 200 },
    });
    const { framesWritten, framesDropped, sendQueueDepth, maxQueue } = track.getWriteStats();

    Publishing at real-time cadence should never drop. A non-zero framesDropped means the producer is outrunning the wire, or that a single write() was larger than maxQueue.

    The timestamp on an audio frame is observability-only. Audio publish is FIFO: the device emits queued samples on its own 10 ms cadence, so the timestamp feeds lastTimestamp and timestampRegressions and does not change what is sent.

    clearBuffer() discards whatever is still queued and not yet sent. Use it when the queued audio has become stale rather than merely late - barge-in, where the speaker is interrupted and the rest of the utterance should never play, is the usual case. Writes after it resume from an empty queue.

    track.clearBuffer();
    

    Send arbitrary string or binary messages. Delivery is reliable and ordered by default.

    const track = createLocalDataTrack({ name: 'chat', ordered: true });
    room.localParticipant.publishTrack(track);
    track.send('hello');
    track.send(Buffer.from([0x01, 0x02]));

    // send() reports the outcome. The promise always resolves - never rejects - so
    // a fire-and-forget send cannot produce an unhandled rejection.
    const result = await track.send('important');
    if (!result.ok) console.warn('send failed:', result.error);

    Messages larger than 64 KB (kMaxMessageSize) are rejected synchronously with a RangeError and never transmitted.

    Pass maxPacketLifeTime (milliseconds) or maxRetransmits (a count) to trade reliability for latency. The two are mutually exclusive, and each must be an integer from 0 to 65535.

    const telemetry = createLocalDataTrack({ name: 'telemetry', maxPacketLifeTime: 500 });
    telemetry.maxPacketLifeTime; // 500
    telemetry.maxRetransmits; // null
    telemetry.reliable; // false

    maxPacketLifeTime and maxRetransmits are number | null, reading back as null when the limit was not set. reliable is true only when neither is set.

    Receive decoded I420 video frames from a remote participant.

    for await (const frame of track.frames()) {
    // frame: {
    // format: 'I420',
    // width, height,
    // y, u, v, // I420Plane: { data: Buffer, stride, width, height }
    // timestamp: number, // microseconds
    // captureTimestamp?: number,
    // rtpTimestamp?: number,
    // frameId: number, // SDK-generated, monotonic per track
    // rotation?: 0 | 90 | 180 | 270,
    // close?(): void, // optional prompt release
    // }
    frame.close?.();
    }
    // The loop ends on unsubscribe or Room disconnect. `break` releases the track.

    // Hint the desired render dimensions to the SFU. Width/height must be positive
    // integers. Only takes effect when the room was connected with a
    // bandwidthProfile that has `contentPreferencesMode: 'manual'`.
    track.setContentPreferences({ renderDimensions: { width: 320, height: 240 } });

    // `isSwitchedOff` is `true` when the SFU has stopped delivering this track
    // (e.g. due to bandwidth-profile constraints). Pair with the
    // `videoTrackSwitchedOff` / `videoTrackSwitchedOn` events on RemoteParticipant.
    track.isSwitchedOff;

    Receive decoded PCM audio frames from a remote participant.

    for await (const frame of track.frames()) {
    // frame: {
    // format: 'PCM_S16LE',
    // sampleRate, channels, frames,
    // pcm: Buffer, // interleaved int16 samples
    // timestamp: number, // microseconds
    // frameId: number, // SDK-generated, monotonic per track
    // close?(): void,
    // }
    }

    Receive string or binary messages from a remote participant.

    track.on('message', (data, track) => {
    /* data is string | Buffer; track is the RemoteDataTrack it arrived on */
    });

    maxPacketLifeTime, maxRetransmits, reliable, and ordered report how the publisher configured delivery. A publisher's limit of 65535 reads back as null, because a subscribed track reports it the same way it reports an unset limit; reliable still distinguishes the two.

    For the precise guarantees - buffer ownership, where frames are dropped, drop policy and ordering, timestamp rules, and the publish invariants - see FRAME_CONTRACT.md.

    I420 planar layout. Each plane is an I420Plane: { data: Buffer, stride, width, height }, where stride is bytes per row (≥ the plane's width, padded for alignment).

    Plane Logical size data size Description
    Y width × height y.stride × height Luminance
    U ⌈width/2⌉ × ⌈height/2⌉ u.stride × ⌈height/2⌉ Chrominance (Cb)
    V ⌈width/2⌉ × ⌈height/2⌉ v.stride × ⌈height/2⌉ Chrominance (Cr)

    Publish and receive use the same planar shape: each of y/u/v is an I420Plane ({ data, stride, width, height }). A received frame can be written straight back out without reshaping.

    Timestamps are plain numbers of microseconds (timestamp). Microsecond resolution stays exact in a JS number for roughly 285 years, and the underlying engine reports microseconds natively. rotation is 0 | 90 | 180 | 270.

    Interleaved 16-bit signed little-endian PCM in a single Buffer.

    • Inputs to LocalAudioTrack.write() are fixed at 48kHz mono — only pcm and frames are accepted.
    • Received AudioFrames include sampleRate, channels, frames, pcm, timestamp (microseconds), and frameId.
    {
    name?: string; // Room name
    videoTracks?: LocalVideoTrack[]; // Tracks to publish on connect
    audioTracks?: LocalAudioTrack[];
    dataTracks?: LocalDataTrack[];
    enableInsights?: boolean;
    enableAutomaticSubscription?: boolean;
    enableDominantSpeaker?: boolean;
    networkQuality?: boolean | { local?: 1; remote?: 0 | 1 };
    preferredAudioCodecs?: ('opus' | 'PCMU')[];
    preferredVideoCodecs?: 'VP8'[];
    videoEncodingMode?: 'auto';
    bandwidthProfile?: BandwidthProfileOptions;
    receiveTranscriptions?: boolean;
    region?: string; // e.g. 'us1', 'au1'
    iceOptions?: IceOptions;
    encodingParameters?: EncodingParameters;
    connectionTimeout?: number; // ms; default 30000, 0 waits indefinitely
    }
    {
    video?: {
    mode?: 'collaboration' | 'grid' | 'presentation';
    maxSubscriptionBitrate?: number; // bits per second
    trackSwitchOffMode?: 'detected' | 'predicted' | 'disabled';
    clientTrackSwitchOffControl?: 'auto' | 'manual';
    contentPreferencesMode?: 'auto' | 'manual';
    };
    }
    {
    maxAudioBitrate?: number; // bits per second
    maxVideoBitrate?: number; // bits per second
    }
    {
    transportPolicy?: 'all' | 'relay'; // 'relay' forces TURN
    iceServers?: IceServer[]; // { urls: string[]; username?: string; credential?: string }
    }
    • Node.js >= 24.0.0
    • OS Linux x86-64, macOS 26+ x86-64
    • Distros Ubuntu 22.04+ and Debian 12+.
    • CPU x86-64 only. There is no arm64 build, so on Apple Silicon Node must run under Rosetta.

    The prebuilt native addon is linked against glibc and requires:

    Requirement Minimum
    glibc 2.34
    libstdc++ (GLIBCXX) 3.4.30
    C++ ABI (CXXABI) 1.3.11

    It also links libX11.so.6, which WebRTC requires unconditionally. Install your distro's X11 client library (libx11-6 on Debian and Ubuntu) even on headless servers.

    On macOS the prebuilt addon requires macOS 26 or later. npm cannot check the macOS version, so on an older release npm install succeeds and the SDK fails when it loads the addon.

    Alpine and other musl-based distros are not supported: the addon is glibc-only. Windows is not supported.

    The example applications listed below demonstrate various ways to use the SDK for audio or video processing. They load credentials from a .env file at the repo root. Copy the template, fill in your credentials, and run:

    cp .env.example .env
    # edit .env: set TWILIO_ACCOUNT_SID / TWILIO_API_KEY / TWILIO_API_SECRET
    node examples/virtual_camera.js [room-name]

    .env is gitignored, so your real credentials are never committed.

    See the examples/ directory:

    Example Description
    virtual_camera.js Decodes an MP4 with ffmpeg and pushes I420 frames to a room.
    video_mirror.js Receives remote video frames and pushes them back as-is.
    audio_push.js Generates a sine wave tone and pushes PCM audio to a room.
    data_channel.js Two participants exchange string and binary messages via data tracks.
    voice_agent.js Bridges room audio to the OpenAI Realtime API for a spoken voice agent (requires OPENAI_API_KEY).
    cv_object_detection.js Runs YOLOX object detection on a participant's webcam and re-publishes the video with bounding boxes.
    cv_face_analysis.js Analyzes a participant's face — presence and an attention estimate (head orientation) — drawn on the video.

    The computer-vision examples (cv_*.js) run local ONNX models via onnxruntime-node and draw with @napi-rs/canvas. These two are large and only these examples need them, so they live in examples/package.json rather than the SDK's own dependencies — install them separately:

    npm install --prefix examples
    

    No cloud service or API key is needed: each example analyzes the first participant's video and expresses its result on a re-published video track. Run them against any room you also join from a browser, publishing your webcam.

    The ONNX model files are not shipped with the repo — download the ones you need and save them to examples/.models/. The examples print these same instructions if a model is missing.

    mkdir -p examples/.models

    # cv_object_detection.js — YOLOX-nano (~3.7 MB)
    curl -L -o examples/.models/yolox_nano.onnx \
    "https://github.com/Megvii-BaseDetection/YOLOX/releases/download/0.1.1rc0/yolox_nano.onnx"

    # cv_face_analysis.js — RTMO-t (a zip containing end2end.onnx, ~27 MB extracted)
    curl -L -o examples/.models/rtmo-t.zip \
    "https://download.openmmlab.com/mmpose/v1/projects/rtmo/onnx_sdk/rtmo-t_8xb32-600e_body7-416x416-f48f75cb_20231219.zip"
    unzip -j examples/.models/rtmo-t.zip '*end2end.onnx' -d examples/.models
    mv examples/.models/end2end.onnx examples/.models/rtmo-t.onnx

    # verify the downloads (the examples also check this on startup)
    echo "c789161ed43c8269fcd4e67c67eeeb4e80c622da2eb296a20bc6007bd18a0b7d examples/.models/yolox_nano.onnx" | shasum -a 256 -c
    echo "20aad6e2e42359cac1c5b4a0b2da00e29bfe91a72a782fdcf287d273a04c1b24 examples/.models/rtmo-t.onnx" | shasum -a 256 -c

    Both models are Apache-2.0 licensed and downloaded from their projects' official channels — YOLOX by Megvii and RTMO (OpenMMLab mmpose). Each example verifies its model's SHA-256 on startup and refuses to run a file that doesn't match.

    See LICENSE.md.