D-ID Client SDK - v3.0.0-0
    Preparing search index...

    Getting started

    This guide takes a blank browser application to a connected agent that answers a question and says a scripted line. It covers the credentials the SDK needs, createAgentManager, connect(), the difference between chat() and speak(), and disconnect(). The SDK runs in the browser only: it needs WebRTC and a <video> element, and there is nothing for it to render in Node.

    1. Log in to D-ID Studio and create an agent — image, voice and knowledge.
    2. In the agents gallery, hover over the agent, open the [...] menu and click </> Embed.
    3. Set the list of domains the agent may be used from, for example http://localhost.
    4. Copy data-agent-id and data-client-key out of the snippet.

    The client key is the credential to use in a page. It is scoped to one agent and to the domains you allowed, which is what makes it safe to ship in front-end code — see ClientKeyAuth. BearerToken and BasicAuth are also accepted by Auth, but they authorize the whole account, so keep them on a server.

    npm i @d-id/client-sdk
    

    The package ships an ES module and a UMD build.

    createAgentManager fetches the agent before it resolves, so agent and starterMessages are readable straight away. Nothing is opened on the wire yet.

    The one callback a video session cannot work without is onSrcObjectReady: it hands you the MediaStream to put on your <video> element. createAgentManager() rejects with a ValidationError when it is missing in any chat mode that streams video — which is every mode except ChatMode.TextOnly, ChatMode.Playground and ChatMode.Maintenance.

    That element is yours to put on the page, and the three attributes are not optional in practice:

    <video id="agent-video" autoplay playsinline></video>
    

    autoplay is what starts playback when the stream arrives — the SDK sets srcObject and nothing else. playsinline keeps iOS Safari from taking the video fullscreen the moment it plays. And browsers block autoplay with sound until the user has interacted with the page, so either start the agent from a click, or add muted and unmute on the first click.

    import * as sdk from '@d-id/client-sdk';

    const videoElement = document.getElementById('agent-video') as HTMLVideoElement;
    const agentManager = await sdk.createAgentManager('agt_fumf1234', {
    auth: { type: 'key', clientKey: 'YOUR_CLIENT_KEY' },
    callbacks: {
    onSrcObjectReady(value) {
    videoElement.srcObject = value;
    },
    },
    });

    The whole options object is read once and never written to, callbacks included. Assigning a handler to the object afterwards has no effect, so give each handler a stable identity and read your changing state from inside it — in React, from a ref.

    connect() creates the stream, the chat and, on Talks (V2) and Clips (V3) agents, the notifications web socket. It resolves once the connection has reached 'connected', by which point onSrcObjectReady has already fired.

    await agentManager.connect();
    

    A second connect() made while the first is still in flight returns that same promise rather than opening a second session, which is what makes it safe in a React StrictMode effect. Once a session exists it rejects instead: call disconnect() first to start a fresh conversation, or reconnect() to keep the current one.

    Two methods do it, and they are not interchangeable.

    • chat() sends the user's message to the agent's LLM and the agent answers in its own words. The answer arrives through onNewMessage, first as partial chunks and then as the full answer, while the video plays it.
    • speak() makes the agent say exactly what you give it, with no LLM involved. Use it for greetings and scripted lines. It takes a SpeakScript — { type: 'text', input } or { type: 'audio', audio_url } — or a plain string as shorthand for a text script.
    await agentManager.speak({ type: 'text', input: `Hi! I'm ${agentManager.agent.name}.` });
    await agentManager.chat('What is the distance to the moon?');

    Render the transcript from onNewMessage rather than from the chat() result: on Expressive (V4) agents the answer travels over the data channel and the resolved ChatResponse is empty. Both promises resolve when the request is accepted, not when the agent has finished speaking — so nothing that waits on them should tear the session down.

    Call disconnect() when the user leaves the conversation or the page, so the session stops consuming credits. After it resolves the manager is reusable — connect() opens a new session with a new chat.

    await agentManager.disconnect();
    

    Leaving the page is the case to wire up first, since nothing else ends the session for you:

    window.addEventListener('beforeunload', () => void agentManager.disconnect());
    
    import * as sdk from '@d-id/client-sdk';

    const videoElement = document.getElementById('agent-video') as HTMLVideoElement;
    let srcObject: MediaStream | undefined;

    const agentManager = await sdk.createAgentManager('agt_fumf1234', {
    auth: { type: 'key', clientKey: 'YOUR_CLIENT_KEY' },
    // Talks (V2) and Clips (V3) agents only; Expressive (V4) agents ignore these. This is
    // the default — see StreamOptions for the codec, session timeout and fluent knobs.
    streamOptions: { compatibilityMode: 'auto' },
    callbacks: {
    onSrcObjectReady(value) {
    srcObject = value;
    videoElement.srcObject = value;
    },
    onVideoStateChange(state) {
    // Legacy streams swap between the agent's idle video and the live stream.
    if (state === 'STOP') {
    videoElement.srcObject = null;
    videoElement.src = agentManager.agent.idle_video ?? '';
    } else {
    videoElement.src = '';
    videoElement.srcObject = srcObject ?? null;
    }
    },
    onConnectionStateChange(state, reason) {
    console.log('connection', state, reason);
    },
    onNewMessage(messages, type) {
    if (type === 'answer') {
    console.log(messages[messages.length - 1].content);
    }
    },
    onError(error) {
    console.error(error);
    },
    },
    });

    await agentManager.connect();
    await agentManager.speak({ type: 'text', input: `Hi! I'm ${agentManager.agent.name}.` });
    await agentManager.chat('What is the distance to the moon?');

    // Both calls above resolve when the request is accepted, not when the agent has spoken:
    // the answer and the video arrive afterwards, through onNewMessage and the video element.
    // So the session ends when the user leaves, not here.
    window.addEventListener('beforeunload', () => void agentManager.disconnect());

    onVideoStateChange is the right signal to swap video sources on a legacy stream — a Talks (V2) agent, or a Clips (V3) agent that did not ask for fluent. A fluent stream sends one video for both states, and every Expressive (V4) session is fluent, so there is nothing to swap; branch on getStreamType() if the same code has to serve both.

    • The four chat modes an application chooses between, and what each one creates on the wire: Chat modes.
    • Running your own functions when the agent's LLM asks for them: Client tools.
    • Microphone, camera, speech-to-text and interrupts on Expressive (V4) agents: Expressive media.
    • Which failures reject and which reach onError: Handling errors.