Getting started — This guide takes a blank browser application to a connected agent that answers a question and says a scripted line.
Chat modes — The chat mode decides how the agent answers: with a streamed video, as text only, or not at all.
Client tools — A client tool is a function that runs in the user's browser when the agent's LLM decides to call it.
Expressive media — An Expressive (V4) session is two-way: the page can publish the user's microphone and camera, interrupt the agent, and send its own data-channel messages.
Handling errors — Every error the SDK raises is a BaseError carrying a kind, and isDIDError is the guard that recognizes one.
Migration guide — @d-id/client-sdk v3 is a breaking release that trims the package's public surface to what the SDK supports.
Agent Manager
createAgentManager — Creates an AgentManager for one agent: its chat, its video stream and its connections.
Agent — An agent's profile, as the Agents API returns it.
AgentAvatar — The avatar an agent speaks through: its rendering tier and the language of its voice.
AgentManager — A live connection to one agent: its profile, its chat, and the video stream it answers on.
AgentManagerOptions — Everything createAgentManager needs to reach an agent and report back to the application.
AnalyticsOptions — What the SDK reports about the session, and where it goes.
StreamType — How the agent's idle and talking video are delivered.
Speak & Scripts
AudioStreamScript — A script that makes the agent lip-sync an audio file you host, with no text-to-speech involved.
SpeakResponse — What speak() resolves with: the video the agent is about to stream.
TextStreamScript — A script that makes the agent say text you supply, synthesized by a text-to-speech provider.
SpeakScript — The script payload speak() accepts: text or audio.
Chat
parseMessageParts — Splits a message's text into the typed parts a UI can render.
ChatResponse — What the Agents API answers a chat request with.
InterruptOptions — What caused the user to interrupt the agent, passed to interrupt().
Message — One message of a chat: what the user asked, or what the agent answered.
MessageSentiment — The sentiment an agent answer was delivered with, as the server reported it.
Rating — A rating stored against one message of a chat, as the Agents API returns it.
RetrievalMetadata — One knowledge citation: the passage of the agent's knowledge base an answer was drawn from.
SubmitFeedbackResponse — What the Agents API stored when end-of-call feedback was submitted.
MessagePart — One renderable piece of a message: a run of text, an image, a video or a link.
ChatMode — How the agent answers: with a streamed video, as text only, or not at all.
Voice
AmazonTtsProvider — Amazon provider details: the provider type and the requested voice id.
AzureOpenAiTtsProvider — Azure OpenAI provider details: the provider type, the requested voice id and an optional voice_config for style, rate and pitch.
ElevenlabsTtsProvider — ElevenLabs provider details: the provider type and the requested voice id. Available to premium users.
MicrosoftTtsProvider — Microsoft Azure provider details: the provider type, the requested voice id and an optional voice_config for style, rate and pitch.
Voice — One voice in D-ID's catalog of text-to-speech voices.
VoiceConfigElevenlabs — How closely an ElevenLabs voice should follow the original it was built from.
VoiceConfigMicrosoft — How a Microsoft Azure or Azure OpenAI voice should deliver the text.
TtsProvider — The provider object a speak script accepts: any of the four text-to-speech providers.
Providers — The text-to-speech engines a voice can be served by.