BYOK is all three jobs or none: it turns on only when a speech-to-text key, an LLM key and a voice
key are all added and checked. Any plan can turn it on.
Each agent can also run its own model from your LLM key. On the agent’s Playground, the
Model list shows every model your key reaches. Leave it on your key’s model, which is the one set
in Settings → Your keys, or pick another for this agent alone.
Each agent speaks with a voice you choose from your own voice provider. Open the agent’s
Voice page: it lists every voice your key can use, and Play speaks a short line in the
agent’s language. The preview runs on your account, so the provider charges you for those few
characters. If a provider hosts its own sample of a voice, that sample plays instead, at no
cost. An agent without a chosen voice speaks with a standard voice from your provider.
On a call, all three run on your accounts: your speech-to-text hears the caller, your LLM answers,
and your voice speaks. If one of your keys cannot be used when a call starts, the call is not
answered on our providers instead.
In the console
Open Settings → Your keys. Each of the three cards has:- Provider. Choose who does that job.
- Model. Every model the provider offers. Before you add a key, the list shows the models the provider offers. Once your key is checked, it shows the models your key reaches. Long lists are grouped by family, such as Flux, Aura-2 and Aura for Deepgram’s voices.
- API key, plus any other field the provider needs. Azure, for example, needs a region.
Providers
Azure Speech takes your key and region. If your Azure resource accepts its key only at its own
address, as Azure AI Foundry resources do, also give its resource endpoint, such as
https://my-resource.cognitiveservices.azure.com.
Any OpenAI-compatible URL takes a base URL that starts with https://. Azure OpenAI’s
https://<resource>.openai.azure.com/openai/v1 is one. If the server does not list its models, send
the model id with the key, and we check the two together with a one-token request.
When a key fails
There is no fallback. If your provider refuses a request (a revoked key, no credit, rate limits), the reply fails. We do not answer on our own models and bill you for it. You get a notification naming the provider and the reason, at most once a day per provider.Customers
If you build on our API with customers, your customers use your keys while BYOK is on. A customer can bring a complete set of three keys of its own, which then overrides yours. Turning BYOK off in your workspace turns it off for all your customers.API
Every request below also works for a customer, with theThinnest-Workspace header.
Status
using is own, developer (a customer using its developer’s keys) or none. models is what the
provider offered this key, and model is the one it runs.
Add or replace a key
kind. One ofstt,llmortts.provider. One of the provider ids in the table above, in lower case. An error lists the ones that do that job.credentials.model. Optional. Without it, a speech key starts on the provider’s usual model. An LLM key starts on the provider’s usual model when it has one; otherwise choose one.- Replacing a key. If you replace a key with one for the same provider, its model is kept.
Change the model, or refresh the list
models list. For a model the provider
launched after you added the key, send "refreshModels": true, alone or together with the new
model.
Turn BYOK on or off
409 until all three keys are added, checked, and each has a model.
Voices
sample is a clip the provider hosts,
when it has one. On Deepgram, a voice is a model, such as aura-2-helena-en or flux-alexis-en.
Preview a voice
audio/mpeg, or audio/wav for Sarvam. Without text, the
voice speaks a short line of ours in the agent’s language, or in language if you send it. Any
text you send is capped at 200 characters, because your provider charges you for every
character. Soniox has no preview here: it speaks only on a live call.
Choose an agent’s model
GET /api/v1/models lists our models in items, as before. On a workspace using its own keys
it also returns byok:
null to go back to your key’s model. GET on the
same path returns the agent’s choice. It also returns stale, the agent’s earlier choice, if your
key no longer reaches it; until you choose again, the agent runs your key’s model.
Choose an agent’s voice
GET /byok/voices lists. Send null to go back to the provider’s
standard voice. GET on the same path returns the agent’s current choice.
To use a different voice for a single call, put it in that call’s overrides, as
"overrides": { "voice": "EXAVITQu4vr4xnSDxMaL" } on POST /api/v1/calls or in a batch.
On a workspace using its own keys, the voice must be one of yours from GET /byok/voices.
A choice belongs to the provider it was made on. If you switch your voice key to another provider,
the agent speaks with the new provider’s standard voice until you choose again. GET shows the old
choice as stale until then.
Remove a key
409 while BYOK is on. Replace the key instead, or turn BYOK off
first.
Changing keys or switching BYOK needs an API key with full access. A read-only key can see the
status.