Voice Configuration
How your agent sounds defines the caller experience. thinnestAI gives you full control over text-to-speech, speech-to-text, interruption behavior, and conversational timing. This guide covers every setting you can tune.The Tasks tab (Tasks & Task Groups) is deprecated and hidden — it’s
superseded by Voice Workflows, which collect structured
data with Collector/Conversation nodes and add branching, knowledge, and tools.
Choosing a TTS Provider
thinnestAI supports 8 text-to-speech providers — from ultra-low-latency engines for real-time voice to specialized providers for Indian languages. Pick the one that fits your use case.Sarvam
India-first voice AI. Best-in-class support for Indian languages and accents. If your agents serve Indian customers, Sarvam is the clear choice. Strengths:- 10+ Indian languages (Hindi, Tamil, Telugu, Kannada, Bengali, Marathi, Gujarati, Malayalam, Punjabi, Odia)
- Native Indian English accents
- Code-switching support (Hindi-English, Tamil-English, etc.)
- Optimized for Indian telecom networks
bulbul:v3 is the current default. It accepts three additional knobs the
older v2 doesn’t:
temperature(0.01–1.0, default0.5) — assistant-style prosody. We default to0.5;0.6is more expressive but adds prosody jitter on long turns.min_buffer_size(30–200chars, default50) — how much text to buffer before the first audio chunk emits. Lower = faster TTFA, more fragmented prosody.max_chunk_length(50–500chars, default150) — splits long sentences for streaming.
bulbul:v3 model — no
extra config needed.
Configuration:
The full v3 roster is ~25 speakers covering Hindi, Tamil, Telugu, Marathi,
Bengali, English, and code-switched accents. Pick voices in the agent
studio’s TTS picker — switching between v2 and v3 automatically filters to
compatible speakers.
Supported Indian languages:
Cartesia
Blazing-fast, high-quality voices with fine-grained emotion control. Great for expressive, dynamic conversations. Strengths:- Sub-100ms latency
- Emotion and speed control mid-sentence
- Streaming word-level timestamps
- Multilingual support
Configuration:
Deepgram TTS
Fast, affordable text-to-speech from the leaders in speech AI. Good balance of quality and cost. Strengths:- Very low latency
- Simple API
- Competitive pricing
- Good for high-volume use cases
Configuration:
ElevenLabs
The most natural-sounding voices with ultra-realistic intonation. Best for customer-facing agents where voice quality is the top priority. Strengths:- Ultra-realistic voices with natural emotion
- Voice cloning (use your own voice)
- 29+ languages
- Extensive voice library
Configuration:
Rime
High-quality, low-latency voices with a focus on accuracy and consistency. Strengths:- Consistent voice quality across calls
- Good pronunciation accuracy
- Fast inference
- Simple integration
Configuration:
Inworld
Specialized for character-driven, immersive voice experiences. Great for agents with strong personas. Strengths:- Character-consistent voices
- Emotional expressiveness
- Real-time voice modulation
- Persona-driven design
Configuration:
Google Cloud TTS
Wide language support with consistent quality. Best for multilingual deployments. Strengths:- 220+ voices across 40+ languages
- Neural2 and Studio voices for premium quality
- SSML support
- Predictable pricing
Configuration:
Azure Speech
Enterprise-grade with fine-grained control. Best for regulated industries. Strengths:- High-quality neural voices
- Custom neural voice training
- SSML with extensive control
- Strong enterprise compliance
Configuration:
Provider Comparison
| Sarvam | ~100ms | 10+ Indian | Yes (best) | Indian market |
| Cartesia (recommended) | ~50-100ms | 10+ | No | Expressive conversations |
| Deepgram | ~100ms | English | No | High-volume, cost-effective |
| ElevenLabs | ~100-200ms | 29+ | No | Premium voice quality |
| Rime | ~120ms | English+ | No | Consistent, reliable |
| Inworld | ~130ms | English+ | No | Character-driven agents |
| Google | ~150ms | 40+ | Hindi, Bengali, Tamil | Multilingual |
| Azure | ~150ms | 60+ | Hindi, Tamil, Telugu | Enterprise, compliance |
Voice Selection and Customization
Selecting a Voice in the Dashboard
- Open your voice agent configuration.
- Go to the Voice tab.
- Select a TTS provider.
- Browse available voices and click Preview to hear samples.
- Adjust settings (speed, pitch, stability) with the sliders.
- Click Save.
Voice Selection via API
Matching Voice to Use Case
Speech-to-Text (STT) Configuration
Speech-to-text converts the caller’s audio into text for your AI agent to process. thinnestAI supports multiple STT providers with native support for Indian and global languages.Deepgram (Default)
Industry-leading speech recognition with the best accuracy and lowest latency. Our default and recommended STT provider. Models:Flux models use Deepgram’s v2 streaming API with native turn detection — the model itself decides when the user has stopped speaking, eliminating the need for a separate VAD silence timer. Pair with Turn Detection Mode → STT Endpointing for the lowest end-to-end latency (~400ms total user-end-of-speech to agent reply).When Flux is selected, the
interim_results, punctuate, and language UI settings are managed by Flux internally and are not forwarded to the STT plugin.Sarvam STT
Best-in-class for Indian languages. If your callers speak Hindi, Tamil, Telugu, or other Indian languages, Sarvam delivers the highest accuracy. Models:Google Speech-to-Text
Strong multilingual support with good accuracy across 125+ languages. Includes a built-in denoiser that filters background noise for cleaner transcriptions in noisy environments like call centers or outdoor calls. Models:"denoiser": true to enable automatic background noise filtering. This is especially useful for phone call audio where ambient noise can degrade transcription accuracy.
Azure Speech-to-Text
Enterprise-grade with custom model training. Best for regulated industries. Models:Cartesia STT
Real-time transcription with word-level timestamps.AssemblyAI
High-accuracy transcription with built-in speaker diarization and entity detection. Models:STT Provider Comparison
| Deepgram (default) | 30+ | Hindi | No | General voice agents |
| Sarvam | 10+ Indian | 10+ Indian | No | Indian languages |
| Google | 125+ | Hindi, Tamil, Telugu+ | No | Multilingual |
| Azure | 100+ | Hindi, Tamil, Telugu+ | No | Enterprise, custom models |
| Cartesia | 10+ | No | No | Timestamps, real-time |
| AssemblyAI | 20+ | No | No | Maximum accuracy |
Multi-Language Support
For agents that handle multiple languages:Turn Detection Mode
Turn detection decides when the user has stopped speaking so the agent can take its turn. Three modes available, each with different latency and accuracy trade-offs.
Ensemble is the default for new agents. Rather than relying on a single signal, it reads what the caller said alongside how long they paused — so it waits through a natural mid-sentence pause (“I’d like to know about…”) and ends the turn promptly on a complete thought (“…your pricing”).
Your choice is respected
Turn Detection Mode is exactly what you set on the Detection tab. Choosing a speech-to-text provider no longer overrides it — the Ensemble model layers on top of any provider, so there’s no need to force a single-signal mode for Sarvam or Deepgram Flux. If you want the lowest possible latency and your provider has a strong native end-of-utterance signal (Deepgram Flux, Sarvam Saaras), pick STT Endpointing explicitly.Configuration via API
STT/TTS Fallback
For production voice agents, you can configure fallback providers that automatically take over if the primary provider fails or times out. This ensures your agent stays responsive even during provider outages.How Fallback Works
- The agent sends audio to the primary provider.
- If the primary provider fails (timeout, error, or degraded quality), the system automatically switches to the fallback provider.
- The switch is seamless — callers experience a brief pause at most, not a dropped call.
Configuring TTS Fallback
Configuring STT Fallback
Recommended Fallback Pairs
Tip: Choose fallback providers that are hosted on different infrastructure than your primary to maximize resilience. For example, pairing a third-party provider with Google Cloud provides good redundancy.
Noise Cancellation
Noise cancellation runs on the audio coming into the agent (the caller’s microphone) before it reaches STT — cleaner input means more accurate transcripts and fewer false interruptions on noisy phone lines. Pick from the Voice → Noise Cancellation dropdown in the agent studio:
Both options are self-hosted (CPU), so they work the same on every
deployment.
None is honored at runtime — pre-fix, an explicit “off”
silently fell back to DTLN.
Hush — install notes
Hush requires a separate model bundle (DeepFilterNet3 weights + the auxiliary speaker-separation ONNX) at$HUSH_MODEL_DIR (default
/app/hush_model). Folder structure expected by the wrapper:
huggingface.co/weya-ai/hush and place under that path
before starting the agent worker. If the bundle is missing, the loader
logs a warning and falls back to DTLN — your agent still gets noise
cancellation, just not Hush-grade.
Configuration via API
dtln | hush | none.
Voice Style Preamble
Every voice agent automatically gets a short “you are speaking, not writing” preamble prepended to its system prompt at runtime. This steers even chat-tuned models (GPT-4o, Sarvam-30B, etc.) toward phone-friendly output:- 1–2 short sentences per reply (~30 words)
- Use contractions (“I’m”, “you’re”)
- Never use markdown, bullets, headings, code blocks, or symbols
- Spell out currency, dates, units (“twenty dollars”, not “$20”)
- Don’t read URLs or email addresses character by character
- Ask one question at a time
**bold**,
bullet lists, or $10 as “dollar sign one zero” on calls. The preamble
catches all of that.
Editing the preamble
Open the agent’s Behavior panel — the Voice Style Preamble card sits at the top of the right column. The textarea is pre-populated with the built-in default text; edit any line to customize for your domain (e.g., legal/medical disclaimer scripts), or click Reset to default to revert. Toggle the switch off if your system prompt already encodes voice style constraints — the preamble will be skipped at runtime.Configuration via API
enabled: false— skip the preamble entirelytext: ""(or omitted) — use the built-in defaulttext: "<custom text>"— use your text instead of the default
Interruption Handling
Interruption handling determines what happens when the caller speaks while the agent is talking.Modes
Allow interruptions (default): The agent stops speaking when the caller starts talking. This feels natural — like a real conversation.Sensitivity Levels
Per-Message Interruption Control
You can control interruption behavior per message in your agent’s system prompt:Silent Timeout Settings
Silent timeouts control what happens when the caller stops speaking.Recommended Timeout Settings
Wait Before Speaking
A small pause between user end-of-turn and the agent starting to speak can make responses feel more human. Counterintuitively, instant replies often sound robotic — a 300-500ms “thinking beat” reads as natural conversation. Default is0 (instant reply, lowest latency); dial up if you want the conversational feel.
When to tune
Configuration via API
on_user_turn_completed_delay in EOU metrics — visible in production logs for verification.
With preemptive generation enabled (default), the LLM call starts on interim transcripts before the user finishes. The
wait_before_speak delay applies between EOU and TTS start — by then the LLM has often already finished, so the wait is the only added latency. Total perceived gap: ~EOU + max(LLM TTFT, wait_before_speak) + TTS TTFB.Voice Formatting and Pronunciation
Pronunciation Overrides
If your agent mispronounces specific words (brand names, technical terms), add pronunciation overrides:Number and Date Formatting
Control how the agent reads numbers and dates:Filler Words and Pauses
Make your agent sound more natural by eliminating awkward silence. thinnestAI supports two types of fillers:Instant Filler Words
Short, natural sounds that play immediately (~200 ms after the user stops speaking) — before the LLM has even produced a token. The real reply slides in behind the filler as it arrives. This masks the LLM + KB latency that would otherwise be perceived as a 1-3 second silence. Where to find it: Voice Configuration → Advanced tab → Filler Words card. Toggle defaults to on for new agents.fillerWordsEnabled— Master toggle. Defaulttruefor new agents.fillerWords— Custom filler words. Leave empty to use language-appropriate defaults:- English —
hmm,uh huh,um,right,okay - Hindi —
hmm,acchha,haan,theek hai,ji - Other —
hmm,aha,um,okay
- English —
fillerWordsMinChars— Skip the filler when the user’s transcript is shorter than this. Default10. Stops “yes” / “no” replies from getting an unnecessary “hmm” in front.fillerMinTtftMs— Collision guard. Only fire the filler if the previous turn’s LLM TTFT was at least this slow (in ms). Default800. On a fast turn (TTFT < 800 ms) the real reply would arrive before the filler audio finishes, causing audio overlap even with cross-fade enabled. The first turn always probes (no prior TTFT data to gate on).
session.say() with allow_interruptions=true and add_to_chat_ctx=false — when the real LLM reply arrives, it preempts the filler instead of queueing behind it, and the filler text never reaches the model’s conversation history.
Filler Words apply in both Cascaded and Speech-to-Speech modes. In S2S, the filler is spoken by the cascaded TTS path (it’s a pre-LLM cue, not part of the realtime model’s stream).
Thinking Phrases
Longer phrases spoken when the agent needs time to process a tool call or complex request:Turn Detection
How the agent decides the caller has stopped talking. Three modes are exposed under Detection in the voice panel.
The semantic model runs locally inside the bot image — no per-call network round-trip, no third-party API. It falls back automatically to LiveKit’s bundled multilingual model, and then to single-method VAD/STT, if the upgraded model can’t load.
Adaptive EOU Confidence
When you pick Ensemble, an extra slider appears: Adaptive EOU Confidence. This is the threshold the semantic model uses to decide whether the caller has actually finished.- Higher (0.85–0.95) — waits for more certainty before ending the turn. Fewer mid-thought cut-offs; the agent is more patient. Good for hesitant callers, technical Q&A, and IVRs that read out long IDs.
- Lower (0.40–0.60) — snappier. The agent responds faster but is more likely to interrupt a thought-pause. Good for terse, transactional chitchat.
- 0.70 (default) is balanced and works for most agents.
Smart Interruption
Under Advanced → Interruption Handling, the Smart interruption toggle uses the same semantic model to recognise filler words and incomplete utterances on the fly. When the caller says “uh huh”, “yeah okay”, or trails off mid-word, the agent ignores it and keeps talking — instead of cutting itself off. Strictly suppresses false interruptions; it never causes the agent to miss a real one. On by default. No-ops gracefully if the semantic model isn’t loaded (the bot falls back to the static never-interrupt phrase list).Proactive Re-engagement
Under Detection → Proactive Re-engagement, an opt-in toggle for one of the most asked-for behaviours in long-form support and onboarding calls. When the caller pauses mid-sentence after a semantically incomplete utterance — “My ticket number is…” (silence) — the agent fires a short, language-aware nudge (“I’m listening, go ahead”) in the caller’s own language, instead of either prematurely answering or sitting silent until the 30-second abandonment poke. Configure:- Pause delay (1.5–8.0 s) — how long the caller can pause mid-thought before the agent nudges. Default 3.0 s.
- Max nudges per pause (1–3) — cap so the agent doesn’t keep prodding. Default 1.
- Explicit phrases (optional) — leave empty for an LLM-generated nudge in the caller’s language. Adding phrases here locks the wording and is not auto-translated.
Testing Your Configuration
After making changes, test thoroughly:- Use the web call test — Click Test Call in the dashboard to hear your changes immediately.
- Test edge cases — Try interrupting, staying silent, speaking quickly, and using unusual words.
- Test on a real phone — Web call audio quality differs from phone audio. Always test over a real phone line before going live.
- Compare providers — Try the same conversation with different TTS providers to find the best fit.
- Get feedback — Have someone unfamiliar with the system test the call and provide honest feedback.
Next Steps
- Inbound Calls — Apply your voice configuration to inbound call handling
- Outbound Calls — Use your configured voice for outbound campaigns
- Call Recording — Record calls to review voice quality over time

