Voice and Language
Yaplet voice agents speak 17 languages with 30 different Gemini voices. This page explains how the voices are grouped, how to preview them, which language to pick, and how the agent's spoken language differs from the two other language settings it is easy to confuse it with.
The 30 Voices
Voices are grouped into seven families by tone and character, and each family contains both masculine and feminine options. Every voice speaks every supported language â picking a voice never locks you into one. The default for a new agent is Kore.
Authoritative
Six voices described as informative, knowledgeable or firm â Charon, Rasalgethi, Sadaltager, Kore, Orus and Alnilam. The steady, credible end of the range.
Bright/Upbeat
Five voices â Zephyr, Puck, Autonoe, Laomedeia and Sadachbia â described as bright, upbeat or lively.
Energetic
Three voices: Fenrir (excitable), Pulcherrima (forward) and Leda (youthful). The fastest, most forward-leaning of the set.
Casual
Five relaxed voices â Aoede, Callirrhoe, Umbriel, Zubenelgenubi and Achird â described as breezy, easy-going, casual or friendly.
Soft/Warm
Four gentler voices: Achernar (soft), Vindemiatrix (gentle), Sulafat (warm) and Enceladus (breathy).
Clear/Smooth
Five neutral, even voices â Iapetus, Erinome, Algieba, Despina and Schedar â described as clear, smooth or even.
Distinctive
Two voices with more character: Algenib (gravelly) and Gacrux (mature).
Picking a Voice
Three guidelines that work for most teams:
- Match the brand tone, not the country. A premium watch brand doesn't change voice between English and German â it changes voice between premium and friendly.
- Listen to the actual greeting. If the agent has a begin message, the preview speaks that text in the selected voice and language, so you hear exactly what your callers will hear. With no begin message set, you get a generic sample sentence in the chosen language instead.
- Run a friend test. Send the greeting recording to two people outside your company. If both say "that's clearly a robot", try a different family. If both say "that's a normal call", you're done.
The 17 Languages
| Language | Code |
|---|---|
| English (US) | en-US |
| English (UK) | en-GB |
| English (AU) | en-AU |
| English (IE) | en-IE |
| Hungarian | hu-HU |
| German | de-DE |
| French | fr-FR |
| Spanish | es-ES |
| Spanish (LATAM) | es-MX |
| Portuguese (BR) | pt-BR |
| Portuguese (PT) | pt-PT |
| Italian | it-IT |
| Dutch | nl-NL |
| Polish | pl-PL |
| Czech | cs-CZ |
| Slovak | sk-SK |
| Romanian | ro-RO |
Three Different Language Settings
"Language" means three separate things in Yaplet, and setting the wrong one is the most common mix-up in this area.
- The agent's spoken language â on the Voice & language tab. This is what the caller hears: it controls the default greeting and the language the agent replies in.
- The brand's language â at Brand â (your brand) â Brand settings â General â Language. This is the language the brand works in, and the fallback for anything you add without a language of its own.
- Each source's own language â set in place on Brand â (your brand) â Knowledge, one per knowledge base or documentation set. It decides how that content is searched; it does not translate anything.
What Happens Mid-Call
A voice agent stays in its configured language for the whole call. The instruction it runs under tells it to reply unmistakably in that language, and that holds even if the caller switches.
- Gemini Live still hears the caller in whatever language they actually speak â what is fixed is the language the agent answers in.
- The search queries the agent writes are produced in the call's language, then translated for the lookup as described above.
- If you need genuinely different languages, give each one its own agent and its own number, and let the caller pick.
Transcribing Non-English Calls
Transcription normally comes from Gemini Live itself. On non-English calls the caller's audio is additionally streamed to Deepgram in parallel, and its text is used for the transcript. This affects the written record only â the spoken side of the conversation is untouched, and there is nothing to configure.