
Configurations
Voice
Voice Agents use TTS (Text to Speech), which generates audio that LLMs generate during the course of a conversation. This is the audio that the end user having the conversation listens to.
EESI platform supports ElevenLabs, OpenAI, Google, Azure Speech, Deepgram, Cartesia, Smallest AI, MiniMax, Sarvam, Rime, Inworld, Camb.ai, xAI, and EESI TTS engines. There are some voices from the providers that we ship by default. You can refer to the providers API documentation to select a voice ID that’s most relevant for your language requirement.
For locally deployed or self-hosted TTS models, EESI also supports Speaches, an OpenAI API-compatible server for speech generation.
If you don’t find your favourite voice, you can always add the voice ID manually.


⌘I