Our voice model, outside the platform
Call the same text-to-speech and speech recognition engine that powers Proto’s AI agents – directly from your own application, on prepaid credits. No plan required – available on Free accounts.
Two endpoints
Both authenticate with a bearer token scoped to your teamspace, and both draw on the same voice-credit balance.
Text-to-speech
Synthesis built from native recordings, not an English voice bent towards a local accent.
- Native speaker recordings, not converted English models
- Voice, gender and accent selectable per language, speed 0.25–4.0
- Up to 5,000 characters a request, returned base64-encoded as MP3 or WAV
Speech recognition
Transcription built for regional accents and sentences that move between a local language and English mid-clause.
- Tuned for regional accents and mixed-language phrasing
- Transcribes up to 15MB of MP3 or WAV audio a request
- Transcript returned with an optional English translation
Call both endpoints from the browser
Type a phrase to hear it spoken, or send a recording to be transcribed – the same requests your own code would make.
Volume pricing, on your balance
The rate steps down as usage grows — size your own volume below rather than reading one headline number that only holds for the first 10,000 credits.
Its own balance
Voice requests draw on prepaid voice credits, not on your platform interaction volume. The two balances are billed and tracked separately.
1 credit a request
Every successful text-to-speech or speech-recognition call costs 1 credit, at any tier. A failed request is not charged.
50 credits free
Every new workspace starts with 50 voice credits, no plan or purchase required to try it.
No plan required
Voice credits can be bought on any plan, from Free upward — there is no tier to reach first.
This covers Voice API credits only – the Proto platform itself (AI agents, users and teams, channels, languages, teamspaces, analytics) is billed separately and isn’t included in the numbers above.
What comes with every credit
Customised voice model for 8 languages
General voice models handle these languages poorly, if at all – so Proto trained its own.
- ASR accuracy
- 93.73%
- Voice quality (MOS)
- 4.00
- ASR accuracy
- 90.66%
- Voice quality (MOS)
- 4.00
- ASR accuracy
- 87.74%
- Voice quality (MOS)
- 3.80
- ASR accuracy
- 70.22%
- Voice quality (MOS)
- 3.60
- ASR accuracy
- 79.40%
- Voice quality (MOS)
- 4.15
- ASR accuracy
- 51.50%
- Voice quality (MOS)
- Not yet published
- ASR accuracy
- Not yet published
- Voice quality (MOS)
- Not yet published
- ASR accuracy
- Not yet published
- Voice quality (MOS)
- Not yet published
Open any language above to see how it was measured against its own published target – including the ones still short of it.
Start with 50 free credits
No plan required. Get API keys and make your first request in minutes.
15‑minute live demo





