OUR VISION 001 / LAUNCH
Voice AI needs to work in the languages people actually speak.
Why we’re building QuantClarity: intelligent routing for voice models, grounded in the languages, conditions, and use-cases that matter to real people.
Voice AI will be a huge component of how we interact with computers in the future. It is a natural interface: you speak, the computer understands, and something gets done. You should not need to learn a new interface every time you want to use a new service.
But whether that experience works still depends heavily on the language you speak.
For low-resource languages like Sindhi and Pashto, spoken by millions of people, the gap is substantial. Speech recognition can struggle to produce a usable transcript. Speech generation can struggle with pronunciation and naturalness. Once you introduce background noise, regional dialects, or a speaker switching between languages, even a promising demo can become an unreliable product.
“Low-resource” describes the data and tooling available for a language. It says very little about how many people need it, or how much they would benefit from software that understands them.
This is where we are starting with QuantClarity.
The best model keeps changing
If you have tried building a voice application, you have probably gone through some version of this process.
You pick a speech recognition model. You test it on a few recordings, and the results look reasonable. Then you try a different dialect. Or a phone call with poor audio. Or someone mixing their native language with English. Performance drops, and you start looking for alternatives.
Another model handles those recordings better, but takes too long to respond. A third is fast enough, but misses names and numbers. You eventually find a combination that works for your use-case.
Then a new model comes out.
Now you need to work out whether it is better, where it is better, and whether the improvement justifies changing your system. The evaluation and integration work starts again.
This is a moving target. Even within a single language, the best choice can depend on the speaker, the recording conditions, and the task. A model that works well for a recorded interview may be a poor fit for a live conversation. The same tradeoffs exist in text-to-speech: pronunciation, naturalness, response time, and cost all matter, and their importance changes with the application.
Yet every team is expected to keep navigating this landscape for itself.
We need intelligent routing for voice
Our vision for QuantClarity is a layer that helps applications use the right voice model for the right situation, and keeps that choice current as models improve.
To do this well, routing needs to be grounded in evaluation. We need to understand how models perform across languages, dialects, noise conditions, and real use-cases. We also need to understand their limits: where they fail, how much latency they introduce, and what it costs to use them.
An application should be able to express what it needs. A live voice assistant might need a response within a tight latency budget. A transcription service might accept a longer wait for better accuracy. A workflow that captures an address or an amount might need an explicit verification step before proceeding.
Those requirements should guide model selection.
As new open and closed source models become available, we should be able to evaluate them against those requirements and adopt them where they improve the experience. Developers should not have to rebuild their voice stack every time the model landscape changes.
Evaluation has to reflect the actual conversation
A single benchmark score cannot tell you whether a model will work for your users.
Word error rate is useful, but it does not capture the full consequence of an error. Getting a conversational filler wrong and getting an account number wrong can have very different outcomes. A transcript can look mostly correct and still cause the application to do the wrong thing.
The questions we care about go further. Does the system understand the speaker’s intent? Does it preserve names, places, and numbers? Can it handle code-switching? Does it respond quickly enough for a conversation to feel natural? Can the user understand the generated speech without effort?
For languages with limited evaluation data, building that evidence is part of the work. It requires representative recordings, people who understand the language and its dialects, and consistent ways to compare results.
Routing is only as useful as the evidence behind it.
Routing also needs to know when it cannot help
There will be situations where none of the available models performs well enough. Routing cannot recover speech that every model misunderstands, or make up for capabilities that do not yet exist.
A useful system needs to recognise those limits and support a way forward. Depending on the use-case, that might mean asking the speaker to repeat something, confirming a critical detail, or handing the interaction to a person.
Sending every request to several models is not a complete answer either. It adds cost and delay, and agreement between models does not guarantee correctness. Additional processing needs to earn its place by improving the outcome.
The end user experiences the whole interaction. They care whether they were understood and whether they could complete what they came to do.
Where QuantClarity starts
We are starting with the gap between what voice models promise and what developers can reliably deliver for underserved languages, including Sindhi and Pashto.
Our direction is to bring model evaluation and routing into a continuous process: understand the available options, select them according to the application’s needs, and reassess those choices as better models arrive.
Over time, this work should also make the remaining gaps clearer. Knowing where every available model fails gives model builders a more useful target for data collection, training, and evaluation.
The ambition is straightforward. A developer building for a Sindhi or Pashto speaker should be able to spend more time on the service they want to provide, with dependable infrastructure helping them navigate the underlying voice models.
And the person using that service should be able to speak in the language they use every day and expect to be understood.
That is the future we want to help build with QuantClarity.
Building for a language that deserves better?
Explore what we’re building