Bryn Labs · the research team behind Superbryn

ReliableVoice AGI.

Voice AI that demos well is easy. Voice AI you can trust with a million real conversations is a research problem. Bryn Labs exists to solve it, with evals that predict production, observability that misses nothing, and agents that learn from every failure.

1M+ production calls studied50+ conversation dimensions measured4 analysis lanes on every call60+ failure modes catalogued
EvalsObservabilitySelf-LearningMemorySimulationSpeech UnderstandingReliable Voice AGI
Why we exist

The last mile of voice AI isn't a feature. It's a science.

Foundation models made voice agents easy to start and hard to trust. A stack of probabilistic systems fails in ways no unit test anticipates: fragmented turns, phantom interruptions, policies that bend under pressure, silence that reads as thought. The gap between a great demo and a dependable agent is where deployments die. That gap is our entire research agenda.

01
Failure hides in the tail

An agent that is 95% right sounds finished and is not. The last 5% is rare, strange, and expensive, and it only shows up at production scale.

02
The stack is non-deterministic

STT, LLM, TTS, and telephony are four probabilistic systems in a trench coat. Every layer can be individually fine and jointly wrong.

03
Trust is earned per call

Callers do not average their experiences. One bad call undoes a hundred good ones, so reliability has to be engineered call by call.

Research pillars

Six threads. One goal.

Reliability is not one problem. It is a stack of them, from the acoustic frame to the final verdict. We work every layer.

01
Evals

Measurement is the product. Scenario suites, adversarial personas, and verdict systems that predict production behavior, not leaderboard behavior.

scenario suites · personas · verdicts
02
Observability

Every call, fully measured: latency, interruptions, policy adherence, outcomes. Computed on all traffic, not a sample, so the tail has nowhere to hide.

dimensions · funnels · alerts
03
Self-Learning

Failure as fuel. Loops that turn a flagged call into a diagnosis, a fix, and a passing re-eval, shrinking the class of mistakes an agent can make twice.

feedback loops · auto-tuning
04
Memory

Continuity across conversations. What an agent should carry between turns, calls, and callers, and what it must provably forget.

context · recall · forgetting
05
Simulation

Reality, rehearsed. Synthetic callers with human impatience: accents, interruptions, dead air, network glitches. Production chaos on demand.

synthetic callers · chaos testing
06
Speech Understanding

Hearing before reasoning. Turn-taking, diarization, voice activity, pronunciation: the acoustic substrate every downstream judgment depends on.

VAD · diarization · turn-taking
Publication timeline

A decade in speech research.

The acoustic-modeling lineage behind the lab, from today's work back through noise-robust recognition at King's College London to subspace Gaussian mixtures at IIT Madras. Year by year, paper by paper.

All publications on Google Scholar →
Current research

On the bench right now.

Updated August 2026
01
Channel-anchored speaker attribution

Recovering who-said-what on stereo calls by anchoring diarization to per-channel voice activity instead of trusting STT speaker labels.

Speech UnderstandingIn production
02
Turn stitching under fragmented STT

Rebuilding true conversational turns when transcription splits them mid-thought, so downstream metrics judge the conversation that actually happened.

Speech UnderstandingIn production
03
Adversarial persona synthesis

Generating simulated callers that probe an agent where it is weakest: impatient, ambiguous, and off-script, at a scale no QA team can match.

SimulationActive
04
Verdict-grounded self-improvement

Closing the loop from a failed call to a prompt fix to a passing re-eval, with humans approving direction instead of hand-writing every change.

Self-LearningActive
05
Policy distillation for runtime guardrails

Compressing an organization’s full policy surface into the small, mode-aware rulesets an agent can actually follow mid-conversation.

EvalsActive
06
Pronunciation quality detection

Catching mispronounced names, numbers, and domain terms in agent speech with phoneme-level acoustic models.

Speech UnderstandingExploratory
07
A failure taxonomy for voice agents

Naming the ways voice agents break: a shared vocabulary of failure modes, detection strategies, and fixes, grown from production evidence.

FoundationsOngoing
Research in production

Superbryn is our lab bench. And our proof.

Most labs publish papers. We ship into Superbryn, the voice-AI reliability platform teams use to test agents before launch and watch them after. Every pillar above runs there against real traffic, which means our research is graded by production, not by benchmarks.

Simulated call testing

Scenario suites and personas run as real phone calls against your agent: before launch, on demand.

Production observability

Every live call analyzed across dimensions the moment it ends. No sampling, no blind spots.

Dimension analysis

Latency, interruptions, sentiment, compliance. Decomposed per turn, comparable across agents.

Metrics that learn

Feedback on verdicts tunes the metrics themselves, so measurement improves with use.

How we work

Lab principles.

01
Measured, or it didn’t happen

No claim ships without an eval behind it. If we cannot measure an improvement, we do not call it one.

02
Production is the benchmark

Benchmarks are rehearsals. The score that counts comes from live traffic, on calls nobody scripted.

03
Failures are the dataset

Every bad call is collected, clustered, and named. A failure mode we can name is a failure mode we can retire.

04
Small lab, full stack

From VAD frames to verdicts, the same people own the whole pipeline. Nothing gets lost between teams.

About the lab

A small lab with production stakes.

Bryn Labs is kept deliberately small: research owned end to end by the people who did the original work, graded by live traffic rather than citations.

NJ
Dr. Neethu Mariam Joy
Co-Founder & CTO
PhD, IIT MadrasPost-Doc, King’s College London
Google Scholar →
Voice agents fail in production because they weren’t built to learn from real-world conditions. We’re changing that.

14+ years of research in speech recognition, noise-robust acoustic modeling, and assistive voice technology. PhD from IIT Madras (2011–2018), followed by postdoctoral research at King’s College London and enterprise NLP work at ZS Associates and Uniphore. Published extensively in IEEE and INTERSPEECH on speaker normalization, low-resource language modeling, and speech recognition for impaired speakers.

+ Behind the founders: the Superbryn engineering team. Every line of lab research is hardened, deployed, and scaled with the builders of the platform.

FAQ

Fair questions.

Anything else, write to support@brynlabs.ai.

What is Bryn Labs?

Bryn Labs is the research team behind Superbryn. We study why voice agents fail in production and build the science that makes them reliable enough to trust with real conversations: evals, observability, self-learning, and memory.

What does “Reliable Voice AGI” actually mean?

A voice system you can hand a real task without supervising every call. Not a model that occasionally dazzles, but a system that holds goals across a conversation, knows when it is failing, and gets measurably better every week it runs. We treat reliability as the capability, not a property you add later.

How does the lab relate to Superbryn, the product?

Superbryn is where our research ships. Every pillar, from simulation to dimension analysis to self-improving metrics, runs inside the platform against real customer traffic. The product is both our lab bench and our proof: research that does not survive production does not survive here.

Do you publish your research?

We publish working artifacts first: our open blueprint library, failure taxonomies, and technical notes as they mature. We favor things teams can use over papers that gather citations.

Can my team work with the lab?

We partner with a small number of teams running voice agents in production as design partners for new evals and observability research. If that is you, write to support@brynlabs.ai.

Are you hiring?

The lab is kept deliberately small. If you have done serious work in speech, evaluation, or agent reliability and want your research graded by production traffic, we want to hear from you.

Work with us

Building a voice agent that can't afford bad calls?

We take on a handful of design partners for new evals and observability research. If your agent handles real customers, we want to study its failure modes and retire them.