Conversation Quality Analyst
listed 19 days ago
Nirdisha removes a role 90 days after it was listed.
- Sector
- AI
- City
- Bengaluru
- Area
- Koramangala and Outer Ring Road
- Experience
- 1 to 3 years
- Role family
- Operations
- Employment type
- FullTime
- Salary
- Not disclosed
- Posted
- 18 Aug 2026 · 2 weeks ago
- Last checked
- 7 Sept 2026
About the role
About Bolna
Bolna is Voice AI infrastructure built for India - and now for the world. We help businesses deploy intelligent voice agents that can call, converse, and convert in any language, at scale. From collections to customer support to sales, our agents handle millions of conversations so humans don’t have to.
We’re a YC F25 company, backed by General Catalyst, with 1,050+ paying customers and growing fast. Our team of ~25 is based in Bengaluru.
The Role
Every voice AI agent Bolna deploys makes real-time judgment calls - when to speak, when to go silent, when a customer is done talking, when to hand off. We’re building automated systems to grade these calls at scale, using LLMs as judges of call quality. But before you trust a model’s judgment, you verify it against a human’s.
That’s this role. You’ll listen to real calls, annotate what actually happened, and check whether our automated systems - LLM-as-judge evals and quantitative signal detection - got it right. It’s precise, high-attention work, and it sits right at the center of how we know our voice agents are actually working.
This is an internship role for someone early in their career who wants hands-on exposure to how a voice AI company builds trust in its own AI.
What You’ll Do
Annotation
Listen to and annotate real customer calls - transcription review, issue tagging, labeling - using tools like Label Studio
Follow (and help sharpen) annotation guidelines for a multilingual environment (Hindi, English, Hinglish, )
Verifying LLM-as-Judge Evaluations
For calls flagged by our automated eval pipeline, verify whether the model’s call was actually correct - for example, confirming whether a detected barge-in (agent/customer talking over each other) genuinely happened by listening to the audio
Mark agreements and disagreements clearly, with reasoning, so we can measure and improve model accuracy over time
All tools needed for this will be provided
Verifying Quantitative Measures
Check system-flagged quantitative signals against the actual call - e.g., confirming whether an “agent interruption” the system detected really occurred at that timestamp
Flag false positives/negatives so we can tighten detection logic
Help identify edge cases that current rubrics or detection logic don’t handle well
Inspecting Calls & Surfacing New Issues
Regularly inspect calls beyond flagged ones to spot new or emerging issues our rubrics and detection systems don't yet cover
Bring these patterns back to the team so rubrics, prompts, and detection logic keep improving
What We’re Looking For
Must-have
Strong attention to detail and the patience to do focused, high-precision work across many calls
Multilingual comfort preferred - Telugu, Tamil, Kannada, Marathi, Gujarati, or Bengali, in addition to English/Hindi
Comfortable learning new tools quickly - Label Studio, dashboards, internal QA apps
Genuine curiosity about AI and voice AI - you want to understand why a call was flagged, not just complete a checklist
The rest of this description is on the employer’s own page.