sviat.infocase studiesllm-judge
Case 01case.ai
Cutting flagged chatbot responses by 30%.
I architected an LLM-as-a-Judge evaluation platform that checks a chatbot’s answers twice: live, before a user sees them, and offline, before every release. Flagged responses fell by 30%.
- Where
- Graphit (now Datasoft Group)
- Role
- Architect · Co-Founder & Lead, Data & AI Platforms
- System
- LLM-as-a-Judge evaluation platform
- Runs
- Live in production + offline before releases
Animated diagram of the LLM-as-a-Judge case: a chatbot answers user questions; a judge model scores every answer before it is served, blocks or reroutes low scores, and sends failures to an evaluation set whose regression tests gate every release. Simulated demo.
- the bot’s sourcesChatbot side — simplified to a RAG bot for this demo
- docs
- tickets
- chats
- users
- chunk · vectoriseChatbot side — simplified for this demo
- vector index · top-kChatbot side — simplified for this demo
- LLM · answerChatbot side — simplified for this demo
- live gate · every answerLive: every answer scored before users see it — low scores blocked or rerouted
- copilot
- document chunk
- query
- retrieved context
- answer
- flagged by judge
Hover or focus a stage for detailsTap a stage for details
The AI pipeline from the homepage, rewired for this case. The chatbot side is simplified; the judge, the eval set and the release loop are what I built. Rates are simulated — the −30% is the real result.
The problem
Too many of a production chatbot’s responses were being flagged.
The platform had two jobs: stop a bad answer before a user sees it, and show that a fix works before it ships.
What I did
01 map
Find the failure patterns
Offline, the judge scores conversations and test sets, and surfaces the failure patterns worth fixing.
↑ stage 04 · Generate
02 ship
Gate every answer live
In production, the judge scores every answer before a user sees it. Low scores are blocked or rerouted instead of served.
↑ stage 05 · Judge → Serve
03 prove
Regression-test every fix
Before each release, fixes are regression-tested against the judge’s test sets, so a fix is measured before it reaches users.
↑ stage 05 · Judge → Serve
Results
- −30%flagged chatbot responses
- Every answerscored live, before a user sees it
- Every releaseregression-tested offline first
Stack & methods
- LLM-as-a-Judge
- Evaluation sets
- Regression testing
- Production gating