sviat.infocase studiesllm-judge

Case 01case.ai

Cutting flagged chatbot responses by 30%.

I architected an LLM-as-a-Judge evaluation platform that checks a chatbot’s answers twice: live, before a user sees them, and offline, before every release. Flagged responses fell by 30%.

Where
Graphit (now Datasoft Group)
Role
Architect · Co-Founder & Lead, Data & AI Platforms
System
LLM-as-a-Judge evaluation platform
Runs
Live in production + offline before releases

case 01 · chatbot_answers → judged_answers

demopausedstatic frame · reduced motionsimulated≈ 1,840 answers/minflag rate ≈ 10%−30% flagged

Animated diagram of the LLM-as-a-Judge case: a chatbot answers user questions; a judge model scores every answer before it is served, blocks or reroutes low scores, and sends failures to an evaluation set whose regression tests gate every release. Simulated demo.

  1. the bot’s sourcesChatbot side — simplified to a RAG bot for this demo
    • docs
    • tickets
    • chats
    • users
  2. chunk · vectoriseChatbot side — simplified for this demo
  3. vector index · top-kChatbot side — simplified for this demo
  4. LLM · answerChatbot side — simplified for this demo
  5. live gate · every answerLive: every answer scored before users see it — low scores blocked or rerouted
    • copilot
  • document chunk
  • query
  • retrieved context
  • answer
  • flagged by judge

Hover or focus a stage for detailsTap a stage for details

The AI pipeline from the homepage, rewired for this case. The chatbot side is simplified; the judge, the eval set and the release loop are what I built. Rates are simulated — the −30% is the real result.

The problem

Too many of a production chatbot’s responses were being flagged.

The platform had two jobs: stop a bad answer before a user sees it, and show that a fix works before it ships.

What I did

  1. 01 map

    Find the failure patterns

    Offline, the judge scores conversations and test sets, and surfaces the failure patterns worth fixing.

    ↑ stage 04 · Generate

  2. 02 ship

    Gate every answer live

    In production, the judge scores every answer before a user sees it. Low scores are blocked or rerouted instead of served.

    ↑ stage 05 · Judge → Serve

  3. 03 prove

    Regression-test every fix

    Before each release, fixes are regression-tested against the judge’s test sets, so a fix is measured before it reaches users.

    ↑ stage 05 · Judge → Serve

Results

  • −30%flagged chatbot responses
  • Every answerscored live, before a user sees it
  • Every releaseregression-tested offline first

Stack & methods

  • LLM-as-a-Judge
  • Evaluation sets
  • Regression testing
  • Production gating

↑↓ navigate↵ selectesc close> terminal

↵ run↑ historytab complete⌫ back

Type > for the terminal