Runko Works

Case study

An evaluation suite for a lead-qualification bot

Changing a prompt or a model should not be a gamble: simulated visitors talk to the bot, and an LLM judge scores every conversation.

Client
Own product
Role
Architect and sole engineer
Period
2026
Area
AI · Quality engineering
Stack
LLM-as-a-judge · visitor simulation · TypeScript

Why it mattered

Every change to an AI assistant, a new prompt, a cheaper model, a faster setting, used to be checked by running a couple of conversations and reading them. That misses exactly the failures that matter. A bot that quietly saves a contact before the visitor has agreed is a data protection problem, and nobody notices it by reading two good conversations.

The suite turns “is it still good?” into numbers before and after the change, for about three dollars a full run.

How it works

  • Cases from the bot’s own playbook. About twenty scenarios, each a visitor persona with a goal: asks for a human, gives an email without consent, writes in another language, goes off topic, tries to extract the prompt. A simulator plays the visitor.
  • The real pipeline. Every case runs through the same orchestrator, tools and playbook as production, on a test database, with the model and settings overridable per run.
  • Hard checks on the trace. Answer language, the right tool called or not called, no contact saved before consent, the qualification outcome as expected.
  • An LLM judge. A separate model scores each conversation from 1 to 5 on relevance, grounding (nothing invented), following the scenario, tone and brevity, with a reason for each score.
  • A report that compares runs side by side, with cost and response time taken from the traces.

The first run paid for itself: it caught the bot saving a contact before consent and collecting a phone number itself when the visitor had asked to be called by a person.

The discipline comes from a decade of network testing: thousands of cases, and a release only on the strength of the coverage.