Why it mattered
Manufacturers hear a lot about AI and see little that looks like their own business. The agency needed demonstrations a machine builder recognises at once: its catalogue, its internal documents, its service desk. They are built for a fictional machine tool maker with a complete set of realistic data, so nothing confidential is involved.
A demo that invents a part number is worse than no demo. The agents had to be tested as seriously as software.
How it works
Three agents, one fictional company
- Product advisor. Recommends workholding from the catalogue for the customer’s machine. Prices appear only after the estimate tool has calculated them.
- Knowledge assistant. Answers employees from internal documents, cites the document and its version, tells an old travel policy from the current one, and says “this is not in the documents” together with who to ask.
- Service agent. Works through error codes and safe checks with an operator, opens a ticket in the CRM and books a service slot. Safety issues, unknown codes and requests for a person go straight to a human.
Five layers of testing, because one green run of a probabilistic system proves almost nothing:
- Acceptance scenarios for every promise the sales page makes.
- An automatic grounding check: any part number, document number, error code, price, email or phone number that appears neither in the agent’s data nor in the visitor’s words is flagged.
- Red teaming: prompt injection, prompt extraction, “I am the director”, fake instructions, pressure around safety, bait such as a non-existent part or a discount.
- Repeatability: the same scenario run several times to catch behaviour that drifts.
- Exploratory conversations, where each next message reacts to the agent’s real answer.