SupportBench
SupportBench is an open testing environment for AI customer-support agents. It uses realistic, synthetic support cases to evaluate whether an agent can understand the problem, use tools correctly, follow policy, take appropriate action, and recover when things go wrong.
Why I started it
Customer support is where AI often meets someone after something has gone wrong. A payment is missing. An account is locked. A service has stopped working. A generic response only adds to the frustration. The customer needs the system to understand what matters, respond with care, and move the problem forward.
That requires more than fluent language. A dependable support agent must find the right evidence, follow policy, use tools safely, recover from mistakes, and know when a person should step in.
SWE-bench gave software-engineering agents a clear test: start with a real issue and a codebase, then determine whether the patch fixes the problem. Support is harder to reduce to a passing test. An answer can be correct while the customer’s problem remains unresolved.
Existing work such as τ-bench tests important parts of customer-service interaction. SupportBench examines the broader system: whether an agent can make sound decisions, take appropriate action, recover when conditions change, and move the problem toward resolution.
The research series
- Stage 3: When escalation replaces action. Published.
- Stage 2: A first local pilot. Published.
- Stage 1: Defining the system we want to evaluate. Published.
- Related essay: The Work of Making AI Matter in the Enterprise.
- Stage 4: A second local baseline across the same complete agent configuration. Next.
- Stage 5: What enterprise readiness for AI agents actually means. Planned.
- Stage 6: Customer support as a proving ground for enterprise AI. Planned.
Stage 1 artifacts
- TinyBench-50 scenario CSV
- Methodology and limitations
- TinyBench-50 v0.1 manifest
- Contribution guide
- Apache-2.0 license
- TinyBench-10 Qwen3 4B local baseline
What it measures
Resolution, policy compliance, tool use, recovery, escalation, and customer-facing explanation in a fixed synthetic environment.
This is independent, early-stage research. It is not a leaderboard or production-readiness certification.