SupportBench
An early-stage, synthetic benchmark for evaluating complete AI-native customer-support systems.
The research series
- Stage 3: When escalation replaces action. Published.
- Stage 2: A first local pilot. Published.
- Stage 1: Defining the system we want to evaluate. Published.
- Related essay: The Work of Making AI Matter in the Enterprise.
- Stage 4: A second local baseline across the same complete agent configuration. Next.
- Stage 4: What enterprise readiness for AI agents actually means. Planned.
- Stage 5: Customer support as a proving ground for enterprise AI. Planned.
Stage 1 artifacts
- TinyBench-50 scenario CSV
- Methodology and limitations
- TinyBench-50 v0.1 manifest
- Contribution guide
- Apache-2.0 license
- TinyBench-10 Qwen3 4B local baseline
What it measures
Resolution, policy compliance, tool use, recovery, escalation, and customer-facing explanation in a fixed synthetic environment.
This is independent, early-stage research. It is not a leaderboard or production-readiness certification.