Productivity Index — Agents
Long-horizon, cross-application tasks across professional workflows — from multi-step legal research to compliance filing — completed without human intervention.
We build benchmarks, evaluation environments, and large-scale expert-annotated datasets — sourced from our marketplace of 25,000+ verified domain specialists — to help AI labs and enterprises measure and improve frontier model performance.
Model capabilities have outpaced publicly available training data. We close that gap with expert-sourced signal at every stage.
High-signal, domain-expert annotations across law, medicine, finance, and research — built to push past the plateau in publicly available training data.
Explainable benchmarks across reasoning, planning, and tool use — scored by verified domain experts, not automated heuristics alone.
Multi-turn, cross-application task environments built with subject-matter experts, used to train and reinforce frontier agents on real-world workflows.
Four benchmark suites tracking how frontier models perform on real, expert-designed work — published openly and updated as new models ship.
Long-horizon, cross-application tasks across professional workflows — from multi-step legal research to compliance filing — completed without human intervention.
Real deliverables across investment research, legal drafting, clinical reasoning, and compliance — benchmarked against verified domain-expert work product.
Software engineering tasks sourced from real repositories and staff-level engineers — measuring correctness, test coverage, and review quality.
Everyday consumer tasks — shopping research, travel planning, form-filling — used to track how close frontier models are to reliable daily assistance.
Every expert on websphere is ranked against domain-specific skill benchmarks — coding, legal reasoning, medical strategy, and frontier-agent tasks — published as an open leaderboard.
How we built a software-engineering benchmark from real repositories and staff-level review standards.
March 24, 2026Measuring long-horizon, cross-application agent tasks across professional services roles.
January 21, 2026Why we added investment banking, consulting, and healthcare roles to our professional benchmark suite.
December 10, 2025Why we think reinforcement learning environments — not static datasets — are the next unit of frontier training data.
September 15, 2025