Research

Frontier data for the AI-native workforce

We build benchmarks, evaluation environments, and large-scale expert-annotated datasets — sourced from our marketplace of 25,000+ verified domain specialists — to help AI labs and enterprises measure and improve frontier model performance.

0 Expert Annotations
0 Verified Experts
0 Domains Covered
0 Published Benchmarks
Our Approach

Data, evaluation, and RL at the frontier

Model capabilities have outpaced publicly available training data. We close that gap with expert-sourced signal at every stage.

01

Frontier data

High-signal, domain-expert annotations across law, medicine, finance, and research — built to push past the plateau in publicly available training data.

02

Model evaluation

Explainable benchmarks across reasoning, planning, and tool use — scored by verified domain experts, not automated heuristics alone.

03

RL environments

Multi-turn, cross-application task environments built with subject-matter experts, used to train and reinforce frontier agents on real-world workflows.

Benchmarks

The AI Productivity Index

Four benchmark suites tracking how frontier models perform on real, expert-designed work — published openly and updated as new models ship.

Productivity Index — Agents

Long-horizon, cross-application tasks across professional workflows — from multi-step legal research to compliance filing — completed without human intervention.

51.2%± 3.8 · top model
RankModelDomainScore
🥇 1Fable 5Agents51.2
🥈 2Opus 4.8Agents49.7
🥉 3Sonnet 5Agents47.3

Productivity Index — Professional

Real deliverables across investment research, legal drafting, clinical reasoning, and compliance — benchmarked against verified domain-expert work product.

69.4%± 2.1 · top model
RankModelDomainScore
🥇 1Opus 4.8Legal69.4
🥈 2GPT-5.4Medical66.8
🥉 3Gemini 3.5Finance64.1

Productivity Index — Code

Software engineering tasks sourced from real repositories and staff-level engineers — measuring correctness, test coverage, and review quality.

65.5%± 6.2 · top model
RankModelDomainScore
🥇 1Sonnet 5Coding65.5
🥈 2Fable 5Coding63.9
🥉 3Opus 4.8Coding61.7

Productivity Index — Consumer

Everyday consumer tasks — shopping research, travel planning, form-filling — used to track how close frontier models are to reliable daily assistance.

56.1%± 3.3 · top model
RankModelDomainScore
🥇 1GPT-5.4Consumer56.1
🥈 2Gemini 3.5Consumer54.0
🥉 3Opus 4.8Consumer52.6
Expert Leaderboard

Public rankings. Real accountability.

Every expert on websphere is ranked against domain-specific skill benchmarks — coding, legal reasoning, medical strategy, and frontier-agent tasks — published as an open leaderboard.

RankContributorDomainScore
🥇 1Research Agents CollectiveAgents97.2
🥈 2SWE Frontier GuildCoding95.8
🥉 3Lex IntelligenceLegal94.1
4MedReason LabsMedical92.6
5QuantStrat NetworkFinance91.3
6Cortex Research GuildAgents90.5
7Northwind SWE CollectiveCoding89.7
8Statute AILegal88.9
9ClinReason GroupMedical87.4
10Alpha Signal LabsFinance86.8
Latest Research

Notes from the benchmark team

Read all research →
Benchmark

Introducing Productivity Index — Code

How we built a software-engineering benchmark from real repositories and staff-level review standards.

March 24, 2026
Benchmark

Introducing Productivity Index — Agents

Measuring long-horizon, cross-application agent tasks across professional services roles.

January 21, 2026
Research

Expanding the Productivity Index

Why we added investment banking, consulting, and healthcare roles to our professional benchmark suite.

December 10, 2025
Perspective

The economy will become an RL environment

Why we think reinforcement learning environments — not static datasets — are the next unit of frontier training data.

September 15, 2025

Want to help build the next benchmark?

We're hiring researchers and engineers to grow the AI Productivity Index — and always looking for verified experts to contribute.