← Task samples
Ankle injury assessment: where care breaks down.
Ankle injury assessment: where care breaks down.
Ankle injury assessment: where care breaks down.
Seven model configurations assess the same injury and complete care in Second Opinion, our clinical care environment. This study traces their scores to the evidence reviewed, decisions made, and work saved.
Seven model configurations assess the same injury and complete care in Second Opinion, our clinical care environment. This study traces their scores to the evidence reviewed, decisions made, and work saved.
56
56
56
selected attempts
selected attempts
52
52
52
scored attempts
7
7
7
model configurations
0
0
0
full-credit attempts
The strongest mean was 52.2 out of 100.
The strongest mean was 52.2 out of 100.
The strongest mean was 52.2 out of 100.
GPT-6 Astra earned the highest mean weighted rubric credit in this snapshot. “Weighted rubric credit” is the share of available scoring points earned, with some criteria contributing more points than others. It is not a medical accuracy percentage or a task pass rate.
GPT-6 Astra earned the highest mean weighted rubric credit in this snapshot. “Weighted rubric credit” is the share of available scoring points earned, with some criteria contributing more points than others. It is not a medical accuracy percentage or a task pass rate.
In the selected 78.5-point Astra attempt, the evaluator credited the fracture site and pattern but found gaps in equipment availability, return-to-running planning, and urgent-care instructions.
In the selected 78.5-point Astra attempt, the evaluator credited the fracture site and pattern but found gaps in equipment availability, return-to-running planning, and urgent-care instructions.
The result makes the distinction concrete: recognizing an injury does not establish that the care workflow is complete.
The result makes the distinction concrete: recognizing an injury does not establish that the care workflow is complete.
Results you can inspect
Results you can inspect
Results you can inspect
Trajectories
Follow each attempt from the evidence it reviewed to the work it saved and the grading it received.
How this comparison was assembled
How this comparison was assembled
How this comparison was assembled
Selection
Selection
Selection
The latest eight selected attempts per model configuration on the same environment image, frozen in the September 14, 2026 snapshot.
The latest eight selected attempts per model configuration on the same environment image, frozen in the September 14, 2026 snapshot.
Scored attempts
Scored attempts
Scored attempts
49 completed attempts plus 3 Flash model failures scored zero. Four Nova environment failures remain unscored and are excluded from its mean.
49 completed attempts plus 3 Flash model failures scored zero. Four Nova environment failures remain unscored and are excluded from its mean.
Task and scoring
Task and scoring
Task and scoring
Revision 15. Fifteen criteria: fourteen with positive weights totaling 100, and one zero-weight safety criterion that can cap the reward.
Revision 15. Fifteen criteria: fourteen with positive weights totaling 100, and one zero-weight safety criterion that can cap the reward.
Evaluator
Evaluator
Evaluator
The recorded evaluator is GPT-5.6 Terra. It judges the task’s declared evidence. These judgments do not establish independent clinical validation.
The recorded evaluator is GPT-5.6 Terra. It judges the task’s declared evidence. These judgments do not establish independent clinical validation.
Interaction
Interaction
Interaction
Agents interact with the environment through tools: reading records, requesting evidence, and saving changes. These runs record tool-based interactions rather than clicks in a graphical application.
Agents interact with the environment through tools: reading records, requesting evidence, and saving changes. These runs record tool-based interactions rather than clicks in a graphical application.
Source of record
Source of record
Source of record
The attempt score CSV, criterion score CSV, exported task specification, and selected saved evidence extracts in the September 14 inspection package.
The attempt score CSV, criterion score CSV, exported task specification, and selected saved evidence extracts in the September 14 inspection package.
The attempt score CSV, criterion score CSV, exported task specification, and selected saved evidence extracts in the September 14 inspection package.
One case. A wider research agenda.
One case. A wider research agenda.
One case. A wider research agenda.
This study examines how evidence gathering, clinical reasoning, and tool use come together in one care workflow. It is one focused view of agent performance; broader evaluation spans additional clinical cases and computer-use workflows. The results here describe this scenario and its recorded tool-based runs.
This study examines how evidence gathering, clinical reasoning, and tool use come together in one care workflow. It is one focused view of agent performance; broader evaluation spans additional clinical cases and computer-use workflows. The results here describe this scenario and its recorded tool-based runs.