↑ BACK TO SHOWCASE

CodeActionBench

Evaluating Agentic Code-as-Policy for Embodied Manipulation

Can general-purpose models turn visual understanding into robot action? We test the code they write, execute, and revise across 25 manipulation tasks.

Results Paper Coming soonCode Coming soon

Leaderboard

Verifier success rate (%)

25 tasks × 9 configurations × 3 attempts · 675 episodes

Detailed leaderboard
  1. GPT-6 Astra (Codex)?GPT-6 Astra (high) runs in the Codex agent, using the same robot tools over MCP.73.33
  2. Claude Opus 549.33
  3. Claude Code (Opus 5)?Default rows use the shared Reference harness. This row uses Claude Code 2.1.212 with vendor-mcp-direct; tasks, D0 tools, budgets and verifier stay fixed, so the complete agent stack is being compared.45.33
  4. Gemini 3.6 Flash20.0
  5. GPT-5.616.0
  6. Qwen 3.8 Max14.67
  7. Grok 4.612.0
  8. Kimi K3 Intl12.0
  9. Claude Sonnet 52.67

Evidence views

Demos

Recorded execution, paired with agent-visible observations. Open a demo for the full trajectory.

Leaderboard

Leaderboard

25 tasks × 3 accepted attempts · N=75 per configuration Submit your results

Sort byAverage Success Rate (ASR): verifier successes divided by scoreable episodes, including the laptop and dual-shoe regrade. Scoreable failures count as zero; pending and INFRA episodes are excluded.Number of the 25 tasks with at least one successful repeat.Number of the 25 tasks successfully completed in all three repeats.Median episode wall-clock time across the agent's 75 accepted attempts. It includes provider latency and is descriptive, not part of rank.Median charged outer tool calls across the agent's 75 accepted attempts. Internal primitives inside run_code are not charged again.
#AgentAverage Success Rate (ASR): verifier successes divided by scoreable episodes, including the laptop and dual-shoe regrade. Scoreable failures count as zero; pending and INFRA episodes are excluded.Number of the 25 tasks with at least one successful repeat.Number of the 25 tasks successfully completed in all three repeats.Median episode wall-clock time across the agent's 75 accepted attempts. It includes provider latency and is descriptive, not part of rank.Median charged outer tool calls across the agent's 75 accepted attempts. Internal primitives inside run_code are not charged again.EvidenceOpen a published per-turn trajectory or compare model trajectories in the Arena.
01GPT-6 Astra (Codex)?GPT-6 Astra (high) runs in the Codex agent, using the same robot tools over MCP.73.3355/7522/2515/258.0 min30
02Claude Opus 549.3337/7519/257/2519.9 min34
03Claude Code (Opus 5)?Default rows use the shared Reference harness. This row uses Claude Code 2.1.212 with vendor-mcp-direct; tasks, D0 tools, budgets and verifier stay fixed, so the complete agent stack is being compared.45.3334/7514/259/2522.7 min36
04Gemini 3.6 Flash20.015/758/253/2513.0 min42
05GPT-5.616.012/759/251/2510.6 min56
06Qwen 3.8 Max14.6711/756/252/2532.5 min36
07Grok 4.612.09/756/251/2515.9 min56
07Kimi K3 Intl12.09/755/250/2545.5 min55
09Claude Sonnet 52.672/752/250/2517.6 min56
Primary metric · verifier successes / scoreable episodes25 tasks · 3 episodes per task · every task weighted equallyTime and charged calls are descriptive resource metrics and never alter rank
Submit your results

Leaderboard submissions are reviewed against the frozen public release identity before publication.

  1. Run all declared model-task episodes with the frozen public task set, tool surface, seed and budget.
  2. Bundle the run identity, verifier results, trajectory manifest and media manifest; include a SHA-256 for the archive.
  3. Open a GitHub issue with the model and system name, paper or report link, archive location and checksum.
Submissions opening soon

Agent arena

Head-to-head verifier comparison

Compare two agents on the same tasks, with outcomes and complete trajectories.

Model Acodex-astra
VS
Model Bclaude-opus-5
17Both covered
5A only
2B only
1Neither covered

Open the arena

Benchmark interface

How agents see, code, and act

Explore the interfaceClose

Agentic Code-as-Policy

Task instructionLift the pot
Recorded agent view 1Recorded agent view 2
Current / retrieved RGB
Evaluated agentMultimodal LLM

Reference harness / Codex CLI / Claude Code

Policy code
Recorded excerptAstra · lift_pot · turn 6
pair2=capture_motion_pair(arm='right',dx=-0.07,dy=0,dz=0)
Tool callsResults / feedback

Robot API

Execute
Recorded head-camera view after the liftPhysical interaction
Outside the interfaceDepth sensingPerception / grasp specialistsPrivileged scene statePre-built skills

The agent decides what to observe, estimate, and execute.

Shared tasks · fixed budgets · hidden physical-outcome verifier

Resources

Results, paper & code

ResultsLeaderboard & trajectories
PaperComing soon
CodeComing soon
@misc{lyu2026codeactionbench,
  title  = {CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation},
  author = {Yiheng Lyu},
  year   = {2026},
  note   = {Release website preview}
}