Overview

Three coding-agent CLIs ran 10 live AWS operations tasks, with and without a task runbook, three repetitions each — 180 graded runs.

10 tasks3 agents2 arms180 runs

Tasks included

  • Attribute a security-group change with CloudTrail
  • Release unattached Elastic IP addresses
  • Investigate and fix Lambda invocation errors
  • Investigate and reconcile terraform configuration drift
See all 10 tasks

Agent profiles

Cursor — best overall
98%$0.20/runsuccess
Fastest, cheapest, and leanest on nearly every task, and it pairs that efficiency with…
Codex — best at adherence
95%adherence
The only agent with a perfect adherence record on the implicit-constraint audit task in…
Claude — best write-ups
5.0 / 5clarity
Its final reports read like an analyst's work product: structured what/who/when…
Cursor — cleanest commands
95%accuracy
Across all 180 runs, 5.2% of Cursor's commands errored (40 of its 60 runs were…

Per-metric comparison

Success

cursor
98%
claude
96%
codex
94%

Adherence

codex
95%
cursor
94%
claude
91%

Cmd accuracy

cursor
95%
claude
93%
codex
90%

Tokens

cursor
183,291
codex
279,038
claude
303,862

Agent time

cursor
59s
codex
110s
claude
121s

Cost

cursor
~$0.204
codex
~$0.463
claude
$0.588

Tasks

Overall takeaways

Methodology & grading

Three coding agents — Claude Code (claude-opus-4-8), Codex CLI (gpt-5.6-sol), and Cursor CLI (Composer 2.5 Fast) — ran ten realistic AWS operations tasks against live sandbox accounts: incident fixes, security remediations, cost cleanups, audits, and forensics. Every task ran in two arms: a skill arm, where the agent receives a task-specific runbook ("skill") describing the expected procedure, and a no-skill arm, where it gets only the task prompt. Three repetitions per agent per arm — 180 runs total. Runs were graded three ways: a deterministic checker that inspects live cloud state after the run, an LLM judge that grades the agent's process and final report against author-set criteria, and human review of every disputed grade. Where the human review changed a machine grade, the run's data is labeled "Manual override" with the reason — several of those corrections turned out to be findings in their own right (see the grading-integrity note at the end).