AWS Operations
Overview
Three coding-agent CLIs ran 10 live AWS operations tasks, with and without a task runbook, three repetitions each — 180 graded runs.
Tasks included
- Attribute a security-group change with CloudTrail
- Release unattached Elastic IP addresses
- Investigate and fix Lambda invocation errors
- Investigate and reconcile terraform configuration drift
Agent profiles
Per-metric comparison
Success
Adherence
Cmd accuracy
Tokens
Agent time
Cost
Tasks
Overall takeaways
Methodology & grading
Three coding agents — Claude Code (claude-opus-4-8), Codex CLI (gpt-5.6-sol), and Cursor CLI (Composer 2.5 Fast) — ran ten realistic AWS operations tasks against live sandbox accounts: incident fixes, security remediations, cost cleanups, audits, and forensics. Every task ran in two arms: a skill arm, where the agent receives a task-specific runbook ("skill") describing the expected procedure, and a no-skill arm, where it gets only the task prompt. Three repetitions per agent per arm — 180 runs total. Runs were graded three ways: a deterministic checker that inspects live cloud state after the run, an LLM judge that grades the agent's process and final report against author-set criteria, and human review of every disputed grade. Where the human review changed a machine grade, the run's data is labeled "Manual override" with the reason — several of those corrections turned out to be findings in their own right (see the grading-integrity note at the end).