Coder Eval
Playwright for coding agents. An open-source, agent-agnostic framework for evaluating and benchmarking AI coding agents and their skills: it runs a real agent — Claude Code, Codex, Antigravity (Gemini), or OpenCode — in a sandbox against declarative YAML tasks, then scores the files and commands the agent actually produced.
The documentation has moved to coder-eval.com/docs.
UiPath/coder_eval stars forks