01
Preparing
Tests for hot AI coding claims
Run it against a real task
AI coding gets a new hot take every few months. People started promoting general workflow skills early in the year; by July, a wave of claims said those skills should be abandoned. Rather than decide who sounds more convincing, the plan is to run the tasks.
Planned test
The test will use short-, medium-, and long-horizon tasks, comparing current flagship and Flash models with equivalent models from when a skill became popular. The same task will run in an opinionated harness such as Codex and a minimal one such as pi. The comparison should show when the claim holds, how many extra tokens it costs, whether success rates move, and how much of the result comes from the harness. These tasks and metrics should eventually become a reusable benchmarking tool across models, skills, and harnesses.
- Short / medium / long tasks
- Flagship / Flash models
- Codex / pi
- Benchmarking tool