Model regression testing
The question
Can newer AI models perform worse against specific engineering environments?
Hypothesis
A newer model can improve generic benchmarks while regressing on a specific codebase, toolchain, or workflow.
Method
Define a fixed local task suite, run candidate models under identical prompts and tool permissions, then compare task success, errors, cost, and output stability.
Outcome
No results yet; this is an IDEA-stage entry.