A Preliminary Study of Type-Error Ablation and AI Coding Agents
Context
For decades, programming language implementors have designed error messages with one consumer in mind: the human programmer. Human-factors research has consistently found that programmers engage with error messages poorly—they skim, miss key information, and are easily overwhelmed. A practical consequence has been a strong design pressure toward brevity so that programmers will actually read them.
Inquiry
AI coding agents are now a second consumer of error messages. Unlike humans, agents will read large quantities of prose; the only practical limit on message size is the context window, and modern context windows are large. This raises a question the programming-language community has not previously had reason to ask: should error-message detail be calibrated differently for AI agents than for humans?
Approach
We investigate this question through a controlled experiment using Shplait, an ML-style statically typed language. We construct a suite of programs containing a single deliberate type error each, and measure how often an AI agent repairs them under ablation: a detailed error context; a proximate error location; a minimal type error; and a dynamic (test suite) error only. An automated oracle uses a combination of a type checker and a test suite to classify each modified program as having a type error, being type-correct but semantically incorrect, or being semantically correct.
Knowledge
We find concrete evidence that more detailed error messages generally improve an agent’s ability to fix type errors. We also see evidence that agents are able to fix programs whose names have all been obfuscated.
Grounding
Our primary experiment comprises 2,400 trials across ten complete runs, using both qwen2.5-coder:14b (a quality open-weight model) and claude-haiku-4.5 (a quality commercial model) via ollama and aider on dedicated GPU hardware. The experimental platform—correct programs, their erroneous counterparts, error modes, and oracle—is described in full.
Importance
If more detailed error messages genuinely help AI agents, then languages designed with only human readers in mind may be providing a suboptimal experience for AI coding. Our results suggest that it is worth considering offering two reporting modes: a terse human mode and a detailed AI mode. More broadly, we believe that language-model-targeted error-message design is an open and important design problem.