Generating Live Examples: An Agentic Approach to In-Situ Examples for Dynamic Program Comprehension
Programming often benefits from dynamic information that complements source code. In particular, program comprehension can take advantage of concrete run-time data whose structure and changes give meaning to abstract identifiers and computational steps.
Obtaining up-to-date run-time information from a part of a system poses a number of challenges: First, a suitable entry point from which to run the code in question is needed. Second, such an entry point might require an explicit setup and expect specific input to reveal the behavior of interest. Third, execution should be repeatable to react immediately to code or input changes.
In example-based live programming (ELP), programmers can attach concrete examples to parts of a program to receive continuous, up-to-date feedback as their program unfolds. To reduce the manual effort involved in example construction, recent approaches either reuse run-time data mined from previous executions or construct synthetic inputs. Reusing observed examples lacks control over their content, whereas synthetic input does not yet scale to larger domain models.
Our approach extends ELP with user-configurable, generated examples that scale to complex object graphs. In contrast to previous generation strategies using large language models (LLMs) with static context, we let an LLM guide the exploration of a code base to construct the required context and validate and refine proposed examples using execution feedback.
We evaluated our example-generation agent on a synthetic dataset of 12 programs with varying complexity, and validated it on 4 real-world projects, demonstrating that the vast majority of functions and branches can be reached through a generated example. We found that execution feedback enabled smaller, locally hosted models to reach the performance of frontier models, while numeric and formal reasoning remained a consistent weakness across LLMs.
We further developed a prototype that integrates our example-generation agent into an ELP system in Visual Studio Code. Our exploratory intra-subject study (n=8) using program comprehension tasks, comparing our prototype to GitHub Copilot as the baseline, revealed that participants trusted their acquired knowledge more when ad-hoc examples were available and used the tool like a lightweight debugger, while Copilot was perceived as providing more accessible, but less reliable predictions rather than execution-grounded observations. However, we found no effect on the correctness of answers to program comprehension questions.
In connection with existing ELP systems and AI assistants, automating example generation can achieve ubiquitous run-time feedback available to both programmers and AI.