Evaluating and rolling back prompts
In AEL, the Agent Engineering Language, prompts are treated as releasable business logic, with evaluations, recorded-model regression tests, drift checks, and promotion with one-step rollback, while active runs stay pinned to their version. This tutorial takes one change to a system prompt from your editor to your deployment, and back.
Step 1: change the prompt
Edit the prompt file, such as prompts/summarize.md. A prompt-only change gives the agent a new version even when no code changes, because an agent's version covers the exact text of every prompt it resolves, its input builders, its types, tools and permissions, its model bindings and their parameters, and its configuration.
Step 2: write the suite
A suite is a set of cases. Each case gives the input; the scripted or recorded model responses, or a live model binding; the expected typed outcome, or the scores that count as success; tool calls that must never happen; and limits for the case. Cover the hard cases too: missing information, contradictory input, invalid model output, tool errors, instructions hidden in retrieved documents and a budget running out.
| Cases | Model responses | Result |
|---|---|---|
| Contract | Scripted. | Must pass exactly. |
| Regression | Recorded from a real model earlier, played back with no network and no credentials. | Must pass exactly. |
| Live | A real model, called live, over a declared sample with limits on time, tokens and cost. | Scores. |
Step 3: run the suite
ael eval --suite support-regression
ael eval compares your working version with a baseline and keeps three kinds of difference apart: contract regressions, drift in live scores beyond what chance explains, and behaviour that no case tests. The command ends with a failure status when a contract or regression case fails. A recording that no longer matches its request, types or tools is marked incompatible, never replayed against the wrong request.
Step 4: compare with the released prompt
ael prompt diff --baseline <release>
ael prompt diff shows how the prompt changed against the released one.
Step 5: plan and promote
ael deploy plan
ael deploy apply --plan <plan>
- The plan shows what changes, prompts, model bindings, tools and their permissions, with a compatibility check for each and the evaluation results.
- Applying stages the new version, checks that it is ready, then switches new runs to it.
- Runs already under way stay pinned to the version they were admitted with, to the end.
- A change that widens a tool's permissions needs a new, authorized plan: a prompt update alone can never widen them.
Step 6: roll back
Rolling back is one step, to the previous compatible version, and the versions a rollback may still need are kept. Re-run your live cases regularly, and after every provider or model change, so that drift shows early. Evaluations and prompt releases describes versions, suites and promotion in full.
Honest limits
- Evaluation results describe how the agent performed on its cases. They are not proof of how a model behaves in every case, and model output is not reproducible run for run.
- Rolling back does not undo an effect that already happened outside your program, such as a message that was sent.
- A hosted model can change without any change on your side; pin model versions where the provider allows it.