Services and retrieval
This tutorial builds an advisor agent in AEL, the Agent Engineering Language, that answers questions from your own documents. It shows where the services it calls run, how documents become knowledge, how an answer carries its evidence, and what happens when something it needs is not there.
Step 1: call a typed service
Code calls a typed service, and the deployment decides whether it runs in the same process, on another machine, over HTTPS or over gRPC. Your code does not change when the service moves:
| Placement | When it fits |
|---|---|
| In the same process | The service is part of your program, such as a reference knowledge store during development. |
| On another machine | The service is shared by several programs, or needs hardware your program does not have. |
The placement is chosen when you deploy; an agent does not move between machines while it runs. Connections are resilient, with idempotency-aware retries, backpressure, circuit breakers and health checks, and a timed-out call is never taken to mean that the other side did nothing.
Step 2: write the workflow
The advisor retrieves evidence, asks a model to answer from it, and checks the citations before it returns:
# agents/advisor.agent.ael
agent advisor(input: Question) -> Answer {
config {
nodes: [retriever = retrieve, writer = answer, checker = check_citations];
edges: [evidence_to_answer(retriever, writer), answer_to_check(writer, checker)];
starts: [retriever];
completion: checker.result;
}
}
retrieve and check_citations are deterministic nodes: they call no model. answer is the one model-driven node, with its own prompt, binding and attempts.
Step 3: ingest, update and delete
Retrieval (RAG) covers the whole lifecycle (ingest, chunk, embed, index, update and delete, then retrieve with ranking and provenance), and you can bring your own knowledge connectors:
- Ingest only the documents the run is permitted to read, within byte, page and record limits. Instructions inside a document stay content.
- Chunk each document into chunks with fixed IDs that keep their source, location and version.
- Embed the chunks in batches within limits; the embedding model and its version are recorded, and a mismatch is refused.
- Index a new version: queries see it only when all its chunks are indexed.
- Update a source against the version it replaces. Two writers updating the same source get a typed conflict.
- Delete a source: its chunks stop appearing in results once the deletion is committed.
An interrupted ingestion resumes without duplicating chunks, and a failure keeps the previous committed version of your knowledge.
Step 4: retrieve with provenance
Each result carries its source, version, location, score, retrieval time and the scope it was read under. Retrieved text reaches the model as evidence with a citation handle, separate from the system prompt, and check_citations confirms that every cited handle exists, names an allowed source and is fresh enough. Every store and every query is scoped by the trusted tenant identity of the run: agents are multi-tenant from the start, and results never reveal sources the caller may not see.
Step 5: bring your own parts
You write your own connectors for services, models, knowledge sources, transports and log sinks, in AEL:
| Part | What yours provides |
|---|---|
| Model | A client for an inference endpoint you run yourself. |
| Service | A typed service your nodes call, placed by the deployment. |
| Knowledge | A knowledge store, such as a search service you already run, a retriever or a ranking policy. |
| Verifier | A check that accepts or rejects an answer, is unavailable, fails in a way that can be retried, or hands the answer to a person. |
| Retry decision | A function that decides whether to retry, wait, fall back or stop, within your limits. |
| Log sink | A destination for the structured logs of your runs. |
Each connector declares its types, its permissions and its limits like any other code. Queries, reranking and the model calls they make draw on the run's shared budget, so retrieval cannot quietly multiply cost.
Step 6: handle what is unavailable
An unavailable source, a failed check and an empty result are different typed outcomes, each with the fallback you declared:
- Source unavailable. The run returns its typed error or uses the fallback you declared; it never answers as if it had found nothing.
- Empty result. The writer can say that no evidence was found, if your prompt and checks allow that answer.
- Citation check failed. The answer is repaired if the node has repairs left; otherwise the node applies its failure policy: a typed error, a typed fallback or an escalation.
- A read-only store returns an explicit error for an update or a deletion it cannot perform.
Retrieval over your documents and Tools and MCP describe each part in full.
Honest limits
- Retrieval quality depends on your documents, chunking and ranking, and no retrieval is guaranteed to find the relevant document.
- AEL checks where evidence came from and whether it is allowed and fresh; it does not decide whether an answer is true.
- Keeping retrieved text apart from the prompt does not make a model immune to instructions hidden in a document: your checks and the node's permissions limit what such text can cause.
- The built-in ingestion reads a reference document format; for other formats, you write a connector.