Studied at commit
4b028fa(2026-10-05,@letta-ai/letta-codev0.34.4). Source paths below are relative to that repository and line numbers refer to that commit. Fact means read in source, tested, or observed at runtime. Interpretation means my reading of intent.
Interactive diagrams (Archify, every node links to source): architecture · execution flow · core abstractions · context flow · memory flow · tool runtime · agent loop · sub-agent flow
Minimal reproduction: experiments/letta-code/
TL;DR
- The agent loop is split across a contract.
- Backend side: the stateful backend runs one model step per run. When the model wants a tool, every call goes out as an
approval_request_messageand the run stops withstop_reason: "requires_approval"(src/backend/dev/provider-turn-executor.ts:496-529). - Client side: the harness classifies permissions, runs the tools on the user’s machine, then sends
{type: "approval", approvals: [...]}as the next request (src/headless.ts:2578-2667). The backend never executes a client tool.
- Backend side: the stateful backend runs one model step per run. When the model wants a tool, every call goes out as an
- One
Backendinterface, typed against the Letta REST client, has two implementations (src/backend/backend.ts:190-369).APIBackendtalks to Letta Cloud.LocalBackendemulates the same run/stream contract in-process, with pi-ai as the provider layer.- The local backend grew out of a test fake. History: facade
7b335668→ local backendf15c80f7→ pi-ai861c37f3.
- Memory (MemFS) is a per-agent git repository compiled into the system prompt from
HEAD. Uncommitted edits never reach the model.- The compiled prompt is cached per conversation and kept byte-stable. A new commit only adds a one-shot
<memory_update>to the next provider call (src/backend/local/local-backend.ts:896-947). - Verified at runtime: later calls in the same conversation no longer carry the update. Compaction, a new conversation, or an explicit recompile picks it up.
- The compiled prompt is cached per conversation and kept byte-stable. A new commit only adds a one-shot
- Context engineering is about prefix stability.
- Volatile facts (time, git status, agent ids, permission mode) ride on the user message as
<system-reminder>parts (src/reminders/engine.ts:548-606). - With a real model the system prompt (29,116 chars) and tool schemas (19 tools, 74,924 chars) were byte-identical across every step of a turn.
- Volatile facts (time, git status, agent ids, permission mode) ride on the user message as
- Long-running conversations use an append-only transcript plus compaction.
- Compaction appends a row and replaces only the in-context id list. The summary re-enters as a user-role
system_alert. - The sliding-window goal ignores the fixed prompt floor (
src/backend/local/compaction.ts:602-630). Verified with a real model: on a small (40–45k) window, a “compacted” request can still fill the window and the turn ends withmax_tokens_exceeded.
- Compaction appends a row and replaces only the in-context id list. The summary re-enters as a user-role
- The execution runtime is client-side and defaults to autonomy.
- The default permission mode is
unrestricted(src/permissions/mode.ts:10). - The kernel sandbox (bwrap/Seatbelt) is opt-in for agent shells and never isolates the network.
- Defense in depth: deny rules and a cross-agent memory guard apply in every mode, plus PreToolUse hooks, secret redaction on every output, and a 32k-char clamp with an overflow file.
- The default permission mode is
- Sub-agents are separate
lettaprocesses, each a headless child speaking the same stream-json protocol.- A fresh child is a new memory-less agent that sees only its brief; a fork gets a copy of the parent conversation. Only the final report returns, as a
<task-notification>. - Verified: in one-shot
letta -pthe parent exits before the report arrives. A long-lived host (bidirectional stream-json, TUI, listener) receives it.
- A fresh child is a new memory-less agent that sees only its brief; a fork gets a copy of the parent conversation. Only the final report returns, as a
Why This Repository Matters
- It is a production coding agent built by the MemGPT/Letta team around a stateful-agent thesis.
- Memory, identity and history live with the agent, not the session.
- The harness runs wherever the user is: laptop, CI, remote “computers”, chat channels.
- It shows a concrete answer to a real deployment problem: state in the cloud, tools on the user’s machine. The answer is a protocol, not an in-process loop. The same client code then also drives a fully local runtime.
- It treats memory as code. Memory is versioned Markdown edited with ordinary file tools and validated by pre-commit hooks. Background “reflection” agents curate it in git worktrees.
- The repository is engineered for AI contributors.
AGENTS.md(symlinked asCLAUDE.md) explains every CI rule with its why, and the checks are enforced:- no
../imports; - named exports and
export function; - a 1,000-line file ratchet;
- zero import cycles;
- layer boundaries;
- mock isolation.
- no
Repository Snapshot
| Item | Value (fact) |
|---|---|
| Package | @letta-ai/letta-code 0.34.4, bin letta → letta.js, a 22 MB Node bundle built with Bun from src/standalone-entry.ts (build.js:76-85) |
| Language / runtimes | TypeScript; dev on Bun (packageManager: bun@1.3.14), published for Node ≥ 22.19 |
| Size | ~305k lines of non-test TS/TSX in src/; 931 test files under src/ (12 of them API-gated integration tests) |
| Largest areas (non-test lines) | cli/ 89k (Ink TUI), channels/ 45k, websocket/ 37k (listener/app-server), tools/ 25k, agent/ 24k, backend/ 22k |
| Key dependencies | @letta-ai/letta-client (REST types), @earendil-works/pi-ai (local provider runtime), @modelcontextprotocol/sdk, ink/react, node-pty |
| History | 3,707 commits; first commit 2025-10-24 |
| Tests | 8,778 unit tests passed, 44 skipped, 7 failed. All 7 failures trace to this container’s environment (see Verification) |
Architecture
The architecture separates where state and reasoning live from where execution happens. The source states this layering in AGENTS.md and enforces it in scripts/check-layer-boundaries.js. The call graph below confirms it.
flowchart LR
user([User / host]) --> entry["letta CLI entry<br/>src/index.ts main()"]
entry --> loop["Client turn loop<br/>headless / TUI / listener"]
loop -->|classify| perms["Permission checker"]
loop -->|approved batch| tools["Tool manager"]
tools --> ws[("Workspace + shell")]
tools -.->|Agent tool| sub["Sub-agent process<br/>child letta"]
tools -->|Edit + git commit| memfs[("MemFS git repo")]
loop -->|stream request| iface["Backend interface"]
iface -->|cloud mode| api["APIBackend"] -->|HTTPS + SSE| cloud["Letta Cloud<br/>server loop, web_search"]
iface -->|local mode| local["LocalBackend<br/>in-process emulation"] -->|1 step per run| llm["LLM providers<br/>via pi-ai"]
cloud -.-> llm
memfs -->|HEAD into prompt| local
Layers (fact; AGENTS.md layer map, enforced by scripts/check-layer-boundaries.js):
cli/ Ink TUI, slash commands ┐
websocket/ listener + app-server (Desktop) ├─ hosts: each owns a turn loop
headless.ts one-shot and bidirectional -p ┘
agent/ domain: send, stream, approvals, memory, skills, sub-agents
tools/ tool definitions + execution manager
backend/ Backend contract; APIBackend (Cloud) and LocalBackend (in-process)
providers/ provider connection helpers permissions/ pure rules utils/ leaf
Three facts define the shape:
- The
Backendis the seam. Every host callsgetBackend()(src/backend/backend.ts:741-744). The mode comes fromresolveBackendMode()(src/backend/backend-mode.ts:24-28).BackendCapabilities(backend.ts:169-183) lets the client feature-gate behaviour such as server-side tool management, remote vs local MemFS, and environment routing. LocalBackendruns inside the same Node process as the client loop. It extendsHeadlessBackend(src/backend/local/local-backend.ts:178), which implements runs, replayable chunks, cancellation and orphan settlement (src/backend/dev/fake-headless-backend.ts:196-776). Only the model call leaves the process.- Client tools are never registered on the server agent. On Letta Cloud, agents are created with only the server tools
web_searchandfetch_webpage(src/agent/create-agent-request.ts:35;src/agent/create.ts:294-297). Client tool schemas travel asclient_toolson every request (src/agent/message.ts:310-337). Commit34367de5(#456, 2026-01-02) introduced this, deleting the older “stub tool registration” path (+178 / −1,154 lines).
Main Execution Flow
Entry point to answer, for letta -p "…" with --backend local. Every step below was also observed at runtime through the logging shim (S1 in Verification).
sequenceDiagram
autonumber
participant U as User
participant H as headless.ts loop
participant P as Permissions
participant T as Tool manager
participant B as LocalBackend
participant S as Local store
participant M as Model (pi-ai)
U->>H: prompt
H->>B: user msg (reminder parts + text) + client_tools
B->>S: settle orphan tool calls, append input
B->>M: cached system prompt + message view + tools
M-->>B: toolcall_end
B->>S: persist assistant message
B-->>H: approval_request_message + stop_reason requires_approval
H->>P: classifyApprovals
P-->>H: allow / deny / ask
H->>T: executeApprovalBatch
T-->>H: tool_return (scrubbed, clamped)
H->>B: type approval with approvals
B->>S: append toolResult
B->>M: next model step
M-->>B: text, no tool calls
B-->>H: assistant_message + end_turn
H-->>U: result
| # | Step | Where (fact) |
|---|---|---|
| 1 | bin → letta.js → standalone-entry.ts → index.ts main() |
package.json:8-10; src/standalone-entry.ts:1-11; src/index.ts:571 |
| 2 | Headless vs TUI: -p, --run, or a non-TTY stdin selects headless |
src/index.ts:809; src/cli/startup-mode.ts:1-18; dispatch src/index.ts:1291-1292 |
| 3 | Build the turn content: shared <system-reminder> parts (session context, agent info, MCP servers, permission mode, memory-git sync, disk space), then preloaded skills, then the prompt |
src/headless.ts:1894-1949; src/reminders/catalog.ts:8-120; src/reminders/engine.ts:548-606 |
| 4 | Turn loop: while (true) → sendMessageStream(conv, input, …, {maxRetries: 0}) |
src/headless.ts:2176-2244; src/agent/message.ts:347-355 |
| 5 | Request body: messages, client_tools, client_skills, streaming, background, include_compaction_messages; secrets scrubbed from outgoing text |
src/agent/message.ts:310-337, 382-440 |
| 6 | Backend run: reject if a run is already active; settle dangling tool calls (unless this is an approval turn); append input; startRun |
src/backend/dev/fake-headless-backend.ts:473-504 |
| 7 | Resolve the system prompt: the cached compiled prompt, plus a <memory_update> if the committed memory revision moved |
src/backend/local/local-backend.ts:462-500, 896-947 |
| 8 | Provider call: preflight compaction check, toPiMessages(view), the transient memory delta as a trailing system message, streamSimple |
src/backend/dev/pi-stream-adapter.ts:535-606 |
| 9 | Map the stream: text/thinking deltas → assistant_message/reasoning_message; toolcall_end → approval_request_message; done → stop_reason (requires_approval / max_tokens_exceeded / end_turn) |
src/backend/dev/provider-turn-executor.ts:427-550 |
| 10 | Persist each chunk to the store and attach the run id; a stream that ends without a stop reason becomes requires_approval or error |
fake-headless-backend.ts:691-776 |
| 11 | Client drain: StreamProcessor accumulates approvals by tool_call_id; end_turn with pending approvals is coerced to requires_approval |
src/cli/helpers/stream-processor.ts:162-220; src/cli/helpers/stream.ts:446-470 |
| 12 | requires_approval → classifyApprovals → one-shot headless denies ask (“Tool requires approval (headless mode)”) → executeApprovalBatch |
src/headless.ts:2578-2667; src/cli/helpers/approval-classification.ts:124-230 |
| 13 | The next input is the approval message; the backend turns it into toolResult messages |
src/headless-response-state.ts:12-21; src/backend/local/local-store.ts:1507-1548 |
| 14 | end_turn → done (a mod turn_end may ask to continue); other stop reasons go through a retry ladder or exit |
src/headless.ts:2524-2575, 2814-2824 |
Fact: the TUI (src/cli/app/use-conversation-loop.ts:744), headless bidirectional mode (src/headless.ts:4272-4600) and the WebSocket listener (src/websocket/listener/turn.ts:314-856) each have their own stop-reason loop. They share sendMessageStream, drainStream*, classifyApprovals, executeApprovalBatch and the recovery classifiers in src/agent/turn-recovery-policy.ts. The list of non-retriable stop reasons is duplicated three times: headless.ts:2814-2823, cli/app/retry.ts:12-21 and websocket/listener/recovery.ts:113-122.
Core Abstractions
classDiagram
class Backend {
<<interface>>
capabilities
createConversationMessageStream()
compactConversationMessages()
recompileConversation()
forkConversation()
}
class APIBackend
class HeadlessBackend {
runs, activeRunByConversation
executeConversationTurn()
}
class LocalBackend {
getOrCompileSystemPrompt()
compactLocalConversation()
}
class HeadlessTurnExecutor {
<<interface>>
execute(input) Stream
}
class ProviderTurnExecutor
class PiStreamAdapter
class LocalStore
Backend <|.. APIBackend
Backend <|.. HeadlessBackend
HeadlessBackend <|-- LocalBackend
HeadlessBackend --> HeadlessTurnExecutor
HeadlessTurnExecutor <|.. ProviderTurnExecutor
ProviderTurnExecutor --> PiStreamAdapter
HeadlessBackend --> LocalStore
Backend (the contract)
- Responsibility: the agent/conversation/run API, plus
capabilities. - Inputs: Letta REST request bodies (
messages,client_tools,client_skills, …). - Outputs: Letta streaming chunks (
assistant_message,approval_request_message,stop_reason,usage_statistics,event_message,summary_message). - Lifecycle: a process-wide singleton from
getBackend();configureBackendMode()swaps it (backend.ts:741-758). - Dependencies:
@letta-ai/letta-clienttypes only. - Source:
src/backend/backend.ts:169-369. - Why it exists: the client loop, TUI and listener can stay identical whether state lives in Letta Cloud or on disk. It also gives tests a fake.
HeadlessBackend / LocalBackend (server emulation)
- Responsibility:
HeadlessBackendprovides the server semantics: runs with ids, one active run per conversation, chunk recording for replay, cancellation, and synthetic results for dangling tool calls.LocalBackendadds the agent semantics: the system prompt compiled from MemFS, the memory delta, compaction policy, the model catalog, and mod hooks for compaction and LLM events.
- Lifecycle: one per process; conversations are loaded from disk lazily.
- Source:
src/backend/dev/fake-headless-backend.ts:196-776src/backend/local/local-backend.ts:178-976
- Why it exists: a fully local runtime without forking the client, plus deterministic executors for tests (
src/backend/dev/headless-turn-executor.ts:94-176).LETTA_LOCAL_BACKEND_EXECUTOR=deterministicswaps the model out (backend.ts:712-722).
HeadlessTurnExecutor / ProviderTurnExecutor / PiStreamAdapter (one model step)
- Responsibility: turn the resolved input into exactly one provider call and map it back to Letta chunks, with bounded recovery inside that single call.
- overflow → compaction, at most 3 (
pi-stream-adapter.ts:64, 826-846) - oversized payload → image elision or compaction (
:848-928) - transient errors → at most 3 retries, with
event_messagechunks (:61, 930-960)
- overflow → compaction, at most 3 (
- Source:
src/backend/dev/headless-turn-executor.ts:13-32src/backend/dev/provider-turn-executor.ts:552-565src/backend/dev/pi-stream-adapter.ts:486-963
- Why it exists: isolates provider specifics. pi-ai owns payload conversion, capabilities and the model catalog.
AGENTS.mdmarks re-implementing that inside Letta Code as “suspicious by default”.
LocalStore (transcript + in-context view)
- Responsibility: append-only
messages.jsonl(pi session-entry format withid/parentId, schema v2), thein_context_message_idslist, the compiled prompt (system-prompt.json), conversation metadata, tool-result repair and clipping on load. - Source:
src/backend/local/local-store.ts:927-1209src/backend/local/local-transcript.ts:24-160
- Why it exists: keeps the full history (what Letta calls “recall memory”) separate from the bounded model view.
Client turn loop (headless.ts, use-conversation-loop.ts, listener turn.ts)
- Responsibility: reminders, send, drain, classify, execute, resume, recover, queue.
- Inputs: user input or queued items (
src/queue/queue-runtime.ts). - Outputs: approval messages and rendered results.
- Why it exists: it is the agent’s execution runtime and UX host. Each host differs in how
askis resolved: a dialog in the TUI,can_use_toolover stdio for the SDK, WebSocket for the listener, auto-deny for one-shot headless.
Tool (defineTool → registry → per-turn snapshot)
- Interface:
ToolAssets {schema, description, modelForm, impl: (args) => Promise<unknown>}(src/tools/define-tool.ts:8-38). - Normalization: the manager normalizes results to
{toolReturn, status, stdout?, stderr?}(src/tools/manager.ts:294-319). - Toolsets: chosen per provider:
default(Claude-style),codex(exec_command/ApplyPatch) orletta(src/tools/toolset-catalog.ts:18-105; auto rulesrc/tools/toolset.ts:61-84). - Snapshots: a per-turn snapshot (
ctx-*, at most 4,096 retained) pins the registry used to execute that turn’s calls (manager.ts:360-405, 735-813).
PermissionChecker
- Decisions:
allow | deny | ask | alwaysAsk. - Modes:
standard | acceptEdits | unrestricted | strict; the default isunrestricted(src/permissions/mode.ts:3-35). - Ordered evaluation (
src/permissions/checker.ts:238-599):- workspace sandbox guard
- cross-agent memory guard
- deny rules
--disallowedTools- alwaysAsk
- mode override
- allow lists
- read-only shell, own-memory writes, reads in the working directory (auto-allowed)
- allow and ask rules
- default: ask
Sub-agent (SubagentConfig + manager)
- Definition: Markdown files with frontmatter (
name,description,tools,model,skills,fork,launchProfile) (src/agent/subagents/index.ts:90-109). - Built-ins:
general-purpose,fork,recall,memory,reflection,init(src/agent/subagents/builtin/*.md). - Launch: a separate process (
src/agent/subagents/manager.ts:180-268, 415-460).
MemFS
- What it is: the agent’s git repository at
~/.letta/agents/<id>/memoryin Cloud mode (src/agent/memory-filesystem.ts:46-57), or$LETTA_LOCAL_BACKEND_DIR/memfs/<id>/memoryfor the local backend (src/backend/local/paths.ts:48-53). - Layouts: v1 =
system/files are core; v2 = root*.mdare core, with aMEMORY.mdindex and deferred child directories (src/agent/memory-format.ts:6-41).
Agent Loop
input ─▶ run started ─▶ model step ─▶ stop_reason?
▲ ├─ requires_approval ─▶ client tools ─┐
└──────────── {type: approval} (new run) ◀───────────────────────┘
├─ end_turn ─▶ done
└─ error / max_tokens_exceeded ─▶ exit (not retried)
- Fact: the loop driver is the client. Each pass through “run started” is a new backend run. The backend has no
whileloop over tool calls;HeadlessBackend.executeConversationTurnperforms exactly one executor call (fake-headless-backend.ts:469-565). - Fact: termination.
end_turnends the loop.max_turnscounts only non-approval inputs (headless.ts:2177-2191).cancelledexits with code 130.- Non-retriable stops exit with an error.
- Fact: recovery happens on both sides.
- Inside a provider call (backend): at most 3 overflow compactions and at most 3 transient retries.
- Before the stream starts (client): approval-pending conflicts → fresh denials; conversation busy → back off 10s·2ⁿ, up to 3 times; transient errors → 1s·2ⁿ with
Retry-After, up to 3 times (src/agent/turn-recovery-policy.ts:391-419). - Mid-stream (client): resume through
streamRunMessages(runId, {starting_after: seq}), at most 20 attempts in headless and 60 in the listener (src/cli/helpers/stream.ts:552-981).
- Fact: parallel tool calls run concurrently but safely (
src/agent/approval-execution.ts:45-118, 427-437).Read,Grep,Glob,ViewImage,Agent,web_searchrun in parallel.Edit/Writeserialize per file path.Bash,exec_command,ApplyPatchand unknown tools share one global lock.- Results keep call order.
- Interpretation: splitting the loop costs one round trip per tool step in Cloud mode. In exchange, the backend never needs access to the user’s machine, and a run can be resumed or replayed by any host (TUI, Desktop, a phone through chat.letta.com) because the state is in the backend.
The agent-loop diagram shows the waiting states (compaction/retry, awaiting approval) and exits.
Context Engineering
Context engineering is central here. The design optimizes for a stable prefix and a bounded, recoverable view.
flowchart LR
subgraph Sources
P[User prompt]; R[Runtime reminders]; O[Raw tool output]; MF[(MemFS HEAD)]; SK[(SKILL.md)]
end
P --> UM[User message]
R -->|system-reminder parts| UM
O --> CL[["clamp 32k + scrub"]]
CL -->|full text| OF[(overflow file)]
CL -->|approval| UM
UM -->|append| TR[(transcript jsonl)]
TR -->|in-context ids| V[Message view]
TR -->|over window − reserve| CP[Compaction] -->|user-role system_alert| V
MF -->|at create / compact| CS[(Compiled prompt, cached)]
CS -->|same bytes| SYS[System prompt]
CS -.->|new commit: memory_update, one call| V
SK -->|name + description| SYS
V --> LLM((Model))
SYS --> LLM
TS[client_tools schemas] --> LLM
1. How context is constructed (local backend, fact). PiStreamAdapter.streamOnce builds a pi-ai Context (src/backend/dev/pi-stream-adapter.ts:594-606):
systemPrompt = cached compiled prompt (base preset + rendered core memory)
+ <available_skills> from client_skills (system-prompt-compilation.ts:391-430)
messages = toPiMessages(in-context view) (orphan tool results removed, trailing assistant stripped)
+ [ <memory_update> as a system message ] (only on the first call after a new commit)
tools = client_tools (every request)
2. What enters (observed, S1). The first request of a fresh agent contained:
- a 29,116-char system prompt with these sections:
- Context Architecture
- Identity
- Existence & Continuity
- Harness Architecture
- Self-evolution
- the
<human>and<persona>core-memory blocks - the
<memory>index <available_skills>listing 19 bundled skills (name + description only)
- one user message of 5 parts: 4
<system-reminder>parts (device/git, agent info and paths, MCP servers, permission mode) followed by the prompt; - 19 tool schemas totalling 74,924 chars.
Bash(12,098) andAgent(10,705) have the longest descriptions.
Fact: the tool schemas outweigh the system prompt about 2.5×, and the usage-reported prompt was about 24k tokens before any conversation. A full sample is in real_model/sample_request_S1.json.
3. What is excluded (fact).
- uncommitted memory (
system-prompt-compilation.ts:81-143readsgit show HEAD:); - bodies of files in v2 child directories, which are only listed as
<directory path=… index=…/>(:154-199); - skill bodies, which load on demand through the
Skilltool and are injected as a user message wrapped in<skill_content>(src/tools/impl/skill.ts:337-379); - tool output beyond the clamp, which goes to an overflow file;
- messages evicted by compaction, which stay in the transcript (recall) and are reachable through the
recallsub-agent; - sub-agent transcripts.
4. Tool results (fact). Every built-in result is scrubbed of secrets (the runtime API key and the agent’s whole secret vault) and clamped (src/tools/manager.ts:2307-2324).
- Per-tool limits: Bash 30k chars (10k head+tail on failure), Read 2,000 lines / 30k chars, Grep 10k, Glob 2,000 files; backstop
TOOL_RETURN_MAX_CHARS = 32_000(src/tools/impl/truncation.ts:11-38). - The full output goes to
~/.letta/projects/<cwd>/agent-tools/<tool>-<uuid>.txt, and the clamped text says[Full output written to: <path>](src/tools/impl/overflow.ts:32-101). - A second clip at 40,000 chars repairs oversized results when a transcript is loaded (
src/backend/local/local-message-projection.ts:14, 469-519).
5. Compaction (fact).
- Preflight: compact when the estimated context exceeds
window − min(16384, 20% of window)(provider-turn-executor.ts:224-272). The comment explains why: pi-ai shrinks the output allowance towindow − context − 4096, so a near-full request otherwise “finishes withlength” instead of raising an overflow. - Post-turn: check again using provider-reported usage (
pi-stream-adapter.ts:751-779). - Default mode
sliding_window(compaction.ts:43-44, 584-650): evict from 30% upward in 10% steps until the kept tail is under(1 − 30%) × window. It cuts only at an assistant message and never separates a pending tool call. - Fallback mode
all: summarize everything except a trailing pending tool call (:652-667). - The summary prompt asks for goals, what happened, identifiers kept verbatim, errors, current state, and “lookup hints” for searching history later (
:61-102). - The summarizer itself retries on overflow with a shrinking transcript, from 120k down to 2k chars (
:30-42, 503-549). - Storage: the result is stored as one appended
compactionrow, andin_context_message_idsbecomes[summary, …kept](local-store.ts:1157-1209). The summary is a user-role message whose text is a JSONsystem_alert(“Note: N messages … have been hidden …”) (compaction.ts:682-708). - Afterwards the system prompt is recompiled, which folds in any committed memory (
local-backend.ts:760-786).
6. Repository context (fact). There is no index or retrieval layer over the user’s code. The agent navigates with Grep/Glob/Read/Bash. The first user message carries the git branch, recent commits and git status as a reminder. Project knowledge persists only if the agent writes it to MemFS or a skill.
7. Sub-agent context (fact, verified S3).
- A fresh child is a new agent with its own system prompt and toolset, no memory blocks (
manager.ts:194-196), and no shared reminders (the reminder catalog has nosubagentmode). Its first request was[system, user(sender reminder + brief)], about 12.5k prompt tokens with 11 tools. - A fork gets a hidden copy of the parent conversation plus a “you are NOT the primary agent” reminder (
src/agent/subagents/fork-conversation.ts:50-106;manager.ts:769-803).
8. Isolation (fact). The parent sees only the child’s final text, via the Agent tool’s output file and a <task-notification> capped at 30k chars (src/tools/impl/task.ts:432-479). In S3b the child’s system prompt never appeared in any parent request.
9. Long-running control (fact + finding).
- Bounded compaction and retries, the summary’s “lookup hints”, and the recall sub-agent keep long conversations workable.
- Finding (verified): the sliding-window goal and its “fits now” check count message tokens only (
compaction.ts:602-630;local-backend.ts:755-758). The system prompt and tool schemas, about 24k tokens here, are left out.- On a 40–45k window this produced compacted requests of 35.1k–37.6k prompt tokens. pi-ai then clamped
max_completion_tokensdown to 1, the model returnedfinish_reason: length, and the turn ended withmax_tokens_exceeded. This happened in S4b (3/3 runs) and in one S4a run. - On 128k+ windows this is unlikely. On small local models (an Ollama 32k window is enough) the floor is a large share of the window.
- On a 40–45k window this produced compacted requests of 35.1k–37.6k prompt tokens. pi-ai then clamped
10. Prefix stability (fact + observation).
AGENTS.mdforbidsrole: "system"notifications in history. Automated context uses user-role<system-reminder>parts instead.- In S1 the system prompt and tool schemas were byte-identical across all three calls. The provider reported 23,680–24,192 of about 24.2k prompt tokens as cached on every S1 call, in all three runs (
provider_usage.json). - The cost of breaking the prefix is visible in the same file. The single S2 call that carried
<memory_update>had only 5,888–6,016 of about 26–27k prompt tokens cached; the next call, back on the original prompt, had about 25–27k cached again. - The cache-reuse header
X-Letta-Response-Stateis sent only for approval-only continuations without new reminders (src/agent/message.ts:51-55, 477-489).
Memory
| Question | Answer (fact unless marked) |
|---|---|
| What is memory? | Committed Markdown in the agent’s MemFS git repo. Root files (v2) or system/ files (v1) are core memory, rendered into the prompt. Indexed child directories and other files are external memory, read on demand. Skills live under skills/. Cloud agents without MemFS still have server-side persona/human blocks (src/agent/memory.ts:16-93). |
| Where is it stored? | Cloud mode: ~/.letta/agents/<id>/memory, synced to ${memfsBase}/v1/git/<id>/state.git (memory-git.ts:228-233). Local backend: $LETTA_LOCAL_BACKEND_DIR/memfs/<id>/memory. |
| Who writes it? | 1. The primary agent, with ordinary Edit/Write/Bash and its own git commit. These are auto-allowed by isOwnMemoryWrite (src/permissions/memory-write-allowance.ts:37-62) and coached by the prompt (src/agent/prompts/letta_root_memfs.md:60-77). 2. The background memory worker and reflection sub-agents, each in a private git worktree merged by the harness under a lock (src/agent/memory-worktree.ts:363-666; src/agent/memory-operation.ts:5-59). 3. A repair worker for conflicts or invalid commits, tried once per distinct state. |
| What validates it? | A pre-commit hook enforces exact name/description frontmatter, a MEMORY.md index per projected directory, skill folders, and tree limits (defaults: depth 2, 20k chars per file, 65,536 chars of core memory). Changing those limits needs LETTA_MEMORY_CONSTRAINTS_UPDATE=1 (src/agent/memory-git-hooks.ts:38-211; src/memory-frontmatter.ts:18-122; src/memory-constraints.ts:302-308). |
| Who retrieves it? | Prompt compilation (core memory). The agent itself reads external files with tools. Reflection reads a 40k-char parent-memory snapshot. |
| How does it enter context? | Compiled into the system prompt at conversation creation, explicit recompile, compaction, or a worker/reflection merge (src/agent/subagents/memory-worker.ts:170). A new commit mid-conversation adds a <memory_update> to the next call only (local-backend.ts:917-939). |
| Scope | Per agent, across all of its conversations and machines. Skills can be per project (.agents/skills), per user (~/.letta/skills) or per agent (in MemFS). Shared-memory repos can be attached to several agents (Cloud). |
| Updates and staleness | The prompt itself says “Editing memory does NOT change your behavior in the current turn” (letta_root_memfs.md:53-56). Verified gap: after the one-shot delta, later calls in the same conversation see neither the new memory nor the delta (probe P3b; real model S2: turn-2 calls reverted to the original prompt hash). The agent still knows about its own edits from the tool calls in its transcript; external edits stay invisible until a recompile. |
| Explicit or implicit? | Explicit: the agent decides what to write. Reflection adds a periodic, out-of-band consolidation pass. The effective default is every 25 steps: DEFAULT_SETTINGS sets reflectionTrigger: "step-count", reflectionStepCount: 25 (src/settings-manager.ts:182-184), and the headless init event of our runs reported step-count/25. The compaction-event default in src/cli/helpers/memory-reminder.ts:51-56 is only reached without those settings, so it is effectively unused. No test pins the fresh-install default. |
Four things that are easy to confuse:
Conversation history backend transcript (Cloud conversation / local messages.jsonl). Immutable, complete, searchable ("recall").
Context what one provider call carries: compiled system prompt + in-context view + tool schemas (+ one-shot delta).
Persistent memory MemFS: committed Markdown in a per-agent git repo; core files are rendered into the system prompt.
External storage overflow files, reflection transcripts (~/.letta/transcripts), memory-worker handoff snapshots, the workspace itself.
Observed with a real model (S2, 3/3 runs). Asked to remember a fact, the agent:
- inspected
$MEMORY_DIRwithgit status; - read
human.md,MEMORY.mdandpersona.md; - edited
human.mdand committed it with its own message (for examplememory: record user's favorite color (teal)).
The <memory_update> reached the next provider call. For this OpenAI-compatible model, pi-ai folded it into the system message, because the model’s compat flags say it does not support mid-conversation system messages (node_modules/@earendil-works/pi-ai/dist/utils/transcript.js:96-104).
In a new conversation the fact was in the compiled prompt, and the model answered “Teal” without tools.
Tools
- Interface and registration (fact): tools are a static table built from
defineTool(...)(src/tools/tool-definitions.ts:122-284). Three other sources add tools: SDK-registered external tools (register_external_tools,src/headless.ts:3919-3935), listener external tools, and mod tools. Serialization order is built-ins → external → mod, with later sources shadowing earlier ones (src/tools/client-tool-serialization.ts:48-78). - Discovery: there is no runtime discovery for built-ins. Tool exposure is a per-model toolset choice. Sub-agent types and skills reach the model through tool descriptions (sub-agent descriptions are injected into the
Agentdescription,src/tools/manager.ts:1185-1187, 1324-1350) and throughclient_skills. - Invocation:
executeTool(name, args, {signal, toolCallId, toolContextId, …})(src/tools/manager.ts:2447-2511) runs a pipeline: resolve the turn snapshot → wait for memory checkouts → modtool_start→ execution-phase mod permission overlay → PreToolUse hooks (block, orupdatedInput) → inject secret env →tool.fn→ flatten → scrub + clamp → PostToolUse feedback → modtool_endoverride (:1966-2433). - Errors: never thrown back to the model. Errors become
status: "error"text such asError executing tool: …,Tool not found: X. Available tools: …,Error: Tool execution blocked by hook. …orInterrupted by user. - Retries: no tool-level retries. Retries apply to provider calls and the stream, not to tools.
- Permissions: see the Permission checker under Core Abstractions. Classification happens before execution. At execution time only mod overlays are re-checked (
manager.ts:2156-2168), so the safety of execution depends on every caller classifying first. - MCP: MCP tools are not
client_tools. The agent callsletta mcp search|schema|callfrom the shell (bundled skillusing-mcp-tools).- Locally configured servers run on the client.
- Servers attached in Letta Cloud run on the server (
POST …/tools/{id}/run,src/backend/api/unified-mcp.ts:286-303). - Interpretation: this keeps the model-facing schema surface constant regardless of how many MCP servers are attached.
- Server tools:
web_search/fetch_webpagerun inside Letta Cloud. Their results arrive astool_return_messageand the client drops the matching approval (stream-processor.ts:162-169).
Runtime / Sandbox
Where reasoning ends and execution begins (fact):
| Step | Owner | Code |
|---|---|---|
| choose toolset, serialize schemas | client | src/tools/toolset.ts:290-439; manager.ts:735-813 |
| inference, emit tool calls | backend (Cloud server or in-process LocalBackend) |
provider-turn-executor.ts:496-529 |
| server tools (web search, cloud MCP) | server | create-agent-request.ts:35; unified-mcp.ts:286-303 |
| classify, ask, execute, scrub, clamp | client | approval-classification.ts; approval-execution.ts:147-441; manager.ts:1966-2511 |
| return results | client → backend, as an approval message |
approval-execution.ts:246-254, 296-301 |
Shell runtime.
- Commands run with
bash -c, orzshon macOS to work around a bash-3.2 heredoc bug (src/tools/impl/bash.ts:122-129). - Processes are detached, and kills target the whole process group: SIGTERM, then SIGKILL after 2,000 ms (
src/tools/impl/shell-runner.ts). - A foreground
Bashcall auto-backgrounds after 10 s when a queue bridge exists. One-shot headless has no bridge, so the call blocks. - At most 32 background processes and tasks (
process_manager.ts:137-143).
Sandbox (fact).
- OS-level filesystem isolation: bwrap on Linux (tmpfs-masks denied roots,
--die-with-parent, no--unshare-net,src/sandbox/bwrap.ts:15-22), and Seatbelt on macOS.AGENTS.md:689-691describes a Windows restricted-token helper, but at this commit the code has no Windows backend (src/sandbox/policy.ts:48allows onlyseatbelt | bwrap). - Where it applies:
- (a) Agent shells, only with
LETTA_FS_SANDBOX=1. Without an available backend it warns and runs unsandboxed (src/sandbox/availability.ts:88-115). - (b) The workspace sandbox requested by a runtime-start (Desktop); this one fails closed (
src/tools/impl/shell-sandbox.ts:62-91). - (c) Memory sub-agents, on by default: the whole child process is wrapped, with writes limited to
~/.lettaand the agent’s own memory (src/agent/subagents/sandbox.ts:20-40;src/permissions/sandbox-policy.ts:255-285).
- (a) Agent shells, only with
- Never isolated: the network,
rgsubprocesses, hook commands, and the parent’s in-processRead/Edit/Write. Those rely on the permission checker’s workspace guard and the cross-agent guard.
Interpretation. The defaults favour an autonomous, always-on agent: unrestricted mode (commit b5b757c8, #2197, “make unrestricted the default”) with opt-in confinement. The invariants that hold in every mode are about other agents’ memory (the cross-agent guard) and secrets (redaction everywhere). They are not about the user’s workspace. That fits the product’s model of “your agent on your machine”, but it shifts the safety burden to deny rules and hooks.
Sub-Agents / Workflow
sequenceDiagram
participant Parent as Parent loop
participant Tool as Agent tool
participant Mgr as Subagent manager
participant Child as child letta process
participant BE as Backend
Parent->>Tool: Agent(prompt, subagent_type)
Tool->>Mgr: launchSubagent
Mgr->>Child: spawn --new-agent --system type, prompt on stdin
Tool-->>Parent: task id + output file (depth 0 returns at once)
Child->>BE: new agent + conversation (or hidden fork)
Child->>BE: its own split loop + tools
Child-->>Mgr: stream-json result event (final text)
Mgr-->>Tool: report
Tool-->>Parent: task-notification, up to 30k chars, as a queued user turn
- Process model (fact): each sub-agent is a headless
lettachild process.- Arguments:
--new-agent --system <type> --output-format stream-json --permission-mode unrestricted, the parent’s allow/deny lists, the child’s--tools, and--max-turns. - The prompt goes in on stdin. The child gets its own process group.
- Environment:
LETTA_CODE_AGENT_ROLE=subagent,LETTA_SUBAGENT_DEPTH= parent + 1, and the parent’s ids (src/agent/subagents/manager.ts:180-268, 330-345, 415-460;src/agent/subagents/subagent-launcher.ts:215-289). - The depth limit is 2.
Agentis only given to full-capability children below the limit (src/agent/subagents/subagent-depth.ts:11-43).
- Arguments:
- Types:
general-purpose: Bash, Read, Edit, Write, todo tools.fork: all tools, the parent’s conversation.recall: a fork with search instructions.memory,reflection,init: memory launch profile, OS-sandboxed.- Custom types from
~/.letta/agents/*.mdand.letta/agents/*.md(src/agent/subagents/index.ts:141-153, 500-590).
- Return contract:
- Depth 0: returns a task id immediately; the report arrives later as a queued
<task-notification>(src/tools/impl/task.ts:937-960, 432-479). - Depth > 0: runs in the foreground and returns the report directly (
task-foreground.ts:22-64). - Reflection, memory and integration children complete silently.
- Depth 0: returns a task id immediately; the report arrives later as a queued
- Observed (S3a/S3b, 3/3 each):
- In one-shot
letta -p, the parent ended its turn with “Launched the subagent… I’ll report as soon as it finishes”. The child’s second provider call never completed. One-shot mode registers no message-queue consumer (the adder is set only in bidirectional mode,src/headless.ts:3596). - In bidirectional stream-json the notification arrived as a second turn and the parent answered “2 lines”. In one run the parent had already read the task’s output file with
Bashbefore the notification arrived.
- In one-shot
- Agent-to-agent messaging:
SendAgentMessageenqueues into another conversation (Cloud routing; by default the parent’s conversation) and returns without waiting. - Workflow: there is also a
Workflowtool with its own skill (workflow-authoring), which runs scripted multi-agent workflows (src/tools/impl/workflow.ts). I did not trace it in depth.
Important Source Files
| File | Why it matters |
|---|---|
src/backend/backend.ts |
the Backend contract, capabilities, APIBackend, mode factory |
src/backend/dev/fake-headless-backend.ts |
run semantics: one active run, settle orphans, persist chunks, missing-stop handling |
src/backend/local/local-backend.ts |
compiled-prompt cache, memory delta, compaction orchestration |
src/backend/dev/provider-turn-executor.ts |
one model step → Letta chunks; tool call → requires_approval; compaction threshold |
src/backend/dev/pi-stream-adapter.ts |
pi-ai Context; bounded overflow/transient/image recovery |
src/backend/local/compaction.ts, local-store.ts, local-transcript.ts |
sliding window, summary packaging, append-only transcript |
src/backend/local/system-prompt-compilation.ts |
MemFS → core memory (from HEAD), skills block |
src/headless.ts |
one-shot and bidirectional turn loops, approval handling |
src/agent/message.ts |
request body (client_tools, client_skills), scrubbing, response-state header |
src/cli/helpers/stream.ts, stream-processor.ts |
draining, approval accumulation, resume |
src/agent/approval-execution.ts, src/cli/helpers/approval-classification.ts |
permission classification and parallel-safe execution |
src/tools/manager.ts, src/tools/impl/truncation.ts, overflow.ts |
execution pipeline, hooks, scrub, clamp, overflow |
src/permissions/checker.ts, mode.ts |
rule order, modes, guards |
src/reminders/engine.ts, catalog.ts |
<system-reminder> construction |
src/agent/subagents/manager.ts, src/tools/impl/task.ts |
sub-agent processes and the report contract |
src/agent/memory-git.ts, memory-worktree.ts, memory-git-hooks.ts |
MemFS sync, worktrees, validation |
src/agent/prompts/letta_root_memfs.md |
the system prompt that teaches the agent its own architecture |
AGENTS.md |
rules with rationale; checked against code above |
5+ Implementation Decisions Worth Learning From
1. Split the agent loop at the tool boundary
- What they did: the backend performs one model step per run. Every client tool call ends the run with
requires_approval. The client executes and resumes with anapprovalmessage.- Approvals and denials share one message shape:
{type: "tool", tool_return, status}or{type: "approval", approve: false, reason}(src/agent/approval-execution.ts:246-301). - The local backend enforces “one active run per conversation” (
fake-headless-backend.ts:473-479). - It repairs interrupted turns by answering dangling tool calls with
Turn did not complete(:480-497;local-store.ts:1550-1580).
- Approvals and denials share one message shape:
- Problem solved: agent state lives in a server, but tools must touch the user’s files and shell, and several hosts (TUI, Desktop, phone, CI) drive the same agent.
- Why it is interesting: it reduces “remote agent, local tools” to a resumable protocol, so any host can continue a run. The same protocol later let them write a fully local backend without touching the client.
- Trade-offs:
- one round trip per tool step in Cloud mode;
- the client must own the retry/resume logic;
- four hosts re-implement the stop-reason loop (the three identical non-retriable stop-reason lists have no parity test);
- tool schemas are re-sent on every request (75k chars here; mitigated by prefix caching).
- Where in source:
provider-turn-executor.ts:496-529;headless.ts:2578-2667;agent/message.ts:310-337. - Reuse: if your tools must run somewhere other than your model loop, make “tool call” a stop reason and “tool result” an input message. Then interruption, approval UIs and remote execution are all the same mechanism.
2. One contract, two backends, and the local one grew out of a test fake
- What they did:
- extracted a
Backendfacade typed against the vendor’s REST client (7b335668, 2026-04-29); - implemented it in-process by turning the test fake
FakeHeadlessBackendinto a real runtime (f15c80f7); - swapped Vercel AI SDK for pi-ai two weeks later (
861c37f3). BackendCapabilitiesflags let the client disable what the local backend can’t do (backend.ts:169-183, 376-390;fake-headless-backend.ts:185-194).- Deterministic executors remain for tests (
headless-turn-executor.ts:94-176).
- extracted a
- Problem solved: a local-first mode with BYO providers, offline use and no cloud account, without forking a 300k-line client.
- Why it is interesting: the most expensive part, client hosts with resume/approval/queue logic, stayed untouched. The emulation is only as wide as the contract.
- Trade-offs:
- the local backend must faithfully emulate server semantics: runs, replay, cancellation, ids, compaction events;
- casts (
as never,as unknown as Stream) mark places where local shapes are coerced into the cloud types; - features that exist only server-side (
web_search, environment routing, server secrets) are capability-gated off.
- Reuse: introduce the protocol seam first, then a fake behind it, then promote the fake. Test against the fake throughout.
3. Memory as git, compiled from HEAD into a cache-stable prompt
- What they did:
- Memory is Markdown in a per-agent git repo. Only committed content is rendered (
system-prompt-compilation.ts:81-143). - The compiled prompt is cached per conversation, keyed by the raw-prompt hash and the memfs revision.
- A new revision does not rewrite the prompt. It adds a one-shot
<memory_update>(local-backend.ts:896-947; introduced for Opus 4.8 in5bca801a, generalized inda372cb3). - Pre-commit hooks validate structure and size (
memory-git-hooks.ts). - Background writers use worktrees merged under a lock.
- Memory is Markdown in a per-agent git repo. Only committed content is rendered (
- Problem solved:
- self-editing memory needs history, diffing, multi-machine sync and conflict handling;
- changing the system prompt on every edit would destroy prompt caching;
- unvalidated self-edits rot.
- Why it is interesting: “commit = publish to context” is a crisp contract the agent can be taught (
letta_root_memfs.md:53-56), and git gives provenance, rollback and sync for free. - Trade-offs: freshness vs cache stability.
- Verified gap (probe P3b, S2): after the one-shot delta, later calls in the same conversation see neither the delta nor the new memory.
- The primary agent’s own edits remain visible through its transcript, but an external edit (worker, another machine) stays invisible until compaction or a new conversation.
- Background merges do recompile, breaking the cache for that conversation.
- Reuse:
- treat committed memory as the single source of truth and render it into a stable prefix;
- if you deliver deltas, persist them into the transcript (or re-send until the next recompile) so they don’t evaporate after one call.
4. Keep the prefix stable: volatile context rides on user messages
- What they did:
- Time, git state, agent ids, permission mode, MCP servers and disk warnings become
<system-reminder>text parts prepended to the user content (reminders/engine.ts:548-606). AGENTS.mdforbids role-systemnotifications in history.- The system prompt teaches the model to treat these tags as instructions.
- Skill-catalog changes arrive as a metadata-only reminder (
client-skills.ts:227-280). - An approval-boundary cache header is sent only when nothing new was added (
message.ts:51-55, 477-489).
- Time, git state, agent ids, permission mode, MCP servers and disk warnings become
- Problem solved: Anthropic and other APIs reject or penalize mid-conversation system messages, and every byte changed in the prefix costs a cache miss.
- Evidence: the S1 system prompt and tool schemas were byte-identical across the turn, and at least 98% of each S1 call’s prompt tokens were served from the provider cache. The one S2 call whose system message changed (the
<memory_update>, folded in by pi-ai) dropped to about 22% cached. - Trade-offs:
- reminders are part of history: session-context and agent-info are sent on the first turn and re-sent after each compaction or working-directory change (
src/reminders/engine.ts:63-77, 291-312;src/reminders/state.ts:88-89); - user-role text carries authority only because the prompt says so;
- every place that displays history must strip them (
src/cli/helpers/backfill.ts:52-67;src/cli/app/system-reminders.ts:8).
- reminders are part of history: session-context and agent-info are sent on the first turn and re-sent after each compaction or working-directory change (
- Reuse: classify context by volatility. Stable things go in the system prompt. Per-turn things go in tagged parts of the turn input. Never edit the prefix to announce state.
5. Append-only transcript, summary as a message, bounded recovery
- What they did:
- The transcript is append-only. Compaction appends one row and replaces
in_context_message_ids. - The summary re-enters as a user-role
system_alertwith “lookup hints”. The sliding window cuts only at assistant messages and protects pending tool calls. - Compaction runs preflight with a 16k reserve and again post-turn from real usage; overflow compaction is capped at 3 per call.
- The summarizer avoids Fable 5’s refusal path by falling back to Opus 4.8 (
compaction.ts:430-455).
- The transcript is append-only. Compaction appends one row and replaces
- Problem solved: long-lived agents accumulate unbounded history, but the history must stay searchable (recall) and the model must keep going.
- Why it is interesting: like deepagents (studied earlier in this series), it never destroys history. Unlike deepagents, the view is a stored id list rather than an event replayed on each call, and the summary lives in the conversation instead of in a file.
- Trade-offs:
- Verified finding: the window goal ignores the fixed prompt floor (
compaction.ts:602-630). On small windows a compacted request can still fill the window and end withmax_tokens_exceeded(S4b 3/3; S4a 1/3). - Summaries are lossy (S4a: 2/3 recalled both facts).
- Verified finding: the window goal ignores the fixed prompt floor (
- Reuse:
- keep the log append-only and the view as data;
- size the compaction goal against the whole request (system + tools + messages);
- re-check after compacting.
6. Client-side runtime with layered, non-bypassable invariants
- What they did:
- Classification happens before execution, with a fixed rule order. Some invariants come before the mode override: the workspace guard, the cross-agent memory guard and deny rules (
checker.ts:278-332). - PreToolUse hooks can block or rewrite arguments.
- Secrets are redacted on the main output channels: built-in tool returns, streamed shell output, overflow files, hook output and outgoing user messages (
src/tools/secret-substitution.ts:202-517). Secrets are injected as env vars and never spliced into commands (:28-56). - Gap (my reading): results of SDK-registered external tools are clamped but not scrubbed (
src/tools/manager.ts:678-684). - A resource-keyed scheduler runs reads in parallel and serializes writes.
- A 32k clamp writes an overflow file.
- Classification happens before execution, with a fixed rule order. Some invariants come before the mode override: the workspace guard, the cross-agent memory guard and deny rules (
- Problem solved: an autonomous agent in
unrestrictedmode still must not read another agent’s memory, leak the user’s API key, or flood its own context. - Trade-offs:
- No re-check at execution: correctness depends on every caller classifying first.
- Bash prefix rules match only the primary command, so
Bash(curl:*)also matchesexport K=… && curl …(permissions-matcher.test.ts:293-307). - The kernel sandbox is opt-in and leaves the network open.
- Reuse: decide which invariants survive “yolo mode” and put them in the classifier ahead of the mode switch. Redact at every sink, not just the tool return.
7. Sub-agents as processes speaking the same protocol
- What they did: a sub-agent is the same CLI started headless.
- A fresh child is a stateless agent with only its brief (verified S3).
- A fork gets a hidden copy of the conversation.
- Only the final report returns, as a queued notification. The depth limit is 2.
- Memory-profile children are OS-sandboxed by default.
- Problem solved: isolation of context, crashes, permissions and filesystem access, plus reuse of every host feature (resume, permissions, logging) for free.
- Trade-offs:
- process spawn cost, and each child re-pays its own system prompt and tools (about 12.5k tokens here);
- verified: one-shot
letta -phas no queue consumer, so a top-level background report is lost when the parent exits.
- Reuse:
- if the agent binary can run headless with a structured stream, use it as its own sub-agent runtime;
- make the “report” the only return channel;
- make sure every host that can launch background children can also receive their results.
8. Learning off the critical path, landed through git
- What they did:
- Reflection: the reflection sub-agent (slash commands
/dream,/reflect;src/agent/reflection-runs.ts:41) reads a normalized transcript slice in a private worktree. By default it runs every 25 steps (a compaction trigger is also available), and only when the parent memory repo is clean (reflection-launcher.ts:770-1000). - Landing: the harness merges the result under a lock. The transcript is marked consumed only on
mergedorno_changes(memory-worktree.ts:235-245, 363-666). - Repairs: conflicts and invalid commits get one automatic repair per distinct state (
memory-conflict-repair.ts).
- Reflection: the reflection sub-agent (slash commands
- Problem solved: continual learning without blocking the user’s turn and without corrupting memory.
- Trade-offs:
- real complexity, about 10 statuses and retry back-off;
- docs drift:
AGENTS.md:762-771describes states (landed,noop,pending_integration) that no longer exist in code; - the primary agent’s own commits don’t take the lock (
memory-operation.ts:25-34).
- Reuse: run memory consolidation as a separate agent on a copy (a branch or worktree). Only advance the “processed up to” pointer after the result lands.
9. A codebase engineered for agent contributors
- What they did:
AGENTS.mdlists each CI-enforced rule with its why:@/imports instead of../(“grep-discoverable”);- named exports and
export function; - a 1,000-line file ceiling with a shrink-only baseline;
- zero cycles;
- layer boundaries;
- mock-isolation checks for Bun’s global
mock.module.
- Why it is interesting: these are search-ergonomics rules for LLM readers, not style preferences. The file-size ratchet exists because “agents commonly inspect large files in slices and miss distant state”.
- Reuse: if agents edit your repo, make conventions mechanical and explain them where the agent will read them.
Minimal Reproduction
experiments/letta-code/ is minilc, a stdlib-only Python package of about 1,100 lines plus a demo. It reproduces the architecture: a split loop over a backend contract, with memory compiled from git. ScriptedModel replaces the LLM so every property can be asserted exactly.
| minilc | reproduces |
|---|---|
backend.py |
Backend contract; LocalBackend run semantics (one active run, one model step per run, tool calls → requires_approval, orphan settlement, bounded overflow/transient recovery); cached prompt + one-shot <memory_update> |
transcript.py, compaction.py |
append-only JSONL + in-context ids; threshold window − min(16k, 20%); sliding window at assistant boundaries; summary as user-role system_alert |
memfs.py |
v2 layout rendered from git show HEAD: only; deferred child directories |
harness.py, reminders.py |
client loop; reminders on the user message; headless denial of ask |
permissions.py, tools.py |
rule order with non-bypassable guard/deny; hooks, scrub, clamp + overflow file; resource-keyed parallelism |
subagents.py |
Agent tool; fresh vs fork context; report-only return; depth limit |
The demo (./run.sh demo) walks through six scenes and checks 11 invariants:
- split loop;
- headless permission denial;
- memory commit → one-shot delta → stale next turn → fresh conversation;
- clamp + overflow;
- sub-agent isolation;
- append-only compaction.
Verification
All commands ran in this session’s container (Linux, Node 22.22, Bun 1.4.2 plus the pinned 1.3.14, Python 3.13, git 2.43).
1. Upstream unit tests (pinned commit, bun install --frozen-lockfile).
for s in 1 2 3 4; do node scripts/run-unit-tests.cjs --shard $s/4; done # CI's command
# → 8,778 pass · 44 skip · 7 fail
Each failure was traced to this container. The failing files were then re-run with the cause removed:
| Failing test(s) | Cause | Evidence |
|---|---|---|
auth/desktop-credentials (1) |
container sets BUN_OPTIONS=--smol, which fork() forwards to Node (“bad option: –smol”) |
passes with BUN_OPTIONS unset |
headless-subagent-stdout-loss (2) |
ambient AWS_* credentials make pi-ai resolve Bedrock (“security token … invalid”) |
passes with AWS_* unset |
utils/startup-log-boundary (1) |
Bun 1.4.2 vs pinned 1.3.14 | passes with Bun 1.3.14 on PATH |
tools/bash-background tree kill (1) |
PID 1 (process_api) does not reap orphans, so the killed grandchild stays a zombie (state Z) and kill(pid, 0) succeeds |
reproduced with a 10-line Node script |
mods/package-installer (2) |
tests make a file read-only (0444) to force a write failure; root ignores file modes | uid=0 wrote to a 0444 file |
2. Build. bun run build produced letta.js (22 MB) in 14 s. All real-model runs use this published Node bundle, not the Bun source path.
3. Runtime probe of the real LocalBackend (upstream_probe/probe_letta_code.ts, no network, capturing executor):
| Probe | Claim | Result |
|---|---|---|
| P1 | tool call → requires_approval; approval → toolResult in the next view |
held (["user","assistant","toolResult"]) |
| P2 | uncommitted memory invisible | held |
| P3a | first call after a commit carries <memory_update>; system prompt unchanged |
held |
| P3b | later calls still see the committed memory | did not hold: both later calls see it nowhere |
| P4a | compaction appends one row, earlier rows byte-identical, view shrinks | held (13 → 14 rows; view 12 → 1) |
| P4b | summary as user-role system_alert; recompiled prompt includes the new memory |
held |
| P5 | dangling tool call settled with “Turn did not complete” | held |
| P6 | second turn during an active run rejected | held |
| P7 | thresholds: 200k → 183,616; 32,768 → 26,215 | held |
4. Real model through the real CLI (real_model/run_letta_real.py).
Setup:
letta.js --backend localdrovedeepseek-v4-1-flashon Volcano Engine Ark.- It ran in headless mode in a throwaway
HOME, with--yolo. - A logging OpenAI-compatible shim sat in front of Ark. Ark’s
/modelsis empty, so discovery needed the shim; the shim also records every request body.
There were 3 full runs, making 40, 38 and 45 provider calls. Per-scenario checks are in results_run{1,2,3}.json. Per-call request shape and provider-reported usage (prompt, cached and max-completion tokens, finish reason) are in provider_usage.json, extracted from the shim logs by extract_usage.py.
| Scenario | Result (3 runs) |
|---|---|
| S1 Read → Write → answer | all 5 checks 3/3. Runs ended requires_approval, requires_approval, end_turn; answer.txt = 4217; system prompt and tool schemas identical across calls; reminders on the user message |
| S2 remember → recall (same conversation) → recall (new conversation) | 5/5 checks, 3/3. The agent committed to MemFS itself; <memory_update> reached the next call; turn-2 calls reverted to the original prompt with no memory; a new conversation answered “Teal” with no tools |
S3a sub-agent in one-shot -p |
6/6 checks, 3/3. The child ran as a separate, memory-less agent (12.5k-token prompt, 11 tools) seeing only its brief; the parent ended before the report |
| S3b sub-agent in bidirectional stream-json | 4/4 checks, 3/3. The report arrived as a <task-notification> turn; the child’s context never appeared in parent requests |
| S4a compaction across turns (45k window) | compaction, append-only transcript and system_alert summary 3/3; recalling both facts in turn C 2/3. In the failed run turn B ended max_tokens_exceeded after a compacted request of 37.6k tokens |
| S4b both reads in one step (40k window) | compaction fired 3/3; answer lost 3/3. In run 1, max_completion_tokens went 11,259 → 11,010 → 1 on the compacted 37,071-token request, and the stream ended with finish_reason: length |
What this shows and doesn’t: the protocol, prefix stability, memory compile and delta behaviour, sub-agent isolation and compaction mechanics all hold with a real model in the published build. S4 deliberately stresses small windows by editing the conversation’s stored context_window_limit. letta model set --model-settings alone does not take effect, because the top-level field wins (fake-headless-backend.ts:146-149). These results come from one model and three runs; they are evidence, not a benchmark.
5. Reproduction.
cd experiments/letta-code && ./run.sh
# [1/4] venv + pip install -e .[test] [2/4] wheel minilc-0.1.0 [3/4] 24 passed [4/4] demo 11/11 PASS
python3 mutation_check.py # 9/9 architectural regressions killed
The mutation check injects one regression at a time into a copy of src/:
- the memory change rewrites the prompt;
- memory compiled from the working tree instead of
HEAD; - destructive compaction;
- a cut between a tool call and its result;
- no orphan settlement;
- reminders written into the system prompt;
- a fresh child that inherits the parent conversation;
- a parallel shell;
- a headless
askthat silently executes.
Each one fails at least one test. The first version of the suite missed “compiled from the working tree”, so a stronger test was added.
6. Diagrams. All 8 Archify diagrams passed finalize --quality showcase --repo-root <clone>: schema validation, verified delivery, strict provenance check and a headless-Chromium browser check. All 1440×900 captures were inspected by eye. See assets/letta-code/archify/README.md.
Known limitations of the verification:
- Letta Cloud was not exercised (no account). Cloud-mode claims rest on client source only; the server-side loop is out of scope.
- The TUI and the WebSocket listener were read in source, not run. Real-model runs used headless one-shot and bidirectional modes.
- The OS sandbox was not exercised:
bwrapis not installed here, and the tests’ sandbox cases are mocked.
What I Would Reuse
- A tool call is a stop reason; a tool result is an input message. This one protocol covers remote execution, approval UIs, interrupts and resumption.
- A backend contract with a fake that can grow into the local runtime.
- Memory as committed files rendered into a stable prefix, with validation hooks and git history.
- Volatility-sorted context: stable prefix, tagged per-turn reminders, no prefix edits to announce state.
- Append-only log + id-list view + summary message + bounded recovery, with the compaction goal sized against the whole request.
- Invariants ahead of the permission-mode switch, and redaction at every sink.
- Sub-agents = the same binary, headless, report-only.
- Things I would change:
- persist the
<memory_update>delta, or keep re-sending it until the next recompile; - include the system prompt and tools in the compaction goal and re-check after compacting;
- give one-shot headless a “wait for background children” option;
- re-check permissions at execution time.
- persist the
Limitations / Open Questions
- Memory freshness (P3b, S2): is the one-call delta intentional, as cache-first policy, or an oversight? The test covering it (
src/backend/local-backend.test.ts:675-730, added inda372cb3) asserts only the first call. Uncertain about intent; not filed. - Compaction vs prompt floor (S4):
assertPromptFloorFitsContextWindow(pi-stream-adapter.ts:361-386) only rejects floors larger than the window, and nothing accounts for the floor in the compaction goal. Worth confirming with maintainers on small local models. - One-shot sub-agents: background reports are dropped in
-pmode; the child is orphaned when the parent exits. Possibly expected for one-shot use, but the tool description still promises a notification. - Cloud mode: the server-side loop (the
letta_agent_v1agent type, which perAGENTS.md:290-291runs onletta_agent_v3.py), server compaction and how Cloud renders MemFS into the prompt are not in this repository. - Docs drift: the reflection worktree states in
AGENTS.md:762-771differ from the code.src/tools/README.md:11, 15still says tools run serially.AGENTS.md:689-691describes a Windows sandbox helper that does not exist in the code. Thecompaction-eventreflection default inmemory-reminder.tsis overridden by the settings defaults. - Not traced in depth: channels (Slack/Telegram/Discord), mods/extensions, the app-server protocol v2,
Workflow, the reflection arena, crons/schedules, LSP.
Further Reading
- Source: letta-ai/letta-code @ 4b028fa, especially
AGENTS.md,src/tools/README.md,src/websocket/listener/AGENTS.md,src/agent/prompts/letta_root_memfs.md. - Commits:
34367de5client_tools spec7b335668API backend facadef15c80f7local backend861c37f3pi-ai migrationb5b757c8unrestricted default5bca801amid-conversation memory update
- Docs: https://docs.letta.com/letta-code/ (memory, MemFS, subagents, permissions).
- Background: MemGPT and Sleep-time compute. Interpretation: background reflection is this paper’s idea applied to a coding agent.
- pi-ai, the provider runtime the local backend delegates to: https://github.com/earendil-works/pi
- Companion studies in this series: deepagents puts middleware on a framework loop, OpenHands uses an event-sourced conversation, and openai-agents-python keeps the loop inside the SDK, with handoffs and sub-agents as tool calls. letta-code instead splits the loop across processes with a protocol.