Studied at commit
71c2da4(2026-10-07,v0.23.1+ 22 commits;pyproject.tomlsaysopenai-agents==0.23.1). Source paths are relative to the repository root, and line numbers refer to that commit. Fact means read in source, tested, or observed at runtime. Interpretation means my reading of intent.
Interactive diagrams (Archify, every node links to source): architecture · execution flow · core abstractions · context flow · memory flow · tool runtime · agent loop · sub-agent flow
Minimal reproduction: experiments/openai-agents-python/
TL;DR
- The SDK is the runtime. Unlike harnesses built on LangGraph, this repo owns its control flow:
AgentRunner._run_implis awhile Trueloop (src/agents/run.py:1026) over a four-state machine,NextStepRunAgain | NextStepHandoff | NextStepFinalOutput | NextStepInterruption(src/agents/run_internal/run_steps.py:164-191). A turn is one logical model call plus its local side effects;max_turns(default 10,src/agents/run_config.py:45) bounds turns, not tokens. - An
Agentowns no control flow. It is a dataclass of instructions, tools, handoffs, guardrails and an output type (src/agents/agent.py:319-418). The Runner resolves dynamic instructions, enabled tools and handoffs again on every turn (run_loop.py:2711-2723; probe P12). - Everything the model can do is a tool call. A handoff is a function tool named
transfer_to_<agent>that switches the current agent (turn_resolution.py:3583-3594);Agent.as_tool()is aFunctionToolthat runs a nestedRunner.run(agent.py:1044-1060). The model sees one uniform contract; the runtime decides what a call means. - Responses-format items are the data model. Runner, sessions,
RunStateand results all storeTResponseInputItems; provider adapters convert at the edge (Converter.items_to_messages,src/agents/models/chatcmpl_converter.py:534). - Context management is mostly opt-in or delegated to the server. By default every tool output is replayed verbatim (probe P9). The SDK offers hooks (
call_model_input_filter,session_input_callback,SessionSettings.limit), an opt-inToolOutputTrimmer, and server compaction (ModelSettings.context_management,OpenAIResponsesCompactionSession). There is no built-in client-side summarizer in the core loop. - Human approval is a serializable pause, not a blocking call.
needs_approvalreturns a result withinterruptions;RunState.to_json()(schema1.20,src/agents/run_state.py:245) can be stored and resumed later; resume finishes the paused turn without a new model call (probe P7, real-model R3). - The newest layer is
SandboxAgent(commit2d665c9a, 2026-04-15): capabilities (Filesystem,Shell,Compaction,Skills,Memory) prepare a per-run execution clone of the agent with extra tools, prompt fragments and sampling params (src/agents/sandbox/runtime_agent_preparation.py:86-157). SandboxMemoryis a two-phase, agent-written file memory. - Checked with real models (deepseek-v4.1-flash and doubao-seed-2.1-pro, 3 runs each): handoffs carry the full transcript including the previous agent’s reasoning item;
as_toolchildren got exactly the one-line brief; and the default redacted tool-error text cut self-correction from 6/6 to 1/6 runs. See Verification.
Why This Repository Matters
- It is the reference implementation of OpenAI’s agent runtime design: a small public surface (
Agent,Runner, tools, handoffs, guardrails, sessions) over a large, defensively engineered core (run_internal/is about 20k lines). - It is a provider-neutral loop with a provider-specific data model. The loop speaks OpenAI Responses items; Chat Completions, LiteLLM and any-llm adapters translate. This makes the portability trade-offs very visible (see probe P17).
- The code base is maintained for coding agents:
AGENTS.mdplus 16 maintainer references in.agents/references/(e.g.runner-lifecycle.md,tool-execution-lifecycle.md) state invariants that every change must preserve. I used them as a map and checked each claim I kept against code. - The commit history shows hardening decisions that are worth copying, for example redacting tool failures by default (
40956e04, #5112) and moving nested handoff history from default-on to opt-in (a776d809#1996 →6ab83d43#2272).
Repository Snapshot
| Item | Value (fact) |
|---|---|
| Package | src/agents/ — 320 Python files, about 135k lines. Largest: run_state.py (5.6k), run_internal/turn_resolution.py (3.9k), tool.py (3.0k), run_internal/run_loop.py (3.0k), run.py (2.8k) |
| Major subpackages (lines) | run_internal/ 20k · sandbox/ 28k · extensions/ 29k (sandbox providers, LiteLLM/any-llm, session backends, experimental Codex tool) · models/ 7.5k · realtime/ 7.2k · mcp/ 4.6k · tracing/ 4.0k · voice/ 2.8k |
| Runtime dependencies | openai>=3.0.0,<4, pydantic>=2.12.2, griffelib, mcp>=1.19.0, websockets, requests (pyproject.toml:9-23) |
| Optional extras | litellm, any-llm, realtime, voice, sqlalchemy, redis, dapr, mongodb, encrypt, and sandbox providers docker, daytona, e2b, modal, blaxel, cloudflare, runloop, vercel |
| Default model | gpt-5.6-luna, overridable by OPENAI_DEFAULT_MODEL (src/agents/models/default_models.py:99-103); OpenAI provider defaults to the Responses API (models/_openai_shared.py:11) |
| History | 2,482 commits; first commit 2025-03-11. Commit rate peaked at 360 in 2026-08 |
| Tests | 414 test files; 12,174 passed with the locked environment (see Verification) |
| Examples | 224 Python files under examples/ |
Architecture
Fact: there is no framework underneath. The layers are all in this repo:
Public API Agent · Runner · tools · handoffs · guardrails · Session · RunState (agent.py, run.py, tool.py, ...)
Run loop _run_impl / start_streaming -> run_single_turn -> turn_resolution (run.py, run_internal/)
Model boundary Model / ModelProvider; Responses, Chat Completions, LiteLLM adapters (models/, extensions/models/)
Execution function tools, MCP, hosted tools, SandboxRuntime + sandbox sessions (tool.py, mcp/, sandbox/)
Cross-cutting tracing spans + processors, usage, hooks (tracing/, usage.py, lifecycle.py)
flowchart LR
app([Application]) -->|"Runner.run()"| runner["Runner<br/>run.py:1026 while True"]
runner -->|each turn| turn["run_single_turn<br/>run_loop.py:2665"]
turn -->|resolve per turn| agent["Agent<br/>config dataclass"]
turn -->|get_response| model["Model adapter<br/>items in, items out"]
model --> llm[(LLM API)]
turn -->|ModelResponse| resolve["turn_resolution<br/>classify + NextStep"]
resolve -->|tool calls| tools["Tool execution"]
tools -.-> mcp[(MCP servers)]
tools -.-> sbx["Sandbox session"]
resolve -->|NextStep| runner
runner <-->|history| session[(Session)]
runner -->|interruption| state[(RunState JSON)]
runner -->|prepare_agent| sbrt["SandboxRuntime"] --> sbx
Richer, source-linked version: architecture.html.
Module boundaries (fact):
| Module | Owns | Depends on |
|---|---|---|
src/agents/run.py |
public Runner, the non-streaming loop, wiring of everything below |
run_internal/*, sandbox.runtime, run_state |
src/agents/run_internal/ |
turn execution, response processing, tool planning/execution, handoffs, session persistence, server-conversation tracking, retries, streaming loop | items, tool, handoffs, models.interface |
src/agents/agent.py, handoffs/, tool.py, guardrail.py |
declarative definitions | run_context, function_schema, strict_schema |
src/agents/models/ |
Model/ModelProvider interfaces, OpenAI Responses (HTTP + WebSocket) and Chat Completions adapters, retry advice |
openai SDK |
src/agents/memory/ |
Session protocol, SQLite and OpenAI Conversations sessions, Responses compaction session |
run_internal.items |
src/agents/sandbox/ |
SandboxAgent, capabilities, manifests, session lifecycle, snapshots, sandbox memory |
runner (Runner.run for memory agents) |
src/agents/mcp/ |
MCP server connections, conversion of MCP tools to FunctionTool |
mcp SDK |
src/agents/tracing/ |
traces/spans, processors, OpenAI exporter (https://api.openai.com/v1/traces/ingest, tracing/processors.py:46) |
none in the loop |
src/agents/realtime/, voice/ |
separate runtimes for realtime sessions and STT→workflow→TTS pipelines | not on the Runner path studied here |
Interpretation: AGENTS.md asks maintainers to keep run.py “focused on orchestration” and put logic in run_internal/. In practice _run_impl is still about 1,800 lines (run.py:623-2400), because every exit path must also handle sessions, guardrails, tracing, sandbox cleanup and resume state.
Main Execution Flow
Entry point traced: await Runner.run(agent, "...", session=session) with a function tool, non-streaming. Runner.run_streamed follows the same steps inside a background task and a second loop (start_streaming, run_internal/run_loop.py:969).
sequenceDiagram
autonumber
participant A as App
participant R as AgentRunner
participant S as Session
participant T as run_single_turn
participant M as Model adapter
participant X as turn_resolution
participant F as Tools
A->>R: Runner.run(agent, input, session)
R->>S: get_items(limit) and save new input
loop until a final output or an interruption
R->>T: run_single_turn (turn += 1)
T->>T: resolve instructions, tools, handoffs
T->>M: get_response(system, items, tools, handoffs)
M-->>T: ModelResponse(output items)
T->>X: process_model_response
X->>F: approvals, guardrails, invoke (concurrent)
F-->>X: outputs or error text
X-->>R: SingleStepResult(next_step)
R->>S: add_items(turn items)
end
R-->>A: RunResult(final_output, new_items, interruptions)
Step by step (each step is a fact read in source unless marked otherwise):
- Public entry.
Runner.run(run.py:273) forwards to the defaultAgentRunner.run(run.py:344-357), which normalizesRunConfig, masks tracing when disabled, and calls_run_impl(run.py:568-605). Errors marked as data-redacted are re-raised without their traceback (run.py:358-378). - Input preparation. For a fresh run,
prepare_input_with_sessionreads the session history (honouringSessionSettings.limit) and returnshistory + new input, plus the items that still need persisting (run_internal/session_persistence.py:414-488). Withconversation_id/previous_response_id, history is not prepended because the server owns it (run.py:707-725). ARunStateinput instead restores context, turn counter and pending approvals (run.py:659-690). - Run-scoped helpers.
SandboxRuntime,PromptCacheKeyResolverand theRunStateitself are created once per run (run.py:855-868). - Loop head. Every iteration (
run.py:1026):- first-turn input guardrails: blocking ones run before sandbox preparation (
run.py:1040-1076); sandbox_runtime.prepare_agent(...)returnsAgentBindings(public_agent, execution_agent)and possibly rewritten input (run.py:1079-1085);- new session input is saved before the first model call (
run.py:1111-1123); current_turn += 1; overmax_turnsraisesMaxTurnsExceededunless an error handler supplies a final output (run.py:1591-1616).
- first-turn input guardrails: blocking ones run before sandbox preparation (
- One turn.
run_single_turn(run_internal/run_loop.py:2665) runs agent-start hooks, gets enabled tools (get_all_tools, which also lists MCP tools), resolves the system prompt and handoffs, and resolves tool-name collisions (run_loop.py:2696-2723). It builds the model input as caller input + replayable generated items with orphan calls pruned (_prepare_turn_input_items,run_loop.py:347-354), or delta-only input when the server manages the conversation. - Model call.
get_new_response(run_loop.py:2800) appliescall_model_input_filter, de-duplicates input, picks the model (get_model,turn_preparation.py:134), resetstool_choiceif this agent already used a tool (maybe_reset_tool_choice,tool_execution.py:561-569), adds a stableprompt_cache_key, and callsmodel.get_response(...)throughget_response_with_retry(run_loop.py:2898-2915;run_internal/model_retry.py:574). Usage is added once per accepted response (run_loop.py:2939). - Classification.
process_model_response(run_internal/turn_resolution.py:2926) turns output items into publicRunItems and executable records: function calls, handoffs, computer/shell/apply-patch actions, MCP approval requests, hosted-tool items. A function call whose name is a handoff tool becomes aToolRunHandoff(turn_resolution.py:3583-3594). An unknown tool raisesModelBehaviorErrorunlesstool_not_found_behavior="return_error_to_model"(turn_resolution.py:3615-3641). - Side effects.
execute_tools_and_side_effects(turn_resolution.py:804) builds a plan, runs approvals and tool guardrails, executes function tools concurrently through_FunctionToolBatchExecutor(tool_execution.py:1547), and appends output items in model order. Then, in this order:- pending approvals →
NextStepInterruption(turn_resolution.py:935-950); - handoffs →
execute_handoffs→NextStepHandoff(:537); tool_use_behaviorsays stop →NextStepFinalOutput(:769-801);- no tools and a message → final output, validated against
output_type(:997-1124); - otherwise →
NextStepRunAgain(:1129-1137).
- pending approvals →
- Loop tail.
run.py:1896-1918replacesoriginal_input(a handoff may have filtered it), setsgenerated_items = pre_step_items + new_step_items, appends session items, and persists the turn (run.py:1932-2000). It then branches onnext_step(run.py:2001,:2171,:2284,:2295). A final output runs output guardrails and returns aRunResult.
Interactive versions: execution-flow.html and agent-loop.html.
Core Abstractions
classDiagram
direction LR
class Agent {
name
instructions str or callable
tools
handoffs
output_type
input_guardrails / output_guardrails
tool_use_behavior
as_tool()
}
class Runner {
run() run_sync() run_streamed()
}
class RunConfig {
model, model_provider
call_model_input_filter
handoff_input_filter
sandbox, tool_execution
}
class RunContextWrapper {
context (never sent)
usage
approvals
}
class Model {
<<abstract>>
get_response()
stream_response()
}
class Tool {
<<union>>
FunctionTool
hosted tools
}
class Handoff {
tool_name
on_invoke_handoff()
input_filter
}
class Session {
<<protocol>>
get_items(limit)
add_items()
}
class RunState {
to_json()
approve() reject()
}
class SandboxAgent
class Capability {
tools()
instructions()
sampling_params()
process_context()
}
Runner --> RunConfig
Runner --> Agent : runs current
Runner --> RunContextWrapper
Runner --> Session
Runner --> RunState : on interruption
Agent --> Tool
Agent --> Handoff
Agent --> Model
Handoff --> Agent : target
Agent <|-- SandboxAgent
SandboxAgent --> Capability
Capability --> Tool : adds
Interactive version: core-abstractions.html.
Agent
- Responsibility: declare what an agent is: instructions (string or
(ctx, agent)callable), tools, MCP servers, handoffs, model and settings, guardrails, output type,tool_use_behavior,reset_tool_choice(agent.py:187-418). - Inputs / outputs: none at runtime; the Runner reads it.
get_system_promptresolves callable instructions against the current context (agent.py:1129-1158);get_all_toolsfilters byis_enabledand appends MCP tools (agent.py:286-316). - Lifecycle: constructed by the application; reused across runs.
clone()is a shallowdataclasses.replace. ASandboxAgentinstance must not run concurrently in two runs (.agents/references/sandbox-runtime-boundary.md). - Why it exists: keeping the agent declarative lets the Runner own all control flow, which is what makes interruption/resume and streaming parity tractable.
Runner / AgentRunner
- Responsibility: the loop, turn accounting, guardrail ordering, handoff switching, persistence, tracing, and result construction (
run.py:273-2827; streaming inrun_loop.py:969-2306). - Inputs: starting agent,
input(string, items, orRunState),context,max_turns,hooks,run_config,error_handlers,session, and server-conversation ids (run.py:275-290). - Outputs:
RunResult/RunResultStreamingwithfinal_output,new_items,raw_responses,last_agent, guardrail results,interruptions, andto_state()/to_input_list(). - Why it exists: a single owner for side-effect ordering. The maintainer reference states the invariants explicitly (one turn increment per logical model call; resume never charges a turn twice; streaming and non-streaming must produce equivalent items) (
.agents/references/runner-lifecycle.md).
SingleStepResult + NextStep*
- Responsibility: the control boundary between “one model response and its local side effects” and “what the loop does next” (
run_steps.py:164-248). - Fields:
original_input,model_response,pre_step_items,new_step_items,next_step, plussession_step_itemswhen a handoff filter hid items from the model but history must keep them. - Why it exists: a closed set of four outcomes is what
RunStatecan serialize. Adding a pausable behavior means adding a step variant with streaming, session, tracing and resume semantics, not a path-localreturn.
Tools: FunctionTool and the Tool union
- Responsibility:
FunctionTool= name, description, strict JSON schema andon_invoke_tool(ctx, json) -> Anyplus policy fields:is_enabled,needs_approval,tool_input_guardrails,tool_output_guardrails,timeout_seconds,failure_error_function,defer_loading,allowed_callers(tool.py:454-600).Toolalso includes hosted tools executed by OpenAI (web/file search, code interpreter, image generation, hosted MCP, tool search) and local action tools (computer, shell, apply_patch, custom) (tool.py:1649-1664). - Why it exists: a single executable contract for local code, MCP tools and sub-agents, with hosted tools kept as data the server executes.
Handoff
- Responsibility: a tool (
tool_name,input_json_schema) whose invocation returns the next agent (on_invoke_handoff), plusinput_filter,nest_handoff_historyandis_enabled(handoffs/__init__.py:126-227). Default name:transfer_to_<snake_case(agent.name)>; the tool output is{"assistant": "<name>"}(:210-219). - Why it exists: routing as a model decision, expressed in the model’s native tool-calling vocabulary.
Model / ModelProvider (model abstraction)
- Responsibility:
Model.get_response(system_instructions, input, model_settings, tools, output_schema, handoffs, tracing, previous_response_id, conversation_id, prompt) -> ModelResponse, plusstream_responseand optionalget_retry_advice(models/interface.py:37-135).ModelProvider.get_model(name)resolves strings;MultiProviderroutes by prefix (openai/,litellm/,any-llm/;models/multi_provider.py:62-75). - Contract worth noting: models must assign a non-empty call id to each tool invocation, stable across resume (
interface.py:38-45). - Why it exists: the loop never sees a wire format. The cost is that the item vocabulary is OpenAI’s, so non-Responses backends lose features (probe P17).
RunContextWrapper
- Responsibility: carries the application’s
contextobject (never sent to the model), the run-wideUsage,turn_input, and the approval records (run_context.py:176-200).ToolContextextends it with call id, tool name and arguments. - Why it exists: dependency injection for tools, hooks and guardrails without putting state into the prompt.
Session
- Responsibility: four async methods,
get_items(limit),add_items,pop_item,clear_session(memory/session.py:53-97). Implementations:SQLiteSession,OpenAIConversationsSession(server-side),OpenAIResponsesCompactionSession(wrapper), and inextensions/memory/async SQLite, SQLAlchemy, Redis, Dapr, MongoDB, encrypted and “advanced” SQLite sessions. - Why it exists: client-side conversation memory with no framework dependency; see Memory.
RunState
- Responsibility: everything needed to resume: current turn and agent, original input, generated and session items, model responses, guardrail results, pending step, tool-use tracker, trace state, sandbox resume state and the context’s approvals (
run_state.py:835-945).approve()/reject()record decisions (:1366-1410);to_json()/from_json()(:1885,:2393). - Schema policy: every bump adds a one-line summary to
SCHEMA_VERSION_SUMMARIES; released versions stay readable; newer versions are rejected by older SDKs (run_state.py:237-310).
SandboxAgent + Capability
- Responsibility:
SandboxAgentaddsdefault_manifest,base_instructions,capabilities(default[Filesystem(), Shell(), Compaction()]) andrun_astoAgent(sandbox/sandbox_agent.py:31-64;capabilities/capabilities.py:7-10). ACapabilityhas five hooks:tools(),instructions(manifest),sampling_params(params),process_context(items),process_manifest(manifest)(capabilities/capability.py:16-70). - Why it exists: a sandboxed agent needs tools bound to a live session, prompt text that depends on the workspace, and model params, all prepared per run. Capabilities are the SDK’s equivalent of middleware.
Agent Loop
stateDiagram-v2
[*] --> prepare: Runner.run(input or RunState)
prepare --> model: turn += 1
prepare --> MaxTurnsExceeded: turn > max_turns
model --> side_effects: ModelResponse
model --> Error: unknown tool / guardrail tripwire
side_effects --> decide: outputs
decide --> prepare: RunAgain or Handoff
decide --> interrupted: Interruption
interrupted --> side_effects: resume with RunState
decide --> [*]: FinalOutput (output guardrails)
Interactive version: agent-loop.html.
- Exit rule (fact): a response with no tool calls and a message is final; if any local tool ran, the loop runs the model again so it sees the outputs (
turn_resolution.py:997-1137).tool_use_behaviorcan short-circuit:"stop_on_first_tool",StopAtTools, or a callable (turn_resolution.py:769-801; probe P10). - Structured output: with an
output_type, the last message text is validated as JSON (agent_output.py:61); invalid JSON raisesModelBehaviorErrorunless anerror_handlersentry supplies output (turn_resolution.py:1044-1095). - Bounding:
max_turns(default 10) counts model invocations across all agents of the run (probe P6). Retries insideget_new_responsedo not count as turns. - Loop guard: after an agent uses a tool,
tool_choiceis reset toNoneso"required"cannot force an endless tool loop (tool_execution.py:561-569; probe P12). - Two loops: the streaming loop (
start_streaming,run_loop.py:969) duplicates the control flow of_run_implwith a queue and background task. Parity is a stated invariant enforced by paired test files (tests/test_agent_runner.py,tests/test_agent_runner_streamed.py). - Error handlers:
error_handlerskeyed bymax_turns,model_refusalandinvalid_final_outputcan turn a terminal error into a final output (run_error_handlers.py:50-55;run.py:1608-1616).
Context Engineering
Context engineering here is mostly about assembly, replay hygiene and hooks. The core loop does not shrink context by itself.
flowchart LR
hist[(Session history)] -->|"prepend, limit N items"| orig[original_input]
new[New input] --> orig
out[Turn outputs] --> gen[generated items]
cfg[Agent config] --> sys[System prompt per turn]
orig -->|SandboxAgent only| pc[process_context]
pc --> asm["prepare input<br/>drop orphan calls"]
gen --> asm
asm --> filt[/"call_model_input_filter"/]
sys --> filt
filt --> req[["Model request<br/>+ tool and handoff schemas"]]
orig -.->|"server-managed: deltas only"| srv[(Responses server state)]
Interactive version: context-flow.html.
- How context is constructed. Per turn, in
run_single_turn/get_new_response(run_loop.py:2665-2915): resolved instructions (get_system_prompt),original_input(caller input, with session history prepended) + generated items converted back to input items (run_item_to_input_item,run_internal/items.py:179-209), tool and handoff schemas, output schema and settings.call_model_input_filterthen seesModelInputData(input, instructions)and may return a replacement (turn_preparation.py:51-93). - What enters the model context.
- instructions; for a
SandboxAgentthey are assembled in a fixed order: SDK base prompt, agent instructions, capability fragments, remote-mount policy, filesystem tree (runtime_agent_preparation.py:173-235). Measured: this prompt was 23,254–23,291 characters, about 16.8k of which is the SDK base prompt (probe P14, R6); - the caller’s input and the session history;
- every generated item: messages, function calls and outputs, handoff outputs, reasoning items, hosted-tool items;
- schemas of enabled tools (MCP tools listed on each turn unless cached,
mcp/server.py:1490) and handoffs.
- instructions; for a
- What is excluded.
RunContextWrapper.context(local state; “never added to model input automatically”,.agents/references/agent-definition-and-run-context.md; the reproduction encodes it astest_context_object_is_never_sent_to_the_model);- approval placeholders and SDK-only metadata (
run_item_to_input_itemreturnsNonefortool_approval_item;strip_internal_input_item_metadata); - orphan function calls in runner-generated history (
drop_orphan_function_calls,items.py:211); - disabled tools and handoffs (
is_enabled), anddefer_loadingtools until Responses tool search loads them; - skill bodies and
MEMORY.mduntil the agent opens them (index and summary only, see Memory); - with
conversation_id/previous_response_id: everything the server already has (OpenAIServerConversationTracker.prepare_input,run_internal/oai_conversation.py:518).
- How tool results are handled. The return value becomes a
function_call_outputviaItemHelpers.tool_call_output_item: strings as-is, structuredinput_text/input_image/input_fileoutputs as content lists, other values stringified or validated againstoutput_type(items.py:845-900). Exceptions go throughfailure_error_function; the default returns the fixed string “An error occurred while running the tool. Please try again.” (tool.py:1980-1985), including for unparsable arguments (probe P4).custom_data_extractorattaches SDK-only data that is never sent. - Truncation and offload. Fact: no default truncation of function-tool outputs. Probe P9 replayed a 5,000-character output verbatim. Bounded outputs exist only where a tool opts in: sandbox
exec_commandtruncates only when the model passesmax_output_tokens(defaultNone,capabilities/tools/shell_tool.py:144-149;util/token_truncation.py:57-61), andShellToolactions and executor results carry an optionalmax_output_length(tool.py:1421-1436).ToolOutputTrimmer(opt-in, commitbc9dbd7d) replaces old outputs with a preview in the model view only (extensions/tool_output_trimmer.py:88-140). There is no offload-to-file mechanism in the core loop. - Summarization / compaction. Three mechanisms, none on by default for a plain
Agent:- server compaction:
ModelSettings.context_management=[{"type": "compaction", "compact_threshold": ...}]is passed to the Responses API (model_settings.py:191-197;models/openai_responses.py:1057). The sandboxCompactioncapability (in the default set) sets it to 90% of the model’s context window, or 240k tokens for unknown models (capabilities/compaction.py:162-208), and drops every item before the newestcompactionitem (:210-225; probe P15); - session compaction:
OpenAIResponsesCompactionSessionwraps a session and callsresponses.compactafter a turn when ≥ 10 candidate items exist, replacing stored history (memory/openai_responses_compaction_session.py:36-69, 91-186; trigger insession_persistence.py:678-740). It defers compaction while local tool outputs are pending; - handoff nesting:
nest_handoff_history=Truecollapses the transcript into one<CONVERSATION HISTORY>message. This is a textual dump, not an LLM summary (handoffs/history.py:83-168; probe P2).
- server compaction:
- Repository context. Not indexed. A
SandboxAgentgets a depth-3 rendering of its manifest in the prompt (_filesystem_instructions,runtime_agent_preparation.py:48-77) and explores withexec_command(“Preferrg”,capabilities/shell.py:16-26). - How sub-agents receive context.
- Handoff (default): the full raw transcript, including the handoff call and its
{"assistant": ...}output (probe P1). With real models the next agent’s input also contained the previous agent’s reasoning item in 6/6 runs (R1). Whether that reasoning is replayed on the wire depends on the adapter: for Chat Completions it is replayed only to DeepSeek-family models (models/reasoning_content_replay.py,default_should_replay_reasoning_content). input_filtermay rewriteinput_history,pre_handoff_items,new_items;input_itemschanges only the model view whilenew_itemsstay in session history (handoffs/__init__.py:158-180;turn_resolution.py:640-720).Agent.as_tool: only the generatedinputstring, or a structured input built byinput_builder(agent.py:721-790; probe P3; R2: 12/12 child calls received exactly one user item).
- Handoff (default): the full raw transcript, including the handoff call and its
- How context isolation works.
as_toolisolation is structural: a freshRunner.runwith a freshToolContext(so approvals do not leak), sharing the application context object and usage accumulator (agent.py:758-790). The parent receivesfinal_output, or the last non-empty message / tool output when the final output is empty (agent.py:1073-1100). Handoffs have no isolation by default; isolation is the filter’s job, and the docstring warns that nesting is “not a redaction mechanism” (handoffs/__init__.py:158-176). Server-managed conversations reject handoff input filters outright (turn_resolution.py:505-534). - How context size is controlled over long runs. By configuration, not by default:
max_turnsbounds steps;SessionSettings.limitbounds items (and sessions also store reasoning items: a real run stored[user, reasoning, message]for one exchange, solimit=2kept only the last reasoning and message — R4);session_input_callback;call_model_input_filter(e.g.ToolOutputTrimmer); server compaction; server-managed state that sends deltas. A stable per-runprompt_cache_keykeeps the prefix cacheable across turns (run_internal/prompt_cache_key.py:17-120). Interpretation: the SDK treats context policy as an application decision and leans on the Responses API for automatic compaction; on other backends the application must supply the policy.
Memory
| Concept | What it is here | Where it lives | Scope | Written by | Read by |
|---|---|---|---|---|---|
| Conversation history | stored input items: user messages, assistant messages, reasoning, tool calls and outputs | a Session backend (SQLite, Redis, SQLAlchemy, Dapr, MongoDB, encrypted wrapper) or the server (OpenAIConversationsSession, conversation_id, previous_response_id) |
session_id / conversation |
the Runner: new input before turn 1, each settled turn after it (run.py:1111-1123, :1932-2000) |
prepare_input_with_session, prepended to the next run |
| Context | one ModelInputData plus schemas |
memory only, rebuilt every turn | one model call | run_single_turn + filters |
the model |
| Persistent memory | sandbox Memory files: memory_summary.md, MEMORY.md, skills/, rollout_summaries/ |
the sandbox workspace (memories/ by default, sandbox/config.py:27-33) |
one workspace; rollouts grouped by conversation/session/group id (run.py:245-257) |
two background agents at session close; the main agent too if live_update=True |
Memory.instructions injects the summary; the agent greps MEMORY.md on demand |
| Local state | RunContextWrapper.context |
process memory | one run | application, tools | tools, hooks, guardrails — never the model |
| External storage | RunState JSON, sandbox snapshots, server conversations |
wherever the app stores them | app-defined | to_json(), sandbox cleanup |
from_json(), sandbox resume |
flowchart LR
run[Runner.run] -->|add_items per turn| sess[(Session)]
sess -->|get_items next run| next[next input]
sess -.->|if wrapped| cmp[responses.compact]
sb[SandboxAgent run with Memory] -->|append segment| roll[(rollout JSONL)]
roll -->|on session close| p1[Phase 1 extract agent]
p1 -->|raw memories| p2[Phase 2 consolidate agent]
p2 -->|edit files| mem[(memory folder)]
mem -->|memory_summary.md| prompt[next prompt]
Interactive version: memory-flow.html.
- Session memory is implicit and complete. Nothing is selected or summarized: whatever the runner generated is appended, and the next run prepends it. Only
limit,session_input_callbackand compaction wrappers change that. - Sandbox memory is explicit and agent-written (
sandbox/memory/):- during a sandbox session each run’s result is appended to
sessions/<rollout_id>.jsonl(memory/manager.py:89-117); flush()is registered as a pre-stop hook of the sandbox session (manager.py:64). At close, a phase-1SandboxAgentwith structured output extracts{rollout_slug, rollout_summary, raw_memory}per rollout (defaultgpt-5.4-mini; rollout truncated to 150k tokens) (memory/phase_one.py:15-126;config.py:48);- a phase-2
SandboxAgent(defaultgpt-5.5, up to 500 turns) consolidates the selected raw memories (≤ 256) into the memory folder with its shell and file tools (memory/phase_two.py:10-37;config.py:45-83); - the read side injects
memory_summary.md(truncated to 15k tokens) into the system prompt with a “quick memory pass” protocol: grepMEMORY.md, open at most 1–2 rollout summaries or skills (capabilities/memory.py:50-90;memory/prompts/memory_read_prompt.md).
- during a sandbox session each run’s result is appended to
- Staleness is handled by prompt policy: “memory is guidance, not truth: current evidence wins”; with
live_update=Truethe agent must fixMEMORY.mdin the same turn when it detects a conflict (memory/prompts.py:36-50). Interpretation: this is the Codex “memories” design moved into a reusable SDK, and it dogfoods the SDK: the memory writers areRunner.runcalls onSandboxAgents. - Skills are procedural memory with progressive disclosure:
SkillsmountsSKILL.mdfolders into the workspace (.agents/by default), puts onlyname: description (file: path)into the prompt, and tells the model to openSKILL.mdwhen a task matches (capabilities/skills.py:621-984). In lazy mode aload_skilltool materializes one skill on demand. Probe P14: the skill body was not in the prompt; R6: both models openedSKILL.mdbefore the changelog in 6/6 runs and followed its marker instruction.
Tools
- Interface:
FunctionTool(tool.py:454);@function_toolderives a strict JSON schema from the signature and docstring (tool.py:2574-2870,function_schema.py). Sync functions run inasyncio.to_thread(tool.py:2783-2785); only async tools support timeouts. - Registration and discovery: static
Agent.tools, filtered per turn byis_enabled(bool or callable), plus MCP tools fetched fromAgent.mcp_serverson every turn (cacheable withcache_tools_list=True,mcp/server.py:1490). Name collisions between tools and handoffs followtool_name_collision_policy(run_config.py:490). Responses tool search can defer tool definitions (defer_loading). - Invocation: planning precedes side effects (
tool_planning.py); function calls in one response run concurrently with outputs kept in model order (probe P13: 0.41 s for 0.4 s + 0.1 s tools);RunConfig.tool_execution.max_function_tool_concurrencybounds local concurrency (run_config.py:136-146). - Result handling: see Context Engineering §4.
tool_use_behaviorcan make a tool result the final output. - Error handling:
- tool exceptions →
failure_error_function(default fixed text;Nonere-raises) (tool.py:640-695); - timeouts →
timeout_behavior="error_as_result"(model-visible text) or"raise_exception"; - unknown tool →
ModelBehaviorErrorby default (probe P5); - model refusal →
ModelRefusalErrorunless handled (turn_resolution.py:1003-1036).
- tool exceptions →
- Retries: none at tool level. Model calls retry through
get_response_with_retrywithModelRetrySettingsand provider retry advice; stateful requests (previous_response_id) and replay-unsafe requests (programmatic tool calling) restrict replays (run_internal/model_retry.py:574-600). - Permissions:
needs_approval(bool or per-call callable) → interruption; tool input guardrails run immediately before invocation (and, withToolExecutionConfig.pre_approval_tool_input_guardrails=True, also before the approval pause,run_config.py:146-150); tool output guardrails run before the output is accepted (tool_execution.py:2732-2805). MCP servers have their ownrequire_approvalpolicies, and hosted MCP approvals surface as interruptions too. - MCP: local servers (stdio, SSE, streamable HTTP) are converted to
FunctionTools (mcp/util.py:266-560);HostedMCPToolis executed by OpenAI (tool.py:1141). Lifecycle of local servers is the caller’s (agent.py:201-210;MCPServerManager). - Browser / computer:
ComputerTooldrives a caller-providedComputerimplementation; there is no bundled browser.
Runtime / Sandbox
Where the reasoning/execution boundary sits (fact). The model only emits output items; process_model_response turns them into execution records; tool_planning / tool_execution decide what runs; the code that runs is either application code (function tools), an MCP server, a sandbox session (capability tools), or OpenAI’s servers (hosted tools). Policy gates (approval, tool guardrails, is_enabled, allowed_callers) sit between classification and invocation.
Interactive version: tool-runtime.html.
| Runtime | Isolation | Source |
|---|---|---|
| Function tools | none: the application’s process | tool.py |
| Hosted tools | OpenAI’s infrastructure; the SDK only records items | turn_resolution.py:3290-3420 |
UnixLocalSandboxClient |
temp workspace on the host; Linux commands run “without OS confinement added by this backend”, macOS uses sandbox-exec without network isolation (examples/sandbox/unix_local_runner.py:1-8) |
sandbox/sandboxes/unix_local.py:231 |
DockerSandboxClient |
container, with persisted network-isolation state | sandbox/sandboxes/docker.py:220 |
| Remote providers | E2B, Modal, Daytona, Runloop, Vercel, Cloudflare, Blaxel | extensions/sandbox/* |
- Session ownership: a caller-provided live session is caller-owned; a session created through
SandboxRunConfig.clientis runner-owned and cleaned up (pre-stop hooks such as memory flush, snapshot persistence, provider deletion) at the end of the run (sandbox/runtime.py:309-320;.agents/references/sandbox-runtime-boundary.md). - Preparation per run:
SandboxRuntime.prepare_agentclones capabilities, ensures the session, validates the workspace scope, binds capabilities, runsprocess_contexton the input and builds the execution clone (sandbox/runtime.py:209-307). Hooks and results keep the public agent (AgentBindings,run_internal/agent_bindings.py:17-38; probe P14last_agent is agent). - Manifest trust: host paths (
LocalDir,LocalFile) need trusted, application-controlled grants; a dictionary manifest cannot authorize host access (sandbox_agent.py:39-43). The trust rules fill 84 lines of.agents/references/sandbox-runtime-boundary.mdand about 2.3k lines ofsandbox/_mount_security.py. - Portability gap (fact, probe P17): the default capabilities only work with the Responses API. With
OpenAIChatCompletionsModel,Filesystem’sapply_patchis aCustomToolthat the Chat Completions converter rejects (UserError: Hosted tools are not supported…), andCompaction’scontext_managementgoes intoextra_argsand is forwarded as a keyword argument (openai_chatcompletions.py:740), which the OpenAI client rejects (TypeError). OnlyShell(+Skills,Memoryread) worked; the real-model R6 runs used that set. I found no mention of this constraint indocs/sandbox/.
Sub-Agents / Workflow
sequenceDiagram
participant R as Runner
participant P as Parent model
participant T as Handoff target
participant W as as_tool wrapper
participant N as Nested Runner
participant C as Child model
R->>P: turn (items + handoffs)
P-->>R: transfer_to_target()
R->>T: same transcript + transfer output
T-->>R: final message (last_agent = target)
R->>P: turn (items + tools)
P-->>R: tool(input = brief)
R->>W: invoke FunctionTool
W->>N: Runner.run(child, brief)
N->>C: [user: brief] only
C-->>N: final message
N-->>W: final_output
W-->>R: function_call_output
R->>P: next turn (parent keeps control)
Interactive version: sub-agent-flow.html.
- Handoffs move control. Only the first handoff in a response is executed; extra ones get the output “Multiple handoffs detected, ignoring this one.” (
turn_resolution.py:574-588). Input guardrails of the target do not run (they belong to the starting agent, first turn only). Themax_turnsbudget is shared. - History policy changed twice (fact, commits): nested history became the default in
a776d809(#1996, 2025-11-17) and was moved back to opt-in in6ab83d43(#2272, 2026-01-20) “while we stabilize nested handoffs”; the field now readsnest_handoff_history: bool = False(run_config.py:378). Agent.as_toolkeeps control. The nested run has its own loop,max_turns, approval scope and resumable state; nested interruptions bubble up through the parent’sRunState(run_state.py:1366-1382:_find_nested_approval_state).on_streamcan forward nested stream events, with a bounded queue (agent.py:606-660).- Orchestration is model-driven. There is no DAG engine. Deterministic orchestration is plain Python around
Runner.run(the docs’ “orchestrating via code”);examples/agent_patterns/shows routing, parallelization and LLM-as-judge built that way. - Experimental extras (not traced in depth):
extensions/experimental/codex/exposes the Codex CLI as a tool;extensions/experimental/hosted_multi_agent/providesOpenAIHostedMultiAgentModel, aModelfor “Responses hosted multi-agent” (beta Responses connection); I did not trace it.
Important Source Files
| File | Why read it |
|---|---|
src/agents/run.py |
Runner; the whole non-streaming loop (_run_impl, :623-2400) |
src/agents/run_internal/run_loop.py |
run_single_turn, get_new_response, and the streaming loop |
src/agents/run_internal/turn_resolution.py |
output classification, handoffs, final-output rules, interrupted-turn resume |
src/agents/run_internal/tool_planning.py, tool_execution.py |
approval planning, concurrent function tools, guardrails, failure conversion |
src/agents/run_internal/run_steps.py |
ProcessedResponse, NextStep*, SingleStepResult |
src/agents/run_internal/session_persistence.py, oai_conversation.py |
client-side history vs server-managed deltas |
src/agents/agent.py |
Agent, as_tool |
src/agents/handoffs/__init__.py, handoffs/history.py |
handoff tools, filters, nested history |
src/agents/tool.py |
tool types, @function_tool, default failure handling |
src/agents/models/interface.py, chatcmpl_converter.py, openai_responses.py |
model boundary and adapters |
src/agents/run_state.py |
resumable state and schema policy |
src/agents/sandbox/runtime.py, runtime_agent_preparation.py, capabilities/ |
sandbox preparation and capabilities |
src/agents/sandbox/memory/ |
two-phase memory generation |
.agents/references/*.md |
maintainer invariants (runner lifecycle, run items, tool execution, sandbox boundary) |
src/agents/testing/model.py |
ScriptedModel, the deterministic model used by the probes |
5+ Implementation Decisions Worth Learning From
1. A closed NextStep state machine as the only control boundary
- What they did: each turn returns a
SingleStepResultwhosenext_stepis one of four variants (run_steps.py:164-248); the loop branches only on that (run.py:2001-2300).RunStateserializes the current step, so an interruption is just a step the loop can return from and later re-enter (resolve_interrupted_turn,turn_resolution.py:1173). - Problem solved: tool execution, handoffs, approvals, guardrails, streaming and persistence all need a single, ordered notion of “what happens next”.
- Why it is interesting: pausability falls out of the design. The maintainer rule “do not bypass this state machine with path-local completion logic” (
runner-lifecycle.md) keeps it that way. - Trade-offs: the streaming loop duplicates the non-streaming one (parity by tests and review, not by construction); every new step type must define streaming, session, tracing and resume behaviour;
_run_implgrew to ~1,800 lines. - Where in source:
run.py,run_internal/run_steps.py,run_internal/turn_resolution.py. - Reuse: make “what next” a small sum type returned by the turn function, and keep the loop a dumb dispatcher over it. Serialize that value, not the call stack.
2. Everything the model can do is a tool call
- What they did: handoffs are function tools named
transfer_to_<agent>(handoffs/__init__.py:214-226, converted like any function at the wire,chatcmpl_converter.py:1040); sub-agents areFunctionTools running a nestedRunner(agent.py:606-1127). The Runner gives meaning to a call by looking it up (turn_resolution.py:3583-3640). - Problem solved: one model contract across providers, and no special prompt protocol for routing or delegation.
- Why it is interesting: “delegate and return” vs “transfer control” is a runtime choice, not a model capability. Real models used both correctly in 12/12 runs (R1, R2).
- Trade-offs: handoff and tool names share one namespace (collision policy needed); only one handoff per response wins; handoff arguments are metadata, not the next agent’s input.
- Reuse: model control transfer as a tool whose executor mutates runner state, and keep delegation as an ordinary tool with a narrow input.
3. One item vocabulary from model to storage
- What they did: the Responses item format is used for model input, run items, session storage and
RunState; adapters convert at the edge (chatcmpl_converter.py:534), and replay helpers repair it (drop_orphan_function_calls, reasoning-ID policy, de-duplication;run_internal/items.py). - Problem solved: resume, sessions, server-managed conversations and provider switching all need the same representation.
- Why it is interesting: history repair becomes a pure function over items (orphan pruning, reasoning items dropped with their call). Reasoning replay policy is explicit per provider (
reasoning_content_replay.py). - Trade-offs: non-Responses backends are second-class: hosted tools,
apply_patchandcontext_managementdo not translate (probe P17); sessions store provider-specific reasoning items (R4) that may not be portable across models. - Reuse: pick one canonical transcript format early, store it everywhere, and put all provider translation in adapters with explicit capability errors.
4. Human approval as a serializable pause
- What they did:
needs_approvalproducesToolApprovalItems andNextStepInterruption; the app callsstate.approve()/reject(rejection_message=...)possibly in another process afterto_json()/from_json(); resume executes approved calls and answers rejected ones without re-calling the model (run_state.py:835-2420;tool_execution.py:1880-1925). Schema changes are versioned with one-line summaries and a backward-read policy (run_state.py:237-310). Sibling tools that need no approval still run in the paused turn. - Problem solved: approvals can take minutes or days; a blocking call would hold a process and a model context.
- Why it is interesting: the pause point is data. Probe P7 and real-model R3 (6/6 approve, 5/5 reject with the custom message relayed to the user) confirm no duplicate execution and no extra model call.
- Trade-offs:
run_state.pyis 5.6k lines, and the schema went from 1.0 to 1.20 since HITL landed in January 2026 (3ce7c24d); agents are code, so they must be re-bound by identity on load; held session writes and nested agent-tool state make the edge cases hard (see schema summaries 1.15–1.20). - Reuse: return “interrupted + state” instead of awaiting a human; version the state format from day one and record what each version adds.
5. Tool failures are observations — redacted by default
- What they did: every function tool is wrapped so exceptions become a model-visible string from
failure_error_function; since40956e04(#5112, 2026-09-21) the default is a fixed message that reveals nothing about the error, even for bad JSON (tool.py:1980-1985).failure_error_function=Nonere-raises; an unknown tool raises by default. - Problem solved: exception text can leak secrets (connection strings, paths, customer data) into prompts, traces and logs.
- Why it is interesting: it is a security default with a measurable capability cost. Measured (R5): the tool raised “order_id must contain digits only, e.g. ‘1042’” for
A-1042. With the default text, the model recovered in 1/6 runs; the other five retried the same id and told the user “the lookup service is returning an error, try again later”. With a custom function that returned the message, it recovered in 6/6 runs on the second attempt. - Trade-offs: safety vs self-correction; developers must write an allow-listed error formatter for validation errors.
- Reuse: convert failures into observations, but separate user-correctable errors (validation, not found) from internal errors, and only expose the former.
6. Context policy as hooks, compaction delegated to the server
- What they did: the core loop replays everything; shrinking is done by
call_model_input_filter(turn_preparation.py:51),session_input_callback,SessionSettings.limit, the opt-inToolOutputTrimmer(commitbc9dbd7d), Responsescontext_managementcompaction, andOpenAIResponsesCompactionSession(commit09443fd0). - Problem solved: one SDK serves chat apps, short tool loops and long sandbox jobs; a single built-in policy would be wrong for most of them.
- Why it is interesting: the final hook sees exactly
ModelInputData(input, instructions)before every call and changes only the per-call view, so stored history stays intact (probe P9; reproduction testtest_trimmer_shrinks_old_tool_outputs_only_in_the_model_view). - Trade-offs: no protection by default — a single large tool output stays in every later request;
limitcounts items (including reasoning), not tokens or turns (R4); automatic compaction effectively requires OpenAI’s Responses API. - Reuse: expose one “last look before the model” hook with the full request, and keep trimming in the view, not in storage.
7. Public agent vs execution agent, prepared by capabilities
- What they did:
SandboxRuntime.prepare_agentclones capabilities per run, binds them to the live session, and builds an execution clone with capability tools, ordered prompt fragments and merged sampling params;AgentBindingskeeps the public agent for hooks, results and handoffs (sandbox/runtime.py:209-307;runtime_agent_preparation.py:86-157;agent_bindings.py:17-38). - Problem solved: tools bound to a live sandbox session cannot live on a reusable agent definition, and identity must not change because an internal clone did the work.
- Why it is interesting: it is middleware without a middleware stack: five hooks per capability, applied once per run (probe P14; reproduction
test_capabilities_are_cloned_per_run). - Trade-offs: capability order and dependencies matter (
required_capability_types); the sandbox prompt is large (~23k characters before any task text); defaults assume the Responses API (P17); aSandboxAgentcannot run in two runs at once. - Reuse: separate the agent you configure from the agent that executes, and derive the latter per run from pluggable capabilities.
8. Memory written by agents, read by progressive disclosure
- What they did: sandbox runs append rollouts; at session close a phase-1 agent extracts structured raw memories and a phase-2 agent consolidates them into
MEMORY.md,memory_summary.md, skills and rollout summaries; only the summary is injected, everything else is searched on demand (sandbox/memory/;capabilities/memory.py). - Problem solved: cross-session learning without a vector store, and without paying for all memory on every call.
- Why it is interesting: memory writing is offline and batch, so the main agent’s latency is unaffected; the reading protocol is explicit about verification and staleness.
- Trade-offs: memory only appears after the sandbox session closes (unless
live_update); quality depends on two extra model runs (cost); default models are OpenAI-specific; I did not run generation end to end with a real model. - Reuse: split memory into a small always-loaded summary plus a grep-able handbook, and write it from transcripts in a background pass.
9. Guardrails trade latency against side-effect safety
- What they did: input guardrails default to
run_in_parallel=Trueand race the first turn;run_in_parallel=Falseblocks before the model call (guardrail.py:72-110;run.py:1040-1076,:1787-1850). - Measured: a blocking guardrail prevented any model call (P11); a slow parallel one let the model call and the tool side effect happen before the tripwire, in both streaming and non-streaming mode (P16). This matches
docs/guardrails.md:36. - Reuse: make the ordering explicit per guardrail, and put irreversible tools behind approvals or tool guardrails rather than relying on agent-level input guardrails.
Minimal Reproduction
experiments/openai-agents-python/ contains miniagents, about 1,700 lines of stdlib-only Python (plus a 170-line demo). It reproduces the architecture, not the product.
| miniagents | reproduces |
|---|---|
run.py::Runner.run + NextStep* |
the loop and its four exits; turn accounting; resume continues the paused turn |
run.py::process_model_response / execute_tools_and_side_effects |
classification by name, approval planning, concurrent tools in model order, handoff, stop_on_first_tool, final output |
run.py::execute_handoff |
first handoff wins; full transcript by default; input_filter and nested history change the model view only |
run.py::RunState |
JSON pause/resume with agents re-bound by name; held session write until the paused turn settles |
agent.py::as_tool |
nested Runner.run on the brief only, shared app context |
tool.py |
strict schema from signature, to_thread, fixed redacted error text |
session.py, trimmer.py |
append-only history with limit; view-only trimming |
capabilities.py |
public vs execution agent; Shell, Skills, Memory, Compaction as capability hooks over a temp-dir workspace |
model.py |
ScriptedModel and a stdlib Chat Completions adapter (items ↔ messages, handoffs as tools) |
The demo (./run.sh demo) runs a support desk: triage hands off to refunds; refunds calls a lookup tool and a policy sub-agent in one turn; the refund tool pauses for approval; the state goes through JSON, is approved and resumes; a second run on the same session sees the history; a blocking guardrail stops an injection attempt; a capability agent answers from its workspace. It asserts 11 invariants.
Verification
All commands ran in this session’s container (Python 3.13.16, uv 0.11.32, Node 22).
1. Upstream test suite (pinned commit, locked dependencies).
uv sync --all-extras --all-packages --group dev --frozen
uv run --frozen pytest -q -n 8 --dist worksteal -m "not serial" # -> 12084 passed, 34 skipped in 157.8s
uv run --frozen python .github/scripts/run_serial_tests.py # -> 90 passed, 4 skipped
The first attempt, from a checkout whose scratch path contained -home-user-, had 62 failures, all in tests/mcp/test_server_errors.py. Those tests assert that the URL credential user:s3cr3t_pw never appears in a rendered traceback; the substring user came from the file paths in the traceback, not from the credential. Re-running from a copy whose path contains no user gave 0 failures. Environment artifact, not a defect.
2. Runtime probe of the real SDK (experiments/openai-agents-python/upstream_probe/probe_openai_agents.py). It drives the real Runner with the SDK’s own agents.testing.ScriptedModel (no network). Result: 17/17 predictions from source matched.
| Probe | Expected (from source) | Observed |
|---|---|---|
| P1 handoff | a tool; full raw transcript to the target | handoffs=[transfer_to_billing], tools []; Billing input [user, function_call, function_call_output]; output {"assistant": "Billing"} |
P2 nest_handoff_history |
one summary message | one assistant message with <CONVERSATION HISTORY> |
P3 as_tool |
child sees only the brief | child input [user: "find the answer"]; parent secret absent; parent got "research result: 42" |
| P4 tool error + bad JSON | fixed text, loop continues | both outputs = default message; hunter2 absent |
| P5 unknown tool | raise; opt-in error text | ModelBehaviorError; opt-in output Tool 'ghost' not found. |
P6 max_turns=2 |
2 model calls | MaxTurnsExceeded after exactly 2 |
| P7 approval | pause, JSON round-trip, one execution | interrupted before execution; schema 1.20; executed once; 2 model calls in total |
P8 session + limit=1 |
history prepended; tail only | run 2 input [user, message, user]; limit=1 kept one item |
| P9 default vs trimmer | verbatim vs preview | 5,002 vs 273 serialized chars |
P10 stop_on_first_tool |
tool output final | "Paris: sunny", 1 model call |
| P11 guardrails | blocking: no model call | blocking 0 calls; parallel tripped after 1 call |
| P12 per-turn resolution | instructions re-resolved; tool_choice reset |
v1, v2; required → None |
| P13 concurrent tools | ~max, model order | 0.406 s; [done-a, done-b] |
P14 SandboxAgent |
capability tools, ordered prompt, compaction param, public agent | tools apply_patch, exec_command, view_image, write_stdin; prompt 23,254 chars in order; skill body absent; memory summary present; context_management threshold 240000; last_agent is agent |
P15 Compaction.process_context |
cut before newest compaction | [compaction, user] |
| P16 slow parallel guardrail | side effect may happen first | send_email executed, then tripwire, in both modes |
| P17 sandbox on Chat Completions | (inferred) Responses-only defaults | default → UserError (apply_patch); Compaction → TypeError (context_management); Shell only → reached the network |
3. Reproduction.
cd experiments/openai-agents-python && ./run.sh
# [1/4] venv + pip install -e .[test] [2/4] wheel built: miniagents-0.1.0-py3-none-any.whl
# [3/4] 40 passed [4/4] demo: 11/11 checks PASS, exit 0
4. Mutation check (mutation_check.py): 11 architectural regressions injected one at a time; 11/11 killed: handoff nests by default (1 test failed), tool exceptions crash the run (1), approval skipped (4), resume re-runs the model (4), paused turn persisted immediately (1), session history not prepended (3), capabilities mutate the public agent (1), tools run sequentially (1), tool_choice never reset (1), guardrails on every turn and agent (1), sub-agent inherits the parent transcript (1). All 40 tests passed after restore.
5. Real-model runs of the real SDK (real_model/run_ark.py). The pinned SDK used OpenAIChatCompletionsModel against Volcano Engine Ark’s OpenAI-compatible endpoint with two models, deepseek-v4.1-flash (deepseek-v4-1-flash-260910) and doubao-seed-2.1-pro (doubao-seed-2-1-pro-260915), 3 runs each. A recording Model wrapper captured every request. Raw results: real_model/results_*.json.
| Scenario | What it tests | Result |
|---|---|---|
| R1 handoff | routing; what the target receives | 6/6 routed to Billing; target input = [user, reasoning, function_call, function_call_output] in 6/6 |
R2 as_tool fan-out |
briefs and isolation | 6/6 called both translators in one turn (2 orchestrator calls); 12/12 child inputs were one user item with only the sentence; customer id leaked 0/12 |
| R3 approval | approve / reject after a JSON round-trip | approve: 6/6 paused, 0 executions before, exactly 1 after; reject: in 5/6 runs the model called the tool (once it asked the user instead); 5/5 relayed the “support ticket” rejection message; 0 executions |
| R4 session | recall; limit=2 |
6/6 recalled “Rust” with full history (4 items); 6/6 answered UNKNOWN with limit=2 (3 items). A separate run showed sessions store [user, reasoning, message] per exchange |
| R5 tool errors | redacted vs visible error | redacted: recovered 1/6; visible: 6/6 on the 2nd attempt |
R6 sandbox (Shell, Skills, Memory) |
skills and memory in a real run | 6/6 opened SKILL.md, then CHANGELOG.md; 6/6 ended with the skill’s marker line; prompt 23,291 chars; in 1 run the model used absolute host paths (/tmp/sandbox-local-…) |
Each scenario took 4–27 s per run. What this shows: the design assumptions held for two current models on small, unambiguous tasks, and the redaction default has a measurable cost. It is n=3 per model, so it is evidence, not a benchmark.
6. Reproduction with real models (real_model/run_mini_ark.py): the same two models drove miniagents through its stdlib Chat Completions adapter: 6/6 runs routed, briefed the policy sub-agent, paused on the refund, resumed after a JSON round-trip with exactly one model call, and executed the refund once.
7. Diagrams. All 8 Archify diagrams passed finalize --quality showcase --repo-root <clone> (schema validation, verified delivery, strict provenance check, headless-Chromium browser check) and visual-check; the light 1440×900 captures were inspected by eye. See assets/openai-agents-python/archify/README.md.
Known limitations of the verification:
- Real-model runs used Chat Completions on non-OpenAI models; the Responses API paths (server-managed conversations, hosted tools, server compaction, WebSocket transport) were verified only by reading source and the upstream tests.
- Sandbox memory generation (phase 1/2) was not run with a real model; its default models are OpenAI models. Only the read side was exercised.
- Streaming was exercised only in probe P16. Realtime, voice, tracing export and the provider sandboxes (Docker, E2B, Modal, …) were not run.
miniagentshas no streaming loop, no tracing, no MCP and no real sandbox isolation.
What I Would Reuse
- A closed “next step” sum type returned by the turn function, with the loop as a dispatcher, so pausing and resuming are data.
- Everything as a tool, with the runtime deciding whether a call executes code, delegates (nested run) or transfers control.
- One canonical transcript format in memory, storage and pause state, with providers behind adapters that fail loudly on missing capabilities.
- Approvals as serializable interruptions, with a versioned state schema and a summary line per version.
- Errors as observations with an explicit exposure policy: expose validation errors, redact internal ones. The R5 numbers argue against a single global default either way.
- A last-look input filter that changes only the per-call view.
- Public vs execution agent, with per-run capabilities contributing tools, prompt and params.
- Things I would change:
- make the default tool-error policy distinguish user-correctable errors;
- give
SessionSettings.limita turn- or token-based mode (it counts reasoning items today); - fail fast, at agent preparation, when sandbox capabilities need the Responses API and the model is not a Responses model.
Limitations / Open Questions
- Responses-only sandbox defaults (P17): is this intended? The error appears only at the first model call, and
docs/sandbox/does not mention it. Uncertain about intent; not filed. - Redacted tool errors (R5): the default trades self-correction for safety. The docstring points to
failure_error_functionas the escape hatch; I found no built-in way to mark an exception as safe to show. Open question. - Reasoning items across handoffs: the target agent’s input contains the source agent’s reasoning items (R1). For Chat Completions the default replays them only to DeepSeek-family models; for Responses models it depends on server rules and
reasoning_item_id_policy. Whether a different model family should ever see another agent’s reasoning is not documented. Partly verified. - Size of the core:
run_state.py(5.6k lines) and_run_impl(~1.8k lines) carry many resume and session edge cases (held writes, nested history ownership, compaction acknowledgement). They are well tested (12k tests), but hard to reason about locally. - Not studied in depth: realtime and voice runtimes, tracing processors beyond the exporter endpoint, the experimental Codex and hosted multi-agent extensions, provider sandboxes, programmatic tool calling, computer use.
Further Reading
- Source: openai/openai-agents-python @ 71c2da4, especially
AGENTS.mdand.agents/references/(maintainer invariants; claims I used were checked against code). - Commits:
40956e04(redact default tool failure details, #5112),a776d809/6ab83d43(nested handoff history on, then opt-in),3ce7c24d(HITL andrun_internal/, #2230),09443fd0(Responses compaction session),bc9dbd7d(ToolOutputTrimmer),2d665c9a(Sandbox Agents),05d6850d(scripted model test utilities). - Docs: https://openai.github.io/openai-agents-python/ (agents, running agents, handoffs, sessions, sandbox agents, guardrails, human-in-the-loop).
- Companion study in this repo: deepagents, a middleware harness on LangGraph that owns no loop, for contrast with this SDK, which owns its loop.