Codex turns: steering, approvals, goals, and recovery
What happens inside one Codex run: how a turn starts and is steered, who answers an approval, how a goal joins many turns into one run, how Stop works, and what the desktop does when a turn loses its child process.
Checked against the source on 30 September 2026
On this page
One thread, one run: start, owner, and steering
A Codex run in HB Code is one reader task that follows one turn, or a chain of turns, on one Codex thread. Everything in this chapter hangs on that reader: approvals reach it through the turn route, Stop talks to it through a control channel, a goal gives it the next turn, and recovery gives it a new route. The process, the actor, and the turn route are in the app-server chapter. The outbox that decides when a message is sent is in the message delivery chapter.
| Step | What the desktop does, and what a failure means |
|---|---|
| Open the thread | Sends thread/ |
| Claim the thread | Writes the owner record for that Codex thread ID: app session, run, and the pid of the Codex child. If it fails: Another run owns the thread. The send stops with “Codex thread … is already running for session …”. |
| Start guard | Records that this run is starting a turn in this session. The run cannot be closed by a finisher until the turn is registered. If it fails: Another turn start owns the session. Nothing was sent. |
| turn/ | Writes the sending ledger entry, then sends turn/ |
| Register the turn | Stores the active turn for the app session: runtime pair, run, thread, turn, and a control channel with 16 slots. Then it wakes the steer lane of the outbox. If it fails: The desktop abandons the turn that Codex already started. |
| Read the turn | The reader takes notifications from the turn route until turn/ |
The owner record is a map in desktop memory with the Codex thread ID as its key. It stops two app sessions that link the same thread from running it at the same time. The same run can claim again without an error. A stale owner is removed when its recorded Codex child is dead. When the record has no pid, it is removed after 45 seconds, and only if the run is no longer active. The guard is released after thread/unload and before the finishers wake the outbox, so the next queued message finds the thread free.
Steering adds a message to the turn that runs now. The desktop sends turn/steer on the runtime pair that owns the active turn. It opens no thread, starts no turn, and claims no owner, because the reader that already runs keeps all three.
| turn/steer field | Value |
|---|---|
| thread | The thread of the active turn. |
| expected | The turn ID the desktop registered. Codex must refuse the steer when another turn is active. |
| client | The operation ID of the outbox entry. The same ID names the message in history later. |
| input | The message text and attachments, built like the input of turn/ |
| Answer | What it means, and what happens to the message |
|---|---|
| turn | Codex took the message into the running turn. The operation ID joins the turn’s list of accepted steers, and the acceptance is saved. |
| JSON-RPC error -32600 | Codex refused the steer. It sends this code when no turn is active, when the active turn has another ID, and when the turn is a review or a compaction. The desktop reads this one code as “not admitted”. The message stays queued. |
| turn | Codex accepted it for a turn the desktop did not expect. Uncertain. The desktop does not send it again. Check the timeline. |
| Any other error or a timeout | The request may or may not have arrived. Uncertain, with the same handling. |
| No active turn in the desktop, or a turn of another run | The desktop did not send turn/ |
Stop and send now needs one more call. A steer that Codex accepted may still wait inside the turn as pending input. The desktop sends turn/detachPendingInputForReplay with the same threadId, expectedTurnId, and clientUserMessageId. The answers detachedForReplay and alreadyDetachedForReplay permit the replay in a new turn. The answer applied means the model already has the message, so no replay starts. The answers recording and staleTarget, and any error, also block the replay.
Who answers an approval
Codex asks for approval with a JSON-RPC request, and it waits for the response. The actor of that Codex child keeps each open request in a map and must find someone who can show it. A request that nobody can show is declined, because a turn that waits for an answer that cannot come would never end.
- Turn routeitem/*/requestApproval
A request for a turn that a run reads goes to that reader. The card has no deadline. It closes with an answer, with the end of the turn, or with serverRequest/resolved.
- Observer claimUNROUTED_APPROVAL_CLAIM_TIMEOUT · 30 s
A request with no route is broadcast. The observer of the root run, or of a watched thread, must claim it within 30 seconds. After a claim the card has no deadline.
- Declined by the desktopdecline
At once with no route and no observer, or with no supported decision. After 30 seconds with no claim. With turn/completed, with an abandoned turn, and when the actor stops.
- Answerrun.chatResolveApproval
Offered to every live Codex child, most recently used first. The actor that holds the request checks request, thread, turn, and item ID and writes the response.
| Card | Request method and decisions offered |
|---|---|
| Command | item/ |
| File change | item/ |
| Permissions | item/ |
- Delivery has two paths. A request for a turn with a live route goes to that turn’s reader, which shows the card in the session. A request with no route is broadcast to observers. Requests from subagent threads take this path, because no reader owns their turns.
- An observer must claim a broadcast request within 30 seconds. The run’s observer first proves that the thread belongs to its root thread, with a limit of 10 seconds for that check. An unclaimed request is declined when the 30 seconds end.
- A request with no supported decision, and a request with no route and no observer, are declined at once.
- Codex can send the same request again after a reconnect. An identical request keeps its place. A different request with the same ID makes the desktop decline that ID and close the card.
- When turn/completed arrives, the actor declines every open request of that turn before the reader sees the completion. Abandoning a turn and shutting the actor down do the same.
- serverRequest/resolved from Codex closes the card as resolved elsewhere. During the initialize handshake no card can exist, so an approval request is declined.
An answer from the desktop window or from a phone takes one path. A phone calls run.chatResolveApproval, and the desktop offers the answer to every live Codex child, most recently used first. Each actor compares the request ID, thread ID, turn ID, and item ID with its map, so exactly one child accepts it. A decision that the request did not offer is refused.
The access level Unrestricted sends approvalPolicy never. If a request still arrives, the reader answers it without a card. It picks the first decision the request offers from acceptForSession, accept, and accept broadly. The other access levels send approvalPolicy on-request.
Goals: many turns in one run
A goal is state that Codex keeps on a thread: an objective, a status, a token budget, and counters. While the status is active, Codex starts the next turn by itself when a turn ends. The desktop does not start these turns. It finds each one and attaches a reader, and the whole chain stays one chat run with one operation ID.
- Goalthread/goal/get · set · clear
State that Codex keeps on the thread: objective, status, token budget, and counters. While the status is active, Codex starts the next turn by itself.
- Attachthread/resume · 3 × 250 ms
After every turn the reader reads the goal. While it is active, the reader resumes the thread and takes the latest turn: a new turn ID, or a turn that is still in progress.
- One runoperation ID
All turns of the chain belong to one chat run with one operation ID. A steer sent between two turns waits for the next turn.
- Stop orderstatus: paused · turn/interrupt
The pause comes first and retries every 250 ms on transport errors. The interrupt has 8 seconds to reach the reader and 8 seconds for the answer. On failure the desktop abandons the turn.
| Method | Use |
|---|---|
| thread/ | Reads the goal. The reader calls it after every turn. |
| thread/ | Sets objective, status, or token |
| thread/ | Removes the goal. |
A goal has the fields objective, status, tokenBudget, tokensUsed, timeUsedSeconds, createdAt, and updatedAt. The status is one of active, paused, blocked, usageLimited, budgetLimited, and complete. The reader follows the chain only while the status is active.
- A goal run starts with thread/goal/set and status active. It sends no turn/start. Codex starts the first turn.
- After each turn the reader sends thread/goal/get. While the goal is active it sends thread/resume with excludeTurns: false and takes the latest turn. A turn counts when its ID is new or when it is still in progress. The reader tries 3 times, 250 ms apart, then reads the goal again.
- A turn that is in progress gets a turn route and is read like any other turn. A turn that already completed is settled from the snapshot in the resume answer.
- Every Codex chat run ends with the same check. A message sent to a thread with an active goal therefore stays one run until the goal leaves the active status. A run that finds no active goal still looks once for a newer turn with thread/resume, and then it ends.
- A steer that arrives between two turns waits. Each new turn wakes the steer lane of the outbox again.
- A goal run unloads its thread with thread/unload at the end. A failed unload there is only logged.
Stop is a ladder of four steps
Stop marks the run as aborted and then works down a ladder. The order matters for goals: a turn that is interrupted while its goal is still active would be followed by a new turn at once.
| Step | Call, limit, and failure |
|---|---|
| Wait for the turn | None. A run that has no registered turn yet is checked every 25 ms. Limit: 45 seconds. If it fails: Stop reports that the cancellation was not acknowledged. The run continues. |
| Pause the goal | thread/ |
| Interrupt | turn/ |
| Abandon | The actor declines the turn’s open approvals, drops the turn route, and writes turn/ |
Codex answers turn/interrupt when the turn has stopped, and it sends turn/completed with the status interrupted. The reader takes that completion as the end of the turn, and the run ends as cancelled. A run that was aborted between turn/start and the registration of the turn is handled by the run itself: it pauses the goal and abandons the new turn, and it retries both every 250 ms on transport errors.
The five-second rule after an error
Codex reports a problem inside a turn with an error notification, and the field willRetry says whether Codex goes on. The error alone never ends the turn. The turn ends with turn/completed, and the reader waits for it, but only for 5 seconds.
- Error that retrieserror { willRetry: true }
The reader ignores it. Codex retries, and the turn goes on.
- Settlement deadlineTURN_FAILURE_SETTLEMENT_TIMEOUT · 5 s
Starts with the first error that won’t retry. The error is saved. turn/completed inside the window ends the turn as failed, whatever status it names.
- Deadline endsFailed
No turn/completed came. The reader fails the turn with the saved error and drops the turn route. It sends no turn/interrupt.
- DrainTURN_POST_TERMINAL_DRAIN_TIMEOUT · 75 ms
After a clean turn/completed the reader takes late item notifications until 75 ms pass with none. Then it closes open approval cards and drops the route.
| What the reader sees | Result of the turn |
|---|---|
| error with will | Nothing changes. Codex retries. |
| error with will | The error is saved and the 5 second deadline starts. A second error replaces the saved one and keeps the first deadline. |
| turn/ | Failed, with the error of the completion, or else the saved error. |
| turn/ | Failed, with the saved error. |
| turn/ | Failed when an error is saved. Otherwise interrupted, and the run ends as cancelled. |
| No turn/ | Failed, with the saved error. The reader drops the turn route. |
| The route closes after a saved error | Failed, with the saved error. No rejoin is tried. |
| turn/ | Completed. The reader keeps reading for 75 ms of silence, then drops the route. |
The reader wakes every 100 ms when no notification arrives. It uses that tick for Stop requests and for the deadline, and a quiet turn is never failed for silence alone. The 75 ms drain exists because Codex can send the last item notifications just after turn/completed. The same 5 seconds apply to a subagent thread: after an error that will not retry, its thread stops counting as active for the child process after 5 seconds.
turn/completed and an error that will not retry are never dropped on a full turn route. They wait in a separate lane in arrival order, and ordinary events for that turn are dropped while one waits.
When the turn route is lost, the desktop asks for the exact turn
The reader has one signal for a lost connection: its turn route closes while no error is saved. The actor closes every route when it stops, for example when the child’s stdout closes. The desktop then asks Codex what became of the turn, and it asks for that exact turn. A turn cannot move to another process. A new Codex child loads the thread from its saved history and reports a turn that was still in progress as interrupted. The answer is inProgress only when the child that answers still runs the turn.
- Lost connectionConnectionLost
The turn route closes and no error is saved. Only this failure starts a rejoin. Any other failure without a final status makes the desktop abandon the turn.
- New clientlease_codex_app_server_for_runtime
The reader waits up to 1 second for the old client to close and leases the same runtime pair. The registry starts a new child if the old one is gone.
- Rejointhread/resume { excludeTurns: false }
The desktop takes the turn with the expected ID. The turn’s recorded access settings, the owner pid, and the observers move to the new client.
- Answerinterrupted · inProgress · completed · failed
A new child reports a turn that was in progress as interrupted. With no error the run fails, or it waits for the next turn of an active goal. A turn is rejoined once.
- Only a lost connection starts a rejoin. Any other failure that is not a final status makes the desktop abandon the turn with turn/interrupt.
- The reader waits up to 1 second for the old client to close. Then it leases a client for the same runtime pair, and the registry starts a new child if the old one is gone.
- The access settings that the old client recorded for the turn move to the new client, and the owner record gets the new pid.
- When the client changed, the observers for subagent approvals and for agent messages are stopped and started again on the new client.
- The rejoin is thread/resume with excludeTurns: false. The desktop takes the turn with the expected ID from the answer. When the answer has no such turn, the run fails and the desktop abandons the turn.
| Status in the answer | What the run does |
|---|---|
| in | The child that answers still runs the turn. The actor opens a new turn route, and the reader registers the turn again and continues. |
| completed | The snapshot settles the run. No route is opened. |
| interrupted with no error | The answer of a new child for a turn that was in progress. A run with an active goal waits for the goal’s next turn. Any other run fails with “Codex app-server turn was interrupted”. |
| interrupted with an error, or failed | The run fails with that error. |
For every answer the desktop first matches the rows on screen to the snapshot of the turn, so the work that Codex saved before the loss stays in the timeline. A turn is rejoined once. When the new route also fails, the run ends with that failure. Between two turns of a goal the reader has no route to lose, so it recovers on transport errors of thread/goal/get and thread/resume instead, at most 2 times in a row.
Child threads, watched threads, rollback
A Codex agent can spawn subagents, and each subagent is a thread of its own below the root thread. No run of the desktop owns their turns. The desktop lists them, saves where their files are, and follows one live only while a window shows it.
| Mechanism | How it works, and its limits |
|---|---|
| Child list | thread/ |
| Child path | thread/ |
| Watched thread | A window that shows a subagent registers a watch for the pair of session and thread. The first watch starts a follower that turns the thread’s unrouted notifications into live rows. The follower stops when the last watch is released and no turn of the thread is live. Limits: Watch ID at most 256 bytes. The thread’s root must be the session’s root, checked within 10 seconds. |
| Follower | Subscribes to the client’s unrouted activity and approvals. When the client goes away it subscribes again. After the child’s turn ends, it matches the live rows to history. Limits: Resubscribes after 1 second. 3 settlement attempts. |
| Rollback | thread/ |
| Live settings | thread/ |
With no watcher nothing follows a subagent thread. Its notifications stay unrouted, and the desktop reads the thread from history when you open it. An approval from a subagent does not need a watch: the observer of the root run claims it, as the approvals section describes.