Compass agent session blobs survive resume
Freezes on merge; later changes supersede by citation, never rewrite.
Issue: RIG-1582. Parent record: compass-agent session persistence (OQ-R3 deferred blobs to this issue). Storage ruling: RIG-4757 option B (Matt, 2026-10-06).
Problem / Intent
Section titled “Problem / Intent”A resumed agent session loses every inline image. The SDK writes image bytes to
a local content-addressed blob directory and keeps only a blob:sha256:<hex>
reference in the session JSONL. Compass tees the JSONL to the Server, but the
blob directory dies with the container. On resume the reference does not
resolve, and the model sees the SDK’s “undecodable image data” text. RIG-1582
makes the bytes survive a resume when the Server has an object store, and makes
the loss explicit (“image unavailable after resume”) when it does not.
Approach
Section titled “Approach”Matt ruled the substance (RIG-4757 option B):
- Blob bytes go through the existing Server
ObjectStoreseam under the keyblobs/<sha256>, with a Postgres index row. - Capture is best-effort.
- With no object store there is no blob persistence, and a resumed agent sees an explicit marker in place of the SDK warning.
Everything below is mechanism.
D1 Capture from the tee
Section titled “D1 Capture from the tee”The agent captures blobs from the committed JSONL lines the tee already sees.
In SDK 18.0.11, #lineFor builds each persisted line through
prepareEntryForPersistence. That step externalizes images synchronously
before the line exists (truncateForPersistence,
session/session-persistence.ts):
* Runs in one synchronous tick ... Image externalization happens via the * synchronous blob-store path (`fs.writeFileSync`), so blob bytes are in the * kernel page cache before the JSONL line referencing them is written.BlobStore.putSync (session/blob-store.ts) names the file by the hash of the
stored bytes:
const hash = new Bun.SHA256().update(data).digest("hex");const blobPath = path.join(this.dir, hash);So each blob:sha256:<hex> in a line reaching TranscriptTeeBackend.append
or writeFull (packages/compass-agent/src/session-tee.ts) names an existing
file under the hash the SDK looks up on load. This holds for any position and
for checkpoint rewrites.
Rejected: an onEntryAppended hook. It sees entries before externalization, so
the agent would have to re-derive the SDK’s position and threshold rules to
reproduce each ref. The SDK’s collab host also overwrites that property.
Mechanics:
- The tee gets one option,
onCommittedLine(line). It runs synchronously after the local write and before the frame is teed. It is wrapped intry/catch, so a throw is logged and never reaches the SDK. SessionBlobUploader.offer(line)matchesblob:sha256:([a-f0-9]{64}). It skips hashes that are done or already queued and enqueues the rest. The queue is bounded atSESSION_BLOB_QUEUE_MAX; it drops the oldest entry and counts the drop.- One worker reads
<blobsDir>/<hash>and sends it, with one upload in flight. - A hash enters the done set only on success,
FailedPrecondition, or oversize (overMAX_SESSION_BLOB_BYTES).FailedPreconditionalso disables the uploader for the lifetime. - Two outcomes do not mark a hash done, and each has its own counter: ENOENT
(
missing) and any other error (failed). There is no retry loop. A later line that still references the hash offers it again. - The done set is not seeded from disk. After a resume, the first checkpoint sends every referenced blob once more. The Server’s own-row check (D3) turns that into a no-op with no PUT.
blobsDircomes from the SDK’sgetBlobsDir()(@oh-my-pi/pi-utils, pinned to the SDK’s exact version), and the resolved path is logged once at boot.isReservedEnvKey(packages/compass-agent/src/cli.ts) also reservesPI_CODING_AGENT_DIRandXDG_DATA_HOME, so the env file cannot move the dir off the Runner’s fixed.omp/agent/blobs.
D2 A dedicated best-effort unary
Section titled “D2 A dedicated best-effort unary”Blob bytes use a new AgentGateway unary, PutSessionBlob{sha256, data}. The
Runner forwards it as RunnerService.RelaySessionBlob. The forwarder copies
the shape of Gateway.Board (go/internal/runner/gateway/board.go):
- it resolves the bound session, and a miss is
CodePermissionDenied; - it sends the session id the Runner owns together with the verbatim request;
- it returns the Server’s code unchanged.
Rejected carriers:
- The durable conversation lane. It is delivered-or-erred with a fatal
latch. Its handler,
Hub.commitFrame(go/internal/runnerhub/relay_comms.go), also accepts only transcript entries. - Inline base64 in the teed lines. A checkpoint body with several images would pass the 16 MiB cap and trip the R4 latch. It would also bloat the PG hot tail and go against the ruling.
Size: AgentGateway reads agentmsg.MaxBytes (16 MiB), and the RunnerService
door reads agentmsg.MaxBytes + 1<<20 (runnerMaxReadBytes), so
MaxSessionBlobBytes = MaxBytes - 1<<20 (15 MiB of raw bytes) fits both. No
mime type is sent; the block’s mimeType stays in the JSONL.
D3 Server: verify, PUT, then index
Section titled “D3 Server: verify, PUT, then index”Hub.RelaySessionBlob(ctx, runnerID, req) takes its guard order from
Hub.CommitConversationFrame (go/internal/runnerhub/relay_comms.go). It adds
runnerSessionCtx scoping, used the way Hub.dropLostSession uses it.
CommitConversationFrame itself does not scope, so do not copy it literally.
Write it as a //nolint:dupl mirror, not a shared helper. The order is:
- No blob store is wired:
CodeUnavailable. - Run
runnerSessionCtx, thenaccountForRunnerSession. A miss isCodeNotFound. - The data is over
MaxSessionBlobBytes, or the hash is not^[a-f0-9]{64}$:CodeInvalidArgument. - Map store errors:
ErrFailedPrecondition(no object store) becomesCodeFailedPrecondition, the “surface absent” code thatHandler.FetchAgentConfiguses.ErrInvalidArgumentbecomesCodeInvalidArgument.- Anything else becomes
CodeInternal.
Store.PutSessionBlob runs these steps in order:
- A nil
objectStorereturnsErrFailedPrecondition. - If
sha256(data)does not match the claimed hash, returnErrInvalidArgument. - If this session’s own index row exists, return nil.
PutSegment("blobs/"+hash, data).- Insert the index row with
ON CONFLICT DO NOTHING.
This is the PUT-before-PG order of flushUpto
(go/internal/store/agent_transcripts.go). A crash between steps 4 and 5 is
repaired by the next attempt’s identical PUT.
The index copies agent_session_archive_segments (0001_init.sql):
CREATE TABLE agent_session_blobs ( session_id TEXT NOT NULL REFERENCES agent_sessions (session_id) ON DELETE RESTRICT, sha256 TEXT NOT NULL CHECK (sha256 ~ '^[a-f0-9]{64}$'), size_bytes BIGINT NOT NULL CHECK (size_bytes > 0), tenant_id TEXT NOT NULL DEFAULT current_setting('compass.tenant_id', TRUE), created_at TIMESTAMPTZ NOT NULL DEFAULT now(), updated_at TIMESTAMPTZ NOT NULL DEFAULT now(), PRIMARY KEY (session_id, sha256));The table ships in a new migration numbered with the next free NNNN at
rebase. verifyChecksum (go/internal/store/store.go) refuses an edited
0001_init.sql, and checkContiguous refuses gaps. Because the 0001 DO
loops ran before this table existed, the migration repeats their work:
- the dynamic
current_schema()GRANT tocompass_appandcompass_system(pgtests use per-test schemas); ENABLEandFORCE ROW LEVEL SECURITY;- the
tenant_isolationpolicy; - the
set_updated_attrigger.
The object key is global, as ruled, so isolation lives in the index:
ReadSessionBlob(ctx, sessionID, sha) checks the session’s row under the
tenant transaction before GetSegment.
D4 Resume: the Runner pulls referenced blobs
Section titled “D4 Resume: the Runner pulls referenced blobs”The Runner pulls blobs inside agentHost.Start (go/internal/runner/host.go)
on an authorized resume. The steps run in this order:
- Fetch secrets (unchanged).
- Scan the body for refs, newest first and deduped.
- Call
FetchSessionBlobs. - Mark absent images (D5).
- Write the resume file.
- Write the blobs.
bindAndStartAgent.
The new RPC is a server stream, RunnerService.FetchSessionBlobs, in the
FetchAgentConfig shape: a header frame, then 512 KiB chunk frames
(configChunkBytes). A ResumeBody field would put every image into one
message on the Sessions stream and stall it.
Authorization copies Hub.BindLifetime
(go/internal/runnerhub/bind_lifetime.go):
AccountForContainermust resolve; a miss isCodePermissionDenied.AccountTenantruns under the system role.- The read runs under
WithTenantwith the session-ownership predicate.
The Server’s response order is:
- One
absentheader for each requested hash that has no index row for this session. - Then the indexed blobs, in request order, until the next one would cross
maxResumeBlobBytes(128 MiB) ormaxResumeBlobs(256). - A failed
ReadSessionBlobskips that blob and is logged.
The Runner enforces the same budget and drops any blob whose recomputed SHA-256
does not match its header. AgentRuntime.WriteAgentFiles, a sibling of
WriteAgentFile (go/internal/runtime/agent.go), writes the survivors in one
exec:
tar -xruns as the agent uid under umask 077 into.omp/agent/blobs;- entries are regular 0600 files named by the computed digest;
agent-image/toolchain.nixaddspkgs.gnutar, since no tar is linked today, and the guest rootfs reuses that image.
Start-latency bound: one stream plus one exec. The fetch runs under
resumeBlobFetchTimeout (30 s) and the budget above. A timeout or any
non-FailedPrecondition error counts as transient, and Start continues. Only
cancellation of the caller’s ctx fails the Start.
Reload reuses the container (agentHost.reloadLocked), so it fetches
nothing.
D5 Marker: the Runner rewrites absent images
Section titled “D5 Marker: the Runner rewrites absent images”If nothing rewrites the body, resolveImageData (session/blob-store.ts) logs
and keeps the ref:
logger.warn("Blob not found for image reference", { hash });return data; // Return the ref as-is; downstream will see invalid base64 but won't crashThe provider clamp then emits
[image omitted: undecodable ${mimeType} data (${reason})]
(replaceUnreadableContent, session/provider-image-budget.ts). That is the
warning the ruling replaces.
The Runner rewrites the resume body in Go before writing it. It replaces each
content image block whose ref is in the absent set with
{"type":"text","text":"[image unavailable after resume]"}. A hash is absent
only in two cases:
- the Server sent an
absentheader for it; - the whole fetch returned
FailedPrecondition(no object store).
Every transient case keeps the ref: a timeout, a stream error, over budget, a digest mismatch, or a Server read failure. This lifetime sees the SDK’s text, and a later resume can still recover the image.
Marking is skipped when the container has already run an agent, because its
blob dir may hold files that were never uploaded. A per-container flag set in
bindAndStartAgent tracks this.
Only lines containing blob:sha256: are decoded (UseNumber, re-encoded with
SetEscapeHTML(false)). Unchanged lines stay byte-identical.
Rejected: an agent-side rewrite before tee init. The agent cannot tell absent
from transient, so a flaky resume would mark stored images and the next
checkpoint would erase their refs. A context extension handler is rejected
too, because emitContext structuredClones every message on each LLM call.
See OQ-2.
Global Constraints
Section titled “Global Constraints”- Size cap. A message is at most
agentmsg.MaxBytes(16 MiB). A blob is at mostMaxSessionBlobBytes(15 MiB raw), and TSMAX_SESSION_BLOB_BYTESmatches it. A larger image is skipped, never split. - Best-effort. No blob path fails a turn, a Start, a resume, or the transcript lane, and the R4 latch never fires for a blob.
- No object store.
CodeFailedPreconditionmeans only “surface absent”. Transient faults areUnavailableorInternal. - Tenant isolation.
agent_session_blobshas the standardENABLE+FORCERLS policy. Every store call runs understore.WithTenant. Reads are keyed by (session, hash) and need that session’s row. - No existence probe.
PutSessionBlobnever asks whether an object exists. Its result and timing depend only on this session’s own row. - Verified bytes. The Server and the Runner both recompute SHA-256. File names come from the computed digest.
- No ids in model-facing text. The marker is the fixed string
[image unavailable after resume]. - Storage boundary. The agent holds no storage credentials (DL-089).
- SDK.
^18.0.11, unpatched. Refs matchblob:sha256:[a-f0-9]{64}. - Schema. New migrations only, numbered at rebase.
- Public repo. Cite only public paths.
T1 proto: blob RPCs (lane proto)
Section titled “T1 proto: blob RPCs (lane proto)”// agent_gateway.proto, service AgentGatewayrpc PutSessionBlob(PutSessionBlobRequest) returns (PutSessionBlobResponse);message PutSessionBlobRequest { string sha256 = 1; bytes data = 2; }message PutSessionBlobResponse {}
// runner.proto, service RunnerServicerpc RelaySessionBlob(RelaySessionBlobRequest) returns (RelaySessionBlobResponse);message RelaySessionBlobRequest { string session_id = 1; PutSessionBlobRequest blob = 2; }message RelaySessionBlobResponse {}
rpc FetchSessionBlobs(FetchSessionBlobsRequest) returns (stream FetchSessionBlobsResponse);message FetchSessionBlobsRequest { string container_name = 1; string session_id = 2; repeated string sha256 = 3; // newest first}message FetchSessionBlobsResponse { oneof frame { SessionBlobHeader header = 1; bytes chunk = 2; } }message SessionBlobHeader { string sha256 = 1; uint64 size_bytes = 2; bool absent = 3; }Also rewrite the ResumeBody comment, which says blobs are “OUT of MVP scope
(RIG-1582)”, and run moon run proto:gen.
- Test:
moon run proto:ciis green.
T2 store: index and blob methods (lane compass-server)
Section titled “T2 store: index and blob methods (lane compass-server)”const MaxSessionBlobBytes = MaxBytes - 1<<20
// go/internal/storetype SessionBlobRow struct{ SHA256 string; SizeBytes int64 }func (s *Store) PutSessionBlob(ctx context.Context, sessionID, sha256Hex string, data []byte) errorfunc (s *Store) ResumeSessionBlobs(ctx context.Context, sessionID string, account AccountID, sha256s []string) ([]SessionBlobRow, error)func (s *Store) ReadSessionBlob(ctx context.Context, sessionID, sha256Hex string) ([]byte, error) // ErrNotFound without this session's rowAdd the migration (D3) and queries/agent_session_blobs.sql (row exists,
insert, ownership-filtered select), then regenerate sqlc. Add
agent_session_blobs to tenantOwned in TestRLSCatalogEnabledAndForced.
Tests (pgtest with the store’s fake ObjectStore):
- A nil object store returns
ErrFailedPrecondition. - A mismatch causes no PUT and no row.
- A repeat call does not PUT again.
- A row is invisible to another tenant.
ReadSessionBlobfor another session’s hash returnsErrNotFound.
T3 hub and handler (lane compass-server)
Section titled “T3 hub and handler (lane compass-server)”type SessionBlobStore interface { AccountTenant(ctx context.Context, account store.AccountID) (store.TenantID, error) PutSessionBlob(ctx context.Context, sessionID, sha256Hex string, data []byte) error ResumeSessionBlobs(ctx context.Context, sessionID string, account store.AccountID, sha256s []string) ([]store.SessionBlobRow, error) ReadSessionBlob(ctx context.Context, sessionID, sha256Hex string) ([]byte, error)}func (h *Hub) SetSessionBlobStore(b SessionBlobStore)func (h *Hub) RelaySessionBlob(ctx context.Context, runnerID string, req *compassv1internal.RelaySessionBlobRequest) (*compassv1internal.RelaySessionBlobResponse, error)func (h *Hub) FetchSessionBlobs(ctx context.Context, runnerID string, req *compassv1internal.FetchSessionBlobsRequest, send func(*compassv1internal.FetchSessionBlobsResponse) error) error// Handler.RelaySessionBlob / Handler.FetchSessionBlobs: runnerSubjectFrom guard, delegate.Wire *store.Store next to SetTranscriptStore.
Tests:
- Every guard maps to its code.
- A pgtest checks that a tenant-B session’s blob row is stamped tenant B.
- Absent headers come first, then header and chunks per blob.
- The budget stops before the blob that would cross it.
- A foreign container gets
PermissionDenied. - A read failure skips that blob, and the stream still ends OK.
T4 Runner gateway forwarder (lane compass-runner)
Section titled “T4 Runner gateway forwarder (lane compass-runner)”type SessionBlobRelay interface { RelaySessionBlob(ctx context.Context, req *connect.Request[compassv1internal.RelaySessionBlobRequest]) (*connect.Response[compassv1internal.RelaySessionBlobResponse], error)}func (g *Gateway) PutSessionBlob(ctx context.Context, req *connect.Request[compassv1internal.PutSessionBlobRequest]) (*connect.Response[compassv1internal.PutSessionBlobResponse], error)Add Deps.SessionBlobs and wire it where Deps.Board is wired.
Tests:
- An unbound container gets
PermissionDenied, with no relay call. - The forwarded request carries the bound session id.
- A Server
FailedPreconditionpasses through unchanged.
T5 Runner resume fetch and write (lane compass-runner)
Section titled “T5 Runner resume fetch and write (lane compass-runner)”type SessionBlob struct{ SHA256 string; Data []byte }type SessionBlobFetch struct{ Blobs []SessionBlob; Absent map[string]struct{} }func (l *ServerLink) FetchSessionBlobs(ctx context.Context, containerName, sessionID string, sha256s []string) (SessionBlobFetch, error)func blobRefsNewestFirst(resumeBody string) []string
// go/internal/runtimetype AgentFile struct{ Name string; Data []byte }func (r *AgentRuntime) WriteAgentFiles(ctx context.Context, id WorkloadID, uid uint32, homeDir, relDir string, files []AgentFile) errorThe FetchSessionBlobs client follows ServerLink.FetchAgentConfig: a chunk
before any header is a contract skew. It enforces the budget, drops digest
mismatches, and runs under resumeBlobFetchTimeout. Add pkgs.gnutar to
agent-image/toolchain.nix. Wire it into agentHost.Start in the D4 order.
Tests:
- Blobs land under
.omp/agent/blobsnamed by the computed digest, in one exec. - A mismatched blob is dropped.
- A timeout or a mid-stream error still starts the agent.
- A fresh Start never fetches.
T6 agent capture and upload (lane compass-agent)
Section titled “T6 agent capture and upload (lane compass-agent)”// transport/index.ts, on RunnerTransportputSessionBlob(req: PutSessionBlobRequest): Promise<void>;interface TranscriptTeeOptions { readonly onCommittedLine?: (line: string) => void }// session-blobs.tsexport const MAX_SESSION_BLOB_BYTES = 15 * 1024 * 1024;export const SESSION_BLOB_QUEUE_MAX = 16;export interface SessionBlobUploader { offer(line: string): void; close(): Promise<void> }export function createSessionBlobUploader(deps: { put: (req: PutSessionBlobRequest) => Promise<void>; blobsDir: string; // getBlobsDir()}): SessionBlobUploader;Wiring:
appendandwriteFullcallonCommittedLineafter the local write.mainpassesofferintocreateTeeSessionStorageand closes the uploader in the existing drain-then-closefinally.- Reserve
PI_CODING_AGENT_DIRandXDG_DATA_HOMEinisReservedEnvKey. - Counters follow
transport/otel-metrics.ts:compass_agent.session_blobs.{uploaded,dropped,oversize,missing,failed}.
Tests:
- An appended line and a checkpoint body each upload a hash once.
- A
failedhash is re-offered by a later line, while an uploaded hash is not. - An oversize blob is counted and marked done.
- ENOENT counts as
missing, notfailed. - Overflow drops the oldest entry.
FailedPreconditiondisables the uploader.- A rejecting
putor a throwingoffernever failsappend.
T7 Runner unavailable marker (lane compass-runner)
Section titled “T7 Runner unavailable marker (lane compass-runner)”const imageUnavailableText = "[image unavailable after resume]"func markAbsentBlobs(resumeBody string, absent map[string]struct{}) (string, int)Start calls markAbsentBlobs between the fetch and the resume-file write. It
skips the call when the container’s started flag is set.
Tests:
- An absent content image becomes the marker.
- A transient (unlisted) hash is left alone.
- Unchanged lines are byte-identical, and large numbers survive.
- A reused container is not marked.
T8 end-to-end resume with an image (lane compass-server e2e)
Section titled “T8 end-to-end resume with an image (lane compass-server e2e)”Extend the go/e2e leg that resumes into a fresh container:
- With an object store, the image bytes are back in the new blob dir.
- Without one, the resumed transcript carries the marker.
- T1 proto: three RPCs,
ResumeBodycomment, regenerate - T2 store: migration, queries, sqlc, blob methods, RLS catalog entry
- T3 hub and handler: relay, fetch, wiring
- T4 Runner gateway:
PutSessionBlobforwarder - T5 Runner resume: fetch client,
WriteAgentFiles, gnutar,Startwiring - T6 agent: tee hook, uploader, transport, reserved env keys, counters
- T7 Runner:
markAbsentBlobsand the reused-container skip - T8 e2e: image survives with an object store; marker without one
Open Questions
Section titled “Open Questions”- OQ-1, retention and erasure (not load-bearing).
blobs/<sha256>keys are shared across sessions and tenants. Deleting one, whether for cost or to erase a tenant’s images at offboarding, needs a system-role reference sweep across every tenant’s index. This record deletes nothing. A follow-up issue should own GC and erasure together. - OQ-2, marker placement (load-bearing; Matt’s call). The default is the Runner marker (D5). It marks only Server-confirmed absences and keeps transient refs recoverable, at the cost of skipping reused containers. The alternative is an agent-side rewrite before the tee storage initializes. It is simpler and covers reused containers, but it cannot tell absent from transient, so one flaky resume durably erases images that are stored. The recommendation is the Runner marker.
- OQ-3, marker coverage (not load-bearing). Upload covers every ref
position. The marker rewrites only
contentimage blocks. Absent refs inimages[], snapcompactframes[],image_url, andimage_generation_call.resultstill show the SDK’s text. Widen it only if those positions show up in real resumes. - OQ-4, budget values (not load-bearing). The defaults are 15 MiB per blob, 128 MiB and 256 blobs per resume, a 30 s fetch, and a 16-entry queue. They are tuning values.