Forking and snapshots
Get a sandbox into a good state once, then branch from it: fork it into parallel attempts, snapshot it for later, or pause it until you need it again.
Fork a running sandbox#
with client.sandbox(template="python") as base:
base.exec(["sh", "-c", "git clone https://github.com/you/app app && cd app && uv sync"])
fork = base.fork()
print(fork.parent_sandbox_id) # base's ID
fork.exec(["sh", "-c", "cd app && uv run pytest -q"])
fork.destroy()fork() returns a new Sandbox. The parent pauses only while its state is captured, then carries on. The fork wakes up with the same memory, files and running processes, and from then on the two are independent: each has its own ID, token and network.
A fork shares its parent's memory pages and disk blocks until it writes to them, so it stores only what it changes. That makes it cheap to fork many times from one prepared machine.
Try several approaches in parallel#
Do the slow setup once, fork a copy per attempt, run them at the same time and keep whichever works.
import asyncio, os
from upstream import AsyncUpstream
async def attempt(sb, patch):
await sb.write_file("/workspace/app/fix.patch", patch)
r = await sb.exec(["sh", "-c", "cd app && git apply fix.patch && uv run pytest -q"])
return r.exit_code == 0
async def main(repo, patches):
async with AsyncUpstream(api_key=os.environ["UPSTREAM_API_KEY"]) as client:
base = await client.sandbox(template="python")
await base.exec(["sh", "-c", f"git clone {repo} app && cd app && uv sync"])
forks = [await base.fork() for _ in patches]
passed = await asyncio.gather(*(attempt(sb, p) for sb, p in zip(forks, patches)))
for sb, ok in zip(forks, passed):
print(sb.sandbox_id, "pass" if ok else "fail")
await sb.destroy()
await base.destroy()Every fork starts from exactly the same state, so differences between attempts come from the patch, not from the setup. The same pattern works for evaluating agents: fork one prepared environment per rollout.
Debug from a failure#
When something fails partway through, fork the sandbox at that moment. The copy holds the failing state, including caches and anything left on disk, so you can investigate it without re-running setup.
result = sb.exec(["sh", "-c", "cd app && uv run pytest -q"])
if result.exit_code != 0:
probe = sb.fork() # the failing machine, exactly as it is now
report = probe.exec(["sh", "-c", "cd app && uv run pytest --last-failed -x -vv"])
print(report.stdout)
probe.destroy()The original keeps running, so an agent can carry on while you, or another agent, dig into the copy.
Save and restore snapshots#
A snapshot keeps a sandbox's state so you can start new sandboxes from it later, after the original is gone.
snapshot = sb.create_snapshot(label="deps-installed")
print(snapshot["id"])
# later, from any client
restored = client.restore_snapshot(snapshot["id"])
restored.exec(["sh", "-c", "cd app && uv run pytest -q"])List snapshots with client.list_snapshots(), or sb.list_snapshots() for one sandbox, and remove them with client.delete_snapshot(id).
Pause instead of destroy#
Pausing suspends a sandbox in place. It keeps its state and resumes where it left off.
sb.pause()
# ... later
sb.resume()To pause automatically instead of destroying when a sandbox times out, create it with on_timeout="pause":
sb = client.sandbox(template="claude-code", on_timeout="pause")Which one to use#
| You get | Use it to | |
|---|---|---|
fork() | A live copy of a running sandbox that shares the parent's pages until it writes. | Branch a prepared machine into parallel attempts, rollouts or a debugging copy. |
clone() | A new, independent copy of the sandbox. | Duplicate a sandbox outright. |
create_snapshot() | Saved state, with an ID and a label. | Return to a known-good point later with restore_snapshot(). |
pause() | The same sandbox, suspended. | Set a sandbox aside without losing it; resume() picks up where it left off. |