ScienceDiscovery
中文 GitHub

Sandbox Execution: services/runner

Runner is a rootless executor. Ordinary Agent execution uses run_shell, including python -m, Python files and Rscript, inside a Bubblewrap/seccomp sandbox with its Agent×Runner Workspace. Each call starts a fresh process; cwd, exports and interpreter memory do not carry across calls. Sandbox network access defaults to none; see §3.1. Runner also manages micromamba environments. Its default listener is 127.0.0.1:4311, with API as its client.

1. Source structure

File Responsibility
server.ts Routes, bearer/HMAC auth, Workspace admission and execution queue, startup preflight
executor.ts Ephemeral Bubblewrap construction, quota/timeout, workspace snapshots
execution-manager.ts Managed Shell lifetime, status/logs/cancel and committed Workspace receipts
kernel-manager.ts, shell-session-manager.ts, session-env-profile.ts Legacy internal primitives; HTTP execution no longer starts persistent workers or injects saved profiles
environment-store.ts micromamba catalog, in-place named environments and revision records
npu-broker.ts Optional Host NPU Broker that starts allowlisted host NPU workloads under Runner control
workloads/ Broker default workload allowlist, Ascend smoke probe, and controlled adapters
seccomp.ts x86_64/aarch64 BPF generated under runner runtime; baseline, network, and no-egress NPU compatibility profiles
egress-gateway.ts Host-side exit for sandbox network access: a UDS HTTP service reused per policy revision, allowed domains and address classification
egress-bridge.ts In-sandbox TCP→UDS bridge script, host interpreter probe, and bwrap bind arguments
request-auth.ts HMAC-SHA256 token/timestamp/body hash with 30-second freshness

2. HTTP surface

GET /health is unauthenticated. Status, environments/revisions/setup, and kernel teardown require bearer auth. /execute and /execute-shell additionally require timestamp/signature headers. When the NPU Broker is enabled, GET /npu/workloads uses bearer auth; GET /npu/jobs?session_id=... and job status/log/result endpoints use bearer auth plus Session checks; POST /npu/jobs and job cancel also require the freshness signature. Reusing an executionId within 60 seconds returns 409.

3. Sandbox construction

Startup checks required Bubblewrap options and executes a probe. Two parts of the sandbox shape can be refused by the environment, and both are settled by probing rather than guessing. /proc is settled first and --disable-userns second, on whichever /proc shape was chosen, so neither can be misdiagnosed as the other.

/proc defaults to --proc /proc, giving the sandbox its own procfs so it sees only its own processes. Docker's default readonlyPaths/maskedPaths make the kernel refuse that mount in the sandbox's own pid namespace (Can't mount proc on /newroot/proc: Operation not permitted); the runner then falls back to --ro-bind /proc /proc and warns. Executions still run, but the sandbox sees the container's process list. The official Compose file keeps the stronger shape with systempaths=unconfined; the fallback is never the default, and privileged is not the way to avoid it.

--disable-userns is added only when a probe proves it usable here, not based on the version or on --help: the option works by writing user.max_user_namespaces, so under LXC and container runtimes that mount /proc/sys read-only it fails and aborts the whole launch even on Bubblewrap 0.8+. The probe runs a minimal sandbox with the option, then without it, which separates an old Bubblewrap that rejects the unknown option from an environment that refuses the sysctl write, and both from a host where no sandbox builds at all. Whenever the option is omitted the runner warns and every other protection — namespaces, seccomp, the mount allowlist — is unchanged, so executions still run. The launcher preflight and the runner share this detection (packages/sandbox-capability), so preflight cannot pass a sandbox the runner then fails to build.

--die-with-parent --new-session --unshare-all --unshare-user [--disable-userns]
--cap-drop ALL
read-only /usr plus system links, /dev; tmpfs /tmp
--proc /proc, or --ro-bind /proc /proc when a fresh procfs is refused
hide host Python/R when managed environments are enabled
read-only selected environment at /opt/science-env
bind Agent×Runner Workspace read-write at /workspace
--clearenv plus runner baseline; cwd selected explicitly for this call
--seccomp 3

Legacy synchronous requests abort on disconnect. Managed Shell Executions survive a client waiting deadline or disconnect; cancellation is explicit through the Execution management endpoint.

3.1 Sandbox network access

Sandbox network access is a system setting. API snapshots it into every Permission Epoch (networkPolicy plus networkAccess, including a content-derived revision) and Runner shapes the sandbox from that snapshot. It is unrelated to the Network proxies settings, which govern the API/Gateway/MCP's own outbound calls and never affect sandbox code.

Mode Sandbox
none (default) Exactly the historical behavior: --unshare-all, no --share-net, baseline seccomp denying every socket syscall, no channel mounted, no outbound environment injected
domain-allowlist Still --unshare-all and still no --share-net. The only exit is a bind-mounted Unix domain socket

The domain-allowlist data path:

sandbox process (own netns, no interface)
  └─ HTTP_PROXY=http://127.0.0.1:18118
       └─ egress bridge (inside the sandbox, on the sandbox's own loopback)
            └─ /run/sciencediscovery/egress.sock (bind mount)
                 └─ egress gateway (in the runner process, runner's own user)
                      └─ allowed domains only → internet

Properties:

3.2 Ascend NPU inside the sandbox

Selected Ascend chips are handed to the sandbox. An earlier revision of this document said they could not be, because a probe inside the bwrap namespace failed with Container ID verify failed (session ct_id=0; device ct_id=...). That was measured while the host's whole /dev was visible, and it is what the driver does in that situation rather than a limit on device passthrough: inside a mount namespace the driver enumerates the cards it can see under the caller's /dev, all-or-nothing, so one card claimed by another tenant fails the call for every card. Exposing only the selected chips removes the condition, and npu-smi info and MindSpore both run inside the sandbox on a 910B3.

What the launch does:

What may be selected is decided per chip by a real probe, not by the host listing: a throwaway sandbox shaped like a real launch binds that one chip and runs npu-smi info in it. The host reports cards as healthy that a sandbox cannot open, so only a chip whose probe succeeded can be ticked, and every execution re-probes the chips it names before launching — a chip claimed by another tenant in the meantime fails the execution by name instead of failing deep inside a framework. Scope is the Ascend 910 series; other chips are listed and refused with the chip name in the reason.

Machine state is read through the driver's own DCMI interface (via the host Python, no compiled addon), falling back to npu-smi info -m plus npu-smi info when that is unavailable. Device identity is the chip logic id — the N in /dev/davinciN — never the card number, because a card can carry more than one compute die.

3.3 Ascend NPU Broker (optional host execution)

Separately from the above, Runner still exposes an opt-in Host NPU Broker for allowlisted host workloads that are not ordinary Agent executions:

The exception is “allowlisted host model job,” not “host shell for the Agent.” NPU deployment variables live in Configuration reference, and the model-visible tool contract lives in Built-in tools.

4. Execution model and quotas

Inspect or modify quotas

curl -s http://127.0.0.1:4310/health | jq '.workspace, .runner.maxWorkspaceBytes, .runner.maxFileBytes, .runner.maxOutputBytes'
curl -s -H "authorization: Bearer $TOKEN" http://127.0.0.1:4310/api/quota-settings

The Web Quotas settings persist values for new executions. Environment seeds are SCIENCE_AGENT_MAX_WORKSPACE_BYTES, SCIENCE_AGENT_MAX_OUTPUT_BYTES, SCIENCE_AGENT_WORKSPACE_MAX_BYTES, and the upload file/request limits; 0 means unlimited for the relevant dimension.

5. Language runtimes

Entry Process
Ordinary run_shell Fresh strict Bash; may launch Python modules/files, Rscript and other selected-environment tools
Legacy Python HTTP execution Ephemeral python3 -I -
Legacy R HTTP execution Ephemeral R --vanilla --slave

Interpreters come from host /usr/bin or managed /opt/science-env/bin.

6. Scientific environments

7. Execution lifetime and migration

HTTP /execute, /execute-shell and /shell-executions reject kernelMode=persistent before creating a Workspace or starting code. Once-scoped permission does not silently downgrade that request. Omit the field or use ephemeral; use a managed Shell Execution for a long-running task.

A resident interpreter could otherwise leave a thread or subprocess writing after a call returned and the Workspace lease was released. On Linux, ephemeral executions end their Bubblewrap PID namespace before the final snapshot and lease release. A managed background Execution instead retains ownership while its workload runs; it is not a reusable interactive Shell.

Historical Session profiles are not injected. Select cwd and environment on every call; put required exports and commands in the same script. Python/R memory is not retained between calls. Notebook-style shared memory is outside this implementation.