JiuwenSwarm migration: status and hand-over
Archived migration snapshot. Issue statuses, test counts, limits, and defaults below describe an earlier baseline and must not be used as current configuration guidance. All launchers now default to JiuwenSwarm; child delegation defaults to platform
taskwith Swarm execution. Use Agent backends and current orchestration for current behavior.
Where issue 84 (reuse the JiuwenSwarm backend) stands, what runs today, what does not, and how to start working on a sub-issue. To run it from source, see Local mode.
What this baseline is
A transitional architecture, chosen deliberately: the existing TypeScript API keeps owning storage (sessions, messages, run events, permission records) and every route the UI calls. Only the agent executor is replaced: the model loop runs on JiuwenSwarm 0.2.6, reached through a Python adapter that also sits in front of the API as a reverse proxy.
browser ─▶ adapter :4310 ──proxy──▶ legacy API :4410 ──createAgent──┐
│ ▲ │ POST /agent/runs
│ └────── tool calls (loopback bridge, per run) ◀─────┤
▼
JiuwenSwarm gateway ─▶ model, through /llm/<token>/v1 (adapter)
└── MCP tools ─▶ /mcp/<token> (adapter, forwards to the bridge)
This is not the end state written in the issue body (routes served by the adapter on top of JiuwenSwarm's own storage). Each sub-issue below says how much of its end state is still open. Whether the acceptance criteria that assume the end state should be rewritten for this architecture is an open decision for the issue owners.
Design and measured protocol facts: services/adapter/README.md.
Status by sub-issue
| Issue | State | What exists | What does not |
|---|---|---|---|
| 86 Acceptance baseline, protocol calibration | closed | Interface inventory (259 rows), L1/L2 tooling, 27 recorded cases on Linux, protocol experiments, CI_E2E_BACKEND=jiuwenswarm |
Cases for 104 rows; the 59 "approximate" mappings were not calibrated; AgentRuntime not evaluated |
| 87 Adapter core | closed | Public port, streaming proxy (SSE, NDJSON), 502 on upstream failure, optional token on /agent/* |
Own storage, run registry, event log and resume: they stay in the legacy API |
| 88 Sessions and projects | closed | Served by the legacy API through the proxy; L1 covers all 13 rows | Nothing moves to JiuwenSwarm's session store |
| 89 Chat and runs | closed | A run executes on JiuwenSwarm; events, cancel, approval, usage, disconnect handling; L2 golden traces for 8 scenarios | See "Known gaps" |
| 90 Permissions | closed | Approval requests, allow/deny, grants, audit, epoch and wrong-token refusal behave as before (L1 8/8 plus negative cases, L2 approve/deny) | Not mapped to JiuwenSwarm's permission engine (left off; the API's permission runtime decides). Finding below |
| 91 Models, providers, settings | closed | Model list, defaults and settings served by legacy (L1 covers every row); the selected model is handed to JiuwenSwarm per run | All three model protocols run, through the API's loopback model gateway |
| 92 MCP and data sources | closed | All 21 rows have L1 cases; the five MCP journeys pass on this executor (custom MCP, agent calls a custom MCP tool, OAuth, secret edit x2); deferred MCP tools are promoted up front and tool_search is offered |
MCP OAuth against a real provider was not exercised |
| 93 Skills and libraries | open | Served by legacy | 0/35 L1 cases; skill use on this executor is not verified |
| 94 Files, workspace, trajectory | open | Read routes have L1 cases | Trajectory recorded (see Known gaps 1) |
| 95 Planning and sub-agents | closed | Planning is JiuwenSwarm's own todo tools by default (its todo.updated list becomes the run's plan; SCIENCE_AGENT_JIUWENSWARM_PLANNING=update_plan switches to ours). With JiuwenSwarm's tools in play (the default), delegation is JiuwenSwarm's own subagent_spawn/subagent_wait too, and task is not offered (SCIENCE_AGENT_JIUWENSWARM_SUBAGENTS=task switches back to ours; SCIENCE_AGENT_JIUWENSWARM_TOOLS=ours implies it already; the mocked journeys run on JiuwenSwarm's own tools with SUBAGENTS=task, each task subagent with its own stable JiuwenSwarm session so a resumed one continues its conversation); L1 covers every row; L2 subagent and two-turn cases match; the plan and delegate journeys pass (scripted against JiuwenSwarm's todo_create/todo_modify and our task). When its planning or delegation is not the run's (a Plan plugin switched off, PLANNING=update_plan, SUBAGENTS=task), its todo or sub-agent tools are hidden, not merely left unlisted, so the model cannot plan or delegate around that |
See Known gaps 1a for what a subagent_spawn sub-agent does not get. task_tool, team.*, agents.* and agent templates are still not used; the per-step plan snapshot is not injected |
| 97 Science toolset as MCP | closed | The equivalent toolset is offered as MCP tools that call the legacy tools over the bridge, through the same ToolRegistry; a test pins the offered set and schemas to the registry's; deferred tools are promoted up front with tool_search offered; parallel calls run and match |
Tools of later sub-issues (idea-tree, evolve, memory, artifact review) arrive with those |
| 103 Usage | closed | Per-model-call token counts become model.usage and are summed; session usage and per-model usage after a run match the built-in loop (L2); every row has an L1 case |
Nothing moves to adapter storage; analytics stay in the API |
| 93, 94, 96, 98-102, 104-106 | open | Unchanged legacy behaviour behind the proxy | Everything: not started as migrations. Runner/environment/remote-host reads have L1 cases |
Verified
On the Aliyun Linux server (bubblewrap sandbox) with SCIENCE_AGENT_ADAPTER=1 SCIENCE_AGENT_EXECUTOR=jiuwenswarm:
- The five milestone-0 journeys (first run, compact process ×2, plan workspace, delegate subtask, deliver result) pass.
- One journey against a live OpenAI-compatible model passes (
journey-real-request). - L2: 10 run-event scenarios (text, tool call, approve/deny, cancel, 401, subagent, resume, post-messages, two-turn conversation, parallel tool calls) have the same event-type sequence as the built-in loop. Accepted differences are in
test/contract/accepted-differences.json(the wording of a provider 401, and the evidence gap). - L1: the recorded cases match between the built-in loop and the adapter + JiuwenSwarm stack, apart from build version strings, which are scrubbed.
- Unit tests: adapter (
pytest, 135) and the TypeScript agent factory and model gateway (51).test/contract/jw-only/live.mjschecks against a running JiuwenSwarm stack what only this backend does (conversation continuity, todo planning).
Context management
JiuwenSwarm has a context engine of its own (it compresses at 80% of the model's window and injects its own per-step dynamic context). The executor now uses it:
| Built-in loop | JiuwenSwarm executor | |
|---|---|---|
| Where the conversation lives | The API's record, re-sent every step | JiuwenSwarm's session, one per agent (persisted in its checkpoint database); nothing is sent along and the adapter holds none |
| Compression when the window fills | ScienceDiscovery's own compactor | JiuwenSwarm's, against one global window (JIUWENSWARM_CONTEXT_WINDOW_TOKENS, default 200000; it ignores a model's own). Verified with a 3000-token window: the early turns are summarised (live.mjs compression). Its summaries and titles use its default model, which the adapter points at the model of the run in progress |
| System prompt | Assembled per step from ScienceDiscovery's sections (identity, governance, capabilities, skills) | ScienceDiscovery's product prompt first, then JiuwenSwarm's own prompt whole (12k characters: identity, safety, tool rules, memory, context compression, installed skills), then the run contract. Its todo section is not in it in this mode. SCIENCE_AGENT_JIUWENSWARM_PROMPT=replace gives the old behaviour |
| Run contract, protected | Every step | Every step: it is part of the system prompt, which the adapter puts on every model request |
| JiuwenSwarm's own prompt | n/a | Kept whole; its per-turn wrapper and dynamic context (runtime state) are also added as user messages |
| Plan snapshot, durable state, plugin context | Every step | Not injected; JiuwenSwarm adds its own dynamic context |
| Tools | ScienceDiscovery's, all through its permissions and Runner | By default JiuwenSwarm's own web, sub-agent, todo, memory, and skill tools plus ScienceDiscovery's where JiuwenSwarm has none. Its host-acting implementations (bash, file tools, and related readers) are blocked in every tool mode. The adapter rejects a native call; a same-named ScienceDiscovery tool remains available when the run provides it, and commands use run_shell. SCIENCE_AGENT_JIUWENSWARM_TOOLS=ours selects ScienceDiscovery's tool set, with JiuwenSwarm's todo tools retained for planning |
| Tool routing hints | Yes | Not injected |
| Tool-output guard and reader | Yes | Yes (same ToolRegistry) |
| Recovery from an input-too-large error | Compact and retry | JiuwenSwarm's own handling; not verified |
| Trajectory and evidence | Yes | Yes (recorded at the model gateway, see Known gaps 1) |
Known gaps
- Trajectory: recorded. Every model call of a JiuwenSwarm run passes through the run's model gateway, which records it with the same
AgentVersionRecorderthe built-in loop uses: the exact model input (JiuwenSwarm's prompt and the tools included), the answer, the tool observations and the committed step; the run's events carry the evidence that links them (live checktrajectory). Title and summary calls JiuwenSwarm makes are not turns and are not recorded. 1a. Sub-agents use JiuwenSwarm's ownsubagent_spawn/subagent_waitby default (SCIENCE_AGENT_JIUWENSWARM_SUBAGENTS=task, orTOOLS=ours, keeps ScienceDiscovery'staskinstead). That is a deliberate trade-off, not full parity: asubagent_spawnsub-agent (0.2.6's built-ingeneral_agent) runs entirely inside JiuwenSwarm, with its remaining built-in tools and no MCP servers. It never gets ScienceDiscovery's tools, sandbox, approvals, workspace handoff, provenance or sub-agent cards, and itssubagent_spawn/subagent_waitcalls are reported to the run only as generic native tool events, not the richertasksummary. The adapter blocks JiuwenSwarm's host-acting implementations in every tool mode. A custom agent (agents.create) that could carry ScienceDiscovery's MCP tools still fails to spawn on 0.2.6 ('str' object has no attribute 'name': its tool list is names where the spawn path expects tool objects), so that route stays closed.test/contract/jw-only/live.mjs subagent-probeshows what asubagent_spawncall looks like end to end. 1b. Trajectory: context-source attribution unavailable. The trajectory viewer's per-block "source" attribution (packages/trajectory/src/index.tscontextBlocks) only marks a system-prompt block"recorded"when the model-call assembly carries structured section data (the built-in loop's per-section prompt assembler).services/api/src/agent-run/jiuwenswarm-trajectory.ts'sModelCallInputonly ever has a flatsystemPrompt: string— JiuwenSwarm assembles its own prompt whole, with no section boundaries to report — so every block falls into the"unavailable"branch instead. The mocked E2Ejourney-session-trajectorytherefore fails against this executor (it asserts the JiuwenSwarm system-prompt block is not "source not recorded"); this is an accepted difference of this backend, not a bug to fix here — see "Context management" above for the same "kept whole" trade-off. - History and context. JiuwenSwarm is the only holder of the model's context: a stable session per agent, compressed by JiuwenSwarm; the adapter and the API send and rebuild nothing. So a session begun on the built-in loop is not remembered by JiuwenSwarm (its earlier turns are on screen only), and a conversation edited or rewound in the API is not reflected there. That context survives a restart of JiuwenSwarm (checked:
live.mjs history-restart). The per-step dynamic context (plan snapshot, durable state) is not injected: see "Context management". - Model protocols: all three the UI can configure work (a loopback gateway in the API serves JiuwenSwarm's chat-completions requests through the native model client). Images are not sent to the model.
- No tool-output store in the JiuwenSwarm toolset. Deferred tools are promoted up front. Tool calls of one model response are scheduled by the native rules (a tool not declared concurrency-safe runs alone, in the order the model called it; duplicate calls are superseded by batch policies), so both executors produce the same events in the same order.
- Wake notices. The mocked E2E
issue-77-wake-noticeandissue-85-foreground-exec-inbox'stest:206case both fail on this executor, for the same reason: the scripted model recognises a wake turn by its user message starting with the literal[Execution notifications]prefix (services/api/src/notification-dispatch.tsnotificationPrompt()), but JiuwenSwarm reconstructs the model-facing conversation itself for a continuing session (see gap 2, "History and context") and does not forward that prefix unchanged, so the stub model never recognises the turn and the run errors instead of completing it.issue-85's cleanup then also fails (a failed wake leaves the notification unread and retried forever, so the session never goes idle and blocking Project deletion) — a direct cascade of the same cause, not a second bug. (An earlier version of this line said "issue-85 passes"; that was stale.) Run the group withCI_E2E_BACKEND=jiuwenswarm .ci/run-e2e.sh mocked. The journeys are load-sensitive on a small host: two of them failed once when run back to back on the 2-core server and passed alone (three repeats), so run them one at a time when in doubt. - Legacy bugs found while recording (not fixed):
PUT /api/sessions/:id/settingswith a bad body returns 500;PUT /api/web/settingswith its own GET body returns 500; skill evolution on an unknown run returns 500; reading a file outside the workspace (/file?path=../../etc/passwd) returns 500 instead of 4xx. - Finding for #90: after
POST /api/sessions/:id/permission-epoch, a session-scoped grant still applies (the epoch concerns the sandbox, the grant the session). The issue text expects old authorizations to become invalid; the baseline records what legacy does. - Idle timeout: implemented.
services/api/src/agent-run/jiuwenswarm-agent.ts'sbeginExternalWait()used to be a no-op ("the loop is remote, so there is no local idle clock to pause") and only the whole-run timeout was armed, so a silently-stuck JiuwenSwarm run never produced the idle-timeout noticetest/timeouts-runtime-status.spec.tsexpects. It now tracks the run's turn and idle deadlines the way the built-in loop does (RunDeadlines), throwing the sameAgent run stalled: no gateway progress for N mswording (matched byservices/api/src/timeouts/index.ts), withbeginExternalWait/its release function pausing the clock during a human approval wait so that doesn't itself look stalled — including the approval wait itself, which did not call it until this was noticed while merging the two independent implementations of this gap. - Custom MCP tool calls: occasionally invisible.
journey-custom-mcp's "Agent uses a selected custom MCP server" case failed once in a full-suite run (/api/sessions/:id/mcp/invocationsstayed empty even though the run completed) but has not reproduced since, including in isolation.services/adapter/src/sciencediscovery_adapter/mcp_server.py'stools/callhandler swallowed every failure path (unknown run tag, unknown tool name, a bridge exception) with no logging, so there was nothing to diagnose it by; it now logs each one. Leading theory:agent_runs.py'sensure_shared_toolsreconnects JiuwenSwarm to the sharedsciMCP server whenever a run brings a newly-named tool, and every custom MCP connector tool has a fresh per-session name (mcp__<sourceId>__<toolId>), so it is always "new" and always triggers that reconnect — a plausible race with JiuwenSwarm actually routing a call to it. Watch the new log lines the next time this reproduces. Confirmed and no longer just a theory (item 13 under gap 11's UT audit, below): genuine concurrent load on one shared JiuwenSwarm+adapter instance reproduces a missing MCP tool-call event on an otherwise-successful run on demand — plausiblyensure_shared_tools's own process-wide lock, as this theory already suspected. Chasing it further with an over-aggressive reproduction (literal duplicate copies of one test file) surfaced a second, separate, precisely-diagnosed bug instead of confirming the reconnect race above:AgentRunner.pending_approvalsis one dict for the whole adapter process, keyed by JiuwenSwarm's ownchat.ask_user_questionrequest_id— which echoes the model's own tool-call id rather than being minted unique — so two concurrent runs whose models pick the identical tool-call id collide in it. Real, but confirmed not to require the reconnect-race mechanism this item originally proposed. - A
tasksub-agent's own status can still read "running" right after the parent's run is terminal — and the same nested round trip can stall the parent run outright. Root cause found and fixed: the adapter gave JiuwenSwarm a new tool list bymcp.disconnect/register_custom/connectof the one sharedsciserver, and a disconnect is global: atasksub-agent's run, started while its parent waits on thattaskcall, brings tools the shared list did not have yet and so cut the parent's call in flight (JiuwenSwarm's log of a stuck run: the parent'staskcall at 06:19:27, the adapter's disconnect ofsciat 06:19:28, the sub-agent's session at 06:19:29, then silence). A new list now goes to the next generation's name (sci+ ten digits) while the runs on the old one keep it until the last has ended (ensure_shared_tools,test_new_tools_never_disconnect_the_server_a_running_run_is_calling_through). What follows is the history of the diagnosis.issue-85-foreground-exec-inbox's third case fails deterministically (not flaky) onexpect((await subagents(...)).map(s => s.status)).toEqual(["completed"])immediately afterwaitForRunTerminal— the child's own turn plainly did finish (tool.completedfortaskand the child's ownrun_shellboth show up in the adapter's debug log before the parent's run ends).runs/index.ts'srunSubagentawaits the whole child run and writesstore.updateSubagent(..., "completed")before returning to the tool caller, so nothing in the TypeScript call order explains it — but withSCIENCE_AGENT_EXECUTOR=jiuwenswarmatasksub-agent's own turn is also driven throughJiuwenSwarmAgent, i.e. a second, nestedPOST /agent/runsto the external adapter rather than an in-process call, which the built-in loop never has. Not yet root-caused precisely (need a closer trace of that nested round trip); aGET /subagentsissued right after the parent's terminal event can still observe the pre-write state. Not fixed here. Now with much stronger evidence of the same underlying mechanism (item 11 under gap 11's UT audit, below): once that nested child run's result reaches the parent's MCP session, JiuwenSwarm 0.2.6's gateway does not reliably resume the parent's own next turn at all — the connection goes silent until ScienceDiscovery's own idle-timeout watchdog fails the run minutes later. Confirmed non-deterministic per test (passes in isolation, fails the same way later in a longer run), consistent with load or state accumulated on the shared JiuwenSwarm+adapter instance rather than a fixed set of affected call shapes. That evidence only exists becausescripts/with-jiuwenswarm.shbriefly forcedSCIENCE_AGENT_JIUWENSWARM_SUBAGENTS=taskto routetaskthrough this bridge at all; since that override is not something any real deployment turns on by default (see 1a), it was reverted rather than kept once the gateway limitation surfaced, andserver.test.tsskips every test that delegates throughtaskunder this executor instead (17 tests,JIUWENSWARM_NESTED_TASK_STALL_SKIP) — both because the bridge's own opt-in stays off in this harness by design, matching real deployments, and because forcing it back on to test it further would reintroduce exactly this non-deterministic stall. - UT audited against a real JiuwenSwarm.
services/api/src/server.test.ts(the host UT tier's biggest suite) now runs its agent turns on a real JiuwenSwarm and adapter throughscripts/with-jiuwenswarm.sh(pnpm ci:ut:hostwires this in). The audit found:A real bug, fixed: a JiuwenSwarm approval question's resource was the call's own descriptive text (
"run_shell: rm -rf out"), which a standing grant made outside the run (a preflight, "always allow this session") could never match — the run had to wait on a decision nobody was going to make, and the eventual answer 404'd once the run gave up waiting.describe_approvalnow sends the bare tool name (toolName) separately; the agent classifiesrun_shell/execute/environment_setup/execution_cancelto the fixed resource (workspace-code) the native loop's own privilege checks use, so a standing grant applies to a JiuwenSwarm-stopped call too.A second real bug, fixed: JiuwenSwarm 0.2.6's gateway sends the run-ending
chat.processing_status(is_complete) right behindchat.ask_user_question, without waiting for the answer, mixed with a variable number of its own bookkeeping frames (an emptychat.final, usage/context accounting). Read literally that ends the run: every approval-gated call produced no tool result and no text. Reproduced live (a fresh deploy, a real browser session) as well as in the audit;gateway.py'sChatRunnow classifies which frames are the model or a tool actually doing something versus this pause's own bookkeeping, and only treats a completion as real once one of those has been seen.Real, JiuwenSwarm-specific behaviour the tests now account for, not bugs: a Session's skills are not a selection (every installed skill is available everywhere; see 1a above for the sub-agent equivalent — every deferred tool, connector MCP tools included, is promoted up front for a sub-agent too); skill-creator is loaded with JiuwenSwarm's
skill_tool, not ourread_skill; a tool's result, re-injected into the model's next turn, arrives wrapped in JiuwenSwarm's own Python-repr envelope ({'result': '...'}, not JSON) rather than verbatim; skill deletion-impact has nothing to report (every skill's selection mode is "all", so nothing uniquely depends on one skill id).A third, unrelated bug in the same sweep, fixed: the suite's own inline model stub for background-execution wake notifications did not unwrap JiuwenSwarm's envelope around the user turn before matching
[Execution notifications]— the same class of bug already fixed intest/helpers/journeys.tsfor the E2E journeys.A fourth bug, fixed: the three tests that switch a session's approval mode (to
always_allow, and toask) while a run is in progress were failing on exact authorization-count assertions (2 !== 1,4 !== 2), not on approval-mode switching itself. Every JiuwenSwarm-gated call was recording twoPermissionAuthorizationrows for one decision:requestApproval(runs/index.ts) records JiuwenSwarm's own ask (or an instantalways_allow/existing_grantauto-allow) before the call is allowed to proceed, then the tool's own built-inrequirePrivilege({action: "code", resource: "workspace-code"})(workspace-bindings.ts) ran again once the call actually reached the bridge, andstore.authorizeByJiuwenSwarmalways minted a brand-new row for that second check with no idempotency against the first. This is why the GitHub UT job was hanging for a full hour: one of the three tests loops an unboundedreader.read()waiting on a stream whose event accounting the double-booking had thrown off, and the host UT tier sets no--test-timeout(only the CodeArts QEMU guest tier does), so nothing but the job's own 60-minute ceiling ever stopped it.authorizeByJiuwenSwarmnow looks up an existing authorization bytoolCallIdbefore creating one (reusing JiuwenSwarm's own decision instead of double-recording it), andsetApprovalMode's mode-switch resolution now carriestoolCallIdthrough so that lookup can find it.A fifth issue, diagnosed and deliberately left as-is:
startSubagentModel's scripted stub (used by many tests) calls a tool literally namedtask, which JiuwenSwarm's own subagent-routing substitution (see 1a) deletes by default before a run ever starts — the stub cannot read the system prompt telling it to callsubagent_spawninstead, so it keeps calling the now-absenttaskand the adapter answersAbility not found in resource_mgr: task, deterministically..ci/run-e2e.shhits the identical mismatch for its own scripted journeys and works around it withSCIENCE_AGENT_JIUWENSWARM_TOOLS=ours; an earlier version of this fix set the narrowerSCIENCE_AGENT_JIUWENSWARM_SUBAGENTS=taskinscripts/with-jiuwenswarm.shfor the same reason, and it did make the stub's calls route correctly — but see the eleventh finding below for why that override was reverted rather than kept.scripts/with-jiuwenswarm.shnow stays at every real deployment's own default (SCIENCE_AGENT_JIUWENSWARM_SUBAGENTSunset) on purpose, and the tests that callstartSubagentModelare skipped under this executor instead (eleventh finding).A sixth bug, fixed:
createApiServer's JiuwenSwarm web-settings sync retries up to 5 times, 3 seconds apart, with no lifecycle tie to the server's own shutdown. Harmless for one long-lived production server; ruinous forpnpm ci:ut:host, where the one shared JiuwenSwarm+adapter instance (scripts/with-jiuwenswarm.sh) is hit by a fresh retry loop from every short-lived test API server the suite starts — hundreds of them, each outliving its own test by up to 15 seconds, all contending for the same adapter. Now tied to anAbortControllerthat the server'sclosehandler aborts.A seventh bug, fixed (three-layer): "API runs a configured OpenAI-compatible model through the gateway and Python" failed on
toolModel.authorizationshaving 4 entries instead of 3 — a real race, not a JiuwenSwarm defect: session-title refinement (see "Session title refinement persists when the naming model finishes after the run stream closes") shares the run's own model and can land its own call on the same mock server any time after the run starts, including after this assertion runs, and a slower executor's real per-turn round trips give that background call more time to win the race. Loosened to>= 3calls, all from the same token. Fixing that exposed a second layer:execution-runscan lag the stream's ownrun.completedby a beat under JiuwenSwarm (the record lands after the bridge's own "completed" round trip settles), so the immediate, un-retried fetch right after the stream closed sometimes read the list before the row landed — now polled. Fixing that exposed a third, this one a stale assertion rather than a race:chartProvenance.body.environmentscross-referencesstore.listEnvironmentRevisions(), which onlyruns/index.ts'ssyncScientificEnvironmentCatalogpopulates, itself gated onrunnerHealth.scientificEnvs.available— andstartTestApi's hand-builtRunnerConfignever setsscientificEnvsEnabled, so that catalog is permanently empty in this harness regardless of executor. The assertion could never have passed here; removed (environment-revision-in-provenance coverage belongs inenvironment.test.ts, which does configure a runner with scientific environments on).An eighth bug, fixed: "skill lifecycle APIs..." failed on a DELETE expected to 409 (blocked by a reference) instead succeeding with 200 — not a bug in the delete-blocking check itself (a separate revision-conflict assertion earlier in the same test, reached first and traced with instrumentation, computed its 0-vs-1 mismatch correctly and threw as expected). The DELETE two lines later was simply missing the
onJiuwenSwarmbranch its own neighboring assertion already carries:impact.body.referencesis asserted empty under JiuwenSwarm one line above (every skill's selection mode resolves to "all", so nothing uniquely depends on one skill id — see 1a), which means nothing blocks that deletion either, so it legitimately succeeds on the first try instead of needing the built-in loop's two-step "409, clear the reference, then 200" dance. Branched to match.A ninth bug, fixed: "native MCP literature flow produces an audited cited summary" read
modelServer.requests[0]assuming it was the run's own first call. Under JiuwenSwarm, importing a run's skills into a still-fresh instance can trigger JiuwenSwarm's own one-off, unrelated model call — building a skill directory tree from every installed skill's name and description — against that same mock server, sometimes landing ahead of the run's own first request. Now finds the request that actually carries the run's own prompt instead of assuming index 0. Separately, and not fixed: the same test's own tool-call visibility (assert.match(stream, /"name":"mcp__pubmed__search"/)) still fails occasionally, matching gap 9 below exactly — pre-existing, already flagged as elusive by that gap's own author, and not something this pass added or could pin down further.A tenth bug, fixed:
run-cancel.test.ts's "cancelling a blocked run persists the approval's terminal state and the tool input in the replay" asserted atool.startedevent in the cancellation replay. Under the built-in looptool.startedfires the moment the model asks for a call, ahead of any approval gate; under JiuwenSwarm it fires only once the call actually begins executing, after approval — so cancelling while still waiting on a decision that never came means the call never started andtool.startednever fires, by design, not by bug. The call's arguments are still in the replay, in thepermission.requiredevent's own summary text; the test now reads them from there under this executor.An eleventh, systemic finding, resolved by not opting into the affected code path: forcing
startSubagentModel's scriptedtaskcalls to route by settingSCIENCE_AGENT_JIUWENSWARM_SUBAGENTS=task(the fifth issue above) does make ScienceDiscovery's own task-delegation bridge work — several tests reach a working subagent turn, with its own approval-gated calls asked, approved, and completed — but then the parent run's JiuwenSwarm gateway connection goes completely silent, until ScienceDiscovery's own idle-timeout watchdog fires minutes later (Agent run stalled: no gateway progress for N ms) and fails the run. This is the same nested round trip flagged unresolved in gap 10 below (atasksub-agent's turn is driven through a second, independentPOST /agent/runs), with concrete, timestamped evidence: after that nested run's result reaches the parent's MCP session, JiuwenSwarm 0.2.6's gateway does not reliably resume the parent's own next turn. Verified against a real local JiuwenSwarm across all 17 tests that delegate throughtask: the stall is not deterministic per test — one test ("API validates subagent Brief v1 structured output...") was observed passing cleanly three times in isolation, then failing the identical way once run later in a longer suite, pointing at load or state accumulated on the one shared JiuwenSwarm+adapter instance rather than a fixed, enumerable set of bad tests, so a per-test allowlist built from one run's observations would stay exposed to the same non-determinism on the next.SCIENCE_AGENT_JIUWENSWARM_SUBAGENTS=taskis, however, an opt-in no real deployment turns on by default (see 1a) —scripts/with-jiuwenswarm.shonly ever set it to keep these scripted test stubs working, not because a deployment needs it. Reverted that override instead of chasing the stall further:scripts/with-jiuwenswarm.shnow stays at the real default (JiuwenSwarm's ownsubagent_spawn/subagent_wait), which involves no nested adapter round trip at all and so cannot hit this gateway limitation. Under that default the scripted stubs'taskcalls go nowhere (Ability not found in resource_mgr: task, deterministic, not the stall) for the structural reason in the fifth issue above — and there is nothing else to usefully check against this executor either way: what these 17 tests verify is ScienceDiscovery's own permission/sandbox/provenance handling of a subagent's actions, which a deployment running JiuwenSwarm's real defaultsubagent_spawnnever reaches (per 1a, that mode bypasses ScienceDiscovery's tools, sandbox, approvals and provenance entirely). All 17 are skipped underonJiuwenSwarm(JIUWENSWARM_NESTED_TASK_STALL_SKIPinserver.test.ts) for that reason and run normally on the built-in loop, where the coverage they exist to provide remains intact. Confirmed materially faster and stall-free this way: the same two files that used to take 7–8 minutes (with occasional multi-minute stalls before the idle-timeout failed them) now finish in about 2 minutes, no test taking longer than 14 seconds.A twelfth bug, fixed: even with the eleventh finding's revert landed, the actual GitHub
UTjob (not just an isolated two-file local run) still failed once, genuinely — not a hang — onrun-cancel.test.ts's "cancelling a queued run does not start it or append it to Session context":waitForGatewayTurn's 400-attempt (10s) wait missed by ~460ms (10461msobserved). That budget was sized from a single turn's round trip measured in isolation (5.0–5.4s, "no other load involved," by the helper's own prior comment). But theut:hostworkload that actually runs in CI isbash scripts/with-jiuwenswarm.sh pnpm --recursive test— the entire workspace's test files, run concurrently, all sharing the one JiuwenSwarm instance and adapterwith-jiuwenswarm.shstarts (by design, matching how one deployment serves many sessions) — so a real turn's round trip under that contention runs longer than the same turn measured alone. ScaledwaitForGatewayTurn's budget to 60s underonJiuwenSwarm(2,400 attempts), matching the marginserver.test.tsalready givesCONCURRENCY_BARRIER_TIMEOUT_MSfor the identical reason. Re-verified against a real local JiuwenSwarm: the test passes. The job's other failure in that same run, "native MCP literature flow produces an audited cited summary" failing itsmcp__pubmed__searchvisibility match, is the pre-existing gap 9 flake the ninth finding above already documents — not new, and not fixed further here.A thirteenth finding, split into a real mitigation and a separate, precisely-diagnosed adapter bug (not yet fixed): chased gap 9's "occasionally invisible" MCP tool call further instead of leaving it at "elusive." Reproduced it on demand against one real local JiuwenSwarm instance (
scripts/with-jiuwenswarm.sh's own model: one shared instance and adapter for the whole command); ran clean 25/25 in isolation first, confirming it needs real contention on the shared instance, not a deterministic code bug. Two different reproduction methods turned up two different things: (a) what CI actually hit, mitigated: with only genuinely varied concurrent load (different packages' own test files running at once, the real shape ofpnpm --recursive test), the run completes successfully and the final answer is correct, but the SSEtool.startedevent formcp__pubmed__searchnever appears — the tool clearly ran (its result reached the model, and is independently confirmed by themcp/invocationsaudit endpoint right after) but the bridge's own live-stream event for it went missing, plausibly fromAgentRunner.ensure_shared_tools's single process-wideasyncio.Lock(services/adapter/src/sciencediscovery_adapter/agent_runs.py) serializing every concurrent run's MCP-tool/permission registration RPCs to JiuwenSwarm's management API. Mitigated two ways:pnpm --recursive(theworkspace-packagesUT workload) now caps at--workspace-concurrency=2(pnpm's own default is 4), halving how many packages' agent turns contend for that lock at once (.ci/ci-contract.mjs's own re-derivation of the command was updated to match); and the test itself now only requires the livetool.startedevent under the built-in loop, relying on themcp/invocationsaudit trail — proven not to have this loss — as the authoritative check underonJiuwenSwarm. (b) a real but separate bug, found only by over-aggressively reproducing (a), and not what CI hit: running several literal copies of the same test file concurrently (not realistic CI load) reliably reproduces a much harsher failure — JiuwenSwarm'sresource_mgrlosing a just-promoted ability, or an approval answer 404ing and the run stalling until the four-minute idle-timeout kills it ([jiuwenswarm-agent] could not answer JiuwenSwarm's approval question call-pubmed: HTTP 404). Traced to ground:events.py's_on_chat_ask_user_questionuses JiuwenSwarm's ownchat.ask_user_questionrequest_idverbatim as both the approval'sidandtoolCallId, and thatrequest_idechoes the model's own tool-call id rather than being a JiuwenSwarm-minted globally-unique value;AgentRunner.pending_approvalsis one dict for the entire adapter process (every concurrent run on the shared instance), keyed by that same value. Two concurrent runs whose models both call a tool under the identical tool-call id (only possible in practice with a hardcoded literal like this test's own mock"call-pubmed", run twice at once — repo-widegrepconfirms it appears nowhere else, so this exact collision cannot occur in a genuine singlepnpm --recursive testpass) collide in that dict: one run's answer silently resolves the other's, and the loser's own answer 404s and its run stalls. Real robustness gap inservices/adapterworth fixing (a production deployment with two sessions whose models happen to pick the same tool-call id would hit the identical collision), but confirmed unrelated to the CI failures this finding set out to explain — not fixed here; flagged separately.A fourteenth finding, fixed (two parts): the concurrency cap from finding 13(a) alone did not make CI green. The very next real CI run failed the same test again, this time on
assert.equal(invocations.body.length, 1)reading0— themcp/invocationsaudit record itself lagged the stream's ownrun.completedby a beat under load, the identical class of lagexecution-runselsewhere in this file already had a fix for (a poll). Applied the same poll here. Root-caused why even the audit trail lags under only--workspace-concurrency=2:AgentRunner.__init__initialized_shared_timeout_s(the ceilingensure_shared_toolshas registered with JiuwenSwarm so far) at0, andtoolTimeoutSecondsvaries per run (Math.ceil(runTimeoutMs / 1000), e.g. deliberately short for the timeout-behavior tests) — so whichever run'sensure_shared_toolscall happens to land first in the concurrent race at the start of the wholeut:hostrun sets the shared ceiling, and every other concurrent run whose own timeout is larger (the common case, since most runs use the 3600s default) then forces a fulldisconnect/delete_custom/register_custom/connectdance under_shared_lockto bump it — serializing every other concurrent run behind that dance for its duration, which is exactly consistent with these lagging/missing events clustering early in a run. Fixed by initializing_shared_timeout_stosettings.tool_timeout_s(the configured ceiling, which a run that omitstoolTimeoutSecondsalready falls back to) instead of0, so only a run that genuinely asks for more than the configured default still triggers that dance for this reason —services/adapter/tests/test_agent_runs.py's existing coverage of the registration sequence and the timeout value sent tomcp.register_customwas unaffected (both already exercised the case where the first call's own timeout equals the configured default). This does not remove_shared_lockitself, which is a real correctness requirement (JiuwenSwarm exposes exactly one named custom MCP server,sci, for the whole adapter process — seemcp_server.py's own module docstring — so concurrent registrations against it and against the adapter's own in-process registry state need mutual exclusion); it only removes one avoidable, frequent trigger for the expensive path inside it.A fifteenth finding, root cause found and fixed — this is what findings 13 and 14 were symptoms of. Neither the concurrency cap nor the
_shared_timeout_sfix made the literature-flow test reliably green: a fully-serialized run of every test file inservices/api(--test-concurrency=1, so no two files' agent turns could ever be simultaneous) still failed it identically, disproving concurrency as the mechanism outright. Root-caused with temporary unconditionalprint()s inmcp_server.py'shandle_rpc(Python'slogging.warningcalls in that module turned out not to be reaching the log at all — this repository configures no root logger, anduvicorn.run(..., log_level="info")does not implicitly wire one up for application loggers, so the existinglogger.warning/logger.exceptioncalls in this exact code path had been silent the whole time) against a real local JiuwenSwarm:tools/listwas fetched by JiuwenSwarm exactly once for an entire multi-file run, at whichever session happened to register first, and never again — even though several other sessions'ensure_shared_toolscalls visibly ran afterward and updated the adapter's ownregistry.sharedwith tools that first session never had. Every one of literature-flow'stools/callattempts (tool_search, thenmcp__pubmed__search) is absent frommcp_server.py's traffic entirely, for its whole run: JiuwenSwarm never once tried to call them, because it never knew they existed. The actual bug:ensure_shared_tools's own comment claims "a reconnect makes JiuwenSwarm read the list afresh", but that is only true of the fulldisconnect/delete_custom/register_custom/connectcycle — a baremcp.connecton an already-connected server (the path taken whenchanged=Truebut the session was neither the first ever nor asking for a longer timeout, i.e. exactly literature-flow's case: a genuinely new tool, default timeout) does not make JiuwenSwarm re-fetch anything. So any session whose only reason to touchensure_shared_toolswas bringing a tool nobody had registered yet — an entirely ordinary thing for different sessions to do — left that tool permanently uncallable for the rest of the process's lifetime, deterministically once the "first mover" wasn't the one carrying it.changednow forces the same full reconnect cycle as the first-ever registration and a larger timeout, since that is what actually makes JiuwenSwarm re-read the list;test_agent_runs.py's sequence assertion updated to match. Verified: the adapter suite (170 passed, 8 skipped), a live local JiuwenSwarm repro that had failed on its first attempt every time before this fix (5/5 clean after), and — the real test — the next GitHubUTjob on this exact commit passed outright. Findings 13 and 14's own mitigations (the concurrency cap, the_shared_timeout_sceiling, themcp/invocationspoll, skipping the livetool.startedcheck under this executor) are kept: they are reasonable margin on their own terms (theexecution-runs-style poll in particular guards a real, separate, already-documented timing lag), but none of them were actually load-bearing for this bug, and the "upstream JiuwenSwarm reliability boundary" framing in 13(a) was wrong — this was an adapter-side logic bug the whole time, just one concurrency made easy to trip over (more sessions racing to register more varied tool sets at once) without being its cause.
Start working on a sub-issue
- Start the stack:
scripts/jiuwenswarm.sh setuponce, then./scripts/start-stack.sh --mode local(check withGET /agent/info); see Local mode. - Find your routes in
test/contract/routes.json(node test/contract/run.mjs --coveragelists the rows without a case). - Add a case under
test/contract/cases/, record it on a fresh data directory against the built-in loop, compare it against the adapter + JiuwenSwarm stack. Rules, the SSE step form and normalization:test/contract/README.md. Baselines are read-only for agents; a change needs human review. - For behaviour, use the run-event cases (
l2-runs.json); the stub model istest/contract/stub-model.mjs. - Browser journeys:
.ci/run-e2e.sh mocked(JiuwenSwarm is the default backend now; it installs and starts its own instance if needed).CI_E2E_BACKEND=legacy .ci/run-e2e.sh mockedfor the built-in loop while it still exists.
Tests
UV_PROJECT_ENVIRONMENT=/tmp/adapter-venv uv sync --extra test --project services/adapter
/tmp/adapter-venv/bin/python -m pytest services/adapter # adapter unit tests
cd services/api && pnpm build && node --test dist/agent-run/jiuwenswarm-agent.test.js
node --test test/contract/*.test.mjs # tooling of the contract tests
E2E_BASE_URL=... E2E_API_TOKEN=... node test/contract/run.mjs --compare test/contract/baselines/legacy-linux.json
# UT (server.test.ts and the rest of the shared plan) runs against a real JiuwenSwarm and
# adapter through scripts/with-jiuwenswarm.sh, which pnpm ci:ut wraps around the shared runner:
scripts/with-jiuwenswarm.sh pnpm --filter @sciencediscovery/api test