ScienceDiscovery
中文 GitHub

Web Search and Web Fetch

This page documents the product-owned web_search and web_fetch tools. They are visible by default in the native executor. The default JiuwenSwarm backend exposes its own free_search/paid_search and fetch_webpage instead; set SCIENCE_AGENT_JIUWENSWARM_TOOLS=ours to use the product tools. The permission, provider, cache, audit, and slash-command details below describe product Web calls; do not assume they apply to Swarm's Web tools. See Agent backends.

Boundary

Product Web is a global base capability, not an MCP Source, and has no Session-level provider override. When the product tools are selected, the model sees web_search and web_fetch. Node owns permission, credentials, cache, CAS, and audit, and calls the vendors itself.

Agent tool → WebBroker ──────────────→ NativeWebProviderClient ──outbound HTTPS──▶ vendor API
             permission/credentials     argument validation + dispatch + 1 MB cap
             cache/CAS/audit            web_fetch additionally validates a public URL (with DNS)

There is no Python sidecar hop: POST /internal/web/invoke and the gateway web router were removed with this change, and the deerflow dependency is gone from the gateway environment. Node remains the product-configuration source of truth, and every vendor host is declared under web.* in config/external-urls.json so mirrored or restricted networks can retarget them.

Search as one aggregated capability

Search is no longer a provider the user picks. web_search walks a fixed list of engines and returns the first one that produces results:

  1. Paid tier, in order: Tavily → Exa → Brave Search API. A provider is attempted only if it is switched on and has a stored API key; an unkeyed provider is skipped without a request.
  2. Free tier, in order: DuckDuckGo → Bing → Brave (public result page). Each engine has its own switch; a switched-off engine is never requested.

Paid before free is deliberate: a keyed vendor gives a stable API contract and better results, and the free engines are there so search still works when no key is configured, a key is exhausted, or one engine starts rate-limiting.

The free engines read the vendors' public result pages, so all three share one failure mode — a layout change yields zero rows. That is treated as a failed attempt, and the aggregation moves to the next engine; only when every candidate fails does the call fail. Bing's organic links are wrapped in bing.com/ck/a redirects, which the client decodes back to the destination so that citations and web_fetch see the real page rather than a tracker.

If no engine is available at all — no paid key and every free engine off — the call fails as INVALID_INPUT pointing at Web settings, without contacting anyone.

Providers and configuration

System configuration → Web providers offers:

Each engine caches under its own route, and a cached answer from any candidate satisfies the query, so a successful free result is not re-fetched through a paid provider on the next call. Every attempt — paid and free, including cache hits — is recorded in one WebInvocation with its engine, tier (paid/free), Jina endpoint, proxy mode, and whether proxying was used, but never a key or proxy URL.

Every provider consumes the proxy policy the broker resolved, applied through the shared proxyDispatcher (which keeps NO_PROXY and protocol selection semantics). With no subprocess left, custom/direct mode needs no process isolation and proxy scope is inherently per-request.

Upgrading from the single-provider settings

Installs configured before aggregation stored one searchProvider (plus ddgsBackend and an optional searchFallbackProvider). Those records are translated on load so spending does not change:

"DDGS" is no longer a user-facing name. The Python ddgs library it referred to was a multi-engine aggregator (its bing backend was disabled upstream in 9.x and silently fell back to auto); the aggregation described above replaces that behaviour with engines this repository owns.

Permission, security, and citation

The native Node loop neutralizes framework tags in remote tool results before they reach history or the UI (see Agent backend). JiuwenSwarm has a separate tool/result path; the native-loop statement is not a guarantee for its Web tools. Outbound sensitive information is an authorization concern before the call; scientific quality still needs later review.

Slash commands

The first version does not support JavaScript browser rendering, authenticated pages, Browserless/Crawl4AI, SearXNG, or automatic MemoryGraph writes.