The server core is synchronous: Snodo.Server.dispatch/3 runs one JSON-RPC request in the calling process against an immutable runtime. Transports add framing, concurrency, cancellation, and delivery around that core.

raw JSON-RPC map
  -> protocol registry and exact profile inspection
  -> selected dialect: context and admission
  -> core operation or negotiated extension route
  -> router callback (tool, resource, prompt, completion)
  -> protocol-neutral result, or an opened subscription
  -> selected dialect: wire and stream shaping

Executor

Snodo.Server.Executor is the transport-neutral execution layer: bounded admission, a bounded queue, deadlines, cancellation tokens, and supervised worker tasks, with cleanup when the submitting process dies. The stdio adapter and the HTTP listener use it; an application can inject its own executor.

Limits on message content

Every transport decodes JSON with the same limits. An integer literal longer than 64 digits is a parse error (-32700). A request id must be a string of at most 256 bytes or an integer in the int64 range (-32600 otherwise), and a progressToken must meet the same bounds (-32602 otherwise). These values are echoed in every response and notification for a request, so an error for a request whose id is out of bounds carries a null id instead.

Resource template matching is linear in the length of the requested URI, with no backtracking, so the message size limit bounds it. Templates themselves are limited to 1,024 bytes and 32 variables when the resource module compiles; see Components.

Stdio

:ok = Snodo.Transport.Stdio.serve(MyServer.runtime())

serve/2 runs until EOF and until admitted requests finish, then returns. Messages are one JSON object per line. Requests run concurrently, responses are written atomically, and notifications/cancelled stops a request with no late response. Output escapes non-ASCII characters as JSON \u escapes, so the bytes do not depend on the device's encoding.

Options include :input and :output devices, :write_timeout (default 5,000 ms; a blocked write is terminal), :max_line_bytes (default 2,000,000; a longer message is refused with -32600 before decoding), :request_timeout, and :executor. A leading UTF-8 byte order mark is ignored.

Standard input is read in chunks that hold at most :max_line_bytes of a line when the VM runs with -noinput, and on OTP 28 and later when stdin is a socket, as Node.js clients provide. Otherwise the VM holds each line in full before the limit applies. Clients that start the server with a pipe, as Python and Erlang clients do, are bounded only with -noinput: elixir --erl -noinput, emu_args for an escript, or vm.args for a release.

Logger output is redirected away from stdout by default, because stdout carries only protocol messages.

Native HTTP listener

children = [
  {Snodo.Transport.StreamableHTTP.Server, runtime: MyServer.runtime(), port: 4000}
]

A dependency-free listener that binds to 127.0.0.1 by default and serves POST /mcp. It answers one request per connection:

  • application/json for ordinary results.
  • text/event-stream when a handler reports progress, and for subscriptions/listen, with keepalive comments and proxy buffering disabled.
  • 405 for GET and DELETE at /mcp. No session IDs are issued.

It checks media types, the mirrored MCP-Protocol-Version, Mcp-Method, and Mcp-Name headers, the Mcp-Param-* headers for a tool's x-mcp-header arguments (see Components), and Origin when present: loopback names by default, :allowed_origin_hosts to change, where an entry with a port ("localhost:3000") pins the port and an Origin with userinfo is refused. :allowed_hosts additionally requires the Host header to name a listed host. It is off by default, because a reverse proxy commonly forwards the public name. A client disconnect cancels the request. Options include :ip, :port, :path, :request_timeout, :read_timeout, :max_header_bytes (32 KB, including the final \r\n\r\n), and :max_body_bytes (2 MB). Body bytes received with the head count only toward :max_body_bytes.

The listener is unauthenticated unless :request_gate is set to a {module, options} implementing Snodo.Transport.StreamableHTTP.RequestGate. The gate sees the parsed method, path, and headers before the body is read. It can serve a separate route, refuse the request, or return trusted identity for Snodo.Authorization. An invalid result or exception fails closed with 500. The listener bounds the gate with :request_gate_timeout; the gate must also bound and clean up any external calls it makes. For OAuth bearer tokens and protected resource metadata, use Snodo.OAuth.ResourceServer.Native from snodo_oauth.

Binding :ip outside loopback logs a startup warning, even with a request gate: the native listener still serves plaintext HTTP. Bind it privately behind a TLS reverse proxy, or use Snodo.Transport.Plug in an HTTP server with TLS. Configure authentication with a request gate or in the Plug pipeline.

The listener also bounds what clients can hold open:

  • :max_connections (default 1,024). A connection accepted at the limit is closed without being read.
  • :head_timeout (default 10,000 ms from accept). The request head must be complete by then; :read_timeout (5,000 ms) still bounds each read.
  • :request_gate_timeout (default 10,000 ms). A gate that does not return in time gets a 504 response.
  • :body_timeout (default 10,000 ms from the end of the head, or from gate admission when a gate is configured). The request body must be read by then; :read_timeout still bounds each read.
  • :max_subscriptions (default 256). A subscriptions/listen stream over the limit is closed at its source and the request gets 503. A slot returns when the connection serving a stream exits, including a client disconnect.
  • :drain_timeout (default 5,000 ms). How long a stopping listener waits for open connections; see Shutdown.

Snodo.Transport.StreamableHTTP.Server.url/1 returns the endpoint URL, which is useful with port: 0 in tests.

Shutdown

When the listener stops, including when its supervisor shuts it down, it drains before exiting:

  1. The listening socket closes. New connections are refused.
  2. A connection whose request head or body is still arriving gets 503 and closes. A request read in full before the drain began runs to its response.
  3. Each open subscriptions/listen stream gets its successful completion response, its source is closed with reason :shutdown, and the connection closes.
  4. Connections still open after :drain_timeout (default 5,000 ms) are killed. Their executor work is cancelled and their subscription sources are closed.
  5. An executor the listener started is stopped. An executor passed with :executor keeps running.

The default drain covers ordinary tool calls and fits, with the rest of an application's shutdown, inside common stop windows such as Docker's 10-second default. A platform with a shorter stop window, such as Fly's default 5-second kill_timeout, needs a smaller :drain_timeout or a longer platform timeout. Raise it toward :request_timeout when requests run longer and the host allows a longer stop.

The listener's child spec sets :shutdown to :drain_timeout plus 5,000 ms so the supervisor waits for the drain. A child spec that overrides :shutdown must keep it above :drain_timeout; a supervisor that kills the listener earlier skips the rest of the drain. The release or platform that stops the VM must allow for the drain as well.

Component limits

The opt-in wrappers in Snodo.Component.Wrap bound individual tools, resources, and prompts after router checks. They apply to direct, stdio, and HTTP requests alike. Their defaults are:

WrapperDefaultDefined failure
Timeouttimeout: 5_000 msComponent timed out
Concurrencylimit: 32 in-flight callbacks per component and keyComponent concurrency limit reached
RateLimitlimit: 60 calls per window_ms: 60_000, max_keys: 10_000 per componentComponent rate limit reached or Component rate key capacity reached

All numeric options must be positive integers; an invalid declaration fails compilation. Concurrency and rate wrappers use one shared bucket by default. Pass key: fn context -> ... end to group by a stable context value such as a verified principal. The core application owns the counters. Concurrency permits return when a callback finishes or its process exits. Rate buckets start with the first call for a key and reset after the configured window; expired keys are pruned when capacity is needed. Counts reset if the core application restarts. A rejected rate call consumes no callback work, while an admitted call counts even if its handler later fails.

The Timeout wrapper runs the callback in a supervised task linked to the request owner. The shared task supervisor admits at most 1,024 concurrent timeout workers and returns Component timeout worker unavailable when full. It stops a task on timeout or owner termination, but cannot undo external side effects. The built-ins return tool isError results, or JSON-RPC error -32603 for resources and prompts. If the core application is not running, they return a defined unavailable error.

Plug and Bandit

The snodo_plug package provides Snodo.Transport.Plug for applications that already run Plug, Bandit, or Phoenix, with their own authentication pipeline, TLS, and timeouts:

children = [
  {Snodo.Server.Executor, name: MyApp.SnodoExecutor, max_concurrency: 32, max_queue: 128},
  {Bandit,
   plug: {Snodo.Transport.Plug, runtime: MyServer.runtime(), executor: MyApp.SnodoExecutor},
   ip: {127, 0, 0, 1},
   port: 4000}
]

It supports the same JSON, progress SSE, and subscription SSE lifecycles. Plug only reveals a disconnect when a write fails, so a request still running after :disconnect_probe_ms (default 5,000) switches to SSE and writes keepalive comments; a failed write cancels the work. :max_subscriptions (default 256) bounds open subscription streams; the count is held in the executor, so Plugs that share an executor share it. :body_timeout (default 10,000 ms from when the Plug starts reading) bounds the whole request body and answers 408; a request that declares Transfer-Encoding gets 411 without being read. Over HTTP/2 the deadline is checked only when an adapter read returns, so it does not bound a client that keeps sending small DATA frames. See the snodo_plug documentation and the application stack for choosing between the native listener and Plug.

Phoenix endpoint

Add snodo_plug and Bandit to the Phoenix application's dependencies. Keep the existing endpoint; start the executor before it in the application supervisor:

children = [
  {Snodo.Server.Executor,
   name: MyApp.SnodoExecutor, max_concurrency: 32, max_queue: 128},
  MyAppWeb.Endpoint
]

Mount the transport in a dedicated router pipeline. The authentication Plug must verify the caller, assign trusted identity with Plug.Conn.assign(conn, :mcp_auth, %{principal: user.id}), and send a 401 or 403 response with halt/1 when access is denied. Do not copy an unverified header into the assign. A verified client-instance identifier may also be assigned to :mcp_cancellation_scope for cross-request cancellation; distinguish client instances even when they share a user. The transport does not authenticate callers itself. For OAuth 2.1 bearer tokens, the snodo_oauth package supplies that plug, the protected resource metadata document, and a scope policy.

pipeline :mcp do
  plug MyAppWeb.Plugs.MCPAuth
end

scope "/" do
  pipe_through :mcp

  forward "/mcp", Snodo.Transport.Plug,
    runtime: MyApp.MCPServer.runtime(),
    executor: MyApp.SnodoExecutor,
    path: "/mcp",
    max_body_bytes: 2_000_000,
    read_timeout: 5_000,
    body_timeout: 10_000,
    request_timeout: 30_000,
    allowed_origin_hosts: ["app.example.com"],
    allowed_hosts: ["app.example.com"]
end

Phoenix's forward removes the /mcp prefix from conn.path_info and adds it to conn.script_name, but leaves conn.request_path as /mcp. Snodo.Transport.Plug compares path: with conn.request_path, so keep the full public path in that option. Replace the example hosts with the public host names used by clients. An Origin header, when present, is checked against allowed_origin_hosts; a value such as "localhost:3000" also restricts the port. This is a host check, not a CORS policy. allowed_hosts checks the request Host header, so include the Host value your reverse proxy forwards. See the snodo_plug options for defaults and response behavior.

The Plug needs the raw request body. Phoenix endpoint plugs run before the router pipelines, and a generated endpoint commonly runs Plug.Parsers there. Replace that parser plug with a conditional wrapper before plug MyAppWeb.Router; keep the application's existing parser options for all other paths:

@parser_opts Plug.Parsers.init(
  parsers: [:urlencoded, :multipart, :json],
  pass: ["*/*"],
  json_decoder: Jason
)

plug :parse_non_mcp
plug MyAppWeb.Router

defp parse_non_mcp(%Plug.Conn{request_path: "/mcp"} = conn, _opts), do: conn
defp parse_non_mcp(conn, _opts), do: Plug.Parsers.call(conn, @parser_opts)

Do not add a body parser to the :mcp router pipeline. The MCP request requires Content-Length, and a request declaring Transfer-Encoding gets 411.

With Bandit's Phoenix adapter, set endpoint HTTP bounds alongside the transport bounds above. For example, merge these options into the application's existing endpoint configuration:

config :my_app, MyAppWeb.Endpoint,
  adapter: Bandit.PhoenixAdapter,
  http: [
    port: 4000,
    http_1_options: [max_header_length: 10_000, max_header_count: 50],
    http_2_options: [enabled: false],
    thousand_island_options: [
      transport_options: [send_timeout: 5_000, send_timeout_close: true]
    ]
  ]

Set connection, header, read, write, and shutdown limits on the endpoint and reverse proxy as well. The transport's body_timeout is checked when an adapter read returns; under HTTP/2, a client sending small DATA frames can keep that read open. Disable HTTP/2 on an endpoint serving untrusted MCP clients, as above, or buffer request bodies at a proxy. Disabling it on the endpoint affects all routes there. The finite send timeout bounds writes that the transport cannot interrupt. The 21_plug_bandit.exs example checks the transport without adding Phoenix to this repository's dependencies.

Other hosts

Snodo.Transport.StreamableHTTP.prepare/3 admits and decodes a Snodo.Transport.StreamableHTTP.Request without running a handler. It returns {:ok, prepared}, or {:response, response} when admission fails. execute/3 runs a prepared request and returns a Response or a StreamResponse. A different HTTP server can translate its requests into that shape. handle/3 runs both steps synchronously.

HTTP client connections

Snodo.Client.connect({:http, url}, opts) keeps up to four pooled requests open for that client and origin. Extra concurrent requests wait up to their request timeout for a free connection. Set :pool_size to change the cap (default 4), :pool_idle_timeout to change how long an unused connection stays open (default 30,000 ms), and :pool_max_requests to retire a connection after a fixed number of requests (default 100). Each option requires a positive integer.

The client reuses only HTTP/1.1 responses with a complete Content-Length body and no Connection: close directive. Event streams, responses framed by socket close or chunked transfer, malformed responses, and failed requests close their sockets. Event streams detach from the pool and use a separate socket while they run. Snodo.Client.close/1 closes idle connections; active requests finish on their own sockets. Response-size and page limits apply on every request, including those sent over reused connections.

Examples

examples/04_stdio_concurrency.exs, 05_http_tools.exs, 21_plug_bandit.exs (from integrations/plug), and 24_client_transports.exs.