Research: The Server Was Fine: Why MCP Health Checks Keep Marking Working Servers Down
Operating MCP at scale, part three: reliability.
On September 30, 2026, Slack's production MCP server answered every tool call a
client sent it, and the client still listed it as "connecting" [1]. The client's
health check had asked the server for ping, the server had replied that it did
not know that method, and the client had read the reply as a dead peer. That was
one of five cases written up in September in which a health check looked at a
working MCP server and marked it down, each on a different codebase [1][2][3][4][5].
This article shows how it happens, why the probe that replaces ping has a trap
of its own, and what a health check looks like when it has to survive a fleet
that speaks two protocol revisions at once.
Let me start with the rule everyone was following.
Clients were told to ping
Three parties meet on every one of these connections: a client or gateway built
to an older revision of the specification, a server that does not serve ping,
and the specification itself, which changed its mind between them.
The 2025-11-25 revision was explicit. A receiver "MUST respond promptly with an empty response," implementations "SHOULD periodically issue pings to detect connection health," and "Multiple failed pings MAY trigger connection reset" [6]. A client that pings every ten seconds and resets after three failures is following that text to the letter.
On a normal day that loop is invisible. The exchange is one line each way:
→ {"jsonrpc":"2.0","id":2,"method":"ping"}
← {"jsonrpc":"2.0","id":2,"result":{}}
The server is marked ready, tools route to it, and nobody thinks about the probe again until a server truly stops answering.
Then the 2026-07-28 revision removed ping, along with logging/setLevel and
notifications/roots/list_changed [7]. A server on the current revision answers
an unknown method with -32601, method not found, and on Streamable HTTP with a
404 carrying that code [8]. Some servers send that reply even when they
negotiated an older revision that still requires ping [1]. Clients written to
the old rule kept pinging, and a refusal reads to them like a failure.
Five working servers, marked down
Slack. The client was penelope, and its maintainer was validating Slack's
hosted server, which had negotiated protocol 2025-06-18 [1]. The client's own
test command passed: 26 tools listed and six real calls succeeded. Its server
list still showed Slack as connecting, with an error ending "Method not found:
ping." Its diagnostic command raised a warning and suggested reading the
logs and restarting the server, which changed nothing, because nothing was
broken. The maintainer filed the issue and closed it the same afternoon [1].
A desktop-automation server. cua-driver 0.22.0 on Windows answered the
same probe like this [2]:
→ {"jsonrpc":"2.0","id":2,"method":"ping"}
← {"jsonrpc":"2.0","id":2,"error":{"code":-32601,"message":"Unknown method: ping"}}
A gateway hosting it as a stdio child counted every reply as a failure and
restarted it on a cycle of 53 to 57 seconds: 131 and 132 disconnects in two hours
on two hosts, with windows in which clients saw zero tools [2]. With no pings
sent, the same process ran for 84 seconds without exiting. The workaround was to
tell the gateway not to ping, and the reporter named its cost: slower detection
of children that really are dead [2]. Here the server was the one out of line,
and the reporter asked it to implement ping. The outcome was the same. A server
that could do all of its work was taken offline roughly once a minute over the
one method it did not serve.
A gateway's own test backend. The clearest record comes from a gateway
project that caught the failure in its own load test before release [5]. Its
next major version probed every backend with ping every ten seconds and
escalated on the third refusal. Its benchmark fixture had never implemented
ping. The log reads:
11:45:20 Health probe was not served method="ping" code=-32601 consecutive=1
11:45:30 Health probe was not served method="ping" code=-32601 consecutive=2
11:45:40 record_failure{reason="health probe unserved"} failures=1..5 threshold=5
11:45:40 Circuit breaker opened backend=workload reason=health probe unserved
11:45:40 Circuit open, rejecting request
About thirty seconds after start, the breaker opened and the backend shed all of
its traffic. The breaker rebuilt the transport, and the rebuilt process still did
not serve ping, so the cycle repeated. Under a 60-second load with 50 virtual
users, tools/call succeeded 48.7% of the time against 100% on the previous
release, with 0.00% HTTP errors: every failure was an HTTP 200 carrying a
JSON-RPC error [5]. An alert built on HTTP error rates would have stayed quiet.
The project fixed it the next day by treating -32601 as proof of life [5].
Two more, with other causes. A virtual-MCP layer probed its backends with
HTTP GET, which Streamable HTTP makes optional. The report's example was
Tableau's MCP server, which accepts POST only. A backend answering 405 or 400,
as the transport allows, was excluded from tool routing while initialize,
ping and tools/list all worked over POST [3]. A gateway
registry's health service skipped notifications/initialized, so its next ping
got 404 Session not found, and a hosted Salesforce server was marked unhealthy
with zero tools [4].
That is three different root causes and one shape. In every case a server answered, the answer was a JSON-RPC error in three cases and an HTTP status in two, and the health check treated an answer as an absence.
Any reply proves the server is alive
The maintainers explained the removal of ping in SEP-2575:
"Client-to-server ping is also removed because any normal RPC call already proves server liveness, and transport-layer mechanisms (HTTP keep-alives, SSE comments, STDIO process status) handle connection-health checks more appropriately." [9]
That reasoning holds, and it is also the fix. A server that sends back -32601
has received the request, parsed it, and written a reply. Only silence, a refused
connection or a timeout says otherwise.
The removal did not create this failure class, since two of the five cases have
nothing to do with ping. It did enlarge it, and it will stay enlarged for as
long as clients and servers from two revisions share a fleet.
The replacement probe can be answered from a cache for an hour
server/discover takes over from ping, and for readiness it is better. It is a
mandatory call that returns supported protocol versions, capabilities and
identity, so one request tells you the peer is alive and which revision it
speaks [10]. In a mixed fleet the second answer is the one you need.
It is also a capability read, and capability reads are cacheable. The caching
page lists server/discover first among the results on which "Servers MUST
include caching hints" [11]. The server/discover page's own example response
carries "ttlMs": 3600000, "cacheScope": "public" [10], and the caching page
defines a public response as one that "Any client, shared gateway, or caching
proxy MAY store and serve the cached response to any user" [11].
Put those three lines together and a shared gateway may answer your liveness
probe from its cache for an hour without the backend being involved. The
changelog's list of cacheable results omits server/discover [7], so a team that
reads only the changelog will not see it. If you own the server, return
ttlMs: 0, which the caching page says "SHOULD be considered immediately stale"
[11]. If you do not, probe the backend directly.
Nothing upstream will catch it for you
Each implementer in those five cases wrote its own health rule, because there was no current one to copy. The official client best-practices page carries no guidance on health checks, timeouts, retries or reconnection [12].
The gateways have built more resilience than their reputation suggests, and the
gap is in the defaults. ContextForge has had exponential backoff with jitter in
tree since July 2025, wired into its gateway and tool services [13]. Its health
checker flips a reachable flag, and the next passing probe brings a backend
back on its own; only a gateway an operator disabled by hand stays down [14]. Its
per-tool circuit breaker, with a half-open trial request, exists as a plugin and
ships with mode: "disabled" in the default configuration [15]. Install it and
change nothing, and you get retry and recovery with no breaker.
Nor will anyone you depend on hand you an availability number to alert against. The MCP Registry working group lists "Registry uptime ≥ 99.9% with automated monitoring and alerting" among its success criteria [16], while the registry's terms of service disclaim any guarantee [17], so that is an objective. TrueFoundry's SLA commits to 99.9% and names "MCP control surfaces" [18], and MintMCP's status page shows an "MCP Gateway" component with a ninety-day uptime bar [19]. Azure API Management documents its MCP feature with no MCP-specific availability commitment [20]. The hosted servers in your critical path show a status light and no number, so the objective is yours to set.
What to do, depending on who you are
Here is the whole argument as a probe specification you can check line by line against the sources:
# MCP backend health probe, for a fleet that mixes 2025-11-25 and 2026-07-28 peers
probe:
method: server/discover # replaces ping; also reports the revision spoken [10]
path: direct-to-backend # never through a shared gateway or caching proxy [11]
alive_if: any_response # a JSON-RPC error or HTTP status is still an answer [9]
treat_as_alive:
- jsonrpc_error: -32601 # unknown method, the current-revision refusal of ping [8]
- http_status: 404 # Streamable HTTP carrying -32601 [8]
dead_only_if:
- connect_failure
- timeout
server_side:
server_discover_ttl_ms: 0 # "SHOULD be considered immediately stale" [11]
alerting:
objective: 0.999 # yours; hosted servers publish none [20]
page: [{window: 1h, short: 5m, burn: 14.4}, {window: 6h, short: 30m, burn: 6}]
ticket: [{window: 3d, short: 6h, burn: 1}] # Google SRE Workbook Table 5-8 [21]
If you write an MCP client or gateway, stop health-checking with ping and
count any reply as proof of life. The gateway in the third case fixed its own
bug in a day by doing exactly that [5]. Any probe that still counts -32601 as a
failure will mark down the next working server that declines.
If you operate a platform, check your gateway's defaults before its feature list, because the breaker you are counting on may ship disabled [15]. Alert on JSON-RPC errors as well as HTTP status, since the third case lost half its tool calls behind a clean HTTP error rate [5]. Set your own objective at the MCP boundary and use the burn rates above, which come from Google's SRE Workbook for a 99.9% objective [21].
If you run an MCP server, return ttlMs: 0 on server/discover so no
intermediary can answer a probe in your place, and serve every method your
negotiated revision still requires. Slack's server negotiated a revision that
requires ping and refused it [1].
What is still unsolved
The protocol has no shared answer to "is this server healthy." The specification removed the old probe for a sound reason, the client best-practices page says nothing about health checks [12], and the replacement probe is cacheable by default. Until that page carries guidance, every client and gateway will keep writing its own rule, and the September list will keep growing.
The limit of this article is that its evidence is five issue reports, most of
them filed by the people who found and fixed the bug. I know of no published
measurement of how many MCP deployments probe with ping today, and no
post-mortem of a production outage caused by one. The five cases show the
mechanism; they cannot tell you how often it is costing anyone traffic.
What it adds up to
In each of the five cases the server was fine and the check was wrong. The
revision removed ping, production servers answer it with an error, and any
answer at all proves the server is alive. Treat it that way, make sure the probe
that replaced ping cannot be answered from a cache, and read your gateway's
defaults before you rely on its features.
Part three of five on operating MCP at scale. Parts one and two cover upgrading a fleet to the 2026-07-28 revision and security; parts four and five cover performance and cost.
Sources
All URLs verified 2026-10-05; issue reports [1], [2], [3] and [5] re-read on 2026-10-06.
edouard-claude/penelope#276, Slack's production server held in "connecting", protocol 2025-06-18. https://github.com/edouard-claude/penelope/issues/276trycua/cua#4001, a server killed and restarted on a cycle. https://github.com/trycua/cua/issues/4001stacklok/toolhive#6497, GET probes excluding conformant backends. https://github.com/stacklok/toolhive/issues/6497agentic-community/mcp-gateway-registry#1817, the skipped initialization and the unhealthy Salesforce server. https://github.com/agentic-community/mcp-gateway-registry/issues/1817MikkoParkkola/mcp-gateway#567, the benchmark fixture, and PR #576, the fix merged 2026-09-19. https://github.com/MikkoParkkola/mcp-gateway/issues/567 and https://github.com/MikkoParkkola/mcp-gateway/pull/576- Ping, revision 2025-11-25, for the periodic-ping recommendation and connection reset. https://modelcontextprotocol.io/specification/2025-11-25/basic/utilities/ping
- Changelog 2026-07-28, for the removal of
pingand the list of cacheable results. https://modelcontextprotocol.io/specification/2026-07-28/changelog - Streamable HTTP, for the
404carrying-32601on an unknown method. https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/streamable-http - SEP-2575, Make MCP Stateless, for the rationale behind removing
ping. https://github.com/modelcontextprotocol/modelcontextprotocol/blob/main/seps/2575-stateless-mcp.md server/discover, for the method and the example response carryingttlMsandcacheScope. https://modelcontextprotocol.io/specification/2026-07-28/server/discover- Caching, for the cacheable-results list, the scope table and
ttlMs: 0. https://modelcontextprotocol.io/specification/2026-07-28/server/utilities/caching - Client best practices, which carries no health check, timeout, retry or reconnection guidance. https://modelcontextprotocol.io/docs/2026-07-28/develop/clients/client-best-practices
- ContextForge retry manager, exponential backoff with jitter. https://github.com/IBM/mcp-context-forge/blob/main/mcpgateway/utils/retry_manager.py
- ContextForge ADR-0009, built-in health checks and automatic reactivation. https://github.com/IBM/mcp-context-forge/blob/main/docs/docs/architecture/adr/009-built-in-health-checks.md
- ContextForge circuit-breaker plugin, and the default plugin configuration that ships it disabled. https://github.com/IBM/mcp-context-forge/tree/main/plugins/circuit_breaker and https://github.com/IBM/mcp-context-forge/blob/main/plugins/config.yaml
- MCP Registry working group charter, the 99.9% success criterion. https://modelcontextprotocol.io/community/working-groups/registry
- MCP Registry terms of service, which disclaim any guarantee of availability. https://modelcontextprotocol.io/registry/terms-of-service
- TrueFoundry service level agreement, naming MCP control surfaces. https://www.truefoundry.com/service-level-agreement
- MintMCP status page, MCP Gateway component. https://status.mintmcp.com/
- Azure API Management, overview of MCP servers. https://learn.microsoft.com/en-us/azure/api-management/mcp-server-overview
- Google SRE Workbook, alerting on SLOs, Table 5-8. https://sre.google/workbook/alerting-on-slos/
Michael Rishi Forrester is the AI Workforce Transformation Lead at Accenture LearnVantage, where he works on Claude and generative AI adoption with governments and forward deployed engineers. With nearly 30 years in technical leadership, workforce transformation, and artificial intelligence, his work centers on enterprise AI governance and adoption and on closing the gap between what executives expect from AI and what AI can actually deliver.