Research: The MCP Cost Nobody Invoices: Tool Definitions Resent on Every Call
Operating MCP at scale, part five: cost.
Five common MCP servers put 55,000 tokens of tool definitions in front of the model before a conversation starts, and a client that loads them all pays for those tokens again on every model call [1]. Across a fleet that can be a six-figure line every month, and nothing on the invoice, in the budget or in the protocol names it as MCP. The team paying for it usually cannot see it, and the team that controls it usually does not know it is a cost.
Here is how it plays out for one imagined fleet, and then why it happens.
Who spends, and who can see it
| Party | What it controls | What it can see |
|---|---|---|
| The budget owner | The AI line in the plan | An invoice "by API key or project, not by business unit, cost center, application, or team" [2] |
| The platform team | Which clients run, which gateway they use, which servers each role gets | Input tokens at the gateway, with tool definitions counted like any other input |
| The client | Which tool definitions go into each model call, and in what order | Nothing about price |
| The server author | How many tools a server has, and how long each description is | Nothing about who loads them, or how often |
| The provider | The price of every token | Input tokens, "including in the tools parameter" [3] |
Every party can see part of the cost. None of them sees the tool surface as a line of its own.
A quiet month
Imagine a company running one million agent tasks a month, with 20 model calls per task. Those volumes are my assumption, and you should swap in your own. Its engineers use Claude Code with five servers like the ones in Anthropic's example, GitHub, Slack, Sentry, Grafana and Splunk: "58 tools consuming approximately 55K tokens before the conversation even starts" [1].
Most of that never reaches the model. Claude Code defers MCP tools by default and discovers them on demand [4], so each call carries a tool search tool "(~500 tokens)" plus "3-5 relevant tools, ~3K tokens" [1]. At Opus 5 list prices, read on 6 October, with one cache write at the start of each task and cache reads after it, that surface costs about $55,000 a month [3]. It is real money, and it is lost inside a much larger inference bill.
The same fleet, one quarter later
Two reasonable changes land.
The platform team routes every client through its own LLM gateway, so it can
enforce budgets and keep logs in one place. It points ANTHROPIC_BASE_URL at the
gateway. Claude Code turns tool search off when that variable points to a
non-first-party host, "since most proxies don't forward tool_reference blocks"
[4]. From that day the full 55,000-token surface rides on every call.
Then security asks for least privilege. Under the 2026-07-28 revision a server's tool set can "vary by the authorization presented on the request", for example "returning only the tools the caller's granted scopes permit" [5]. The team turns it on. Each distinct permission set is now a distinct prompt prefix, with its own cache to warm.
The tasks are the same, the model is the same, and so is the number of calls. Assume caching behaves as it did before. On the arithmetic below, the surface line goes from about $55,000 a month to about $866,000, and the per-user tool lists push it higher by an amount nobody has published a way to estimate. The budget owner sees input tokens rise under the gateway's API key [2]. The gateway sees more input tokens and has no field that says which of them were tool definitions. Neither change was a mistake, and the fix for the first one is a single environment variable, which comes later.
Where the money goes
The arithmetic fits on one line, so you can redo it with your own numbers:
monthly cost = tasks per month × model calls per task × surface tokens × price per million ÷ 1,000,000
Prices are Anthropic's list prices, read on 6 October [3]. Opus 5 input is $5 per million tokens, a five-minute cache write is $6.25, and a cache read is $0.50. Opus 5.5 cache reads are $0.20.
| Tool surface per call | Where the count comes from | Opus 5, all cache reads | Opus 5, one write then 19 reads per task | Opus 5, uncached | Opus 5.5, all cache reads |
|---|---|---|---|---|---|
| ~3,500 tokens, deferred | Tool search "(~500 tokens)" plus "3-5 relevant tools, ~3K tokens" [1] | ~$35,000 | ~$55,000 | ~$350,000 | ~$14,000 |
| 55,000 tokens, all loaded | Anthropic's five-server example [1] | ~$550,000 | ~$866,000 | ~$5,500,000 | ~$220,000 |
| 134,000 tokens, all loaded | "tool definitions consume 134K tokens before optimization" [1] | ~$1,340,000 | ~$2,110,000 | ~$13,400,000 | ~$536,000 |
Worked through for the 55K row: 1,000,000 × 20 × 55,000 is 1.1 trillion tokens. At the $0.50 cache-read price that is $550,000; at the $5 uncached price, $5,500,000. The middle column charges one cache write at the start of each task and reads after it, which averages about $0.79 per million and lands near $866,000. That column is where the imagined fleet sat in both of its months.
Two cautions on the inputs. Anthropic published the counts in November 2025, and its pricing page now says Claude 4.7 and later models use a tokenizer that "produces approximately 30% more tokens for the same text" [3]. I did not re-count. The table also leaves out tool results, which are input on every later turn too. One transcript passed between two tools "could mean processing an additional 50,000 tokens" [6], and I found no fleet-scale measurement of results.
The one measurement of MCP's share against a no-MCP baseline comes from a vendor measuring its own server. Twilio ran the Cline agent with and without its MCP server, repeating each task at least ten times. Cost rose about 27.5% on average, with "~28.5% more cache reads and ~53.7% more cache writes", and tasks finished about 20.5% faster [7].
A cached prefix is cheap until it moves
Caching is what takes the 55K row from $5.5 million to $550,000, and it holds only while the prefix is identical. "Modifying tool definitions (names, descriptions, parameters) invalidates the entire cache" [8]. On a 55K surface on Opus 5, a write instead of a read costs about $0.32 more per call.
An upstream release is cheap. Each distinct surface in each workspace pays one new write, plus rewriting the cached history of conversations in flight. For a fleet with a few dozen distinct surfaces that is tens of dollars per release. On the Claude API, Claude Platform on AWS and Microsoft Foundry the cache is isolated per workspace, so a fleet split across workspaces warms each one separately [8].
The expensive case is a prefix that never settles, and MCP has two ways to
produce one. The first is tool order. The 2026-07-28 revision says servers
"SHOULD return tools from tools/list in a deterministic order" to keep caches
warm [9]. A server that shuffles on every connection breaks the cache on every
connection. The client can fix this on its own: it assembles the tools array,
so sorting by server and tool name gives a stable prefix whatever the servers
return.
The second is the per-caller tool list from the imagined quarter. A fleet that scopes tools per user can hold as many cached surfaces as it has permission combinations, which moves it from the all-reads column toward the middle one. Tighter access control costs cache hit rate, and nothing published measures by how much.
One recent change helps. With Anthropic's inline-tools-2026-09-15 beta, a client can
"add a tool, or change a tool's definition, partway through a conversation
without editing tools", after which "The cached prefix still matches" [8]. A
client that applies a list_changed notification that way keeps its cache.
Nothing between the API key and the server names the line
The closest thing to MCP attribution is in Claude Code's telemetry. Its cost
counter carries mcp_server.name, the "MCP server whose tool result this request
consumed" [10]. That attributes the cost of reading a tool's output to the
server that produced it. It does not split tool-definition tokens by server, which
is the line this part prices. Configured server names are replaced with
"custom" unless OTEL_LOG_TOOL_DETAILS=1 is set, which also exports tool
parameters and inputs [10]. Naming your internal servers in cost data means
exporting what was sent to them.
The protocol carries no cost and no payer. The 2026-07-28 tools page has no price,
cost or budget field [5]. _meta reserves traceparent, tracestate and
baggage for OpenTelemetry context [11], with no key for a user or cost center,
and baggage is asserted by the caller, so it can correlate but should not decide
who pays. The authenticated carrier is the token: OAuth token exchange, and the
Enterprise-Managed Authorization extension, which puts the organization's identity
provider in charge of server access [12]. I found no published convention that
joins the principal in that token to a cost record.
OpenTelemetry does not fill the gap yet. Its GenAI conventions define token counts and no cost attribute [13]. A 2024 proposal to add one was withdrawn after a maintainer asked, "How would instrumentation code know the cost of token?" [14]. Payment inside MCP has not moved either: SEP-2007 was closed in June as dormant [15], and a paid-tools issue built on x402 was closed as not planned on 28 September [16].
Gateways count a tool's own cost only if an operator types in its price. LiteLLM
takes a fixed default_cost_per_query per server with per-tool overrides, or a
hook that sets a cost per call [17]. Bifrost has a per-client tool_pricing
map, "cost per execution" [18]. Neither page says whether tool cost counts
against the same budget as inference.
Four levers, and what each one costs you
| Lever | Measured effect | Evidence |
|---|---|---|
| Tool search or deferred loading | Definition tokens from ~72K to ~3.5K, about 95% (my arithmetic) | Anthropic's own comparison [1]; default in Claude Code [4] |
| Code execution in place of direct tool calls | 77.4% fewer total tokens; output up 120%, latency up 7% | AIMultiple, independent, two tasks, GPT-4.1 [19] |
| Prompt caching | Reads at 0.1x base input on Opus 5; a five-minute write "pays off after one cache read" | Published price [3] |
| Deterministic tool order | Keeps the prefix stable | Specification SHOULD [9]; no measured effect found |
Deferred loading is the lever the imagined fleet lost, and it is the cheapest to
get back. Claude Code behind your own gateway keeps tool search off unless you set
ENABLE_TOOL_SEARCH [4]. On the Messages API, tools load in full unless marked
defer_loading [3]. The cost of deferring is that the model has to search for
a tool before it can call one.
Code execution's saving depends on what the removed tokens cost. Repricing AIMultiple's token counts at Opus 5 rates (my arithmetic), cost per run falls about 73% when input is uncached and about 35% when input is billed at the cache-read rate. The extra output tokens cost $25 per million either way. A fleet that already caches well gets the smaller number, and pays for it in latency.
Caching and ordering cost nothing to adopt, but they reward the opposite of what per-caller tool lists ask for. That trade between access control and cache hit rate is the one this part cannot price for you.
What to do, depending on who you are
If you own the AI budget, ask the platform team for one number: definition tokens per model call, per client. Put it and your own task volume into the formula above, and carry it as its own line. Until someone measures it, it is invisible inside "input tokens."
If you run the clients and the gateway, check whether every client and agent
runtime defers tools, including Claude Code behind your own gateway. Give each
role only the servers it needs. Sort the tools array in the client. Count how
many distinct permission sets your per-caller tool lists produce, because each is
a cache you pay to warm. Enter prices for the tools that wrap paid APIs, or they
cost zero in your spend data. Derive the payer from the authenticated token at
each hop and use baggage only to correlate.
If you write MCP servers, return tools/list in a deterministic order and
keep descriptions short. Every token you add to a description is billed on every
call of every client that loads your server.
What nobody has priced yet
Three lines on an MCP budget still have no independent figure. Building the
control plane has only vendor estimates from companies selling the alternative,
such as Zuplo's "2-4 engineers working for 3-6 months" for a production server
[20], with no sample or method. Reviewing servers has none at all: part two
counted 39,617 servers in the official registry on October 5, and I found
published effort figures for building servers and none for reviewing them. The
third is re-review, which a list_changed from an approved server may oblige,
and I found no figure for how often approved servers change.
The limit of this part is that its dollar figures rest on published token counts, list prices and an assumed volume. I have seen no fleet's invoice. My guess is that the review cost is larger than the cache cost of the same change, and nobody, including me, has measured it. If you run an MCP review process, count the hours from your first server and publish them.
What it adds up to
The largest MCP cost you can price today is the tool surface your clients resend on every model call. It is billed as ordinary input tokens under whoever's key made the call, so no invoice, budget or protocol field names it. It is also the cost you can cut fastest. Find out which clients load every definition, price the surface you carry, keep it short and stable, and put the number in front of whoever owns the AI budget.
Part five of five on operating MCP at scale. Parts one to four cover operational excellence, security, reliability and performance.
Sources
All URLs verified 2026-10-06. Vendor posts are marked.
- Anthropic, "Introducing advanced tool use on the Claude Developer Platform", 2025-11-24: 55K, 72K and 134K counts, tool search figures. https://www.anthropic.com/engineering/advanced-tool-use
- FinOps Foundation, "Tokenomics: Managing AI Value in SaaS Model Token Costs". https://www.finops.org/wg/token-economics-saas/
- Anthropic pricing: Opus 5 and Opus 5.5 rates, cache multipliers and break-even,
defer_loading, tokenizer note. https://platform.claude.com/docs/en/about-claude/pricing - Claude Code MCP documentation: tool search default and the
ANTHROPIC_BASE_URLfallback. https://code.claude.com/docs/en/mcp - MCP specification 2026-07-28, tools: per-authorization tool sets, no cost field, annotations untrusted. https://modelcontextprotocol.io/specification/2026-07-28/server/tools
- Anthropic, "Code execution with MCP", the 50,000-token transcript example. https://www.anthropic.com/engineering/code-execution-with-mcp
- Noah Mogil, "Performance Testing of Twilio Alpha's MCP Server", Twilio, 2025-04-10 (vendor). https://www.twilio.com/en-us/blog/developers/twilio-alpha-mcp-server-real-world-performance
- Anthropic prompt caching: tool-definition invalidation, workspace isolation,
inline-tools-2026-09-15. https://platform.claude.com/docs/en/build-with-claude/prompt-caching - MCP specification 2026-07-28 changelog, deterministic tool ordering. https://modelcontextprotocol.io/specification/2026-07-28/changelog
- Claude Code monitoring: cost counter MCP attributes, redaction. https://code.claude.com/docs/en/monitoring-usage
- MCP specification 2026-07-28,
_metaand OpenTelemetry trace context. https://modelcontextprotocol.io/specification/2026-07-28/basic/index - MCP Enterprise-Managed Authorization extension. https://modelcontextprotocol.io/extensions/auth/enterprise-managed-authorization
- OpenTelemetry GenAI attribute registry. https://github.com/open-telemetry/semantic-conventions/blob/main/docs/registry/attributes/gen-ai.md
- OpenTelemetry semantic conventions issue #1062, token cost. https://github.com/open-telemetry/semantic-conventions/issues/1062
- SEP-2007, payment support, closed as dormant 2026-06-24. https://github.com/modelcontextprotocol/modelcontextprotocol/pull/2007
- MCP issue #3393, paid MCP tools with x402, closed as not planned 2026-09-28. https://github.com/modelcontextprotocol/modelcontextprotocol/issues/3393
- LiteLLM MCP cost tracking. https://docs.litellm.ai/docs/mcp_cost
- Bifrost MCP client schema,
tool_pricing. https://github.com/maximhq/bifrost/blob/main/core/schemas/mcp.go - Şevval Alper, "Code Execution with MCP", AIMultiple. https://aimultiple.com/code-execution-with-mcp
- Zuplo, "Build vs Buy MCP Server Infrastructure", 2026-04-22 (vendor). https://zuplo.com/learning-center/build-vs-buy-mcp-server-infrastructure
Michael Rishi Forrester is the AI Workforce Transformation Lead at Accenture LearnVantage, where he works on Claude and generative AI adoption with governments and forward deployed engineers. With nearly 30 years in technical leadership, workforce transformation, and artificial intelligence, his work centers on enterprise AI governance and adoption and on closing the gap between what executives expect from AI and what AI can actually deliver.