Skip to content

The MCP context window problem, and a fix

The MCP context window problem: each listed tool sits in the model's prompt on every turn. Elaichi lists no connected tools; the model searches for them.

Uday Gajavalli Updated 7 min read
Diagram of an AI client tool list collapsing from many entries into two meta-tools

Why does a long tool list fill the MCP context window?

Every tool an MCP server lists takes room in the model's context window before the user types a word. That is the MCP context window problem, and it grows with every account a company connects. MCP (Model Context Protocol) is the standard way an AI assistant calls tools in other apps. When a client connects, it asks the server for its tools with tools/list. The MCP specification gives each entry a name, a description and an inputSchema, a JSON Schema for every argument the tool accepts. The client hands those definitions to the model. OpenAI's function-calling guide is plain about the cost: tool definitions "count against the model's context limit and are billed as input tokens".

The numbers get large quickly. Anthropic's documentation puts a typical setup of five MCP servers at about 55,000 tokens of tool definitions "before Claude does any work".

Two costs follow, and they differ in kind. The first is budget: context spent on tool definitions is context not spent on the ticket, the thread or the contract under review. The second is selection. Anthropic names tool selection accuracy as the second problem that grows with a tool library, and OpenAI advises keeping the number of initially available functions small "for higher accuracy". Duplicates make it worse. The specification warns that a client aggregating several servers may "encounter naming collisions" and suggests prefixing tool names with a server identifier. A prefix tells two Notion workspaces apart by label. Nothing in it says which one holds the customer contracts.

Why does Elaichi never list connected tools, however few?

Because the list would grow with every account a company connects, and one endpoint for a whole company connects a lot of them. Elaichi serves a single organization-wide endpoint, POST /mcp, with two kinds of tool behind it. Control-plane operations are named elaichi__{resource}__{operation} and manage connections, members, restrictions and the audit trail. A member's connected third-party tools are named {account}__{tool}.

In Elaichi, connected tools are never listed one by one, however few there are. There is no threshold where listing switches on. The model looks a tool up with search_tools and calls what it found through execute_tool, and those two are the only way in. The list the model reads therefore does not grow with the number of connected accounts, whether a company connects three of the 600+ connectors or forty.

What does the model see in the tool list?

Two meta-tools, search_tools and execute_tool, plus the control-plane operations, which stay listed individually. A list holding nothing but elaichi__ operations does not mean nothing is connected. Search covers connected tools only and never returns a control-plane operation.

Administration stays in plain view because it is a fixed set that does not grow with connections. A search step in front of it would add a round trip to the most common requests.

The working loop changes shape. Instead of scanning a flat list, the model writes a query, reads back a short ranked set and calls one tool. It pays the schema cost only for what the search returned.

Results share the same window, so Elaichi caps them too: 32,000 characters per string, 200 items per array, 200 keys per object and a nesting depth of 8. A cut sets elaichi_truncated, so the model knows the answer was trimmed rather than complete.

How does search rank connected tools?

By matching words, not meaning. Ranking is purely lexical over three fields: the tool name, the description and the connector label. An exact name match scores 5, a name prefix in either direction 3, the connector label 2 and the description 1.

A relevance floor sits on top. It is a minimum share of the query that a tool must match before search returns it at all. A tool must account for at least half of the query's own weight, measured with inverse document frequency (IDF), which counts rare words for more than common ones. Without a floor, a search always returns something, and a tool from an app the user never mentioned is worse than nothing, because the model calls it. The Cal.com search that returned a Notion tool is the failure that produced the rule.

Two smaller rules matter too. The words set, connection, frozen and more are dropped from description scoring, because they sit on every merged or frozen tool. Before that, a query containing "connection" scored every merged tool the same and handed the model an arbitrary account. Ties break on codepoint order, never locale collation, so a server's language settings cannot decide which tools make the cut.

Account labels stay scorable on purpose. Name two Notion accounts notion-legal and notion-marketing, and a query mentioning legal ranks the right one first. Name them notion and notion-2, and the ambiguity has moved out of the model and into your naming.

Does the search step open a second path into your systems?

No. The execute_tool wrapper is only a naming indirection. It unwraps to the same tool name and the same arguments and falls through the identical gates. There is no separate execution path and no privilege in the wrapper.

The tool:execute permission gates the whole endpoint ahead of every scope. Without it, tools/list comes back empty and a call returns an in-band error naming the permission. Guest, Auditor and Billing Admin do not hold it. OAuth, the sign-in standard that issues a scoped grant instead of a shared password, adds a second ladder: mcp:read, mcp:write, mcp:destructive and mcp:tools. A tool classified forbidden is reachable under no scope, and a connected tool whose method is a delete still needs mcp:destructive. Restrictions are then checked by one resolver at every place a member can reach a connector or tool, including listing, connecting, advertising, executing, the final outbound-request check and opening a stored tool file.

One limit, stated plainly. The prompt-injection write gate lives in the Elaichi agent window and does not apply to a raw tool call. On the endpoint itself, what holds is role-based access control (RBAC) per operation, the forbidden classification, output redaction, OAuth scope limits and audit logging of each call that reaches execution.

Can restrictions keep tools out of the context window?

Yes. A restriction decides which connectors and which individual tools a target may reach, and a tool it withholds stays out of the tool list and cannot be called. Search names it, flagged restricted, with no description or schema, so it spends a name and not a schema.

In Elaichi, restriction targets are role or user only; the organization default is the absence of a rule, which means allow everything. A rule aimed at a member can only narrow what their role allows, and never replaces or loosens a role rule. Within each layer, allow rules combine, block rules combine, and a block always beats an allow. One trap: the allowlist stage engages on the presence of an allow rule, not its contents, so an allow rule that names nothing denies everything. Blocks and allows also match a tool differently, which matters before your first allowlist.

Frozen parameters trim the schema rather than the list. A frozen key is removed from the schema the model reads, so it costs no context. Its value is merged over whatever the caller sends, so passing the key cannot un-freeze it.

A restriction or role change takes effect within about two minutes. Grant revocation, member removal and suspension take effect on the next call.

What does the audit trail record for a searched call?

The execution, not the wrapper. Elaichi keeps one entry for each connected-tool call that reaches execution (a call refused earlier writes none). Each entry records the account the call actually reached, taken from the execution rather than the intent, so a reviewer never has to unwrap execute_tool to see what ran.

That answers the first question after an unexpected change, which is which of two Notion workspaces the agent wrote to. The entry also carries the operation, its classification, whether it was approved, how it ended and, for a failure, an error code only. It logs the one path argument that names the object, as the target id, and nothing else about the arguments.

Discovery becomes a search problem, and search can miss. Ranking is lexical, so a query phrased in words that appear in none of the tool name, the description or the connector label will not match. The relevance floor makes the honest outcome an empty result rather than a confident wrong one. Empty is the better failure, and the user still notices it.

There is also a round trip. Search, then execute, is two calls where a flat list needed one. A model working through a long task pays that latency each time it reaches for a new kind of action.

When a tool seems to be missing, Elaichi gives the model a fixed order of checks:

  1. Does the connection exist?
  2. Is its status active? If not, reconnect it.
  3. Is the tool in the connection's tool list? If not, check whether a restriction covers the connector.
  4. Is mcp:tools on the OAuth grant?
  5. Connected tools are never in the list, so call search_tools for the name.

Most misses are permission gaps rather than typos. The specification lets the tool set "vary by the authorization presented on the request", and Elaichi filters what a caller can see by role and by the grant's scopes. A connection that is pending or needs_reauth contributes no tools at all. It is absent, not present and failing.

When is a flat tool list still the right answer?

When the server is small and has one job. A single-purpose MCP server with six tools and one user can list everything, and search would add a round trip that buys nothing. Anthropic's guidance draws the same line. It calls standard tool calling without search "a better fit when you have fewer than 10 tools", and says the same when every tool is used in every request. If that is your situation, wait. The case against adding a gateway yet makes the argument.

Search starts to earn its round trip once a second team, a second account or a second client shows up. At that point the question stops being how many tools the model can hold. It becomes which tools this person should have been offered at all. Search does not answer that. Roles and restrictions do.

To see the shape of a governed endpoint, read how the control plane fits together, browse the connector catalog, or look at what each team does with it. More on the protocol itself sits in How MCP works.

FAQ

Frequently asked questions

Why do MCP tools use up the context window?

Because an AI client passes every listed tool's name, description and argument schema to the model as part of the prompt, and that text counts as input on every request. A long list spends context before any work starts, and it gives the model more near-identical names to choose between. Searching for a tool instead of listing every one keeps the prompt small however many apps are connected.

Does Elaichi ever list connected tools one by one?

No. In Elaichi, connected tools are never listed one by one, however few there are. There is no threshold at which listing switches on. The model looks a tool up with search_tools and calls it through execute_tool. Control-plane operations are the exception: they stay listed individually, and search_tools never returns one.

Does execute_tool bypass permissions, scopes or restrictions?

No. In Elaichi, execute_tool is only a naming indirection. It unwraps to the same tool name and the same arguments and passes the same gates: the tool:execute permission, the OAuth scopes, the forbidden classification, and restriction checks at every place a member can reach a connector or tool, including listing, connecting, advertising, executing, the final outbound-request check and opening a stored tool file. There is no separate execution path and no privilege in the wrapper.

Do restricted or frozen tools still use context window space?

No. A tool withheld by a restriction in Elaichi stays out of the tool list, so it costs no schema. Search names it, flagged restricted, with no description or schema. Frozen parameters shrink the schema instead: a frozen key is removed from the schema the model reads, and its value is merged over the caller's arguments at execution, so passing the key cannot un-freeze it.

What happens when tool search cannot find a tool?

Elaichi returns an empty result rather than a weak match from another app, and gives the model a fixed order of checks: whether the connection exists, whether it is active, whether the tool is in the connection's tool list or restricted, and whether mcp:tools is on the OAuth grant. A missing tool is more often a permission gap than a typo, because what a caller can see is filtered by role and by the grant's scopes.

Put agents to work on your own systems

14 days on Gold, no credit card. Start with one app and one team.

Works with
Claude ChatGPT Cursor and any other MCP client, or the Elaichi Agent.
When the trial ends
Nothing is deleted. Connections, roles and the audit log stay where they are, so subscribing picks up exactly where you left off.