How does MCP tool search work in Elaichi?
MCP tool search in Elaichi picks a tool in two steps. First it scores every connected tool the caller can reach by matching the words of the query against the tool's name, its description and its app's label. Then it refuses any tool that covers less than half of what the query asked for, counting rare words for more than common ones. That second step is the relevance floor. It is why a search for one app's tool comes back empty instead of returning a confident match from another app.
Search carries this much weight because it is the only way in. In Elaichi, connected tools are never listed one by one, however few there are. The model sends a query to search_tools, reads back a short ranked list and runs its choice through execute_tool. The pool it searches gets large fast. The Asana and Xero connector pages each list well over a hundred tools, so a member with a few accounts connected can be choosing among hundreds.
Keep two things apart. The MCP specification defines how a client lists a server's tools with tools/list, which may be paginated, and calls one with tools/call. Its tools page defines no search and no ranking. In Elaichi, search_tools and execute_tool are ordinary entries in that list, a choice made in the application layer rather than a protocol feature. Another server could list everything, paginate, or rank differently. Anthropic offers developers a version of the same idea for Claude, a tool search tool that searches a catalog instead of loading every definition up front.
Two things stay fixed around the search. Control-plane operations stay listed individually, and search_tools never returns one, because it ranks only the connected third-party tools. The other fixed point is that execute_tool is purely a naming indirection. It resolves to the same tool name and arguments and passes the same authorization gates as a direct tools/call, with no separate path and no extra privilege. So search decides only which connected tools the model gets to see. Whatever it returns is the model's whole picture of what those accounts can do.
What went wrong when a Notion tool answered a Cal.com question?
A user with only Notion connected called search_tools with list_all_cal_com_schedules, a guess at a Cal.com tool name. The top result was list_all_notion_users.
Ranking is purely lexical, which means it matches words, not meaning. It reads three fields: the tool name, the description and the connector label. The weights are fixed. An exact name-token match scores 5, a name-token prefix in either direction 3, a connector-label match 2, and a description-token match 1. In that query, list and all scored 5 each against list_all_notion_users. The tokens cal, com and schedules scored nothing, because nothing in a Notion-only tool set contains them. The generic half of the query carried the result, and the specific half was silently discarded.
A ranking with no floor always returns something. It sorts a list and hands back the top of it. When the query names an app the user has not connected, the top of the list is a tool from a different app with a plausible verb in its name. That is worse than an empty result, because a model does not treat a weak match as weak. It calls it.
How does search decide a result is too weak to return?
It compares what a tool matched with everything the query asked for, weighting rare words more. A tool must account for at least half of the query's own weighted mass to be returned at all. That rule is the relevance floor.
The weighting is inverse document frequency (IDF), a standard measure in information retrieval. Stanford's Introduction to Information Retrieval explains why it exists. Some terms "have little or no discriminating power", so a rare term gets a high IDF and a frequent term usually a low one. The book's example is a collection about the auto industry, where "auto" appears in almost every document. In tool names, list plays that part.
Elaichi measures rarity over the tools the caller can actually see, which is the pool left after restrictions, not the full catalog. The word list appears in nearly every tool name in any pool, so its weight is low. The tokens schedules, cal and com appear in almost nothing a Notion-only user can reach, so their weight is high. The weights are computed per query against the current pool rather than stored.
Run the incident through it. The tool list_all_notion_users covers only the two cheapest tokens and none of the expensive ones. That is well under half of the query's weight, so it never clears the floor. The caller gets an empty result, and the model reports that no matching tool exists. That is the correct answer: the user has no Cal.com connection, so there is no schedule to list.
Half is the line where a candidate has matched at least as much of the query's informative content as it missed. Below that line, no raw score makes it an answer to the question.
Why measure a match against the query instead of a fixed score?
Because a fixed cutoff stops working as queries and catalogs change, while a share of the query's own weight stays fair for every query. The table compares three ways to add a floor.
| Approach | How it works | Where it fails or holds |
|---|---|---|
| Absolute score cutoff (for example, "return nothing scoring under 6") | A fixed threshold on the raw weighted score | Scores are not comparable across queries. A six-token query gathers more raw score than a one-token query, and a query that names the connector adds 2 to every candidate from that connector. Tune for long queries and short ones return empty; tune for short ones and long ones pass junk. The cutoff also drifts as the pool grows with each new connector. |
| Embedding or cosine-similarity rerank | Vector similarity over tool name and description embeddings | It still always returns a ranked list, so a top-k over embeddings has the same no-floor failure with better-sounding scores. It is also weak on this exact case: "list records from a SaaS app" sits close in embedding space whether the app is Notion or Cal.com. It adds index upkeep on every connector edit, fork or new connection, latency inside the tool-call budget, and results that are harder to explain afterward. |
| Weighted-share floor (relative, per query) | A candidate must cover at least half of the query's own IDF-weighted token mass | It normalizes by construction: the same rule applies to a one-word search and a six-token guessed tool name. It holds after the tenth connector is added, because both the candidate's coverage and the query's total are recomputed against the live pool on every call. |
A reranker, however good, is an upgrade to the scoring underneath. It does not remove the need for a floor, because any function that puts candidates in order will hand back a top result even when every candidate is irrelevant. The floor is a separate mechanism: a threshold on whether to return anything at all, checked after scoring.
Other MCP servers may reasonably choose a threshold other than half, combine a lexical floor with a semantic fallback, or show a confidence score instead of cutting to empty. The claim here is narrower than calling this the only correct design. Whatever the scoring, a floor has to be relative to the query's own weight, not a fixed score. Otherwise it breaks under the two-sided pressure of short exact-name queries and long guessed ones sharing one endpoint.
Which two other ranking bugs did Elaichi fix?
Two, and both share the floor's root cause: tokens that look like signal and are not. The floor is one of three corrections in Elaichi's ranker.
Description stop-words. The words set, connection, frozen and more are left out of description scoring, because they appear in the description of every merged or frozen tool, whatever the tool does. A merged tool is one tool that reaches several accounts of the same app. Information retrieval calls such words stop words: words so common they are of little value in selecting a match. Before the fix, a query containing "connection" gave every merged tool the same description score. Ties broke on name order, so the model was handed an arbitrary account's tool. Account labels and frozen field values stay scorable on purpose, because a caller might legitimately search for them.
Tie-break by codepoint, never locale collation. When two tools score the same, Elaichi orders them by raw character codes, not by locale-aware sorting. Advertised tool names use exactly the characters that locale-aware comparison reorders, and the ranked list is cut to a limit before the model sees it. A locale-dependent tie-break would therefore decide which tools the model sees at all, depending on the server's locale. Codepoint order is identical everywhere the ranker runs. The MCP specification asks for the same property in listings: servers "SHOULD return tools in a deterministic order".
What does refusing weak matches cost you?
Some recall. In information retrieval, precision is the share of returned results that are relevant, and recall is the share of relevant results that get returned. The same text notes that the two "clearly trade off against one another". Lexical matching has no synonyms. Search "calendar" against a connector whose tools all say "schedule", and the only path to a match runs through the description or the connector label, if either contains the word. With a floor, that near-miss stops being a weak result and becomes no result. Genuine synonym queries will sometimes come back empty where a synonym-aware system would have found the tool.
The two failures do not cost the same. An empty result is visible: the caller can rephrase, and the model reports it found nothing. A wrong result stays invisible until it has run against a live third-party account, and the failure shows only after the side effect. The floor gives up some recall to cut silent wrong-tool runs. The case for it is strongest where tools write: a wrong read wastes a call, while a wrong write changes a live record.
How do restrictions affect tool search?
They shrink the pool before ranking sees it, which makes search more accurate as well as safer. A restricted tool is withheld from the tool list and cannot be called, so it never competes with the tools the caller may run. Search ranks it separately and names it, flagged restricted, with no schema. Where the client shows Elaichi's search results as a card, the card says how many matches are restricted for the person and names up to five of them. Elaichi's resolver is the code that decides allow or deny. It runs at every place a member can reach a connector or tool, including listing, connecting, advertising, executing, the final outbound-request check and opening a stored tool file.
The practical effect shows in two members. One, whose role reaches two connectors, searches a small, clean pool. IDF weights there discriminate sharply, few candidates share tokens, and the floor rarely settles a close call. The other, with everything connected, searches a pool where many tools compete on generic verbs like list, get and create. The weights flatten, and the floor does more of the work. Restrictions are a security control first, and mechanically they are a precision control on search too.
One timing detail: when you change a restriction, search results reflect it within about two minutes. Grant revocation, member removal and suspension are faster. In Elaichi, removing or suspending a member revokes every live grant in the same transaction as the membership change.
When does a small server not need tool search at all?
When it has a handful of tools and lists them all. If you build your own MCP server with eight tools and return every one from tools/list, the model reads the whole list on every request. No ranker sits in the path, and a relevance floor solves a problem you do not have. The weighting and the stop-word list are irrelevant to a small, fixed, hand-written tool list. Stay there as long as the tool count lets you.
The floor starts to matter when two conditions hold together:
- The server serves tools through search rather than a full list, so a ranker is in the path at all. Elaichi does this for every connected tool.
- The caller is a model composing a tool name from memory or inference, not a person choosing from a visible, complete list.
Both conditions hold the moment a company connects its first SaaS account behind one endpoint rather than a server per member for each client. That is why the floor is an endpoint-level guarantee and not a per-connector setting.
For what Elaichi's endpoint serves once several accounts sit behind it, see the product overview and a single connector's tool surface such as Zendesk. For the team-by-team view of who searches for what, start with the use-cases index.