Claude CodeMcpAi ToolingDeveloper ExperienceTokens

We Replaced 7 MCPs With curl and Saved 99% of the Context

MCP servers dump every tool's schema into Claude's context at session start, whether you use them or not. Seven of ours cost 240k tokens. We replaced them with lightweight connectors that keep a one-liner each and load the full instructions only when called. Total: 1,500 tokens.

Pratham BhatiaPratham BhatiaSDE-1
Oct 3, 2026
6 min read
We Replaced 7 MCPs With curl and Saved 99% of the Context
Fig. 01 — A dispatch on claude code
Share this article

Most Claude Code setups connect to external tools through MCP servers. Slack, Vercel, Figma, ClickUp. Each one gets its own server, and Claude can call any of its tools at any time. It works.

The problem is what it costs before you type anything.

What an MCP server actually puts in your context

MCP uses JSON-RPC. When Claude Code starts, it calls tools/list on every connected server. The server replies with every tool it offers: the name, a description, and the full JSON Schema for every input parameter. Claude needs all of that to know what it can call and how.

How an MCP server loads into Claude's context

Vercel alone has over 200 tools. Each tool carries a schema with nested objects, enums, and descriptions. That single server puts roughly 94,000 tokens into the context. Slack adds 54,000. Figma adds 43,000. ClickUp, Fathom, AWS, Chrome DevTools: by the time all seven are connected, the context starts at about 240,000 tokens.

That is 24% of the 1M context window, used up before you ask a single question.

It compounds from there. This block is re-sent on every turn. Ask Claude to fix a typo in a README, and those 240,000 tokens ride along with the request. They are part of every API call for the rest of the session, billed as input tokens every time.

What we actually use per session

We tracked it for a month across our team. A typical session touches one or two of these tools. A Vercel deploy check. A Slack message. A Figma frame export. Nobody uses all seven in one session, yet all seven load every time.

  • Loaded every session. ~240,000 tokens of tool schemas.
  • Actually used per session. One or two tools, maybe 3,000 to 5,000 tokens of relevant instructions.

The replacement: connectors that load on demand

We wrote a lightweight connector for each app. Each connector keeps only a one-line description in context at session start. The full instructions (auth setup, API endpoints, curl helpers, gotchas) load only when Claude actually needs that app.

The connectors call the same REST APIs the MCP servers called internally, written as curl instead of routed through JSON-RPC.

How a connector loads into Claude's context

How a connector works

A connector is a single file of instructions. It holds the curl commands for one app's REST API: the auth header, a small helper function, and a table of the endpoints we use, plus the quirks we hit along the way. When a task needs that app, Claude reads the file and runs the curl commands from it.

There is no server to run and no schema to register. The file is the whole connector, and it is only read when the task calls for it.

The numbers

Token cost per session start: MCP vs connector

AppMCP tokensConnector (always)Connector (on use)
Vercel93,700~250~4,000
Slack54,200~140~5,000
Figma43,300~260~3,000
ClickUp27,500~240~4,000
Fathom~8,000~250~3,500
AWS~7,500~230~4,500
Chrome DevTools~5,600~180~3,000
Total~239,800~1,5003,000–5,000 per app

The always-on context dropped from ~240,000 to ~1,500 tokens. That is 184 times smaller, or 0.15% of the context window instead of 24%.

When you actually use an app, its full instructions load for that session. A session that checks a Vercel deploy and reads a Slack channel uses about 1,500 + 4,000 + 5,000 = 10,500 tokens. Still 23 times less than the MCP baseline of 240,000.

How auth works without an MCP server

Each MCP server handles its own authentication. Without them, the connectors need to do it themselves. We store every token in the macOS Keychain and retrieve it at call time.

  • Vercel, Figma, Fathom, ClickUp. Claude opens the token page in the browser, asks you to paste the personal access token, tests it with a curl call, and saves it to the Keychain only if it works.
  • Slack. A built-in script reads your existing signed-in session from Chrome or the Slack desktop app. It extracts the session token, checks it, and saves it. You never copy a token. macOS may ask to allow Keychain access.
  • AWS. Uses your existing aws CLI profiles and SSO login. No token to save.

First-run setup happens automatically the first time you use each connector. After that, it works until a token expires.

How to install

The plugin is open source. Three commands inside Claude Code:

/plugin marketplace add prathambhatia/custom-connectors
/plugin install custom-connectors@custom-connectors
/reload-plugins

Connectors show as custom-connectors:slack, custom-connectors:vercel, custom-connectors:figma, and so on. Each one replaces the corresponding MCP.

Source: github.com/prathambhatia/custom-connectors

The real cost is the context you don't see

The token numbers from MCP servers are easy to miss. Nobody shows you a "240,000 tokens loaded" banner at session start. The tools work, so there is no error to investigate. The cost shows up as a faster context fill, earlier compactions, and a higher bill at the end of the month.

Once we measured it, the fix took a weekend. Same APIs, same auth, same results, loaded when needed instead of always.

Pratham Bhatia

Pratham Bhatia

SDE-1

Continue reading

All articles →
Jev 1.13: Fastest, Cheapest, Twelfth
model evaluation

Jev 1.13: Fastest, Cheapest, Twelfth

We tested Jev 1.13 against 10 other models on two real jobs. It was the cheapest and the fastest. On accuracy it came 12th out of the 19 runs.

Rachit SrivastavaSep 20 · 13 min