I have been building tools in the weeds (like SocketBooks) and leading teams at Buddy Payroll for years now and I've gotten to the point where I think I am obsessed with mental models used to solve a problem. When it comes to software or protocol, we can be seduced to use a piece of software that we are familiar with or heard about, but may not be the most suitable for the architectural requirements of the task - I like to call it “Solution Bias”.
With AI agents, it's now a clash between two formats: Model Context Protocol (MCP) and Command Line Interface (CLI). But, as I think about recent Scalekit benchmarks and so much detective work that Webmaster Ramos has done, I've realised we're asking the wrong question. There's not two tools to choose from, there's two currencies to choose from.
1. You Don't Choose a Tool, You Choose a Currency
Agentic architecture is about balancing time to engineer and number of input tokens per run.
When you pay to "hire" an MCP server, that's as much as you can pay, since no humans are involved. It's a low QPS (queries per second) but high velocity play. In minutes, you can start an official server for MCP; your agent will know how to manage dozens of services. It's just for a prototype that will go on to hit five different APIs by tomorrow morning.
On the other hand, when you're developing for scale, with a high quality of queries per second (QPS), agent runs a million times a month - that convenience becomes a large recurring tax. One front-end cost of purchasing time to define whitelists, rich descriptions and runtime context in a purpose built CLI wrapper or native SDK is that the effort is required up front. However, it will save continuous token costs. As Webmaster Ramos aptly stated: The title of the question is wrong: ‘MCP or CLI' would imply it's the same purpose, but in reality it's a choice between two currencies.
2. The hidden costs of the "Schema Dump"
Being a UX certified developer, I've been thinking quite a bit about the cognitive load of the human, but also the models we're trying to guide. The biggest issue with the MCP performance is the "schema dump".
The benchmarks indicate the GitHub Copilot MCP server is a ‘fat' server that displays descriptions for 93 different tools in the context window when the server is launched. This is ~55,000 tokens before the agent reads the user's prompt! When a model is given a huge encyclopedia, and it only needs to know how to open one door.
They conceived another method than AWS, and constructed a thin layer of abstraction over their CLI, but they do have token overhang. Overall, the data from Ramos shows that the optimised CLI is 43% to 60% faster than the official AWS MCP server for tasks when performing the input tokens operation. Why? A target execution layer just conveys to the model exactly what it's supposed to know for the very specific task it is meant to do.
3. Reliability is a Local Game
I've never seen a payroll system that is 99% reliable. You need 100%. The transport layer is important when viewing agent performance.
CLI (Local Execution): 100% success rate in controlled benchmarks. You are running in a local subprocess, avoiding the "middleman". There are no timeouts or handshakes for remote servers with TCP. In the case of MCP (Network-Dependent), the performance of the benchmark tests is approximately 72% (most of the time with exception 'ConnectTimeout').
If an agent uses a protocol mediated access layer you are creating a network dependency in a reasoning loop. With remote protocol layers, especially for a developer who is running either a local build or deployment pipeline, a zero network dependency guarantee can be difficult to achieve.
4. This is the Inner and Outer Loop Hybrid.
Let's switch from either/or to Inner vs. Outer Loop. This is similar to the unix philosophy of doing one thing at a time but being able to orchestrate at a high level.
The Outer Loop (MCP): It's where you can implement higher-level orchestration, multi-tenant SaaS environments, cross-service integration and interactive tool discovery. An MCP service allows the ability to create, validate and evolve structured data from the conversational interface while the model is being developed or authored in active status. If there is structure/content, use a native SDK or a CLI wrapper (Inner Loop). Using MCP adds an unneeded protocol handshake, and brings in the token bloat for serving content, running local builds, or for automated deployment pipelines. The CLI and SDK provide execution with no context bloat and low latency and deterministic reliability.
In creative workflows, together they are the fluid orchestrator, authoring tool, and the high speed distribution engine during production.
5. Closing the "Context Gap" and the Batching Secret
Of all the most enlightening aspects of the Ramos benchmark, the one most clearly stated was Hypothesis 5: That it wasn't smart MCP was, it was just that it had a better watch.
The AWS MCP server sent the Date header of the HTTP protocol to the model.The Date header was leaked from the AWS MCP server to the model in the HTTP protocol. Otherwise an LLM might try to run time-dependent commands using an old time stamp. Model's clock is synced in the system prompt and a CLI's reasoning parity back to MCP.
Also, if you're looking to optimise your execution layer, consider using “Effect B” – “Batch Calling”. In MCP servers, models are able to send an array of commands in one turn. In a CLI or custom wrapper you can do this in your node or python using asyncio.gather where the agent can send out multiple commands at once and get them back all at once.
Runtime context (provided by the runner, not by the tool):
- Current UTC time: [Dynamic ISO Timestamp]
- Default region: [e.g., us-east-1]
- Identity: [User ARN/Role]
- This account is real and live; commands return real data.
By adding dynamic context and supporting array schemas for batching, you achieve up to 60% input token savings while keeping the agent just as "smart" as an MCP-integrated one.
Conclusion: Hiring the Right Tool for the Job
At the end of the day, SDKs, the command-line and protocol servers are all in different levels of abstraction.
The MCP protocol excels in providing standardised governance, multi-tenant orchestration and flexible tool discovery in interactive workflows, with its “Outer Loop”. CLI and native SDKs are king of the "Inner Loop" (prior to NATR), offering low latency, low token use, and reliability in the execution of known tasks at scale.
In designing your next agentic system, try to ignore the hype cycle. Question yourself: What is my QPS/Engineering Labor overwhelm point?
When shipping a production level agent, don't be afraid to prevent Solution Bias from you building the lean execution wrapper your unit economics deserve. Invest an hour or so of engineering effort in optimising the inner loop – your token budget, latency measures, and users will thank you.
