DevOps Engineer
8
8Most articles about the Model Context Protocol (MCP) open the same way: an AI agent needs external tools, and MCP provides the connection.
But what happens if we remove the AI entirely?
We built a cloud migration cost-estimation pipeline that connects to three MCP servers, makes hundreds of tool calls, and produces a priced, citable output. For its first several months, it contained no AI code at all — no model, no agent, no chat window, no inference bill.
Strip away the AI terminology and MCP is three things.
Every MCP message is a JSON object with fields such as jsonrpc, method, params, and id, where the id lets the client match responses to requests. No proprietary binary encoding, no gRPC schema compiler, no OpenAPI generator — just the well-established JSON-RPC 2.0 convention.

The protocol is identical either way. Only the transport changes.
Tools dominate in practice, because they are what let a client invoke functionality on the server.

The client opens the session and agrees to a protocol version, asks what tools exist, then invokes one. If you have ever written an HTTP or JSON-RPC client, none of this is new.
One part of MCP clearly reflects its origins: a tools/call response returns a list of content blocks.

The result is text because the design assumes a model will eventually read it. Without an LLM or a sustainable LLMOps practice, the adaptation is trivial — pull out the text and parse it:

That's the whole adaptation. The protocol isn't AI-specific; only its result format is AI-flavoured.
The problem was cloud migration cost estimation: take a customer's Azure cloud inventory — potentially thousands of resources spanning VMs, storage, managed databases, instance sizes, operating systems, and service tiers — and produce a realistic AWS cost estimate.
That needed two kinds of information:
What is the AWS equivalent of this source service, and can we cite a source for that mapping? Guessing isn't an option when a customer's architect reviews the result.
Which configuration values are valid for that AWS service today?
Hardcoding them creates a maintenance problem, because AWS continuously adds instance families, changes service configurations, and retires older options.
Our first implementation held these mappings in Python dictionaries:
It worked, but it went stale almost immediately: every change in the provider's available configurations could mean a code change. We needed something better.
We found MCP servers offering exactly the information we needed, and connected the pipeline to three:

Three completely different implementations — a hosted HTTP service, an npm package, a locally bundled server — behind one interface. That uniformity turned out to be one of the biggest benefits.
There is no framework here. A stdio MCP client is a subprocess handle, a JSON encoder, and a dictionary mapping request IDs to responses.
Starting the server. The client launches the server as a child process with stdin and stdout as pipes, and sends the server's logs to stderr so they can't contaminate the protocol stream. From then on, stdout is a wire, not console output.
The handshake. The client sends an initialize request naming the protocol version it intends to speak, its capabilities, and who it claims to be; the server replies with its own. A one-way notifications/initialized message confirms the session is live. After two messages, the connection is usable.
Calling a tool. Every call after that is the same shape — a tool name and an arguments object:

_rpc writes one line of JSON to stdin, reads stdout until it finds the line whose id matches, and returns it. That ID matching is the only real bookkeeping in the client, and it's what makes the thing safe to reason about: a response is never assumed from ordering; it is claimed by identity.
Making it readable. Everything else is thin naming. search_services, get_service_fields, validate_estimate, and export_estimate each pass a tool name and a small dictionary to call, so the pipeline reads in the vocabulary of cost estimation rather than JSON-RPC — and a tool rename is a one-line change in one place.
The whole client was roughly 150 lines of Python for all three servers: no agent framework, no MCP host application, no model runtime, no inference layer.
The AWS Pricing Calculator was the biggest surprise. There's no traditional public pricing API providing everything we need from the calculator — but there is an MCP server that exposes it.
That changes the equation. When a vendor ships an MCP server before a clean public API, MCP stops being an AI integration choice and becomes the integration interface itself.
tools/list returns the available tools with their input schemas. Better still, the Pricing Calculator server exposes get_service_fields, which returns the input fields and valid options for a service — instance types, database editions, deployment options. Instead of maintaining our own list of valid AWS configurations, we ask the calculator:
The candidate list comes from the vendor's own data. A new instance family gets picked up on the next catalog refresh with no code change; a retired one stops being selected. Our hardcoded table becomes a fallback, not the source of truth.

All three support the same operations — initialize, tools/list, tools/call — so the client needs to know nothing about their implementations. Adding a fourth is cheap.
This is probably the most important reason we chose this architecture.
Same inventory in, same estimate out. Not "approximately the same." Exactly the same. Three properties follow from that.
We cache resolutions by resource signature: type, SKU, tier, OS. An inventory of thousands of resources usually holds only a few dozen unique configurations, so it collapses to a small set of signatures. After the first run, an unchanged inventory can mean almost zero MCP calls — good for cost, speed, and reliability.
If a server is unavailable, we get a clear error such as TimeoutError — far safer than a plausible-looking but wrong instance type:
rather than:

The pipeline doesn't depend on which model version, IDE or host application was involved, or on how an agent interpreted the task. The runbook drives the MCP servers directly.
The architecture was straightforward; the details below were not obvious.
For stdio servers, stdout must carry only JSON-RPC messages — a stray Loaded 350 services corrupts the stream. We sent server logs to stderr and treated stdout as the protocol channel:

Concurrent requests over a single subprocess interfere with each other, so we lock around the write/read sequence and correlate strictly on the JSON-RPC id. Without it, two threads interleave their requests, and both wait indefinitely.
We send the protocol version explicitly at initialization rather than accepting whatever the server prefers.
npx -y package@latest is convenient, but the server can change underneath your application, and it needs npm registry access at runtime — a problem inside a locked-down container. For production, know exactly which dependencies are downloaded and when.
Some endpoints publish AAAA records. With no functional IPv6 route, connections sit in SYN-SENT until they time out before IPv4 is attempted. We forced IPv4 in the affected runtimes.
A response may arrive as content → text → JSON string rather than content → JSON object, so the client extracts the text blocks, joins them, parses the result — and handles servers that return prose instead of JSON.
Eventually we did add a model. What matters is where. It does three jobs.
1. Reviewing the mapping. The model gives a verdict on each source-to-AWS row. Advisory only — it never changes the price — and cached, so editing one row doesn't reprocess the table.
2. Handling free-text requests. For something like "map this workload to something serverless," the model uses the same MCP tools:

Its proposal isn't automatically trusted; the calculator's validation has to confirm the configuration is valid.
The model proposes; the deterministic gate disposes.
3. Writing the report. The model generates the prose summary. The numbers are already calculated — it's only explaining them.
The model never does the arithmetic. Not one dollar figure in the final output comes from model inference.
Both layers use the same MCP client and the same MCP servers.
Starting without AI didn't prevent us from adding it later — it made the integration easier. By the time the model arrived, we already had verified MCP integrations, validated tool schemas, caching, deterministic validation, and clear failure handling.
The model was operating on a controlled interface rather than reaching directly into an unpredictable collection of external systems. That is a much safer architecture.
MCP servers are increasingly available for systems where REST APIs are incomplete, hard to authenticate against, undocumented, absent entirely, or where configuration options change frequently.
And the interface stays remarkably simple — initialize, tools/list, tools/call — over a process pipe or HTTP. The caller can be an AI agent, a Python script, a CI pipeline, a scheduled job, or a traditional application. It doesn't have to be an LLM.
MCP was designed in an AI-driven ecosystem, but the protocol underneath doesn't require AI. Our pipeline makes hundreds of MCP calls, caches most of them, validates its configurations, and produces the same result every time — and for a long time there was no model anywhere in it.
The interesting lesson isn't that MCP works without AI. It’s that MCP can be treated as an integration protocol first and an AI protocol second.
AI may be the reason MCP became popular. But you don't need a model's permission to use it.