Most MCP servers are shaped like the API underneath them, not like the job an agent is actually trying to do. That happens almost by default: whether you hand-write tools or generate them from a spec, the easy move is one tool per API operation. The fix isn’t a different generator — it’s deciding, operation by operation, whether an agent should see the API’s shape at all, or a shape built around the job it’s there to finish.

Why “one tool per endpoint” is the default, not a decision

An API already has the raw material for MCP tools: each operation has a method, a path, parameters, and a response shape. Mapping each one to a tool is mechanical, which is why every generator — from a hand-rolled script to an OpenAPI-to-MCP converter — defaults to it. Nobody designs a 1:1 mapping on purpose; it’s just what falls out of automating the conversion.

That default optimizes for coverage, not usability. It guarantees every operation the API supports is reachable. It says nothing about whether an agent calling those tools in sequence will actually finish the job correctly, or whether the model can tell which of several similar-looking tools to reach for — the selection problem we cover in how many tools an MCP server should expose.

A job, not an endpoint, is the right unit to design around

Take a billing API with separate endpoints for create_invoice, add_line_item, and send_invoice. Mirrored 1:1, an agent asked to “invoice the client for this work” has to call three tools in the right order, carry the invoice ID between them, and notice if step two silently failed before calling step three. Every one of those is a place the sequence can go wrong without the model realizing it.

A tool shaped around the job — one create_and_send_invoice call that takes a client and line items and does all three steps server-side — removes that failure mode entirely. The agent makes one decision instead of three, and there’s no intermediate state for it to lose track of. This is the same instinct behind good tool naming and descriptions: write for what the caller is trying to accomplish, not for what the backend happens to call its internal methods.

What you give up when you combine steps

Composite tools aren’t free. A few real costs worth weighing before you collapse a sequence into one call:

  • Less flexibility. If an agent sometimes needs to create an invoice without sending it — to let a human review it first — a single fused tool can’t do that. You either keep both the atomic and composite versions, or you lose the partial case.
  • Harder retries. A tool that does three things server-side needs to behave correctly if the model calls it twice after a timeout. Our piece on MCP tool idempotency covers this in general; it gets more pressing here, because “did it partially run” is a harder question for a three-step tool than a one-step one.
  • Murkier errors. If step two of three fails inside a composite tool, the error has to say which step failed and what state the job is left in — not just that “something went wrong.” See our guide on MCP tool error handling for what a model actually needs in that message to recover instead of retrying blindly.

When not to combine

A few cases where the atomic, per-endpoint tool is the better call:

  • The steps are genuinely independent. If a lookup is useful on its own — checking an invoice’s status, say — fusing it into a bigger workflow tool just hides a capability agents will want standalone.
  • One step is destructive and the others aren’t. A tool that reads a customer record and then deletes one is a worse design than two separate tools, because you can’t put the delete behind its own approval gate if it’s welded to a read. Keep destructive actions as their own tool so a gateway, or a human, can treat them differently.
  • The sequence genuinely varies by caller. If some integrations need steps in a different order, or need to stop partway, a fixed composite tool can’t represent that. Atomic tools plus a clear description of the usual order is more honest than a workflow tool that only covers the common case.

Where this leaves auto-generated servers

If you’re generating a server from an OpenAPI spec — by hand with a converter, or with gate’s MCP Builder, which reads a spec or docs URL and drafts one tool per operation, capped at 60 and each one flagged read or write for you to review — you’re starting from the atomic, endpoint-shaped end of this spectrum by construction. That’s a reasonable starting point, and for a lot of APIs it’s where you should stop: most operations really are independent jobs. But if you notice an agent consistently calling the same three tools in the same order to get anything done, that’s the signal to write a small amount of custom code on top that exposes the job directly, rather than asking every agent that connects to rediscover the right sequence on its own.

Whichever shape you land on, review the result the same way: run it through gate’s free security scanner to check the tool descriptions aren’t hiding anything, and decide per tool — atomic or composite — whether it needs to sit behind an approval step before an agent can call it unsupervised.

The short version: an API endpoint and an agent’s job are not the same unit, and defaulting to the endpoint shape just because it’s what the generator produces is a decision you’re making without noticing. Combine steps an agent always does together; keep steps separate when one is destructive, independently useful, or varies by caller. Either way, the question is the same one that should drive every tool on the server: what is the agent actually trying to get done.

The bottom line

One tool per API operation is the cheapest server to generate, not the easiest one for an agent to use well. Browsing the servers in gate’s MCP server directory, the pattern holds: the ones that read as well-designed usually have a tool list that lines up with the jobs people actually hire the integration to do, not a 1:1 mirror of the vendor’s API reference. Start from the job, let the API shape the implementation, not the interface.