MCP already has pagination — for the wrong thing. The protocol defines a cursor / nextCursor pattern for tools/list, resources/list, and prompts/list, so a client can page through a server’s catalog of tools without pulling them all in one response. What it has no built-in mechanism for is a single tool call that returns a lot of data: a search that matches ten thousand rows, a list-records call against a CRM with no upper bound, a log query that could return anything from ten lines to ten thousand. That gap is yours to design around, and getting it wrong is one of the quieter ways an MCP server becomes unusable.

Why this is a different problem from listing tools

List-primitive pagination is a protocol-level concern: it’s about how a client discovers what a server offers, and it’s handled the same way for every server that implements it. A tool result is different. The server doesn’t know, ahead of time, how big the underlying answer to “find all orders from last month” will be — that depends entirely on the caller’s data. And the consumer of that result isn’t a UI rendering a scrollable list; it’s a model, reading the whole thing into its context window in one shot.

That second part is what makes this worth solving deliberately instead of shrugging and returning everything. A model doesn’t page through a wall of text the way a person scrolls a page — it reads all of it, every time, and pays for it in context. A single oversized tool result can crowd out the rest of a conversation, get silently truncated by the client, or just make the model worse at using what it got. See our post on how many tools an MCP server should expose for the same budget problem one level up, at the tool-list stage instead of the result stage.

The patterns that actually work

None of these are exotic. They’re the same techniques REST APIs have used for years, adapted to a tool call instead of an HTTP request:

  • Cursor or offset arguments on the tool itself. Give the tool its own cursor (or page / offset) and limit input parameters, and return a cursor for the next page in the result alongside the data. This mirrors the list-primitive pattern closely enough that a model already primed on MCP tends to use it correctly without much prompting.
  • A hard cap with a clear signal that it was hit. Return at most N results and say so explicitly — “showing 50 of 3,241 matches, narrow your filter or call again with an offset” — rather than silently truncating. A model that doesn’t know it saw a partial answer will confidently reason from an incomplete one.
  • Push filtering into the tool’s input schema. The best fix for a huge result is often not paginating it but not producing it: date ranges, status filters, and a required search term turn “list all orders” into something that returns a workable size by default.
  • Summarize before you return raw rows. For some tools, the useful answer isn’t the full dataset at all — it’s an aggregate (count, sum, top N) with an option to drill in. A tool that returns “142 orders totaling €38,004, top 5 by value: …” is often more useful to a model than the same 142 rows in full.

Whichever pattern you pick, keep the mechanism visible in the tool’s description and input schema, not just in a comment in your code. The model only knows a tool paginates if the description says so and the schema exposes the cursor argument — the same principle we cover in writing MCP tool descriptions.

What doesn’t work

Returning everything and letting the client figure it out is the most common failure mode, because it’s the path of least resistance when you’re first building a tool against a small test dataset. It works right up until someone points the tool at a real account with real volume, and then it either times out, gets truncated somewhere in the transport, or blows the context budget for the rest of the conversation.

The other common mistake is inventing a second tool just to page through the first one’s results — a get_more_results tool with no arguments that somehow remembers where the last call left off. That relies on server-side state tied to a specific conversation, which sits awkwardly against MCP’s move toward a stateless core, and it’s fragile the moment a client retries a call or runs two searches in parallel. A cursor argument the model passes back explicitly is more verbose but far more robust, for the same reason we describe in MCP tool idempotency: state that lives implicitly on the server is state that breaks the moment a call doesn’t happen exactly once, in exactly the order you assumed.

If you’re generating a server from an API

This is exactly the gap we flagged in turning an API into an MCP server: a paged list endpoint in a REST API or OpenAPI spec doesn’t automatically become a well-behaved paginated tool. An auto-converted tool usually just forwards the API’s own page or cursor parameter as-is, which is a reasonable start, but it’s worth checking after conversion that the tool’s description actually tells the model when and how to use it, and that there’s a sane default limit instead of the API’s own default (which is often “no limit” or a number picked for a web UI, not a model’s context window). If you use a spec-based generator, this is one of the first things worth reviewing in the drafted tool list before you approve it.

Where this meets governance: a paginated tool is still a tool, and a “list everything” call against a shared system is often the one worth a second look — not because pagination is a security feature, but because a tool that can return an unbounded amount of someone else’s data is exactly the kind of call a gateway’s per-tool rules and activity log are meant to make visible, regardless of how many pages it took to get there.

The bottom line

MCP’s cursor-based pagination covers discovering a server’s tools, resources, and prompts — it says nothing about what a single tool call returns. For any tool whose result size depends on the caller’s data, build in a cursor or limit argument, cap the default, say clearly when a result was cut short, and push filtering into the input schema before you resort to paging at all. See gate’s MCP server directory for how a range of real servers handle large result sets today.