Point your AI assistant at your invoices: shipping an MCP server
FoxyInvoice now speaks MCP β the Model Context Protocol. Connect Claude Desktop, Claude Code, VS Code, or Cursor to your workspace and ask it to invoice a client in plain English. This post is the whole design: how auth works when the client is a robot, why the tools are thin wrappers over the REST handlers, and the one trick that makes agent retries harmless.
What shipped
A JSON-RPC endpoint at /api/v1/mcp that speaks the Model
Context Protocol's Streamable HTTP transport, exposing seven tools:
| Tool | What it does |
|---|---|
list_clients | Search clients by name or email |
create_client | Create a client |
list_products | List the catalog |
list_invoices | List invoices, filter by status/client |
get_invoice | One invoice, all lines, computed totals |
create_invoice | Create a draft invoice |
send_invoice | Email a draft to its client |
The conversation looks like this:
βInvoice Globex for 10 hours of consulting at $150/hour, due in 30 days.β β the agent callslist_clientsto find Globex's id, callscreate_invoice, and reports back: βCreated INV-2026-0007 for $1,608.75 ($1,500 + $108.75 CA tax), due November 7. It's a draft β want me to send it?β
Every number in that sentence came from the server, not the model. That is the design constraint the rest of this post hangs off.
Auth: a robot is not a browser session
The SPA authenticates with short-lived JWTs behind httpOnly cookies. That is exactly right for a browser and exactly wrong for an assistant you configure once: a token that dies every 15 minutes breaks every MCP client config on earth. But a long-lived credential needs a story for "how do I make it stop".
The answer is personal access tokens β the same shape GitHub chose:
- Generated in Settings β AI assistants, shown once, stored only as a SHA-256 hash. The database can't leak what it doesn't have.
- Presented as
Authorization: Bearer foxy_β¦on every call. No cookies, no handshake state, no session id β the server is stateless, so any HTTP client that can POST JSON can drive it. - Revocable in one click, checked on every request.
- A token acts as its user β live. There is no permissions snapshot in the token. Every request re-resolves the user's current roles from the database, so disabling a user or changing their roles takes effect on the very next tool call. Nothing to propagate, no cache to invalidate, no "I removed them from the workspace and their integration still worked" bug class.
JWTs are deliberately not accepted at the MCP endpoint β a short-lived browser credential pasted into an assistant config would be a support ticket factory.
Tools are thin, the core is shared
The tempting way to build this is a parallel implementation: tool
handlers that re-do the queries the REST handlers already do. The correct
way is to make the REST handlers thin and share what's underneath. We
extracted the bodies of the client, product, invoice, and send handlers
into *_core functions that take the authenticated user and the
request, and both surfaces call them:
REST handler ββ
βββΊ clients::create_core(state, auth, body)
MCP tool call ββ
Which means everything the REST API guarantees, the tools inherit for free, because it is literally the same code:
- Permission checks β a token can never do more than
its user can.
create_invoicerunsinvoice:create; no permission, no draft. - Tenant scoping β every query filters by the token's tenant. Cross-tenant probing returns not-found, not data.
- Server-owned money math β the agent passes
qtyandunitPriceand nothing else. The server assigns theINV-YYYY-NNNNnumber, computes per-line tax from the client's jurisdiction and the tenant's nexus rules, rounds per the invoice rules, writes the audit trail. The model never states an amount, because a confidently wrong total on an invoice is not a bug, it's a liability. - Quotas β
send_invoicecounts against the same monthly send limit as the UI.
Sending is a separate tool, on purpose
A draft costs nothing. An email is irreversible. So creating and sending
are different tools with different names, and create_invoice
always produces a Draft β there is no "send too" parameter.
The send tool's description tells the agent to confirm with the user
first. Is a tool description binding? No. But models follow it
remarkably well, and the real backstop is structural: an agent has to make
a second, deliberate, differently-named call to reach a human being's
inbox. Accidents need two mistakes instead of one.
The retry problem, solved with an idempotency key
Agents retry. A dropped connection mid-create_invoice means
the model will try again β and without protection, "invoice the client"
becomes two invoices. Every transport error would be a coin flip on a
duplicate.
So create_invoice accepts an optional requestId
(any UUID the agent picks and reuses across retries of the same logical
create). The server stores it as the invoice's client-supplied id; a
replay with the same requestId returns the
existing invoice instead of inserting a duplicate. First call
creates, retry returns what the first call created. This is the same
mechanism the offline-first SPA already uses β when you create a client on
a plane and it syncs later, a retried sync can't duplicate it either. One
idea, two surfaces.
Errors an agent can act on
MCP separates protocol errors from tool errors, and we lean on that:
business failures β client not found, validation failed, quota exhausted,
no permission β come back as successful tool calls with
isError: true and a plain-language message. The model reads
"Client has no email address β cannot send invoice", tells the user, and
offers to add one. Only genuine protocol garbage (malformed JSON-RPC,
unknown method) is a JSON-RPC error. The difference is an assistant that
recovers by itself versus one that says "an error occurred".
The protocol, honestly assessed
MCP is young and moving fast, and we took a position on it: implement the spec's Streamable HTTP transport statelessly and skip the rest. No session ids, no server-initiated SSE streams, no stdio. What that buys:
- The endpoint is documented in one page β initialize negotiation, tools/list, tools/call, and a 401 when the bearer is missing.
- Statelessness composes with idempotency: any request can die and be retried, because there is no session to lose.
- Server-side, it's one axum handler β the same codebase weight class as any other endpoint, not a subsystem.
We'll track the spec as it settles. The stateless subset is the part every client already agrees on.
Try it
If you have a workspace: Settings β AI assistants β Generate token, paste the one-liner it gives you into Claude Code (or the JSON into Claude Desktop), and ask your assistant to invoice someone. The full reference β every tool, the token lifecycle, the wire format β lives at foxyinvoice.com/docs/mcp.