Skip to content
AI

A2A protocol: how to design an agent API other agents can call

A2A v1.0 shipped in March 2026. Six design steps for an A2A server: Agent Card, outcome skills, eight task states, bindings, per-call auth, versions.

October 6, 2026 · 10 min read

Three dark tiles on the left, the middle one lit amaranth, linked to a card of skill rows pinned to a large closed dark box on the right; an amaranth line runs from the lit tile to the lit skill row, and a row of small fading squares runs back to a lit tile beside it

By the end of this guide you will have the design of an A2A server for an agent you already run: a published Agent Card, skills another agent's model can choose between, a task lifecycle mapped to your own job states, an update channel that fits your infrastructure, authorization on every call and a plan for versions. The Agent2Agent (A2A) protocol is an open standard that lets one AI agent discover another, delegate a task to it and collect the result, without either side seeing the other's prompts, memory or tools.

A2A reached its first stable release, v1.0, in March 2026. Google announced it in April 2025 and gave it to the Linux Foundation in June 2025. In August 2026 it became a hosted project of the Agentic AI Foundation, the same neutral home as MCP. The Linux Foundation counts more than 150 supporting organizations and production deployments inside Azure AI Foundry and Amazon Bedrock AgentCore. The protocol is stable enough to design against. The design decisions are still yours, and they are what this guide covers.

Where does A2A sit next to MCP?

The A2A documentation describes MCP as vertical and A2A as horizontal. MCP gives one agent more tools: a database, a calendar API, a file store. A2A connects that agent to other agents it does not control. A scheduling agent inside your SaaS uses MCP to read your calendar service and A2A to accept work from a customer's procurement agent. Most products that expose an agent will run both.

The difference that matters for design is opacity. An MCP client calls your tools one by one and sees each result. An A2A client hands you a goal and gets back a task, status updates and artifacts. Your prompts, your tool chain and your intermediate reasoning stay private. If you have not built the MCP side yet, start with what an MCP server does for a SaaS.

What you need before starting

  • An agent that already completes one job end to end inside your product, with its own tools.
  • A durable job store, a table or a queue that survives restarts. A2A tasks can run for hours, and callers read them back by id.
  • An identity provider that speaks OAuth 2.0 or OpenID Connect. Agent Cards declare API key, HTTP auth, OAuth 2.0, OpenID Connect and mutual TLS schemes.
  • One of the official SDKs: Python, Go, Java, JavaScript, C#/.NET or Rust. The spec requires SDKs to handle version negotiation, which is the part you least want to write by hand.

Step 1: Describe each skill as an outcome a caller wants

Skills are what the calling agent's model reads to decide whether to delegate to you. Each skill in the Agent Card carries an id, a name, a description, tags, examples and the media types it accepts and returns. That text is a prompt for a model you have never met. Write it like documentation for a stranger: what the skill does, what it needs, what it returns.

Group skills by outcome. "Reconcile an invoice against its purchase order" is one skill. The functions behind it (fetch the invoice, match the lines, flag the variance) are your implementation, and A2A's opaque execution keeps them out of the card. Keep the list short. Every extra skill is one more choice the caller's model can get wrong. Give each skill at least one example in plain language and, when the input is structured, one in JSON, as the specification's own sample card does.

Step 2: Publish the Agent Card at the well-known address

An A2A server must publish an Agent Card. The standard place is https://your-domain/.well-known/agent-card.json; registries and direct configuration also work. A minimal v1.0 card for one skill looks like this:

{
  "name": "Invoice Reconciliation Agent",
  "description": "Matches supplier invoices against purchase orders and flags variances.",
  "version": "2.1.0",
  "provider": { "organization": "Example SaaS", "url": "https://example.com" },
  "supportedInterfaces": [
    { "url": "https://agents.example.com/a2a/v1", "protocolBinding": "JSONRPC", "protocolVersion": "1.0" }
  ],
  "capabilities": { "streaming": true, "pushNotifications": true, "extendedAgentCard": true },
  "securitySchemes": {
    "oidc": { "openIdConnectSecurityScheme": { "openIdConnectUrl": "https://auth.example.com/.well-known/openid-configuration" } }
  },
  "securityRequirements": [{ "schemes": { "oidc": { "list": ["openid"] } } }],
  "defaultInputModes": ["application/json", "text/plain"],
  "defaultOutputModes": ["application/json"],
  "skills": [
    {
      "id": "reconcile-invoice",
      "name": "Reconcile an invoice",
      "description": "Compares one supplier invoice with its purchase order. Returns matched lines, variances and a pass or hold decision.",
      "tags": ["finance", "invoices", "reconciliation"],
      "examples": ["Reconcile invoice INV-2291 against PO-7781."],
      "inputModes": ["application/json", "application/pdf"],
      "outputModes": ["application/json"]
    }
  ]
}

Four fields carry most of the design weight.

  • supportedInterfaces lists every endpoint in order of preference. Clients take the first one they support, so put your best-tested binding first.
  • capabilities must be honest. If a client calls an optional feature the card does not declare, the server has to return an error. Declare streaming only once streaming works.
  • securitySchemes tells callers how to get a token before the first request.
  • extendedAgentCard lets you keep the public card short and serve tenant-specific skills only to authenticated callers through GetExtendedAgentCard.

Serve the card with Cache-Control and an ETag derived from its version, as the spec recommends. When callers need to verify where a card came from, sign it with JSON Web Signature over its canonical JSON form (RFC 7515 and RFC 8785).

Step 3: Map your job states to the task lifecycle

A task in A2A v1.0 moves through eight states. Write down which of your internal job states lands in each one before you write any handler.

  • Submitted and working: accepted, then in progress.
  • Input-required: the agent needs more information from the caller. The caller answers with a new message carrying the same task id.
  • Auth-required: the agent needs a credential or a human approval before it continues.
  • Completed, failed, canceled and rejected: terminal. A terminal task accepts no further messages.

Three rules from the spec shape the rest. Task ids are generated by the server, and a client cannot create a task with its own id. Results belong in artifacts; messages carry conversation and status, so a caller looking for output reads the artifacts. And a quick answer that needs no tracking can come back as a single message with no task at all. We use rejected for work the agent decides not to do (out of scope, against policy) and keep failed for work that broke, because the caller's next move differs: retry elsewhere, or retry later.

The contextId groups related tasks into one conversation. Agents may expire contexts, and the spec asks you to document that policy. Write the expiry into the skill descriptions or the documentation URL, so a caller knows how long a follow-up stays valid.

Step 4: Choose a binding and an update channel

v1.0 defines three bindings that must behave identically: JSON-RPC 2.0, gRPC and HTTP+JSON. The operations share names across all three: SendMessage, SendStreamingMessage, GetTask, ListTasks, CancelTask, SubscribeToTask. In the REST binding they become paths such as POST /message:send and GET /tasks/{id}. Pick the binding your gateway and observability already handle. On a web stack that usually means JSON-RPC or HTTP+JSON over HTTPS.

Callers can follow a task three ways: polling GetTask, streaming, or push notifications to a webhook they register. Our default is streaming for tasks that finish while the caller waits, and push for anything long enough that the caller may disconnect. Push is where the spec is most specific about security:

  • Send the configured credentials with every webhook call.
  • Time out webhook requests after 10 to 30 seconds and retry with exponential backoff.
  • Reject webhook URLs that point at localhost, link-local or private ranges such as 10.0.0.0/8 and 192.168.0.0/16. Without that check, your agent becomes a request-forgery tool against its own network.
  • On the receiving side, process notifications idempotently. Duplicates are allowed by design.

Step 5: Authorize every call against the caller

The spec requires an authorization check on every operation, before any query that could reveal whether a resource exists. ListTasks returns only the tasks the caller may see, even when the request carries no filter. GetTask on another tenant's task answers as if the task did not exist. Map the caller's token to a tenant at the edge, and pass that tenant into every store lookup.

Use the auth-required state for authority the agent does not hold yet: a downstream OAuth token, or a human sign-off before a refund above a threshold. The task pauses, the caller arranges the credential and the work resumes.

Treat every incoming part as untrusted. A delegating agent's text reaches your model, so the prompt-injection defenses you use on user input apply here too. File references must be validated before you fetch them, for the same request-forgery reason as webhooks. The controls we described for hardening an MCP server for enterprise (SSO, audit trail, a gateway in front) carry over almost unchanged.

Step 6: Version the protocol and the card separately

Two version numbers live in an A2A deployment. The protocol version is Major.Minor, sent by clients in the A2A-Version header and declared per interface in the card. The card's own version field is your agent's release, and it changes whenever skills change.

A request with an empty A2A-Version header must be treated as 0.3, and a version the interface does not serve returns VersionNotSupportedError. That matters because v1.0 broke compatibility with 0.3: message/send became SendMessage, enum values moved to SCREAMING_SNAKE_CASE, and protocolVersion moved from the card into each interface. If 0.3 callers already depend on you, serve both versions as separate interfaces and retire 0.3 on a date you publish.

How to verify it works

  1. Fetch the card with curl. Expect 200, valid JSON, Cache-Control and an ETag.
  2. Send a message with no token. Expect an authentication error, not a task.
  3. Send a valid message. Expect a task with a server-generated id in the submitted or working state.
  4. Read that task with a token from another tenant. Expect a task-not-found error.
  5. Cancel the same task twice. Both calls must leave the same state.
  6. Open a stream and confirm it closes when the task reaches a terminal state.
  7. Send A2A-Version: 9.9. Expect VersionNotSupportedError.
  8. Register a webhook at http://127.0.0.1. Expect a refusal.

Common failures and fixes

  • Skills named after internal functions. Callers pick the wrong one or none. Rename them after outcomes and add examples.
  • Output returned inside status messages. Callers that read artifacts find nothing. Move results into artifacts and keep messages for questions and progress.
  • Capabilities declared before they work. A card that says streaming: true without a working SubscribeToTask breaks clients that trust it. Turn the flag on the day the feature ships.
  • 0.3 and 1.0 names mixed on one endpoint. Clients negotiate one version per interface. Split the interfaces.
  • contextId used as identity. It groups conversations and proves nothing about the caller. Authorization comes from the token.

Going further

A2A covers the conversation between agents; each agent still needs its own tools and its own governance. Read what the Linux Foundation handoff changed for MCP for the governance side, and 12 mistakes teams make on their first multi-agent system for what goes wrong once several agents share the work.

Sources

Frequently asked questions

Related articles

Studio

Start a project.

We write about what we build. Tell us what you want to build.