AIA2A protocol: how to design an agent API other agents can call
A2A v1.0 shipped in March 2026. Six design steps for an A2A server: Agent Card, outcome skills, eight task states, bindings, per-call auth, versions.

By the end of this guide you will have the design of an A2A server for an agent you already run: a published Agent Card, skills another agent's model can choose between, a task lifecycle mapped to your own job states, an update channel that fits your infrastructure, authorization on every call and a plan for versions. The Agent2Agent (A2A) protocol is an open standard that lets one AI agent discover another, delegate a task to it and collect the result, without either side seeing the other's prompts, memory or tools.
A2A reached its first stable release, v1.0, in March 2026. Google announced it in April 2025 and gave it to the Linux Foundation in June 2025. In August 2026 it became a hosted project of the Agentic AI Foundation, the same neutral home as MCP. The Linux Foundation counts more than 150 supporting organizations and production deployments inside Azure AI Foundry and Amazon Bedrock AgentCore. The protocol is stable enough to design against. The design decisions are still yours, and they are what this guide covers.
Where does A2A sit next to MCP?
The A2A documentation describes MCP as vertical and A2A as horizontal. MCP gives one agent more tools: a database, a calendar API, a file store. A2A connects that agent to other agents it does not control. A scheduling agent inside your SaaS uses MCP to read your calendar service and A2A to accept work from a customer's procurement agent. Most products that expose an agent will run both.
The difference that matters for design is opacity. An MCP client calls your tools one by one and sees each result. An A2A client hands you a goal and gets back a task, status updates and artifacts. Your prompts, your tool chain and your intermediate reasoning stay private. If you have not built the MCP side yet, start with what an MCP server does for a SaaS.
What you need before starting
- An agent that already completes one job end to end inside your product, with its own tools.
- A durable job store, a table or a queue that survives restarts. A2A tasks can run for hours, and callers read them back by id.
- An identity provider that speaks OAuth 2.0 or OpenID Connect. Agent Cards declare API key, HTTP auth, OAuth 2.0, OpenID Connect and mutual TLS schemes.
- One of the official SDKs: Python, Go, Java, JavaScript, C#/.NET or Rust. The spec requires SDKs to handle version negotiation, which is the part you least want to write by hand.
Step 1: Describe each skill as an outcome a caller wants
Skills are what the calling agent's model reads to decide whether to delegate to you. Each skill in the Agent Card carries an id, a name, a description, tags, examples and the media types it accepts and returns. That text is a prompt for a model you have never met. Write it like documentation for a stranger: what the skill does, what it needs, what it returns.
Group skills by outcome. "Reconcile an invoice against its purchase order" is one skill. The functions behind it (fetch the invoice, match the lines, flag the variance) are your implementation, and A2A's opaque execution keeps them out of the card. Keep the list short. Every extra skill is one more choice the caller's model can get wrong. Give each skill at least one example in plain language and, when the input is structured, one in JSON, as the specification's own sample card does.
Step 2: Publish the Agent Card at the well-known address
An A2A server must publish an Agent Card. The standard place is https://your-domain/.well-known/agent-card.json; registries and direct configuration also work. A minimal v1.0 card for one skill looks like this:
{
"name": "Invoice Reconciliation Agent",
"description": "Matches supplier invoices against purchase orders and flags variances.",
"version": "2.1.0",
"provider": { "organization": "Example SaaS", "url": "https://example.com" },
"supportedInterfaces": [
{ "url": "https://agents.example.com/a2a/v1", "protocolBinding": "JSONRPC", "protocolVersion": "1.0" }
],
"capabilities": { "streaming": true, "pushNotifications": true, "extendedAgentCard": true },
"securitySchemes": {
"oidc": { "openIdConnectSecurityScheme": { "openIdConnectUrl": "https://auth.example.com/.well-known/openid-configuration" } }
},
"securityRequirements": [{ "schemes": { "oidc": { "list": ["openid"] } } }],
"defaultInputModes": ["application/json", "text/plain"],
"defaultOutputModes": ["application/json"],
"skills": [
{
"id": "reconcile-invoice",
"name": "Reconcile an invoice",
"description": "Compares one supplier invoice with its purchase order. Returns matched lines, variances and a pass or hold decision.",
"tags": ["finance", "invoices", "reconciliation"],
"examples": ["Reconcile invoice INV-2291 against PO-7781."],
"inputModes": ["application/json", "application/pdf"],
"outputModes": ["application/json"]
}
]
}Four fields carry most of the design weight.
supportedInterfaceslists every endpoint in order of preference. Clients take the first one they support, so put your best-tested binding first.capabilitiesmust be honest. If a client calls an optional feature the card does not declare, the server has to return an error. Declare streaming only once streaming works.securitySchemestells callers how to get a token before the first request.extendedAgentCardlets you keep the public card short and serve tenant-specific skills only to authenticated callers throughGetExtendedAgentCard.
Serve the card with Cache-Control and an ETag derived from its version, as the spec recommends. When callers need to verify where a card came from, sign it with JSON Web Signature over its canonical JSON form (RFC 7515 and RFC 8785).
Step 3: Map your job states to the task lifecycle
A task in A2A v1.0 moves through eight states. Write down which of your internal job states lands in each one before you write any handler.
- Submitted and working: accepted, then in progress.
- Input-required: the agent needs more information from the caller. The caller answers with a new message carrying the same task id.
- Auth-required: the agent needs a credential or a human approval before it continues.
- Completed, failed, canceled and rejected: terminal. A terminal task accepts no further messages.
Three rules from the spec shape the rest. Task ids are generated by the server, and a client cannot create a task with its own id. Results belong in artifacts; messages carry conversation and status, so a caller looking for output reads the artifacts. And a quick answer that needs no tracking can come back as a single message with no task at all. We use rejected for work the agent decides not to do (out of scope, against policy) and keep failed for work that broke, because the caller's next move differs: retry elsewhere, or retry later.
The contextId groups related tasks into one conversation. Agents may expire contexts, and the spec asks you to document that policy. Write the expiry into the skill descriptions or the documentation URL, so a caller knows how long a follow-up stays valid.
Step 4: Choose a binding and an update channel
v1.0 defines three bindings that must behave identically: JSON-RPC 2.0, gRPC and HTTP+JSON. The operations share names across all three: SendMessage, SendStreamingMessage, GetTask, ListTasks, CancelTask, SubscribeToTask. In the REST binding they become paths such as POST /message:send and GET /tasks/{id}. Pick the binding your gateway and observability already handle. On a web stack that usually means JSON-RPC or HTTP+JSON over HTTPS.
Callers can follow a task three ways: polling GetTask, streaming, or push notifications to a webhook they register. Our default is streaming for tasks that finish while the caller waits, and push for anything long enough that the caller may disconnect. Push is where the spec is most specific about security:
- Send the configured credentials with every webhook call.
- Time out webhook requests after 10 to 30 seconds and retry with exponential backoff.
- Reject webhook URLs that point at localhost, link-local or private ranges such as 10.0.0.0/8 and 192.168.0.0/16. Without that check, your agent becomes a request-forgery tool against its own network.
- On the receiving side, process notifications idempotently. Duplicates are allowed by design.
Step 5: Authorize every call against the caller
The spec requires an authorization check on every operation, before any query that could reveal whether a resource exists. ListTasks returns only the tasks the caller may see, even when the request carries no filter. GetTask on another tenant's task answers as if the task did not exist. Map the caller's token to a tenant at the edge, and pass that tenant into every store lookup.
Use the auth-required state for authority the agent does not hold yet: a downstream OAuth token, or a human sign-off before a refund above a threshold. The task pauses, the caller arranges the credential and the work resumes.
Treat every incoming part as untrusted. A delegating agent's text reaches your model, so the prompt-injection defenses you use on user input apply here too. File references must be validated before you fetch them, for the same request-forgery reason as webhooks. The controls we described for hardening an MCP server for enterprise (SSO, audit trail, a gateway in front) carry over almost unchanged.
Step 6: Version the protocol and the card separately
Two version numbers live in an A2A deployment. The protocol version is Major.Minor, sent by clients in the A2A-Version header and declared per interface in the card. The card's own version field is your agent's release, and it changes whenever skills change.
A request with an empty A2A-Version header must be treated as 0.3, and a version the interface does not serve returns VersionNotSupportedError. That matters because v1.0 broke compatibility with 0.3: message/send became SendMessage, enum values moved to SCREAMING_SNAKE_CASE, and protocolVersion moved from the card into each interface. If 0.3 callers already depend on you, serve both versions as separate interfaces and retire 0.3 on a date you publish.
How to verify it works
- Fetch the card with
curl. Expect 200, valid JSON,Cache-Controland anETag. - Send a message with no token. Expect an authentication error, not a task.
- Send a valid message. Expect a task with a server-generated id in the submitted or working state.
- Read that task with a token from another tenant. Expect a task-not-found error.
- Cancel the same task twice. Both calls must leave the same state.
- Open a stream and confirm it closes when the task reaches a terminal state.
- Send
A2A-Version: 9.9. ExpectVersionNotSupportedError. - Register a webhook at
http://127.0.0.1. Expect a refusal.
Common failures and fixes
- Skills named after internal functions. Callers pick the wrong one or none. Rename them after outcomes and add examples.
- Output returned inside status messages. Callers that read artifacts find nothing. Move results into artifacts and keep messages for questions and progress.
- Capabilities declared before they work. A card that says
streaming: truewithout a workingSubscribeToTaskbreaks clients that trust it. Turn the flag on the day the feature ships. - 0.3 and 1.0 names mixed on one endpoint. Clients negotiate one version per interface. Split the interfaces.
- contextId used as identity. It groups conversations and proves nothing about the caller. Authorization comes from the token.
Going further
A2A covers the conversation between agents; each agent still needs its own tools and its own governance. Read what the Linux Foundation handoff changed for MCP for the governance side, and 12 mistakes teams make on their first multi-agent system for what goes wrong once several agents share the work.
Sources
Frequently asked questions
They solve different problems, so one does not replace the other. An MCP server lets an AI agent call your product's tools one at a time, with the caller's model doing the planning. An A2A server lets another agent hand your agent a whole goal and get back a tracked task and its results, while your prompts and tools stay private. If customers' agents mostly need to read and write data, MCP is enough. Add A2A when your own agent does work worth delegating, such as a multi-step reconciliation or a review that can pause for human approval.
No. Google created A2A and announced it in April 2025, then donated it to the Linux Foundation in June 2025 with AWS, Cisco, Microsoft, Salesforce, SAP and ServiceNow as founding members. Since August 2026 it is a hosted project of the Agentic AI Foundation, which also hosts MCP. The normative definition is a Protocol Buffers file in a public repository under the Apache 2.0 license, and changes go through the project's open governance.
No. A REST API is deterministic: the same request returns the same shape every time, which is what integrations, reports and billing depend on. An A2A agent answers goals in natural language or structured data, and its path to the answer can change between model versions. Keep the REST API as the contract for systems, and put the A2A agent beside it for work that needs judgment. Many A2A agents call the same REST API internally through their own tools.
Yes. An A2A server can act as a client toward downstream agents, and chains of agents are a normal pattern. Three things need design. Credentials: your agent calls the next one with its own token, not the original caller's, and records whose request it serves. Time: each hop adds latency, so long chains should use push notifications rather than holding streams open. Loops: put a hop limit or a trace id in message metadata, so two agents cannot delegate the same task back and forth.
Related services
Related articles
AI
AIWebMCP in 2026: the browser API that lets agents use your SaaS UI
WebMCP lets a web page register tools that browser AI agents call directly. Chrome 149 runs it as an origin trial. What it changes for SaaS interfaces.
AI