Anthropic Claude API: A Practical Guide

Anthropic Claude, LLM RAG, Use Cases

What Is Anthropic Claude?

Claude is a large language model (LLM) developed by Anthropic, named after Claude Shannon, the father of information theory. Claude is built with a focus on safety and helpfulness, and it’s the model family behind Claude.ai, Claude Code, and the Claude API covered in this guide.

Claude models power applications ranging from text summarization and content creation to code generation and autonomous agents. The current model lineup, as of this writing, includes Claude Opus 5, Claude Sonnet 5, and Claude Haiku 4.5, each right-sized for a different balance of intelligence, speed, and cost, plus limited-availability models like Claude Fable 5 and Claude Mythos 5 for the most demanding, long-running agentic work.

This is part of a series of articles about Anthropic Claude.

What Is the Claude API?

The Claude API, developed by Anthropic, allows developers to integrate Claude’s language models into their own applications for tasks like text generation, summarization, document analysis, and conversational agents. It’s a RESTful API at api.anthropic.com that gives you programmatic access to Claude models directly, rather than through the Claude.ai chat interface.

It’s built for scale: pay-as-you-go pricing by default, with usage tiers that increase automatically as your organization’s usage history grows, and custom volume pricing available for enterprise deployments. Access to Claude models is also available through Amazon Bedrock and Google Cloud Vertex AI, for teams that want to stay inside their existing cloud provider’s billing and compliance boundary.

Features of the Claude API:

  • Text and code generation: create detailed responses, generate and debug code, and more.
  • Large context windows: current-generation models support context windows up to 1 million tokens, enough to hold an entire codebase or a long document set in a single request.
  • Model Context Protocol support: connect Claude directly to remote MCP servers through the built-in MCP connector, without writing your own MCP client.
  • Advanced tool use: let Claude call external tools and APIs, run sandboxed code, search and fetch live web content, and read and write files through the Files API.
  • Prompt caching: reuse previously processed context across requests to cut costs and latency.
  • Skills and Memory: teach Claude reusable procedures and best practices, and let it store and consult information across a session through a dedicated memory file.
  • Claude Managed Agents: a set of composable APIs for building and deploying longer-running agents at scale, billed separately from the standard Messages API.
  • Enterprise-grade security: SOC 2 Type II compliance, and a HIPAA Business Associate Agreement available on request for the Claude API and certain Claude Enterprise configurations, with zero data retention required for some surfaces. See the note on HIPAA further down before treating this as a simple checkbox.
  • Official SDK support: client libraries for Python and TypeScript, with community SDKs for other languages.

Claude API Pricing

Claude API pricing is pay-as-you-go, billed per million tokens (MTok) processed, and it varies by model. As of this writing, here’s the current lineup:

Model Input Output Best for
Claude Opus 5 $5 / MTok $25 / MTok Complex agentic coding and enterprise work
Claude Haiku 4.5 $1 / MTok $5 / MTok Fastest, most cost-effective model for lightweight tasks
Claude Fable 5 $10 / MTok $50 / MTok Long-running, complex agentic work

*Sonnet 5 pricing is introductory through August 31, 2026; standard pricing of $3 input / $15 output per MTok applies from September 1, 2026.

Prompt caching adds its own pricing on top of base rates: a 5-minute cache write costs 1.25x the base input price, a 1-hour cache write costs 2x, and a cache read (hit) costs just 0.1x the base input price, so caching pays for itself after one or two reads depending on the cache duration you choose.

Other pricing to know about:

  • Batch API: 50% off both input and output tokens for asynchronous, non-time-sensitive processing.
  • Web search: $10 per 1,000 searches, plus standard token costs for the content returned.
  • Web fetch: no additional charge beyond standard token costs for the fetched content.
  • Code execution: free when used alongside web search or web fetch; otherwise billed by execution time, with 1,550 free hours per organization per month, then $0.05 per hour per container.
  • Claude Managed Agents: billed on both tokens (at standard Model pricing rates) and session runtime, at $0.08 per session-hour while a session is actively running.

For the full, current breakdown, including cloud platform pricing on AWS and Google Cloud, always check Anthropic’s official pricing page before budgeting, since per-model rates and introductory pricing windows change over time.

Claude API Vs. Claude.ai Subscription Pricing

It’s worth being precise about a common point of confusion: the Claude API’s pay-as-you-go pricing above is separate from the Claude.ai and Claude Code subscription plans (Free, Pro, Max, Team, Enterprise). Those are flat monthly or annual plans for using Claude through the chat interface or Claude Code with a personal account, not for building your own application. If you’re integrating Claude into a product, the API’s per-token pricing is what applies; the consumer subscription tiers aren’t a substitute for API access, and API usage isn’t included in a Pro or Max subscription.

Claude API Rate Limits

Anthropic enforces rate limits on the Claude API to manage capacity and prevent misuse. Limits are defined by usage tier, and organizations are placed on a tier automatically based on usage history and account standing, moving to higher tiers over time as they use the API responsibly, rather than through a fixed manual upgrade process. New organizations typically start on an entry-level Evaluation tier with lower limits while account history is established.

Rate limits are enforced along three dimensions, tracked independently:

  • Requests per minute (RPM): how many API calls your organization can make in a 60-second window.
  • Input tokens per minute (ITPM): how many input tokens can be processed per minute.
  • Output tokens per minute (OTPM): how many output tokens can be generated per minute.

Cached tokens from prompt caching don’t count against your ITPM limit the same way fresh input tokens do, which is one of the reasons caching can meaningfully raise your effective throughput without needing a higher tier. Exact RPM/ITPM/OTPM values differ by model and tier and are visible for your organization in the Claude Console, or programmatically through the Rate Limits API; check there for current numbers rather than relying on a fixed published table, since limits are tuned per organization and change over time.

If you exceed a rate limit, the API returns a 429 error, with response headers showing your current usage and when the limit resets. There’s also a separate, fixed request size limit to be aware of: 20 MB per request on Amazon Bedrock, 30 MB on Google Cloud, with Claude Platform on AWS using the same limits as the direct Claude API; exceeding this returns a 413 request_too_large error regardless of your usage tier.

Getting Started with Anthropic Claude API

To begin using the Claude API, follow these steps to set up your environment and make your first API call.

Start with the Workbench

Before diving into code, it’s worth starting with the Workbench, a web-based interface within the Claude Console that lets you experiment with Claude’s capabilities interactively.

  1. Log into the Claude Console and navigate to the Workbench.
  2. Ask Claude a question by typing it into the User section, for example, “Why is the Earth round?”
  3. Run the query and observe the response. You can adjust the response with a System Prompt, for example asking Claude to answer only in bullet points.
  4. Convert your Workbench session into code by clicking Get Code, which generates Python or TypeScript code that replicates your session.

Install the SDK

Anthropic provides official client SDKs for Python (3.7+) and TypeScript (4.5+).

Python:

  1. Create a virtual environment: python -m venv claude-env
  2. Activate it: source claude-env/bin/activate on macOS/Linux, or claude-env\Scripts\activate on Windows.
  3. Install the SDK: pip install anthropic

Set Your API Key

Each API call requires an API key for authentication. The SDK expects this key to be set as an environment variable named ANTHROPIC_API_KEY.

macOS and Linux:

    export ANTHROPIC_API_KEY='My-API-key'

Windows:

    set ANTHROPIC_API_KEY='My-API-key'

Treat your API key like any other credential: don’t commit it to source control, and use separate keys per environment or service where possible so you can revoke one without affecting the rest.

If you’re calling the API directly over HTTP rather than through an SDK, the key goes in the x-api-key request header (the SDKs handle this for you automatically from the ANTHROPIC_API_KEY environment variable):

curl https://api.anthropic.com/v1/messages \

  --header "x-api-key: $ANTHROPIC_API_KEY" \

  --header "anthropic-version: 2023-06-01" \

  --header "content-type: application/json" \

  --data '{"model": "claude-sonnet-5", "max_tokens": 1024, "messages": [{"role": "user", "content": "Hello, Claude"}]}'

Claude API Examples

Basic Request and Response

import anthropic

client = anthropic.Anthropic()

message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Hey, Claude"}
    ],
)

print(message.content)

This sends a single user message and returns Claude’s reply as a structured response object, with the reply text in message.content alongside usage information showing input and output token counts, which you can use to track cost as you build.

Multi-Turn Conversations

The Claude API is stateless: each request must include the full conversation history you want Claude to have context on.

import anthropic

client = anthropic.Anthropic()

message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Hey, Claude"},
        {"role": "assistant", "content": "Hey there!"},
        {"role": "user", "content": "Explain what an LLM is."}
    ],
)

print(message.content)

This lets you build up multi-turn conversations by appending each new user and assistant turn to the messages list on every subsequent call.

System Prompts

A system prompt sets Claude’s role or constraints for the whole conversation, separate from the user’s messages:

message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    system="You are a concise technical writer. Answer in three sentences or fewer.",
    messages=[
        {"role": "user", "content": "What is prompt caching?"}
    ],
)

Vision

Claude can process both text and images in a single request. Images need to be base64-encoded and included in the message content, and Claude supports JPEG, PNG, GIF, and WebP formats.

import anthropic
import base64
import httpx

image_url = "https://upload.wikimedia.org/wikipedia/commons/example.jpg"
image_media_type = "image/jpeg"
image_data = base64.b64encode(httpx.get(image_url).content).decode("utf-8")

client = anthropic.Anthropic()

message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image",
                    "source": {
                        "type": "base64",
                        "media_type": image_media_type,
                        "data": image_data,
                    },
                },
                {"type": "text", "text": "What's in this image?"}
            ],
        }
    ],
)

print(message.content)

Claude API and the Model Context Protocol

The Claude API’s MCP connector lets you connect Claude directly to remote MCP servers from the Messages API, without running your own MCP client. Instead, you list one or more servers in an mcp_servers array, along with an mcp_toolset entry per server controlling which tools are enabled:

client = anthropic.Anthropic()

response = client.beta.messages.create(
    model="claude-opus-5",
    max_tokens=1000,
    messages=[{"role": "user", "content": "What tools do you have available?"}],
    mcp_servers=[
        {
            "type": "url",
            "url": "https://example-server.modelcontextprotocol.io/sse",
            "name": "example-mcp",
            "authorization_token": "YOUR_TOKEN",
        }
    ],
    tools=[{"type": "mcp_toolset", "mcp_server_name": "example-mcp"}],
    betas=["mcp-client-2025-11-20"],
)

A few things worth knowing: the server has to be publicly reachable over HTTP (Streamable HTTP or SSE); local stdio servers aren’t supported through this route. You can allowlist or denylist individual tools per server rather than exposing every tool a server offers. And this is a distinct feature from connecting MCP servers in Claude Code or Claude Desktop, covered in our Claude MCP guide, which walks through the CLI-based claude mcp add workflow for interactive use.

Challenges of Using the Claude API Directly

While the Claude API is powerful, implementing it in real-world applications often comes with real friction, especially for teams looking to move quickly:

  • Backend infrastructure requirements. You’ll need to build and maintain systems to handle API requests, authentication, and data flow.
  • Manual prompt management. Prompts often live in code, making them harder to update, test, and standardize across use cases.
  • Handling rate limits and errors. Managing retries, failures, and usage limits adds complexity, especially as you scale past a single tier.
  • Inconsistent outputs. Without structured prompt systems, results can vary across teams or use cases.
  • Scaling across teams and workflows. What works for one use case can be difficult to replicate across different departments or processes.

The Claude API provides a strong foundation, but many teams find that building and maintaining everything around it takes significant engineering time.

From API Calls to Real Workflows

Using the Claude API effectively goes beyond making individual calls. Most real use cases require coordinating multiple steps into a complete workflow:

  • Collecting input (a user query, a document, a data source)
  • Structuring prompts for consistent output
  • Processing and formatting responses
  • Triggering actions (sending a message, updating a system, storing data)
  • Integrating with existing tools like CRMs, databases, or communication platforms, often through MCP servers rather than one-off custom integrations

Managing these steps by hand gets complex quickly, especially as a workflow grows in scope and importance. Instead of building everything from scratch, many teams look for a platform that handles the orchestration layer.

Turn Claude API Ideas Into Real Workflows with Obot

Claude’s API is powerful, but turning it into production-ready workflows takes time, infrastructure, and ongoing management.

Obot is an open source platform that helps bridge that gap. Admins can deploy Obot and connect it to Claude and to the tools your team already uses, then publish agents and workflows for people to interact with. It includes native support for the Model Context Protocol, so connecting Claude to internal systems doesn’t mean writing and maintaining a custom integration for every tool.


To get started, deploy Obot via Docker:
docker run -d --name obot -p 8080:8080 \
  -v /var/run/docker.sock:/var/run/docker.sock \
  ghcr.io/obot-platform/obot:latest

Then visit http://localhost:8080.

Obot helps you:

  • Reduce the need for custom backend development
  • Integrate Claude with your existing tools through governed MCP connections
  • Build and deploy AI-powered workflows faster, with centralized access control and audit logging

Building apps with the Claude API

Using Otto8 to build agents and workflows

Otto8 is an open source agent platform that works with different GenAI model providers, including Claude to develop, run and share AI agents and workflows.  It provides a full platform for building copilots, assistants and automated workflows. Out of the box it includes a full RAG capability for ingesting data, websites and pdfs, as well as a framework for integrating with web services and APIs.  Admins can deploy and run an Otto8 server, and then publish agents and workflows for users to interact with.   The platform includes native OAuth 2.0 authentication for user services, so users can authenticate into their instances of apps.

To get started, just deploy Otto8 via docker:

docker run -d -p 8080:8080 -e "OPENAI_API_KEY=<OPEN AI KEY>" ghcr.io/otto8-ai/otto8:latest

Then visit http://localhost:8080.

The otto8 CLI can be installed via brew on MacOS or Linux:

brew tap otto8-ai/tap
brew install otto8

You can also start by downloading the binary directly from the latest releases on GitHub.

Claude API Examples

Basic Request and Response

To make a basic request to the Claude API and receive a response, you can use the following code snippet. This example shows how to send a simple message to Claude and receive a reply.

Python example:

    import anthropic

    # Initialize the Claude client
    client = anthropic.Anthropic()

    # Send a basic message to Claude
    message = client.messages.create(
        model="claude-3-5-sonnet-20240620",
        max_tokens=1024,
        messages=[
            {"role": "user", "content": "Hey, Claude"}
        ],
    )

    # Print the response
    print(message.content)

Response example:

    {
      "id": "msg_01XFDUDYJgAACzvnptvVoYEL",
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "text",
          "text": "Hey there!"
        }
      ],
      "model": "claude-3-5-sonnet-20240620",
      "stop_reason": "end_turn",
      "usage": {
        "input_tokens": 12,
        "output_tokens": 6
      }
    }

In this example, the user sends a message saying “Hello, Claude,” and the API responds with “Hello!” This basic interaction demonstrates how to send a simple prompt and receive a straightforward reply.

Multiple Conversational Turns

The Claude API supports multiple conversational turns by maintaining a stateless API, meaning that each request must include the full conversation history. This allows developers to build complex interactions over time.

Python example:

    import anthropic

    # Initialize the Claude client
    client = anthropic.Anthropic()

    # Send a series of messages to create a conversation
    message = client.messages.create(
        model="claude-3-5-sonnet-20240620",
        max_tokens=1024,
        messages=[
            {"role": "user", "content": "Hey, Claude"},
            {"role": "assistant", "content": "Hey there!"},
            {"role": "user", "content": "Explain what is an LLM."}
        ],
    )

    # Print the response
    print(message.content)

Response example:

    {
      "id": "msg_018gCsTGsXkYJVqYPxTgDHBU",
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "text",
          "text": "Sure, here's a description of a large language model (LLM)..."
        }
      ],
      "stop_reason": "end_turn",
      "usage": {
        "input_tokens": 50,
        "output_tokens": 500
      }
    }

This example builds a conversation with Claude, where the user asks multiple questions, and Claude responds accordingly, maintaining the context of the conversation.

Customizing the Response

You can influence Claude’s response by pre-filling part of the response content. This is particularly useful when you want to guide Claude towards a specific type of answer, such as multiple-choice responses.

Python Example:

    import anthropic

    # Initialize the Claude client
    client = anthropic.Anthropic()

    # Send a multiple-choice question and guide Claude's response
    message = client.messages.create(
        model="claude-3-5-sonnet-20240620",
        max_tokens=1,
        messages=[
            {"role": "user", "content": "What is the Greek for fear? (A) Arachnea, (B) Philosophia, (C) Phobia"},
            {"role": "assistant", "content": "The answer is ("}
        ]
    )

    # Print the response
    print(message)

Response example:

    {
      "id": "msg_01Q8Faay6S7QPTvEUUQARt7h",
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "text",
          "text": "C"
        }
      ],
      "model": "claude-3-5-sonnet-20240620",
      "stop_reason": "max_tokens",
      "stop_sequence": null,
      "usage": {
        "input_tokens": 50,
        "output_tokens": 1
      }
    }    

In this example, Claude is prompted with a multiple-choice question about the Latin name for ants, and the response is guided to produce a single character, “C,” indicating the correct choice.

Vision

Claude can process both text and images in a request. For images, you need to encode them in base64 and include them in the API call. Claude supports various image formats like JPEG, PNG, GIF, and WebP.

Python example:

    import anthropic
    import base64
    import httpx

    # Download and encode the image
    image_url = "https://en.wikipedia.org/wiki/Wolf#/media/File:Eurasian_wolf_2.jpg"
    image_media_type = "image/jpeg"
    image_data = base64.b64encode(httpx.get(image_url).content).decode("utf-8")

    # Initialize the Claude client and send an image for analysis
    client = anthropic.Anthropic()

    message = client.messages.create(
        model="claude-3-5-sonnet-20240620",
        max_tokens=1024,
        messages=[
            {
                "role": "user",
                "content": [
                    {
                        "type": "image",
                        "source": {
                            "type": "base64",
                            "media_type": image_media_type,
                            "data": image_data,
                        },
                    },
                    {"type": "text", "text": "What type of animal is in this image?"}
                ],
            }
        ],
    )

    # Print the response
    print(message)

Response example:

    {
      "id": "msg_01EcyWo6m4hyW8KHs2y2pei5",
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "text",
          "text": "This image shows a wolf, specifically a Eurasian wolf at Polar Zoo in Bardu, Norway. The wolf is shown walking on the snow, with its paws partially covered by the snow. The image is focused on capturing the general idea and features of a wolf."
        }
      ],
      "model": "claude-3-5-sonnet-20240620",
      "stop_reason": "end_turn",
      "stop_sequence": null,
      "usage": {
        "input_tokens": 2000,
        "output_tokens": 100
      }
    }

From API Calls to Real Workflows

Using the Claude API effectively goes beyond making individual API calls. In practice, most use cases require coordinating multiple steps into a complete workflow.

For example, a typical AI-powered workflow might include:

  • Collecting input (user query, document, or data source)
  • Structuring prompts for consistent output
  • Processing and formatting responses
  • Triggering actions (sending a message, updating a system, storing data)
  • Integrating with existing tools like CRMs, databases, or communication platforms

Managing these steps manually can quickly become complex—especially as workflows grow in scope and importance.

Instead of building everything from scratch, many teams look for ways to streamline this process.

Platforms like Obot make it easier to connect models like Claude with your existing tools and turn individual prompts into end-to-end workflows – without needing to manage every layer of infrastructure manually.

Turn Claude API Ideas Into Real Workflows with Obot

Claude’s API is powerful, but turning it into production-ready workflows takes time, infrastructure, and ongoing management.

Obot helps bridge that gap by giving you a flexible platform to:

  • Reduce the need for custom backend development
  • Integrate Claude with your existing tools
  • Build and deploy AI-powered workflows faster

FAQ

What is the Claude API?

The Claude API is Anthropic’s RESTful interface for accessing Claude models programmatically, letting developers integrate Claude into their own applications for tasks like text generation, document analysis, and autonomous agents.

How much does the Claude API cost?

Pricing is pay-as-you-go per million tokens, and varies by model: Claude Haiku 4.5 is the cheapest at $1/$5 per MTok (input/output), Claude Sonnet 5 is $2/$10 introductory through August 2026, and Claude Opus 5 is $5/$25. Additional charges apply for features like web search. Always check Anthropic’s current pricing page, since rates and introductory windows change.

Is the Claude API the same as a Claude Pro or Max subscription?

No. Pro, Max, and Team are flat-rate subscriptions for using Claude.ai or Claude Code with a personal account. The Claude API is a separate, pay-as-you-go product for building your own applications, billed per token rather than by subscription.

How do I get a Claude API key?

Create an account in the Claude Console, generate an API key from your account settings, and set it as the ANTHROPIC_API_KEY environment variable, or pass it directly to the SDK client when you initialize it in code.

What are the Claude API’s rate limits?

Limits are enforced per organization across requests per minute, input tokens per minute, and output tokens per minute, and your organization is placed on a usage tier automatically based on usage history. Exact numbers vary by model and tier and are visible in the Claude Console.

What are the Claude API’s rate limits?

Limits are enforced per organization across requests per minute, input tokens per minute, and output tokens per minute, and your organization is placed on a usage tier automatically based on usage history. Exact numbers vary by model and tier and are visible in the Claude Console.

Is the Claude API HIPAA compliant?

Not by default, and not automatically on every plan. Anthropic offers a HIPAA Business Associate Agreement for the Claude API and for certain Claude Enterprise configurations, but it requires a signed BAA, specific configuration such as zero data retention, and excludes some newer API features. Consumer plans (Free, Pro, Max, Team) aren’t covered. Treat “HIPAA options” as something you configure and contract for, not a default feature.

Can the Claude API connect to MCP servers?

Yes, through the built-in MCP connector, which lets you list remote, publicly reachable MCP servers directly in a Messages API request. This is separate from connecting MCP servers in Claude Code or Claude Desktop for interactive use.

Bottom Line

The Claude API gives you direct, programmatic access to Claude’s current models, from Haiku 4.5 for fast, cheap tasks to Opus 5 and Fable 5 for complex, long-running agentic work. Getting a basic integration running takes an afternoon. Turning that into a reliable, governed, production workflow across a team is where most of the real engineering effort goes, and that’s where a platform like Obot can take on the infrastructure so you can focus on the workflow itself.