Claude API 429 Too Many Requests: Why It Happens and How to Fix It

Marcus
Marcus
Proxy Network Analyst

A Claude API 429 Too Many Requests response means the request reached Anthropic but could not be served under the current usage limits. That does not always mean your application simply sent too many requests in one minute. A 429 can come from request or token rate limits, a sudden acceleration in traffic, a usage-tier monthly spend cap, or a separate Fast Mode limit.

The first troubleshooting step is therefore not to retry blindly. Read the error body and response headers, especially retry-after, then decide whether the request should wait, reduce throughput, or stop retrying until account-level access resumes.

Quick Answer

For a normal Claude API rate-limit 429, read the retry-after header and wait at least that long before sending the next request. If the response has no retry-after header and the Messages API error contains error.details.error_code = enforced_spend_limit_reached, the organization has reached its usage-tier monthly spend cap. Repeated retries will continue to fail until access resumes or the limit is raised. A proxy or a different API key does not increase an organization-level Claude API rate limit.

Key Takeaways
  • Claude API 429 responses use the rate_limit_error error type.
  • Messages API limits can apply to requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM).
  • A normal rate-limit 429 includes retry-after; retries sent earlier will fail.
  • A usage-tier spend-cap 429 has no retry-after and can be identified by enforced_spend_limit_reached on the Messages API.
  • The Anthropic Python SDK retries 429 responses twice by default, so custom retry loops should account for SDK-level retries.
  • Short bursts and sudden traffic growth can trigger 429 responses even when long-term average traffic looks acceptable.

What Does Claude API 429 Actually Mean?

Anthropic classifies HTTP 429 as rate_limit_error. For the Messages API, normal limits are measured in requests per minute, input tokens per minute, and output tokens per minute. Anthropic also documents acceleration-limit 429s when an organization's traffic increases sharply, and usage-tier spend caps can return the same 429 status with a different diagnostic pattern.

Possible Cause What to Look For Should You Retry Immediately? First Action
RPM limit 429 with retry-after; request-related limit headers No Wait for retry-after and reduce request bursts
ITPM limit 429 with token-limit headers No Reduce uncached input throughput or wait for capacity
OTPM limit 429 while output throughput is high No Throttle output-heavy workloads and wait for reset
Acceleration limit 429 after a sharp increase in traffic No Ramp traffic more gradually
Usage-tier monthly spend cap No retry-after; enforced_spend_limit_reached on Messages API No; repeated retries will fail Wait for access to resume or request a higher limit
Fast Mode limit 429 with retry-after and Fast Mode rate-limit context No Wait for capacity or use standard speed when appropriate
Table 1: Claude API 429 responses need to be classified before deciding whether and how to retry.

The important distinction is that the HTTP status alone is not enough. Anthropic's current rate-limit documentation specifies different recovery behavior for a normal rate limit and a monthly spend-cap 429.

Claude Console Rate limits page showing requests per minute, input tokens per minute, and output tokens per minute
Figure 1: An example Claude Console Rate limits view showing separate request, input-token, and output-token limits. Actual limits depend on the organization, workspace, model, and tier.

Check retry-after Before Retrying

For a normal Messages API rate-limit 429, Anthropic returns a retry-after response header with the number of seconds to wait. Requests sent before that delay expires will fail again. Treat the server-provided delay as the minimum retry time instead of replacing it with a hard-coded sleep value.

429 retry decision
  1. Read retry-after.
  2. If it exists, wait at least that long before the next attempt.
  3. If it is missing, inspect the error body before assuming the header was lost.
  4. If the Messages API error contains enforced_spend_limit_reached, stop retrying until account-level access changes.
  5. If neither condition explains the failure, log the request ID and the relevant rate-limit headers for further diagnosis.

This prevents a costly failure mode: an application receives a spend-cap 429, assumes every 429 is transient, then keeps retrying a request that cannot succeed yet.

Read the Claude Rate-Limit Headers

Claude API responses include rate-limit headers that show the active limit, remaining capacity, and reset time. These values are more useful than copying a fixed quota from an old tutorial because Anthropic can apply different limits by tier, model class, workspace, or service configuration.

Header What It Tells You
retry-after Seconds to wait before retrying a normal rate-limited request
anthropic-ratelimit-requests-limit Maximum requests allowed within the active request-limit period
anthropic-ratelimit-requests-remaining Remaining request capacity before the request limit is reached
anthropic-ratelimit-requests-reset RFC 3339 time when request capacity is fully replenished
anthropic-ratelimit-input-tokens-* Input-token limit, remaining capacity, and reset time
anthropic-ratelimit-output-tokens-* Output-token limit, remaining capacity, and reset time
anthropic-ratelimit-tokens-* The most restrictive token limit currently in effect
Table 2: Claude rate-limit headers reveal the active constraint and when capacity becomes available again.

The reset headers are timestamps, while retry-after is a delay in seconds. When a normal 429 includes both, use retry-after for the next retry and keep the reset values for monitoring and capacity planning.

Catch Claude API 429 Errors in Python

The current Anthropic Python SDK maps HTTP 429 responses to anthropic.RateLimitError. During diagnosis, temporarily disabling SDK retries makes it easier to inspect the first 429 response instead of waiting for the client's built-in retry attempts.

import anthropic
from anthropic import Anthropic

client = Anthropic(max_retries=0)

try:
    message = client.messages.create(
        model="claude-opus-5-5",
        max_tokens=64,
        messages=[
            {"role": "user", "content": "Reply only with OK"}
        ],
    )
    print(message.content[0].text)

except anthropic.RateLimitError as exc:
    headers = exc.response.headers
    body = exc.response.json()

    retry_after = headers.get("retry-after")
    request_id = headers.get("request-id")
    error = body.get("error", {})
    details = error.get("details", {}) or {}

    print("status:", exc.status_code)
    print("error type:", error.get("type"))
    print("retry-after:", retry_after)
    print("request-id:", request_id)
    print("detail code:", details.get("error_code"))

Do not log the API key, authorization header, or full credentials alongside this diagnostic output. The request ID, error type, retry-after, detail code, timestamp, and selected rate-limit headers are normally enough to identify the branch.

If the code is running against an older Anthropic SDK release, update the SDK before relying on current exception behavior and transport details. The current Python SDK documentation lists RateLimitError for HTTP 429 and documents the retry controls used below.

Understand Anthropic SDK Automatic Retries

Anthropic's official SDKs retry several transient failures automatically. In the Python SDK, connection errors, HTTP 408, 409, 429, and server errors at 500 or above are retried two times by default with a short exponential backoff. Anthropic also states that the SDK honors retry-after when the header is present.

This matters when application code also implements its own retry loop. Five application-level attempts do not necessarily mean five HTTP attempts if every call can trigger SDK retries underneath.

from anthropic import Anthropic

# Useful while diagnosing the first raw 429 response.
client = Anthropic(max_retries=0)

Use max_retries=0 as a troubleshooting control, not as a universal production recommendation. Production clients usually benefit from bounded retries, but retries should remain aware of server-provided delays and should not multiply uncontrollably across SDK, queue, worker, and application layers.

Rate Limit vs Monthly Spend Cap

A normal rate limit and a usage-tier monthly spend cap can both return HTTP 429, but they require different responses.

Normal rate-limit 429

A normal Messages API limit returns rate_limit_error with retry-after. Wait for the specified delay, then retry at a lower or smoother throughput.

Usage-tier spend-cap 429

When an organization reaches its service-configured monthly usage-tier spend cap, Anthropic pauses API usage and returns HTTP 429. On the Messages API, the error body includes:

{
  "type": "error",
  "error": {
    "type": "rate_limit_error",
    "details": {
      "error_code": "enforced_spend_limit_reached"
    }
  },
  "request_id": "req_..."
}

This 429 does not include retry-after. Anthropic explicitly notes that SDK automatic retries also fail until access resumes. The correct action is to check the organization's Billing and Rate limits pages, wait for the stated access-resumption time, or request a higher limit when appropriate.

Do not confuse this with a spend limit that your own organization administrator sets below the tier cap. Anthropic currently documents those user-configured organization or workspace spend-limit failures as HTTP 400 invalid_request_error, not the usage-tier 429 described above. Claude Code workspace limits have separate behavior and can return 429 in specific cases.

RPM vs ITPM vs OTPM

Messages API throughput is not controlled by one generic “requests per minute” number. Anthropic applies separate request and token constraints, so reducing only concurrency may not solve a token-throughput bottleneck.

Limiter What It Measures Typical Fix
RPM Requests per minute Queue requests, smooth bursts, reduce concurrency
ITPM Input tokens per minute Reduce repeated uncached context, use prompt caching where appropriate
OTPM Actual output tokens generated per minute Throttle output-heavy jobs or distribute work more evenly
Table 3: RPM, ITPM, and OTPM can each become the active Claude API constraint.

Prompt caching can reduce ITPM pressure

For most current Claude models, cached input tokens read from prompt cache do not count toward ITPM. Tokens written to cache and uncached input still count. Repeated system instructions, tool definitions, large context documents, and conversation history are therefore good candidates for prompt caching when the workload naturally reuses them.

One exception matters: Anthropic currently documents different cache-read rate-limit behavior for Claude Haiku 3.5. Avoid assuming that one caching rule applies to every model generation.

max_tokens does not set OTPM usage

Anthropic evaluates OTPM using the output tokens actually generated. The max_tokens request parameter does not itself count toward the OTPM limit, so lowering max_tokens is not a guaranteed fix unless it also reduces real output generation.

Avoid Acceleration-Limit 429s

A workload can encounter 429 responses after a sharp increase in usage even when the long-term average appears to fit within published limits. Anthropic calls these acceleration limits and recommends ramping traffic gradually while maintaining more consistent usage patterns.

This also explains why a service can behave normally during a low-volume test and then return 429 errors immediately after a large deployment or job fan-out. Build a controlled ramp instead of moving from a handful of requests to full production concurrency in one step.

Short bursts matter for ordinary rate limits too. Anthropic notes that a nominal per-minute request allowance can be enforced over shorter intervals, so a burst concentrated into one or two seconds can fail even if the minute-level average looks acceptable.

Reduce 429s Without Slowing Everything Down

The goal is not to serialize every API call. A better design keeps throughput high while preventing bursts from exceeding the active request or token bucket.

Practical controls
  • Queue bursty work: release requests at a controlled rate instead of starting every task simultaneously.
  • Limit concurrency by workload: separate short interactive calls from large context or output-heavy jobs.
  • Watch remaining headers: throttle before *-remaining reaches zero.
  • Use prompt caching for repeated context: reduce uncached input pressure where the model and workload support it.
  • Use Message Batches for suitable asynchronous jobs: batch processing has separate limits and can be a better fit for non-interactive bulk work.
  • Ramp new traffic gradually: avoid sudden step changes that can trigger acceleration limits.

Rate limits are applied separately by model class, so capacity planning should be based on the models actually used by the workload rather than one global assumption.

Claude API 429 vs 529

HTTP 429 and 529 can both look like temporary capacity problems, but Anthropic assigns them different meanings. A 429 is a rate_limit_error tied to organization usage, spend, acceleration, or a service-specific rate limit. A 529 is overloaded_error, meaning the API is temporarily overloaded.

Question 429 529
Official error type rate_limit_error overloaded_error
Typical scope Your organization, workspace, usage pattern, or service-specific limit Anthropic API capacity
First check Error details, retry-after, rate-limit headers, Console Usage Retry behavior and service status if the issue persists
Does changing a proxy raise the limit? No No
Table 4: Claude API 429 and 529 require different diagnoses even though both may be retryable in some situations.

Anthropic's current API error documentation also notes that official SDKs retry transient failures, including rate limits and 5xx errors, twice by default.

When a Proxy Is Not the Fix

If Anthropic returns a structured HTTP 429 with a request ID, the request has already reached the API. Changing an HTTP or SOCKS5 proxy does not increase the organization's Claude API rate limit, monthly spend cap, or token capacity.

If the actual symptom is a proxy authentication error, TLS problem, connection timeout, or a request that never receives an Anthropic HTTP response, use the Claude API proxy setup and transport checks instead. If the response is HTTP 401 or 403, move to the Claude API authentication and permission checks.

Check Claude Console Before Changing Code

The Claude Console Usage page provides separate rate-limit charts for input and output tokens. Anthropic documents the input chart as showing hourly maximum uncached input tokens per minute, the current input-token rate limit, and cache rate. The output chart shows hourly maximum output tokens per minute and the current output-token limit.

Claude Console Usage page showing token usage and rate-limited requests
Figure 2: Claude Console Usage can show requests blocked by rate limits alongside API token usage.

Use those charts together with the response headers. The headers tell you what happened to a specific request; the Console shows whether the issue is a recurring capacity pattern.

Recommended troubleshooting order
  1. Confirm the HTTP status is 429.
  2. Record the request ID.
  3. Read retry-after and the rate-limit headers.
  4. Inspect the error body for enforced_spend_limit_reached.
  5. Check Claude Console Usage and Rate limits.
  6. Only then adjust queues, concurrency, caching, model distribution, or account limits.

Frequently Asked Questions

What causes a Claude API 429 error?
A Claude API 429 can be caused by request-per-minute limits, input- or output-token limits, a sharp acceleration in organization traffic, a usage-tier monthly spend cap, or a service-specific limit such as Fast Mode. Read the response headers and error body to identify the branch.
How long should I wait after a Claude API 429?
For a normal rate-limit 429, wait at least the number of seconds in the retry-after header. Earlier retries will fail. If the 429 is a usage-tier spend-cap response, retry-after is absent and ordinary retries will not restore access.
Why does my Claude 429 response have no retry-after header?
On the Messages API, a usage-tier monthly spend-cap 429 does not include retry-after. Check whether error.details.error_code equals enforced_spend_limit_reached. Also consider whether an intermediary removed a header, but do not assume that before checking the error body.
What is the difference between RPM, ITPM, and OTPM?
RPM limits request count, ITPM limits input-token throughput, and OTPM limits output-token throughput. A workload can stay below RPM and still hit a token limit if requests contain large uncached prompts or generate substantial output.
Does the Anthropic Python SDK automatically retry 429 errors?
Yes. The current Python SDK retries 429 responses two times by default with a short exponential backoff, and Anthropic states that official SDKs honor retry-after when it is present. Use max_retries to change or temporarily disable that behavior.
Can prompt caching reduce Claude API rate-limit errors?
It can reduce ITPM pressure when a workload repeatedly reuses cacheable context. For most current Claude models, cached input tokens read from prompt cache do not count toward ITPM, while uncached input and cache-creation tokens still do. Check model-specific behavior before relying on this optimization.
Can changing an API key or proxy fix a Claude API 429?
Not when the 429 comes from an organization-level rate limit or usage-tier spend cap. A different network route does not increase API quota. Use proxy troubleshooting only when the request is failing at the connection or transport layer instead of returning a structured Anthropic 429.
What is the difference between Claude API 429 and 529?
HTTP 429 is rate_limit_error and generally relates to organization usage, spend, acceleration, or a service-specific limit. HTTP 529 is overloaded_error and indicates that the Anthropic API is temporarily overloaded.
Why am I still getting 429 even though my average traffic is below the limit?
Limits can be enforced over shorter intervals, so bursts may fail even when the minute-level average appears acceptable. Anthropic also documents acceleration-limit 429s after sharp increases in organization traffic. Smooth the request rate and ramp new workloads gradually.

Final Thoughts

A Claude API 429 should be treated as a classification problem before it becomes a retry problem. Start with retry-after, the rate-limit headers, the request ID, and the error details. A normal rate-limit response usually needs controlled waiting and smoother throughput; a usage-tier spend-cap 429 needs an account-level change or time for access to resume.

Once the cause is clear, optimize the correct layer: queue bursts for RPM, reduce uncached repeated context for ITPM, control output-heavy workloads for OTPM, ramp traffic gradually for acceleration limits, and avoid retry loops when Anthropic explicitly indicates that retries cannot succeed yet.

About the author
View all articles
Marcus
Marcus
Proxy Network Analyst

Marcus is a network infrastructure analyst specializing in proxy configuration, IP routing, browser connectivity, and network troubleshooting. His work focuses on diagnosing HTTP/SOCKS proxy connections, authentication failures, DNS behavior, firewall rules, and IP routing across browser and automation environments.

Service areas
Proxy Testing , IP Diagnostics,Network Troubleshooting & Reliability

You may be interested in

Perplexity API Key Not Working? Diagnose 401, 403, and Access Errors

Perplexity API Key Not Working? Diagnose 401, 403, and Access Errors

If a Perplexity API key is not working, diagnose the request that actually ran rather than using a successful browser login as proof that API access is healthy. A 401, 403, 429, 5xx, and a DNS or TLS failure belong to different troubleshooting branches. Preserve the HTTP status, error body, endpoint, runtime, and timestamp before retrying. Then reproduce the problem with the smallest documented request so application code, streaming, SDK wrappers, gateways, and retry middleware do not hide the failing layer. Quick Answer When a Perplexity API key is not working, verify the official endpoint, Bearer authorization header, key source,...

Marcus

Marcus

Proxy Network Analyst

Claude API proxy setup in Python with a proxy server between Python code and the Claude API

How to Use a Proxy with Claude API in Python

A Python application that calls the Claude API normally uses the network route available to the process that runs it. When you need a specific outbound route for development, fixed-egress testing, or an approved network environment, the Anthropic Python SDK can send requests through an explicit proxy instead of relying on the machine's default connection. The current Anthropic Python SDK uses httpx2 for its HTTP layer and lets you customize that layer with DefaultHttpxClient. Anthropic directly documents an HTTP proxy configuration, while HTTPX2 also provides optional SOCKS proxy support. That makes it possible to use either an HTTP proxy or...

Clark

Clark

IPWeb Technical Researcher

Claude API 401 and 403 errors troubleshooting cover with API request panel, key icon, account card, and security shield

Claude API Authentication Errors: How to Fix 401 and 403 Responses

A Claude API authentication error is easier to diagnose when you separate authentication, billing, permission, rate-limit, and transport failures before retrying. A 401 authentication_error and a 403 permission_error are different decisions, and neither should be treated as a generic “blocked API” response. Start with the smallest valid Claude API request, preserve the HTTP status, error type, request ID, workspace context, and timestamp, and keep the API key itself out of logs, screenshots, chat, and support tickets. Quick Answer For Claude API 401 and 403 errors, verify the authentication header, API-key state, workspace selection, and permission context before changing the network....

Marcus

Marcus

Proxy Network Analyst

Ready to scale your data operations?
Join 10,000+ teams using IPWeb to power their web data collection. Start free today.

Strictly anti-abuse

Fraud, automated operation, and unauthorized use are prohibited.

Enterprise-level services

For legitimate commercial and technical use cases only

Risk control and restrictions

Abnormal behavior may trigger service restrictions or termination.

Compliance data use

Data acquisition and use must comply with relevant regulations.

Privacy protection first

The collection or misuse of sensitive personal information is strictly prohibited.

All services are subject to《the Usage Policy》