> ## Documentation Index
> Fetch the complete documentation index at: https://docs.protodesk.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

> Read rate-limit headers, control concurrency, and recover from throttling.

Protodesk limits request bursts to protect API availability. Use the response headers to pace your integration and handle `429` responses even when normal traffic is well below the limit.

## Read the response headers

| Header | Meaning |
| - | - |
| `RateLimit-Limit` | Capacity of the applicable request bucket |
| `RateLimit-Remaining` | Whole request tokens left after this request |
| `RateLimit-Reset` | Seconds until the bucket refills to capacity if no more requests consume it |
| `Retry-After` | Seconds to wait before retrying a throttled request |

`RateLimit-Reset` is a duration, not a Unix timestamp. The bucket refills continuously; it is not a fixed minute-long window. Treat these headers as optional and keep a bounded backoff fallback when they are absent.

## Understand the budgets

The current server defaults separate reads, writes, and heavier operations:

| Class | Sustained refill | Burst capacity |
| - | - | - |
| Reads: `GET`, `HEAD`, and `OPTIONS` | 20 requests/second | 40 |
| Writes: other methods, except heavier routes | 10 requests/second | 20 |
| Heavier operations, including OAuth client registration | 2 requests/second | 5 |

These are server defaults, not a guaranteed throughput allocation. Follow the headers returned by the service and allow for additional provider or infrastructure limits.

Authenticated REST requests share a budget per bearer credential and request class. Requests without a bearer credential use the client IP. MCP and public OAuth routes use an IP budget; MCP requests sent with `POST` consume the write budget even when the requested tool only reads data. Multiple clients behind one public IP can therefore affect each other.

## Recover from a 429

1. Pause the affected work queue for at least `Retry-After` seconds when present.
2. Add a small random delay so queued requests do not all resume together.
3. Retry only if the operation is safe to repeat. Preserve the original idempotency key and request for supported writes.
4. Limit attempts and total elapsed time. Surface repeated throttling to the operator instead of retrying forever.

When `Retry-After` is absent, use exponential backoff with jitter. See [Errors and retries](/guides/errors-and-retries) for deciding which requests to repeat.

## Reduce unnecessary requests

Bound concurrent requests across workers that share a credential. Follow [cursor pagination](/guides/pagination), cache data that rarely changes, and use [webhooks](/guides/webhooks) for background updates. Poll message delivery with increasing intervals and a deadline rather than continuously refreshing it.

An HTTP `413` means the request body is too large; waiting will not fix it. Reduce the payload according to the endpoint's requirements. AI credit exhaustion and channel-provider restrictions also need their own resolution rather than an HTTP rate-limit retry loop.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.