429 responses even when normal traffic is well below the limit.
Read the response headers
RateLimit-Reset is a duration, not a Unix timestamp. The bucket refills continuously; it is not a fixed minute-long window. Treat these headers as optional and keep a bounded backoff fallback when they are absent.
Understand the budgets
The current server defaults separate reads, writes, and heavier operations:
These are server defaults, not a guaranteed throughput allocation. Follow the headers returned by the service and allow for additional provider or infrastructure limits.
Authenticated REST requests share a budget per bearer credential and request class. Requests without a bearer credential use the client IP. MCP and public OAuth routes use an IP budget; MCP requests sent with
POST consume the write budget even when the requested tool only reads data. Multiple clients behind one public IP can therefore affect each other.
Recover from a 429
- Pause the affected work queue for at least
Retry-Afterseconds when present. - Add a small random delay so queued requests do not all resume together.
- Retry only if the operation is safe to repeat. Preserve the original idempotency key and request for supported writes.
- Limit attempts and total elapsed time. Surface repeated throttling to the operator instead of retrying forever.
Retry-After is absent, use exponential backoff with jitter. See Errors and retries for deciding which requests to repeat.
Reduce unnecessary requests
Bound concurrent requests across workers that share a credential. Follow cursor pagination, cache data that rarely changes, and use webhooks for background updates. Poll message delivery with increasing intervals and a deadline rather than continuously refreshing it. An HTTP413 means the request body is too large; waiting will not fix it. Reduce the payload according to the endpoint’s requirements. AI credit exhaustion and channel-provider restrictions also need their own resolution rather than an HTTP rate-limit retry loop.