Rate limits
Limits apply at two levels: per API key, and per account. Whichever is stricter wins.
Defaults
| Scope | Limit | Adjustable |
|---|---|---|
| Generation endpoints | 30 req / min | Yes, per key |
| Read endpoints (history, credits, providers) | 60 req / min | No |
| Login | 10 req / min | No |
| Concurrent generations | Plan-dependent | By upgrading |
Set a lower rate_limit_per_minute on any key at creation time to bound a specific integration below the account default.
Headers
Every response carries the current window state:
X-RateLimit-Limit: 30
X-RateLimit-Remaining: 24
X-RateLimit-Reset: 1754769600
Retry-After: 12 # only on 429
Read X-RateLimit-Remaining and slow down before you hit zero, rather than treating 429 as the signal.
Working within the limits
- Queue, don't spray. Push generation jobs onto a worker queue with bounded concurrency rather than firing them from request handlers.
- Back off with jitter. Retrying every client at exactly
Retry-Afterre-synchronises the stampede. Add randomness — see errors. - Separate keys per workload. A batch job on its own key with its own limit cannot starve interactive traffic.
- Batch at the prompt level. Ten classifications in one call beats ten calls, and usually costs fewer tokens.
Long-running requests
Text and image calls complete inline; allow a client timeout of at least 120 seconds. Video and speech return immediately and are polled, so they never hold a connection open. Do not set an aggressive timeout on generation endpoints — an abandoned request still consumes upstream capacity and settles against your balance.
Higher limits
Sustained throughput above the defaults is available on higher plans. If you need a specific ceiling, get in touch with your expected requests per minute, average token counts and which models you intend to call.