Rate Limits
Understand and handle API rate limits
Rate limits keep the platform fair and stable for everyone. Every plan includes generous limits, and clients should handle 429s with automatic retries.
Plan limits
| Plan | Price | Requests/mo | Tokens/mo | Seats |
|---|---|---|---|---|
Free | $0 | 100 | 200K | 1 |
Starter | $15/mo | 1,000 | 1.5M | 1 |
Pro | $49/mo | 8,000 | 12M | 1 |
Team | $149/mo | 40,000 | 60M | 5 |
Enterprise | Custom | Custom | Custom | Unlimited |
Start immediately with 100 requests and 200K tokens per month. All 60+ models are available on every paid plan. Free/Starter/Pro are individual plans; Team/Enterprise are separate workspace plans for teams, not an upgrade path from Pro.
Two limits, not one
| Limit | Scope | What happens |
|---|---|---|
| Monthly quota (table above) | Per user / per team owner's plan | 429 with code: rate_limit_exceeded. No Retry-After — it clears at the start of the next month. |
| Short-term burst | Per client IP, not per plan | POST /v1/ai/*: 20 requests / minute. Every other endpoint: 100 / minute. 429 with a Retry-After header. |
A retry loop that only waits for Retry-After will spin against the monthly
quota. Check the status body: if code is rate_limit_exceeded and there is no
Retry-After, you are out of monthly quota — back off long, or upgrade.
Rate limit headers
Burst-limited responses carry the standard headers:
Successful chat requests also return your monthly-quota position:
Handling rate limits
Implement exponential backoff for retries · monitor rate limit headers proactively · cache responses when possible · batch requests where you can · contact support for higher limits
