Outerstellar Outerstellar

AI Providers

AI API Congestion Calendar

Track AI provider rate limits, token-per-minute quotas, and peak-hour throttling across DeepSeek, Claude, OpenAI, GLM, Qwen, Moonshot Kimi, Mistral, and Google Gemini. Each provider's window is sourced to its official docs, an AMA, or a community observation — see the Sources list on every card.

Your time: --:--:-- UTC · --:-- UTC · ·
Overlay regional peak hours:

Live AI Provider Congestion Schedule (24-Hour UTC)

Congestion Levels: Low — Normal quotas Medium — Slightly reduced High — Significantly reduced Critical — Heavily restricted
DeepSeek
--
GLM (Zhipu AI)
--
Qwen (Alibaba Cloud)
--
Moonshot (Kimi)
--
Mistral
--
Claude (Anthropic)
--
Gemini (Google)
--
OpenAI (ChatGPT/GPT-5.x)
--
0h
3h
6h
9h
12h
15h
18h
21h

Hover any segment to see the precise UTC hour range, congestion level, and throttle percentage. The vertical line marks the current UTC time across all providers.

This chart is community-maintained.

The hourly windows above are best-effort estimates compiled from public documentation, official AMAs, and developer observations. They can be wrong, and providers change quotas without notice. If you spot a stale window, a missing provider, or have a verified correction — please send it to the community:

AI Provider Rate Limits & Peak Hours

Each provider has an official rate-limit policy — see the linked docs in Sources on every card. The 24-hour colored windows reflect when each provider's compute and quota budgets tighten most aggressively, with the exact mechanism documented per-provider.

DeepSeek API Peak Hours

Chinese AI lab offering R1 reasoning and V3 general-purpose models, accessible via API and chat.

No documented time-of-day peak window — congestion tracked via 'Server Busy' status
08–18 US/EU + Asia daytime (UTC) 18–24 Asia evening + US late (UTC)

DeepSeek's official support guidelines explicitly use a 'Server Busy' error state to mask cluster-wide capacity saturation. When the cluster hits peak load, the documented mitigation is to toggle off 'Deep Think' (R1) — reasoning models consume vastly more compute and hit the saturation first — and use the standard V3 model, which is allocated a separate, higher-capacity processing pool. DeepSeek's own concurrency limits (verified) are 500 for v4-pro and 2,500 for v4-flash. Specific peak-hour windows are not published — the visual blocks above illustrate global business-hour load.

GLM (Zhipu AI) API Peak Hours

Zhipu AI's GLM-5 / GLM-4.7 series — strong bilingual (Chinese/English) models, 1M-token context on advanced tiers.

Official peak window: 14:00–18:00 UTC+8 (= 06:00–10:00 UTC) — 3× quota multiplier
06–10 GLM official peak (14:00–18:00 UTC+8) — 3× quota multiplier 10–14 Asia late afternoon — 2× off-peak-daytime multiplier 14–18 Asia evening taper 18–22 Asia evening (2× multiplier)

Zhipu's official GLM developer documentation states that the platform applies a peak-hour quota multiplier: token usage is deducted at 3× the normal rate during the official 14:00–18:00 UTC+8 peak window (= 06:00–10:00 UTC) and 2× during other Asia daytime hours. Newer reasoning models (GLM-5.2, GLM-5-Turbo) are most affected. Concurrent requests are also hard-capped — even paid Max plans queue simultaneous workflows during peak.

Qwen (Alibaba Cloud) API Peak Hours

Alibaba's Qwen family of models — strong multilingual performance with Qwen-Max and Qwen-Plus tiers.

No documented time-of-day peak window — usage follows Alibaba Cloud infrastructure events
00–24 Event-driven (Double 11, CNY): quota reallocated to retail

Qwen API is integrated with Alibaba Cloud's broader infrastructure. Qwen does not document a fixed daily peak window — the verified impact is event-driven: during Double 11 (Singles' Day), Chinese New Year, and other Alibaba e-commerce events, model API quotas can drop up to ~70% for non-enterprise tiers as compute is re-allocated to retail search/recommendation workloads. The visual block above illustrates the event-driven nature, not a fixed clock-time pattern.

Moonshot (Kimi) API Peak Hours

Moonshot AI's Kimi K2 / K2.7 Code models — long-context specialists popular in the Chinese market.

Documented peak: 09:00–22:00 CST (= 01:00–14:00 UTC)
01–05 CST 09:00–13:00 (UTC) — documented peak 05–09 CST 13:00–17:00 (UTC) — documented peak 09–14 CST 17:00–22:00 (UTC) — documented peak

Moonshot's official Kimi Code + API documentation governs usage with a strict rolling 5-hour window plus weekly and monthly caps — when request frequency spikes during China high-traffic periods, rate limiting automatically freezes the quota until the 5-hour window rolls forward. Free tier (Tier 0) is hard-limited to 1 concurrent request + 3 RPM + 1.5M TPD, so any rapid clicking during errors exhausts the quota instantly. The visual windows above mirror the documented 09:00–22:00 CST peak.

Mistral API Peak Hours

French AI lab with Mistral Large, Small, and Codestral — strong European focus with EU-regulated API access.

No documented time-of-day multiplier — per-second rate limits only

No time-of-day congestion restrictions documented.

Mistral's La Plateforme API enforces per-second rate limits without a documented time-of-day quota multiplier. The visual block above does not imply specific peak windows — verified behaviour is per-second concurrency caps and 429 retries. On-premise self-hosted deployments are unaffected. Live uptime per surface area is published at the Mistral Status page (linked below).

Claude (Anthropic) API Peak Hours

Anthropic's Claude Opus / Sonnet / Haiku family.

Documented peak window: 13:00–19:00 UTC weekday (= 5 AM–11 AM PT / 1 PM–7 PM GMT)
13–19 Weekday peak (UTC) — 5-hour budget burns faster 19–23 Weekday evening US/EU overlap 23–13 Off-peak / weekends — limits reset faster

Anthropic's official policy (r/Anthropic AMA, July 2025): tier-based rate limits PLUS a 5-hour rolling session window PLUS a weekly cap (Aug 28, 2025 onward) — Max 20x users get 240–480 weekly Sonnet 4 hours and 24–40 weekly Opus 4 hours. The 5-hour budget burns significantly faster during US business hours (5 AM–11 AM PT / 1 PM–7 PM GMT) when compute demand spikes — Anthropic frames this as 'limited compute during peak hours', not a per-hour quota multiplier. Anthropic uses a queue-based delay system for Sonnet rather than outright 429 rejections; Opus is often unavailable for lower-tier accounts during peak. Pro ($20) and Max ($100/$200) plans all expose these windows to the same peak-hour compression.

Gemini (Google) API Peak Hours

Google's Gemini models — multimodal with generous free-tier quotas on AI Studio.

Documented recommendation: schedule heavy jobs off-peak + weekends

No time-of-day congestion restrictions documented.

Google's official Vertex AI documentation describes three PayGo tiers (Standard / Priority / Flex) and explicitly recommends that developers schedule token-heavy tasks and background workflows during off-peak hours and on weekends — Standard PayGo tiers experience heavily elevated throttling during high-demand daytime windows. Google AI Studio free tier is the first to throttle; Vertex AI paid tiers maintain guaranteed throughput. Gemini Flash models are less affected than Pro/Ultra at peak. The visual block above represents the documented baseline — no fixed clock-time peak is publicly published.

OpenAI (ChatGPT/GPT-5.x) API Peak Hours

OpenAI's GPT-5.x, o-series reasoning models, and the GPT-5.5 'mini' fallback for ChatGPT subscribers.

Documented rolling cap: 160 msgs / 3 hours (ChatGPT Plus); auto-downgrade to mini

No time-of-day congestion restrictions documented.

Official OpenAI Help Center (help.openai.com/en/articles/11909943): ChatGPT Plus and ChatGPT Go subscribers get up to 160 messages per 3-hour window with base models (GPT-5.5) — after the threshold is hit, chats are silently switched to the 'mini' variant of the model until the window resets. Plus tier uses an adaptive 'smart throttle' that tightens the cap during global business hours when server load spikes; users hitting the threshold are also returned temporary 'too many requests' cooldowns. GPT-4o mini / o4-mini maintain higher quotas than full-size models during the same window. Tier-1 free-tier API keys face the steepest reduction (90%+) at peak. The visual block represents the documented rolling-3-hour cap — no fixed clock-time peak is published by OpenAI.

Understanding AI Provider Congestion Levels

Low
Provider API returns full advertised quotas. No throttling; suitable for production batch jobs, training pipelines, and CI workloads.
Medium
Tokens-per-minute (TPM) and requests-per-minute (RPM) reduced by 25–35%. Acceptable for interactive workloads; expect occasional 429 retry-after headers.
High
Quota cut by 45–60%. Requests may queue. Avoid long-context or batch completions; switch to alternate providers if latency budgets are tight.
Critical
Quota cut by 65–80% or higher. Lower-tier API keys may be denied outright. Only enterprise or paid-tier accounts maintain throughput.

AI Provider Throttling — Key Questions

When is the best time to call the DeepSeek API to avoid rate limits?

DeepSeek's lowest-congestion window is roughly 22:00–11:00 UTC (overnight Asia through European morning). Avoid the 13:00–17:00 UTC wave when China daytime overlap US morning. DeepSeek's documented mitigation for the "Server Busy" state is to turn off "Deep Think" (R1) — reasoning requests hit the saturation first — and fall back to V3, which sits on a higher-capacity processing pool.

Why does OpenAI ChatGPT Plus silently downgrade me to a "mini" model?

ChatGPT Plus and ChatGPT Go are officially capped at 160 messages per 3-hour window with the base model (GPT-5.5). Once you cross that line, chats are silently switched to the model's "mini" variant until the window resets — this is an OpenAI Help Center feature, not a bug. The adaptive throttle tightens further during global business-hour load spikes. To preserve full-model access during crunch time, switch to API key billing or downgrade usage.

When does Anthropic Claude hit the "5-hour window" wall hardest?

Anthropic frames the rolling 5-hour session window in conjunction with a weekly cap (introduced 28 Aug 2025). Community observations and one specific 430-upvote complaint thread indicate the 5-hour budget burns significantly faster during US business hours, roughly 13:00–19:00 UTC (5–11 AM Pacific / 1–7 PM GMT). Anthropic frames this as limited compute during peak hours rather than a stated time-based multiplier. Max 20x users get 240–480 Sonnet 4 hours/week and 24–40 Opus 4 hours/week.

When does Zhipu's GLM apply its 3× peak-hour quota multiplier?

Zhipu's official GLM developer documentation states the platform applies a peak-hour token quota multiplier: usage is deducted at 3× normal rate during the official 14:00–18:00 UTC+8 window (≈ 06:00–10:00 UTC) and 2× during other Asia daytime hours. The flagship GLM-5.2 and GLM-5-Turbo reasoning models are the most affected. Concurrent request caps are additionally hard-limited even on Max paid plans.

How can I avoid AI provider throttling without changing providers?

Three reliable strategies: (1) batch off-peak — schedule large jobs in each provider's lowest-quota-multiplier window; (2) downshift the model — GPT-4o mini, Claude Haiku, Gemini Flash, and GLM Flash tiers are dramatically less throttled; (3) add retry-after jitter — respect the Retry-After response header and back off exponentially to avoid compounding the queue.