DeepSeek API Peak Hours
Chinese AI lab offering R1 reasoning and V3 general-purpose models, accessible via API and chat.
No documented time-of-day peak window — congestion tracked via 'Server Busy' status
08–18 US/EU + Asia daytime (UTC)
18–24 Asia evening + US late (UTC)
DeepSeek's official support guidelines explicitly use a 'Server Busy' error state to mask cluster-wide capacity saturation. When the cluster hits peak load, the documented mitigation is to toggle off 'Deep Think' (R1) — reasoning models consume vastly more compute and hit the saturation first — and use the standard V3 model, which is allocated a separate, higher-capacity processing pool. DeepSeek's own concurrency limits (verified) are 500 for v4-pro and 2,500 for v4-flash. Specific peak-hour windows are not published — the visual blocks above illustrate global business-hour load.
GLM (Zhipu AI) API Peak Hours
Zhipu AI's GLM-5 / GLM-4.7 series — strong bilingual (Chinese/English) models, 1M-token context on advanced tiers.
Official peak window: 14:00–18:00 UTC+8 (= 06:00–10:00 UTC) — 3× quota multiplier
06–10 GLM official peak (14:00–18:00 UTC+8) — 3× quota multiplier
10–14 Asia late afternoon — 2× off-peak-daytime multiplier
14–18 Asia evening taper
18–22 Asia evening (2× multiplier)
Zhipu's official GLM developer documentation states that the platform applies a peak-hour quota multiplier: token usage is deducted at 3× the normal rate during the official 14:00–18:00 UTC+8 peak window (= 06:00–10:00 UTC) and 2× during other Asia daytime hours. Newer reasoning models (GLM-5.2, GLM-5-Turbo) are most affected. Concurrent requests are also hard-capped — even paid Max plans queue simultaneous workflows during peak.
Qwen (Alibaba Cloud) API Peak Hours
Alibaba's Qwen family of models — strong multilingual performance with Qwen-Max and Qwen-Plus tiers.
No documented time-of-day peak window — usage follows Alibaba Cloud infrastructure events
00–24 Event-driven (Double 11, CNY): quota reallocated to retail
Qwen API is integrated with Alibaba Cloud's broader infrastructure. Qwen does not document a fixed daily peak window — the verified impact is event-driven: during Double 11 (Singles' Day), Chinese New Year, and other Alibaba e-commerce events, model API quotas can drop up to ~70% for non-enterprise tiers as compute is re-allocated to retail search/recommendation workloads. The visual block above illustrates the event-driven nature, not a fixed clock-time pattern.
Moonshot (Kimi) API Peak Hours
Moonshot AI's Kimi K2 / K2.7 Code models — long-context specialists popular in the Chinese market.
Documented peak: 09:00–22:00 CST (= 01:00–14:00 UTC)
01–05 CST 09:00–13:00 (UTC) — documented peak
05–09 CST 13:00–17:00 (UTC) — documented peak
09–14 CST 17:00–22:00 (UTC) — documented peak
Moonshot's official Kimi Code + API documentation governs usage with a strict rolling 5-hour window plus weekly and monthly caps — when request frequency spikes during China high-traffic periods, rate limiting automatically freezes the quota until the 5-hour window rolls forward. Free tier (Tier 0) is hard-limited to 1 concurrent request + 3 RPM + 1.5M TPD, so any rapid clicking during errors exhausts the quota instantly. The visual windows above mirror the documented 09:00–22:00 CST peak.
Mistral API Peak Hours
French AI lab with Mistral Large, Small, and Codestral — strong European focus with EU-regulated API access.
No documented time-of-day multiplier — per-second rate limits only
No time-of-day congestion restrictions documented.
Mistral's La Plateforme API enforces per-second rate limits without a documented time-of-day quota multiplier. The visual block above does not imply specific peak windows — verified behaviour is per-second concurrency caps and 429 retries. On-premise self-hosted deployments are unaffected. Live uptime per surface area is published at the Mistral Status page (linked below).
Claude (Anthropic) API Peak Hours
Anthropic's Claude Opus / Sonnet / Haiku family.
Documented peak window: 13:00–19:00 UTC weekday (= 5 AM–11 AM PT / 1 PM–7 PM GMT)
13–19 Weekday peak (UTC) — 5-hour budget burns faster
19–23 Weekday evening US/EU overlap
23–13 Off-peak / weekends — limits reset faster
Anthropic's official policy (r/Anthropic AMA, July 2025): tier-based rate limits PLUS a 5-hour rolling session window PLUS a weekly cap (Aug 28, 2025 onward) — Max 20x users get 240–480 weekly Sonnet 4 hours and 24–40 weekly Opus 4 hours. The 5-hour budget burns significantly faster during US business hours (5 AM–11 AM PT / 1 PM–7 PM GMT) when compute demand spikes — Anthropic frames this as 'limited compute during peak hours', not a per-hour quota multiplier. Anthropic uses a queue-based delay system for Sonnet rather than outright 429 rejections; Opus is often unavailable for lower-tier accounts during peak. Pro ($20) and Max ($100/$200) plans all expose these windows to the same peak-hour compression.
Gemini (Google) API Peak Hours
Google's Gemini models — multimodal with generous free-tier quotas on AI Studio.
Documented recommendation: schedule heavy jobs off-peak + weekends
No time-of-day congestion restrictions documented.
Google's official Vertex AI documentation describes three PayGo tiers (Standard / Priority / Flex) and explicitly recommends that developers schedule token-heavy tasks and background workflows during off-peak hours and on weekends — Standard PayGo tiers experience heavily elevated throttling during high-demand daytime windows. Google AI Studio free tier is the first to throttle; Vertex AI paid tiers maintain guaranteed throughput. Gemini Flash models are less affected than Pro/Ultra at peak. The visual block above represents the documented baseline — no fixed clock-time peak is publicly published.
OpenAI (ChatGPT/GPT-5.x) API Peak Hours
OpenAI's GPT-5.x, o-series reasoning models, and the GPT-5.5 'mini' fallback for ChatGPT subscribers.
Documented rolling cap: 160 msgs / 3 hours (ChatGPT Plus); auto-downgrade to mini
No time-of-day congestion restrictions documented.
Official OpenAI Help Center (help.openai.com/en/articles/11909943): ChatGPT Plus and ChatGPT Go subscribers get up to 160 messages per 3-hour window with base models (GPT-5.5) — after the threshold is hit, chats are silently switched to the 'mini' variant of the model until the window resets. Plus tier uses an adaptive 'smart throttle' that tightens the cap during global business hours when server load spikes; users hitting the threshold are also returned temporary 'too many requests' cooldowns. GPT-4o mini / o4-mini maintain higher quotas than full-size models during the same window. Tier-1 free-tier API keys face the steepest reduction (90%+) at peak. The visual block represents the documented rolling-3-hour cap — no fixed clock-time peak is published by OpenAI.