Qwen3.7 Max is Alibaba's flagship - and as of July 2026 the highest-ranked Chinese model on the Artificial Analysis Intelligence Index at 56.6 (global #5, domestic #1), with a 1M token context, 35 hours of unattended autonomous execution, and an API at $2.50 in / $7.50 out per million tokens - roughly a third of what Claude Opus 4.7 charges on output. If you've been writing off Chinese labs as "good enough for the domestic market," this is the model that should make you stop. Here's the unvarnished deep dive: what it does, what it costs, where it loses, and whether you should care.
What Is Qwen3.7 Max?
Vendor: Alibaba Cloud (阿里云), built by the Qwen team at the Tongyi Lab. Released May 20, 2026 at the Alibaba Cloud Summit, as the top of the Qwen3.7 family (which also ships a cheaper, vision-capable Qwen3.7-Plus). The official positioning is "面向智能体时代的新一代旗舰模型" - an agent-era flagship. It runs on a trillion-parameter MoE architecture, carries a 1M token context window, and supports dual thinking/non-thinking modes so you can trade depth for latency.
Two details worth flagging. First, it went from public unveiling to live API on Bailian (Alibaba Cloud's model platform) in two days - May 22. Second, it's API-only. Previous Qwen models shipped open weights; this one doesn't. The open-source community grumbled, and fairly - but the capability leap is the reason they closed it.
Core Capabilities
No hand-waving - here are the numbers, cross-checked against Artificial Analysis, Arena, and Alibaba's official release notes from July 2026, not vendor marketing.
- Artificial Analysis Intelligence Index: 56.6 - a 4.8-point jump over Qwen3.6 Max Preview (51.8), and ahead of Gemini 3.5 Flash (55.3).
- MMLU-ProX: 87%, #1 globally - ahead of Claude Opus 4.5 (85.7%). Multilingual professional knowledge, and Qwen3.7 Max tops the whole board.
- GPQA Diamond: 92.2 - frontier-grade reasoning.
- Terminal-Bench 2.0 Terminus: 72.3 - domestic SOTA, multiple sub-scores (SWE-Pro 60.4, SWE-Multilingual 78.4, SciCode 52.7)刷新全球最高分.
- MCP-Mark: 64.6. Deep-Planning: 63.1. SpreadSheetBench: 84.5. General-agent and office-automation benchmarks, all top-tier.
- Kernel Bench L3: 2.37x median speedup, 100% pass rate. It optimizes GPU kernels on its own.
- WMT24++ 85.9, MAXIFE 89.3, IFBench 81.2 - multilingual translation and instruction-following, global lead.
The headline demo, though, isn't a benchmark. On a brand-new chip platform, Qwen3.7 Max ran for 35 hours straight, made 1,158 tool calls, and self-evolved a GPU kernel to 10x the original inference speed - no human in the loop. That's not a score; that's a proof of existence for long-horizon agents. Most frontier models tap out well before that.
Honest caveats: it's text-only. No vision. (Vision goes to Qwen3.7-Plus.) Knowledge cutoff is mid-2025, no real-time web access. And past 500K tokens the middle of the context starts to soften - the million-token window is real, but it isn't magic.
Pricing
API on Alibaba Cloud Bailian, list price: 12 yuan per million input tokens, 36 yuan per million output tokens. The 2026 promo knocks that to 50% off - 6 yuan in / 18 yuan out. Cache hits drop to 0.6 yuan/MTok; batch inference stacks another 50% off. One million tokens free for new accounts. There's also a Token Plan monthly subscription for teams that want flat-rate billing with priority scheduling.
In USD terms (international deployment): $2.50 in / $7.50 out per million tokens. Compare GPT-5.4 at ~$17.50/MTok and Claude Opus 4.7 at ~$15 in / $25 out. On output, Qwen3.7 Max is roughly a third of Opus 4.7's rate - and that's before the promo and cache discounts kick in. For high-volume Chinese workloads, the gap gets absurd fast.
The catch for US/EU users: it runs through Alibaba Cloud's infrastructure. If data residency or vendor geography matters to your compliance team, that's a real consideration, not a footnote.
Leaderboard Performance
As of the July 2026 snapshots:
- Artificial Analysis Intelligence Index: #5 globally, #1 among Chinese models (56.6).
- MMLU-ProX: #1 globally (87%).
- Code Arena: #4 at 1541 Elo - the first Chinese model ever to crack that table, slotting between Claude Opus 4.6 Thinking and Opus 4.6.
- Arena math: #7. Software & IT: #9. Expert arena: #9.
- Arena overall (text): #13 - solid but not top-tier.
- Vision Arena: #16 - highest-ranked Chinese lab, though remember Max itself is text-only; this reflects the broader Qwen3.7 family.
The pattern is telling. On pure knowledge and reasoning (MMLU-ProX, GPQA), it's at or near the global top. On the Arena chat vote - which rewards vibe and conversational polish - it's mid-pack. On coding and agentic tasks, it's the strongest Chinese model by a clear margin and genuinely competitive globally. Read that as: a serious workhorse, not a charming conversationalist.
How It Compares
Same modality (frontier text/reasoning), July 2026:
| Model | AA Index | Pricing (in/out per MTok) | Best at | |-------|----------|---------------------------|---------| | Claude Fable 5 | ~60 (#1) | $10 / $50 | Long reasoning, English prose | | GPT-5.6 Sol | ~58.9 (#2) | $5 / $30 | Math, tools, balanced | | Claude Opus 4.7 | 57.3 | $15 / $25 | Agentic, long docs | | Qwen3.7 Max | 56.6 (#5, #1 CN) | $2.50 / $7.50 | Chinese, cost, autonomy | | Gemini 3.5 Flash | 55.3 | $1.50 / $9 | Speed, multimodal | | DeepSeek V4 | high GPQA | ~$0.14 / $0.28 | Pure volume |
On Artificial Analysis's separate coding and agentic indices (June 2026), the pecking order is GPT-5.5 (Coding 74.9, Agentic 74.1), Claude Opus 4.8 (Coding 56.7, Agentic 77.8), then Qwen3.7 Max (Coding 50.1, Agentic 66.6) - the highest Chinese model on both, ahead of Kimi K2.6 and GLM-5.0.
The honest read: Qwen3.7 Max is not beating Fable 5 or GPT-5.6 on raw intelligence. It's roughly 3-4 points behind on the AA Index, and the coding-index gap to GPT-5.5 (50.1 vs 74.9) is real. But it costs a fraction of what they cost, and on agentic persistence - 35 hours, 1158 tool calls - it's in a class of basically one. Developer Paul Couvert's verdict after hooking it to Hermes Agent and OpenCode: it can "basically replace GPT-5.5 and Opus 4.7" for a lot of dev work. That's not vendor talk; that's a working dev with receipts. In one Atomic Chat test - building a self-training Tetris AI - Qwen3.7 Max did it for $1.32 in tokens and outperformed both Western flagships.
Who Should Use Qwen3.7 Max
- Chinese-first content: legal contracts in Chinese, financial-report analysis, long-form drafting. More natural than any Western frontier model, at a tenth of the price.
- Long-horizon agents: anything you'd otherwise babysit. The 35-hour autonomous run isn't a parlor trick.
- Cost-bound dev work: high-volume coding tasks where Opus-level prices would sink the budget.
- Enterprise in China: compliance, data residency, native integration with Alibaba Cloud's 400+ ecosystem services.
Don't use it for:
- Multimodal / vision tasks - it's text-only. Use Qwen3.7-Plus.
- Real-time web-grounded Q&A - no live search, mid-2025 cutoff.
- The absolute hardest English reasoning - Fable 5 and GPT-5.6 still edge it. If that last 3% matters, pay for it.
The mental model: Qwen3.7 Max is the workhorse you run all day, not the specialist you call for the one impossible problem.
x-rush: 顶级模型 + 智能路由
Here's the part that matters if you're actually trying to use these models instead of just reading about them. x-rush 接入 Qwen3.7 Max 等顶级大模型 - alongside Claude Fable 5, GPT-5.6, Gemini 3.5, DeepSeek V4, Kimi K3, Grok 4.5 - and 智能路由到最符合任务的模型. A Chinese legal brief routes to Qwen3.7 Max; a hard English reasoning prompt routes to Fable 5; a speed-sensitive one routes to Gemini 3.5 Flash; a cost-bound volume job routes to DeepSeek V4. The router reads language, difficulty, modality, and how agentic the task is, then picks the model that's actually best for that job - not a fixed default.
And the pool isn't frozen. 接入随世界潮流随时更新. When Qwen3.7 Max shipped in May and jumped the AA Index by 4.8 points, the router learned to send more Chinese work its way. When Gemini 3.5 Flash cut prices, the cost curve for fast text rerouted through it. When a new model meaningfully leads on quality, speed, or price, it enters the pool - the frontier moves, and the router moves with it.
That's the whole pitch: you stop betting on a single model, and start getting the best model for every task, updated as the world updates.
How to Use It on x-rush
The text workbench is where Qwen3.7 Max lives in practice. Open the Text workbench, drop in your prompt - a Chinese contract review, a long-form draft, a multi-step agent task - and the router decides whether Qwen3.7 Max is the right call or whether a faster, cheaper, or smarter model gets you the same answer for less. You don't pick the model. You describe the task; the router picks.
That's the point of the abstraction. You get Qwen3.7 Max's ceiling on Chinese and autonomy when you need it, and you don't pay for it when you don't.
The Bottom Line
Qwen3.7 Max is, by every independent measure I can find, the strongest Chinese model shipping in July 2026 - #5 globally and #1 domestically on the Artificial Analysis Intelligence Index, #1 on MMLU-ProX, the first Chinese model in Code Arena's top tier, and the only frontier model that's publicly demonstrated 35 hours of unattended autonomous work. It's also text-only, mid-pack on chat-vibe Arena, and a few points behind Fable 5 and GPT-5.6 on raw intelligence. None of that is a contradiction; it's a model with a clear shape - the cheap, Chinese-native, agent-persistent one.
If your work is Chinese-first, long-horizon, or cost-bound, Qwen3.7 Max is the model you want in the room. If your work is English-shaped, latency-sensitive, or needs vision, you want something else - and the smart move in 2026 isn't to pick one, it's to route. That's what x-rush does: 接入顶级大模型,智能路由到最符合任务的,随世界潮流随时更新。
Pick the model for the job. Or let a router do it for you.
Try the text workbench - it runs on x-rush's smart-routed model pool, Qwen3.7 Max included. The frontier, and every model that dethrones it next, is already in the pool.