Alibaba Releases Qwen3.7-Max Agent Model With 69.7 on Terminal-Bench 2.0

Alibaba's Qwen team released Qwen3.7-Max, a proprietary agent foundation model that posts a 69.7 on Terminal-Bench 2.0-Terminus, sustains 35-hour autonomous runs across 1,000+ tool calls, and is benchmarked to run inside Claude Code, OpenClaw, and Qwen Code.

Alibaba Releases Qwen3.7-Max Agent Model With 69.7 on Terminal-Bench 2.0

Alibaba's Qwen team released Qwen3.7-Max on Wednesday, positioning it as a proprietary agent foundation model rather than a chat completion model. The release is closed-source and API-only, available through Alibaba Cloud Model Studio with a 1 million-token context window. The headline number: Qwen3.7-Max scores 69.7 on Terminal-Bench 2.0-Terminus, ahead of DS-V4-Pro Max at 67.9, and ties Anthropic's Opus 4.6 Max on SWE-Verified at roughly 80.

The benchmark sweep

Qwen3.7-Max posts top-tier scores across the agent-relevant evaluation set: Terminal-Bench 2.0-Terminus, SWE-Pro, SciCode, MCP-Mark, GPQA Diamond, HMMT Feb 2026, and IMOAnswerBench. The point Alibaba is making with this spread is not that any single benchmark is dispositive — none of them are — but that the same weights post competitive numbers across coding, science reasoning, agent tooling, and math at the frontier band. That is the standard most-recent flagship releases from OpenAI, Anthropic, Google, and DeepSeek are converging on.

The harness story

The most interesting design choice is that Qwen3.7-Max is tested and tuned to run inside multiple external agent harnesses, including Anthropic's Claude Code, OpenClaw, Qwen's own Qwen Code, and custom in-house harnesses. Most labs ship a flagship model bound to their own runtime; Alibaba is shipping a model whose performance the company is willing to be measured on inside other vendors' agent loops. That posture follows from a real bet: agent reliability is a function of weights plus tool-use training plus harness compatibility, and Alibaba is asserting that its weights port cleanly enough across runtimes to compete on the harness the developer already prefers.

The 35-hour run

Alibaba's headline demo: an internal kernel-optimization task on a new chip platform that Qwen3.7-Max ran for 35 continuous hours, executing more than 1,000 tool calls and iterative code modifications, ending with a roughly 10x speedup on the target kernel. Long-horizon autonomous execution is the part of the agent stack most labs are still publicly weak on — sessions tend to either fail closed or drift over multi-hour timescales. A repeatable 35-hour run with non-trivial output is a useful upper-bound datapoint, even allowing for vendor-curated framing.

The closed-weights pivot

Qwen has historically been one of the most open Chinese model families — Qwen2 and Qwen3 weights are on Hugging Face and have been adopted across the open-source agent stack. Qwen3.7-Max is closed. Alibaba's framing is that agent reliability requires post-training and tool-routing work that does not translate to open weights, so the highest-capability tier will sit behind the Cloud Model Studio API. That tracks with what DeepSeek, Moonshot, and Zhipu are signaling on their flagship tiers — the open-weights story stays in the mid-tier, the frontier tier ships closed. The strategic shape of Chinese AI is starting to look more like the U.S. cohort than it did six months ago, which matters as Anthropic's multi-cloud chip portfolio and OpenAI's IPO position both stake bigger claims on the agent layer.

Alibaba is buying exposure to that layer as well as building it: in July 2026 it led the extension that brought AI video startup AIsphere's Series C to $439 million, its second check into the maker of PixVerse in under a year.

Comments

Get tomorrow's roundup. Free.

One email each morning. Sneakers, sports, culture, tech.

Link copied