RouterHub

Qwen3.7-Max and Qwen3.7-Plus Are Now Available on RouterHub

RouterHub Team · Updated 2026-07-24

Qwen3.7-Max (qwen/qwen3.7-max) and Qwen3.7-Plus (qwen/qwen3.7-plus) are now available on RouterHub.

Qwen3.7-Max is the text-first option for complex reasoning and coding work. Qwen3.7-Plus is the balanced multimodal option for workflows that combine long context with text, images, or video. RouterHub lets teams evaluate both through the same unified access layer they already use for other leading AI models.

Which Qwen3.7 model should you start with?

Start with Qwen3.7-Max when the workload is text-first and the strongest reasoning or coding performance is the priority. Start with Qwen3.7-Plus when the workflow needs visual context, long documents, tool use, structured output, or multimodal input.

Model RouterHub model ID Input → output Context / max output Start here when
Qwen3.7-Max qwen/qwen3.7-max Text → text 1M / 64K tokens Complex reasoning, coding, and multi-step agent tasks
Qwen3.7-Plus qwen/qwen3.7-plus Text, image, video → text 1M / 64K tokens Long-context, multimodal, and structured workflow tasks

Qwen3.7-Max for complex text and reasoning work

Qwen3.7-Max is the flagship text model in the Qwen3.7 family. It is designed for demanding tasks that require reasoning, code generation, multilingual understanding, or sustained multi-step execution.

Good first tests include:

  • Complex coding and debugging tasks
  • Technical research and synthesis
  • Multi-step agent workflows
  • Long-form analysis and document processing
  • Multilingual support and translation workflows

For production teams, the best evaluation starts with real prompts rather than a generic leaderboard. Use a task with a clear success criterion, compare Qwen3.7-Max with the model currently in production, and record output quality, latency, retries, and total usage.

What benchmark scores has Qwen3.7-Max published?

Qwen has published the following Qwen3.7-Max evaluation results. These are provider-published results, not benchmarks independently reproduced by RouterHub.

Benchmark Qwen3.7-Max result What it helps evaluate
Artificial Analysis Intelligence Index 56.6 Broad frontier-model capability comparison
GPQA Diamond 92.4 Graduate-level scientific reasoning
HLE 41.4 Difficult multidisciplinary reasoning
HMMT 2026 Feb 97.1 Competition mathematics
IMOAnswerBench 90.0 Mathematical problem solving
Apex 44.5 Advanced reasoning
Terminal Bench 2.0-Terminus 69.7 Terminal-based agent execution
SWE-Verified 80.4 Repository-level software engineering
SWE-Pro 60.6 More difficult software engineering tasks
SWE-Multilingual 78.3 Multilingual software engineering

Benchmark setups, agent scaffolds, tool access, and task distributions vary. Treat these results as useful evaluation signals, not a production decision by themselves. The practical test is to compare Qwen3.7-Max with your current model on the prompts, tools, latency requirements, and success criteria that matter to your team.

Qwen3.7-Plus for long-context multimodal workflows

Qwen3.7-Plus is the balanced model in the Qwen3.7 family. On RouterHub, it accepts text, image, and video input, with a 1M-token context window, function calling, and structured output support.

That makes Qwen3.7-Plus a practical starting point for:

  • Analyzing long reports, contracts, and knowledge bases
  • Reviewing screenshots, product images, and visual documents
  • Building assistants that need structured JSON output
  • Video and image understanding tasks
  • Multi-turn workflows that need to retain more context

Alibaba Cloud has not published a matching public benchmark scorecard for Qwen3.7-Plus. Its practical value is best evaluated through workload fit: how well it handles your visual inputs, document context, tool calls, and structured outputs at an acceptable operating profile.

How should teams evaluate both models?

  1. Choose a real workflow with a measurable quality bar.
  2. Keep the prompt, tools, context, and success criteria fixed.
  3. Run Qwen3.7-Max and Qwen3.7-Plus where the input type allows it.
  4. Review answer quality, completion time, tool behavior, retries, and usage.
  5. Select a default model for the workload and retain the other for cases where the task profile changes.

For example, a coding assistant may begin with Qwen3.7-Max for difficult technical tasks, while a document-review product may use Qwen3.7-Plus when images, video, or very long context are part of the input.

Why access Qwen3.7 through RouterHub?

Model launches should not require teams to rebuild their integration or create a new operating workflow every time a provider updates its lineup.

RouterHub gives teams a unified way to access, compare, and operate across leading AI models. You can evaluate Qwen3.7 alongside the models already used in your product, then make a deliberate routing decision based on your own quality, latency, and cost requirements.

That keeps model experimentation close to production reality: one access layer, a clearer evaluation process, and more flexibility when workloads change.

Frequently asked questions

Is Qwen3.7-Max available on RouterHub?

Yes. Qwen3.7-Max is available on RouterHub as qwen/qwen3.7-max for teams evaluating complex text, reasoning, coding, and multilingual workloads.

Is Qwen3.7-Plus available on RouterHub?

Yes. Qwen3.7-Plus is available on RouterHub as qwen/qwen3.7-plus for multimodal workflows involving text, images, or video.

What is the difference between Qwen3.7-Max and Qwen3.7-Plus?

Qwen3.7-Max is the flagship text model for demanding reasoning and coding tasks. Qwen3.7-Plus is a balanced multimodal model for long-context workflows that can include text, images, and video.

Should I select a model based only on benchmark scores?

No. Use benchmark scores to form an initial hypothesis, then test representative production tasks. Compare quality, completion time, tool behavior, retries, and total usage before setting a production default.

Start testing Qwen3.7 on RouterHub

Compare Qwen3.7-Max and Qwen3.7-Plus in your existing RouterHub workflow.