Qwen3-Max Thinking beats Gemini 3 Pro and GPT-5.2 on Humanity's Last Exam (with search)
Qwen3-Max Thinking beats Gemini 3 Pro and GPT-5.2 on Humanity's Last Exam (with search)
Chinese AI and tech firms continue to impress with their development of cutting-edge, state-of-the-art AI language models.
Today, the one drawing eyeballs is Alibaba Cloud's Qwen Team of AI researchers and its unveiling of a new proprietary language reasoning model, Qwen3-Max-Thinking.
You may recall, as VentureBeat covered last year, that Qwen has made a name for itself in the fast-moving global AI marketplace by shipping a variety of powerful, open source models in various modalities, from text to image to spoken audio. The company even earned an endorsement from U.S. tech lodgings giant Airbnb, whose CEO and co-founder Brian Chesky said the company was relying on Qwen's free, open source models as a more affordable alternative to U.S. offerings like those of OpenAI.
Now, with the proprietary Qwen3-Max-Thinking, the Qwen Team is aiming to match and, in some cases, outpace the reasoning capabilities of GPT-5.2 and Gemini 3 Pro through architectural efficiency and agentic autonomy.
The release comes at a critical juncture. Western labs have largely defined the "reasoning" category (often dubbed "System 2" logic), but Qwen’s latest benchmarks suggest the gap has closed.
In addition, the company's relatively affordable API pricing strategy aggressively targets enterprise adoption. However, as it is a Chinese model, some U.S. firms with strict national security requirements and considerations may be wary of adopting it.
The Architecture: "Test-Time Scaling" Redefined
The core innovation driving Qwen3-Max-Thinking is a departure from standard inference methods. While most models generate tokens linearly, Qwen3 utilizes a "heavy mode" driven by a technique known as "Test-time scaling."
In simple terms, this technique allows the model to trade compute for intelligence. But unlike naive "best-of-N" sampling—where a model might generate 100 answers and pick the best one — Qwen3-Max-Thinking employs an experience-cumulative, multi-round strategy.
This approach mimics human problem-solving. When the model encounters a complex query, it doesn't just guess; it engages in iterative self-reflection. It uses a proprietary "take-experience" mechanism to distill insights from previous reasoning steps. This allows the model to:
Identify Dead Ends: Recognize when a line of reasoning is failing without needing to fully traverse it.
Focus Compute: Redirect processing power toward "unresolved uncertainties" rather than re-deriving known conclusions.
The efficiency gains are tangible. By avoiding redundant reasoning, the model integrates richer historical context into the same window. The Qwen team reports that this method drove massive performance jumps without exploding token costs:
GPQA (PhD-level science): Scores improved from 90.3 to 92.8.
LiveCodeBench v6: Performance jumped from 88.0 to 91.4.
Beyond Pure Thought: Adaptive Tooling
While "thinking" models are powerful, they have historically been siloed — great at math, but poor at browsing the web or running code. Qwen3-Max-Thinking bridges this gap by effectively integrating "thinking and non-thinking modes".
The model features adaptive tool-use capabilities, meaning it autonomously selects the right tool for the job without manual user prompting. It can seamlessly toggle between:
Web Search & Extraction: For real-time factual queries.
Memory: To store and recall user-specific context.
Code Interpreter: To write and execute Python snippets for computational tasks.
In "Thinking Mode," the model supports these tools simultaneously. This capability is critical for enterprise applications where a model might need to verify a fact (Search), calculate a projection (Code Interpreter), and then reason about the strategic implication (Thinking) all in one turn.
Empirically, the team notes that this combination "effectively mitigates hallucinations," as the model can ground its reasoning in verifiable external data rather than relying solely on its training weights.
Benchmark Analysis: The Data Story
Qwen is not shy about direct comparisons.
On HMMT Feb 25, a rigorous reasoning benchmark, Qwen3-Max-Thinking scored 98.0, edging out Gemini 3 Pro (97.5) and significantly leading DeepSeek V3.2 (92.5).
However, the most significant signal for developers is arguably Agentic Search. On "Humanity's Last Exam" (HLE) — the benchmark that measures performance on 3,000 "Google-proof" graduate-level questions across math, science, computer science, humanities and engineering — Qwen3-Max-Thinking, equipped with web search tools, scored 49.8, beating both Gemini 3 Pro (45.8) and GPT-5.2-Thinking (45.5) .

Qwen3-Max key benchmarks. Credit: Alibaba Cloud Qwen Team on X
This suggests that Qwen3-Max-Thinking’s architecture is uniquely suited for complex, multi-step agentic workflows where external data retrieval is necessary.
In coding tasks, the model also shines. On Arena-Hard v2, it posted a score of 90.2, leaving competitors like Claude-Opus-4.5 (76.7) far behind.
The Economics of Reasoning: Pricing Breakdown
For the first time, we have a clear look at the economics of Qwen's top-tier reasoning model. Alibaba Cloud has positioned qwen3-max-2026-01-23 as a premium but accessible offering on its API.
Input: $1.20 per 1 million tokens (for standard contexts <= 32k).
Output: $6.00 per 1 million tokens.
On a base level, here's how Qwen3-Max-Thinking stacks up:
Model | Input (/1M) | Output (/1M) | Total Cost | Source |
Qwen 3 Turbo | $0.05 | $0.20 | $0.25 | Alibaba Cloud |
Grok 4.1 Fast (reasoning) | $0.20 | $0.50 | $0.70 | xAI |
Grok 4.1 Fast (non-reasoning) | $0.20 | $0.50 | $0.70 | xAI |
deepseek-chat (V3.2-Exp) | $0.28 | $0.42 | $0.70 | DeepSeek |
deepseek-reasoner (V3.2-Exp) | $0.28 | $0.42 | $0.70 | DeepSeek |
Qwen 3 Plus | $0.40 | $1.20 | $1.60 | Alibaba Cloud |
ERNIE 5.0 | $0.85 | $3.40 | $4.25 | Qianfan |
Gemini 3 Flash Preview | $0.50 | $3.00 | $3.50 | |
Claude Haiku 4.5 | $1.00 | $5.00 | $6.00 | Anthropic |
Qwen3-Max Thinking (2026-01-23) | $1.20 | $6.00 | $7.20 | Alibaba Cloud |
Gemini 3 Pro (≤200K) | $2.00 | $12.00 | $14.00 | |
GPT-5.2 | $1.75 | $14.00 | $15.75 | OpenAI |
Claude Sonnet 4.5 | $3.00 | $15.00 | $18.00 | Anthropic |
Gemini 3 Pro (>200K) | $4.00 | $18.00 | $22.00 | |
Claude Opus 4.5 | $5.00 | $25.00 | $30.00 | Anthropic |
GPT-5.2 Pro | $21.00 | $168.00 | $189.00 | OpenAI |
This pricing structure is aggressive, undercutting many legacy flagship models while offering state-of-the-art performance.
However, developers should note the granular pricing for the new agentic capabilities, as Qwen separates the cost of "thinking" (tokens) from the cost of "doing" (tool use).
Agent Search Strategy: Both standard
search_strategy:agentand the more advancedsearch_strategy:agent_maxare priced at $10 per 1,000 calls.Note: The
agent_maxstrategy is currently marked as a "Limited Time Offer," suggesting its price may rise later.
Web Search: Priced at $10 per 1,000 calls via the Responses API.
Promotional Free Tier:To encourage adoption of its most advanced features, Alibaba Cloud is currently offering two key tools for free for a limited time:
Web Extractor: Free (Limited Time).
Code Interpreter: Free (Limited Time).
This pricing model (low token cost + à la carte tool pricing) allows developers to build complex agents that are cost-effective for text processing, while paying a premium only when external actions—like a live web search—are explicitly triggered.
Developer Ecosystem
Recognizing that performance is useless without integration, Alibaba Cloud has ensured Qwen3-Max-Thinking is drop-in ready.
OpenAI Compatibility: The API supports the standard OpenAI format, allowing teams to switch models by simply changing the
base_urlandmodelname.Anthropic Compatibility: In a savvy move to capture the coding market, the API also supports the Anthropic protocol. This makes Qwen3-Max-Thinking compatible with Claude Code, a popular agentic coding environment.
The Verdict
Qwen3-Max-Thinking represents a maturation of the AI market in 2026. It moves the conversation beyond "who has the smartest chatbot" to "who has the most capable agent."
By combining high-efficiency reasoning with adaptive, autonomous tool use—and pricing it to move—Qwen has firmly established itself as a top-tier contender for the enterprise AI throne.
For developers and enterprises, the "Limited Time Free" windows on Code Interpreter and Web Extractor suggest now is the time to experiment. The reasoning wars are far from over, but Qwen has just deployed a very heavy hitter.
More Articles
Google announces a new protocol to facilitate commerce using AI agents
Google said that merchants can now offer discounts to users directly in AI mode results.
Meta acquires Manus: what it means for your enterprise AI agent strategy
Meta acquires Manus: what it means for your enterprise AI agent strategy
Black box AI isn’t enough: Why enterprise consulting is moving to grounded models
"Grounded AI is non-negotiable, because accuracy isn’t optional when we’re doing million-dollar transformation projects within the SAP ecosystem."
Simplifying the AI stack: The key to scalable, portable intelligence from cloud to edge
To unlock the next wave of AI innovation, the industry must pivot decisively away from siloed development and toward streamlined, end-to-end platforms.
Late Kim Sae Ron's Last Film Suffers Terrible Fate
The movie Before We Knew, starring the late Kim Sae Ron, quickly fell out of the box office rankings shortly after its release, contrary to early expectations. According to the Korea…
Entertainment Stocks Lose Billions Of Dollars, Concerns About World Tour Disruption Arise
Amid the K-entertainment stock market shaking due to the war between the U.S., Israel, and Iran, the combined market capitalization of the four major entertainment companies has evaporated by more…
Famous Producer Indicted For Forcibly Molesting A Colleague
A famous variety show producer (PD) accused of forcibly molesting a junior colleague at work has been sent to trial. On the 27th, the Seoul Western District Prosecutors’ Office announced…
Flexiquiz Review 2025: The Ultimate Online Quiz Maker
Creating quizzes doesn’t have to be a daunting task. Whether you’re a teacher needing to assess students, a business seeking training tools, or a hobbyist hosting a fun quiz night, having the right platform can make all the difference. Enter Flexiquiz, a robust and user-friendly online quiz maker. With easy-to-use features, advanced customization options, and budget-friendly pricing, Flexiquiz is shaping up to be the quiz maker of choice for 2025. This review will explore its key features, benefits, pricing, and why it stands out among other tools. What is Flexiquiz? Flexiquiz is an online platform designed to create quizzes, exams, and interactive tests for professional and personal use. Think back to those multiple-choice tests you used to take in school, but this time, you’re in control as the quiz creator! From live quiz games to formal assessments, Flexiquiz offers the flexibility to cater to a wide range of needs. Whether for fun or functionality, it’s easy to use and packed with features that make quiz creation both simple and engaging. Why Choose Flexiquiz? Flexiquiz is designed by Logic Software with a mission to simplify online assessments for various users, ranging from educators to enterprises. Here’s what makes it special: Effortless quiz creation tailored to your needs. Options to host secure, formal exams or casual live quizzes. Continuous updates and new features to enhance functionality. Top Benefits of Flexiquiz 1. Ease of Use Flexiquiz’s intuitive interface guides users through the quiz creation process. No prior experience? No problem! You can create quizzes in minutes with step-by-step instructions. 2. Variety of Question Types Gone are the days of boring assessments. Flexiquiz offers diverse question types such as multiple choice, true/false, fill-in-the-blank, and matching. Add images or videos to further engage your audience. 3. Customizable Designs Personalize your quizzes with logos, colors, and designs that match your brand or style. Flexiquiz even supports custom CSS for advanced users seeking more detailed customization. 4. Automated Grading Save valuable time with automatic grading. Flexiquiz evaluates responses instantly and provides participants with immediate scores. 5. Detailed Reporting Gain insights with Flexiquiz’s in-depth reports, which track individual scores, group performance, and even question-specific analytics. 6. Secure Features With options like access codes and secure browser settings, Flexiquiz ensures that both casual and formal assessments remain private and fair. 7. Live Quiz Hosting Inject fun and interactivity into your quizzes with Flexiquiz’s live quiz mode. Display leaderboards to track participant scores in real-time. Best Features of Flexiquiz Quiz Maker Flexiquiz’s core feature is easy to use and versatile. With support for multimedia integration, you can craft engaging quizzes appropriate for any audience. Custom CSS For tech-savvy users, Flexiquiz allows styling through custom CSS, helping to create visually unique quiz presentations. Exam Builder If you’re designing formal tests, Flexiquiz has you covered. Implement time limits, randomize questions, and prevent cheating with secure settings. Custom Grading You’re not limited to traditional right-or-wrong scoring. Flexiquiz lets you assign partial credit, weigh scores, and provide personalized feedback. Live Quizzes Leverage the interactive live quiz feature to host gamified sessions, whether for classrooms, corporate training, or just a fun night with friends. Flexiquiz Pricing Flexiquiz offers a range of pricing plans tailored to meet different needs, ensuring flexibility and affordability for users. The free plan includes basic features such as unlimited quizzes, 100 responses per month, and access to essential customization options. For more advanced features, the premium plans provide additional benefits, such as higher response limits, advanced reporting, integrations, and branding options. These plans include: Starter Plan: Ideal for individuals, with expanded response limits and basic reporting. Team Plan: Perfect for small groups, offering shared accounts, team management, and collaboration tools. Business Plan: Designed for larger organizations, including priority support, API access, and enhanced security options. Enterprise Plan: Customized for extensive needs, with dedicated account managers and bespoke solutions. Flexiquiz offers plans to suit varying needs and budgets. Pros and Cons Pros Supports multiple question types. Highly customizable with options for advanced design changes. Affordable pricing plans. Detailed analytics and reporting. Cons Occasional glitches. Higher learning curve for advanced features like custom CSS. Limited design templates out-of-the-box. Alternatives to Flexiquiz If Flexiquiz doesn’t quite fit your requirements, here are some alternatives worth considering: ProProfs Quiz Maker for education and training purposes. Typeform for sleek, visually engaging quizzes. SurveyMonkey for those who need combined survey and quiz capabilities. Google Forms for a free, no-frills option. Kahoot! for gamified quizzes tailored to live interactive sessions. Real-Life Use Case Our team recently needed a training solution for onboarding new employees. Instead of lengthy presentations, we used Flexiquiz to create an interactive quiz with videos, images, and custom feedback for each question. Here’s what stood out: Straightforward setup: Crafting the quiz took less than an hour. Interactive elements: Videos and images kept participants engaged. Automatic grading: Results were instantly available, saving us hours of manual review. Flexiquiz turned a tedious process into a dynamic, engaging learning experience! Final Thoughts Flexiquiz is a feature-packed, easy-to-use quiz maker ideal for educators, businesses, and quiz enthusiasts. Whether you’re hosting live sessions or designing formal tests, it offers tools to enhance engagement and save time. With affordable pricing, advanced features like custom CSS, and a free trial to start, Flexiquiz is a fantastic option for anyone seeking a flexible online quiz platform. Try it today and streamline your quiz-making experience! Frequently Asked Questions Can I use Flexiquiz for free? Yes! Flexiquiz offers a free plan with basic features, though it limits the number of questions and responses. What customization options are available? Flexiquiz allows branding with custom logos and colors. Advanced users can even apply custom CSS for detailed design changes. Are live quizzes supported? Absolutely! With Flexiquiz, you can host live quizzes with real-time leaderboards for interactive engagement. Can educators use Flexiquiz for exams? Yes, the tool includes features like randomized questions, time limits, and browser security to ensure integrity in exams.