Model comparison
How close are DeepSeek and Qwen to Claude?
In regions where Claude, GPT and Gemini are unavailable, SOVEREIGN offers models such as DeepSeek and Qwen. Below are our own measurements on 40 automatically graded tasks and public benchmarks — including where Claude is still stronger.
Measured on 2026-09-27
Open SOVEREIGN →Measured result
- Qwen 3.8 27B scored 92.5% of Claude Opus 4.8's result on our tasks (37 vs 40 solved) at about 5× lower cost. This is not parity: Claude Opus 4.8 solved more tasks. See the table below for where the gap is.
- In our run Claude Opus 4.8 was ahead in: Coding, Instruction following.
- Not measured yet: Claude Opus 5, Claude Sonnet 5, DeepSeek V4 Pro, DeepSeek V4.1 Flash, Qwen 3.8 Max. Their API route was unavailable during this run (no credits/quota). We don't publish estimates in place of measurements.
Our eval: 40 tasks, graded automatically
Every task is checked by code, not by a person or another AI: unit tests for programs, exact answers for math, strict rules for instructions and text. Mostly in Russian, a few in Uzbek and English.
| Model | Coding15 tasks | Math & logic10 tasks | Instruction following8 tasks | Russian text & translation7 tasks | Overall | Median latency | Cost of 40 tasks | Route |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.8reference | 100%(15/15) | 100%(10/10) | 100%(8/8) | 100%(7/7) | 100%(40/40) | 3.6 s | $0.16 | rsi |
| Qwen 3.8 27B | 93.3%(14/15) | 100%(10/10) | 75%(6/8) | 100%(7/7) | 92.5%(37/40) | 0.5 s | $0.030 | groq |
| Claude Opus 5 | not measured yet | |||||||
| Claude Sonnet 5 | not measured yet | |||||||
| DeepSeek V4 Pro | not measured yet | |||||||
| DeepSeek V4.1 Flash | not measured yet | |||||||
| Qwen 3.8 Max | not measured yet | |||||||
Cost = tokens used × the OpenRouter list price for the same model. Latency depends on the provider's hardware (Groq is unusually fast), not only on the model.
Per-task results
| Task | Claude Opus 4.8 | Qwen 3.8 27B |
|---|---|---|
| js-plural-ru · ru | pass | pass |
| js-palindrome-ru · ru | pass | pass |
| js-merge-intervals · ru | pass | pass |
| js-roman · ru | pass | pass |
| js-parse-duration · ru | pass | pass |
| js-flatten · ru | pass | pass |
| js-format-rub · ru | pass | pass |
| js-eval-rpn · en | pass | pass |
| py-top-k-words · ru | pass | pass |
| py-inn · ru | pass | fail |
| py-spiral · ru | pass | pass |
| py-lcs · ru | pass | pass |
| py-to-snake · uz | pass | pass |
| py-dijkstra · ru | pass | pass |
| py-wrap-text · ru | pass | pass |
| math-trains · ru | pass | pass |
| math-divisors · ru | pass | pass |
| math-dice · ru | pass | pass |
| math-distinct-digits · ru | pass | pass |
| math-sisters · ru | pass | pass |
| math-modpow · ru | pass | pass |
| math-walls-uz · uz | pass | pass |
| math-sequence · ru | pass | pass |
| math-weekday · ru | pass | pass |
| math-factorial-en · en | pass | pass |
| if-json-object · ru | pass | pass |
| if-json-sorted · ru | pass | pass |
| if-three-sentences · ru | pass | pass |
| if-keywords · ru | pass | fail |
| if-no-letter-o · ru | pass | fail |
| if-numbered-list · ru | pass | pass |
| if-answer-in-russian · en | pass | pass |
| if-uppercase · ru | pass | pass |
| wr-en-ru · ru | pass | pass |
| wr-ru-en · ru | pass | pass |
| wr-uz-ru · ru | pass | pass |
| wr-ru-uz · uz | pass | pass |
| wr-business-email · ru | pass | pass |
| wr-summary · ru | pass | pass |
| wr-product · ru | pass | pass |
Public benchmarks
Figures copied from the linked pages on 2026-09-27. In places the model versions differ from the ones we measured — check the exact model name in each row.
Artificial Analysis Intelligence Index v4.3.2
independentComposite intelligence index (10 evaluations), points.
| Model | Score |
|---|---|
| Claude Opus 5.5 (max with fallback) | 58 |
| Qwen3.8 Max (0902) | 45 |
| DeepSeek V4.1 Flash (max) | 39 |
| Claude Sonnet 5 (max) | 38 |
| DeepSeek V4 Pro 0813 (max) | 36 |
| Qwen3.8 27B (xhigh) | 34 |
Source: artificialanalysis.ai/leaderboards/models · 2026-09-27
LMArena — Text
independentRating from blind human votes (Elo), text.
| Model | Score |
|---|---|
| claude-opus-5-high | 1491 ± 4 |
| claude-opus-4-8-high | 1480 ± 4 |
| qwen3.8-max | 1479 ± 5 |
| deepseek-v4.1-flash-max | 1477 ± 8 |
| deepseek-v4-pro-high-20260813 | 1464 ± 7 |
| claude-sonnet-5-high | 1462 ± 4 |
| qwen3.8-27b | 1438 ± 6 |
Source: arena.ai/leaderboard/text · 2026-09-25
Humanity's Last Exam (Artificial Analysis)
independentHardest expert questions, % correct.
| Model | Score |
|---|---|
| Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | 61% |
| Qwen3.8 Max (0902) | 43% |
| Claude Sonnet 5 (Adaptive Reasoning, Max Effort) | 41% |
| DeepSeek V4 Pro 0813 (Reasoning, Max Effort) | 41% |
Source: artificialanalysis.ai/models/comparisons/claude-opus-5-5-vs-deepseek-v4-pro, artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-qwen3-8-max · 2026-09-27
Terminal-Bench 4.0 (Artificial Analysis)
independentAgentic coding in a terminal, % solved.
| Model | Score |
|---|---|
| Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | 60% |
| Qwen3.8 Max (0902) | 39% |
| Claude Sonnet 5 (Adaptive Reasoning, Max Effort) | 14% |
| DeepSeek V4 Pro 0813 (Reasoning, Max Effort) | 14% |
Source: artificialanalysis.ai/models/comparisons/claude-opus-5-5-vs-deepseek-v4-pro, artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-qwen3-8-max · 2026-09-27
SWE-bench Verified
self-reported by vendorsFixing real GitHub issues, % resolved.
| Model | Score |
|---|---|
| Claude Opus 4.8 | 88.6% |
| Claude Sonnet 5 | 85.2% |
| DeepSeek-V4-Pro-Max | 80.6% |
| Qwen3.7 Max | 80.4% |
| DeepSeek-V4-Flash-Max | 79.0% |
Source: llm-stats.com/benchmarks/swe-bench-verified · 2026-09-27
Methodology and limitations
- 40 tasks, 1 run per model, on 2026-09-27. Temperature 0 where supported.
- Models are called through the same provider routes the app itself uses (OpenRouter, RSI, Groq, OmniRoute); the route actually used is shown in the table.
- Code is checked with hidden unit tests in an isolated process (no network, with a timeout) — all tests must pass.
- Text tasks check only alphabet, length and required terms — style and fluency are not judged.
- This is a small, focused test, not a full benchmark: a single run can vary, and 40 tasks cannot cover everything.
- Per public benchmarks, Claude's top models are still ahead at agentic coding in a terminal (Terminal-Bench) and on the hardest expert questions (Humanity's Last Exam).
- Cost and latency are indicative: real prices differ on resellers and free tiers.
FAQ
Do DeepSeek or Qwen work exactly like Claude?
No, and we don't claim that. On many everyday tasks they are close — see the measured numbers above — but not identical, and Claude is ahead on some tasks.
Why can't I choose Claude in my region?
Some model providers don't serve certain regions. We only offer models whose providers allow your region.
Which model should I use for code?
Check the Coding column and pick the best-scoring model available to you. For long agentic work, public benchmarks show a larger gap to Claude.
How often is this page updated?
When models are added or updated, we rerun the eval and update the date above.