AI Model Research · Benchmark & Market Signal · 2026 Snapshot

AI Model Benchmark & Market Signal
2026 Snapshot

เปรียบเทียบ 16 โมเดลชั้นนำอย่างเป็นกลาง
ครอบคลุม Benchmark, Platform Runtime, Market Signal
4 partially verified · 3 fully verified · 9 evidence gap

16 Models Sources: OpenRouter + Artificial Analysis Snapshot: 2026 ภาษาไทย + EN Terms
⚠️ วิธีอ่านหน้านี้:
ข้อมูลอ้างอิงจาก OpenRouter (Platform Runtime) และ Artificial Analysis (Benchmark, Estimated) — ควรตรวจสอบ source หลักก่อนตัดสินใจ

ความหมายของป้ายกำกับ:
Benchmark — quality/performance indicators (Artificial Analysis)
Platform Runtime — price, context, provider-route data (OpenRouter)
Market Signal — adoption/availability proxy, not absolute quality or market share
Limited / Preview — ข้อมูลไม่ครบถ้วน หรือโมเดลยังอยู่ในช่วง Preview/Beta
Estimated — ค่าประมาณการจาก source ที่ไม่ใช่การวัดโดยตรง
Confirmed   Estimated   Limited Data   Partially Verified
Executive Summary

ภาพรวมที่ใช้ตัดสินใจได้จริง

สรุป insight สำคัญจากการเปรียบเทียบ 16 โมเดล — เน้นสิ่งที่แตกต่างจริง ไม่ใช่คำอธิบายกว้าง ๆ

🏆
Frontier Leader (Verified)
Claude Opus 4.6
Intelligence 53 verified · $5/$25 · legacy-stable baseline. GPT-5.5 (Intelligence 60) and Claude Opus 4.7 (Intelligence 57) are partially verified but not yet comparison-eligible.
⚡
Fastest Throughput (Verified)
Gemini 2.5 Flash
198 tok/s verified · $0.30/$2.50 · highest verified throughput in this snapshot
💰
Best Value (Verified)
Gemini 2.5 Flash
Intelligence 21, $0.30/$2.50, 198 tok/s — best verified cost/speed ratio. Other models pending evidence.
🧠
Best for Coding
Pending evidence
No verified model currently holds a coding-specialist role. Qwen3.7 Max is a candidate pending verification.
📚
Stable Production Baseline
Gemini 2.5 Flash
Verified Intelligence 21, $0.30/$2.50, 198 tok/s — stable production baseline with complete evidence
🎯
Agentic Capability
Pending evidence
Kimi K2.6 (Intelligence 54, partially verified) is the leading agentic candidate. DeepSeek V4 Pro remains evidence-gap. No comparison-eligible agentic winner yet.
🚀
Highest Intelligence (Comparison-Eligible)
Claude Opus 4.6
Intelligence 53 comparison-eligible · $5/$25. GPT-5.5 (Intelligence 60, partially verified) is not yet comparison-eligible pending OpenRouter pricing confirmation.
⚠️
Evidence Status
9 evidence gap · 4 partial
9 of 16 models lack benchmark metrics. 4 models now partially verified (GPT-5.5, Claude Opus 4.7, Gemini 3.5 Flash, Kimi K2.6) with AA intelligence (throughput/TTFT documented in source matrix, excluded from dashboard). 3 fully verified. Partially verified models excluded from comparative filters and winner cards until OpenRouter pricing confirmed.
Key Observations
Frontier intelligence verified from AA
GPT-5.5 (Intelligence 60), Claude Opus 4.7 (Intelligence 57), and Gemini 3.5 Flash (Intelligence 55) now have verified AA metrics. GPT-5.5 leads overall but is not yet comparison-eligible. Claude Opus 4.6 (Intelligence 53) remains the highest comparison-eligible frontier model.
Evidence gap narrowing
4 of 16 active models now partially verified with AA intelligence (GPT-5.5, Claude Opus 4.7, Gemini 3.5 Flash, Kimi K2.6). AA throughput/TTFT documented in source matrix but excluded from dashboard pending OpenRouter runtime normalization. 9 models remain evidence-gap. Only Claude Sonnet 4.6, Claude Opus 4.6, and Gemini 2.5 Flash have complete verified evidence with OpenRouter pricing. Partially verified models excluded from derived scoring, comparative filters, and winner cards (score_eligible=false, comparison_eligible=false).
Expanded Chinese model coverage (pending evidence)
DeepSeek V4 Flash + V4 Pro, MiniMax M2.7, Step 3.5 Flash, Qwen3.7 Max, Kimi K2.6 — broader representation across flash, balanced, and agentic segments. All pending benchmark verification.
Verified pricing gap
Claude Opus 4.6 at $25/M output vs Gemini 2.5 Flash at $0.30/$2.50 — 10x output premium (verified). Gemini 2.5 Flash remains the cost-efficient verified baseline. gpt-oss-120b pricing pending route verification.
Market Signal

Market Signal View

Adoption and availability proxy based on route status, market tier, segment, source coverage, and caveat flags. This is not a full market-share or token-volume ranking.

Segment Breakdown
DERIVED FROM MODELS DATA · adoption & availability proxy by segment
Segment Models Count Signal Interpretation
CAVEAT
Market Signal here is a proxy layer. It reflects current model categorization, route maturity, source coverage, and caveat visibility in this dashboard. It does not represent full market share, monthly token volume, or independent quality ranking.
Benchmark View

กราฟเปรียบเทียบ Benchmark Metrics

ข้อมูล Pricing/Context จาก OpenRouter (Confirmed, Platform Runtime) — Intelligence Index, Throughput, TTFT จาก Artificial Analysis (Estimated, Benchmark Quality) — 4 partially verified (data not shown) · 3 fully verified · 9 evidence gap — สลับ metric ได้ด้วยปุ่มด้านล่าง

Intelligence Index (Artificial Analysis)
ESTIMATED FROM ARTIFICIAL ANALYSIS · ↑ สูงกว่า = ฉลาดกว่า
* Intelligence Index จาก Artificial Analysis — 4 partially verified (GPT-5.5, Claude Opus 4.7, Gemini 3.5 Flash, Kimi K2.6), 3 fully verified, 9 evidence gap. Partially verified models not comparison-eligible — see Watchlist section. ไม่ใช่ absolute truth · ควรตรวจสอบจาก source หลักก่อนตัดสินใจ
Market Signal View

ตารางเปรียบเทียบครบ

ข้อมูล Pricing, Runtime, Provider-route availability — proxy layer of adoption/availability, not absolute model quality or market share

Confirmed (OpenRouter — Pricing/Context)
Estimated (Artificial Analysis — Intelligence Index, Throughput, TTFT)
Limited Data (evidence gap)
Partially Verified (AA intelligence verified, OpenRouter pricing pending)
กรอง:
Model ↕ Vendor ↕ Intelligence* Input $/M ↕ Output $/M ↕ TTFT* ↕ Throughput ↕ Context ↕ Open/Closed Tools Cache Modalities Best For ↕ Segment
* Intelligence Index, Throughput, TTFT จาก Artificial Analysis (Estimated) — ไม่ใช่ benchmark แบบตัดสินขาด · Pricing/Context จาก OpenRouter (Confirmed) · Partially Verified models have AA intelligence but not yet comparison-eligible (OpenRouter pricing pending) · Context = Max Tokens (input+output) จาก OpenRouter
Model Segmentation

Model Cards

ข้อมูลสรุปแต่ละโมเดล พร้อม feature สำคัญ — จัดกลุ่มตาม capability และ price tier

How to Choose

จะเลือกโมเดลอย่างไรให้เหมาะกับงาน

แยกตาม use case 10 กรณี — Best Choice + Runner-up + เหตุผล

Watchlist / Candidates

Models Pending Full Verification

These models have partially verified AA intelligence scores but are not comparison-eligible (score_eligible=false, comparison_eligible=false). They are excluded from winner cards, runner-up slots, recommendation logic, and comparative ranking until OpenRouter pricing and runtime data are confirmed.

Methodology

วิธีการวิเคราะห์และข้อจำกัด

BENCHMARK
Intelligence Index และ Agentic Index ทุกโมเดลอ้างอิงจาก Artificial Analysis (Estimated) — ไม่ใช่ Confirmed benchmark ที่วัดโดยตรงใน dashboard นี้ · ควรตรวจสอบจาก source หลักอีกครั้งก่อนตัดสินใจ
PLATFORM RUNTIME
ข้อมูล Pricing/Context window อ้างอิงจาก OpenRouter (Confirmed, Platform Runtime). Throughput และ TTFT ใน dashboard นี้มาจาก Artificial Analysis (Estimated, Benchmark Quality) — ไม่ใช่ค่า p50 จาก OpenRouter provider · ค่าจริงอาจต่างกัน
MARKET SIGNAL
Market Signal = adoption & availability proxy based on platform presence, provider/runtime availability, route visibility, and source confidence. It does not represent full market share, token-volume ranking, or absolute model quality.
LIMITED / PREVIEW
Models marked Limited Data (confidence 'L') have insufficient benchmark evidence — intelligence, throughput, TTFT, or pricing fields may be unavailable. Values may change after source verification. 9 models remain evidence gap, 4 partially verified, 3 fully verified. See evidence gap list in docs/model-benchmark-source-matrix.md.
ESTIMATED
ค่า Intelligence Index ทุกโมเดลเป็น Estimated จาก Artificial Analysis — ไม่มีค่า Confirmed สำหรับ Intelligence ในชุดข้อมูลนี้ · ห้ามตีความเป็นข้อสรุปตายตัว
DATA SOURCES & SNAPSHOT CAVEAT
Pricing/Context จาก OpenRouter (Confirmed, Platform Runtime) · Intelligence Index, Throughput, TTFT จาก Artificial Analysis (Estimated, Benchmark Quality). Benchmark and runtime values are treated as a working 2026 snapshot. Market Signal is a proxy layer based on currently available platform/runtime indicators in this file, not a full OpenRouter monthly usage-volume dataset.
MISSING DATA POLICY
ข้อมูลที่ไม่มีจากแหล่ง primary จะแสดงเป็น "Limited Data" — ห้ามสร้างตัวเลขขึ้นเอง · PR 4C update: 9 of 16 active models have evidence gaps. 4 models partially verified (GPT-5.5, Claude Opus 4.7, Gemini 3.5 Flash, Kimi K2.6) with AA intelligence (throughput/TTFT documented in source matrix but excluded from dashboard). 3 models fully verified (Claude Sonnet 4.6, Claude Opus 4.6, Gemini 2.5 Flash) with complete evidence. Partially verified models excluded from comparative filters and winner cards.
OPEN WEIGHTS DEFINITION
Open = weights เปิดเผยสาธารณะ (DeepSeek, GLM, gpt-oss-120b) · Closed = Proprietary API เท่านั้น (Anthropic, Google, OpenAI frontier, xAI). MiniMax M2.7 and Qwen3.7 Max open/closed classification pending verification.
Sources

แหล่งข้อมูล

OPENROUTER - PLATFORM MEASUREMENTS
Metric Scope: Pricing, Context Window, Runtime Throughput, Latency p50
Model Pages & Grouped Sources:
Traceability: per-model cards distinguish direct model pages from grouped lookup references when a stable direct route is not available
● Confirmed Source
SUPPORTING RESEARCH & CAVEATS
Supplementary Notes: PR 4B model refresh — 13 of 16 active models require benchmark source verification
Reporting Conventions:
- Claude Opus 4.6 retained as active_legacy_stable (superseded by Opus 4.7, but continuing production use)
- gpt-oss-120b route status pending verification (not benchmark-eligible until confirmed)
- คะแนน Intelligence ถูกคุมให้คงป้าย Limited / Estimated ทุกจุด
● Methodology Caveat