ระบบ Confidence Label
ทุก claim ในรายงานนี้ติดป้ายระดับความเชื่อมั่น เพื่อให้ใช้ตัดสินใจได้อย่างปลอดภัย:
1. Executive Summary
Hermes Agent คือ open-source AI agent ของ Nous Research ที่วางตัวเป็น "agent infrastructure" มากกว่า single-purpose assistant จุดต่างหลักคือ closed learning loop (สร้าง/ปรับ skill เองจากประสบการณ์) + persistent cross-session memory + multi-platform gateway + multi-backend deployment
ข้อสรุปเชิงกลยุทธ์:
- โครงการนี้ ของจริงและ active สูงมาก — ไม่ใช่ vaporware ✅ VERIFIED ยืนยัน ~203k stars, 12,883+ commits, 1,547 contributors, MIT license — ค่า ณ 27 มิ.ย. 2026
- จุดขายด้าน architecture (self-host, provider-agnostic, learning loop) ยืนยันได้จาก primary source
- จุดขายด้าน performance: MoA มีตัวเลข HermesBench ใน docs ทางการแล้ว (MoA 0.8202 > Opus 4.8 0.7607 > GPT-5.5 0.7412) แต่ยังเป็น benchmark ของผู้พัฒนาเอง 🟡 VENDOR-CLAIM — ห้ามใช้เป็นข้อสรุปว่า "ชนะ frontier models" แบบ decision-grade
- เหมาะกับ use case ส่วนตัว/ทีมเล็กที่ self-host ได้ ยังไม่ใช่ production-grade enterprise tool ที่ proven
2. ข้อมูลผลิตภัณฑ์ ✅ VERIFIED
| รายการ | ข้อมูล | สถานะ |
|---|---|---|
| ผู้พัฒนา | Nous Research | ✅ VERIFIED |
| License | MIT (open-source สมบูรณ์) | ✅ VERIFIED |
| วันเปิดตัว | 25 กุมภาพันธ์ 2026 | ✅ VERIFIED (หลายแหล่งตรง) |
| เวอร์ชันล่าสุด | v0.17.0 (v2026.6.19), release 19 มิ.ย. 2026 | ✅ VERIFIED (GitHub releases) |
| แพลตฟอร์ม | macOS 12+, Windows 10/11 (native), Linux/any distro, Termux | ✅ VERIFIED |
| ภาษาหลัก | Python 82.5%, TypeScript 13.6% | ✅ VERIFIED |
| GitHub stars | ~203k–204k (ณ 27 มิ.ย. 2026) | ✅ VERIFIED with timestamp |
| Commits / Contributors | 12,883+ commits · 1,547 contributors (ณ 27 มิ.ย. 2026) | ✅ VERIFIED with timestamp |
| Open issues / PRs | 5k+ / 5k+ | ✅ VERIFIED |
3. Features หลัก ✅ VERIFIED
ทั้งหมดในตารางนี้ยืนยันจากเว็บทางการและ README:
| Feature | รายละเอียด |
|---|---|
| Lives Everywhere | Gateway เดียวต่อ Telegram, Discord, Slack, WhatsApp, Signal, Email, CLI — one agent, one memory, every surface |
| Persistent Memory | จำข้าม session, auto-generate skills, FTS5 session search + LLM summarization, Honcho dialectic user modeling |
| Closed Learning Loop | สร้าง skill หลังงานซับซ้อน, skill ปรับปรุงตัวเองระหว่างใช้, periodic nudges ให้เก็บความรู้ — กลไกจริง: ไม่ใช่ fine-tuning weights แต่เป็น prompt-driven text manipulation: agent เขียน/แก้ markdown skill files (agentskills.io format) ใน ~/.hermes/skills/ ผ่าน skill_manage tool ใช้ Progressive Disclosure Pattern (Level 0–2) เพื่อคุม token |
| Scheduled Automation | Built-in cron scheduler สั่งด้วยภาษาธรรมชาติ ส่งผลไปแพลตฟอร์มใดก็ได้ รันแบบ unattended |
| Subagents / Parallelize | Spawn subagent แยก + Python RPC scripts ทำ pipeline แบบ zero-context-cost |
| Browse / Multimodal | Web search, browser automation, vision, image gen, TTS, multi-model reasoning |
| 6 Terminal Backends | local, Docker, SSH, Singularity, Modal, Daytona (Modal/Daytona = serverless hibernate-on-idle) — README/docs ระบุ 6, homepage บางจุดแสดง 5; ใช้คำว่า "up to six" |
| 40+ / 60+ tools | docs บางหน้าระบุ "60+ built-in tools" บางหน้าใช้ "40+" — ยึดเป็น lower bound พร้อมวันที่ docs snapshot |
| MCP Integration | ต่อ MCP server ใดก็ได้ |
| Provider-agnostic | 300+ models ผ่าน Nous Portal / OpenRouter / NIM / Kimi / GLM / OpenAI / endpoint เอง — hermes model สลับได้ ไม่ lock-in |
| OpenClaw migration | hermes claw migrate นำเข้า SOUL.md, memories, skills, allowlist, API keys อัตโนมัติ |
4. Mixture of Agents (MoA) — จุดที่ต้องระวังที่สุด
สิ่งที่ ✅ VERIFIED
- MoA เป็นฟีเจอร์จริง อยู่ใน docs ทางการ ทำงานเป็น "virtual model provider" เลือกเป็น preset ได้
- กลไก: reference model ตอบก่อน → ส่ง output เป็น private context → aggregator model สังเคราะห์คำตอบจริง (reference models ได้ trimmed view: system prompt + tool transcript ถูกตัดออก)
- Config จริงจาก docs (default preset): เป็น 2-model preset — reference =
gpt-5.5, aggregator =claude-opus-4.8 - MoA ออกแบบให้ prompt cache ของบทสนทนาหลักไม่แตก — การเลือก MoA preset เป็น model selection ปกติ ไม่ทำลาย cached prefix
- อ้างอิงงานวิจัยจริง: MoA pattern (Wang et al., arXiv:2406.04692)
reference_models/aggregator ได้ผ่าน YAML แต่ raw source tools/mixture_of_agents_tool.py ยังมีโครงสร้าง fixed-pipeline + legacy constants — เป็นช่องว่างระหว่าง abstraction layer กับ execution layer ที่ทีมกำลัง unify; ให้ยึด docs ทางการเป็นหลัก และระบุว่า implementation จริงอาจยังไม่ตรง docs 100%
การเลือก MoA preset เดียว = ส่งข้อมูลออกไป ≥2 frontier providers ต่อ 1 task (gpt-5.5 reference + opus-4.8 aggregator) โดยผู้ใช้อาจไม่รู้ตัว สำหรับงาน client/government ที่มีข้อกำหนด data residency / PDPA นี่ไม่ใช่แค่ overhead เรื่อง cost — มันคือ potential compliance breach ที่ทำให้เสียสัญญารัฐได้ ต้องถือเป็น hard constraint: ห้ามใช้ MoA (หรือ cloud model ใดๆ) กับข้อมูล sensitive ของ client โดยไม่มี data-flow review ก่อน
HermesBench Scores 🟡 VENDOR-CLAIM
มีตัวเลขใน docs ทางการแล้ว (ต่างจากฉบับก่อนที่ระบุว่า "ยังไม่เผยแพร่"):
docs ระบุว่า MoA ชนะ strongest component (Opus 4.8) ราว ~6 points — การตีความตัวเลข "8% / 11%" ของทีม (X, 26 มิ.ย.) สอดคล้องกับ relative uplift: MoA vs Opus = +0.0595 (≈ +7.8% relative), MoA vs GPT-5.5 = +0.0790 (≈ +10.7% relative) — แต่ถ้าตีความเป็น "percentage points" จะผิด
ทำไมยังคง label 🟡 VENDOR-CLAIM:
- HermesBench เป็น benchmark ของผู้พัฒนาเอง — ยังไม่พบ independent/public leaderboard ที่มี methodology, task set, judge model, reproducibility script และ submission process ครบ
- ข้อเท็จจริงเชิงโครงสร้าง: MoA preset ใช้ Claude Opus 4.8 เป็น aggregator เอง ดังนั้นนี่คือ "Opus-aggregated ensemble ชนะ Opus เดี่ยว" ไม่ใช่ model ใหม่ชนะ Opus แบบ standalone — wording ที่แม่นคือ "ตาม HermesBench ที่ Nous เผยแพร่ใน docs, MoA preset ที่ใช้ Opus 4.8 เป็น aggregator เหนือ gpt-5.5 reference ทำคะแนนสูงกว่า Opus 4.8 และ GPT-5.5 เดี่ยว ใน benchmark ของทีมเอง"
- ⚠️ COMPUTE-FAIRNESS: HermesBench ไม่ได้ระบุ inference-time compute / token budget ที่ MoA ใช้เทียบ baseline — MoA ใช้ ≥2–3x compute (reference + aggregator + loop) ถ้าให้ Opus เดี่ยวใช้ compute เท่ากัน (best-of-N, self-consistency, extended thinking) อาจชนะ MoA; claim ที่ defensible คือ "MoA > Opus ที่ compute-matched — ยังไม่พิสูจน์"
- ⚠️ CITATION laundering: การอ้าง Wang et al. (arXiv:2406.04692) รองรับแค่ pattern ของ MoA ไม่ได้รองรับตัวเลขเฉพาะของ HermesBench — ต้องแยกให้ชัดว่า "pattern มีงานวิจัยรอง แต่ result ของ Hermes ไม่มี independent backing"
4b. MoA Cost & Governance Risk
✅ VERIFIED MoA ไม่ใช่ single-call path — reference model ถูกเรียกก่อน แล้วจึงเรียก aggregator; ถ้า aggregator เรียก tools, agent loop รอบถัดไปทำซ้ำ ทำให้ model-call count คูณตามจำนวน reference + aggregator ต่อ iteration
สูตรประเมิน cost คร่าวๆ:
MoA cost/task ≈ Σ(reference model I/O) + aggregator model I/O + tool-loop repeat
🟠 COMMUNITY มี GitHub issue (#38952) ขอให้ MoA configurable เพราะ implementation ช่วงหนึ่ง hardcoded frontier models และกังวลเรื่อง cost (อ้างถึง $2+/call) — ใช้เป็น risk signal ไม่ใช่ verified defect ของเวอร์ชันล่าสุด
Cost control ที่แนะนำ:
| Control | เหตุผล |
|---|---|
| ห้ามตั้ง MoA เป็น default model สำหรับ cron/gateway ทุกงาน | model-call count เพิ่มตาม references + aggregator |
| แยก profile/API key สำหรับ MoA | ลด blast radius ของ cost + credential exposure |
| ตั้ง budget alert ที่ provider/OpenRouter/Nous Portal | governance baseline |
| ใช้ cheap auxiliary models สำหรับงานย่อย | docs รองรับ auxiliary model slots |
| ใช้ MoA เฉพาะงาน high-value / hard reasoning | วาง MoA เป็น "premium reasoning mode" ไม่ใช่ default |
5. Use Cases จริง 🟠 COMMUNITY
ทั้งหมดมาจากผู้ใช้/blog รอง ยังไม่ cross-confirm ระดับ primary:
| ด้าน | การใช้งาน | ผลลัพธ์ที่อ้าง |
|---|---|---|
| Automation / Monitoring | ตั้ง cron บน VPS ตรวจ GitHub repo, สรุป PR/issue, ส่ง daily digest เข้า Telegram/Slack | ลดการพึ่ง n8n/Zapier |
| Coding / Orchestration | Hermes เป็น orchestrator สั่งงานจาก Telegram → spawn Claude Code/Codex แก้โค้ดใน Docker sandbox | จำ convention โปรเจกต์ผ่าน MEMORY.md/USER.md |
| Research / Knowledge | ต่อ MCP เข้าเอกสารองค์กร, inbox triage, ค้นนโยบายผ่าน Slack | Programmatic tool calling รวมหลายขั้นใน inference เดียว |
| Content Generation | Long-form writing, แตก subagent ตรวจข้อมูล | 🔴 UNVERIFIED เคส "นิยาย 79,000 คำ" = anecdote ชั้นเดียว |
| Personal Assistant | จัดการ email/home automation, draft replies ด้วย memory | HN user: "rules my personal life" |
6. ข้อจำกัด / จุดอ่อน
- Context Bloat ✅ VERIFIED — release notes + community ยืนยันว่า system prompt + 30+ tool schemas + memory injection ทำให้ first-turn token บวมเกิน 10,000 tokens ส่งผลให้โมเดลเล็ก (<14B) รวน/หลอนคำสั่งง่าย — ทีมแก้ด้วยการ trim default skill set + relevance gate ในเวอร์ชันล่าสุดแล้ว
- CLI หนัก/ช้าตอนเริ่ม 🟠 COMMUNITY — background processes เยอะ
- Cross-domain generalization จำกัด 🟠 COMMUNITY — skill ที่สร้างเองใช้ได้ดีกับงานซ้ำเดิม แต่ข้ามสายงานยังไม่สมบูรณ์
- Self-overwrite / hallucinated skill edits 🟠 COMMUNITY (theme verify ได้ เลข issue unverified) — ระบบที่ agent แก้ skill file เองมีรายงานว่าบางครั้ง เขียนทับ user-set instructions หรือสร้าง correction ที่หลอน — กระทบ accountability โดยตรง
- Prompt Injection risk ✅ VERIFIED — มีระบบ permission/approval จริง แต่ถ้าเปิด YOLO mode เสี่ยงสูง
- Setup ซับซ้อนสำหรับ beginner — Docker/permissions/sandboxing
- ยัง early สำหรับ production — บาง feature เป็น alpha + velocity สูง = breaking changes บ่อย
6b. Security & Enterprise Readiness Checklist
✅ VERIFIED Hermes มี security controls หลายชั้นใน docs ทางการ — user authorization, command approval modes, container isolation, MCP env filtering, SSRF protection, context-file scanning, cross-session isolation/path-traversal hardening, input sanitization และ production checklist
| Area | Checklist | Label |
|---|---|---|
| Runtime isolation | ใช้ Docker/Modal/Daytona/Singularity กับงานที่แตะ external content / client data; เลี่ยง local backend กับ untrusted workload | Low |
| Approval mode | ใช้ manual/smart approval; ห้าม YOLO/approval-off ใน env ที่มีข้อมูลจริง | Low |
| Gateway access | ตั้ง explicit allowlist สำหรับ Telegram/Discord/Slack/Email; ห้าม allow-all บน public gateway | High |
| Secrets | ห้าม forward host secrets เข้า container โดยไม่จำเป็น; review docker_forward_env | High |
| MCP credentials | แยก credentials ต่อ MCP server; เลี่ยง broad env passthrough; ตรวจ redaction/logging | High |
| SSRF / web access | คง allow_private_urls: false บน public deployment; block private IP / metadata endpoints | High |
| Context injection | ทดสอบ AGENTS.md / SOUL.md / README injection + context-file scanner ก่อนใช้จริง | High |
| File permissions | กำหนด terminal.cwd เป็น workspace เฉพาะ; run as non-root; จำกัด mount path | High |
| Cost abuse | แยก API key สำหรับ gateway/cron; ตั้ง quota/budget; log model calls | High |
| Audit evidence | เก็บ config snapshot, version, logs, allowlist, provider bill, test cases | High |
| Enterprise assurance | ต้องมี independent review/pen-test ก่อนใช้กับข้อมูลลูกค้า/รัฐ | Critical |
Attack surface เพิ่มเติมที่ต้องระวัง
-
✅ VERIFIED Honcho customer-facing recall leak (issue #40170):
เมื่อใช้ Honcho เป็น memory provider + เปิด customer-facing gateway (เช่น WhatsApp) ระบบ inject
<memory-context>block ที่มี operator observations (contacts, infrastructure, prior sessions) เข้า reply ที่ส่งถึง end customer — ข้อมูล operator รั่วถึงลูกค้าโดยตรง โค้ดอยู่ที่agent/conversation_loop.py→ ถ้าจะใช้ Hermes เป็น customer-facing bot ต้องปิด Honcho หรือ verify การ isolate ก่อน Critical - 🟠 COMMUNITY (theme verify ได้ เลข unverified) Subagent constraint inheritance: มีรายงานว่า subagent ไม่ inherit parent behavioral constraints อัตโนมัติ — ถ้าใช้ delegation กับงาน sensitive ต้อง set constraint ที่ระดับ subagent เอง High
-
✅ VERIFIED มีเครื่องมือช่วย:
hermesมี supply-chain audit (OSV.dev scan) + secret redaction (เปิด default) + config validation — ใช้เป็น baseline ได้
6c. Risk-Ownership: Continuity, Exit Cost & Accountability
Maintainer / Vendor Continuity Risk Critical
- Hermes ทั้ง stack ขึ้นกับ Nous Research องค์กรเดียว ซึ่งเป็น research lab ไม่ใช่ infra company
- Monetization เดียวคือ Nous Portal — ถ้า Portal ไม่ทำเงิน, Nous pivot, หรือถูก acquire → project เสี่ยงถูกลด priority กลางสัญญา client
- 🟡 VENDOR-CLAIM bus factor: แม้มี 1,547 contributors แต่ commit หลักกระจุกที่ teknium1 + core team — ตรวจ contributor distribution จริงก่อนพึ่งพาระยะยาว
Hidden Lock-in ผ่าน State Format High
- lock-in จริงไม่ได้อยู่ที่ license (MIT แก้ปัญหานั้นแล้ว) — อยู่ที่ memory format, skill format, SOUL.md/MEMORY.md convention, learning-loop state ที่ไม่มี portable standard
- เมื่อทีมสะสม skills + persistent memory เป็นพันรายการ การ migrate ออก จาก Hermes ทำไม่ได้ง่าย — ต้องประเมิน exit path ก่อน adopt
- 🟡 soft lock-in ผ่าน Portal: convenience ของ OAuth เดียว + 300 models + Tool Gateway ทำให้ทีม default ไปใช้ Portal แล้วถอนยาก
Accountability Gap ของ Self-Generated Artifacts Critical
- Closed learning loop = agent สร้าง/แก้ skill เอง → skill ที่ generate เองอาจมี logic ผิด → reused เงียบๆ ข้ามหลาย client task → ตรวจย้อนหลังไม่ได้ว่าตัดสินใจมาจาก skill เวอร์ชันไหน
- สำหรับ consulting ที่ต้องรับผิดต่อ deliverable นี่คือ non-deterministic liability ที่ไม่มี audit trail
- Mitigation: เปิด trajectory saving + version control skill directory (git) + แยก user-locked skills ออกจาก agent-editable
Total Cost of Ownership จริง High
- "$5 VPS / idle ≈ 0" คือ infra cost ที่ถูกที่สุด ไม่ใช่ TCO จริง
- TCO จริง = maintainer time ของทีม 3 คน ในการ track breaking changes (velocity สูงมาก: v0.16→v0.17 = ~1,475 commits, 235k insertions) + re-audit security ทุก release ที่สำคัญ
Exit Cost / Reversibility Critical
ยังไม่มีใครประเมินว่า "ถ้า adopt แล้วต้องถอย" ต้นทุนเท่าไร — ต้องกำหนด exit criteria + portable data export ก่อน commit ทีมเข้า workflow รอบ Hermes
Competitive Displacement Risk High
🟠 COMMUNITY agent space เคลื่อนเร็วมาก — adopt วันนี้อาจ legacy ใน 12 เดือน ออกแบบ adoption ให้ reversible + provider-agnostic เพื่อกัน lock-in กับ trajectory ที่อาจถูกแซง
7. เปรียบเทียบคู่แข่ง
7.1 vs OpenClaw (คู่เทียบตรงที่สุด) 🟠 COMMUNITY
| ด้าน | Hermes | OpenClaw | หมายเหตุ |
|---|---|---|---|
| Core strength | Self-improving loop | Skill catalog ใหญ่ (ClawHub) | ต่างปรัชญา |
| Memory | ลึกกว่า (FTS5 + user modeling) | ดี แต่ตื้นกว่า | Hermes เด่น |
| Setup | ซับซ้อน (2–4 ชม.) | ง่าย (<30 นาที) | OpenClaw เด่น |
| Stability | ดีขึ้นเรื่อยๆ | รายงาน hang บ่อย | ขึ้นกับแหล่ง |
| Ecosystem | เล็กแต่ technical | ใหญ่/mature | OpenClaw เด่น |
hermes claw migrate นำเข้าจาก OpenClaw โดยตรง — สะท้อนว่า Hermes positioning ตัวเองเป็น upgrade path จาก OpenClaw อย่างเป็นทางการ นี่คือ signal เชิงกลยุทธ์ที่หนักแน่นกว่าคำรีวิว
7.2 vs ChatGPT Agent / Claude / Manus / AutoGPT
- เทียบได้ (architecture): Hermes เด่นกว่าเรื่อง self-host, provider-agnostic, learning loop, multi-platform gateway — ✅ VERIFIED เพราะคู่แข่งส่วนใหญ่เป็น closed/cloud product
- เทียบไม่ได้ (performance/quality): 🔴 UNVERIFIED ไม่มีหลักฐานว่า reasoning/tool-use ของ Hermes ชนะ Claude/GPT — อย่าสรุป
8. Business Model / Pricing ✅ VERIFIED
- ตัว agent = ฟรี (MIT) รันด้วย API key ของตัวเองได้เต็มที่
- Nous Portal = subscription เสริม (tier Free/Plus/Super/Ultra) ให้ 300+ models + Tool Gateway (web search/image/TTS/browser) ผ่าน OAuth เดียว — เป็น ทางเลือก ไม่บังคับ
- รันได้ถูกมาก: $5 VPS, หรือ serverless (Modal/Daytona) ที่ idle cost ≈ 0
- 🟠 ราคา tier เฉพาะยังไม่ระบุตัวเลขในแหล่งที่ตรวจ — ต้องดู portal.nousresearch.com โดยตรง
9. ข้อสรุปสำหรับการตัดสินใจ
จุดแข็งที่ defensible (ใช้อ้างได้)
- ✅ VERIFIED Open-source MIT แท้ + self-host + provider-agnostic = ไม่มี vendor lock-in
🟡 CONDITIONALLY TRUE "ควบคุมข้อมูล": ระดับ data sovereignty ขึ้นกับ deployment mode, model provider, Tool Gateway, browser/web access — อย่าเขียนว่า "ควบคุมข้อมูล 100%" เว้นแต่ระบุ architecture ที่ปิด external data path ชัดเจน - ✅ VERIFIED Architecture maturity จริง (gateway, cron, subagents, MCP, up to 6 backends, learning loop)
- ✅ VERIFIED Adoption signal จริง (~203k stars, #1 OpenRouter token usage) — เป็น usage signal ที่หนักแน่น แต่ไม่พิสูจน์ quality/retention โดยลำพัง
จุดที่ต้องระวัง (อย่าใช้เป็นจุดขาย)
- 🟡 VENDOR-CLAIM "แซง frontier models" = vendor-claim — HermesBench มีตัวเลขแล้ว (MoA 0.8202 vs Opus 0.7607 vs GPT-5.5 0.7412) แต่เป็น benchmark ของผู้พัฒนาเอง + ใช้ Opus เป็น aggregator ยังไม่มี independent leaderboard
- 🟠 COMMUNITY Production-readiness ยังไม่ proven, security controls มีจริงแต่ต้อง sandbox + ยังไม่มี public audit
- 🔴 UNVERIFIED ตัวเลขประสิทธิภาพเฉพาะจาก community (113ms, 79k คำ, token counts) = ยังไม่ verify
- 🟡 MoA = premium reasoning mode ที่มี cost overhead จริง (extra reference calls) — อย่าตั้งเป็น default
คำแนะนำการใช้งาน
ใช้ Hermes เป็น internal presales accelerator บน non-sensitive data — สรุป TOR/BOQ, แตก subagent ตรวจเอกสาร, draft architecture doc, ต่อ MCP เข้า Obsidian งานกลุ่มนี้ ไม่แตะ client data sensitive และ leverage จุดแข็งจริง (memory + subagent + MCP) ได้เต็มที่ ความเสี่ยง compliance ต่ำ เริ่มได้เลยไม่ต้องรอ
เขียนเชิง "open-source agent ที่มาแรง + architecture น่าสนใจ" ได้ แต่ต้องใส่ 3 caveat ก่อน publish: (1) compute-fairness ของ MoA, (2) แยก citation laundering ของ Wang et al., (3) frame stars ว่า "จริงแต่ตรวจ organic-ness ไม่ได้"
อย่าเพิ่ง — ต้องผ่าน self-audit ก่อน: data-flow review (กัน MoA/cloud egress), security checklist section 6b, exit-cost plan, accountability/audit trail สำหรับ self-generated skills เมื่อผ่านครบแล้วค่อยขยับเข้า production ทีละ scope
ภาคผนวก: แหล่งอ้างอิง Primary
- เว็บทางการ:
hermes-agent.nousresearch.com - GitHub repo:
github.com/NousResearch/hermes-agent(MIT, v0.17.0) - MoA docs + HermesBench:
hermes-agent.nousresearch.com/docs/user-guide/features/mixture-of-agents - Security docs:
hermes-agent.nousresearch.com/docs/user-guide/security/ - 🟠 COMMUNITY MoA cost issue: GitHub issue #38952
- ประกาศ MoA claim: Nous Research X, 26 มิ.ย. 2026
- MoA research base: Wang et al., arXiv:2406.04692
- ✅ VERIFIED Honcho recall leak: GitHub issue #40170
- v0.17.0 release stats: ~1,475 commits / ~800 PRs since v0.16.0
GitHub issue บางเลขที่ reviewer ภายนอกอ้าง verify ไม่ได้ในการตรวจครั้งนี้ จึงรับเฉพาะ theme และ label เป็น "เลข unverified" — ไม่อ้างเลข issue เหล่านั้นต่อสาธารณะ
Final Positioning
Hermes Agent คือ open-source agent infrastructure ที่ active สูงและ architecture น่าสนใจมาก — แต่ benchmark/performance claim โดยเฉพาะ MoA ต้องจัดเป็น vendor-claim จนกว่าจะมี independent leaderboard ที่ reproducible
วาง Hermes เป็น architecture/adoption story ไม่ใช่ benchmark-winner story — และเพิ่ม risk-ownership story (continuity, exit cost, data-egress, accountability) ซึ่งเป็นจุดที่ profile คุณควรเด่นที่สุด
นี่ทำให้รายงานแข็งกว่า content ทั่วไปที่ขาย hype และรักษามาตรฐาน verify-first ได้ครบ
Balance: ไม่ได้แปลว่า "ห้ามใช้" — wedge ที่ปลอดภัย (internal presales บน non-sensitive data) คือ recommended-yes ที่เริ่มได้เลย ความระมัดระวังอยู่ที่ client/government data เท่านั้น ไม่ใช่ blanket block