Every personal AI agent, compared on evidence.
PAEval compiles independent evaluations, hands-on tests and official documentation for personal agents worldwide in one traceable place. Every number links to the source that measured it.
Personal agents
All tracked products with every independent score available today, each on its source's own scale.
| # | 1–10 | % | 0–100 | ||
|---|---|---|---|---|---|
| Grok BotxAI (via Cursor) · United States | 7.3 | 30% | — | 2of 35 records | |
| InstinctInstinct (Spear Street Technology) · United States | 8.5 | 40% | — | 4of 16 records | |
| Meta MuseMeta · United States | 9.3 | 20% | — | 1of 31 records | |
| ArkClawByteDance · China | — | — | 92.61 | 0of 0 records | |
| AsaplyAsaply Inc. · United States | 7.5 | — | — | 2of 17 records | |
| AutoClawZhipu AI · China | — | — | 93.23 | 0of 0 records | |
| CaddyCaddy Technologies Inc. · United States | 7.3 | — | — | 2of 16 records | |
| DuMateBaidu · China | — | — | 93.44 | 0of 0 records | |
| Gemini SparkGoogle · United States | Listed, unscored | 60% | — | 2of 22 records | |
| KimiClawMoonshot AI · China | — | — | 96.01 | 0of 0 records | |
| MaxClawMiniMax · China | — | — | 94.79 | 0of 0 records | |
| MiMoClawXiaomi · China | — | — | 96.74 | 0of 0 records | |
| OllieConfabulation Corporation · United States | 8.3 | — | — | 3of 17 records | |
| OpenAI dotsOpenAI · United States | 8.4 | — | — | 1of 32 records | |
| PallyPally Technologies · United States | 8.3 | — | — | 2of 17 records | |
| QwenPawAlibaba · China | — | — | 92.97 | 0of 0 records | |
| ShuffleDNFT · United States | 7.8 | — | — | 2of 15 records | |
| StepClawStepFun · China | — | — | 91.31 | 0of 0 records | |
| sznNomadic Futures · United States | 8.2 | — | — | 3of 16 records | |
| TomoMapo Labs · United States | 7.7 | — | — | 2of 15 records | |
| TownTown.com · United States | 8.3 | — | — | 2of 16 records | |
| Wenxin Assistant (task mode)Baidu · China | — | — | 97.62 | 0of 0 records | |
| WorkBuddy (Tencent)Tencent · China | — | — | 96.61 | 0of 16 records | |
| ArkClaw-LiteByteDance · China | — | — | Mar batch only | 0of 0 records | |
| ArkClaw-ProByteDance · China | — | — | Mar batch only | 0of 0 records | |
| AutoGLM (hosted application)Zhipu / Z.ai · China | — | — | — | 2of 15 records | |
| Claude (including Cowork execution)Anthropic · United States | — | — | — | 2of 24 records | |
| CoPawAlibaba · China | — | — | Mar batch only | 0of 0 records | |
| Coze 3.0 / Coze Agent (formerly Coze Space)ByteDance · China | — | — | — | 0of 14 records | |
| Cue by ManusManus · Singapore | — | — | — | 2of 19 records | |
| Doubao Phone Assistant (consumer version)ByteDance · China | — | — | — | 2of 14 records | |
| DuClawBaidu · China | — | — | Mar batch only | 0of 0 records | |
| Hermes AgentNous Research · Open source / community | — | — | — | 0of 15 records | |
| HONOR YOYO AgentHonor · China | — | — | — | 0of 13 records | |
| Lindy(Personal Lindy + Teammate)Lindy · United States | — | — | — | 1of 22 records | |
| Manus 2.0 / Manus StudioManus · Singapore | — | — | — | 2of 29 records | |
| Microsoft Copilot Cowork for personal accounts(Preview)Microsoft · United States | — | — | — | 2of 20 records | |
| MiniMax Agent/MiniMax Code Work ModeMiniMax · China | — | — | — | 1of 12 records | |
| OpenClawOpenClaw Foundation · Open source / community | — | — | — | 1of 16 records | |
| Perplexity Computer (Cloud)Perplexity · United States | — | — | — | 2of 21 records | |
| Perplexity Personal Computer (local access)Perplexity · United States | — | — | — | 2of 20 records | |
| QClawTencent · China | — | — | Mar batch only | 0of 0 records | |
| Qwen App (task execution / work assistant)Alibaba · China | — | — | — | 1of 14 records | |
| QwenWorkAlibaba / DingTalk · China | — | — | — | 0of 14 records |
Independent leaderboards
Three evaluators that test finished products on real tasks. Each publishes a different scale, sample and product set, so read each chart on its own.
Capability coverage
What vendors document and third parties observed, by capability domain. A filled cell is evidence of availability, not of task success.
| Product | Autonomy & background work | Execution environments | Calls, messages & identity | Connectors & tools | Memory & personalization | Channels & interfaces | Research & deliverables | Transactions & bookings | Collaboration & parallelism | Reliability & recovery | Permissions, security & data |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Meta Muse | |||||||||||
| OpenAI dots | |||||||||||
| Grok Bot | |||||||||||
| Claude (including Cowork execution) | |||||||||||
| Manus 2.0 / Manus Studio | |||||||||||
| Gemini Spark | |||||||||||
| Perplexity Computer (Cloud) | |||||||||||
| Microsoft Copilot Cowork for personal accounts(Preview) | |||||||||||
| OpenClaw | |||||||||||
| Qwen App (task execution / work assistant) | |||||||||||
| WorkBuddy (Tencent) |
Popular comparisons
Find agents that…
What independent tests say about personal AI agents (October 2026)
Most agents have never been independently tested, and the evaluators that do test them barely overlap.
Comparisons, capability pages and the first report
- Added 75 head-to-head comparison pages, generated only for products that share an independent evaluator or the same product type.
- Added alternatives pages for every profiled product, 11 capability pages and a self-hosted agents page, all built from source-backed records.
- Added a page per independent evaluator and the report 'What independent tests say about personal AI agents (October 2026)'.
No vendor-reported rankings
Vendor-reported benchmark results are labelled and kept out of scored columns. Model-level benchmarks (a model inside a harness) are listed separately from product-level evaluations. Editorial feature comparisons are cross-checks, not scores.
Frequently asked questions
What is a personal AI agent?
A personal AI agent acts for one person over time rather than only answering questions: it can carry out tasks across apps, accounts and devices, remember context, and keep working in the background. PAEval covers persistent assistants, delegated work agents, device-integrated agents and self-hosted agents; chat-only assistants and developer coding agents are out of scope unless they offer personal delegation.
Which personal AI agents score highest in independent evaluations?
It depends on the evaluator, and their scores cannot be combined. On Assistant Benchmark (1–10 scale, updated 8 Oct 2026), Meta Muse has the top score of 9.3 among 12 scored products, based on 8/15 dimensions. In micro1 PersonalAgentBench, Gemini Spark completed 6 of 10 workflows with trusted completion (60%, 95% interval 31–83%); the intervals of all tested products overlap, so the ranking is not statistically separated. In SuperCLUE-XClaw's Aug 2026 batch of Chinese-language tasks, Wenxin Assistant (task mode) leads with 97.62/100; products within one point share a tier.
Does PAEval rank personal AI agents?
Not yet. The PAEval Index is in development and no PAEval score or ranking is published. The site shows each independent evaluator's results on that evaluator's own scale, next to documented capabilities, and never averages scores across evaluators.
Where does the PAEval data come from?
From 339 cited sources: independent product-level evaluations, hands-on reviews, and vendors' official documentation, pricing and policy pages. PAEval tracks 44 products, 30 with full profiles, and 539 capability records. Each record keeps its source URL, date, evidence type and any region or plan conditions. Unknown is recorded as unknown, never as unsupported.
Can I download or cite the data?
Yes. Products, scores, capability records and sources are downloadable as JSON and CSV from https://paeval.com/about#data. Cite as: PAEval, snapshot 2026-10-09, https://paeval.com. When quoting a score, also cite the original evaluator.