Personal Agent Evaluation · Global evidence hub

Every personal AI agent, compared on evidence.

PAEval compiles independent evaluations, hands-on tests and official documentation for personal agents worldwide in one traceable place. Every number links to the source that measured it.

44products tracked · 30 fully profiled
3independent product-level score sources
539source-backed capability records
339cited sources across EN · ZH · JA · FR · DE
1
Native scales, never blendedEach third-party score stays on its publisher's scale with its sample, date and coverage beside it.
2
Features are not performanceWhat a vendor documents and what an evaluator measured are shown side by side, never merged.
3
Unknown is not unsupportedMissing evidence stays visibly missing. Nothing is scored as zero by absence.

Personal agents

All tracked products with every independent score available today, each on its source's own scale.

Compare products →
PAEval Index: in development. We are building our own PAEval Index for personal agents. No PAEval score or ranking is published yet; it will appear here, together with its full method, once it is ready. Until then, this page shows independent third-party results on their own scales, alongside documented capabilities.
#Type1–10%0–100
Grok BotxAI (via Cursor) · United StatesIndividual/Team Agent7.3#11 · 7/1530%95% CI 11–60—2of 35 records
InstinctInstinct (Spear Street Technology) · United StatesUniversal Personal Execution Agent8.5#2 · 12/1540%95% CI 17–69—4of 16 records
Meta MuseMeta · United StatesPersistent personal agent9.3#1 · 8/1520%95% CI 6–51—1of 31 records
ArkClawByteDance · ChinaClaw-style agent (SuperCLUE category)——92.61#4 · 5/50of 0 records
AsaplyAsaply Inc. · United StatesLocal consumption/errand transaction type…7.5#10 · 2/15——2of 17 records
AutoClawZhipu AI · ChinaClaw-style agent (SuperCLUE category)——93.23#4 · 5/50of 0 records
CaddyCaddy Technologies Inc. · United StatesMessage-Based Universal Personal/Home…7.3#12 · 9/15——2of 16 records
DuMateBaidu · ChinaClaw-style agent (SuperCLUE category)——93.44#4 · 5/50of 0 records
Gemini SparkGoogle · United StatesPersistent personal agentListed, unscored60%95% CI 31–83—2of 22 records
KimiClawMoonshot AI · ChinaClaw-style agent (SuperCLUE category)——96.01#2 · 5/50of 0 records
MaxClawMiniMax · ChinaClaw-style agent (SuperCLUE category)——94.79#3 · 5/50of 0 records
MiMoClawXiaomi · ChinaClaw-style agent (SuperCLUE category)——96.74#1 · 5/50of 0 records
OllieConfabulation Corporation · United StatesHome personal assistant8.3#6 · 13/15——3of 17 records
OpenAI dotsOpenAI · United StatesPersistent personal agent8.4#3 · 7/15——1of 32 records
PallyPally Technologies · United StatesMulti-channel universal personal execution…8.3#5 · 13/15——2of 17 records
QwenPawAlibaba · ChinaClaw-style agent (SuperCLUE category)——92.97#4 · 5/50of 0 records
ShuffleDNFT · United StatesCreation/entertainment companion…7.8#8 · 10/15——2of 15 records
StepClawStepFun · ChinaClaw-style agent (SuperCLUE category)——91.31#5 · 5/50of 0 records
sznNomadic Futures · United StatesUniversal Life Execution Agent8.2#7 · 13/15——3of 16 records
TomoMapo Labs · United StatesPersonal goals/life companionship and…7.7#9 · 11/15——2of 15 records
TownTown.com · United StatesWork-based personal agents and team…8.3#4 · 7/15——2of 16 records
Wenxin Assistant (task mode)Baidu · ChinaClaw-style agent (SuperCLUE category)——97.62#1 · 5/50of 0 records
WorkBuddy (Tencent)Tencent · ChinaDesktop personal work agent——96.61#2 · 5/50of 16 records
ArkClaw-LiteByteDance · ChinaClaw-style agent (SuperCLUE category)——Mar batch only0of 0 records
ArkClaw-ProByteDance · ChinaClaw-style agent (SuperCLUE category)——Mar batch only0of 0 records
AutoGLM (hosted application)Zhipu / Z.ai · ChinaHosted cloud-device agent———2of 15 records
Claude (including Cowork execution)Anthropic · United StatesGeneral-purpose work agent———2of 24 records
CoPawAlibaba · ChinaClaw-style agent (SuperCLUE category)——Mar batch only0of 0 records
Coze 3.0 / Coze Agent (formerly Coze Space)ByteDance · ChinaHosted personal work agent———0of 14 records
Cue by ManusManus · SingaporeIndependent identity, multi-agent…———2of 19 records
Doubao Phone Assistant (consumer version)ByteDance · ChinaDevice-integrated personal agent———2of 14 records
DuClawBaidu · ChinaClaw-style agent (SuperCLUE category)——Mar batch only0of 0 records
Hermes AgentNous Research · Open source / communitySelf-Hosted Personal Agent———0of 15 records
HONOR YOYO AgentHonor · ChinaDevice-integrated personal agent———0of 13 records
Lindy(Personal Lindy + Teammate)Lindy · United StatesIndividual/Team Agent———1of 22 records
Manus 2.0 / Manus StudioManus · SingaporeGeneral-purpose work agent———2of 29 records
Microsoft Copilot Cowork for personal accounts(Preview)Microsoft · United StatesPersonal Execution Agent (Preview)———2of 20 records
MiniMax Agent/MiniMax Code Work ModeMiniMax · ChinaHybrid general-purpose work agent———1of 12 records
OpenClawOpenClaw Foundation · Open source / communitySelf-Hosted Personal Agent———1of 16 records
Perplexity Computer (Cloud)Perplexity · United StatesCloud execution agent———2of 21 records
Perplexity Personal Computer (local access)Perplexity · United StatesCloud + local execution agent———2of 20 records
QClawTencent · ChinaClaw-style agent (SuperCLUE category)——Mar batch only0of 0 records
Qwen App (task execution / work assistant)Alibaba · ChinaConsumer ecosystem personal agent———1of 14 records
QwenWorkAlibaba / DingTalk · ChinaHybrid personal and team work agent———0of 14 records
Default order: number of independent scored sources, then alphabetical — not a quality ranking. Each source column uses its own scale and is never blended with the others.Which sources are scored →

Independent leaderboards

Three evaluators that test finished products on real tasks. Each publishes a different scale, sample and product set, so read each chart on its own.

All leaderboards →
Assistant Benchmark (independent) · US

Assistant Benchmark

1–10 mean of scored dimensions

025810Meta Muse8/15 dimsMeta Muse: 9.39.3Instinct12/15 dimsInstinct: 8.58.5OpenAI dots7/15 dimsOpenAI dots: 8.48.4Ollie13/15 dimsOllie: 8.38.3Pally13/15 dimsPally: 8.38.3Town7/15 dimsTown: 8.38.3szn13/15 dimsszn: 8.28.2Shuffle10/15 dimsShuffle: 7.87.8Tomo11/15 dimsTomo: 7.77.7Asaply2/15 dimsAsaply: 7.57.5Caddy9/15 dimsCaddy: 7.37.3Grok Bot7/15 dimsGrok Bot: 7.37.3
Updated 8 Oct 2026. Coverage differs by product.Full leaderboard →
SuperCLUE · CN

SuperCLUE-XClaw

0–100 weighted total of 5 dimensions

80859095100Wenxin Assistant…BaiduWenxin Assistant (task mode): 97.6297.62MiMoClawXiaomiMiMoClaw: 96.7496.74WorkBuddy (Tencent)TencentWorkBuddy (Tencent): 96.6196.61KimiClawMoonshot AIKimiClaw: 96.0196.01MaxClawMiniMaxMaxClaw: 94.7994.79DuMateBaiduDuMate: 93.4493.44AutoClawZhipu AIAutoClaw: 93.2393.23QwenPawAlibabaQwenPaw: 92.9792.97ArkClawByteDanceArkClaw: 92.6192.61StepClawStepFunStepClaw: 91.3191.31
Aug 2026 batch. Axis starts at 80. Ties within 1 pt.Full leaderboard →

Capability coverage

What vendors document and third parties observed, by capability domain. A filled cell is evidence of availability, not of task success.

Full matrix →
ProductAutonomy & background workExecution environmentsCalls, messages & identityConnectors & toolsMemory & personalizationChannels & interfacesResearch & deliverablesTransactions & bookingsCollaboration & parallelismReliability & recoveryPermissions, security & data
Meta Muse
OpenAI dots
Grok Bot
Claude (including Cowork execution)
Manus 2.0 / Manus Studio
Gemini Spark
Perplexity Computer (Cloud)
Microsoft Copilot Cowork for personal accounts(Preview)
OpenClaw
Qwen App (task execution / work assistant)
WorkBuddy (Tencent)
Independently observedDocumentedPartly disclosed/testedLimited / conditionalPlannedHistorical / unconfirmedNot supportedNot disclosedUnknownNo record yetIncludes independent evidence
Report · 10 Oct 2026

What independent tests say about personal AI agents (October 2026)

Most agents have never been independently tested, and the evaluators that do test them barely overlap.

Latest update · 10 Oct 2026

Comparisons, capability pages and the first report

  • Added 75 head-to-head comparison pages, generated only for products that share an independent evaluator or the same product type.
  • Added alternatives pages for every profiled product, 11 capability pages and a self-hosted agents page, all built from source-backed records.
  • Added a page per independent evaluator and the report 'What independent tests say about personal AI agents (October 2026)'.

Full changelog →

What PAEval is not

No vendor-reported rankings

Vendor-reported benchmark results are labelled and kept out of scored columns. Model-level benchmarks (a model inside a harness) are listed separately from product-level evaluations. Editorial feature comparisons are cross-checks, not scores.

How sources are classified →

Frequently asked questions

What is a personal AI agent?

A personal AI agent acts for one person over time rather than only answering questions: it can carry out tasks across apps, accounts and devices, remember context, and keep working in the background. PAEval covers persistent assistants, delegated work agents, device-integrated agents and self-hosted agents; chat-only assistants and developer coding agents are out of scope unless they offer personal delegation.

Which personal AI agents score highest in independent evaluations?

It depends on the evaluator, and their scores cannot be combined. On Assistant Benchmark (1–10 scale, updated 8 Oct 2026), Meta Muse has the top score of 9.3 among 12 scored products, based on 8/15 dimensions. In micro1 PersonalAgentBench, Gemini Spark completed 6 of 10 workflows with trusted completion (60%, 95% interval 31–83%); the intervals of all tested products overlap, so the ranking is not statistically separated. In SuperCLUE-XClaw's Aug 2026 batch of Chinese-language tasks, Wenxin Assistant (task mode) leads with 97.62/100; products within one point share a tier.

Does PAEval rank personal AI agents?

Not yet. The PAEval Index is in development and no PAEval score or ranking is published. The site shows each independent evaluator's results on that evaluator's own scale, next to documented capabilities, and never averages scores across evaluators.

Where does the PAEval data come from?

From 339 cited sources: independent product-level evaluations, hands-on reviews, and vendors' official documentation, pricing and policy pages. PAEval tracks 44 products, 30 with full profiles, and 539 capability records. Each record keeps its source URL, date, evidence type and any region or plan conditions. Unknown is recorded as unknown, never as unsupported.

Can I download or cite the data?

Yes. Products, scores, capability records and sources are downloadable as JSON and CSV from https://paeval.com/about#data. Cite as: PAEval, snapshot 2026-10-09, https://paeval.com. When quoting a score, also cite the original evaluator.