Real apps, real uncertainty.
Live accounts, changing content, permissions, CAPTCHAs, and network conditions are part of the evaluation.
A real-device benchmark for everyday mobile GUI work.
MobileWorld-Real evaluates whether a mobile GUI agent can complete real-life tasks on live Android devices. Human-written requests run across real apps, accounts, content, and networks—where interfaces change, pop-ups interrupt workflows, and information may be missing.
Everyday tasks, evaluated where people actually use them.
Scroll horizontally to inspect the full figure.

MobileWorld-Real spans 409 tasks, 104 apps, and seven areas of everyday mobile use. The figure shows representative tasks, difficulty coverage, and the long-tailed app distribution.
Scroll horizontally to inspect the full figure.Live accounts, changing content, permissions, CAPTCHAs, and network conditions are part of the evaluation.
Tasks include long-horizon execution, comparison and ranking, deep app entry points, pop-up recovery, and cross-app coordination.
We develop an agent system named AutoJudge to review real-device execution at the trajectory level. AutoJudge examines the task instruction and complete action–screenshot trace, assigns pass, failed, or environment error with a concise rationale, and returns an auditable final outcome.