MobileWorld-Real

A real-device benchmark for everyday mobile GUI work.

MobileWorld-Real evaluates whether a mobile GUI agent can complete real-life tasks on live Android devices. Human-written requests run across real apps, accounts, content, and networks—where interfaces change, pop-ups interrupt workflows, and information may be missing.

409Human-written end-to-end tasks
104Live Android apps
7Everyday-use domains
HELD OUTTasks and trajectories excluded from training

BENCHMARK PROFILE

Everyday tasks, evaluated where people actually use them.

Scroll horizontally to inspect the full figure.

MobileWorld-Real spans 409 tasks, 104 apps, and seven areas of everyday mobile use. The figure shows representative tasks, difficulty coverage, and the long-tailed app distribution.

MobileWorld-Real spans 409 tasks, 104 apps, and seven areas of everyday mobile use. The figure shows representative tasks, difficulty coverage, and the long-tailed app distribution.

Scroll horizontally to inspect the full figure.
01

Real apps, real uncertainty.

Live accounts, changing content, permissions, CAPTCHAs, and network conditions are part of the evaluation.

02

Multi-step everyday work.

Tasks include long-horizon execution, comparison and ranking, deep app entry points, pop-up recovery, and cross-app coordination.

03

Auditable outcomes.

We develop an agent system named AutoJudge to review real-device execution at the trajectory level. AutoJudge examines the task instruction and complete action–screenshot trace, assigns pass, failed, or environment error with a concise rationale, and returns an auditable final outcome.