Qwen-UI-Agent

Towards Next-Generation Real-World Centric Foundation GUI Agent

One agent that thinks, searches, and acts across mobile, desktop, and the web to complete real-world, long-horizon tasks.

MAI-UI Team ·Alibaba Token Hub

Performance comparisons on MobileWorld, MobileWorld-Real, and AndroidDailyPerformance comparisons on OSWorld-Verified, WebArena, and ScreenSpot-Pro

WHAT IT CAN DO

Built to complete real work across GUI interfaces.

MOBILE GUI USE
CAPABILITY 01

Mobile GUI Use

Optimized for everyday tasks in real-device mobile GUI use, the model navigates changing Android apps, accounts, content, and network conditions to search, compare, schedule, shop, and coordinate reliably.

01 / 06

PERFORMANCE

End-to-End GUI Tasks

MOBILE GUI USE
GUI-Only Success rate (%)

MobileWorld

Qwen-UI-AgentAlibaba
Seed 2.1 ProByteDance Seed
GPT-5.6 SolOpenAI
Claude Opus 4.8Anthropic
Qwen 3.7 PlusAlibaba
Gemini 3.1 ProGoogle
Qwen-UI-AgentClosed-source ModelOpen-weight Model
REAL-DEVICE MOBILE GUI USE
Success rate (%)

AndroidDaily

Qwen-UI-AgentAlibaba
Seed 2.1 ProByteDance Seed
Gemini 3.1 ProGoogle
Claude Opus 4.8Anthropic
GPT-5.6 SolOpenAI
Qwen 3.7 PlusAlibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model

GUI Grounding

GUI GROUNDING
Grounding score (%)

ScreenSpot-Pro · No zoom

Qwen-UI-AgentAlibaba
GUI-Owl-1.5Alibaba
UI-Venus-1.5Ant Group
Qwen 3.7 PlusAlibaba
MAI-UIAlibaba
Seed 2.1 ProByteDance Seed
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

ScreenSpot-Pro · Zoom-in

Qwen-UI-AgentAlibaba
Seed 2.1 ProByteDance Seed
GUI-Owl-1.5Alibaba
Qwen 3.7 PlusAlibaba
UI-Venus-1.5Ant Group
MAI-UIAlibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

SS-V2

Qwen-UI-AgentAlibaba
Seed 2.1 ProByteDance Seed
Qwen 3.7 PlusAlibaba
MAI-UIAlibaba
UI-Venus-1.5Ant Group
GUI-Owl-1.5Alibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

MM-GUI-L2

Qwen-UI-AgentAlibaba
MAI-UIAlibaba
Seed 2.1 ProByteDance Seed
Qwen 3.7 PlusAlibaba
UI-Venus-1.5Ant Group
GUI-Owl-1.5Alibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

OSW-G-R

Qwen-UI-AgentAlibaba
Qwen 3.7 PlusAlibaba
Seed 2.1 ProByteDance Seed
MAI-UIAlibaba
UI-Venus-1.5Ant Group
GUI-Owl-1.5Alibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

UI-Vision

Qwen-UI-AgentAlibaba
Qwen 3.7 PlusAlibaba
Seed 2.1 ProByteDance Seed
UI-Venus-1.5Ant Group
MAI-UIAlibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model

Author-reproduced result: the baseline was independently evaluated in the authors’ environment rather than copied from the model provider’s report.

Broader Capabilities

Real-world tasks demand more than interface interaction—they also require knowledge, multimodal reasoning, instruction following, and tool use. Qwen-UI-Agent gains strong GUI capabilities without becoming a narrow GUI-only model, preserving the base model’s general reasoning and agentic strengths for broader tasks.

01

Foundational reasoning

Multimodal understanding, knowledge, mathematics, and instruction following.

Swipe horizontally to compare all models →

BenchmarkQwen-UI-AgentQwen3.5-27BUI-Venus 30B-A3BGUI-Owl 32BOpenCUA-72B
MMMU-Pro72.473.532.439.531.0
RealWorldQA83.183.175.376.766.4
CharXiv-RQ77.776.844.750.939.6
MathVision82.882.036.850.626.6
AI2D_TEST91.191.984.384.878.9
MMLU-Pro86.586.065.673.958.8
IFEval (prompt-level strict)90.290.481.384.570.6
02

Broader agentic work

Tool use, terminal tasks, multi-turn service work, coding workflows, and deep-research retrieval on BrowseComp (BC) and BrowseComp-ZH (BC-ZH).

Swipe horizontally to compare all models →

BenchmarkQwen-UI-AgentQwen3.5-27BUI-Venus 30B-A3BGUI-Owl 32BOpenCUA-72B
Tau2-Bench89.989.222.76.114.4
Terminal-Bench 2.0 · Avg 550.141.13.20.09.0
Claw-Eval · Avg 373.566.930.629.626.4
Claw-Eval · Pass@351.841.25.55.50.5
BFCL-v474.271.319.832.728.3
SkillsBench · Avg 528.024.90.50.30.0
QwenClawBench · Avg 344.248.56.45.111.4
BrowseComp (BC)64.161.0
BrowseComp-ZH (BC-ZH)75.062.1
EVALUATION NOTE

All scores were independently reproduced in the authors' evaluation environment. Some harness, judge, simulator, runtime, or task-subset settings differ from official evaluations and are documented in the technical report.

DEMOS

Choose a demo domain05 domains
Available workflowsReal-device Mobile Use
Demo case 01 / 05

Mobile GUI Use

Recipe research + e-shopping

Task instruction (translated from Chinese)
I’m planning to make “passion-fruit sour-soup beef” tonight. Search Douyin for the most-saved photo-and-text post, save it, and remember the ingredients I need to prepare. Then, in the Hema app, purchase all the ingredients mentioned in the post—excluding seasonings—select delivery for 18:45 today, and place the order.

Complete everyday workflows across changing apps on physical mobile devices.

The agent first extracts a recipe and its ingredients from Douyin, then carries that information into Hema to complete a time-constrained grocery order.

CITATION

BIBTEX · CITATION
@article{zhou2026qwen_ui_agent,
  title={{Qwen-UI-Agent} Technical Report: Toward Next-Generation Real-World Centric Foundation {GUI} Agents},
  author={Zhou, Hanzhang and Tong, Panrong and Zhang, Xu and Kong, Quyu and Cai, Chenglin and Xia, Tianyu and Zhang, Gongjie and Zhang, Jianan and Li, Long and Chen, Long and Wang, Lei and Dai, Gaole and Li, Pengxiang and Chen, Liangyu and Wang, Yue and Hoi, Steven},
  journal={arXiv preprint arXiv:2607.28227},
  year={2026}
}