Qwen-UI-Agent

Towards Next-Generation Real-World Centric Foundation GUI Agent

One agent that thinks, searches, and acts across mobile, desktop, and the web to complete real-world, long-horizon tasks. State-of-the-art performance on mobile. Frontier-level performance across desktop and web.

MAI-UI Team ·Alibaba Token Hub

Performance comparisons on MobileWorld, MobileWorld-Real, and AndroidDailyPerformance comparisons on OSWorld-Verified, WebArena, and ScreenSpot-Pro

WHAT IT CAN DO

Built to complete real work across GUI interfaces.

MOBILE GUI USE
CAPABILITY 01

Mobile GUI Use

Optimized for everyday tasks in real-device mobile GUI use, the model navigates changing Android apps, accounts, content, and network conditions to search, compare, schedule, shop, and coordinate reliably.

01 / 06

PERFORMANCE

End-to-End GUI Tasks

MOBILE GUI USE
GUI-Only Success rate (%)

MobileWorld

Qwen-UI-AgentAlibaba
Seed 2.1 ProByteDance Seed
GPT-5.6 SolOpenAI
Claude Opus 4.8Anthropic
Qwen 3.7 PlusAlibaba
Gemini 3.1 ProGoogle
Qwen-UI-AgentClosed-source ModelOpen-weight Model
REAL-DEVICE MOBILE GUI USE
Success rate (%)

AndroidDaily

Qwen-UI-AgentAlibaba
Seed 2.1 ProByteDance Seed
Gemini 3.1 ProGoogle
Claude Opus 4.8Anthropic
GPT-5.6 SolOpenAI
Qwen 3.7 PlusAlibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model

GUI Grounding

GUI GROUNDING
Grounding score (%)

ScreenSpot-Pro · No zoom

Qwen-UI-AgentAlibaba
GUI-Owl-1.5Alibaba
UI-Venus-1.5Ant Group
Qwen 3.7 PlusAlibaba
MAI-UIAlibaba
Seed 2.1 ProByteDance Seed
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

ScreenSpot-Pro · Zoom-in

Qwen-UI-AgentAlibaba
Seed 2.1 ProByteDance Seed
GUI-Owl-1.5Alibaba
Qwen 3.7 PlusAlibaba
UI-Venus-1.5Ant Group
MAI-UIAlibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

SS-V2

Qwen-UI-AgentAlibaba
Seed 2.1 ProByteDance Seed
Qwen 3.7 PlusAlibaba
MAI-UIAlibaba
UI-Venus-1.5Ant Group
GUI-Owl-1.5Alibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

MM-GUI-L2

Qwen-UI-AgentAlibaba
MAI-UIAlibaba
Seed 2.1 ProByteDance Seed
Qwen 3.7 PlusAlibaba
UI-Venus-1.5Ant Group
GUI-Owl-1.5Alibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

OSW-G-R

Qwen-UI-AgentAlibaba
Qwen 3.7 PlusAlibaba
Seed 2.1 ProByteDance Seed
MAI-UIAlibaba
UI-Venus-1.5Ant Group
GUI-Owl-1.5Alibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model
GUI GROUNDING
Grounding score (%)

UI-Vision

Qwen-UI-AgentAlibaba
Qwen 3.7 PlusAlibaba
Seed 2.1 ProByteDance Seed
UI-Venus-1.5Ant Group
MAI-UIAlibaba
Qwen-UI-AgentClosed-source ModelOpen-weight Model

Author-reproduced result: the baseline was independently evaluated in the authors’ environment rather than copied from the model provider’s report.

Broader Capabilities

Real-world tasks demand more than interface interaction—they also require knowledge, multimodal reasoning, instruction following, and tool use. Qwen-UI-Agent gains strong GUI capabilities without becoming a narrow GUI-only model, preserving the base model’s general reasoning and agentic strengths for broader tasks.

01

Foundational reasoning

Multimodal understanding, knowledge, mathematics, and instruction following.

Swipe horizontally to compare all models →

BenchmarkQwen-UI-AgentQwen3.5-27BUI-Venus 30B-A3BGUI-Owl 32BOpenCUA-72B
MMMU-Pro72.473.532.439.531.0
RealWorldQA83.183.175.376.766.4
CharXiv-RQ77.776.844.750.939.6
MathVision82.882.036.850.626.6
AI2D_TEST91.191.984.384.878.9
MMLU-Pro86.586.065.673.958.8
IFEval (prompt-level strict)90.290.481.384.570.6
02

Broader agentic work

Tool use, terminal tasks, multi-turn service work, coding workflows, and deep-research retrieval on BrowseComp (BC) and BrowseComp-ZH (BC-ZH).

Swipe horizontally to compare all models →

BenchmarkQwen-UI-AgentQwen3.5-27BUI-Venus 30B-A3BGUI-Owl 32BOpenCUA-72B
Tau2-Bench89.989.222.76.114.4
Terminal-Bench 2.0 · Avg 550.141.13.20.09.0
Claw-Eval · Avg 373.566.930.629.626.4
Claw-Eval · Pass@351.841.25.55.50.5
BFCL-v474.271.319.832.728.3
SkillsBench · Avg 528.024.90.50.30.0
QwenClawBench · Avg 344.248.56.45.111.4
BrowseComp (BC)64.161.0
BrowseComp-ZH (BC-ZH)75.062.1
EVALUATION NOTE

All scores were independently reproduced in the authors' evaluation environment. Some harness, judge, simulator, runtime, or task-subset settings differ from official evaluations and are documented in the technical report.

DEMOS

Choose a demo domain05 domains
Available workflowsReal-device Mobile Use
Demo case 01 / 05

Mobile GUI Use

Recipe research + e-shopping

Task instruction (translated from Chinese)
I’m planning to make “passion-fruit sour-soup beef” tonight. Search Douyin for the most-saved photo-and-text post, save it, and remember the ingredients I need to prepare. Then, in the Hema app, purchase all the ingredients mentioned in the post—excluding seasonings—select delivery for 18:45 today, and place the order.

Complete everyday workflows across changing apps on physical mobile devices.

The agent first extracts a recipe and its ingredients from Douyin, then carries that information into Hema to complete a time-constrained grocery order.

CITATION

BIBTEX · TEMPLATE
@misc{qwenuiagent2026,
  title  = {Qwen-UI-Agent Technical Report:
            Toward Next-Generation Real-World-Centric
            Foundation GUI Agents},
  author = {MAI-UI Team},
  year   = {2026},
  note   = {Alibaba Token Hub}
}