How to test an iOS app with an AI agent
Let a model drive your iOS app and never let it judge. Mobster decides pass or fail from accessibility tree assertions on one settled screen, with proof.
A model is good at the part of an end-to-end test that breaks most often: getting through the flow. “Complete onboarding, don’t allow notifications, reach the paywall” survives a renamed button and a new screen in the middle. A model is a poor judge of its own run. An agent that says “done” is reporting what it meant to do.
mobster verify gives each job to the right worker. A model may drive. The verdict comes from assertions you wrote before the run, checked in code against the app’s accessibility tree. In the words of the Checks reference: what the model says about its own run is never an input to the verdict.
That’s the difference from AI assertions that ask a model. Maestro’s assertWithAI, for one, uploads a screenshot to an LLM with your assertion and takes back true or false, and Maestro marks it experimental (its docs, read 8 October 2026). A Mobster verdict is a fact about the screen, with the frame that shows it.
An assertion is a fact about the screen
An assertion is a small object with one kind key. Mobster checks it against the tree WebDriverAgent reports for the app, with no model, OCR or pixel comparison.
| Kind | Form | Holds when |
|---|---|---|
| Text present | text: "Choose your plan" | A shown element’s label or value contains the text |
| Text absent | no_text: "Loading" | No shown element’s label or value contains it |
| Element shown | visible: {selector} | At least one shown element matches |
| Element absent | absent: {selector} | No shown element matches |
| Value | value: {selector} with equals: X | Exactly one shown element matches, and its value equals X |
| Count | count: {selector} with equals, at_least or at_most | The number of shown elements that match meets the bound |
A selector matches on id (the accessibility identifier), label, value, role, enabled and selected, and every field you give must match. label and value take an exact string or a /regex/.
A check file puts the flow and the assertions together. This one is written for the paywall of Daybreak, the sample app in the CLI’s repository:
version: 1name: A new user sees three plans on the paywallapp: bundle: dev.mobster.daybreak path: .mobster/build/Build/Products/Debug-iphonesimulator/Daybreak.appsteps: - Complete onboarding as a new user. Don't allow notifications. - Reach the paywall.expect: - text: Choose your plan - count: { id: /^plan_/ } equals: 3 - value: { id: plan_annual } equals: $39.99 / year - visible: { label: Restore Purchases, role: button } - no_text: Loadingbudget: { max_seconds: 180, max_usd: 0.25 }An unknown key at any level is an error, so a typo can’t quietly weaken a check. Text is normalized before it’s compared (curly quotes, dashes and invisible marks), so a label that says Don’t Allow matches Don't Allow. text and no_text never join two elements: “$39.99” in one and ”/ year” in the next don’t make “$39.99 / year”.
One settled screen decides
A flaky verdict usually comes from reading a screen mid-animation. Mobster decides on a settled screen, with every assertion on the same read:
- It reads the tree, waits 0.3 s and reads again. Two reads with the same fingerprint are settled.
- When a settled read satisfies every assertion, that read decides.
- Otherwise it reads again every 0.5 s until the timeout, 5 s by default (
--assert-timeout, 0 to 60). - At the timeout the last read decides, and the result says whether the screen was still changing.
So a no_text: Loading that holds on one read and a text: Plans that holds on another never combine into a pass. Content behind a full-screen cover or scrolled out of view doesn’t count as shown, and an agent scrolls to it with swipe and until, which stops as soon as an assertion holds.
Four verdicts, and two rules that keep a pass honest
The exit code is the verdict, so a script can branch on it:
| Verdict | Exit | Means |
|---|---|---|
passed | 0 | Every assertion held on one settled read, with the app in front |
failed | 1 | An assertion didn’t hold, or the app wasn’t running in front at the end |
needs_review | 2 | Nothing deterministic decided the run |
couldnt_run | 3 | The check is invalid, or this Mac couldn’t run it |
A crash fails the run as app_not_running, whatever the assertions say. Then two rules stop a pass that proves nothing:
- No assertions ends
needs_review. There is nothing to decide by. - Held before the flow ends
needs_reviewtoo. Right after launch, before any step, Mobster takes one read. If every assertion already holds on it, the check doesn’t show that the flow works. Without this rule, a paywall check that only asserts “Choose your plan” would pass with onboarding broken, whenever the app happens to open on the paywall.
Three ways to drive, one way to judge
| Mode | Who performs the steps | Model calls |
|---|---|---|
| Launch-only | Nobody. Mobster launches the app, opens a deep link if given, and checks | None |
| Key-less | Your coding agent, through mobster mcp | None by Mobster |
| Smart | Mobster, on your OpenAI or Anthropic key | One per agent turn, with the run’s spend capped by --max-usd (default $0.25, at most $1.00) |
Launch-only and key-less runs send nothing anywhere. In the key-less mode, your agent’s model drives and Mobster only judges: Give Claude Code a real iPhone and Test your iOS app with Codex walk through that loop. Checks run on simulators Mobster manages, and mobster verify --device checks an app you already installed on your USB iPhone without clearing its data.
The proof comes with the verdict
Every run writes a folder with result.json, report.html, the run’s check.yaml, its events and its frames. The report is one HTML file with no external requests. It shows each assertion with what was actually on screen, each step with its frame, and the verdict frame beside the accessibility overlay: every element Mobster read outlined, and each asserted element in green when it held or red when it failed.
A failure says what it found instead. Daybreak’s planted missing-plan bug turns ForEach(Plan.all) into ForEach(Plan.all.prefix(2)). The paywall check on that build, on 28 September 2026:
✗ failed The paywall shows three plans with Annual at $39.99 a year (18.2 s)
✓ text "Choose your plan"
✗ count id=/^plan_/ == 3: found 2: plan_weekly, plan_monthly
✗ value id=plan_annual == "$39.99 / year": not found; closest: "plan_monthly"
✓ visible label=Restore Purchases role=button
✓ no_text "Loading"When text or an element isn’t found, the line lists up to 5 near misses from the screen. That is what an agent needs to fix the right code, and what you need to trust the verdict without rerunning it.
How long it takes
Measured on 28 September 2026 on an M5 Max with macOS 26.5.1, Xcode 26.4 and the iOS 26.4 simulator, running Daybreak’s paywall check with five assertions:
| Run | Took | n |
|---|---|---|
| First run: create and boot the simulator, build and start WebDriverAgent, install, check | 77.7 s and 103.6 s | 2 |
| Warm run: simulator and WebDriverAgent up, the app reinstalled, five assertions | 6.8 s and 8.6 s | 2 |
A warm run spends about 1.5–2.5 s installing, 1.3–1.5 s launching and 1.4–1.7 s asserting. Simulators stay booted between runs, so the next run starts warm.
Make your app checkable
- Give the controls your checks depend on an identifier:
.accessibilityIdentifier("plan_annual")in SwiftUI. Anidselector then survives copy changes and translations. - Mark selection as a trait:
.accessibilityAddTraits(.isSelected). Theselectedselector reads it. - Write each expectation so it only holds after the steps.
Write your first check
Mobster CLI is free. Install it, build your app for the simulator, and run a launch-only check, which needs no key:
curl -fsSL https://mobster.dev/install.sh | sh
mobster sim doctor --fix
xcodebuild -scheme <Scheme> -destination 'generic/platform=iOS Simulator' \
-derivedDataPath .mobster/build CODE_SIGNING_ALLOWED=NO build
mobster verify --app .mobster/build/Build/Products/Debug-iphonesimulator/<App>.app \
--text "Welcome" --save first-checkGetting started with verify takes you from there to checks with steps.
Maestro vs XCUITest vs Mobster says when to use which, and Mobster vs Maestro and Mobster vs XCUITest compare them row by row.
Questions
Does Mobster use a model to decide whether a test passed?
No. A model may perform the steps, but the verdict comes from assertions checked against the app's accessibility tree on one settled read. What the model says about its own run is never an input to the verdict.
What does needs_review with held_before_flow mean?
Every assertion already held right after launch, before any step ran. Such a check doesn't show that the flow works, so add an expectation that only holds after the steps.
Do I need accessibility identifiers for Mobster checks?
No. Selectors also match labels, values, roles and enabled or selected state. Identifiers make a check survive copy changes and translations, so add them to the controls your checks depend on.
Can AI write UI tests for an iOS app?
Yes. A coding agent can write a Mobster check: the steps in plain English and the assertions that must hold, saved as a YAML file you commit. Mobster runs it and decides pass or fail from those assertions, so the agent writes the test but never grades it.
What are AI agents good for in iOS e2e testing?
Getting through the flow, the part of a mobile app test that breaks most often: a model can follow “complete onboarding, reach the paywall” past a renamed button or a new screen. Judging the result is the part to keep from it, so Mobster decides pass or fail from assertions on the accessibility tree.
Sources
We read each of these on 8 October 2026. Reviewed by Andy Guo on . Corrections are welcome at the GitHub repo.