Guide · agent testing on iPhone, iPad and Duo
How AI Agents Test iPhone Apps in the Simulator (and Why They Miss Landscape, iPad and Duo)
An agent can build and run your iPhone app in the simulator, but the default tools cannot rotate, fold or pose it. What it skips, and how to give it control.
An AI agent that builds your iPhone app will tell you the feature is done. What did it check before saying that? Usually: the code compiles, and the app launches in the simulator in the state it was already in. That is the whole test.
This gap can let agent-built apps ship bugs that a human tester would catch in a minute.
What the default tools let an agent do
Apple's developer tools, the same ones Xcode and a command-line agent use, let an agent compile the app, install it on a simulator, launch it, take a screenshot and read logs. Claude Code's desktop app adds a simulator pane that can drive a simulated device: tap, type, look at the screen. Its documented limits (simulated devices only, local sessions only, and not yet with Xcode 27) are on using an AI coding agent with Xcode. All of that covers "does it run". It does not cover "does it work in the states a user will put it in".
What the default tooling does not give the agent is control over the device's state: rotate the simulator to landscape, fold or unfold an iPhone Duo, set the hinge angle, put the device face up or face down, and step through those states to watch how the layout responds. There is no control for it, so the agent does not do it, and it does not know it skipped anything.
Why this rarely showed up before
Most iPhone apps are portrait-only. For years the one state the agent tested was the only state that mattered, so nobody noticed the tests were one-state tests.
Three things changed that. iPad apps have always needed both orientations and a resizable window, and agents building "universal" apps were shipping layouts nobody had seen in landscape. The iPhone Duo added a folded cover screen, an open screen, and a hinge angle the app can respond to (what Duo changes for an existing app lists the seven updates). And once agents started building whole apps rather than single functions, the untested states became the user's problem instead of the developer's.
On Android the picture is different: the emulator tooling has supported orientation and fold control for longer, so an agent building the Android side of the same app can already test those states. The gap is on the Apple side.
What "properly testing a Duo feature" means
Take one feature: a screen that shows a list on the cover display and expands to a two-column layout when the phone is opened. To test it the way a tester would, the agent has to look at the cover screen with the phone closed, open the phone and check the two columns, rotate to landscape in both states, and, if the design reacts to the hinge angle, step the hinge through a few angles and check that nothing jumps or clips.
With the default tools the agent can do none of that on its own. It builds the feature, runs it in whatever state the simulator was left in, and reports success. You open the phone and Save and Cancel are off the screen.
Giving the agent the controls
Modaal exposes simulator controls to the agent as tools: start and stop the simulator, rotate it, fold and unfold, set the hinge angle in small steps, set the pose. The agent uses them the same way it uses the compiler: it builds the feature, then puts the simulator through each state and reads the screenshot before it reports done.
That is what the Xcode-versus-Modaal Duo comparison video shows. Same model, same prompt. On one side the agent has the default tools and stops at the state it started in; on the other it folds, rotates, checks each state and fixes what it finds. The difference is not the model. It is whether the agent can reach the states.
If you are in Claude Code alone
You are the tester. After every feature that touches layout, rotate the simulator yourself, and if you target iPad or Duo, check every state by hand. Write that into the project rules so the agent at least reminds you. Or hand the agent the controls and let it do what it would do anyway if it could.
Frequently asked questions
It can build the app, install it on a simulator, launch it, take screenshots and read logs, and the desktop app's simulator pane can tap and type on a simulated device. What the default tooling does not give it is control over the device's state: rotation, folding, hinge angle and pose.
Because it tested the feature in the one state the simulator was already in. With no control to rotate the device, it never saw the landscape layout, and it does not know it skipped a state.
Looking at the cover screen with the phone closed, opening it and checking the open layout, rotating to landscape in both states, and, if the design reacts to the hinge angle, stepping the hinge through several angles to check nothing jumps or clips.
Less so. Android emulator tooling has supported orientation and fold control for longer, so an agent building the Android side of the same app can already test those states. The gap is on the Apple side.
Modaal exposes simulator controls to the agent as tools: start and stop the simulator, rotate it, fold and unfold, set the hinge angle in small steps, set the pose. The agent puts the simulator through each state and reads the screenshot before it reports the feature done.
Be the tester. After every feature that touches layout, rotate the simulator yourself; if you target iPad or Duo, check every state by hand, and write that into the project rules so the agent reminds you.
Let the agent test every state before it says done.
Same model, same prompt. In Modaal the agent folds, rotates and checks each state itself. Free plan for one project.
Keep reading
- iPhone Duo apps: what changes
The seven updates an existing app needs.
- Using an AI coding agent with Xcode
What the simulator pane can and cannot do, from Anthropic's docs.
- Building iOS apps with Claude Code
The six things that break, and what people do about them.
- Why your vibe-coded app breaks on every new feature
The architecture and test loop behind autonomous testing.