Guide

    Prompting your agent for native mobile UI: a designer’s actual workflow

    AI agents write competent Swift and produce competent, generic screens. The gap between that and an app people love is a craft gap, not a code gap — and it closes with a workflow, not a magic prompt. This is the one we use at Modaal, taught live in our Buildcamp and used to ship a real app to the App Store.

    A disclosure and a promise. The disclosure: this is Modaal's own method, written up from a live design session of our Buildcamp and from the two documentation guides where we've systematised it. The promise: everything below was used on a real app. Memory Lane, a private memory-box app, went from idea to the App Store with around fifty features, built the way this article describes — one person building function, a designer layering craft on top; the split the whole workflow is designed around.

    The core belief behind the workflow, and you should know it before deciding whether to adopt it: the future belongs to crafted applications — apps with a human's taste and emotional intent behind them, not raw AI-generated design. AI didn't lower the bar for shipping an app; it lowered the bar for shipping a generic one. Which makes the craft the differentiator, and the interesting question becomes: how does a designer — or a non-designer willing to learn — get craft out of an agent?

    Eight principles. The first one is the strangest, so let's start there.

    Four App Store screenshots of Memory Lane, the private social app built with Modaal
    The finished app on the App Store — Memory Lane, private social, built with the workflow below. View Memory Lane on the App Store

    1. Iterate your design in HTML — even though you’re building native Swift

    This is the counterintuitive core of the whole workflow. Your app is native SwiftUI. Your design iteration should not be.

    Here is the reasoning, straight from the session: it's much easier and faster to iterate on designs using HTML than doing it in native Swift code. An agent can regenerate an HTML mock of a screen in seconds, you can open it instantly, click through it, and ask for five variants of a card style without waiting on a build. Design iteration is a volume game — you want dozens of cheap rounds, not a handful of expensive ones — and HTML is where rounds are cheapest. Claude's design tooling generates and edits it natively.

    So the shape of the workflow is: explore and decide in HTML, then transfer the decisions — not the HTML — into your native app. The transfer mechanism is tokens (principle 6), which is what makes this safe: you are never asking an agent to "convert HTML to Swift" wholesale; you are asking it to apply a set of frozen decisions to native code it already knows how to write.

    One practical trick from the session that makes HTML exploration even better: when you have a design you like, download the whole HTML project archive and keep it. It becomes the reference artifact you hand to Modaal in the transfer step — an unambiguous "make it look like this," better than any verbal description.

    2. Build function first, then design as a layer

    Memory Lane was built ugly on purpose. Elena's first pass was deliberately minimal — black and white, no styling effort at all — because her goal was a functioning app: every flow working, the backend real, the product decisions made. In her words, the aim is to make it work first, with no emotion behind any screen. Only then did the design pass begin.

    Why this order works with agents: design decisions made before the product is real get remade anyway. Screens get cut, flows get merged, the PRD changes while you build — Elena descoped continuously during the functional pass. Every hour of styling invested in a screen that later dies is an hour lost, and with an agent, restyling a finished app is cheap in a way it never was before. The design layer lands once, on screens that have earned their existence.

    Honesty from the same session: this is a workflow, not a law. We've since experimented with a design-first build for another app, where a short brief goes to design before any code exists and the design leads. If you're a designer, leading with design may suit you. But if you're a non-designer building with an agent, function-first has a second benefit — by the time you reach the design pass, you know your product well enough to have taste about it.

    Memory Lane running on an iPhone simulator inside Modaal, next to a native widgets feature plan
    Memory Lane inside Modaal during the functional pass — feature plans on the right, the live app on the left.

    3. Feed references — and say exactly what to take from each

    Agents are excellent imitators and terrible mind-readers. The single highest-leverage prompting habit in this entire article is how you hand over reference material.

    Collect a small set of screens whose mood matches your app — we browse Mobbin for real app screens and 60fps.design when the question is motion. A handful, not a wall: the session's hard-won rule is several screens, not lots of screens. Then — this is the part almost everyone skips — tell the agent precisely what to take from each one: "from this screen, the shadows and the rounded corners; from this one, the typography and its hierarchy." An unannotated reference is an invitation to grab everything, including the parts you didn't want.

    Our docs add two refinements. Bring one or two anti-references too — "not this: too neon, too dense" — because what you don't want is information the agent can't infer. And before any generation, use the describe-back prompt: "First describe each reference back — what you see and what you'll take from it. Don't generate anything yet." If the agent's read of your references is wrong, you want to find out before it builds forty screens on the misunderstanding.

    If you can't name what you like about a reference, ask the agent to disassemble it for you — have it describe the corners, the type, the palette. That's how a non-designer borrows a designer's vocabulary, one screen at a time.

    4. Fundamentals before layout — and one north-star screen before everything

    The most common failure mode when prompting for design is asking for everything at once: "redesign my app to look like these screenshots." The session's method is the opposite — a strict order, one decision class at a time.

    Fundamentals first: colors, then corner radii, then typography, then spacing. Apply one, check it everywhere, then the next. The Memory Lane transfer followed exactly this ladder — colors applied and verified before radii were even mentioned. Each step is small enough to review properly, and when something looks wrong you know exactly which decision caused it.

    Then one screen — the north-star. Every app has one screen that carries its DNA; in Memory Lane it was the lane itself, the timeline of memories — get that screen right and the rest of the app follows from its decisions. Iterate the north-star hard (two to four rounds is normal), then apply the resulting style to a second, structurally different screen. If the style survives contact with a different layout, it's a system; the third screen is nearly free, because by then all the real decisions are made.

    And iterate with reactions, not specifications. You don't need to know that a color is oversaturated — you need to say "warmer," "cooler," "airier," "denser," "quieter," "louder," "softer," "sharper." Our docs formalise this as the four dials, and there's a diagnostic prompt to match: ask the agent to rate the current screen on each dial, then turn the ones that feel wrong. Reaction vocabulary is how people without design training steer design — and it works because the agent translates the dial into the twelve specific changes you couldn't have named.

    5. Speak in roles and systems, not in colors and pixels

    A lesson from the same session, courtesy of a real app review. A co-parenting calendar had one color doing three jobs — marking one parent's days, one child's events, and serving as the app's primary accent. Every use was individually reasonable; together they made the calendar unreadable, because color was the only signal and it meant three things.

    The fix is a habit: give every color a role, and address it by role. "Accent," "surface," "text-primary," "parent-A," "child-marker" — never "the green one." Prompt in roles and the agent maintains the system; prompt in colors and it decorates. Two corollaries from the review: when two roles compete for attention, deliberately mute one — your days should outshine your co-parent's, not fight them; and never let color carry meaning alone, because for colorblind users a red/green distinction may simply not exist. Add a name, a stripe, a shape.

    The same role-thinking applies across the board, and our design-prompts guide is essentially a phrasebook for it: iOS type sits on a fixed scale (34pt large titles down to 11pt captions) — ask for "one step up," not "a bit bigger." Weights progress regular → medium → semibold → bold. Tap targets are 44×44pt minimum. Text contrast is a checkable WCAG AA standard, not a vibe, and non-text elements need 3:1. Every one of those is a prompt the agent can verify mechanically — which is exactly the kind of instruction agents execute best.

    6. Freeze the decisions into tokens — then transfer to native, one category at a time

    This is where the HTML exploration becomes a native app without chaos, and it's the most systematised part of the workflow.

    When the north-star screens are right, freeze: ask the agent to populate a design-system file with every decision made so far. Ours produces three artifacts — design-system.md (the intent and rules, in prose), DesignTokens.swift (every visual value, as the single source of truth), and the north-star screenshots. One trick from the docs that prevents a whole failure class: keep those files out of the project until the freeze, and mark unfilled values bright magenta — so if the agent ever styles from an empty template, you see it instantly instead of shipping it.

    Then the transfer into Modaal, and the ladder from principle 4 applies again: Plan mode on, explain the situation — "I've redesigned in HTML and need this brought into our native app" — then apply tokens by category: colors first, build, check; then radii; then typography; then spacing. Only when the fundamentals are verified do you move to layout, screen by screen: hand Modaal the screenshot and the HTML file and ask it to rebuild the layout using the tokens already in the project. Memory Lane's transfer ran exactly this sequence, and the step-by-step discipline is what made it uneventful.

    Two closing habits keep the system honest. Periodically run a consistency health-check — "check all screens for consistency of spacing, corner radii, shadows, typography" — because agents drift, and a ten-second audit prompt catches it. And when writing the handoff from your design tool to Modaal, have the agent write that prompt itself — it knows what it decided better than your paraphrase of it. From then on, new screens are born in Modaal directly, inheriting tokens and components automatically; the HTML phase was scaffolding, and it comes down.

    7. Animations last — and the slider trick for people without motion vocabulary

    Micro-animations are fine-tuning, not foundation. The session was unambiguous: set the fundamentals, land the design, and only then walk the app asking "where would a small, pleasant motion add feeling?" In Memory Lane those moments are tiny — a tick, a little hand animation on the sharing screen, a moment in onboarding — and they cost little while changing how the app feels disproportionately.

    The problem with prompting for motion is vocabulary. You know you want it "bouncier," but bounciness is a number you've never heard of. The solution is the best single prompting trick in this article: ask the agent to put adjustable sliders on the animation's attributes, right in the HTML preview. Speed, damping, bounce — as draggable controls next to the running animation. You drag until it feels right, read the numbers off the sliders, and hand exactly those values to Modaal. You steered motion design without ever learning its language — the notation came to you.

    For illustrated animations, the honest note from the session: hand-crafted motion is still the hardest thing here for non-designers. The working routes: Lottie's free animation library (you can open the JSON and re-color it to your palette), or generating frames and asking the agent to sequence them — Memory Lane's onboarding animation was built that way, from a handful of generated shots. Describe motion by feel — spring or ease, snappy or gentle — and remember the system setting that outranks all of it: respect Reduce Motion.

    8. Design for delight — the layer above usability

    The last principle is the reason the others matter. Elena teaches it from the product-delight framework of Nesrine Changuel, who literally wrote the book on it: products climb a pyramid — functional, reliable, usable, pleasurable — and delight lives at the top, where joy and surprise arrive together. Her three pillars: remove friction, anticipate needs, exceed expectations. A story she tells makes the third concrete: a ride-share app noticing an unplanned five-minute stop and checking in — are you okay? — with an emergency button ready. Nobody expects that; that's the point.

    For an AI-built app this layer is not decoration — it's the moat. Function is now cheap for everyone, which means everyone's floor is the same and the ceiling is emotional. The honest self-assessment from the session: Memory Lane so far removes friction and anticipates needs; exceeding expectations is still on the roadmap. That's the right way to use the framework — as an audit of where your app stops on the pyramid, and a prompt for what its next layer of craft should be.

    Which returns to where this article started. An agent will happily generate you a functional, reliable, even usable app. The last two layers — the crafted feel, the moment that surprises — come from a human with taste directing the agent deliberately. That's what all eight principles are for.

    Frequently asked questions

    Not with one prompt — with a sequence. Collect a few reference screens and tell the agent exactly what to take from each ("the shadows from this one, the type hierarchy from that one"). Have it describe the references back before generating anything. Then decide fundamentals one at a time — colors, corner radii, typography, spacing — before touching layout, and perfect one north-star screen before styling any other. The order is the prompt: agents asked for everything at once produce generic screens; agents steered through small decisions produce yours.

    Speed of rounds. Design iteration is a volume game, and an agent regenerates an HTML mock in seconds — five variants of a card style cost nothing, while the same exploration in native code means builds between every look. The workflow that follows: explore and decide in HTML, freeze the decisions into design tokens, then transfer the tokens — not the HTML — into the native app category by category. You get HTML’s iteration speed and end with real SwiftUI built on a design system.

    Use reactions, not specifications. "Warmer or cooler, airier or denser, quieter or louder, softer or sharper" — the four dials — steer an agent effectively because it translates the feeling into the specific changes you couldn’t name. For motion, the trick is even better: ask the agent to put draggable sliders on the animation’s attributes in an HTML preview, drag until it feels right, and read off the numbers. And ask the agent to disassemble screens you admire — it will teach you the vocabulary one reference at a time.

    A token file is the single source of truth for every visual value in your app — each color by role, each radius, each type size. It matters doubly with agents: it is how design decisions survive the transfer from an HTML exploration into native code, and it is what keeps fifty screens consistent when the thing writing them does not remember last week. Prompt against tokens ("use our accent role"), run periodic consistency audits, and have the agent scan for hardcoded values that snuck past the system.

    Yes — with a workflow, not a magic prompt, and with honest limits. References plus the fundamentals-first ladder gets a non-designer surprisingly far; the four dials and role-based vocabulary cover the steering; the slider trick covers motion. What remains genuinely hard without a designer is original illustration and signature animation — the session’s honest note — where the working routes are Lottie’s free library re-colored to your palette, or generated frames sequenced by the agent. And the delight layer at the very top benefits from human taste more than any other.

    Last. Micro-animations are fine-tuning: set the fundamentals, land the design on real screens, then walk the app asking where a small motion would add feeling — a tick, a transition, an onboarding moment. Added early they get rebuilt with every layout change; added last they take little time and disproportionately change how the app feels. Describe them by feel (spring or ease, snappy or gentle), tune them with sliders in an HTML preview, and always respect the user’s Reduce Motion setting.

    Start free. Ship native.

    One project, unlimited prompts. No card.

    Keep reading