aiflutterarchitecturewebdev

The Hardest Feature in an AI App Builder Is Saying 'Done'

Generating the code is the easy part. Building Code Horizon — an AI platform that takes a prompt to a live URL — the hard problem was knowing when a build is actually finished. A 13-state status machine, why we chose polling over streaming, and regenerating only what changed.

Ayan ParvaizAyan Parvaiz9 min read
Code Horizon — an AI software platform that turns a prompt into a working, deployed app, built in Flutter web by Ayan Parvaiz

Ask a model to "build me a restaurant finder app" and code comes out. In 2026 that is not impressive any more — it is a commodity, and a dozen products do it.

The interesting problems all start after the code exists. Did it compile? Did the compiled thing get hosted anywhere? Does the URL actually open? What will the infrastructure cost per month? And the one that turned out to be hardest:

When exactly are you allowed to tell the user it is ready?

I have spent the last year building Code Horizon — a Flutter web platform that takes a one-line idea through a seven-step pipeline and out the other side as a deployed, browsable application. Almost every architectural decision in it traces back to that single question.

The lie in the green checkmark

Here is the failure mode that shaped the whole system.

The generator finishes. The code compiles. Compilation returned success, so the obvious thing to do — the thing most builders do — is flip the status to done, show a green tick and render the preview link.

Then the user clicks the link and gets a 404.

Because compiling and being reachable are two different events with a gap between them. The build artifact still has to be registered with hosting, and that registration can fail on its own: a name collision, a rate limit, a transient error on the hosting API. The code is perfect. The URL is dead.

That gap has a name in our status enum, and it is the single most important state in the system:

build_complete

It looks like success. Every instinct says treat it as success. It is roughly the halfway point.

`build_complete` does not mean done. It means the code compiled and nothing is serving it yet.

Once you accept that, the design follows: a build is a transaction, not a boolean. It is not finished until every stage that the user's experience depends on has finished.

Thirteen states and one property

The status enum has thirteen members. What makes it work is not the number, it is that every member carries a single property: isTerminal.

queued
   ↓
building ─────────────► build_failed        (terminal)
   ↓
build_complete   ← the trap: looks done, is not
   ↓
registering ──────────► registration_failed (terminal after retries)
   ↓                            │
completed  (terminal)           └── auto-retry, max 3
   ↑
   └── only here is the preview URL rendered

isTerminal answers exactly one question — may the client stop polling? — and by encoding it on the enum rather than in a chain of if statements scattered across controllers, there is exactly one place to be wrong.

Two rules fall out of it:

The preview URL is never rendered before `completed`. Not at build_complete, not at registering. The user cannot click a link that does not resolve, because the link does not exist yet.

`registration_failed` is not final on first sight. Hosting registration fails for boring, transient reasons, so it auto-retries up to three times before the user is told anything. Most failures never reach a human. The ones that do are real.

That second rule is worth dwelling on. The naive version surfaces every error immediately, which feels honest and is actually worse — you have taught the user that your product fails often, when what actually happened is that an API returned a 503 and would have worked 800ms later.

Why polling, and not streaming

There is no SSE stream and no WebSocket in this system. Build status is plain HTTP polling.

I want to defend that, because "real-time streaming build logs" is the more impressive-sounding sentence and I could have written it.

A stream is a stateful connection, and stateful connections have to be babysat. Reconnect on network change. Reconcile on browser refresh. Handle the tab going to background where timers get throttled. Deal with the user having the same project open in two tabs. Every one of those is a bug I would have had to find, in a system where the underlying job already takes tens of seconds.

Polling is a request that either returns or does not. Refresh the browser and it just resumes. Close the tab and nothing leaks. When something is wrong you can read it in the network tab.

We run two loops at different rates, and the difference between them is a product decision, not a technical one:

WhereIntervalWhy
Workspace10 secondsThe user is on the screen watching progress
Background20 secondsThe user pressed "notify me" and walked away

The second loop lives in a global controller that survives navigation and keeps its subscriptions in browser storage. That is the part users actually notice: start a build, go back to the dashboard, browse other projects — and the completion notification still finds you.

The lesson generalises. Reach for streaming when the value is in the individual token arriving. Token-by-token AI responses, yes. A job that takes forty seconds and has thirteen possible states, no. A ten-second poll is indistinguishable from real-time to a human waiting on a build, and it costs a fraction of the complexity.

Seven steps, one screen

The pipeline is seven stages, and the user can chat with the AI at every one of them:

#StepWhat it produces
1Core IdeaA raw sentence becomes a structured product description
2MVPA feature list, each item can be kept or removed
3User FlowScreen-by-screen navigation map
4PreviewA real build, running in a device frame
5CodeBrowsable file tree of the generated source
6ArchitectureStack, integrations, and a monthly cost estimate
7PublishLive subdomain, custom domain, GitHub, share link

It is deliberately not a multi-page wizard. All seven live as tab state on a single route, and each tab holds its own chat thread. Switch back to step two and you are back in that step's conversation, not a fresh one.

The reason is that refinement is not linear. Real users reach the preview, realise the feature list was wrong, and want to go back — without losing the reasoning that got them there. A wizard punishes that. Tabs with independent conversations invite it.

There is one small detail at the very front I am unreasonably fond of: a gibberish filter before submission. Under five characters, rejected. The same character four times in a row, rejected. An implausible vowel-to-consonant ratio, rejected.

It is fifteen lines of code and it exists for an economic reason, not a UX one. Every build spends real tokens. Letting a keyboard mash start a pipeline burns money and gives the user a garbage result they will judge you for. The cheapest place to stop a bad build is before it starts.

Do not regenerate what did not change

Early on, editing anything upstream re-ran the whole pipeline. Change one feature in the MVP list and core idea, user flow and build all regenerated. It was slow and it was expensive, and both of those costs landed on the user.

Now a build returns flags instead: which upstream sections its output has invalidated. Those tabs get an "Update available" badge, and tapping it regenerates that section only. Editing a section directly through chat clears its flag too.

StepCan go stale?
Core IdeaYes
MVPYes
User FlowYes
Preview, Code, Architecture, PublishNo — these are outputs, not inputs

The framing that made this click: the first three steps are inputs, the rest are outputs. Inputs can be invalidated by things that happen downstream. Outputs are just regenerated. Once you draw that line, which sections need a stale flag stops being a judgement call.

Making token cost legible

Every AI action has a real cost behind it, so usage is metered in credits. The balance is split into three types, and the split is the business model in miniature:

TypeSourceJob
dripFree, dailyA reason to come back tomorrow
bonusIncluded with a planWhat the paid tier is actually worth
purchasedOne-off packsA pressure valve for heavy users

Two decisions around that matter more than the split itself.

Estimate before spending. Expensive actions ask the backend what they will cost before running. Above a threshold, the user gets a confirmation dialog. Nobody discovers the price after the fact.

Errors carry actions. "Insufficient credits" on its own is a dead end. Ours returns what to do about it: upgrade, buy a pack, or wait until tomorrow — because tomorrow the drip refills. That third option is not a consolation prize. It is the reason a free user is still around next week.

The balance itself lives in one global place, refreshes on its own every couple of minutes, and is always visible in the nav bar. Money in a product should never require a page load to check.

Your code, your GitHub

Generated code you can only look at is a hostage situation. The push-to-GitHub flow is what makes "no lock-in" a real claim rather than a marketing line: OAuth in a popup, the popup signals back to the app when authorisation completes, the user names a repository and picks private or public, and the code lands in their account.

The connection is tracked per project, not globally, so one user can put different projects in different repositories. And it can be disconnected at any time.

Versions are a fear feature

Every significant build creates a new version, and versions belong to a family under a root id. Old previews stay openable, changelogs are readable, and any earlier state can be restored.

This is not really a version control feature. It is an answer to a specific anxiety: what if I say the wrong thing in chat and ruin what I already have?

You cannot. The previous version is still there. Users explore far more aggressively once they know that, and aggressive exploration is what makes the product feel good.

What I would tell another team building this

  • Model the build as a transaction, not a boolean. The gap between "compiled" and "reachable" is where users lose trust.
  • Put `isTerminal` on the enum. One place to decide whether the work is over.
  • Retry the transient class of failures silently. Surfacing a 503 that would have cleared in a second teaches users your product is flaky.
  • Do not stream what you can poll. Streaming is right when the token itself is the value; a long job with a state machine is not that.
  • Separate inputs from outputs. It turns "should this be regenerated?" from an argument into a lookup.
  • Never render a link you have not confirmed resolves. A dead link costs more trust than a slow one.

The generation was never the hard part. Any team can get a model to emit code that looks right.

The hard part is the honesty layer on top — refusing to show a green tick until the user's app genuinely opens. That is one product decision, and thirteen enum values, two polling loops, a retry policy and a background notifier all exist to serve it.


I build production Flutter apps and AI-powered products. If you are shipping something like this, my portfolio is at [ayan-parvaiz.web.app](https://ayan-parvaiz.web.app).

Keep reading