flutteraiopenaidartmobile

Streaming AI Responses in Flutter: SSE, Cancellation and Multi-Model Fallback

How Jarvis — a Flutter AI assistant shipping for 3.5 years in 19 languages — streams OpenAI and DeepSeek responses token by token. SSE parsing in Dart, the UTF-8 bug that corrupts non-English text, cancellation, and falling back between models.

Ayan ParvaizAyan Parvaiz8 min read
Jarvis on iPhone — the assistant picker, a streaming AI chat and image generation, a Flutter AI app shipping for three and a half years

Every Flutter + ChatGPT tutorial ends at the same place: await http.post, decode the JSON, setState. It works, you get an answer, the demo looks fine.

Then you ship it and users tell you the app is broken, because they stared at a spinner for eleven seconds.

I have been running Jarvis — a Flutter AI assistant with multi-personality chat and image generation — on the App Store and Play Store since January 2023. It is on version 1.0.51 now, localized into 19 languages including right-to-left scripts. Three and a half years of continuous releases teaches you things that a weekend demo cannot.

The single biggest one is that streaming is not a nice-to-have. It is the difference between an app that feels dead and an app that feels alive. Here is how it actually works in Dart, including the bug that silently corrupts text in every language that is not English.

Why streaming changes the app, not just the UI

A full response from a large model takes somewhere between three and fifteen seconds. That is the number you cannot change.

What you can change is time to first token — usually under a second. Streaming lets you show that first word immediately and keep painting the rest as it arrives.

The user's perception is not "this took nine seconds." It is "it started answering right away." Same total time, completely different app.

There is a second benefit nobody mentions: when a user can see the answer going wrong at word four, they hit stop. That is a cancelled request instead of a full completion, and on a per-token bill that adds up.

The transport: server-sent events

OpenAI-compatible APIs stream over SSE (server-sent events), not WebSockets. This confuses people coming from real-time chat, so it is worth being clear about why.

SSE is one-way and rides on a normal HTTP response. You send one request, the server holds the connection open and writes chunks. No handshake, no separate socket lifecycle, no reconnect logic. For "one question, one long answer" it is exactly the right shape.

The wire format is plain text:

data: {"choices":[{"delta":{"content":"Hello"}}]}

data: {"choices":[{"delta":{"content":" there"}}]}

data: [DONE]

Blank-line separated, each event prefixed with data: , terminated by a literal [DONE].

Reading the stream in Dart

http.post will not help you — it waits for the whole body. You need Client.send() with a StreamedResponse:

dart
final client = http.Client();

final request = http.Request('POST', Uri.parse('$baseUrl/chat/completions'))
  ..headers.addAll({
    'Authorization': 'Bearer $apiKey',
    'Content-Type': 'application/json',
    'Accept': 'text/event-stream',
  })
  ..body = jsonEncode({
    'model': model,
    'messages': messages,
    'stream': true,
  });

final response = await client.send(request);

Now the important part.

The bug that corrupts every non-English language

Here is the version almost everyone writes first:

dart
// Broken. Do not ship this.
await for (final chunk in response.stream) {
  final text = utf8.decode(chunk);   // <-- the bug
  // ...parse text
}

It works perfectly in English and destroys text in most other languages.

response.stream emits arbitrary byte chunks. The network does not care about character boundaries. A three-byte character — Bengali, Arabic, Chinese, an emoji — can be split across two chunks: two bytes at the end of one, one byte at the start of the next.

utf8.decode on an incomplete sequence either throws or gives you a replacement character. Your user sees আম become আ and then a black diamond.

We shipped this bug. It was invisible in English testing and reported from Bangladesh and Saudi Arabia within a week.

The fix is to decode the stream, not each chunk. A single utf8.decoder transform carries the partial bytes across chunk boundaries for you:

dart
final lines = response.stream
    .transform(utf8.decoder)      // stateful across chunks — this is the fix
    .transform(const LineSplitter());

One transform on the stream. Never utf8.decode inside the loop.

Buffering: chunks are not events

The second trap is assuming one network chunk equals one SSE event. It does not. You can receive half an event, or three events in one chunk.

LineSplitter handles most of this, but you still need to skip blanks, strip the prefix and watch for the terminator:

dart
final buffer = StringBuffer();

await for (final line in lines) {
  if (line.isEmpty) continue;
  if (!line.startsWith('data: ')) continue;

  final payload = line.substring(6);
  if (payload == '[DONE]') break;

  try {
    final json = jsonDecode(payload);
    final delta = json['choices'][0]['delta']['content'] as String?;
    if (delta != null) {
      buffer.write(delta);
      yield buffer.toString();
    }
  } on FormatException {
    // A truncated payload. Skip it rather than killing the stream.
    continue;
  }
}

That try/catch is not defensive padding. Under a bad mobile connection you will get a partial JSON payload, and an uncaught FormatException there ends the answer mid-sentence with no error the user understands.

Cancellation, or how to leak money

The user taps stop. The user navigates back. The user backgrounds the app mid-answer.

If you do not cancel, the request keeps running on the server and you keep paying for tokens nobody will ever read. Worse, the callback fires into a disposed widget and you get a setState() called after dispose() crash.

Two things have to happen, and both matter:

dart
StreamSubscription<String>? _sub;
http.Client? _client;

void _stop() {
  _sub?.cancel();       // stop listening
  _client?.close();     // actually tear down the connection
  _sub = null;
  _client = null;
}

@override
void dispose() {
  _stop();
  super.dispose();
}

Cancelling the subscription alone is not enough. Without client.close() the socket stays open and the server keeps generating.

Then a rule that took us an embarrassingly long time to adopt: persist the partial answer. If a user stops at word forty, those forty words cost real money and are often still useful. Save what you have to the conversation instead of discarding it.

Multi-model fallback

Jarvis routes across more than one provider. The reason is not benchmark scores, it is availability — when one provider has a bad hour, an app with a single hard-coded model is simply down.

The rule that matters is a first-token deadline, separate from the overall timeout:

dart
Stream<String> completion(List<Message> messages) async* {
  for (final model in _fallbackChain) {
    try {
      yield* _stream(model, messages).timeout(
        const Duration(seconds: 8),   // time to FIRST token
      );
      return;                          // succeeded — stop trying
    } on TimeoutException {
      continue;                        // next model in the chain
    } on HttpException {
      continue;
    }
  }
  throw const AiUnavailable();
}

Eight seconds of silence means something is wrong upstream. Fall through to the next model rather than making the user wait out a 60-second socket timeout.

One caveat worth knowing before you build this: once tokens have started arriving you cannot silently switch models. The user has already read the first half. Retrying from scratch makes text rewrite itself on screen, which reads as a bug. Fall back before the first token, surface an error after it.

Do not rebuild on every token

This one is a Flutter-specific performance trap.

A fast model emits tokens quicker than the screen refreshes. If every token calls setState on a page holding a full conversation, you are rebuilding the entire list sixty-plus times a second, and any Markdown rendering re-parses the whole string every time.

Two fixes, in order of impact:

Isolate the rebuild. The streaming message should be its own widget listening to its own ValueNotifier. Nothing above it in the tree needs to rebuild while tokens arrive.

Coalesce the tokens. Buffer incoming deltas and flush on a timer rather than on arrival:

dart
Timer? _flush;
final _pending = StringBuffer();

void _onToken(String delta) {
  _pending.write(delta);
  _flush ??= Timer(const Duration(milliseconds: 50), () {
    _text.value += _pending.toString();
    _pending.clear();
    _flush = null;
  });
}

Fifty milliseconds is invisible to a reader and cuts rebuilds by an order of magnitude on the low-end Android devices where most of the world runs your app.

And leave the typing indicator up until the stream closes, not until the last token lands. Otherwise it flickers off during every natural pause in generation.

Conversation memory without a vector database

Jarvis gives every persona its own memory, and people assume that means embeddings and a vector store.

For a chat assistant it usually does not. Context windows are large enough now that the practical strategy is much simpler:

  • Keep the persona's system prompt pinned at the top, always.
  • Keep the last N turns verbatim.
  • When the conversation outgrows the budget, summarise the oldest turns into a short running summary and drop the raw text.

You are trading a little fidelity for a bounded, predictable token cost per message — which is the number that decides whether a subscription app makes money. Reach for a vector database when users need retrieval across documents, not to remember what they said four messages ago.

The checklist I run on any streaming AI feature

  • utf8.decoder applied to the stream, never utf8.decode per chunk
  • Tested with Bengali, Arabic and emoji, not only English
  • Malformed SSE payloads skipped, not fatal
  • client.close() on cancel, not just subscription.cancel()
  • Streams cancelled in dispose()
  • Partial answers persisted when the user stops
  • First-token deadline separate from the request timeout
  • No model switching after the first token has rendered
  • Streaming message isolated so the conversation list does not rebuild
  • Tokens coalesced on a timer, not applied one by one
  • Typing indicator tied to stream close, not last token

What three and a half years actually taught me

The streaming code was the interesting part. It was not the hard part.

The hard part was that the platform moves underneath you. In that time we migrated the entire push stack from FCM v1 to v2 without breaking existing subscribers, absorbed breaking OS changes every six months, and shipped the whole thing into 19 languages where right-to-left layout turned out to be a genuinely different UI, not a mirrored one.

An AI wrapper is a weekend. An AI product is the three years after the weekend — the migrations, the subscription edge cases, the bug reported only by users whose alphabet you do not read.

Build the streaming properly and you have a good demo. Build the rest properly and you have something that is still shipping in three years.


I build production Flutter apps and AI-powered products. If you are shipping something like this, my portfolio is at [ayan-parvaiz.web.app](https://ayan-parvaiz.web.app).

Keep reading