Skip to content

Zuni, a real-time voice AI assistant

A personal voice AI assistant I built end to end: a native iPhone app with natural, interruptible conversation and a backend that lets the AI act on real accounts.

Client
My own product
Role
Sole engineer, iOS and backend
Year
2026

In brief#

Zuni is a personal voice AI assistant I designed and built end to end: a native iPhone app and a backend that lets the AI act on real accounts such as email, calendar, files and GitHub. Version 1.0 of both shipped in August 2026, about two weeks after the first commit. The hard part was not the AI. It was making conversation feel natural: fast replies, interruptions and an honest conversation history.

The challenge#

Most voice assistants feel like walkie-talkies: you speak, wait, then listen. I wanted a personal assistant that felt like talking to a person, and that could take repetitive coordination off my plate: triaging email, checking my calendar, finding a free slot or a nearby place, and drafting messages and posts in my voice.

The bar for success was simple: approving a draft from Zuni had to be faster than doing the task myself. An earlier experiment built on phone shortcuts showed the limits quickly, with no conversation memory and no visual feedback, which pushed me to build a native app.

What I delivered#

  • A native iPhone app and a backend, both tagged version 1.0, with separate development and production apps that install side by side.
  • Natural, full-duplex conversation. Speech is recognised while I talk and the reply plays as it streams in. I can interrupt mid-sentence. Echo cancellation stops it hearing its own voice, interruption triggers on recognised speech rather than loudness, and short acknowledgements fill slow moments instead of dead air.
  • More than twenty-five tools it can use on my behalf, across Gmail, Google Calendar, Google Drive, contacts, notes, weather, nearby places, image search, GitHub and LinkedIn.
  • Rich visual answers alongside the voice: weather, places with maps and photos, and image results as cards.
  • Native iPhone integration: Siri ("Ask Zuni"), the Share Sheet, and understanding photos I show it.
  • Memory of the current conversation and of what matters across conversations.

Timeline#

  • August 2026: backend and iPhone app started within days of each other.
  • Late August 2026: version 1.0 tagged on both, with production running.
  • Early September 2026: refinements to visual answers and tools; about 690 commits across both codebases by then.
  • September 2026: hosting retired to free the server for another product.

How I approached it#

  • A clear split of responsibilities. The phone handles capture and playback. The backend owns the conversation, the AI, the tools and memory. One persistent connection per voice session carries speech in and speech out.
  • A latency budget, measured per turn. Replies are spoken in short pieces as soon as each is ready, the speech connection stays warm between turns, and every turn is timed. I benchmarked models on time to first token and chose one built for live conversation, keeping slower reasoning for background work.
  • Research before patching. When interruption kept triggering on the assistant's own voice, I stopped making small fixes, researched how production voice systems handle echo, and redesigned the voice pipeline in a planned sequence of steps.
  • Clear safety boundaries. Zuni never touches money. Sharing and deleting files needs confirmation, a newly spoken email address is read back before it is used, and every tool call is recorded in an audit log.
  • Secure account access. Each account connects over OAuth, tokens are encrypted before they are stored, and API keys live only on the backend, never in the app.
  • Tested as it was built. Over 770 automated backend tests, alongside tests in the iPhone app.

Hard problems solved#

  • Interruptions that ignore the assistant's own voice. A loudness-based detector was the wrong design. Echo cancellation and interrupting on recognised speech fixed it.
  • Truthful history after an interruption. When I cut in, the conversation records only what I actually heard, not the whole planned reply.
  • Long sentences cut off mid-thought. The speech service was ending turns too early on long utterances. Tuning end-of-speech detection fixed it.
  • One bad record breaking a whole session. A malformed tool-call record broke every later turn in a conversation. A structural cleanup of the history fixed it.
  • Answers that claimed images were on screen. Routing picture requests to the image tool up front, before the model answers, stopped it describing images it had never fetched.

Results#

  • Model time to first token cut from a median of 3.4 seconds to 0.9 seconds, about 2.7 times faster.
  • Speech synthesis starts in 0.75 seconds on a warm connection, against 2.4 seconds when connecting cold.
  • Version 1.0 complete on both the app and the backend, backed by over 770 automated backend tests.

Status#

Zuni was built as a personal assistant and a working demonstration of real-time voice AI. Version 1.0 is complete. The backend is not running publicly today: in September 2026 I retired its hosting to give the server to another product, with everything backed up and ready to restore.

What I learned#

The AI was the easy part. The real engineering time went into latency, interruption and keeping the conversation history honest. Those details decide whether a voice assistant feels like a person or a walkie-talkie.

Have a project in mind?

Share your goals, timeline and any constraints, whether you need it built or want expert advice. I read every message myself, as the engineer who would build it, and reply with a clear view on approach, scope and next steps.

Hire me