Work/2026

    Real-Time Voice Companion

    Press one button and talk. It listens, thinks and answers out loud, with no typing and no menus.

    Role
    Sole engineer
    Client
    Independent founder, private contract
    Built with
    OpenAI Realtime · Next.js · Semantic VAD · Netlify
    The voice companion interface: a short explanation and a single Talk button
    The whole interface. Voice-first means there is almost nothing to look at.

    The interesting problem

    Realtime voice is easy to demo and hard to make comfortable. The hardest issue here was self-interruption: the app's own voice coming out of a phone speaker was picked up by its own microphone, and the system treated that as the user interrupting.

    A cough in the room could stop it mid-sentence.

    What actually fixed it

    Several detection settings were tried — different voice-activity modes, sensitivity thresholds, and far-field noise reduction. Each reduced the problem; none removed it, because the microphone stays open while the model speaks and no server-side setting fully solves that.

    The reliable fix was a mute control the user holds. Not clever, but it works every time, which the clever options did not.

    That is worth stating plainly: the honest answer was a simpler interface, not a better parameter.

    The realtime voice loop
    1. 01

      Microphone

      stays open while the model speaks

    2. 02

      Voice detection

      decides when you have finished

    3. 03

      Realtime model

      listens and answers as audio

    4. 04

      Speaker

      audible to the same microphone

    • The loop is the problem: the app's own voice, out of a phone speaker, re-enters its own microphone and reads as an interruption.
    • Tuning detection modes, thresholds and far-field noise reduction each reduced it. None removed it. A user-held mute control did.

    Live demo