<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Voice-Agent on Armstrong Yan</title><link>https://yanqian.github.io/topics/voice-agent/</link><description>Recent content in Voice-Agent on Armstrong Yan</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Thu, 13 Aug 2026 14:13:57 +0800</lastBuildDate><atom:link href="https://yanqian.github.io/topics/voice-agent/index.xml" rel="self" type="application/rss+xml"/><item><title>The Future of Hey Jarvis: When AI Needs an Entry Point</title><link>https://yanqian.github.io/posts/publish/building-hey-jarvis-future/</link><pubDate>Thu, 13 Aug 2026 14:13:57 +0800</pubDate><guid>https://yanqian.github.io/posts/publish/building-hey-jarvis-future/</guid><description>&lt;p&gt;In &lt;a href="https://yanqian.github.io/posts/publish/building-hey-jarvis/" &gt;the first article in this series&lt;/a&gt;, I explained that Hey Jarvis grew out of a small need. When my wife and I were talking before bed and a question came up, we wanted to ask it aloud and get an answer without reaching for a phone.&lt;/p&gt;
&lt;p&gt;I went on to build the Pipeline, Realtime, and Mac App, working through a series of challenges involving &lt;a href="https://yanqian.github.io/posts/publish/building-hey-jarvis-voice-interaction/" &gt;making voice interaction truly work&lt;/a&gt; and &lt;a href="https://yanqian.github.io/posts/publish/building-hey-jarvis-mac-product/" &gt;turning a demo into a Mac product&lt;/a&gt;.&lt;/p&gt;</description></item><item><title>The Challenges of Hey Jarvis, Part I: Making Voice Interaction Actually Work</title><link>https://yanqian.github.io/posts/publish/building-hey-jarvis-voice-interaction/</link><pubDate>Thu, 13 Aug 2026 13:59:31 +0800</pubDate><guid>https://yanqian.github.io/posts/publish/building-hey-jarvis-voice-interaction/</guid><description>&lt;p&gt;In &lt;a href="https://yanqian.github.io/posts/publish/building-hey-jarvis/" &gt;the previous article&lt;/a&gt;, I wrote about how quickly the first Hey Jarvis Pipeline—a serial voice-processing flow—met my original requirements.&lt;/p&gt;
&lt;p&gt;I would say, “Hey Jarvis.” It would record my question, send it to the AI, and read the answer aloud.&lt;/p&gt;
&lt;p&gt;On a flowchart, the voice assistant looked finished.&lt;/p&gt;
&lt;p&gt;But once I started using it, I discovered that the hardest parts of voice interaction often had little to do with what the system recognized or how it responded. The more fundamental questions were harder to answer: When does it start listening? When does it stop? Can it distinguish my voice from its own? And when the conversation ends, who gets the microphone back?&lt;/p&gt;</description></item><item><title>I Built Hey Jarvis: It Started with a Bedtime Question</title><link>https://yanqian.github.io/posts/publish/building-hey-jarvis/</link><pubDate>Thu, 13 Aug 2026 13:55:06 +0800</pubDate><guid>https://yanqian.github.io/posts/publish/building-hey-jarvis/</guid><description>&lt;p&gt;My wife and I often talk before going to sleep.&lt;/p&gt;
&lt;p&gt;Sometimes we talk about what happened that day. Other times, one small detail leads us into history, technology, health, or some fact that has suddenly crossed our minds. Inevitably, we arrive at a question neither of us can answer, but both of us want answered right away.&lt;/p&gt;
&lt;p&gt;The natural thing should be to simply ask.&lt;/p&gt;
&lt;p&gt;Instead, one of us usually reaches for a phone, unlocks it, opens an app, taps a text box, and reformulates the question. By the time the answer appears, our original conversation has already been interrupted.&lt;/p&gt;</description></item></channel></rss>