<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Realtime on Armstrong Yan</title><link>https://yanqian.github.io/tags/realtime/</link><description>Recent content in Realtime on Armstrong Yan</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Thu, 13 Aug 2026 13:59:31 +0800</lastBuildDate><atom:link href="https://yanqian.github.io/tags/realtime/index.xml" rel="self" type="application/rss+xml"/><item><title>The Challenges of Hey Jarvis, Part I: Making Voice Interaction Actually Work</title><link>https://yanqian.github.io/posts/publish/building-hey-jarvis-voice-interaction/</link><pubDate>Thu, 13 Aug 2026 13:59:31 +0800</pubDate><guid>https://yanqian.github.io/posts/publish/building-hey-jarvis-voice-interaction/</guid><description>&lt;p&gt;In &lt;a href="https://yanqian.github.io/posts/publish/building-hey-jarvis/" &gt;the previous article&lt;/a&gt;, I wrote about how quickly the first Hey Jarvis Pipeline—a serial voice-processing flow—met my original requirements.&lt;/p&gt;
&lt;p&gt;I would say, “Hey Jarvis.” It would record my question, send it to the AI, and read the answer aloud.&lt;/p&gt;
&lt;p&gt;On a flowchart, the voice assistant looked finished.&lt;/p&gt;
&lt;p&gt;But once I started using it, I discovered that the hardest parts of voice interaction often had little to do with what the system recognized or how it responded. The more fundamental questions were harder to answer: When does it start listening? When does it stop? Can it distinguish my voice from its own? And when the conversation ends, who gets the microphone back?&lt;/p&gt;</description></item></channel></rss>