How full-duplex WebRTC streaming, acoustic tokenizers, and edge inference make human-speed conversational voice possible.