🎙️ Realtime-Venus-Audio
Talk to Realtime-Venus, a 9B
audio-visual interaction model from Ant Group & Tsinghua (MiniCPM-o 4.5 / Qwen3-8B
backbone). This demo runs the audio checkpoint's turn-based chat path: speech in →
text answer and native speech out, decoded with the bundled Token2wav vocoder.
Project page · Paper · GitHub
Examples (official clips from the Realtime-Venus repository)
Inputs longer than 60 s are truncated. The assistant speaks with the reference voice bundled in the checkpoint (assets/HT_ref_audio.wav). Full-duplex streaming, proactive interaction and asynchronous delegation need the Realtime-Venus-Harness runtime and are not part of this single-turn demo.