🎙️ Realtime-Venus-Audio

Talk to Realtime-Venus, a 9B audio-visual interaction model from Ant Group & Tsinghua (MiniCPM-o 4.5 / Qwen3-8B backbone). This demo runs the audio checkpoint's turn-based chat path: speech in → text answer and native speech out, decoded with the bundled Token2wav vocoder.

Project page · Paper · GitHub

Examples (official clips from the Realtime-Venus repository)

Inputs longer than 60 s are truncated. The assistant speaks with the reference voice bundled in the checkpoint (assets/HT_ref_audio.wav). Full-duplex streaming, proactive interaction and asynchronous delegation need the Realtime-Venus-Harness runtime and are not part of this single-turn demo.