Vladyslav Didyk

NVIDIA PersonaPlex

Full-duplex speech model: listens while it talks, with a configurable voice and role.

Natural-sounding phone or kiosk agent. Today a research model, better as a demo than as production.

Research release, last pushed March 2026, needs a strong GPU. Use Pipecat for paid work.

Code: MIT (file LICENSE-MIT). Model weights (nvidia/personaplex-7b-v1): NVIDIA Open Model License Agreement; the model card also lists CC-BY-4.0 as additional information for the underlying Kyutai Moshi weights.

Yes: code is MIT and the model card states 'This model is ready for commercial use', but the weights are under the separate NVIDIA Open Model License (gated on Hugging Face: you must log in, share contact information and accept the terms), so the client deployment must comply with that licence, not just MIT.

Needs a GPU: model card lists NVIDIA Ampere (A100) / Hopper (H100) as supported and 'NVIDIA A100 80 GB' as test hardware; no minimum VRAM is stated anywhere I opened (README offers --cpu-offload if GPU memory is insufficient). 7B speech-to-speech model based on Moshi. Hugging Face token and licence acceptance required. Last push 2026-03-02 (research code release, low activity). Voice conditioning: get consent for any real person's voice. No vendor-hosted product found.

Same buyers as Pipecat

Natural-sounding phone or kiosk agent. Today a research model, better as a demo than as production.

none found

None named

Catalog