Models useful for discourse and dialogue research

If you want to add other models or find errors, please create GitHub issues or pull requests (Edit this file.). If you don’t have an account on GitHub, please email at resources@sigdial.org.

Name Category Language Brief Description Paper
Moshi Full-duplex spoken dialogue model English Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec. Défossez et al., 2024
VoiceActivityProjection Self-supervised learning of Turn-taking Events trained with English data Voice Activity Projection is a Self-supervised objective for Turn-taking Events. Ekstedt and Skantze, 2022