Corpora and datasets for discourse and dialogue reserach

Parts of the contents of the list are extracted from the papers using LLMs, so they might be wrong. If you find errors, please create GitHub issues or pull requests (Edit this file.). If you don’t have an account on GitHub, please email at resources@sigdial.org.

Parts of this list have been adapted from A Survey of Available Corpora for Building Data-Driven Dialogue Systems, with permission; see the survey website for reference and please cite the paper if useful.

We also referred to the survey paper On the Need for Thoughtful Data Collection for Multi-Party Dialogue: A Survey of Available Corpora and Collection Methods. We would like to thank the authors.

We would also like to thank David Traum who provided the information.

Name Language Modalities Data Types Task/Domain Participants Size Ave. # of Turns Brief Description Paper
Let’s go & DSTC1 English Speech Audio Bus schedules Human-System 171K dialogues N/A Telephone conversations between real users and bus information systems Raux et al. 2006
Georgetown University Multilayer corpus (GUM) English Mixed (text and speech) text, markup and transcripts 24 spoken and written genres Human-Human ~300K tokens ~55 utterances per document A multilayer English corpus of 24 spoken and written genres annotated for RST and PDTB discourse relations, subtyped coreference and bridging anaphora, entity and proposition salience, multiple summatization, UD syntax and more Zeldes et al. 2025
Georgetown Chinese Discourse Treebank Mandarin Chinese Mixed (text and speech) text, markup and transcripts 5 spoken and written genres Human-Human ~63K tokens ~54 utterances per document A multilayer Chinese corpus of 5 spoken and written genres annotated for RST discourse relations and dependencies, UD syntax and more Peng et al. 2022
DSTC2 English Speech Transcripts and ASR results Restaurant search Human-System 15K dialogues, 3.7M words 7.88 Telephone conversations between hired users and restaurant search system Henderson et al, 2014
MultiWoz 2.0 English Text Text Multiple domains (restaurant, hotel, etc.) Human-Woz 8.5K dialogues, 115K turns, 1.5M tokens 13.18 A fully-labeled collection of human-human written conversations spanning over multiple domains and topics Budzianowski et al., 2018
HCRC MapTask Corpus English Face-to-face Audio, video (not available) direction giving Human-Human 128 dialogues, 174K words, 18hrs A set of 128 dialogues that has been recorded, transcribed, and annotated for a wide range of behaviours, and has been released for research purposes. Anderson et al., 1991
AMI Corpus English face-to-face close-talking and far-field microphones, individual and room-view video cameras, projection, a whiteboard, individual pens. Face-to-face meetings Multi-party human 175 dialogues, 900K words, 100hrs A multi-modal data set consisting of 100 hours of meeting recordings Carletta et al, 2005
Ubuntu Dialogue Corpus English IRC chat text Chat on Ubuntu Human-Human 930K dialogues, 100M words 7.71 Dialogues extracted from Ubuntu chat stream on IRC Lower et al, 2015
DailyDialog Dataset English Text Text Daily communication Human-Human 13K dialogues, 1.5M words 7.9 DailyDialog is a high-quality multi-turn dialogue dataset that covers conversations about daily life. It is manually labeled with communication intention and emotion information, making it useful for training and evaluating dialogue systems. Li et al. 2017
Persona Chat English Chat text Text Open domain Human-Human 11K dialogues, 162K utterances A chit-chat dataset where paired Turkers are given assigned personas and chat to try to get to know each other. Zhang et al., 2018
Schema-Guided Dialogue Dataset English Text Text 16 domains Human-System 16K dialogues, 330K turns The dataset consists of conversations between a virtual assistant and a user ranging over a variety of domains including Travel, Events, Payment, Media, Restaurants, Weather etc. Annotations for natural language understanding, dialogue state tracking, policy learning, natural language generation and user simulation learning are also included. Rastogi et al., 2020
EmoWOZ English Text Text Multiple domains (restaurant, hotel, etc.) Human-Woz More than 11K dialogues 14.63 A large-scale open-source dataset for emotion recognition in task-oriented dialogues with n 83K emotion annotations of user utterances Feng et al. 2022
Schema-Guided Dialogue (SGD) English text text 26 services across 16 domains including alarms, banks, buses, calendar events, flights, homes, hotels, media, movies, music, payment, rental cars, restaurants, ridesharing, services, trains, travel, messaging, and weather Simulated user-system interactions Over 16,000 dialogues, 329,964 turns 20.44 The SGD dataset is designed to support the development of conversational interfaces that can handle multiple domains and services, particularly in scenarios with zero-shot learning where models encounter unseen services or APIs. It uses a schema-guided approach where intents and slots are dynamically provided, facilitating easier integration of new services without retraining. Rastogi et al., 2020
Internet Argument Corpus 2.0 English text text Online forums and debates on social and political topics Human-Human 24,000 posts, 11,079 threads, 3452 authors, 56M tokens Varies, data includes multiple posts per thread The IAC 2.0 is an expanded dataset designed to support research on many different aspects of social language and dialogue structure, particularly in online forums on social and political topics. It features an SQL schema for organizing dialogues from several platforms into a structured database format. Abbott et al., 2016
The Settlers of Catan Corpus English text text Game strategy and conversation Human-Human 21 games annotated, ca. 2000 dialogue turns, ca. 40 games collected Includes ‘a few dozen self-contained bargaining conversations’ per game A corpus of online chats between agents playing The Settlers of Catan, a competitive win–lose game involving negotiations. The corpus aligns players’ conversations with the state of the game, focusing on negotiation dialogues and strategic interactions. Afantenos et al., 2012
Let’s Go Public corpus English speech audio Public transportation Human-System 627 dialogues, 9162 turns 14.6 The corpus contains dialogues from the Let’s Go Public spoken dialog system, which provides bus schedule information during off-peak hours. It includes transcribed calls from the general public, featuring interactions influenced by various user attitudes and environmental conditions. Raux et al., 2005
Dialog State Tracking Challenge English speech text Bus timetable information Human-System 15K transcribed and labeled human-computer dialogs Varies by dataset; e.g., TRAIN1A: 14.7, TEST4: 10.9 A corpus of 15,000 human-computer dialogue interactions used for evaluating dialogue systems, specifically focusing on the task of dialog state tracking. The corpus contains dialogs from various dialog systems interacting with real users, collected under the Spoken Dialog Challenge hosted by Carnegie Mellon University. Williams et al., 2013
Carnegie Mellon Communicator English speech audio Travel planning (air transportation, hotel reservations, car rentals) Human-System N/A N/A The Carnegie Mellon Communicator system assists users in creating complex travel itineraries through a conversational interface. It utilizes schemas to manage dialogues, aiming to support problem-solving activities by providing information, proposing solutions, and highlighting potential constraint violations. Rudnicky et al., 1999
ATIS Spoken Language Systems Pilot Corpus English speech audio, text Air travel information Human-Woz 41 sessions, 1041 utterances 25.4 utterances per session The ATIS corpus is designed for developing and evaluating speech systems that understand spontaneous speech, focused on air travel information. Hemphill et al, 1990
RITEL Corpus French speech audio Open-domain Human-System 582 dialogs, 5360 user queries, 6 hours of user speech 9 The RITEL Corpus is a Human-Computer open-domain question answering spoken dialog corpus that includes orthographically transcribed and annotated dialogues focusing on specific entities and topics. It involves a real interaction system rather than a Wizard-of-Oz setup. Rosset and Petel, 2006
The MATCH corpus English speech audio Healthcare, appointment scheduling Human-Human 447 dialogues, 6237 turns 14.0 The MATCH corpus is a linguistically annotated corpus collected to study the interaction between older and younger users with simulated spoken dialogue systems. It focuses on the effects of cognitive ageing on users’ interactions and was designed to develop technologies to help older users live independently. Georgila et al, 2010
Frames English text text Travel Human-Human 1369 dialogues, 19986 turns 15 Frames is a corpus of human-human dialogues collected in a Wizard-of-Oz setting to study complex dialogue flows and decision-making behaviour. The dialogues involve users trying to book travel packages with constraints, exploring options and making selections, facilitated by assistants who manage these requests. El Asri et al., 2017
Multi-Domain In-Car Assistant Dialogue Dataset English text text Calendar scheduling, weather information retrieval, point-of-interest navigation Human-Woz 3,031 dialogues; 2,425 training, 302 validation, 304 test dialogues 5.25 This dataset contains dialogues across three domains relevant to in-car personal assistant tasks. Each dialogue is grounded in a knowledge base, making it suitable for developing architectures that reason over world knowledge. Eric et al., 2017
The Walking Around Corpus English speech audio Pedestrian navigation and spatial cognition Human-Human 36 dialogues, detailed transcripts Multiple tasks involved The corpus consists of experimentally parameterized collection of spontaneous spoken dialogues, focusing on lexical choice and variability during direction-giving tasks. It involves participants communicating over mobile phones while one navigates a campus based on directions from a stationary partner. Brennan et al., 2013
Intelligence Squared Debates (IQ2 Debates) English speech text Various (e.g., foreign policy, health, technology) Human-Human 108 debates, average 12,801 words and 117 turns per debate 117 A corpus of transcripts from Oxford-style debates held in the US, covering a wide range of topics with experts debating motions before a live audience. The dataset tracks conversational dynamics and strategies used to sway audience opinions. Zhang et al., 2016
Idiap Wolf Database English multimodal audio, video role-playing game, competitive Human-Human 7.3 hours of recordings, 50 day-phase games, 36 participants N/A The Idiap Wolf Database consists of audio-visual recordings from a competitive role-playing game where players have deceptive and non-deceptive roles. The unique aspect of this corpus is its focus on group behavior and deception in a controlled game setting. Hung and Chittaranjan, 2010
ICSI Meeting Recorder Dialog Act (MRDA) Corpus English speech audio, text natural meetings Human-Human 75 meetings, approx. 72 hours of speech, 180,218 dialog act tags N/A A corpus of hand-annotated dialog acts and adjacency pairs from naturally occurring multi-party meetings recorded at the ICSI. It includes over 180,000 dialog act tags across approximately 72 hours of meetings, focusing on complex discourse phenomena. Shriberg et al., 2004
The Trains 93 Dialogues English speech audio Task-oriented dialogues involving a planning assistant and manufacturing and shipping goods Human-Human 98 dialogues, 5900 turns, 55000 words Approximately 60.2 A corpus of task-oriented dialogues set in the Trains domain where a user collaborates with a planning assistant to accomplish tasks involving manufacturing and shipping goods in a railroad freight system. Includes audio files, time-aligned word and phoneme transcriptions. Heeman and Allen, 1995
ICT Rapport Datasets English multimodal audio, video Narrative task involving retelling events from a sexual harassment awareness video Human-System 131 participants N/A The Rapport Agent is designed to elicit rapport from human participants within a dyadic narrative task. It utilizes real-time analysis of acoustic properties of speech and speaker gestures to generate nonverbal feedback like nods and posture shifts. Gratch et al., 2007
D64 Multimodal Conversational Corpus English multimodal text, audio, video General conversation Human-Human N/A N/A A corpus designed to observe conversational behavior as closely as possible to natural interaction, including elements like gaze, posture, and simultaneous movements. The data, collected in a domestic setting, includes extensive video, audio, and motion-capture records. Oertel et al., 2013
Cardiff Conversation Database (CCDb) English audiovisual audio, video Natural conversations Human-Human 30 conversations, 300 minutes of audio-video data Approximately 10 per conversation (estimated from 5-minute average duration per conversation) A unique 2D audiovisual database containing natural conversations between pairs of people, annotated for speaker activity, facial expressions, head motion, and non-verbal utterances. Aubrey et al., 2013
4D Cardiff Conversation Database (4D CCDb) English multimodal 3D video (4D), audio Natural, dyadic conversations Human-Human 17 minutes, 34 sequences N/A The 4D CCDb is the first 4D (3D Video) audio-visual database containing natural conversations between pairs of people. It includes fully annotated speaker and listener activities such as conversational facial expressions, head motion, and verbal/non-verbal utterances. Vandeventer et al., 2015
Group Affect and Performance (GAP) Corpus English multimodal audio, text Group interaction and decision-making Human-Human 13 group meetings, 104.45 minutes of recordings N/A The GAP corpus contains meeting audio, transcriptions, annotations, decision-making performance, as well as group member influence, post-meeting ratings of satisfaction, and demographics. It is designed to stimulate research on the computational analysis of small group meetings. Braley and Murray, 2018
MULTISIMO Corpus English multimodal text, audio, video Collaborative group interactions in a quiz solving task Human-Human 23 sessions, approximately 4 hours total N/A The MULTISIMO Corpus involves collaborative group interactions where participants work together to solve quiz questions. It includes multimodal data from different cameras and microphones, synchronized and complemented by personality test results and experience assessment surveys. Koutsombogera and Vogel, 2018
Movie-DiC English text text Multiple genres (action, crime, drama, thriller, etc.) Human-Human 132,229 dialogues, 764,146 turns 5.78 A dialogue corpus extracted from movie scripts for studying semantic and pragmatic aspects of human communication in various contexts and styles. Banchs, 2012
Cornell Movie-Dialogs Corpus English text text Movie scripts Human-Human 220,579 conversational exchanges from 617 unique titles 5 or more exchanges per pair A large set of imagined conversations derived from movie scripts, providing a rich resource for studying linguistic coordination and stylistic convergence in fictional dialogues. Danescu-Niculescu-Mizil and Lee, 2011
Conversation Dialog Corpora from Television and Movie Scripts English text text Television shows and movies Human-Human 1,042,288 dialog pairs (raw), 86,719 dialog pairs (after filtering) N/A This corpus contains conversation pairs extracted from television and movie scripts. The dialogues are filtered to ensure they are between two speakers, using a method called tri-turn filtering and semantic similarity filtering. The final corpus includes 86,719 high-quality query-response pairs. Nio et al., 2014
TVD: a reproducible and multiply aligned TV series dataset English text text, audio, video TV Series (The Big Bang Theory and Game of Thrones) Human-Human 132 episodes of TBBT, 5 episodes of GoT (manual transcripts), 17 TBBT and 10 GoT episodes (subtitles), 17 TBBT and 10 GoT episodes (automatic transcripts), outlines and summaries for multiple episodes N/A The TVD dataset is built around two TV series, The Big Bang Theory and Game of Thrones, and includes multiple tracks such as manual and automatic transcripts, multilingual subtitles, episode outlines, and various metadata. The dataset is designed for tasks like summarization, scene retrieval, and speech retrieval. Roy et al., 2014
Annotated Corpus of Film Dialogue for Learning and Characterizing Character Style English text text Film dialogue from multiple genres (drama, thriller, crime, comedy, action, romance, adventure) Human-Human 862 film scripts, 664,000 lines of dialogue, 9,599,000 tokens N/A A corpus of film dialogue collected from the IMSDb archive, annotated for linguistic structures and character archetypes, used to learn character models of linguistic style. Walker et al., 2012a
SubTle Corpus English, Portuguese text text Horror, Sci-fi, Western, Romance Human-Human SubTle - Portuguese: 2,930,173 I-R pairs; SubTle - English: 3,454,480 I-R pairs Varies by genre, average ranges from 419 to 580 I-R pairs per subtitle file A corpus of Interaction-Response pairs extracted from subtitles files, created to help dialogue systems deal with Out-of-Domain interactions. Ameixa and Coheur, 2013
OPUS Multiple languages (over 90 languages) text text Multiple domains (legislative texts, administrative texts, movie subtitles, software localization, newspaper texts) Human-Human Over 40 billion tokens, 2.7 billion parallel units (aligned sentences and sentence fragments) N/A A growing language resource of freely accessible parallel corpora and related tools, used for various applications including machine translation, translation studies, and cross-linguistic corpus studies. Tiedemann, 2012
NPS Internet Chatroom Conversations English text text General chat, open to any topic Human-Human 10K posts, 45K tokens N/A The corpus consists of online chat dialogues collected from various chat rooms, annotated with lexical, syntactic, and discourse information. It was developed to support natural language processing applications such as author profiling, entity identification, and social network analysis. Forsyth and Martell, 2007
Twitter Conversations Corpus English text text Open-domain (Twitter conversations) Human-Human 1.3 million conversations 2 (majority of conversations have only 2 posts) A large corpus of 1.3 million Twitter conversations, enabling the study of open-domain dialogue acts and structure in a new medium. Ritter et al., 2010
Twitter Triple Corpus English text text Social Media (Twitter) Human-Human 127M triples N/A (Context + Message + Response as triples) A large-scale corpus mined from Twitter, used for training context-sensitive response generation models. The corpus consists of triples representing context, message, and response. Sordoni et al., 2015
NUS SMS Corpus English, Chinese text text General SMS communication Human-Human 57,824 messages N/A A public SMS corpus focusing on English and Mandarin Chinese SMS messages, collected through crowdsourcing methods. Chen and Kan, 2013
Settlers of Catan Strategic Conversation Corpus English text text Game negotiation (Settlers of Catan) Human-Human 21 games annotated with approximately 2000 dialogue turns Varies per game, approximately a few dozen per game A corpus of online chat negotiations during the game The Settlers of Catan, focusing on strategic conversation and negotiation dialogues. Afantenos et al., 2012
Cards corpus English text text Task-oriented (card game in a maze-like environment) Human-Human 744 transcripts, 23,532 utterances, 137,323 words 31.63 The Cards corpus is built from a two-person online video game where players collaborate to complete a task. The game records everything, allowing for detailed study of player utterances, context, and strategies in a simple, controlled environment. Djalali et al., 2012
Agreement by Create Debaters (ABCD) English text text Online discussion forums (e.g., createdebate.com) Human-Human 10K discussions, 200K posts approximately 20 turns per discussion A large corpus derived from the Create Debate website, containing over 10,000 discussions with more than 200,000 posts annotated for agreement, disagreement, or neutrality. Rosenthal and McKeown, 2015
Internet Argument Corpus (IAC) English text text Political debate and discourse Human-Human 390,704 posts in 11,800 discussions N/A A corpus for research on deliberation and debate, containing argumentative discourse from the online debate site 4forums.com. It includes posts on various political and social topics with annotations for topic, stance, and various dialogic and argumentative markers. Walker et al., 2012b
Multi-Party Chat (MPC) Corpus English text text Online chat environments Human-Human 7317 turns, 58175 words Approximately 520 per session A corpus of multi-party online conversations collected in a chat-room environment to model social phenomena such as agenda control, influence, and leadership in online interactions. Shaikh et al., 2010
Ubuntu Chat Corpus Multiple languages (English, Chinese, Russian, Brazilian Portuguese, Spanish, Italian, Polish, Swedish) text text Technical support for Ubuntu OS Human-Human 11 channels, 40M+ messages, 2.9GB (compressed to 0.6GB) Average message length varies across channels (21.7 to 57.6 characters) The Ubuntu Chat Corpus is a large, publicly available corpus consisting of IRC chat logs from various Ubuntu support channels. It includes messages in multiple languages and covers technical discussions related to Ubuntu OS. Uthus and Aha, 2013
The Movie Dialog Dataset English text text Movies Human-Human ∼75k movie entities, ∼3.5M training examples Varies by task A set of four tasks designed to evaluate different prerequisite qualities of end-to-end dialog systems, focusing on the movie domain. These tasks include question-answering, recommendation, QA+recommendation dialog, and Reddit discussion. Dodge et al., 2015
Cooperative Vision-and-Dialog Navigation (CVDN) English multimodal text, image Navigation in simulated, photorealistic home environments Human-Human 2050 dialogues, 7k navigation trajectories 6 A dataset of over 2k embodied, human-human dialogues situated in simulated, photorealistic home environments for studying vision-and-dialog navigation tasks. Thomason et al., 2020
Talk The Walk English multimodal text, audio Navigation in NYC neighborhoods Human-Human 10,310 dialogues 62 Talk The Walk is a large-scale dialogue dataset grounded in action and perception, where a ‘guide’ and a ‘tourist’ communicate to achieve the goal of navigating the tourist to a target location in New York City. De Vries et al., 2018
Japanese Emotion-Tagged Dialogue Corpus Japanese text text Twitter dialogues Human-Human 3,828 dialogues, 13,806 utterances 3.6 A Japanese dialogue corpus annotated with expressed and experienced emotions for each utterance, collected from Twitter. Ide and Kawahara, 2022
MultiWOZ 2.1 English text text Multiple domains (hotel, taxi, restaurant, etc.) Human-Woz 10K dialogues, over 115K turns 11.5 MultiWOZ 2.1 is a multi-domain dialogue dataset with corrections in state annotations and dialogue utterances, building on the original MultiWOZ 2.0. It includes system and user dialogue acts and offers a benchmark for dialogue state tracking models. Eric et al., 2019
MultiWOZ 2.2 English text text Multiple domains (Restaurant, Hotel, Attraction, Taxi, Train, Hospital, Bus, Police) Human-Woz 10K dialogues, 115K turns N/A MultiWOZ 2.2 is an updated version of the MultiWOZ dataset, with corrections to dialogue state annotations, redefined ontology, and additional slot span annotations. It is used as a benchmark for dialogue state tracking in task-oriented dialogues across multiple domains. Zang et al., 2020
MultiWOZ 2.3 English text text Multiple domains (Train, Taxi, Hotel, Restaurant, Attraction, Hospital, Bus, Police) Human-Woz 10K dialogues, 2.5M tokens unknown MultiWOZ 2.3 is a multi-domain task-oriented dialogue dataset with enhanced annotation corrections and co-reference annotation. Han et al., 2021
MultiWOZ 2.4 English text text Multiple domains (e.g., restaurant, hotel, taxi) Human-Woz 2,000 dialogues, 14,000 turns N/A MultiWOZ 2.4 is an updated version of the MultiWOZ 2.1 dataset. It includes refined annotations in the validation set and test set to improve the evaluation of dialogue state tracking models, focusing on task-oriented dialogues across multiple domains. Ye et al., 2022
JMultiWOZ Japanese text text travel-related domains (tourist attractions, accommodation, restaurants, shopping facilities, taxis, weather) Human-Woz 4,246 dialogues, 61,186 turns, 1.1M tokens 14.4 A large-scale Japanese multi-domain task-oriented dialogue dataset focused on travel-related domains. Ohashi et al., 2024
RealPersonaChat (RPC) Japanese text text General chit-chat conversations Human-Human 14K dialogues, 421K utterances, 5.55M tokens 30.09 A large-scale realistic dialogue corpus in Japanese that includes the actual personas and personality traits of the interlocutors. It is the world’s largest corpus of dialogue data that includes personas and personality traits. Yamashita et al., 2023
DIHANA Spanish speech audio Train services (nationwide trains in Spain) Human-Woz 900 dialogues, 6,278 user turns, 9,129 wizard turns, 48,243 words 7.0 Spontaneous speech dialogues for train service queries using the Wizard of Oz technique, focused on information retrieval for nationwide trains in Spain. Benedí et al, 2006
Wizard of Wikipedia English text text Open-domain (various topics including commuting, music festivals, Arnold Schwarzenegger, etc.) Human-Human 22.3K dialogues, 201.9K turns 9.0 Open-domain dialogues grounded with knowledge retrieved from Wikipedia, focusing on conducting knowledgeable discussions. Dinan et al., 2018
FoCus (Call For Customized conversation) English text text Geographical landmarks Human-Machine 14,452 dialogues, 173,424 utterances 11.99 The FoCus dataset contains conversations about geographical landmarks, where the machine provides customized and knowledgeable responses by grounding the dialogue in both Wikipedia knowledge and user persona. Jang et al., 2022
MPCHAT English multimodal text, image Episodic memory-based dialogues sourced from Reddit Human-Human 15K multi-turn dialogues, 42,531 utterances by 25,877 users 2.83 (approx.) A multimodal persona-grounded dialogue dataset where personas reveal speakers’ episodic memories using both text and images. Ahn et al., 2023
DuLeMon Chinese text text Open-domain dialogue with a focus on long-term persona memory Human-Chatbot 27,501 dialogues 16.2 DuLeMon is a dataset designed for studying long-term memory conversation tasks in Chinese. It focuses on the active construction and utilization of the user’s persona in long-term interactions, with explicit annotation of persona-related information in each dialogue. Xu et al., 2022b
MSPD (Multi-Session Personalized Dialogue) Korean text text Personalized conversations, including daily, knowledge-based, empathetic, and personalized dialogues Human-Human-System 13,469 episodes, 53,880 sessions, 601,062 utterances 11.15 A Korean Multi-Session Personalized Dialogue dataset designed to enable models to generate personalized responses grounded on user persona attributes, focusing on natural and engaging conversation across multiple sessions. Kwon et al., 2023
BlendedSkillTalk English text text Multiple domains (personal background, knowledge, empathy) Human-Human 5k conversations, 56k utterances 11.2 BlendedSkillTalk is a dataset designed to evaluate a model’s ability to blend multiple conversational skills—knowledge, empathy, and personal background—within a single conversation. Smith et al., 2020
Empathetic Dialogues English text text Emotional situations in personal conversations Human-Human 25K dialogues, 24,850 conversations 4.31 A dataset of 25k conversations grounded in emotional situations, designed to improve empathetic dialogue generation. Rashkin et al., 2019
PEC (Persona-based Empathetic Conversations) English text text Multiple domains (happy, offmychest) Human-Human 355K conversations Training set has 6 most recent turns per conversation A large-scale, multi-domain dataset for persona-based empathetic conversations collected from Reddit, focusing on the impact of persona on empathetic responses. Zhong et al., 2020
PersonaMinEdit English text text Persona-grounded dialogues Human-Human Multiple human references N/A PERSONAMINEDIT is a dataset designed to evaluate persona-grounded minimal editing, focusing on editing dialogue responses to improve persona consistency while maintaining coherence with the dialogue history. Wu et al., 2021a
Inadequate-Tiny-ConvAI2 (IT-ConvAI2) English text text Dialogue generation domain Human-Human 1,595 conversations N/A IT-ConvAI2 is a dataset that emphasizes the out-of-predefined persona (OOP) problem in personalized dialogue generation. It is built by removing query-related personas from the original ConvAI2 dataset. Liu et al., 2022
LiveChat Chinese text text Live streaming, multi-party conversations Human-Human 1.33M dialogues, 9.4M utterances 7.1 A large-scale personalized dialogue dataset automatically constructed from live streaming videos, containing detailed persona profiles and multi-party conversations. Gao et al., 2023
PER-CHAT English text text Open-domain Human-Human 1.5M dialogues, 300K user profiles Single-turn dialogues PER-CHAT is an open-domain single-turn dialogue dataset consisting of 1.5M conversations and 300k user profiles collected from Reddit. It includes detailed personalization information such as user profiles and comment histories, making it suitable for generating personalized responses in dialogue systems. Wu et al., 2021b
Pchatbot Chinese text text Open-domain (Weibo), Professional domain (Judicial forums) Human-Human 198.88M dialogues, 397.75M utterances 26.21 for PchatbotW, 2.95 for PchatbotL Pchatbot is a large-scale Chinese conversation dataset dedicated to the development of personalized dialogue models, containing two subsets collected from Weibo and Judicial forums respectively. The dataset includes anonymized user IDs and timestamps to enable personalized dialogue modeling. Qian et al, 2021
Multimodal EmotionLines Dataset (MELD) English multimodal text, audio, video emotion recognition in conversations Human-Human 1,433 dialogues, 13,000 utterances 9.6 MELD is a multimodal multi-party conversational emotion recognition dataset that includes text, audio, and visual data from the TV series Friends. It is designed for emotion recognition in conversations. Poria et al., 2019
Multi-Party Dialogue Dataset (MPDD) Chinese text text Social interactions, Interpersonal relationships Human-Human 4,142 dialogues, 25,548 utterances 6.168 MPDD is a Chinese multi-party dialogue dataset annotated with emotion and interpersonal relationship labels on each utterance. The dialogues are sourced from TV series scripts and are designed to facilitate the analysis of emotions and relationships in social dialogues. Chen et al., 2020
RobotSlang Benchmark English text text, audio, video Robot Localization and Navigation Human-Human 169 dialogues, nearly 5k utterances, 1k minutes of robot camera and control streams 28 A benchmark of human-human cooperative trials for controlling a physical robot through natural language dialogues, focusing on localization and navigation tasks. Banerjee et al., 2020
TEACh (Task-driven Embodied Agents that Chat) English multimodal text, actions (environment interactions) Household tasks in a simulated environment Human-Human 3,047 dialogues 13.67 TEACh is a dataset of over 3,000 human-human dialogues where a Commander with oracle task knowledge communicates with a Follower to complete household tasks in a simulated environment. The dataset supports studies on embodied intelligence, including language grounding, dialogue understanding, and task execution. Padmakumar et al., 2021
Minecraft Dialogue Corpus English text text Collaborative building in Minecraft Human-Human 509 dialogues, 15,926 utterances, 113,116 tokens 30.7 A collection of 509 human-human written dialogues and game logs for a collaborative building task in a Minecraft-based environment, where one player instructs another to build a structure. Narayan-Chen et al., 2019
DialFRED English multimodal text, audio, video Household tasks (navigation and object manipulation) Human-Agent 53K task-relevant questions and answers N/A DialFRED is a dialogue-enabled embodied instruction following benchmark that allows an agent to actively ask questions and use the information in the response to better complete household tasks. It is built by augmenting the ALFRED benchmark and includes a human-annotated dataset with 53K task-relevant questions and answers. Gao et al., 2022
Dialog State Tracking Challenge 3 (DSTC3) English speech text, audio Tourist information (restaurants, pubs, coffee shops) Human-System 2,275 dialogs, 17,677 turns N/A The third Dialog State Tracking Challenge (DSTC3) focused on evaluating the ability of trackers to generalize to new entities, such as new slots and values not present in the training data. The challenge involved human-computer dialogs in the tourist information domain, covering restaurants, pubs, and coffee shops in Cambridge, UK. Henderson et al., 2014
Friends TV Show Emotion Corpus English text text TV Show Transcripts Human-Human 12,606 utterances, 897 scenes, 97 episodes 14.05 A corpus comprising transcripts from the TV show Friends, annotated with seven emotions on consecutive utterances in multiparty dialogues. Zahiri and Choi, 2017
Hazumi Japanese multimodal text, audio, video, posture, physiological data chit-chat (food, travel, etc.) Human-WoZ 214 dialogues (15 to 20 minutes), 18,162 exchanges 84.9 A multimodal dialogue corpus with various manual annotations, including those provided by five third-party annotators as well as those given by the participants themselves. The corpus also includes physiological data. Komatani and Okada, 2021
KokoroChat Japanese Text (role-play) Text Psychological counseling Human-Human (trained counselor role-play) 6,589 dialogues ~91.2 utterances per dialogue A high-quality, human-collected Japanese psychological counseling dialogue dataset where trained counselors simulate both client and counselor in one-hour text-based sessions, with detailed client feedback per session (20 rating items). Qi et al., 2025
Switchboard Telephone Speech Corpus (Switchboard-1) English Speech (telephone conversations) Audio, transcripts Open-domain conversational speech Human-Human Approximately 2,400 dialogues (~260 hours of speech; ~3 million words) ~6 minutes per dialogue (i.e., ~12 turns typical) — average not explicitly given Spontaneous two-speaker telephone conversations across roughly 70 topics, fully transcribed and time-aligned, with speaker demographics and call metadata recorded for speech technology and linguistic research Godfrey et al., 1992
CALLHOME American English Speech (LDC97S42) English Speech (telephone conversations) Audio (2-channel μ-law at 8 kHz), with optional transcripts (LDC97T14) Open-domain personal telephone conversations Human-Human 120 dialogues (~30 minutes each; ~60 hours total) N/A (unspecified average turns) Unscripted telephone calls between native speakers, mostly family or friends, fully recorded and documented for ASR research. Canavan et al., 1997
CALLFRIEND American English-Non-Southern Dialect (LDC96S46) English Speech (telephone conversations) Audio (2-channel μ-law at 8 kHz) Open-domain conversational speech Human-Human 60 dialogues, each 5–30 minutes (up to ~30 minutes each) N/A (not specified) Unscripted telephone conversations between native speakers of non-Southern American English, with metadata such as speaker demographics and call quality, collected for language identification research Canavan & Zipperlen, 1996
Corpus of Spoken Professional American-English (CSPA) English Speech transcripts Text (transcripts) Professional domain: academic meetings and press conferences Human-Human (various professional speakers) ~2 million words across two sub-corpora of ~1 million words each (17 files) N/A Transcripts of unscripted spoken interactions—mainly faculty council and committee meetings, and White House press conferences—minimally coded to retain hesitations and disfluencies. Barlow, 2000
COLT – The Bergen Corpus of London Teenage Language English Speech (audio recordings with transcripts) Audio, orthographic and prosodic transcripts, POS tagging Spontaneous teenage talk (informal, conversational) Human-Human (peer teenage conversations) ~500,000 words from recordings by 31 teenagers N/A Spontaneous conversational language of 13–17-year-old London teens captured via walkman devices, transcribed and POS-tagged for sociolinguistic and discourse analyses. Stenström et al., 2002 (COLT project)
Dependency Dialogue Act Corpus English Text (multi-party dialogues) Text transcripts with dialogue-act annotations (Dependency Dialogue Acts framework) Classroom discussions, board games, and online game chat (multi-genre) Human-Human multi-party interactions 33 dialogues, over 9,000 utterance units N/A (not specified separately) A dense annotation of multi-party conversational data across four genres—physics and engineering classroom discussions, board game interactions, and online game chat—using the Dependency Dialogue Acts framework, with double annotation and adjudication for high consistency. Cai et al., 2025
British National Corpus (BNC) English (British) Mixed (text-based spoken and written)—not dialogue per se Text (written samples, transcribed speech) Multiple domains (e.g., newspapers, fiction, conversations, academic, letters) Mixed participants (various genres of text and spontaneous spoken contributions) ~100 million words total; ~10 million words are spoken (from various types, including conversation) N/A A large-scale balanced corpus of late-20th-century British English, encompassing both written texts and transcribed spoken data (including some conversation), intended for general-purpose linguistic research but not focused on dialogue corpus specifically. Leech et al., 1990s
COLT – The Bergen Corpus of London Teenage Language English Speech Audio, Text Teenage casual talk (London) Human-Human ~500 K words (≈ half a million words) Varies (3 to 39 turns per conversation) Spontaneous conversations recorded by teenage recruits (aged 13–17) using Walkman, then orthographically transcribed, edited, and POS-tagged for linguistic research Stenström et al., 2002
Idiap Wolf Corpus English Multimodal (audio-visual) Audio, Video Competitive role-playing game (Werewolf-style group interaction) Human-Human (multi-party) Undisclosed exact size (volunteers in role-playing sessions) Varies (triadic or multi-party conversations in sessions) Natural conversational data of volunteers engaged in a competitive role-playing game, captured in an audio-visual corpus to explore group behavior Hung & Chittaranjan, 2010
Teams Corpus English Speech Audio, Text, Video, Questionnaire data Cooperative board-game conversation Human-Human (multi-party, 3–4 participants) Over 47 hours of recordings from 62 teams (213 participants) Varies per session (game-based multi-party dialogue) Audio, video, aligned transcripts, and questionnaire data collected from teams playing the cooperative board game Forbidden Island™, designed to study acoustic-prosodic and lexical entrainment in multi-party spoken dialogues Litman et al., 2016
Critical Role Dungeons and Dragons Dataset (CRD3) English Text Text Open-ended role-playing game dialogue (Dungeons & Dragons) Human-Human (multi-party, fixed group of players and a Dungeon Master) 159 episodes; 398,682 turns High (varies per episode; dataset spans full gameplay episodes) Transcribed unscripted live-streamed Dungeons & Dragons sessions featuring storytelling through collaborative dialogue; includes abstractive summaries mined from Fandom wiki Rameshkumar & Bailey, 2020
Michigan Corpus of Academic Spoken English (MICASE) English (American English) Speech Audio, Text Academic spoken events (lectures, seminars, meetings, advising, study groups) Human-Human ~1.8 million words (~200 hours across 152 speech events) Varies by event (unspecified average) Spoken academic interactions recorded at the University of Michigan across diverse academic contexts and departments, transcribed and annotated for linguistic study Simpson-Vlach & Leicher, 2006
Canal9 Political Debate Corpus English Multimodal (Speech + Video) Audio, Video, Text annotations Political debates (public broadcast debates) Human-Human (multi-party + moderator) 70 debates; ≈43 hours of recordings Varies by debate (multi-party structure) Public political debates annotated richly for social interaction features—including speaker turns, agreement/disagreement, roles, shot segmentation, and speaker identity—recorded in broadcast studio settings Vinciarelli et al., 2009
Interview English Text Text News interview transcripts (media dialog) Human-Human ≈ 105K conversations Varies (not specified; multi-turn interviews) Transcribed news interview dialogues gathered from media transcripts, annotated with speaker roles for each turn to support conversational modeling Majumder et al., 2020
MediaSum English Text Text Media interviews from NPR and CNN Human-Human ≈ 463.6K transcripts Varies per interview (not specified) Transcribed interviews from radio (NPR) and TV (CNN) with associated summaries or topic descriptions, making it a large-scale dataset for dialogue summarization Zhu et al., 2021
Corpus of American Soap Operas (SOAP) English Text (script transcripts) Text Soap opera scripts (American television) Human-Human (scripted dialogues) ~100 million words from over 22,000 transcripts Varies per episode (not specified) A vast compilation of transcripts from ten popular American soap operas (early 2000s), offering rich examples of everyday-styled, multi-party scripted dialogue for linguistic study
Serial Speakers English Multimodal (Speech + Video) Audio (speech turns), Text (encrypted turns via subtitles), Video (shots) TV serials (Breaking Bad, Game of Thrones, House of Cards) Human-Human (multi-party dialogues in TV series) 155 episodes (exact word/turn counts not specified) Varies per episode (multi-party scripted dialogues) Annotated dataset of episodes from three popular American TV serials with speech-turn boundaries, speaker labels, scene and shot boundaries, recurring shots, and interacting speaker annotations; text content encrypted but recoverable via users’ own subtitle files Bost et al., 2020
MEISD English Text, Speech, Vision (multimodal) Text, Audio, Video Multiple domains (TV-series dialogues) Human-Human (multi-party) 1,000 dialogues (from 10 TV series) Varies (multi-party dialogues; average not specified) A balanced multimodal dialogue dataset annotated with multiple emotions, emotion intensities, and sentiment per utterance, collected from ten popular TV shows across genres, with textual, audio, and visual modalities for emotion and sentiment analysis. Firdaus et al., 2020 (COLING)
NPS Chat Corpus English Text Text (chat logs annotated) Online chat / Internet-mediated communication Human-Human (chat) Not specified Not specified A chat corpus annotated with lexical (POS), syntactic, and discourse labels (chat dialog-act), intended to support statistical NLP applications like author profiling and entity identification. Forsyth & Martell, 2007
Molweni English Text Text (chat logs with questions and annotations) Technical support chats (Ubuntu IRC) Human-Human (multi-party) 10,000 dialogues, 88,303 utterances ~8.82 A multiparty dialogue-based MRC dataset with discourse dependency annotations (modified SDRT) and both answerable and unanswerable questions, derived from Ubuntu IRC logs. Li et al., 2020
Pushshift Reddit Dataset English Text Text (Reddit submissions and comments) Open-domain social media (Reddit) Human-Human (multi-participant threads) ~651M submissions, ~5.6B comments (2005–2019) Varies (thread-level discussions; average not specified) A large, continuously updated repository of Reddit data—historical submissions and comments—provided via dumps and an API for research, archiving, and social media analysis. Baumgartner et al., 2020
Reddit Domestic Abuse Dataset English Text Text (Reddit posts and comments) Domestic abuse discussions on social media Human-Human (submissions and responses) 1,336 abuse posts; 17,020 non-abuse posts Varies (thread-level posts; average not specified) A classification dataset of Reddit submissions labeled as abuse (e.g., “domestic-violence”, “survivors-of-abuse”) versus non-abuse (e.g., “advice”, “anger”, “casual-conversation”) to support detection of domestic abuse discourse online. Schrading et al., 2015 (EMNLP)
ISL Meeting Speech Part 1 (ISL-MC1) English Speech (audio recordings of meetings) Audio (multi-channel WAV files); Transcripts (orthographic text) Meeting domain (natural and artificial meetings across various scenarios) Human-Human (multi-participant meetings) 18 meetings, ~10 hours of speech (105 audio files) Varies—average meeting duration ~34 minutes; participants ~5 per meeting Multi-channel microphone recordings of real and staged meetings collected at CMU (2000–2001), with orthographic transcriptions, speaker turn timestamps, and annotations of spontaneous speech phenomena and disfluencies. Burger et al., 2002 (ICSLP)
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse Parsing in Conversations Hindi + English (code-mixed: Hinglish) Multimodal Audio, Text (transcriptions) Multiple customer-support domains (e-commerce, pharmaceutical, stock broker applications, e-marketplace, education) Human-Human (two-party call-center dialogues) 799 dialogues, 8,811 utterances, ~79,867 words ~11.03 utterances per dialogue A real-world, code-mixed (Hindi/English) multimodal corpus of customer call-center interactions across multiple domains, annotated at the span level with nine discourse relations, forming directed discourse graphs—reflecting genuine noisy ASR and diarization conditions. Shukla et al., 2025 (Findings ACL)
KwaiChat Multiple (multilingual: 4 languages) Multimodal (video-driven dialogue) Video, text dialogue content (comments, replies), metadata (domains, topics) Multimedia discussions: video-based interactions around shared videos Human-Human (multi-participant dialogues via video comments/replies) 93,209 videos, 246,080 dialogues N/A A massive dataset of human-to-human, video-driven multicultural multi-participant dialogues collected via a video-sharing platform, annotated across diverse dialogue types, domains, languages, and topics—designed to support multilingual dialogue generation over rich video context. Shi et al., 2025
MLDR English Multimodal (text and image) Text utterances and images (interleaved), with multi-granularity semantic annotations and query-fragment pairs Fine-grained fragment retrieval in multi-modal long-form dialogues; open domain (daily life, work and technology, health and emotion, consumption, mobility and travel, etc.) Human-System (synthetically generated long-form dialogues from short dialogue sources using Qwen3-235B) Dialogues averaging 25.45 turns each, covering 3 topics per dialogue; also includes a real-world WeChat-based test set of 580 dialogue samples (avg. 75.38 turns) with 1,250 query-dialogue pairs 25.45 (MLDR); 75.38 (WeChat test set) MLDR (Multi-modal Long-form Dialogue Retrieval) is the longest-turn multi-modal dialogue retrieval dataset to date, constructed by combining and extending short multi-modal dialogues into long-form, multi-topic conversations averaging 25.45 turns across three distinct topics. It supports fine-grained fragment retrieval tasks with multi-granularity annotations and diverse query types (multimodal, utterance-only, image-only, and negative samples), and is complemented by a real-world WeChat-based test set of 580 dialogues (avg. 75.38 turns) with 1,250 annotated query-dialogue pairs. Bi et al. 2026
Sympatheia-18k English Speech Synthetic speech audio (TTS-generated), text query–response pairs, valence–arousal (VA) metadata Emotion-conditioned empathetic spoken dialogue (open domain) Human-System ~18K spoken query–response pairs: ~12K emotional split (~1K per emotion across 12 emotions) and ~6K neutral split (500 neutral queries × 12 emotion-conditioned responses) Sympatheia-18k is a synthetic emotion-conditioned spoken dialogue corpus comprising approximately 18,000 speech query–response pairs anchored to 12 discrete emotions represented as continuous valence–arousal coordinates. It contains an Emotional split pairing affect-rich user queries with emotion-appropriate responses, and a Neutral split pairing emotionally neutral queries with 12 differently emotion-conditioned responses to support explicit affect-control training. Dindar et al. 2026
BEA-Dialogue+ Hungarian Speech Audio, Transcripts Conversational ASR / Dialogue transcription Human-Human 200 hours total (train: 183.41h, dev: 7.83h, eval: 8.70h); 27,484 segments; 1,728,495 words ~10.4 utterances per 30-second segment (train) BEA-Dialogue+ is a 200-hour conversational Hungarian speech corpus derived from BEA database recordings, expanding the earlier BEA-Dialogue corpus (85 hours) by relaxing speaker-disjointness constraints for experimenters and dialogue partners while preserving full separation of primary speakers. It provides transcribed multi-speaker natural conversations segmented into ~30-second units, benchmarked for dialogue ASR using Whisper and FastConformer models with SOT-based fine-tuning. Gedeon et al. 2026
AppTek Call-Center Dialogues English (14 accents: Australian, Canadian, Chinese, British, Scottish, Welsh, Irish, Indian, Mexican, Singaporean, African American Vernacular, General US American, Southern US American, South African) Speech Audio (16 kHz, 16-bit PCM WAV, split-channel), Verbatim transcripts Call-centre / customer service (16 service-oriented scenarios: agriculture, aviation, banking, delivery, energy, entertainment, finance, food, health, hospitality, insurance, real estate, retail, technology, telecommunication, travel) Human-Human (role-played agent–customer pairs) 128.6 hours of speech, 1,746 single-channel recordings, 156 speakers, ~8–11 hours per accent ~10.4 minutes per session (session length range: 5–15 minutes) A long-form English ASR evaluation corpus of spontaneous, role-played call-centre agent–customer conversations spanning 14 English accents (covering varieties from Australia, Canada, China, UK, Ireland, India, Mexico, Singapore, the US, and South Africa) and 16 service-oriented domains. Recordings were made via VoIP, manually transcribed verbatim by professional annotators with multi-stage quality assurance, and released under CC BY-SA 4.0. Beck et al. 2026
SuSuInterActs Mandarin Chinese Multimodal (Speech, Full-body Motion, Facial Expression, Text) Speech audio, full-body motion capture (6D rotation, 63 joints at 20 FPS), facial expression blendshapes (51-dim ARKit), dialogue text with behavior annotations Expressive interactive dialogue / role-playing conversational agent (single virtual character) Human-System (professional actors performing scripted multi-turn dialogues as a single virtual character, SuSu) 21,133 clips, 36.9 hours total; 2,656,484 motion frames; 12,367 samples with facial blendshape data 6.3 s average duration per sample; 18.7 Chinese characters per utterance SuSuInterActs is a multimodal Chinese dialogue corpus captured via optical motion capture, featuring 21K clips (37 hours) of synchronized speech audio, full-body motion (including hands), and facial expressions for a single virtual character (SuSu). Each utterance is annotated with facial expression and body action labels, supporting multi-turn, role-conditioned conversational motion generation research. Jin et al. 2026
Portal Dialogue Corpus English Multimodal (speech/audio, video, game state data) Audio recordings, screen/video recordings, manually-corrected transcripts, game engine demo files (player positions, orientations, object locations/velocities), dialogue act annotations, task/subtask completion timestamps Collaborative puzzle-solving in a 3D cooperative video game (Portal 2) Human-Human (18 pairs, 36 participants) 11.5 hours of gameplay, 24.5K total utterances, 109K words 1,365 utterances per session (average) The Portal Dialogue Corpus is a multimodal corpus of spoken human dialogue collected from 18 pairs of players (36 participants) playing the cooperative mode of the video game Portal 2. It comprises 11.5 hours of gameplay featuring 24.5K utterances, and includes player audio, screen video recordings, game state data, manually-corrected transcripts, and multi-layer dialogue act annotations, supporting study of complex situated linguistic phenomena such as spatial reference, clarification and repair, and ad-hoc convention formation. Tomlin et al. 2025
MDC-R+SDRT (merged corpus) English Text Text transcripts with reference/ambiguity annotations and SDRT discourse structure annotations (including clarification questions) Collaborative Minecraft building task (instruction-giver / instruction-follower) Human-Human 101 dialogues, 3,343 utterances, 29,174 tokens, 7,600 markables; 182 clarification questions and 218 confirmation questions in the subset A merged corpus combining two existing annotations of the Minecraft Dialogue Corpus (MDC-R and MSDC) into a single MMAX-format resource, providing aligned reference/referential-ambiguity annotations and SDRT discourse structure (including clarification and confirmation questions) over 101 task-oriented collaborative building dialogues. The corpus is intended to support research on the relationship between referential ambiguity and clarification requests. Madge et al. 2025
DocTalk English Text Synthesized multi-turn dialogues (user utterances LLM-generated; assistant utterances derived from Wikipedia text) Multi-topic information-seeking (Wikipedia-grounded) Human-System (synthetic) 730,707 conversations; ~8 billion tokens; mean 82.2 turns per conversation 82.2 DocTalk is a large-scale, synthetically constructed multi-turn pre-training dialogue corpus derived from English Wikipedia articles via a graph-based pipeline. Each conversation spans multiple related Wikipedia documents and features topical shifts, with assistant utterances drawn directly from Wikipedia text and user utterances generated by an LLM, yielding over 730k long multi-topic information-seeking dialogues. Lee et al. 2025
SITT Dataset Polish, English Text Text (dialogues with expert multi-label annotations of social influence categories and techniques) Social influence and manipulation detection in conversational text Human-Human (sourced from existing datasets and GPT-4o-generated dialogues) 746 dialogues, annotated with 58 techniques across 9 categories; 2,177 expert annotation assignments 6.46 The SITT Dataset is a 746-dialogue corpus annotated by 11 experts with 58 fine-grained social influence techniques organized into 9 categories (the Social Influence Technique Taxonomy, SITT). Dialogues were sourced from the MentalManip and CToMPersu datasets as well as GPT-4o-generated examples, originally annotated in Polish and translated into English, and used to benchmark LLMs on hierarchical multi-label social influence detection. Mieleszczenko-Kowszewicz et al. 2025
French OSCE Dialogue Dataset French Speech (audio recordings) and Text (automatic transcripts) Audio (WAV), automatic transcripts (ASR + diarization), manually corrected transcripts (subset), OSCE station sheets (physician, patient, evaluator), dialogue annotations (specialty, consultation type, objectives) Medical OSCE (Objective Structured Clinical Examination) training: doctor-patient simulations covering history-taking, diagnosis, breaking bad news, patient education, and more, across 12+ medical specialties Human-Human (sixth-year medical students role-playing physician, patient, and evaluator) 240 recorded dialogues (30 hours audio) across 23 OSCE stations; 192 OSCE station sheets; additionally 792 LLM-generated dialogues (1.22M words) across 11 stations A French corpus of 240 recorded and transcribed OSCE (Objective Structured Clinical Examination) training dialogues, collected from 99 sixth-year medical students across 23 clinical stations, totalling 30 hours of audio. The dataset also includes 192 structured OSCE station sheets and 792 synthetically generated dialogues produced by a controllable LLM-based pipeline, providing a resource for developing and evaluating virtual patient systems for French medical education. Bonzi et al. 2026
PeerMathDial English Speech (audio/video recorded, manually transcribed) Transcripts (manually reviewed ASR transcripts), dialogue act annotations, student self-report survey data Peer collaborative math problem solving (middle school, small-group) Multi-party human (student-student with occasional teacher intervention) 55 dialogues, 6,406 turns, 27 students 116.5 PeerMathDial is the first dataset of peer Collaborative Problem Solving (CPS) dialogues collected from authentic middle-school mathematics classrooms (grades 6–8), containing 55 small-group sessions from 27 students totalling 6,406 turns. The corpus is annotated with a corpus-grounded, LLM-assisted dialogue act taxonomy covering six functional dimensions of collaborative interaction, and is complemented by student self-report surveys on confidence, collaboration, and leadership. Yue et al. 2026
HEALTHDIAL Multilingual (Arabic, Chinese, English, Spanish) Speech and Text Audio recordings (user speech, WAV), ASR transcriptions, human post-edited transcriptions, LLM-generated system responses, machine-generated system speech (TTS), knowledge snippet annotations, speaker demographic and sociolinguistic metadata Health information seeking (knowledge-grounded, RAG-based spoken dialogue) Human-System 6,000 dialogues (1,500 per language), 41,988 dialogue turns, 163 hours of user speech, 208 hours of machine-generated system speech, 12,045 unique WHO knowledge snippets ~7 turns per dialogue (41,988 turns / 6,000 dialogues) HEALTHDIAL is a large-scale, multilingual, multi-parallel spoken dialogue dataset comprising 6,000 information-seeking dialogues across Arabic, Chinese, English, and Spanish, grounded in WHO health content. User utterances are recorded by native speakers representing diverse dialects and language varieties, with each speaker annotated for demographic and sociolinguistic variables, supporting benchmarking of RAG-based spoken dialogue systems. Hu et al. 2026
CPCD (Chinese Psychological Counseling Dataset) Chinese (Mandarin) Text Synthetic multi-session counseling dialogues, student profiles, temporal stress event graphs, session memory summaries Campus psychological counseling; long-horizon multi-session mental health support for college students Human-System (student agent and counselor agent simulation) 100 student profiles, 90,000 counseling dialogue units, ~11.45 million Chinese characters CPCD is a synthetic Chinese long-horizon dialogue dataset for college psychological counseling, constructed using the Psy-Chronicle framework. It contains 100 student profiles, 90,000 counseling dialogue units (~11.45M Chinese characters) spanning semester-length stress event trajectories, with explicit correspondences among student profiles, temporal stress event graphs, cross-session counseling dialogues, and structured memory summaries. To the authors’ knowledge, it is the first publicly available long-horizon dialogue dataset for Chinese college psychological counseling, accompanied by CPCD-Bench for evaluating session-level response, long-horizon memory recall, and temporal-causal reasoning. Gou et al. 2026
IndicMedDialog Multilingual (English, Assamese, Bengali, Gujarati, Hindi, Marathi, Punjabi, Tamil, Telugu, Urdu) Text Text (synthetic and template-based multi-turn medical dialogues with disease labels and optional patient pre-context) Medical consultation / differential diagnosis (symptom elicitation and diagnosis across 12 disease categories spanning 8 organ systems) Human-System (simulated physician–patient) 2,980 parallel multi-turn dialogues (English), yielding 29,800 language-specific dialogue instances across 10 languages; average 5.7 turns per dialogue (combined MDDial + Synthetic split) 5.7 (MD+SYN combined); 4.9 (MDDial base); 6.6 (Synthetic) IndicMedDialog is the first parallel multi-turn medical dialogue dataset spanning English and nine Indic languages (Assamese, Bengali, Gujarati, Hindi, Marathi, Punjabi, Tamil, Telugu, and Urdu). It extends the MDDial corpus with LLM-generated synthetic consultations covering 12 disease categories, translated using TranslateGemma, verified by native speakers, and refined via a script-aware post-processing pipeline; each dialogue includes optional patient pre-context (age, gender, allergies, pre-existing conditions) to support personalised symptom elicitation. Nigam et al. 2026
BSDD (Biomedical Streaming Dialogue Dataset) English Text Text (LLM-simulated multi-role research discussion transcripts with proactive intervention point annotations) Biomedical research team discussions (cancer, Alzheimer’s disease, sepsis); proactive intervention detection and generation Human-System (multi-role LLM-simulated dialogues among Pharmacologist, Medicinal Chemist, Bioinformatician, and Clinical Physician personas, grounded in PubMed literature) 3,206 dialogues; 14,590 sampled rounds (train: 2,726 dialogues / 13,630 rounds; validation: 240 dialogues / 480 rounds; test: 240 dialogues / 480 rounds) 20 rounds per dialogue BSDD is a benchmark of LLM-simulated multi-role biomedical research discussion dialogues grounded in PubMed articles (2024), annotated with proactive intervention points (positive/unlabeled/negative labels). It is designed to train and evaluate systems that determine when and how a proactive AI assistant should intervene in streaming scientific team meetings. Wu et al. 2026
Syn-TurnTurk Turkish Text Synthetic text dialogues with turn-taking annotations (floor transfer offsets, overlaps, silences) Turn-taking prediction in spoken dialogue Human-Human (synthetic/LLM-generated two-person dialogues) 1,625 dialogues, 12,560 speaker turns (turn changes); 37,680 labeled samples (12,560 positive, 25,120 negative) 7.73 Syn-TurnTurk is a synthetic Turkish dialogue dataset generated using five Qwen large language models across 79 diverse topics, designed to support turn-taking prediction research. Dialogues incorporate human-like speech features such as overlaps, strategic silences, and everyday interjections, with annotated floor transfer offsets and turn boundary labels. Bayrak et al. 2026
CPB-Bench English, Chinese (Bilingual) Text Text (dialogue transcripts annotated with challenging patient behavior labels) Medical consultation / clinical dialogue safety evaluation Human-Human (real and role-played doctor–patient consultations) 692 multi-turn dialogues (91 information contradiction, 65 factual inaccuracy, 275 self-diagnosis, 261 care resistance instances); 352 true negatives also included 26.43–109.21 (varies by source dataset) CPB-Bench (Challenging Patient Behaviors Benchmark) is a bilingual (English and Chinese) benchmark of 692 multi-turn medical consultation dialogues annotated with four clinically grounded categories of challenging patient behaviors: information contradiction, factual inaccuracy, self-diagnosis, and care resistance. It is built by annotating patient utterances from four existing medical dialogue datasets (SIMORD, MediTOD, MedDG, IMCS-21) using GPT-4o filtering followed by human adjudication, and is designed to evaluate LLM responses under realistic, imperfect patient inputs. Li et al. 2026
Fin-Vault English Text Text (multi-turn user-bot dialogues annotated with politeness categories and demographic region labels) Financial advisory (banking, credit/debit card management, insurance, stock investments, loans, taxation, budgeting, trading, personal finance) Human-System 1,417 annotated multi-turn dialogues, 4,006+ utterances, vocabulary size 3,398 ≥3 turns per conversation; avg. ~234.88 queries per conversation (likely a dataset-level stat); avg. ~145.33 words per conversation Fin-Vault is a domain-specific multi-turn financial conversational dataset comprising 1,417 annotated user-bot dialogues sourced from online financial forums (e.g., Reddit r/personalfinance, Bogleheads). Dialogues span 10 financial domains including banking, credit/debit cards, insurance, and investments, and are annotated with politeness categories (Polite, Neutral, Impolite) and demographic region labels. Das et al. 2025
DISPLACE-M Hindi (with code-switching to Indian English and regional dialects: Haryanvi, Bhojpuri, Magahi) Speech Audio recordings, manual transcripts (verbatim, Devanagari script), speaker diarization annotations (RTTM), topic labels, clinical dialogue summaries Frontline healthcare / medical consultations (community health worker – care seeker interactions) Human-Human (non-physician health workers and healthcare seekers) ~55 hours total annotated audio; 40 hours development set (25 h diarization/ASR dev + 15 h topic/summarization dev) + 15 hours blind evaluation DISPLACE-M is an annotated corpus of spontaneous, real-world medical conversations between frontline community health workers (ASHA/Anganwadi workers) and healthcare seekers recorded in rural and semi-urban India. The dataset features noisy, overlapping, code-mixed Hindi speech across ~55 hours and supports four tasks: speaker diarization, ASR, topic identification, and dialogue summarization. Dhanya et al. 2026
Multifaceted Skill-of-Mind English Text Text (multi-turn dialogues annotated with conversational skill explanations and skill labels, derived from 12 source datasets via GPT-4 annotation) Multi-domain social dialogue including chitchat, counseling, task-oriented, long-term conversation, negotiation, and persuasion Human-Human 99,997 dialogues (≈100K); 38+ conversational skill categories; 109,591 total skill-of-mind annotations Multifaceted Skill-of-Mind is a multi-turn conversation dataset of ~100K dialogues annotated with skill-of-mind labels: a free-text rationale (explanation) and one or more conversational skills selected from a hierarchical taxonomy of 38+ skills (covering Interpersonal, Memory & Knowledge Management, Cognitive & Problem-Solving, Communication & Listening, and Task-Oriented categories). The dataset is derived from 12 existing source dialogue datasets spanning diverse social contexts (demographics, persona, rules of thumb) and interactive scenarios (long-term, counseling, task-oriented), with annotations generated by GPT-4 using a perspective-taking prompting approach. Lee et al. 2024
AVCC (Audio-Visual Conversation Corpus) English Multimodal (audio and video) Audio (monaural), Video (single-camera face tracks), Voice activity annotations Multiparty turn-taking / conversational dynamics Human-Human (multi-party: 2- and 3-speaker settings) ~31 hours (30h 52m) of video; 17h 31m two-speaker, 13h 21m three-speaker; 22h 28m training, 8h 24m validation The Audio-Visual Conversation Corpus (AVCC) is a ~31-hour dataset of unedited, single-camera, static third-party-perspective multiparty conversations collected from YouTube and Twitch livestreams, covering both two- and three-speaker settings. It is specifically designed for causal turn-taking modeling, preserving natural conversational flow (mutual silences, overlaps, hesitations) without editorial jump cuts, with manually refined speaker voice activity annotations. Qi et al. 2026
M³C (Multimodal Multi-Session Multi-Party Conversation) English Multimodal (text, image, and audio) Machine-generated text dialogues, image captions (from COCO), audio captions (from AudioCaps and Clotho), multimodal memory summaries Open-domain conversation with simultaneous visual and auditory inputs in shared multi-party, multi-session settings Human-System (model-to-model; four speakers per episode, multi-party per session) 54K episodes (34K train, 8K validation, 12K test); 16K sessions; 2.5M turns; 24K images; 73K audio clips M³C is a machine-generated multimodal conversation dataset featuring four speakers across three consecutive multi-party sessions, where all participants simultaneously experience synchronized visual (images) and auditory (audio) inputs in a shared spatial and temporal context. It supports research on open-domain, multi-session, multi-party dialogue with both “eyes and ears” modalities, and includes multimodal memory summaries linking cross-session references. Jang et al. 2025
MSP-Conversation English Speech Audio, time-continuous emotional annotations (valence, arousal, dominance), speaker diarizations Speech emotion recognition; naturalistic multi-party conversational speech from podcasts Human-Human (multi-party, sourced from publicly available podcasts) 310 conversations, 908 conversation parts, 77 hours 26 minutes; >450 speakers; 12,555 speaking turns overlapping with MSP-Podcast MSP-Conversation is a large-scale naturalistic speech corpus of 310 multi-party podcast conversations (77+ hours) annotated with time-continuous emotional traces for valence, arousal, and dominance, collected using a joystick-based annotation tool (CARMA) with at least six raters per segment. The corpus includes detailed manual speaker diarizations and overlaps with a subset of the MSP-Podcast corpus to enable direct comparison between in-context (time-continuous) and out-of-context (utterance-level) annotation methods. Martinez-Lucas et al. 2026
TTS Conversational Corpus with Interjections English Speech Audio recordings Conversational speech for customer-care voice agents; includes dialog act tags and interjections Human (single professional voice actor) ~7 hours of audio (~4K sentences); ~6 hours conversational, ~1 hour non-conversational expressive material A single-speaker corpus of US English conversational speech recorded by a professional voice actor, designed for training conversational TTS systems. The corpus features two-part scripted dialogues annotated with ten dialog act tags and seven interjection types, targeting customer-care interaction scenarios. Fernandez et al. 2022
STAR Pre-training Corpus English Text Text (natural language utterances paired with SQL queries, synthesized context-dependent multi-turn conversations) Context-dependent text-to-SQL parsing (cross-domain, multi-turn) Human-System ~480K context-dependent text-to-SQL conversations ~5 (based on example; not explicitly stated as an average) A large-scale synthesized pre-training corpus of ~480K high-quality context-dependent text-to-SQL conversations, constructed by combining single-turn question-SQL pairs from Spider, SParC, and CoSQL with BART-based utterance generation and ~100 manually crafted follow-up grammar templates. Designed to support pre-training of tabular language models for multi-turn text-to-SQL parsing across 200 databases and 138 domains. Cai et al. 2022
ECC (English Conversation Corpus) English Speech Audio, Transcripts, speaker labels, sentence boundaries Conversational speech / daily conversations for English language learning Human-Human 24 hours of speech, 66 conversational videos, ~962 conversations (training set), 28,837 sentences (training set) 30.4 sentences per conversation The English Conversation Corpus (ECC) consists of 24 hours of speech collected from 66 conversational YouTube videos originally produced for second-language English learning. It is annotated with transcriptions, speaker labels, and sentence boundaries, covering conversations performed by 2–9 speakers with an average of 30.4 sentences per conversation. Li et al. 2022
Contrack English Text Text (chat transcripts) with entity reference annotations (people and locations), including coreference, grammatical gender, plurality, and group membership labels Open-domain social conversation; entity context tracking (slot tagging, coreference resolution, plural mention resolution, entity linking) Human-Human (crowdworker pairs, or single crowdworker playing both roles) 7,245 conversations, 85,538 turns; avg. 5.8 entities and 15.2 references per conversation 11.8 Contrack is a large-scale human-human casual conversation corpus for entity-centric context tracking, containing 7,245 scenario-seeded conversations annotated with people and location entity references, including coreference links, grammatical gender, plurality, and group membership information. It is designed to support unified modeling of subtasks such as slot tagging, coreference resolution, plural mention resolution, and entity linking. Rückert et al. 2022
OpenAssistant Conversations (OASST1) Multilingual (35 languages, predominantly English and Spanish) Text Text messages, quality ratings (Likert-scale and binary labels), preference rankings, conversation trees Open-domain assistant-style conversations; LLM alignment (SFT and RLHF) Human-Human (crowd-sourced prompter and assistant roles, with optional synthetic messages) 161,443 messages (91,829 prompter, 69,614 assistant) across 66,497 conversation trees (10,968 complete); 461,292 quality ratings; 8,576 synthetic messages OpenAssistant Conversations (OASST1) is a large-scale, human-generated and human-annotated assistant-style conversation corpus collected via worldwide crowd-sourcing from over 13,500 volunteers. It comprises 161,443 messages in 35 languages organized into conversation trees, annotated with 461,292 quality ratings and preference rankings, intended to support research on LLM alignment via supervised fine-tuning and reinforcement learning from human feedback. Köpf et al. 2023
DinG (Dialogues in Games) French Speech (recordings with manual transcriptions) Audio recordings, manual transcriptions with timecode alignment, question-type annotations Multi-party board game interaction (Catan); bargaining and resource negotiation Multi-party human (3–4 players per game) 10 games, 23,575 turns, 2,528 questions; ~702 minutes of recorded dialogue 2,357.5 turns per game (range: 476–3,572) DinG is a corpus of manual transcriptions of real-life, spontaneous, oral multi-party dialogues among French-speaking players of the board game Catan, comprising 10 recorded games (~702 minutes total, 23,575 speech turns). Transcriptions include timecode alignment, speaker disambiguation, overlap marking, and question-type annotations (yes/no, wh-, disjunctive, phatic, completion suggestion), distributed under CC BY-SA 4.0. Boritchev et al. 2022
StreamDial Chinese (with translated English, French, and Korean versions in preparation) Text Synthesized multi-turn task-oriented dialogues with structured session quadruplets including user persona, agent persona, conversational blueprint, and dialogue history Vertical service domains: Automotive (vehicle consultation/sales), Restaurant (discovery/reservation), Hotel (search/booking) Human-System (simulated via multi-agent LLM synthesis grounded in real streaming media signals) 87,498 dialogue sessions; 1,497,320 turns 17.11 StreamDial is a large-scale, multi-domain task-oriented dialogue dataset synthesized from publicly available streaming media (live streams and short videos) using the STREAM framework. It covers Automotive, Restaurant, and Hotel service domains, with each session structured as a quadruplet ⟨user persona, agent persona, conversational blueprint, dialogue history⟩ capturing realistic service behaviors such as requirement mining, constraint conflicts, negotiation, and recovery. Xue et al. 2026
RealReasoning English Text Text (synthetic multi-turn task-oriented dialogues with manually annotated reasoning questions and ground-truth labels) Logical reasoning (math word reasoning and commonsense reasoning) grounded in realistic task-oriented scenarios Human-System (LLM-based user agent and assistant agent) 500 dialogues, 2,398 total turns 4.796 RealReasoning is a synthetically generated multi-turn task-oriented dialogue dataset designed to benchmark LLM logical reasoning in realistic scenarios. Each instance comprises a multi-turn dialogue (generated by LLM agents using a trilevel optimization framework) paired with a manually annotated reasoning question (either math word reasoning or commonsense reasoning) and a ground-truth label. Zhu et al. 2026
SLURP-TN Tunisian Arabic dialect Speech Audio recordings, transcripts, SLU annotations (intent, slot filling) Spoken Language Understanding; multi-domain (Emails, Weather, News, Books/Takeaway, Alarm, General) Human (read speech by native speakers) 4,165 utterances; ~5 hours total audio (train: 2h 46m / dev: 44m / test: 1h 3m); 2,677 train / 595 dev / 893 test segments SLURP-TN is a multi-domain spoken language understanding corpus for the Tunisian Arabic dialect, created by having 55 native speakers manually translate and record utterances from six domains of the English SLURP dataset. It provides audio in three acoustic conditions (clean, noisy, headphone) at 48 kHz, with SLU annotations (intent detection and slot filling), rich speaker metadata (gender, age, regional dialect), and extensive code-switching, making it suitable for SLU, ASR, speaker identification, and text-to-speech research. Elleuch et al. 2026
KoCC-TTS Korean Speech Audio, Text (transcripts) Task-oriented dialogue; Korean call-center TTS prosody evaluation Human-Human (manager–customer call center interactions) 50 utterances (high-quality human-curated samples) KoCC-TTS (Korean Call-Center TTS) is a curated evaluation dataset of 50 high-quality utterances drawn from authentic Korean call-center manager–customer conversations, designed to benchmark TTS systems on transcription robustness and conversational prosody in task-oriented Korean speech synthesis. Shin et al. 2026
MultiATIS and MultiSNIPS English Text Text (utterances with intent and slot annotations) Spoken language understanding — multi-intent detection and slot filling (task-oriented dialogue) Human-System MultiATIS: 19,760 utterances (18K train, 1K dev, 1K test); MultiSNIPS: 50,000 utterances (45K train, 2.5K dev, 2.5K test) MultiATIS and MultiSNIPS are two synthetic multi-intent spoken language understanding datasets constructed by concatenating single-intent utterances from ATIS and SNIPS using BERT’s next sentence prediction (NSP) head to ensure semantic coherence. Each utterance contains 1–3 intents (sampled with probabilities 0.3/0.5/0.2) and is annotated with intent labels and BIO-format slot tags, yielding more naturalistic multi-intent samples than prior random-concatenation datasets. Li et al. 2026
RealMem English Text Synthesized multi-session dialogues with structured memory points, schedules, and natural user queries Long-term project-oriented human–agent interaction across 11 scenarios (e.g., travel planning, fitness, financial planning, mental health support, code architecture design) Human-System (simulated via dual User Agent and Assistant Agent) 2,000+ cross-session dialogues; 14,028 total dialogue turns; 1,415 questions; 5,072 total memory items; avg. context length 269,190 tokens per user 6.8 turns per session (avg. 205 sessions per user) RealMem is a benchmark for evaluating LLM memory systems in long-term, project-oriented interactions, comprising over 2,000 cross-session dialogues across eleven realistic scenarios. It is constructed via a three-stage synthesis pipeline (Project Foundation Construction, Multi-Agent Dialogue Generation, and Memory and Schedule Management) and evaluates agents on four query types: Static Retrieval, Dynamic Updating, Proactive Alignment, and Temporal Reasoning. Bian et al. 2026
TOD-ProcBench Multilingual (English, Arabic, Chinese, French, German, Hindi, Spanish) Text Text: multi-turn dialogue transcripts, complex natural language condition-action instruction documents, instruction-violating synthetic responses Task-oriented dialogue / customer support (trip booking, banking, healthcare, e-commerce); instruction-following evaluation Human-Human (sourced from ABCD), with LLM-generated instruction-violating responses 55 instruction documents (one per user intent); 1,004 test conversations; 6,953 partial conversations (Task 1); 3,964 balanced compliant/non-compliant examples (Task 2); 3,310 partial conversations (Task 3) TOD-ProcBench is a multilingual benchmark derived from the ABCD dataset for evaluating LLMs’ ability to follow complex, fine-grained condition-action natural language instructions in multi-turn task-oriented dialogues. It provides 55 complex instruction documents across 7 languages and three instruction formats (Nested If-Then, Flattened If-Then, Flattened JSON), supporting three tasks: instruction retrieval and next-action prediction, compliance evaluation, and compliant response generation. Ghazarian et al. 2025
TACT (TOD-And-Chitchat Transition) English Text Text (multi-turn dialogues with mode transition annotations and intent labels) Mixed task-oriented and open-domain chitchat dialogue with mode transitions (e.g., restaurant/train booking, smart home, etc.) Human-System (synthetically constructed via LLM augmentation of MultiWOZ 2.2 and SLURP) 9,936 dialogues (two variants: TACTMultiWOZ with 7,199 dialogues and TACTSLURP with approx. 2,737 dialogues implied; 50 intents; 12 unique dialogue flow patterns) 16.42 TACT is a dataset for transition-aware dialogue modeling that integrates task-oriented dialogue (TOD) and open-domain chitchat within single sessions. Built on MultiWOZ 2.2 and SLURP, it features structurally diverse mode-transition flows (TCT, CTC, TCTCT, etc.) with an average of ~2 mode switches and ~1 recovery per dialogue, supporting both user- and agent-driven transitions and the training of proactive conversational agents. Yoon et al. 2025
M-EDESConv & M-TESC English, Japanese Text, Speech Text dialogues with emotional validation labels (M-EDESConv); spoken dialogue transcripts with emotional validation labels (M-TESC) Emotional support conversation; emotional validation in dialogue Human-Human M-EDESConv: ~126,553 utterances (46,002 validating, 80,551 non-validating), ~120k total turns; M-TESC: 3,080 utterances (1,052 validating, 2,028 non-validating) M-EDESConv is a ~120k-utterance English–Japanese multilingual corpus derived from EmpatheticDialogues and ESConv, annotated for emotional validation phenomena via hybrid manual and automatic annotation. M-TESC is a multilingual (English–Japanese) spoken-dialogue test set derived from the TUT Emotional Storytelling Corpus, also annotated for validation timing, supporting research on validating response identification, validation timing detection, and validating response generation. Pang et al. 2026
ConversationGoT-120h English Speech Audio, Transcripts, Hierarchical behavior state annotations (communicative function labels, interaction behavior labels), evidence-based rationales Open-domain full-duplex conversational behavior modeling (turn-taking, backchannel, interruption, communicative function recognition) Human-Human (real subset from Candor); Human-System/Synthetic (synthetic subset with TTS voices) 120 hours of two-person dialogues (60 hours real, 60 hours synthetic); 720 synthetic samples averaging ~5 minutes each; 1-second resolution annotations ConversationGoT-120h is a 120-hour causal streaming benchmark dataset for second-level conversational behavior modeling in full-duplex spoken dialogue. Each one-second segment is annotated with a hierarchical conversational behavior state consisting of a high-level communicative function (Constative, Directive, Acknowledgment, Commissive), a low-level interaction behavior (Continuation, Turn-taking, Interruption, Backchannel, Silence), and an evidence-grounded rationale generated under strictly causal (past-only) conditions, supporting research on streaming behavior perception and interpretable reasoning in duplex dialogue systems. Zhou et al. 2026
EChat-200K English Speech Audio (synthesized speech queries and responses with paralinguistic labels including emotion, age, gender, and sound events) Empathetic spoken dialogue with paralinguistic cue recognition and response generation Human-System ~200K speech-to-speech conversations (~0.2K hours) EChat-200K is a speech-to-speech empathetic dialogue corpus containing approximately 200K conversations rich in paralinguistic information (emotion, age, gender, sound events), constructed using DeepSeek for text generation and CosyVoice2 for speech synthesis. It includes both single-label and multi-label empathetic data, with a subset incorporating real audio input queries to reduce overfitting to synthetic speech. Geng et al. 2025
GoT-Duplex Hybrid Corpus English Speech Two-channel audio, ASR transcripts, hierarchical speech act labels (high-level and low-level), per-second Graph-of-Thoughts rationale annotations Full-duplex conversational behavior detection and reasoning; covers turn-taking, backchannels, interruptions, and continuation in two-speaker dialogue Human-Human (real subset from Candor); Human-System/Synthetic (simulated two-speaker dialogues generated via GPT-4o + CosyVoice2 TTS) Synthetic: 28,000 clips, 192 hours, 37,100 rationale entries; Real (Candor subset): 118 hours A hybrid corpus for training and evaluating conversational behavior reasoning in full-duplex spoken dialogue systems, combining controllable synthetic two-speaker dialogues (28,000 clips, 192 hours, generated via GPT-4o and CosyVoice2 TTS) with a curated 118-hour subset of the real Candor corpus. Each one-second segment is annotated with hierarchical speech act labels (high-level: constative/directive/commissive/acknowledgment; low-level: turn-taking/interruption/backchannel/continuation) and human-validated Graph-of-Thoughts rationale text, totaling 37,100 rationale entries. Pan et al. 2025
Spoken DialogSum English Speech (synthesized), Text Synthetic multi-speaker audio, dialogue transcripts with disfluencies and backchannels, factual summaries, emotion-rich summaries, utterance-level emotion/pitch/speaking-rate labels, speaker age and gender labels Spoken dialogue summarization (factual and emotion-rich), paralinguistic attribute prediction (emotion, age, gender) Human-Human (simulated; TTS-synthesized multi-speaker dialogues) 13,460 dialogues, 251,575 utterances, ~160 hours of audio Spoken DialogSum is the first large-scale spoken dialogue corpus aligning synthetic multi-speaker conversational audio with both factual and emotion-rich summaries, plus utterance-level labels for speaker emotion, pitch, speaking rate, age, and gender. It is built by LLM-based style transfer and backchannel insertion of DialogSum scripts, followed by expressive TTS synthesis, yielding 13,460 emotion-diverse dialogues (~160 hours) suitable for spoken dialogue summarization and paralinguistic understanding research. Lu et al. 2025
MNSC (Multitask National Speech Corpus) Singlish (Singapore English, including code-switching with Mandarin, Malay, and Tamil) Speech Audio recordings, orthographic transcripts, synthesized QA pairs, human-annotated dialogue summaries and QA test sets, paralinguistic metadata (gender, accent) Automatic Speech Recognition (ASR), Spoken Question Answering (SQA), Spoken Dialogue Summarization (SDS), Paralinguistic Question Answering (PQA) Human-Human Approx. 10,000 hours of audio; sentence-level ASR train sets: ~2.3M and ~2.5M samples; dialogue-level train sets: up to ~104K samples per part; human-verified test sets: 3K–6K samples for ASR/PQA, 100 samples per subtask for SQA/SDS (800 total human-annotated QA/summarization test samples) MNSC is a standardized, multitask spoken Singlish corpus derived from Singapore’s National Speech Corpus (NSC), featuring standardized train/test splits and human-verified test sets across four tasks: Automatic Speech Recognition (ASR), Spoken Question Answering (SQA), Spoken Dialogue Summarization (SDS), and Paralinguistic Question Answering (PQA). It is the largest well-organized resource for Singlish-specific spoken language processing, covering monologue and dialogue speech with code-switching across English, Mandarin, Malay, and Tamil. Wang et al. 2025
J-CHAT Japanese Speech Audio, ASR transcripts with word-level alignment, speaker diarization labels (turn durations and speaker IDs) Open-domain spontaneous spoken dialogue (YouTube videos and podcasts) Human-Human 76,036 hours; 5,424,514 dialogues (YouTube: 11,017 hrs, 1,015,109 dialogues; Podcast: 65,019 hrs, 4,409,405 dialogues) 10.10 overall (YouTube: 7.58; Podcast: 10.68) J-CHAT (Japanese Corpus for Human-AI Talks) is a 76,000-hour open-source Japanese spoken dialogue corpus automatically constructed from YouTube and podcast audio using language identification, speaker diarization, background-music removal (Demucs), and ASR transcription. It provides train/valid/test splits, per-turn speaker and timing labels, and ASR transcripts with subword-level alignment, designed to support end-to-end spoken dialogue system development. Nakata et al. 2024
SaSLaW Japanese Multimodal (Speech, Audio, Video) Close-talking microphone speech, binaural ear-mounted microphone audio (hearing), head-mounted egocentric video, impulse responses, ambient noise recordings Spontaneous face-to-face dialogue in varied audio environments (noisy, moderate, quiet); designed for environment-adaptive text-to-speech synthesis Human-Human (two-person dyads; three male-male pairs and one female-female pair) 4 speaker pairs; approximately 30 minutes of recorded speech per pair on average; train/test splits of 299/49 utterances (spk01) and 443/64 utterances (spk06) reported for two analysed pairs 5–8 turns per conversation SaSLaW is a spontaneous Japanese dialogue speech corpus capturing synchronised first-person (egocentric) audio-visual recordings of what each speaker speaks (close-talking microphone), hears (binaural ear-mounted microphone), and sees (head-mounted camera) during face-to-face conversations conducted under varying real-world noise conditions. It also includes impulse responses and ambient-noise-only recordings to support reproducible evaluation of environment-adaptive TTS models. Take et al. 2024
MultiDialog English Multimodal (Speech, Video/Face, Text) Audio recordings, face video recordings, text transcripts, emotion annotations Open-domain face-to-face spoken dialogue (based on TopicalChat topics: fashion, politics, books, sports, general entertainment, music, science & technology, movies) Human-Human 8,733 dialogues, 187,859 utterances, ~340 hours 21.51 utterances/dialogue (approx. 11.0 turns/dialogue) MultiDialog is the first large-scale multimodal (audio, video, and text) spoken dialogue corpus, consisting of approximately 340 hours of parallel audio-visual recordings of 8,733 open-domain human-human conversations. Derived from the TopicalChat dataset, it features 12 speakers recorded with simultaneous speaker and listener video streams, per-utterance emotion annotations (7 categories), and is designed to support research in face-to-face dialogue systems, talking face synthesis, and emotion-conditioned multimodal generation. Park et al. 2024
PxCorpus (PxSLU + PxDialogue) French Speech Audio recordings, manual transcripts, semantic slot/intent annotations, dialogue act annotations Spoken drug prescription via goal-oriented human-system dialogue Human-System 1,981 recordings, 903 dialogue sessions, 3,675 dialogue turns, 22,440 tokens, 14,068 slot-label instances, ~262 minutes (4h) of speech 3.83 dialogue turns per session (3,675 turns / 959 sessions) PxCorpus is the first publicly available spoken drug prescription corpus in French, collected from 55 participants (physicians, medical experts, and non-experts) interacting with a smartphone-based spoken dialogue system for e-prescribing. It is distributed in two parts: PxSLU (for spoken language understanding, with slot and intent annotations in CoNLL format) and PxDialogue (utterances in full dialogic context with additional dialogue-level annotations), supporting development and evaluation of SLU and dialogue policy models. Kocabiyikoglu et al. 2023
CALLS Japanese Speech Audio recordings, dialogue text with emotion labels Empathetic spoken dialogue in customer center settings: complaint handling and attentive listening Human-Human (simulated; single female operator speaker recorded, customer voices not recorded) 3,272 operator utterances (6.5 hours of recorded speech), 3,312 customer utterances (text only); 820 complaint handling dialogue lines + 600 attentive listening dialogue lines 4–10 turns per dialogue (complaint handling); 4 turns (attentive listening) CALLS (Complaint handling and Attentive Listening Lines Speech) is a Japanese empathetic dialogue speech corpus covering simulated customer-center phone calls in two subsets: situation-oriented complaint handling and positive attentive listening. It features a single female speaker acting as an operator, with emotion-labelled utterances recorded at 48 kHz, designed to extend empathetic dialogue speech synthesis (EDSS) to polite and formal dialogue domains. Saito et al. 2023
Foodie (IterChat) English Text Text dialogues with annotated preference slots (StateGain and PreferenceExtraction labels) Food preference extraction in task-oriented dialogue Human-System 3,500 samples Foodie (IterChat) is a synthetically generated dataset for food-domain user preference extraction in task-oriented dialogues, constructed using the IterChat framework. Each sample pairs a history preference state with a single-turn dialogue, annotated with StateGain and PreferenceExtraction labels; GPT-4 was used to generate dialogues and annotations were validated by experienced human annotators. Wang et al. 2025
Disc3D English Multimodal (3D point clouds, RGB-D images, text) 3D scene scans (point clouds, RGB-D frames), automatically generated dialogue/QA text 3D scene understanding: scene/view/object captioning, visual grounding, and five object-centric QA tasks (object size, absolute distance, relative distance, object count, attribute recognition) Human-System (automated pipeline using MLLMs and LLMs; human revision for test split only) 2.08 million samples across 25K hybrid (real and synthetic) 3D scenes Disc3D is a large-scale, automatically curated 3D scene dialogue dataset comprising over 2 million multi-task dialogue samples across 25K hybrid real and synthetic 3D scenes. Generated via a fully automated pipeline centred on Discriminative Object Referring, it spans scene, view, and object captioning, visual grounding, and five object-centric QA tasks, with explicit resolution of viewpoint and object referring ambiguities. Wei et al. 2025
Safety Reasoning Multi-Turn Dialogue Dataset English Text Text (multi-turn dialogues with human-annotated malicious intent labels, severity levels, and model-generated Chain-of-Thought safety reasoning) Safety / adversarial multi-turn jailbreak detection Human-System 2,177 multi-turn dialogues A human-annotated dataset of 2,177 multi-turn dialogues derived from known LLM jailbreak attack strategies (e.g., ActorAttack, Chain of Attack), targeting GPT-4 series models. Each dialogue turn is labeled with malicious intent indicators, one or more of 37 predefined malicious categories across 7 high-risk domains, severity levels (0–10), and Claude 3.7 Sonnet-generated Chain-of-Thought safety reasoning, designed to train safety reasoning moderators against multi-turn adversarial attacks. Kuo et al. 2025
PersonaTAB Dialog Dataset English Speech Audio, Transcripts (ASR), timestamps, response type labels (turns, backchannels, interjections), laughter annotations, emotion/sentiment labels, Big Five personality labels Personality prediction from fully-duplex telephonic speech dialogues Human-Human 95 conversations, 190 speakers (subset of Fisher corpus, folder 000) A dialogue dataset derived from a subset of the Fisher telephonic speech corpus, preprocessed via an automatic pipeline to add word-level timestamps, laughter event labels, response-type annotations (turns, emotive/cognitive backchannels, interjections), emotion/sentiment labels, and Big Five personality trait labels assigned through LLM inference and human evaluation. Each conversation is a 12-minute two-channel phone call between two speakers. Inoue et al. 2025
PROASSIST English Multimodal (text dialogues aligned with egocentric video) Synthetic text dialogues with timestamped assistant and user utterances, derived from annotated egocentric videos Proactive task guidance across multiple procedural domains: cooking, object manipulation, assembly, and laboratory tasks Human-System (synthetic user–assistant dialogues generated via LLM pipeline) 30,135 dialogues spanning 478.7 hours of video (train/validation/test splits); sourced from six egocentric video datasets PROASSIST is a large-scale synthetic dialogue dataset for proactive task guidance, created by an automated LLM-based pipeline that synthesizes multi-round assistant–user dialogues from timestamped annotations of egocentric procedural videos. The dataset spans six source video collections (Ego4D-Goalstep, EpicKitchens, HoloAssist, Assembly101, EgoExoLearn, WTaG) and covers cooking, manipulation, assembly, and laboratory domains, with dialogues annotated for assistant intent, response type, and task progress summaries. Zhang et al. 2025
MemeCMD Mandarin Chinese Multimodal (text and image/meme) Text dialogues, meme images with MLLM-generated annotations Open-domain multi-turn conversation with contextually retrieved memes (news-based and role-based scenarios) Human-System (dual-agent LLM-generated dialogues) 34,758 dialogue turns total; Meme Library of 6,023 annotated meme images; dialogues in 6 sub-datasets (News-based and Role-based, each at 6, 12, and 18 turns) 6, 12, or 18 turns (depending on sub-dataset) MemeCMD is an automatically generated Chinese multi-turn dialogue dataset combining a MLLM-annotated meme library of 6,023 images with dialogues auto-generated by dual GPT-4 agents across diverse news-based and role-based scenarios. A retrieval framework with adaptive threshold decay ensures contextually appropriate and naturally spaced meme insertion, yielding 34,758 dialogue turns enriched with multimodal meme responses. Wang et al. 2025
Interaction Dialogue with Privacy Multilingual (English, Mandarin Chinese) Text Text (user queries with annotated privacy phrases and corresponding privacy information summaries) Privacy detection in real-name user interactions with large language models (open-domain and task-oriented dialogues) Human-System 249,683 user queries from 33K dialogues; 154K annotated privacy phrases (85,320 English phrases from 97,659 queries; 68,910 Chinese phrases from 151,988 queries) A large-scale multilingual dataset of user queries drawn from LLM interaction and human-human dialogue corpora (ShareGPT, CrossWOZ, DuConv, LCCC-base), automatically annotated with privacy phrase spans and natural-language privacy information summaries using a GPT-4o-based pipeline. Designed to support development and evaluation of local privacy detection models for real-name user interactions with LLMs. Zeng et al. 2025
MultiTalk Bilingual (Chinese and English) Speech Synthesized speech audio, dialogue scripts with paralinguistic and emotional annotations Multi-party multi-turn conversational speech synthesis; diverse topics (family, health, education, environment, career, technology, entertainment, others) Human-System (agent-synthesized multi-party dialogues with 30 distinct characters) 4,437 utterances; 100,773 tokens total (57,737 CN + 43,036 EN); ~32,441 seconds total audio (14,086s CN + 18,355s EN) 4.44 utterances per dialogue (4.81 CN, 4.41 EN) MultiTalk is a bilingual (Chinese and English) multi-party, multi-turn speech dialogue dataset generated using the DialogueAgents hybrid agent-based framework, featuring 30 distinct characters, rich paralinguistic and emotional annotations, and diverse topic coverage. It is designed to support research on conversational speech synthesis and dialogue systems. Li et al. 2025
SeniorTalk Mandarin Chinese Speech Audio recordings, transcripts with timestamps, speaker demographic metadata (age, gender, regional origin), accent intensity labels, paralinguistic event markers (laughter, noise, music) Spontaneous conversation among super-aged seniors (75–85 years); topics include health, leisure, retirement life, diet, and others Human-Human 55.53 hours, 101 conversations, 202 speakers, 60,029 utterances SeniorTalk is a Mandarin spontaneous conversational speech dataset comprising 55.53 hours from 101 natural dialogues involving 202 speakers aged 75–85, recruited from 16 provinces across China. It features rich multi-dimensional annotations (speaker demographics, temporal segmentation, overlapping speech, transcriptions, accent intensity, and paralinguistic markers) to support speaker verification, speaker diarization, speech recognition, and speech editing tasks targeting super-aged seniors. Chen et al. 2025
DeepDialogue English Multimodal (Text and Speech) Text dialogues, Synthesized emotional speech audio (two TTS variants: XTTS-v2 with emotion conditioning and Orpheus with implicit emotion) Open-domain multi-turn conversation spanning 41 domains (e.g., travel, cooking, philosophy, science) with explicit emotional progressions across 20 distinct emotion categories Human-System (LLM-LLM pairs simulating two conversational agents) 40,150 dialogues, 241,825 turns, 488+ hours of audio (per TTS variant) 6.1 DeepDialogue is a large-scale multimodal dataset of 40,150 high-quality multi-turn dialogues generated by pairing 9 LLMs (4B–72B parameters) across 41 domains and 20 distinct emotions with coherent emotional progressions. All dialogues are accompanied by synthesized emotional speech in two variants (explicit emotion-conditioned XTTS-v2 and implicit Orpheus), totalling over 480 hours of audio per variant, making it the first large-scale open-source multimodal dialogue dataset with turn-level emotional consistency. Koudounas et al. 2025
Diamante Mandarin Chinese Text Text (human-annotated dialogues with model-assisted candidate selection, revision, or rewriting) Open-domain chit-chat Human-WoZ (human annotators assisted by PLATO-XL dialogue model) 6,838 dialogues, 98,115 utterances 14.35 utterances per dialogue (approx.) Diamante is a Chinese open-domain chit-chat dataset collected via a human-in-the-loop annotation process in which annotators select, revise, or rewrite model-generated candidate responses (produced by PLATO-XL) to build high-quality multi-turn conversations. The dataset also captures implicit human preference signal (ranked responses) to support preference-aligned training, and covers 26 topic categories including Society, Entertainment, and Education. Lu et al. 2022
PSYDIAL Korean Text Synthetic dialogues (LLM-generated text) Personality-based chit-chat (Extraversion dimension of the Big Five personality model) Human-System (simulated; two LLM-generated virtual characters) 2,932 dialogues 8.16 PSYDIAL is the first Korean dialogue dataset focused on personality-based dialogues, generated via a five-step LLM prompting pipeline (personality setting, profile selection, dialogue generation, filtering, and regeneration). It covers four personality pairings based on the Extraversion/Introversion dimension of the Big Five model, with an average of 8.16 turns per dialogue and average utterance token length of 33.25 syllables. Han et al. 2024
Self-Feeding Chatbot Deployment Datasets English Text Text chat logs, user satisfaction ratings, textual feedback utterances Open-domain chit-chat (persona-based conversation) Human-System Three datasets: (1) deployment chat logs (513k messages); (2) satisfaction ratings (42k examples: 1k train, 500 valid, 1k test initial + 40k additional train); (3) textual feedback on bot errors (62k examples: 60k train, 1k valid, 1k test) 10 turns per conversation (20 utterances including initial prompt) for feedback collection conversations Three datasets collected during deployment of a self-feeding chit-chat agent on a crowdsourcing platform: (1) human-bot deployment chat logs (513k messages), (2) crowdsourced user satisfaction ratings (1–5 scale) for bot responses (42k examples), and (3) natural-language corrective feedback provided by users when the bot detected its own errors (62k examples). All datasets are available via the ParlAI platform. Hancock et al. 2019
APT Database English Text Text (synthetically generated empathetic dialogues with appraisal-theory-based emotion decompositions) Empathetic response / emotional support dialogue Human-System (ChatGPT-generated synthetic dialogues) 9,663 dialogues; 19,896 responses; 7 major emotion categories; 23 emotion subcategories; 230 factors; 2,415 situations ~2 turns per dialogue (short dialogues) The APT Database is a synthetically generated empathetic response resource constructed using ChatGPT guided by appraisal theory. It covers 7 major and 23 subcategories of emotions, 230 influencing factors, and 2,415 situations, yielding 9,663 short empathetic dialogues designed to support retrieval-augmented empathetic response generation. Hu et al. 2024
CSConv & RoleCS Chinese Text Text (transcripts of real customer-agent dialogues rewritten by LLMs, with support strategy annotations; plus LLM-synthesized role-playing dialogues) Customer support / customer service (banking and financial services topics: account & transaction management, product consultation, technical support, complaints & dispute resolution, marketing & promotions, risk management, financial consulting) Human-System (real customer–agent dialogues rewritten by LLM for CSConv; LLM role-playing agents simulating customer and supporter for RoleCS) CSConv: 1,855 dialogues, 50,587 utterances (rewritten), avg. 27.27 utterances/dialogue; RoleCS: 11,232 dialogues, 263,580 utterances, avg. 23.47 utterances/dialogue CSConv: 27.27; RoleCS: 23.47 CSConv is an evaluation dataset of 1,855 real-world Chinese customer–agent conversations rewritten by LLMs to reflect deliberate use of 12 COPC-grounded support strategies across 5 conversational stages, with expert annotations. RoleCS is a complementary synthetic training dataset of 11,232 strategy-rich dialogues generated via a multi-role LLM role-playing framework aligned with the same Customer Support Conversation (CSC) framework. Zhu et al. 2025
STAMPsy Mandarin Chinese Text Text (multi-turn dialogues annotated with counselor helping skills, spatiotemporal state stamps, dialogue goal types, and knowledge graph triples) Psychological counseling — covering five dialogue types: task-oriented dialogue for diagnosis, knowledge-grounded dialogue, conversational recommendation, empathetic dialogue, and question answering Human-System (LLM-simulated therapist–client dialogues, reviewed and annotated by clinical psychologists) 5,006 dialogues, 50,423 utterances (24,762 user / 25,661 bot), 8 counselor helping skill categories ~16.51 goals per dialogue; avg. 3.74 distinct dialogue-type goals per dialogue STAMPsy is the first Chinese spatiotemporal-aware mixed-type dialogue dataset for psychological counseling, containing 5,006 multi-turn conversations annotated with five dialogue types (diagnosis, knowledge-grounded, conversational recommendation, empathetic dialogue, QA), eight counselor helping skills, and spatiotemporal state stamps linking dialogues to time, location, and weather context. Each dialogue includes at least three distinct dialogue-type goals and is grounded in a psychological knowledge graph constructed under the 9-Box Case Conceptualization Model. Wang et al. 2024
STICKERCONV English Multimodal (text and image/sticker) Text, Sticker images Multimodal empathetic dialogue Human-System (LLM-agent simulated) 12,931 dialogue sessions; 67,505 sticker usages (5,800 unique stickers); 70,048 turns (train+val+test); 2,000 user personality profiles 5.49 turns per session STICKERCONV is the first multimodal empathetic dialogue dataset, comprising 12,931 dialogue sessions with interleaved text and sticker (image) responses generated by an LLM-based multi-agent system (Agent4SC). It covers 2,000 diverse user personality profiles, 5,800 unique stickers, and an average of 5.22 stickers and 5.49 turns per session, serving as a benchmark for multimodal empathetic response generation. Zhang et al. 2024
CaSiNo English Text Text (chat dialogues), persuasion strategy annotations, pre/post-survey responses (demographics, personality traits, satisfaction, opponent likeness) Negotiation (campsite neighbors negotiating for food, water, and firewood packages) Human-Human 1,030 dialogues, 846 unique participants; annotated subset (CaSiNo-Ann): 396 dialogues, 4,615 utterances 11.6 utterances per dialogue (avg. 22 tokens per utterance) CaSiNo (Camp Site Negotiation) is a corpus of 1,030 human-human negotiation dialogues collected via Amazon Mechanical Turk, in which pairs of participants role-play as campsite neighbors negotiating for food, water, and firewood packages. Dialogues are annotated with 9 persuasion strategies spanning cooperative to self-interested behaviors, and accompanied by participant demographics, personality traits, and post-negotiation satisfaction and opponent-likeness ratings. Chawla et al. 2021
MoPHES Multi-turn Counseling Dialogues Dataset Mandarin Chinese Text Text (synthetic multi-turn counseling dialogues and mental health condition labels) Psychological counseling / mental health support (anxiety and depression) Human-System 34,381 multi-turn dialogues (dialogue dataset); 6,046 labeled samples (mental conditions dataset); benchmark: 200 mental condition samples + 100 dialogue samples 5.00 A Chinese multi-turn psychological counseling dialogue dataset constructed by transforming 34,827 single-turn QA pairs (sourced from PsyQA and EmoLLM) into 5-turn dialogues via GPT-4o-mini prompting, along with a 6,046-sample mental conditions dataset labelled for anxiety and depression severity. An accompanying benchmark with 200 mental-condition samples and 100 dialogue samples supports automatic evaluation of mental state prediction and multi-turn counseling dialogue quality. Wei et al. 2025
MeDial-Speech English Speech Audio (WAV), manual speech transcriptions, ASR transcriptions, speaker role annotations, Audacity segment timing files Medical consultations covering four health conditions: Lewy body dementia, heart failure, shoulder pain, and angina Human-WOZ (robot-patient via teleoperated Wizard-of-Oz) and Human-Human (doctor-patient) 581 dialogues, 11,197 turns, 264,451 words, 111.4 hours of speech, 12.6 GB 22.48 MeDial-Speech is a spoken medical consultation dataset collected in realistic environments from both robot-patient (Wizard-of-Oz teleoperation) and doctor-patient dialogues, covering four health conditions: Lewy body dementia, heart failure, shoulder pain, and angina. It includes 111+ hours of audio with manual transcriptions and speaker role annotations, and is accompanied by a dialogue benchmark (sentence selection) for evaluating LLMs on medical conversational AI tasks. Cuayáhuitl et al. 2026
Synthetic Dutch Medical Dialogues Corpus Dutch Text Synthetic text dialogues Medical consultations (nephrology) Human-Human (simulated doctor–patient via LLM generation) 9 dialogues; mean 867 words and 39 turns per dialogue 39 A corpus of nine synthetic Dutch doctor–patient medical dialogues in the nephrology domain, generated using a Dutch fine-tuned LLM (ChocoLlama) with real clinical conversation transcripts as linguistic and structural reference. Dialogues cover topics including symptoms, medication use, lifestyle, and laboratory results, and were evaluated with both quantitative metrics and qualitative review by native Dutch speakers and medical practitioners. Kuan et al. 2026
EMSDialog English Text Synthetic multi-speaker dialogue transcripts annotated with speaker roles, turn-level topics, and diagnosis labels Emergency Medical Services (EMS) — conversational diagnosis prediction Multi-party human roles (medic, partner, patient, bystanders, dispatcher) — synthetically generated (Human-System simulation) 4,414 dialogues; 43 diagnosis classes; average 5.2 speaker roles per dialogue 114.3 utterances per dialogue EMSDialog is a large-scale synthetic dataset of 4,414 multi-speaker EMS conversations grounded in real-world Electronic Patient Care Reports (ePCRs), generated via a multi-LLM-agent pipeline with rule-based factual and topic-flow verification. Each dialogue is annotated with 43 EMS protocol diagnosis labels, turn-level speaker roles (medic, partner, patient, bystanders), and topic labels following official EMS clinical topic-flow guidelines. Ge et al. 2026
PRMB English Text Text (simulated and real-world CBT counseling dialogues, progressive session summaries, pairwise preference pairs, Best-of-N response sets) Cognitive Behavioral Therapy (CBT)-based psychological counseling, multi-session long-horizon dialogue Human-System (real CBT counselor responses vs. LLM-generated responses; simulated client personas) 13,893 prompts from 118 CBT cases (6 sessions each); 6,948 pairwise preference pairs; 6,945 Best-of-N queries; over 12K pairwise preference instances total PRMB is a benchmark for evaluating reward models in long-horizon, multi-session CBT-based counseling dialogue. It spans 6 therapeutic sessions and 21 diverse negative experience categories, comprising ~13,893 prompts derived from 118 validated CBT cases, with ~6,948 pairwise preference pairs and ~6,945 Best-of-N queries generated by ten state-of-the-art LLMs, incorporating both pairwise and Best-of-4 preference evaluations. Zhou et al. 2026
Psy-Insight English, Mandarin Chinese (Bilingual) Text Transcripts of face-to-face counseling dialogues with multi-task labels (psychotherapy method, emotion, strategy, topic) and explainable annotations (turn-level reasoning and observation, session-level background, guidance, and summary) Mental health counseling Human-Human (therapist–client, face-to-face counseling) 951 sessions (520 English, 431 Chinese); 189 cases (114 English, 75 Chinese); 11,984 turns total (6,208 English, 5,776 Chinese) 46 turns/session (English); 77 turns/session (Chinese) Psy-Insight is the first bilingual (English and Chinese), explainable multi-task dataset of real face-to-face mental health counseling dialogues, collected from books and blogs. Dialogues are annotated at both turn level (therapist reasoning, client emotion/observation, strategy) and session level (psychotherapy method, topic, background, guidance, summary) to support multi-task learning and chain-of-thought fine-tuning of LLMs for mental health support. Chen et al. 2025
MedSynth English Text Synthetic dialogue-note pairs (doctor-patient dialogues and SOAP-format clinical notes) Medical documentation; Dialogue-to-Note (Dial-2-Note) and Note-to-Dialogue (Note-2-Dial) generation in primary care and related specialties Human-System (LLM-generated role-playing agents simulating doctor-patient interactions) 10,035 dialogue-note pairs covering 2,001 unique ICD-10 codes 47 turns per dialogue (avg); dialogues avg 932 tokens / 55 sentences; notes avg 621 tokens / 23 sentences MedSynth is a large-scale synthetic dataset of 10,035 medical dialogue-note pairs covering 2,001 ICD-10 codes, generated by a multi-agent GPT-4o pipeline informed by real-world disease distributions from a US insurance claims database. Dialogues simulate doctor-patient encounters and are paired with SOAP-structured clinical notes, providing an open, privacy-compliant resource for training and benchmarking medical documentation models. Mianroodi et al. 2025
EmoDoctor English Text Text (LLM-rewritten patient queries with negative emotions and corresponding empathetic/soothing doctor responses) Healthcare / Medical consultation with emotional support Human-Human (original data sourced from HealthCareMagic.com and iCliniq.com; rewritten by LLM) ~110K training dialogues (approximately 60K Empathetic Response entries + 50K Emotional Question + Soothing Response entries); 7K test dialogues 1 (single-turn dialogues) EmoDoctor is an emotionally-augmented medical dialogue dataset constructed by using large language models to rewrite real-world doctor-patient conversations from HealthCareMagic.com and iCliniq.com. It comprises two subsets: ~60K Empathetic Response (ER) entries where doctor responses are rewritten to express empathy and compassion, and ~50K Emotional Question + Soothing Response (EQ+SR) entries where patient queries are infused with one of five negative emotions (fear, anxiety, embarrassment, frustration, distrust) and doctor responses are rewritten to soothe those emotions while retaining medical knowledge. Tsai et al. 2025
PatientSim English Text Simulated multi-turn doctor-patient dialogue transcripts, structured clinical profiles Medical history-taking and differential diagnosis in emergency department settings (covering myocardial infarction, pneumonia, urinary tract infection, intestinal obstruction, and cerebral infarction) Human-System (LLM-simulated patient interacting with doctor LLM or human doctor) 170 clinical profiles × 37 personas; 108 dialogues used for persona evaluation, 52 dialogues for factual accuracy/plausibility evaluation PatientSim is an open-source, persona-driven patient simulator for generating realistic multi-turn doctor-patient consultation dialogues. It combines 170 structured clinical profiles derived from MIMIC-IV and MIMIC-IV-ED real-world data with 37 distinct patient personas defined along four axes (personality, language proficiency, medical history recall level, and cognitive confusion level), supporting evaluation and training of medical dialogue systems and serving as an educational tool for healthcare. Kyung et al. 2025
LCMDC (Large-scale Chinese Medical Dialogue Corpora) Mandarin Chinese Text Text (patient consultations and doctor responses scraped from an online medical platform) Medical triage and consultation (coarse-grained department triage, fine-grained disease diagnosis, and open-ended medical Q&A) Human-Human (patient queries and doctor responses) Three sub-datasets: Coarse-grained Triage dataset (439,630 samples, 14 categories); Fine-grained Diagnosis dataset (199,600 samples, 120 categories); Medical Consultation dataset (472,418 Q&A pairs) LCMDC is a large-scale Chinese medical dialogue corpus comprising three sub-datasets collected from the “Quick Doctor” online medical platform: a coarse-grained triage dataset (~440K patient consultations across 14 departments), a fine-grained diagnosis dataset (~200K entries across 120 diseases), and a medical consultation dataset (~472K question-answer pairs. It is designed to support medical triage classification and open-ended medical dialogue generation research. Wang et al. 2024
SPADE Dialogue Datasets English Text Text (human-written and LLM-generated customer service dialogues) Customer service (hotel booking); Machine-Generated Text detection Human-System and Human-Woz (source); augmented to Human-LLM and LLM-LLM variants 14 datasets, each containing 616 dialogues, derived from 616 refined MultiWOZ 2.1 hotel dialogues A collection of 14 synthetic dialogue datasets for Machine-Generated Text (MGT) detection, produced via five structured prompt-based data augmentation frameworks (Missing Sentence Completion, Next Response Generation, Goal-to-Dialogue, Paraphrase, and End-to-End Conversation) applied to refined MultiWOZ 2.1 hotel-booking dialogues. Datasets span Partial-Chatbot and Full-Chatbot categories, generated using GPT-3.5 and Llama 70B, and are benchmarked against eight MGT detection models. Li et al. 2025
Physician Intent Trajectories Dataset (Aci-bench annotation) English Text Transcripts with physician intent annotations Medical / Clinical: physician intent classification and next intent prediction in doctor-patient dialogues Human-Human 207 dialogues, 5,541 doctor-patient turns (5,292 usable samples after processing) ~26.8 turns per dialogue (5,541 turns / 207 dialogues) A fine-grained physician intent annotation layer over the Aci-bench doctor-patient dialogue dataset, covering 207 role-played clinical dialogues and over 5,000 turns labeled with 20 intent classes organized under the SOAP framework (Subjective, Objective, Assessment, Plan). Annotations were verified by approximately 90 medical experts recruited via the Prolific crowd-sourcing platform, achieving 81.13% annotation accuracy. Röhr et al. 2025
Chinese Customer Service Dialogue Intent Clustering Dataset Mandarin Chinese Text (transcripts of audio calls) Transcripts (ASR from customer service calls), human-annotated intent cluster labels Customer service intent clustering; domains include banking, telecommunications, and insurance Human-Human (customer and service agent) 8,184 dialogues; 55,085 unique sentences; 1,507 human-annotated intent clusters A large-scale Chinese dialogue intent clustering dataset derived from audio transcriptions of over 100,000 real-world customer service calls across banking, telecommunications, and insurance domains. The dataset comprises 55,085 unique sentences annotated into 1,507 intent clusters (885 domain-specific, 622 out-of-domain) by 15 human experts using an “Action-Objective” naming convention, and is notable for its high semantic diversity and inclusion of noisy, out-of-domain queries. Hong et al. 2024
FineMed English, Chinese Text Synthetic instruction-response pairs, including common responses and long-form (o1-style) reasoning responses; DPO preference pairs Medical question answering and dialogue, spanning 5 primary medical specialties and 29 subspecialties (e.g., internal medicine, surgery, obstetrics and gynecology, pediatrics, otorhinolaryngology) Human-System (synthetically generated via LLM pipeline) ~300,000 SFT instruction-response pairs; ~33,000 DPO preference pairs FineMed is a large-scale synthetic medical SFT dataset generated from internet medical corpora (FineFineWeb), comprising ~300,000 high-quality instruction-response pairs across 5 primary and 29 secondary medical categories, with quality and complexity filtering via an LLM-as-a-judge framework. It also includes ~33,000 DPO preference pairs pairing long-form o1-style reasoning responses with common responses, designed to support multi-stage supervised fine-tuning and preference optimization of medical LLMs. Yu et al. 2025
CSDS Mandarin Chinese Text Text (dialogue transcripts, abstractive summaries, extractive key utterance annotations) Customer service dialogue summarization (e-commerce) Human-Human 10,701 dialogues (9,101 train / 800 dev / 800 test); 30,000+ dialogue-summary pairs across three summary types ~26 turns (25.11–26.00 across splits) CSDS is a fine-grained Chinese customer service dialogue summarization dataset built on real-world e-commerce conversations. Each dialogue is annotated with three types of abstractive summaries—an overall summary, a user-oriented summary, and an agent-oriented summary—all organized by topic structure via QA-pair annotation, along with key utterance indexes as extractive references. Lin et al. 2021
ABCD (Action-Based Conversations Dataset) English Text Text (fully labeled dialogues with action annotations, subflow labels, slot-value annotations, and Agent Guidelines) Customer service / online retail task-oriented dialogue with policy-constrained procedural actions Human-Human 10,042 dialogues, 177,407 turns (train split), 1,626,160 tokens (train split); 55 user intents, 30 domains, 231 slots 22.1 ABCD is a fully-labeled human-to-human customer service dialogue dataset containing 10,042 conversations grounded in explicit company policy guidelines. It features 55 distinct user intents requiring unique policy-constrained action sequences, and supports two novel tasks: Action State Tracking and Cascading Dialogue Success. Chen et al. 2021
Customer Service Dialogue Summarization Dataset Mandarin Chinese Text (ASR transcripts) Dialogue transcripts (from ASR) and human-written abstractive summaries Customer service (E-commerce call centre); topic-oriented dialogue summarization Human-Human (customer and service agent) 18,860 dialogues, 953K utterances (train: 17,189 / dev: 820 / test: 851) ~50.5 utterances per dialogue (overall avg across splits ~1,285 tokens per dialogue) A real-world Mandarin Chinese spoken dialogue dataset collected from the call centre of an E-commerce company, comprising ~18.86K dialogues automatically transcribed from audio via ASR (CER 9.3%) and paired with agent-written abstractive summaries covering the customer’s problem and the agent’s solution. Average dialogue length is ~1,285 tokens and average summary length is ~54 tokens. Zou et al. 2021
TeleSalesCorpus Chinese (grounded in real-world Chinese telemarketing; dialogues generated in simulation) Text Text (multi-turn dialogues with dialogue state annotations) Telemarketing / goal-driven persuasive sales dialogue Human-System (simulated: LLM-based User Agent vs. LLM-based Sales Agent, orchestrated by a Dialogue Manager) 2,000 dialogues TeleSalesCorpus is the first large-scale, real-world-grounded dialogue dataset for telemarketing, consisting of 2,000 high-fidelity multi-turn sales conversations generated via a state-aware three-agent LLM simulation seeded from anonymized real-world sales interactions. The corpus captures complex business rules, promotional objectives, customer objections, and diverse conversational states across the full telemarketing dialogue lifecycle. Zhang et al. 2025
CToMPersu English Text Text (synthetic multi-turn persuasive dialogues with mental state annotations) Persuasive dialogue, multi-domain (e.g., travel, health, technology, business; 35 domains total) Human-System (multi-agent LLM simulation with role-separated persuader and persuadee agents) 6,275 dialogues across 35 domains CToMPersu is a large-scale, multi-domain, multi-turn persuasive dialogue dataset constructed using ToMMA, a multi-agent framework guided by causal Theory of Mind. It enforces double-blind role separation between persuader and persuadee agents and includes structured mental state annotations (generative and preventative belief/desire components) to ensure causal Theory-of-Mind consistency across 6,275 dialogues spanning 35 domains. Zhang et al. 2025
PersuasionForGood English Text Text (chat transcripts), persuasion strategy annotations, participant demographic and psychological survey data Charity donation persuasion (Save the Children) Human-Human 1,017 dialogues; 300 annotated dialogues (ANNSET); 4,313 annotated sentences; 1,285 participants; 8,141 unique tokens 10.43 A human-human persuasion dialogue dataset collected on Amazon Mechanical Turk in which one participant (persuader) attempts to convince the other (persuadee) to donate part of their task earnings to a charity (Save the Children). A subset of 300 dialogues is annotated with 10 persuasion strategy categories, and all participants completed pre- and post-task surveys capturing demographic and psychological profiles (Big-Five personality, Moral Foundations, Schwartz Portrait Value, Decision-Making style). Wang et al. 2019
FaRM (Fact to Misinform) English Text Text (factual multiple-choice questions paired with systematically generated persuasive misinformation) Factual QA with persuasive misinformation; LLM robustness evaluation Human-System 1,952 entries (drawn from 1,500 questions across BoolQ, Natural Questions, and TruthfulQA; each question paired with control statements and three types of rhetorical appeals, each with three unique persuasive messages) Up to 4 turns per persuasive conversation FaRM is a dataset of straightforward factual multiple-choice questions (sourced from BoolQ, Natural Questions, and TruthfulQA) paired with systematically GPT-4-generated persuasive misinformation, including control statements and logical, credibility, and emotional rhetorical appeals. It is designed to evaluate LLMs’ susceptibility to belief change under multi-turn persuasive dialogue containing misinformation. Xu et al. 2023
SoMi-ToM English Multimodal (text, video, images) First-person game screenshots, third-person perspective videos with subtitles, multi-agent dialogue transcripts, action commands, system feedback Theory of Mind evaluation in embodied multi-agent social interactions (Minecraft crafting tasks with collaboration and obstruction social dynamics) Human-System (LVLM agents controlling Minecraft characters) 35 tasks; 35 third-person perspective videos; 363 first-person perspective images; 1,225 expert-annotated multiple-choice questions (1,050 first-person, 175 third-person) SoMi-ToM is a multimodal benchmark for evaluating multi-perspective Theory of Mind (ToM) in embodied multi-agent social interactions, built from LVLM agent interactions in Minecraft. It covers diverse crafting goals and social relationships (collaboration and obstruction), supporting both first-person real-time state inference and third-person goal and behavior inference, with 1,225 expert-annotated multiple-choice questions across 35 tasks. Fan et al. 2025
Multilingual Persuasive Dialogue Dataset (RPG) Multilingual (English, Spanish, French, Italian, German) Text Text (video game dialogue lines, automatically extracted and labelled) Persuasion detection in dialogue, extracted from role-playing video games (Neverwinter Nights, Knights of the Old Republic 1 & 2) Human-System (player–NPC dialogue options) 41,484 sentences total after tokenization (7,903 persuasive, 33,581 non-persuasive); multilingual parallel across 5 languages A multilingual, parallel dataset of persuasive and non-persuasive dialogue sentences automatically extracted and labelled from three BioWare RPGs (Neverwinter Nights; Knights of the Old Republic 1 & 2), covering English, Spanish, French, Italian, and German. Each instance is labelled as persuasive or non-persuasive based on in-game developer tags (e.g., [Persuade]), and sentences are aligned across all five languages via shared string IDs. Pöyhönen et al. 2022
CVLUE Chinese Multimodal (text and image) Images, captions, question-answer pairs, referring expressions, visual dialogues Chinese culture-centric vision-language understanding: image-text retrieval, visual question answering, visual grounding, and visual dialogue Human annotators (crowdsourced) ITR: 30,009 images (17,920 train / 3,116 valid / 8,973 test); VQA: 24,102 images (14,362 / 2,571 / 7,169); VG: 18,119 images (10,769 / 1,965 / 5,385); VD: 6,662 images (3,975 / 651 / 2,036); 92 object categories across 15 semantic fields Up to 10 Q&A turns per visual dialogue (VD task) CVLUE is a Chinese vision-language understanding evaluation benchmark whose images were collected from the Chinese Internet by native Chinese speakers, ensuring cultural representativeness. It covers four tasks—image-text retrieval, visual question answering, visual grounding, and visual dialogue—across 92 object categories from 15 semantic fields that reflect Chinese culture. Wang et al. 2024
Non-Cooperative GuessWhat?! Corpus English Text Text (yes/no/n/a dialogue turns in a visual question-answering game) Visual dialogue — non-cooperative goal-object identification (GuessWhat?! game variant) Human-System (human non-cooperative answer-players vs. autonomous question-player) 3,746 dialogues; ~2.7K unique images; ~2.8K unique objects; ~8.1K questions; ~2.3K unique words 4.99 questions per dialogue A corpus of 3,746 non-cooperative dialogues built on top of the GuessWhat?! visual dialogue game, in which human crowdworkers play as deceptive answer-players attempting to mislead an autonomous question-player away from the correct goal object. The dataset supports research on detecting and modeling non-cooperative conversational behavior. Sicilia et al. 2022
TouchStone English Multimodal (text and image) Images, questions, fine-grained human image annotations, and model-generated dialogue responses Evaluation of large vision-language models across five ability categories: basic descriptive ability, visual recognition, visual comprehension, visual storytelling, and multi-image analysis Human-System 908 questions across 27 subtasks and 5 major categories TouchStone is a comprehensive visual dialogue evaluation dataset consisting of open-world images and manually annotated questions covering five major categories of vision-language abilities and 27 subtasks, ranging from basic recognition and description to literary creation and multi-image analysis. It is designed for automated evaluation of large vision-language models using LLMs as judges, by converting image content into fine-grained textual annotations. Bai et al. 2023
MOD (Meme incorporated Open-domain Dialogue) Mandarin Chinese Multimodal (text and image/meme) Text utterances, Internet meme images, emotion annotations per meme-bearing utterance Open-domain conversation with Internet memes Human-Human 45,174 dialogues, 606,014 utterances, 307 unique Internet memes 13.42 A large-scale Chinese multimodal open-domain dialogue dataset in which Internet memes are incorporated into multi-turn conversations. Each meme-bearing utterance is annotated with a corresponding emotion label, and the dataset includes a “hard” test split featuring memes unseen during training to evaluate model generalisation. Fei et al. 2021
CTA Interesting Facts Dataset English Text Annotated text facts with relevance, interestingness, and feature-level labels (conciseness, specificity, novelty, relevance, informativeness); entity links; source URLs Cooking / Conversational Task Assistance Human-System 1,379 annotated interesting facts; 420 unique entities; 606 expert-annotated positive instances (from a 750-fact manually annotated subset) A dataset of 1,379 task-specific interesting facts for the cooking domain, extracted from online sources and annotated for relevance, interestingness, and five interestingness features (conciseness, specificity, novelty, relevance, informativeness) grounded in socio-psychological theories of human interest. Designed to support research on user engagement in Conversational Task Assistants (CTAs). Vedula et al. 2024
Syn-Multi English Speech Synthesized audio, transcripts, intent and slot annotations Task-oriented dialogue (Restaurant and Movie domains) Human-System Combined Restaurant (11,234 turns, 1,116 training dialogues) and Movie (3,562 turns, 384 training dialogues) domains; 3 intents, 12 slot types, 21 user dialogue act types Syn-Multi is a synthetic multi-turn end-to-end spoken language understanding dataset built by applying a Transformer text-to-speech model to existing text-only Restaurant and Movie domain dialogue data, combining them into a single audio corpus with intent and slot annotations for multi-turn E2E SLU research. Wei et al. 2021
LOTUSDIS Thai Speech Audio (multi-channel single-device recordings), transcripts with speaker labels and overlap masks Far-field meeting transcription / conversational ASR Multi-party human (3 participants per session) 114 hours total (multi-channel); ~20 hours unique sessions; 90 sessions; 86 unique speakers; 160,915 utterances (train+dev+test) LOTUSDIS is the first publicly available Thai far-field conversational speech corpus, comprising 114 hours of spontaneous, unscripted multi-party meeting recordings captured simultaneously by nine independent single-channel devices spanning six microphone types at distances from 0.12 m to 10 m. It includes utterance-level transcripts with speaker labels and overlap masks, and is partitioned into standard train/dev/test splits with a reproducible Whisper-based ASR baseline. Tipaksorn et al. 2025
DenseVisDial English Multimodal (text and image) Text (questions, answers, relevance-annotated reference answer sets); Images Visual dialogue — answering sequences of questions about images Human-Human 123,287 training images, each with up to 10 Q&A exchanges and 100 candidate answers per question (covering ~1.2M Q&A pairs); automatic reference sets (Σ) constructed for the entire VisDial v1.0 dataset 10 DenseVisDial is an extended annotation of the VisDial v1.0 dataset in which semi-automatically constructed sets of relevant reference answers are provided for every question–image pair across the full dataset. The reference sets are built using a semi-supervised CCA-based method seeded from sparse human relevance annotations and validated via Amazon Mechanical Turk, and are released together with a revised generative evaluation scheme based on NLP consensus metrics (CIDEr, METEOR, BERT, FastText). Massiceti et al. 2020
VoiceAgentBench Multilingual (English, Hindi, Bengali, Marathi, Tamil, Telugu, Malayalam) Speech Synthetic spoken audio (TTS-generated queries with diversity-based voice conversion), structured ground-truth tool invocation annotations Agentic tool use for voice assistants: single-tool invocation, single-tool with retrieval, parallel tool calling, sequentially dependent tool calling, multi-turn dialogue-based tool calling, and adversarial safety evaluation Human-System 6,134 spoken queries varies by category (single-turn to multi-turn; multi-turn subset from API-Bank) VoiceAgentBench is a comprehensive speech benchmark for evaluating Speech Language Models (SpeechLMs) on realistic agentic tasks, comprising 6,134 synthetic spoken queries across English and six Indic languages (Hindi, Bengali, Marathi, Tamil, Telugu, Malayalam). It spans six evaluation categories—single-tool invocation, single-tool with retrieval, parallel tool calling, sequentially dependent tool calling, multi-turn dialogue-based tool calling, and safety evaluations—with speaker diversity simulated via a farthest-point-sampling strategy on ECAPA-TDNN speaker embeddings for TTS voice conversion. Jain et al. 2025
SimsConv English Text Text (LLM-generated multi-turn role-playing dialogues with character profiles, scene descriptions, emotions, and conversation topics) Persona-driven role-playing / customisable character simulation Human-System (LLM-generated character-to-character interactions) 68 customised characters, 1,360 scenes, 13,971 multi-turn dialogues 10.3 SimsConv is a role-playing dialogue dataset featuring 68 freely customisable fictional characters defined by pre-defined aspects (career, aspiration, traits, skills) and expanded personal/social profiles. Characters interact across 1,360 real-world scenes with 13,971 multi-turn dialogues guided by 16 emotion types and 18 conversation topics, generated using GPT-4 with human verification. Yang et al. 2024
CharacterBot Lu Xun Dataset Mandarin Chinese Text Text (essay collections, multiple-choice QA pairs, generative QA pairs, style transfer pairs) Character persona simulation; literary style transfer and ideological comprehension based on Lu Xun’s essays Human-System 638 essays across 17 collections; 1,914 multiple-choice QA instances; 1,914 generative QA instances; 1,907 style transfer instances A Chinese-language dataset derived from 17 essay collections (638 essays) by the renowned writer Lu Xun, used to train and evaluate deep character persona simulation. It comprises three fine-tuning task subsets — multiple-choice question answering, generative question answering, and style transfer — each generated via GPT-4o and validated by human annotators to capture Lu Xun’s linguistic style and ideological depth. Wang et al. 2025
CharacterEval Chinese Text Text (dialogues with character utterances, behaviors, and scene descriptions; character profiles) Role-playing conversation; characters derived from Chinese novels and scripts Human-System 1,785 multi-turn role-playing dialogues; 11,376 examples; 77 characters; split into 6,811 training and 4,564 test examples 9.28 turns per conversation (avg. 369.69 tokens per conversation) CharacterEval is a Chinese benchmark dataset for evaluating Role-Playing Conversational Agents (RPCAs), comprising 1,785 multi-turn dialogues featuring 77 characters from diverse Chinese novels and scripts. Dialogues were extracted using GPT-4, filtered through human quality control, and augmented with detailed character profiles from Baidu Baike; evaluation covers 13 metrics across 4 dimensions including conversational ability, character consistency, role-playing attractiveness, and personality back-testing. Tu et al. 2024
CHILDES Filler-Gap Dependency Annotated Dataset English Text Transcripts with automated filler-gap dependency annotations (construction type and extraction site labels) Child language acquisition; filler-gap dependency detection in child-directed speech and child speech Multi-party human (children and adult caregivers) 2,841,084 utterances across 50,327 transcripts from 57 corpora An automatically annotated version of 57 English CHILDES corpora (North American English), in which utterances from child-directed and child speech are labelled for three filler-gap dependency constructions (matrix wh-questions, embedded wh-questions, and relative clauses) and their extraction-site subtypes (subject, object, adjunct, etc.), produced by a hybrid constituency- and dependency-parsing detection tool. Zhou et al. 2026
MERLIon CCS English, Mandarin Chinese (code-switching) Speech Audio (Zoom video call recordings), manual linguistic transcriptions with language-level timestamps Parent-child shared book reading (child-directed speech); language identification and language diarization Human-Human (parent–child pairs) 305 recordings, 112 parent-child pairs, ~57 hours total (28h36m dev + 28h47m eval); ~25h English speech, ~5h Mandarin speech; ~80K language segments MERLIon CCS is a first-of-its-kind Zoom video call audio corpus of English-Mandarin code-switching child-directed speech, collected from parent-child shared book reading sessions in Singapore home environments. It comprises 305 recordings from 112 parent-child pairs (over 30 hours annotated), featuring spontaneous in-the-wild code-switching in Singaporean English and Mandarin accents, manually annotated with fine-grained language-level timestamps by multilingual transcribers, and released to support language identification and language diarization research. Chua et al. 2023
WikiRole English, Mandarin Chinese Text Text (self-simulated multi-turn role-play dialogues) Character-based role-play dialogue Human-System Train: 3,902 roles (Chinese: 3,184; English: 3,902), 7,086 sessions, 36,164 turns; Test: 100 roles, 100 sessions, 498 turns WikiRole is a large-scale, multilingual, multi-turn role-play dialogue dataset constructed via the DITTO self-alignment method, in which an instruction-following LLM simulates role-play conversations grounded in character profiles collected from Wikidata and Wikipedia. Covering 3,902 characters in English and Chinese, it is approximately ten times larger in number of roles than previously available role-play datasets; the training split is self-generated by seed LLMs, while the held-out test split (100 roles, 498 turns) is generated using GPT-4-Turbo. Lu et al. 2024
CHILDES-UD2LF (Adam & Hagar CDS Corpora) English, Hebrew Text Transcripts with Universal Dependencies (UD) syntactic annotation and transduced sentential logical forms (LFs) Child-directed speech; child language acquisition Human-Human ~17K English utterances (Brown’s Adam corpus, ~80% of child-directed utterances); ~24K Hebrew utterances (Berman’s Hagar corpus, all child-directed utterances) Two corpora of child-directed speech (CDS) drawn from CHILDES — Brown’s Adam corpus (English, ~17K utterances) and Berman’s Hagar corpus (Hebrew, ~24K utterances) — annotated with cross-linguistically consistent Universal Dependencies (UD) syntactic structures and automatically transduced sentential logical forms (LFs), enabling comparative and computational studies of child language acquisition. Szubert et al. 2021
IND (Integrative Negotiation Dataset) English Text Text (dialogues with intent annotations) Integrative negotiation in online marketplace (price and product-bundle negotiation for electronic goods) Human-System (semi-automated GPT-J generation with human-in-the-loop post-editing) 4,163 dialogues (train: 3,330 / test: 500 / validation: 333); 57,393 utterances total 13 A dataset of integrative negotiation dialogues for the online marketplace domain, where products are modelled as bundles of items and negotiation covers price, addition/removal of bundle items, and delivery. Dialogues were generated via few-shot prompting of GPT-J using intent-action simulation, followed by human expert post-editing and quality filtering. Ahmad et al. 2023
Prosody-Text Aligned Child-Directed Speech Corpus English Speech, Text Audio, Orthographic transcripts, Word-level and phone-level forced alignments, Prosodic features (88-dimensional eGemaps feature vectors) Child-directed speech / language acquisition Human-Human (caregiver–child naturalistic interaction) ~414,500 child-directed utterances (Brent: ~154,700; Providence: ~259,800); language model training sentences: Brent 106,647, Providence 134,690, combined 267,337 A large-scale multimodal corpus of child-directed speech constructed from the Brent and Providence portions of the English CHILDES corpus, augmented with automatic word-level and phone-level text-audio alignments (via Montreal Forced Aligner) and automatically extracted 88-dimensional eGemaps prosodic feature vectors for each spoken word token. The corpus is used to investigate prosodic features as predictors of the age of acquisition of words. Frermann et al. 2017
D0T English Text Synthetic dialogues with silver-standard dialogue state annotations and slot descriptions Zero-shot dialogue state tracking across 1,000+ diverse task-oriented domains Human-System (LLM-generated synthetic dialogues) 1,003 domains, 5,015 dialogues, 100,471 turns, 487,460 slot-value pairs, 173,572 unique slot names, 2,061,332 tokens 20.0 D0T (Diverse 0-shot Tracking) is a fully automatically generated synthetic dataset for training zero-shot dialogue state tracking models, covering an unprecedented 1,000+ task-oriented domains. Each dialogue is generated using an LLM-based pipeline and annotated with silver-standard dialogue state updates and natural language slot descriptions. Finch et al. 2024
FineDialFact English Text Text (dialogue responses, atomic facts, Wikipedia evidence passages, factual labels) Fine-grained dialogue fact verification / hallucination detection Human-System 8,750 atomic facts total: 695 human-annotated (370 from HybriDialogue, 325 from OpenDialKG) and 8,055 GPT-4o-annotated (3,978 from HybriDialogue, 4,077 from OpenDialKG) FineDialFact is a benchmark for fine-grained dialogue fact verification, constructed by extending two public knowledge-grounded dialogue datasets (HybriDialogue and OpenDialKG). Dialogue responses are decomposed into atomic facts, each independently annotated (by humans and GPT-4o) with one of three labels—Supports, Refutes, or Not Enough Information—using Wikipedia as an external knowledge source. Chen et al. 2025
RECAP English Text Synthetically generated user-agent dialogue transcripts with human-vetted annotations Intent rewriting for agentic planning; task-oriented dialogue across domains including cooking, programming, health, flights, and restaurants Human-System (simulated user–agent dialogues generated via LLM) 810 validated conversation instances Short (~3 utterances), Medium (~7 utterances), Long (~12 utterances) RECAP (REwriting Conversations for Agentic Planning) is a benchmark of 810 synthetically generated and human-vetted user–agent dialogues designed to evaluate intent rewriting for downstream agentic planning. It covers five intent-understanding challenge types (shifted intent, noisy input, underspecified intent, multi-intent, perfect intent) across five domains and three conversation lengths. Mitra et al. 2025
DUO (Dialogue dataset with User subjective and Objective evaluations) English Text Text dialogues with subjective user evaluations and objective third-party annotations (preference, stylistic similarity, consistency, empathy/engagingness ratings on 5-point Likert scale) Open-domain dialogue (empathetic communication via EmpatheticDialogues setting; knowledge-grounded conversation via Wizard of Wikipedia setting) Human-System 314 dialogues (157 ED + 157 WoW); 96 third-party annotated dialogues (50 ED + 46 WoW); ~6,450 total utterances; ~107,797 total tokens ~10 turns per dialogue (~20–21 utterances) DUO is a multi-turn open-domain human-bot dialogue dataset collected via Amazon Mechanical Turk, featuring both subjective evaluations (preference, stylistic similarity, consistency, empathy/engagingness) from dialogue-participating users and objective evaluations from independent third-party annotators, across two settings (EmpatheticDialogues and Wizard of Wikipedia) and two dialogue systems (GPT-4o and Llama-3.1-70B-Instruct) under three style-control conditions. It is designed to support analysis of the relationship between stylistic similarity and user preferences in open-domain dialogue. Numaya et al. 2025
ISCO-800 English Text Text (user background profiles including name, occupation, educational background, personality, interests and hobbies, career history) Open-domain dialogue; user background diversity for proactive chatbot training Human-System (LLM-generated user agent profiles) 800 user background profiles across 40 occupation groups 5 ISCO-800 is a user background dataset containing 800 diverse user profiles spanning 40 sub-major occupation groups drawn from the ISCO-08 classification, designed to construct user agents for training and evaluating proactive open-domain chatbots. Each profile (~50–100 words) includes name, occupation, educational background, personality, interests, hobbies, and career history, generated via GPT-4 to ensure realism and diversity, and is split into training (500), validation (100), and test (200) sets. Wang et al. 2025
ImplexConv English Text Text (synthetically generated multi-session dialogues with persona traits and implicit reasoning scenarios) Long-term personalized open-domain conversation with implicit reasoning (question answering over multi-session history) Human-System 2,500 examples; ~255,000 total sessions; ~600,000 persona traits; avg. ~2,000 turns and ~60,000 tokens per example ~2,000 turns per example (across ~100 sessions) ImplexConv is a large-scale, long-term multi-session dialogue dataset comprising 2,500 examples, each containing approximately 100 conversation sessions, designed to study implicit reasoning in personalized dialogues. Unlike existing datasets, it uniquely incorporates both opposed and supportive implicit reasoning scenarios, where relevant persona information is embedded in subtle, syntactically or semantically distant connections rather than explicit statements, making it a challenging benchmark for retrieval-based and long-context models. Li et al. 2025
REALTALK English Multimodal (text and image) Text (messaging app transcripts), Images Long-term open-domain chit-chat / social conversation Human-Human 10 conversations, ~894 turns/conversation, ~21.9 sessions/conversation, ~17,110 tokens/conversation; 728 annotated memory-probing QA pairs; 600 annotated speaker events 894.4 REALTALK is a 21-day corpus of authentic human-human messaging app dialogues collected from 10 participant pairs (native US English speakers, aged 18–25), each spanning approximately 21 daily sessions and over 16,000 words per conversation, including shared images. The dataset is annotated with 728 memory-probing QA pairs (multi-hop, temporal reasoning, and commonsense) and speaker-level life events, supporting benchmarks for persona simulation and long-term memory evaluation in open-domain dialogue. Lee et al. 2025
SODA-Eval English Text Text (dialogue turns, GPT-4-generated issue annotations, overall quality scores on a 1–5 Likert scale, and natural language explanations) Open-domain dialogue quality evaluation Human-System (GPT-3.5-generated dialogues annotated by GPT-4, with human validation) 122,648 turn-level assessments across 10,000 dialogues (split: 85,876 train / 24,535 validation / 12,237 test) SODA-Eval is a large-scale open-domain dialogue evaluation benchmark built on the GPT-3.5-generated SODA dataset, containing over 120K turn-level quality assessments across 10K dialogues. Each annotation, produced by GPT-4, covers issue detection (coherence, commonsense, repetition, engagement, etc.) and an overall 1–5 quality score with natural language explanation, validated by human annotators. Mendonça et al. 2024
SynCPKL English Text Text (synthetic dialogue–persona knowledge fact pairs with binary relevance labels) Commonsense persona knowledge linking for open-domain dialogue Human-Human (sourced from PersonaChat) 39,802 examples (two variants: SynCPKL-H and SynCPKL-T, each 39,802 examples) 5 utterances per dialogue window SynCPKL is a synthetic dataset for training commonsense persona knowledge linkers, generated via the SynCPKL Pipeline using GPT-3.5-Turbo. Each example pairs a dialogue context window (5 utterances from PersonaChat) with a persona commonsense fact triple (head, relation, tail) from the PeaCoK knowledge graph, labeled for relevance to the target speaker. Lin et al. 2024
CGDIALOG+ English Text Text transcripts with human-annotated causal relations between dialogue history utterances and responses, plus pairwise human preference judgements Open-domain dialogue evaluation (emotional support conversation, multi-session chat, dialogue-based reading comprehension) Human-Human 2,444 history-response pairs (694 ESConv + 800 MSC + 950 DREAM); 9,970 utterances; 1,800 pairwise human preference annotations CGDIALOG+ is an extension of the CGDIALOG dataset that provides human-annotated causal relations between dialogue history utterances and responses across three domains (ESConv, MSC, DREAM), totalling 2,444 history-response pairs. It also includes 1,800 pairwise human preference judgements over responses from multiple dialogue systems, intended to facilitate development and evaluation of automatic dialogue response quality metrics. Feng et al. 2024
EDA (Emotional Dialogue Acts) English Multimodal (text, audio, video) Transcripts, emotion labels, dialogue act labels Conversational emotion and dialogue act recognition Human-Human 23,747 utterances (10,039 from IEMOCAP; 13,708 from MELD) The EDA corpus enriches two existing multimodal conversational emotion datasets (IEMOCAP and MELD) with automatically generated dialogue act labels using an ensemble of recurrent neural annotators trained on the Switchboard Dialogue Act corpus. It enables joint analysis of emotion and dialogue act co-occurrences in conversation. Bothe et al. 2020
SQPsychConv English Text Synthetic therapist-client dialogue transcripts Mental health counseling (Cognitive Behavioral Therapy for depression and anxiety) Human-System (LLM-simulated therapist and client agents) Seven dataset variants, each containing 2,090 conversations; utterance counts range from ~64,238 to ~101,694 per variant (e.g., SQPsychConvmistral: 98,342 utterances; SQPsychConvllama3.3: 101,694 utterances) 15.5–24.6 turns per dialogue (varies by model variant; e.g., llama3.3: 24.6, mistral: 23.1, command: 17.5) SQPsychConv is a collection of synthetic therapist-client counseling dialogues generated by the SQPsych pipeline, which conditions dual-agent LLM role-play on real-world structured client metadata and standardized psychological questionnaires (HAM-D, HAM-A, BDI) grounded in Cognitive Behavioral Therapy principles. Seven variants are produced using different open-weight LLMs (23B–123B parameters), each yielding 2,090 multi-turn conversations covering major depressive disorder and control groups, validated by human clinical experts and automatic LLM-based evaluation. Vu et al. 2025
DOTS English Text Text (simulated dialogues with slot schemas and dialogue state labels) Task-oriented dialogue across diverse domains (e.g., travel booking, garden planning, library book checkout); training split covers 787 domains, test split covers 25 domains across 10 multi-domain scenarios Human-System (LLM-simulated user and agent, with human guidance and correction for test set) Training: 2,771 dialogues, 88,240 turns, 787 domains, 6,810 slots, 44,120 values; Test: 300 dialogues, 7,844 turns, 25 domains, 208 slots, 3,922 values Training: ~31.8 turns/dialogue; Test: ~26.1 turns/dialogue DOTS is a fully automatic LLM-based task-oriented dialogue simulation dataset with ground-truth slot schemas and dialogue state labels spanning diverse task domains. The training split is generated automatically using GPT-4o/GPT-4o-mini; the test split was generated via the same pipeline with 10 handcrafted scenarios, then manually cherry-picked and corrected by human experts to support Slot Schema Induction (SSI) evaluation on novel, non-leaked domains. Finch et al. 2026
CoDEl-BR Brazilian Portuguese Speech and Text Audio (.flac), transcripts (YouTube automatic and Whisper ASR), topic annotations, candidate metadata (gender, race, party affiliation, election result) Electoral debate analysis; discourse and argumentation analysis, stance and sentiment detection, polarization modeling, topic modeling Multi-party human (candidates, journalist moderators, journalist narrators, and questioners) 2,943 transcript segments, ~32 hours of audio, 318,085 words, 22 debates, 13 Brazilian state capitals, 142 unique speakers (28 candidates) CoDEl-BR (Corpus de Debates Eleitorais Brasileiro) is a corpus of transcripts and audio recordings from 22 second-round mayoral debates held in 13 Brazilian state capitals during the 2024 municipal elections. It was constructed via a semi-automated multimodal pipeline and is enriched with dual ASR transcriptions (YouTube and Whisper), LLM-extracted topic annotations, and candidate demographic metadata (gender, race, party affiliation, election result). Gomes et al. 2026
Clinical ASR Impact Benchmark (Clinician-Annotated Subset) English Text Aligned ground-truth and ASR-hypothesis utterance pairs with clinician-assigned clinical impact labels Clinical dialogue: post-operative cataract consultations and general-practice (primary care) consultations Human-System 298 annotated utterance pairs (from 42 calls); Metrics Subset: 278 pairs A clinician-annotated benchmark of ASR transcription errors drawn from two doctor–patient dialogue datasets (proprietary Dora cataract consultations and open-source Primock57 primary-care mock consultations), each paired ground-truth/ASR utterance labelled by expert clinicians on a three-point clinical impact scale (No / Minimal / Significant Impact). The Primock57 clinical subset and accompanying code are publicly released; the Dora subset remains proprietary. Ellis et al. 2026
SPEECHMENTALMANIP English Speech (synthetic TTS audio) Synthetic multi-speaker TTS audio, text transcripts, manipulation labels (binary presence + tactic), human re-annotations Mental manipulation detection in spoken dialogue Human-Human (synthetic rendering of scripted movie dialogues) 2,915 dialogue transcripts rendered as audio; 609 manipulative and 90 non-manipulative clips used for evaluation (699 total evaluation clips); 100 dialogues human re-annotated in both text and audio modalities SPEECHMENTALMANIP is a synthetic multi-speaker speech benchmark for mental manipulation detection, created by augmenting the text-based MENTALMANIP dataset (movie dialogue snippets) with high-quality, voice-consistent Text-to-Speech rendered audio via a two-phase TTS pipeline using ElevenLabs voices. It enables direct comparison between text and speech modalities for detecting and attributing manipulative tactics (11 categories) in dialogue, and includes human re-annotations for a 100-dialogue subset under both modalities. Chen et al. 2026
InstructionVidDial English Multimodal (text, image, and video) Text dialogues, user-uploaded images, instructional video moments Instructional plan guidance (Cooking and DIY tasks) Human-System 6,760 dialogues, ~114K dialogue turns InstructionVidDial is a multimodal conversational dataset for instructional plan guidance, extending TastyVidDial with both Cooking and DIY plans (sourced from the COIN dataset). Dialogues are semi-automatically generated and augmented with plan-grounded visual question answering (pVQA) turns, where user-uploaded images are aligned with instructional video moments and plan steps. Glória-Silva et al. 2026
StoryMI English Text Text (simulated multi-turn dialogues with MI behavioral code annotations) Motivational interviewing (MI) psychotherapy; mental health counseling across 13 DSM-5 symptom domains Human-System (LLM-simulated therapist and client agents) 6,000 dialogues, 113K+ utterances, grounded in 1,000 questionnaire–story pairs; covers 12 MI codes and 13 symptom domains 13.3–25.6 turns (varies by model) StoryMI is a dataset of 6,000 simulated multi-turn motivational interviewing (MI) dialogues grounded in 1,000 questionnaire-derived situational stories covering 13 DSM-5 symptom domains and 12 MI behavioral codes. Dialogues are generated by a multi-LLM-agent framework featuring therapist, client, and interaction-manager agents, and are annotated with MI codes (MISC/MITI scheme) enabling macro-level counseling strategy evaluation. Meng et al. 2026
CogDialogue-QA Chinese Text Text (simulated teacher-student tutoring dialogues) STEM tutoring dialogues (Mathematics, Physics, Chemistry, Biology, Geography) for secondary school education Human-System (GPT-4o-simulated teacher and student agents) 23,209 teacher-student interaction turns; covers all relations and learning objectives in CogNet-KG (494 knowledge points, 7,920 relations/edges) CogDialogue-QA is a high-quality simulated tutoring dialogue dataset constructed from CogNet-KG, a cognitively-structured educational knowledge graph spanning five STEM subjects across secondary school. Dialogues are procedurally generated using GPT-4o, with teacher questioning strategies adaptively guided by cognitive relations and learning objectives in CogNet-KG, reflecting pedagogical principles such as Zone of Proximal Development and adaptive scaffolding. Yu et al. 2026
LongMP-Bench English Multimodal (text and image) Text (synthesized dialogues), Images (persona identity images and retrieved contextual images) Multimodal persona understanding in long-term personalized dialogue Human-System (synthetic user personas interacting with dialogue agents) 150 conversations, 2,105 sessions, 22,998 turns, 1,366 images; QA tasks: 8,978 questions across 6 subtypes; Response generation: 404 turns 153.32 turns per conversation (10.93 turns per session) LongMP-Bench is a benchmark for evaluating multimodal persona understanding in long-term dialogues, featuring 150 synthetically generated users with visually consistent and dynamically evolving personas across extended multi-session conversations. It includes QA and response generation tasks covering persona tracking, multimodal reasoning, and personalized response generation, with human refinement for quality assurance. Li et al. 2026
DraDDP English Multimodal (Text, Video, Audio) Text (subtitles/utterances), Video, Audio Multi-party dialogue discourse parsing (dependency structure and relation type identification) Multi-party human (fictional characters from TV drama) 495 dialogue segments, 6,374 utterances, 9.1 hours of parallel video content 12.88 DraDDP (Drama-based Dialogue Discourse Parsing) is the first publicly available English multimodal multi-party dialogue discourse parsing dataset, constructed from Season 1 of the American TV series Friends. It contains 495 dialogue segments with 6,374 utterances and 9.1 hours of parallel video content, annotated with discourse dependency structures and 16 SDRT-based relation types covering rich multi-party interaction scenarios. Liu et al. 2026
Live-Aid Chinese Multimodal (video, audio, text) Video clips, audio streams, timestamped viewer comments (danmaku), human-annotated temporally aligned video responses, dialogue summaries, ASR transcripts E-commerce live streaming (viewer–host interaction); covers 44 product categories Human-Human (multi-party: live stream viewers and host) 8,053 video sessions; 80,037 dialogue turns (53,319 text turns + 26,718 video turns); 53,319 danmaku messages; 1,100+ hours of video; 46,819 unique users; sourced from 1,763 live streams 9.94 turns per session Live-Aid is the first large-scale Chinese interleaved live-streaming dialogue dataset with human-annotated, temporally aligned video responses, spanning over 1,100 hours across 8,053 video sessions from e-commerce live streams. It features high-density viewer danmaku tightly coupled with real-time audio-visual evidence, and is accompanied by an agent-enhanced benchmark covering eight evaluation tasks across multimodal understanding, dialogue modeling, and temporal reasoning. Lei et al. 2026
OpenDialog English, Mandarin Chinese Speech Audio, Speaker-attributed transcripts (ASR) Spoken dialogue (open domain, in-the-wild) Human-Human 6,833 hours total (5,074 hours English, 1,759 hours Chinese) OpenDialog is the first large-scale open-source spoken dialogue dataset derived from in-the-wild speech data, comprising 6.8k hours of two-speaker dialogues in English and Mandarin Chinese. It was constructed via a multi-stage pipeline including VAD, speaker diarization, ASR with WhisperD, LLM-based dialogue classification, and DNSMOS-based quality filtering, and is intended for training spoken dialogue generation models. Zhu et al. 2026
J-Shuwa Japanese Sign Language (JSL) and Japanese Video (Sign Language) and Text Video clips with aligned Japanese text subtitles (hard-coded subtitles and closed captions) Sign Language Translation (JSL-to-Japanese) Human (Deaf signers recorded in YouTube videos) 197,742 parallel JSL-Japanese sentence pairs; ~300 hours of video; segmented from 6,322 videos; ~31K unique vocabulary items J-Shuwa is a large-scale Japanese Sign Language (JSL)–Japanese parallel corpus constructed from YouTube videos containing both hard-coded subtitles and closed captions. It comprises approximately 197K video-text sentence pairs (~300 hours), making it the largest publicly available JSL dataset, and is intended to support sign language translation and a broad range of JSL research tasks. Mo et al. 2026
INSURE-Dial English Text De-identified ASR transcripts (text only; no audio released), structured JSON phase annotations with span boundaries, ask/answer role flags, and compliance labels Insurance pharmacy benefit verification calls (U.S. healthcare); covers IVR navigation, patient identification, coverage status, drug formulary/restrictions/copay checks, and agent identification Human-System (AI-initiated outbound calls with live insurance representatives) 1,050 calls (50 real, 1,000 synthetic); 48,191 turns; 618,407 tokens; 2,100 drug queries 45.9 overall (71.2 for real calls; 44.6 for synthetic calls) INSURE-Dial is the first public benchmark for compliance-aware auditing of insurance benefit-verification phone calls. It comprises 50 de-identified real AI-initiated calls with live insurance representatives and 1,000 synthetically generated calls, all annotated with a phase-structured JSON schema covering ordered audit phases (IVR, greeting, patient identification, coverage status, drug formulary/restrictions/copay checks, and agent CRN), with span boundaries, ask/answer role flags, and Information/Procedural compliance labels supporting two evaluation tasks: Phase Boundary Detection and Compliance Verification. Kulkarni et al. 2026
PsyChainD Mandarin Chinese Text Text (synthetically generated multi-turn counseling dialogues) Psychological counseling across 10 DSM-5 personality archetypes and 86 counseling subtopics (e.g., love problems, family, self-growth, relationships, work) Human-System (LLM-simulated client and counselor agents) 10,456 dialogues 18.52 PsyChainD is a large-scale Chinese psychological counseling dialogue dataset of 10,456 synthetically generated multi-turn sessions, constructed using the PsyChain chain-of-agents framework. Each dialogue is grounded in one of 10 DSM-5 personality archetypes paired with diverse life scenarios, and features stage-structured therapeutic progression, safety monitoring, and expert supervisory guidance across 86 counseling subtopics. Feng et al. 2026
MentalBench-100k / MentalAlign-70k English Text Text (therapeutic conversation contexts, human therapist responses, LLM-generated responses, and human expert + LLM judge ratings on 7 attributes) Mental health support / therapeutic dialogue (crisis helplines, counselling, one-turn support interactions) Human-Human (original counselling dialogues); Human-System (LLM-generated responses paired with human contexts) MentalBench-100k: 10,000 authentic one-turn therapeutic conversations, each paired with 9 LLM-generated responses (100,000 response pairs total); MentalAlign-70k: 70,000 ratings across 1,000 conversations × 10 responses × 7 attributes, from 4 LLM judges and human experts 1 (single-turn) MentalBench-100k consolidates 10,000 authentic single-session therapeutic conversations from three real-world clinical and counselling datasets (MentalChat16K, EmoCare/Psych8k, CounselChat), each paired with responses from 9 diverse LLMs, yielding 100,000 response pairs. MentalAlign-70k provides 70,000 ratings on seven cognitive and affective attributes (Cognitive Support Score and Affective Resonance Score), comparing four LLM judges against clinical human experts across 1,000 conversations, enabling reliability analysis via the Affective–Cognitive Agreement Framework (ICC with bootstrap confidence intervals). Badawi et al. 2026
BOULDER English Text Synthetically generated dialogue histories with tool calls and results, isolated reasoning task prompts, automatically verifiable answers Travel-related task-oriented dialogue (trains, hotels, restaurants, attractions); covers arithmetic, spatial, and temporal reasoning Human-System 800 test examples (100 per task × 8 tasks), each presented in isolated and dialogue-based variants (3 main setups: baseline, dialogue, dialogue-concise) BOULDER (Benchmarking of Usefulness of LLMs in Dialogue-Embedded Reasoning) is a dynamic benchmark of 800 synthetically generated travel-related task-oriented dialogue examples covering eight reasoning tasks (arithmetic, spatial, temporal) across four domains. Each problem instance is presented in both an isolated and a multi-turn dialogue-based variant, enabling controlled comparison of LLM reasoning performance in and out of dialogue context. Kartáč et al. 2026
OlaBench Chinese (industrial deployment context; language not explicitly stated but implied by ByteDance/Chinese platform context) Text Text dialogues (multi-turn customer service sessions, including tool-call traces and retrieved QA pairs); human expert annotations for risk and hallucination labels Industrial intelligent customer service (ICS) spanning retrieval-augmented generation (RAG), workflow-based, and agentic settings; sub-domains include account services, identity & compliance, social ecosystem, and content & features Human-System OlaBench-Core: 3,000 dialogues (768 RAG + 952 Workflow + 1,280 Agent); OlaBench-Risk: 1,000 dialogues (618 + 228 + 138 + 16); OlaBench-Hall: 1,000 dialogues (481 + 321 + 144 + 54); totalling ~5,000 benchmark instances across three subsets OlaBench-Core: 3.5–3.8 turns (RAG 3.5, Workflow 3.6, Agent 3.8); OlaBench-Risk: 3.5–5.1 turns; OlaBench-Hall: 3.7–6.6 turns OlaBench is a real-world industrial customer-service benchmark derived from live deployment data, evaluating models across six dimensions: Dialogue Quality, Policy Compliance, Tool Calling, Critical Business Risk, Hallucination, and Latency. It covers three system paradigms (RAG, Workflow, Agent) and includes human-verified annotations for safety-critical risk and hallucination subsets. Gao et al. 2026
DebtBench English Text Synthetic persona profiles and simulated collector-debtor dialogues with strategy annotations Debt collection negotiation Human-System (simulated collector agent vs. LLM-based debtor agent) 11,000 debtor personas (10,000 train / 1,000 test); dialogues simulated per persona DebtBench is the first public, persona-enriched debt collection negotiation benchmark, constructed via a three-stage synthesis pipeline that distills behavioral patterns from 1,000 real collector–debtor conversations into privacy-preserving synthetic personas and dialogues. Each of the 11,000 debtor personas is characterized along four multi-dimensional axes—background, personality traits, cognitive attributes, and life-grounded scenario—capturing the rich behavioral heterogeneity (emotional expression, cognitive limitations, diverse linguistic styles) observed in real-world debt negotiation. Yang et al. 2026
SAD English Text Text (Reddit discussion threads with stance and argumentation strategy annotations) Multi-turn argumentative dialogue / debate (open-domain controversial topics) Human-Human 392,822 dialogue examples, 722,812 utterances, 20,619 topics 3.69 SAD (Strategic Argumentative Dialogue) is a large-scale dataset of real-world multi-turn argumentative dialogues derived from the Reddit r/ChangeMyView community, covering 20,619 controversial topics. Each utterance is annotated with a stance (support/oppose) and up to five argumentation strategy labels (Question, Causality, Example, Analogy, Statement), supporting strategy-conditioned argument generation tasks. Liu et al. 2026
JobNego and ResNego English Text Text (synthetic negotiation dialogues annotated with Emotion-aware Negotiation Strategy-informed Chain-of-Thought (ENS-CoT) rationales) Emotionally intelligent negotiation: job interview negotiation (JobNego) and resource allocation negotiation (ResNego) Human-System (Wizard-of-Oz seed dialogues; full datasets generated via ChatGPT prompting) JobNego: 840 dialogues (504 train / 168 dev / 168 test), 12,492 utterances total; ResNego: 1,648 dialogues (988 train / 330 dev / 330 test), 20,187 utterances total JobNego: ~16.1 (train), ~13.1 (dev), ~12.9 (test); ResNego: ~13.9 (train), ~10.0 (dev), ~9.6 (test) JobNego and ResNego are two synthetic negotiation dialogue datasets annotated with interpretable Emotion-aware Negotiation Strategy-informed Chain-of-Thought (ENS-CoT) rationales, covering 12 emotion categories and 12 emotion-aware negotiation strategies. JobNego contains job interview negotiations between a candidate and an employer, while ResNego contains resource allocation negotiations in a camping setting; both are generated via ChatGPT few-shot prompting with human expert quality verification. Kajare et al. 2026
SUMM-RE (EDU-segmented) French Speech, Text (transcripts) Audio recordings, manual transcripts, automatic transcripts, elementary discourse unit (EDU) segmentation annotations Meeting conversations (event planning); discourse segmentation into elementary discourse units (EDUs) Multi-party human (2–4 participants per session) ~100 sessions (~20 hours manually transcribed and annotated; ~80 hours automatically transcribed and segmented); ~73 meetings in dev/test with manual annotation; average ~534 EDUs per dialogue 534 EDUs per dialogue (average) A large French corpus of multiparty meeting dialogues annotated with elementary discourse unit (EDU) segmentation, built on the SUMM-RE corpus. It comprises approximately 20 hours of manually transcribed and discourse-annotated conversations and 80 hours of automatically transcribed and discourse-segmented data, covering event-planning discussions among 2–4 participants. Prévot et al. 2025
BRAGE Norwegian (Bokmål) Text (ASR transcripts of phone calls) Transcripts of customer service phone calls, annotated with product category labels Customer service dialogue classification; telecommunications product category identification Human-Human (customer and customer service agent) 300 dialogues BRAGE is a private benchmark of 300 transcribed Norwegian customer service phone calls from a telecommunications provider, annotated with eight product category labels. It is designed to evaluate zero-shot classification capabilities of large language models using the same codebook instructions provided to human annotators. Riess et al. 2025
DSLC7 Audio-Visual Task-Oriented Dialogue Dataset Japanese Multimodal (Speech, Video/Avatar, Visual display) Audio recordings (system and user), Video recordings (system CG avatar and Travel Viewer; frontal video of user), system input/output logs, subjective evaluation questionnaires (13 metrics per dialogue) Tourist spot selection (task-oriented travel planning dialogue between human users and dialogue systems) Human-System 257 dialogues, 1,865 minutes; 94 human evaluators, 9 systems (preliminary round); 6 additional dialogues (final round) A large-scale competition-based dataset of Japanese audio-visual task-oriented dialogues collected during the 7th Dialogue System Live Competition (DSLC7). It comprises 257 dialogues (1,865 minutes) between nine diverse dialogue systems and 94 human evaluators performing a tourist spot selection task, including audio and frontal video of both participants, system logs, and per-dialogue subjective evaluations across 13 metrics. Sato et al. 2025
CS-Sum Multilingual (Mandarin-English, Tamil-English, Malay-English) Text Text (code-switched dialogues with human-annotated English summaries) Dialogue summarization (code-switching) Human-Human 3,238 dialogues total: 1,320 EN-ZH, 1,000 EN-TA, 918 EN-MS CS-Sum is the first benchmark for code-switched (CS) dialogue-to-English summarization, covering three language pairs: Mandarin-English (EN-ZH), Tamil-English (EN-TA), and Malay-English (EN-MS). Dialogues were created by native-speaking university students translating English dialogues from DialogSum and SAMSum into naturalistic code-switched conversations, each paired with a human-annotated English summary. Suresh et al. 2025
MonoTODia English (translated from German) Text Annotated task-oriented dialogues (generated from e-mails), slot-value annotations, dialogue act annotations Travel booking (hotel, flight, package deals) Human-System 1,850 dialogues (1,500 train, 150 validation, 200 test); test split has crowd-worker gold-standard annotations MonoTODia is a task-oriented dialogue dataset for travel booking, generated by translating real-world German e-mail requests (from a travel agency company) into annotated multi-turn dialogues using fine-tuned LLMs. The test split features crowd-worker gold-standard slot and dialogue act annotations; train and validation splits use LLM-predicted annotations refined on the gold-standard test data. Steindl et al. 2025
PicPersona-TOD English Multimodal (text and image) Text (dialogues with DST and dialogue policy labels), Images (user face/persona images), Google Maps reviews, Wikipedia entries Task-oriented dialogue across 18 service domains (e.g., restaurant, hotel, attraction, train, taxi, bus, movie, home); personalized response generation Human-System 18,148 dialogues, 18 service domains 17.23 PicPersona-TOD is the first task-oriented dialogue dataset that incorporates user face images as visual personas, enabling personalized system responses tailored to user-specific factors such as age, formality, and emotional context. It is constructed via an automated GPT-4o-based pipeline combining MultiWOZ-2.2 and SGD dialogues with FFHQ user images, Google Maps reviews, and Wikipedia entries, and includes DST and dialogue policy labels. Lee et al. 2025
ArtGenEval-GPT++ English Text Synthetic dialogues (LLM-generated text) Art museum tour guidance and visitor engagement (art domain chatbot interactions) Human-System ~12,500 dialogues spanning 821 artworks from 384 artists across 26 art styles ArtGenEval-GPT++ is a synthetically generated dataset of approximately 12,500 dyadic and group multi-turn dialogues between museum visitors and a chatbot tour guide/expert, generated using GPT-4. The dataset covers diverse visitor profiles (age, gender, ethnicity, knowledge level, emotional state) and museum scenarios, and is designed for training and fine-tuning context-aware, personalized conversational agents in the art domain. Rachidi et al. 2025
DiaSafety-CC English Text Text (dialogue context-response pairs with safety labels and free-text rationales) Dialogue safety evaluation with cross-cultural annotation Human-System 1095 dialogues (single-turn context-response pairs), annotated by 6 raters (3 from Nigeria, 3 from India) 1 DiaSafety-CC is a cross-cultural reannotation of the DiaSafety English dialogue safety test set, in which three raters each from Nigeria and India provide Safe/Unsafe labels and free-text reasons for 1,095 single-turn context-response dialogues spanning five safety categories. The dataset enables cross-cultural analysis of dialogue safety annotation disagreements between Western and non-Western annotator groups, and includes rater demographic metadata. Ajayi et al. 2025
ShopDial English Text Text (synthetic dialogues generated via bottom-up LLM-based pipeline) E-commerce conversational question answering (shopping companion / customer service across six product categories: vacuums, diapers, sofas, TV, food, clothing) Human-System (simulated customer–virtual assistant) 6,000 dialogues 8.03 Shopping Companion Dialogues (ShopDial) is a synthetic, knowledge-grounded task-oriented dialogue dataset for e-commerce conversational QA, generated via a bottom-up pipeline (BUSY) that first produces factually grounded QA pairs from a product database and then connects them into coherent multi-turn conversations. It covers six shopping categories (vacuums, diapers, sofas, TV, food, clothing) and includes “unknown” turns and negative user feedback to reflect realistic interactions. Qian et al. 2025
DSLCMM Japanese Multimodal (speech, video/facial images, system gesture/facial expression commands) Audio, Video, Transcripts (user utterances), System command logs (gestures and facial expressions), Subjective evaluation scores Open-domain casual conversation and situated scenario-based dialogue (Human-Machine) Human-System 1,747 dialogues; 32 multimodal dialogue systems; 90,261 total utterances (across all subsets); ~143 hours of recorded dialogue DSLCMM is a Japanese multimodal human-machine dialogue corpus built from data collected across two editions (DSLC5 and DSLC6) of the Dialogue System Live Competition series. It comprises 1,747 dialogues between 32 multimodal dialogue systems and human users, including user/system speech and video recordings, system gesture and facial expression command logs, user utterance transcriptions, and subjective user evaluation scores on a 5-point Likert scale. Higashinaka et al. 2025
KoED Korean Text Text (reconstructed and handcrafted empathetic dialogues with multi-label emotion annotations) Empathetic dialogue; cross-cultural emotion recognition and empathetic response generation Human-Human 1,360 dialogues across 34 emotion categories (1,280 culturally adapted from English ED + 80 handcrafted for Korean-specific emotions ‘jeong’ and ‘han’); 40 dialogues per category; 555 unique primary/supplementary emotion label combinations KoED (Korean Empathetic Dialogues) is a culturally-reconstructed empathetic dialogue benchmark extending the English EmpatheticDialogues (ED) dataset. Rather than direct translation, dialogues were meticulously adapted to authentic Korean cultural contexts and supplemented with 80 handcrafted dialogues for uniquely Korean emotional concepts (‘jeong’ and ‘han’), with multi-label emotion annotations covering 34 emotion categories. It is designed exclusively as a zero-shot evaluation benchmark for assessing cross-cultural empathetic understanding in LLMs. Lee et al. 2025
Stephanie Dataset English Text Text (LLM-generated step-by-step dialogues with persona information) Open-domain social conversation (persona-based chit-chat) Human-System (simulated two-party dialogues generated via LLM) 5,457 dialogues A high-quality step-by-step dialogue dataset derived from the PERSONA-CHAT training set, generated using the Llama3-70b model with a dual learning strategy and a further-split post-editing method. Unlike single-step dialogue datasets, each dialogue consists of multiple short, consecutive messages per speaker turn, designed to mimic the natural flow of human instant-messaging conversations. Yang et al. 2025
AkaCE (Akan Cinematic Emotions) Akan Multimodal (Speech, Video, Text) Audio, Video, Text transcripts, Emotion labels, Word-level prosodic prominence annotations Emotion recognition in conversation; movie dialogues Multi-party human 385 dialogues, 6,162 utterances, 4,477 turns, 117,305 words, 308 speakers, 21 movies 11.62 AkaCE is the first multimodal emotion dialogue dataset for an African language (Akan), containing 385 emotion-labeled dialogues and 6,162 utterances sourced from 21 Akan-language movies, covering audio, visual, and textual modalities. It is also the first prosodically annotated African language dataset, featuring word-level prosodic prominence annotations alongside seven-class emotion labels and gender-balanced speaker representation (308 speakers). Sasu et al. 2025
LlamaPIE Semi-Synthetic Dialogue Dataset English Text Semi-synthetic dialogues with user profiles, memory/event contexts, proactive assistant responses, and silence/timing markers Proactive in-ear conversational assistance during human-human dialogue (reminders and social guidance) Human-Human with embedded AI assistant (simulated via Claude generation) 8,892 dialogues total (3,128 Synthetic + 2,758 SODA-based + 3,006 PerLTQA-based) ~22–23 speaker turns per dialogue; ~3.7–4.0 assistant turns per dialogue A semi-synthetic dataset of multi-party dialogues constructed to train proactive in-ear assistants, where each example includes a user profile, two contextual memory events, a timestamped conversation with silence markers, and concise 1–3 word assistant responses. Dialogues are generated using Claude and grounded in real conversational contexts from the SODA and PerLTQA datasets. Chen et al. 2025
PersonaLens English Text Simulated multi-turn task-oriented dialogues, user profiles (demographic information, preferences, past interaction summaries), task specifications, situational contexts Task-oriented assistance across 20 domains including Alarm, Books, Buses, Calendar, Events, Finance, Flights, Games, Hotels, Media, Messaging, Movies, Music, Rental Cars, Restaurants, Services, Shopping, Sports, Train, and Travel Human-System (LLM-simulated user agent interacting with LLM AI assistants) 122,133 dialogues; 1,500 user profiles; 111 tasks across 20 domains (86 single-domain, 25 multi-domain) 20 turns max for single-domain tasks, 30 turns max for multi-domain tasks PersonaLens is a benchmark for evaluating personalization in task-oriented conversational AI assistants, featuring 1,500 diverse user profiles with rich demographic information, preferences, and interaction histories, and 122,133 simulated dialogues spanning 111 tasks across 20 domains. It includes a LLM-based user agent for realistic dialogue simulation and a judge agent for automated assessment of personalization, task completion, and response quality. Zhao et al. 2025
EmoCare English Text Text (multi-turn dialogues with fine-grained problem type annotations, user scenarios, and seeker profiles) Emotional support conversation (mental health, empathetic dialogue) Human-System (LLM-simulated seeker and supporter roles, reviewed by psychology experts) 2,574 dialogues, 42,770 utterances, 45 fine-grained problem categories 16.61 EmoCare is a large-scale emotional support conversation (ESC) dataset constructed via a systematic fine-grained problem augmentation method, expanding problem type coverage from 13 coarse categories to 45 fine-grained categories across emotional, interpersonal, and behavioral domains. Dialogues feature diverse real-world scenarios and detailed seeker profiles, and were reviewed and validated by professional psychology experts. Shi et al. 2025
CliniDial English Multimodal (speech/audio, transcripts, video from two camera angles, physiological signals) Audio recordings, transcriptions, simulated patient physiological signals (9 signal types), video from 2 camera angles, behaviour-code annotations Clinical operation teamwork and team reflection; behaviour code classification (Seek, Evaluate, Plan, Implement, None) Multi-party human (anesthesiologist trainees, support staff, confederate surgeon; 6 participants per session) 22 sessions; 6,500 turns; 49,900 words; ~6,900 annotated utterances 311 turns per session CliniDial is a naturally occurring multimodal dialogue dataset collected from simulations of medical operations, featuring audio, transcripts, dual-angle video, and 9 simulated patient physiological signals. Dialogues are annotated with team-reflection behaviour codes (Seek, Evaluate, Plan, Implement, None) to study teamwork dynamics during clinical procedures. Deng et al. 2025
DICE-BENCH English Text Text (synthesized multi-party, multi-round dialogues with function-call annotations) Tool/function-calling evaluation in multi-round, multi-party dialogue (everyday scenarios such as weather checking, car rental, hotel booking, etc.) Human-System (simulated multi-party dialogues with 2–4 agent personas and a virtual AI assistant) 1,607 dialogue instances across 4 rounds; 124 tools, 270 tool-graph edges DICE-BENCH is a benchmark framework for evaluating LLM tool/function-calling capabilities in realistic multi-round, multi-party dialogues. It comprises 1,607 synthesized conversation instances (covering 1–4 rounds and 2–4 participants) built via a tool dependency graph and a multi-agent persona system, validated through automated, rule-based, and human filtering stages. Jang et al. 2025
CharacterCraft Chinese Text Text (multi-turn role-playing dialogues extracted from Chinese novels and revised via iterative augmentation-reconstruction) Character role-playing / persona-consistent dialogue (fictional characters from Chinese novels) Human-System 21,392 multi-turn dialogues, 121,418 utterances, 369 unique characters 5.68 CharacterCraft is a large-scale, high-quality Chinese role-playing dataset constructed by extracting character dialogues from novels using a fine-tuned dialogue extraction model, then applying an iterative augmentation-reconstruction method to reduce the literary-reality language gap. It covers 369 fictional characters across 21,392 multi-turn sessions and 121,418 utterances, and is accompanied by a reference-guided LLM-as-a-judge evaluation framework. Yin et al. 2025
InteractSpeech English Speech Audio (dual-track speech), speaker timestamps, interaction event annotations, transcripts Spoken dialogue interaction: interruptions, backchannels, turn-taking, gap and pause detection Human-System 150 hours of speech; 90K text utterances; 148h training set, 2h in-domain test set, 1h OOD test set InteractSpeech is a 150-hour English multi-turn spoken dialogue dataset designed to train and evaluate spoken dialogue models on nuanced real-time interactional phenomena such as interruptions, backchannels, gaps, and pauses. It combines 112 hours of synthetically generated dual-track speech (produced via LLM-based text generation and advanced TTS) with 38 hours of filtered real-world conversational data, all annotated with precise speaker timestamps and over 10 types of interaction events. Chen et al. 2025
AuraDial Chinese Text Text (single-turn dialogues and multi-turn dialogue sessions) AI psychological counseling / mental health support Human-System 399K samples total (300K+ single-turn dialogues, 90K+ multi-turn dialogue sessions); 447M total characters; avg. instruction length 252.2 chars, avg. response length 654.2 chars AuraDial is a large-scale, human-centric Chinese dialogue dataset for AI psychological counseling, comprising over 300,000 single-turn dialogues and 90,000 multi-turn dialogue sessions. Instructions are primarily sourced from real-world user queries on public Chinese counseling platforms, and counselor responses are generated via a novel rephrasing-based synthesis pipeline designed to produce empathetic, human-like replies. Zhang et al. 2025
CATCH Counseling Dialogue Dataset Chinese Text Synthesized multi-turn counseling dialogues with turn-level chain-of-thought (Memory-Driven Dynamic Planning CoT) annotations Mental health counseling (Single-Session Therapy, SST) Human-System (client self-reports from Yixinli platform; dialogues synthesized via LLM-based multi-agent framework) 233 multi-turn counseling dialogues, 6,898 response entries with MDP CoT A synthesized dataset of multi-turn mental health counseling dialogues generated using the CATCH framework, which applies a Progressive Dialogue Synthesis strategy grounded in Single-Session Therapy (SST) principles. Each counselor turn is annotated with an explicit Memory-Driven Dynamic Planning (MDP) chain-of-thought capturing memory enhancement, global planning, and strategy reasoning, derived from client self-reports collected from the Yixinli platform. Chen et al. 2025
RealCBT English Text (transcripts from video) Transcripts of video-recorded CBT counseling sessions, with metadata annotations (client problem, client gender, client attitude) Cognitive Behavioral Therapy (CBT) counseling sessions Human-Human (counselor–client dyads) 76 dialogues; 190,714 total words (82,436 client words, 108,278 counselor words); avg. 2,516 words per session; total duration 1,224.67 minutes RealCBT is a dataset of 76 authentic Cognitive Behavioral Therapy (CBT) dialogue transcripts collected from publicly available videos on YouTube and Vimeo, manually reviewed and corrected, and annotated with client problem, gender, and attitude metadata. It is released to support research on emotional dynamics and the evaluation of synthetic therapy data. Wang et al. 2025
ASD-iLLM-8k Mandarin Chinese Speech (audio recordings), Text (transcripts) Audio recordings (WAV, 16kHz) and multi-turn dialogue transcripts derived via automatic and manual transcription, with annotated child unresponsive states Clinical autism intervention (Applied Behavior Analysis topic dialogue intervention for autistic children) Human-Human (20 clinicians and 74 autistic children across 6 treatment centers) 8,035 multi-turn topic dialogues (287 real + 7,748 GPT-4.1-synthesised); derived from 64.2 hours of audio; 100-dialogue held-out test set 13.55 (doctor turns per dialogue); 10.17 (child turns per dialogue) ASD-iLLM-8k is the first publicly available Chinese multi-turn dialogue dataset for clinical autism intervention, constructed from 64.2 hours of real clinical recordings collected at six treatment centres involving 20 clinicians and 74 autistic children, then augmented with GPT-4.1-synthesised dialogues across 27 subtopics. Dialogues follow Applied Behavior Analysis (ABA) principles and include annotated child unresponsive states, covering 10 main intervention topic areas such as self-care, social interaction, and cognition. Lai et al. 2025
MDSEval English Multimodal (text and image) Text dialogues, images, generated summaries, human quality judgments Multimodal dialogue summarization meta-evaluation Human-Human 198 dialogues, 990 summaries (5 per dialogue), human annotations across 8 quality aspects 17.1 MDSEval is the first meta-evaluation benchmark for Multimodal Dialogue Summarization (MDS), consisting of 198 image-sharing dialogues curated from PhotoChat and DialogCC, each paired with five MLLM-generated summaries and human judgments across eight quality dimensions (including multimodal coherence, coverage, faithfulness, and topic progression). Dialogues are selected using a novel Mutually Exclusive Key Information (MEKI) filtering criterion to ensure genuine cross-modal summarization challenge. Liu et al. 2025
MINDS (Multilingual Interactions with Norm-Driven Speech) Mandarin-English and Spanish-English (bilingual) Text Transcripts of bilingual dialogues annotated for social norm category and adherence/violation status Cross-cultural social norm classification and adherence detection in multi-turn dialogue Human-Human (two-party bilingual: one foreign-language speaker, one English speaker) 31 dialogue sessions, 835 utterances MINDS is a bilingual, multi-annotated dialogue corpus comprising 31 multi-turn conversational sessions across Mandarin-English (16 sessions) and Spanish-English (15 sessions) speaker pairs. Each utterance is annotated for social norm category and adherence/violation status by multiple human annotators, enabling cross-cultural and realistic norm expression modeling. Sahu et al. 2025
SEER English Text Text transcripts with span-level emotion evidence annotations, sentence-level emotion category labels, and valence labels Emotion evidence detection; identifying text spans that express emotion in real-world spoken discourse Human annotation of existing speech corpora (MSP-Podcast and MuSE transcripts) 1200 sentences total: 200 single sentences (Task 1) and 200 passages of 5 consecutive sentences / 1000 sentences (Task 2) SEER (Span-based Emotion Evidence Retrieval) is a benchmark of 1,200 real-world sentences with new span-level emotion evidence annotations, sentence-level categorical emotion labels, and valence labels, derived from transcripts of the MSP-Podcast and MuSE corpora. It supports two tasks: single-sentence emotion evidence identification (200 sentences) and multi-sentence emotion evidence identification across 5-sentence passages (200 passages, 1,000 sentences). Sampath et al. 2025
SQLWOZ English Text Text (dialogues with SQL-based dialogue state annotations and API call logs) Task-oriented dialogue for travel guidance (restaurant, hotel, attraction, taxi, train booking) with complex user requirements Human-System (LLM-simulated user and dialogue agent) 22,955 dialogues, 294,214 turns, 96,048 API calls; split into train/dev/test (18,365/2,295/2,295) 12.8 SQLWOZ is a task-oriented dialogue dataset built on the MultiWOZ ontology (5 domains, 30 slots) in which user requirements are represented as SQL statements rather than slot-value pairs, enabling four categories of complex constraints: multiple values, excluded values, preferred/prioritized values, and conditional values. Dialogues are generated automatically via GPT-4o-based user and agent simulators and validated for goal fulfilment and SQL correctness. Xu et al. 2025
3MDBench English Multimodal (text and image) Medical images, generated textual symptom descriptions, simulated multi-turn doctor-patient dialogues Medical telemedicine consultation and diagnosis across 34 diagnoses in 5 medical domains (e.g., dermatology, throat/mucosae) Human-System (multi-agent: LLM-based Doctor Agent, temperament-driven Patient Agent, Assessor Agent) 2,996 cases (images with associated textual complaints); dialogues capped at 28 utterances; 34 diagnoses across 5 domains ~14–15 utterances per dialogue (varies by model and temperament; e.g., 13.32–17.48 average utterances reported) 3MDBench is an open-source benchmark for simulating and evaluating Large Vision-Language Model (LVLM)-driven telemedicine consultations. It comprises 2,996 multimodal cases (medical images paired with generated textual symptom descriptions) across 34 diagnoses, featuring a temperament-driven Patient Agent (sanguine, choleric, melancholic, phlegmatic) and an Assessor Agent that evaluates both diagnostic accuracy and consultation/communication quality via adapted Mini-CEX criteria. Sviridov et al. 2025
DeepWell-Adol Mandarin Chinese Text Text (multi-turn dialogues: human expert-written and automatically generated) Adolescent positive mental health and wellbeing promotion (emotion regulation, academic & career development, social & interpersonal relationships, lifestyle & environmental adaptation, personal growth & self-identity) Human-WoZ (human expert-written seed dialogues; LLM-generated coach–adolescent dialogues) 1,795 multi-turn dialogues (925 expert-written + 870 computer-generated); 1,337 used for fine-tuning after format filtering 6.88 (expert-written); 13.18 (computer-generated) DeepWell-Adol is a domain-specific Chinese multi-turn dialogue corpus grounded in positive psychology and coaching, designed to promote positive mental health and wellbeing among adolescents. It comprises 925 human expert-written seed dialogues and 870 automatically generated dialogues produced via a two-stage scenario-based augmentation framework (DeepSynergy), covering five adolescent mental health themes. Qiu et al. 2025
MSE Conversational Dataset English Text Text (simulated doctor-patient conversation transcripts derived from written questionnaire responses, with human-generated reference summaries) Mental health assessment (Mental State Examination); dialogue summarization Human-System (participant responses to a structured 12-item MSE questionnaire transformed into simulated doctor-patient dialogues) 405 dialogues, 9720 utterances 24 turns per dialogue (12 doctor questions + 12 patient responses) A dataset of 405 simulated doctor-patient conversations derived from a 12-item Mental State Examination (MSE) questionnaire administered to university students, covering diverse mental health aspects such as mood, social life, memory, and stress. Each conversation is paired with a human-generated reference summary, supporting research on automated mental health assessment and dialogue summarization. Sahu et al. 2025
MMD-Eval (Multi-turn Medical Dialogue Evaluation) Chinese Text Structured medical records, annotated doctor–patient dialogue turns (intent and dialogue-state labels), multi-turn doctor–patient dialogues generated via a task-oriented dialogue system interacting with medical LLMs Medical consultation / clinical diagnosis Human-System (task-oriented dialogue system simulating patients; medical LLMs acting as doctors) 2,636 structured medical records; ~20,000 annotated sentences for dialogue system training (9,176 training + 2,294 validation for intent recognition; 7,576 training + 1,894 validation for slot filling); 1,000 professional-doctor-annotated test instances MMD-Eval is an interactive evaluation benchmark for assessing the proactive communication and diagnostic capabilities of medical LLMs via multi-turn simulated doctor–patient consultations. It pairs a task-oriented dialogue system (trained to act as a patient) with 2,636 structured medical records spanning multiple clinical departments, enabling automatic generation of multi-turn dialogue data and evaluation along dimensions of communication competence and clinical diagnostic competence. Liu et al. 2025
CoPrUS-MultiWOZ English Text Text (synthetically augmented dialogue transcripts with miscommunication and repair utterances) Task-oriented dialogue (multi-domain booking: hotel, restaurant, train, attraction, taxi) Human-WOZ (base data), synthetically augmented via LLM ~1,900 modified dialogues (18% of MultiWOZ 2.1) CoPrUS-MultiWOZ is an augmented version of the MultiWOZ 2.1 task-oriented dialogue dataset in which nearly 1,900 dialogues have been post-hoc enriched with synthetic miscommunication turns (misunderstandings, non-understandings, and vaguely related questions) and corresponding repair utterances generated via a two-step LLM prompting pipeline (CoPrUS), aiming to make benchmark dialogues more realistic by going beyond the “happy path.” Steindl et al. 2025
CRISP Bilingual (English and Chinese) Text Text dialogues with sentence-level supportive strategy annotations and cognitive distortion type labels Cognitive Restructuring (CR) psychotherapy — identifying and restructuring negative thoughts/cognitive distortions arising from mental health issues across 10 categories and 54 sub-categories of situations Human-System (LLM self-play simulating therapist and help-seeker, distilled from GPT-4o) 22,063 dialogues, 796K+ utterances 36.48 CRISP is a large-scale, high-quality bilingual (English and Chinese) dialogue dataset for Cognitive Restructuring (CR) psychotherapy, distilled from GPT-4o using the CRDial framework. It contains 22,063 multi-stage multi-turn supportive dialogues with fine-grained sentence-level supportive strategy annotations and cognitive distortion type labels, covering 54 sub-categories of mental health situations, designed to train conversational LLMs for CR-based psychotherapy. Zhou et al. 2025
PoSum-Bench English, French Text Text (conversation transcripts with LLM-generated summaries and positional bias annotations) Conversational summarization (formal meetings, casual dialogues, customer service interactions) Human-Human 2,773 dialogues (2,273 English, 500 French); English: ~1,505,323 words, ~43,893 turns; French: ~198,500 words, ~26,500 turns Varies by subset: ICSI 166, MeetingBank 310, DialogueSUM 10, QMSUM 46, SummEdits 36, TweetSum 5, DECODA (FR) 53 PoSum-Bench is a bilingual (English and French) benchmark for evaluating positional bias in LLM-based conversational summarization, aggregating and curating 2,773 dialogues from six English corpora (ICSI, QMSum, DialogueSUM, MeetingBank, SummEdits, TweetSum) and one French corpus (DECODA), spanning formal meetings, casual conversations, and customer service interactions. It includes LLM-generated summaries from ten instruction-tuned models and provides a novel sentence-level semantic similarity metric for reference-free quantification of leading and recency bias. Sun et al. 2025
CoALM-IT English Text Text (instruction-tuning samples covering dialogue state tracking, function/API calls, and multi-turn ReAct-style reasoning) Task-oriented dialogue, function/API calling, multi-turn conversational agents Human-System 311,583 samples, 211,184,321 tokens (SNIPS: 13,028 samples; Hammer: 13,819 samples; ToolAce: 202,500 samples; SGD ReAct/CRA: 82,236 samples) CoALM-IT is a multi-task instruction-tuning dataset combining task-oriented dialogue state tracking (SNIPS), single- and multi-turn function calling (Hammer, ToolAce), and a novel Conversational ReAct API (CRA) component derived from the SGD dataset using GPT-4o. The CRA subset is the first multi-turn TOD dataset to explicitly incorporate ReAct-style intermediate reasoning steps (Thought–Action–Observation) alongside API calls, yielding 82,236 samples across hotel booking, restaurant reservation, and other task-oriented domains. Acikgoz et al. 2025
SHARE English Text Text (dialogue transcripts extracted from movie scripts, with persona summaries, personal event summaries, mutual events, and shared memory annotations per utterance) Open-domain long-term dialogue with shared memory Human-Human (movie script character pairs; person-person) 3,216 episodes; 17,679 sessions; 119,087 utterances 6.74 utterances per session; 5.50 sessions per episode SHARE is an open-domain long-term dialogue dataset constructed from 1,201 movie scripts, containing dyadic multi-session conversations annotated with persona information, personal events, mutual events, and implicitly extractable shared memories between speakers. Over 61% of episodes contain at least one shared memory, supporting research on engaging and sustainable long-term dialogue systems. Kim et al. 2025
HiCUPID English Text Synthetic dialogues and QA pairs (GPT-4o-generated), user metadata (personas, profiles, schedules) Personalized AI assistant; open-domain conversational personalization Human-System (synthetic user–assistant dialogues) 1,500 synthetic users; 100,000 dialogues (50,000 train + 10,000 Test 1 + 10,000 Test 2 per split breakdown); Train: 1,250 users, 50,000 dialogues, 40,000 QA pairs; Test 1: 1,250 users, 10,000 QA pairs; Test 2: 250 users, 10,000 dialogues, 10,000 QA pairs; avg. dialogue history ~17,256 tokens per user 10 turns per persona dialogue; 1 turn per profile/schedule dialogue HiCUPID (Conversations with User Personal Information Dataset) is a synthetic, GPT-4o-generated benchmark for training and evaluating LLMs as personalized assistants. Each of 1,500 synthetic users is defined by 25 persona dimensions, 5 profile attributes, and 10 schedules, with personal information revealed implicitly across multi-turn dialogue histories; the benchmark includes single-info and multi-info QA pairs and a Llama-3.2-based automated evaluation model aligned with human preferences. Mok et al. 2025
BridgeKG English Text Text transcripts with conversational grounding act annotations and grounded knowledge item annotations in JSON-LD format Information-seeking dialogues across five knowledge domains: geography, history, media, nutrition, and sports Human-Human 26 dialogues, 669 turns, 250+ conversational grounding annotations, 127 grounded knowledge item annotations ~25.7 turns per dialogue BridgeKG is a dialogue corpus of 26 human information-seeking conversations spanning five knowledge domains (geography, history, media, nutrition, sports), annotated with conversational grounding acts (explicit, implicit, clarification) and grounded knowledge items represented as knowledge graph structures in JSON-LD format. It is designed to support research on conversational grounding and knowledge identification in dialogue systems. Schneider et al. 2024
Self-Emotion Dialogue Dataset English Text Text (GPT-4 generated dialogues with and without self-emotion, paired) Empathetic open-domain conversation with self-emotion blending Human-System (LLM-simulated agent pairs) 20,605 dialogues (train: 14,274 / val: 2,762 / test: 3,569), paired with and without self-emotion A paired GPT-4-generated dialogue dataset derived from the EmpatheticDialogues corpus, containing conversations both with and without self-emotion (speaker emotional states caused by out-of-context life events). Each dialogue pair shares the same conversational context but differs in whether a self-emotion condition (random event style) is provided to the responding agent. Zhang et al. 2024
Placement Game Dialogue Dataset English Text Text (chat transcripts) Collaborative 2D object placement; negotiation of goal state Human-Human 71 games (2 rounds each) A dataset of human-human dialogues collected via an online two-player 2D object placement game, in which pairs of players must negotiate—without a pre-defined target—how to arrange five movable objects identically on their respective boards. The corpus is used to study balanced vs. asymmetric collaboration strategies and associated task performance. Jeknic et al. 2024
RoleCraft-GLM Dataset Mandarin Chinese Text Text dialogues with emotion annotations and character profiles Personalized role-playing with non-celebrity, everyday personas Human-System 27,259 multi-turn dialogues; 39,422 instructions; 157,742 responses; 20 characters 14.64 A Chinese conversational dataset for personalized role-playing featuring 20 diverse, non-celebrity everyday personas, each with detailed character profiles and emotion annotations drawn from Ekman’s ten-category emotion taxonomy. Dialogues are sourced from social media interactions, film and television scripts, and customer service logs, and are designed to support emotionally nuanced, character-consistent dialogue generation. Tao et al. 2024
RecomMind Japanese Text Text (dialogue transcripts), seeker internal state annotations (knowledge and interest at entity level, first- and second-person), external knowledge annotations, questionnaire responses Movie recommendation Human-Human 1,201 dialogues, 21,014 utterances, 52,586 knowledge-annotated entities, 52,246 interest-annotated entities, 739 movies 17.5 RecomMind is a Japanese movie recommendation dialogue dataset in which both the seeker and the recommender annotate each entity mentioned in the dialogue with the seeker’s level of knowledge and interest (High/Neutral/Low) from first- and second-person perspectives, respectively. It is designed to support analysis and modeling of how seeker internal state influences recommendation success. Kodama et al. 2024
SportsVD English Multimodal (video and text) Video clips, dialogue transcripts (YouTube comments and replies) Sports event commentary; opinion-based video-grounded dialogue (basketball and football game highlights) Human-Human 5,114 videos, 39,097 dialogues, 195,460 sentences 2.23 SportsVD (Sports-domain Video-dialogue Dataset) is an event-content-oriented multimodal dialogue dataset collected from YouTube, pairing sports game highlight videos (basketball and football) with opinion-based multi-turn conversations reconstructed from user comments and replies. It is the first video-dialogue dataset focused on complex sports events and opinion-based (rather than purely fact-based Q&A) responses, with dialogues frequently requiring external knowledge about players and teams. Cheng et al. 2024
DialogCC English Multimodal (text and image) Text dialogues, Images Open-domain image-sharing social dialogue Human-Human (sourced from crowdsourced text-only dialogue datasets, images aligned automatically) 83,209 dialogues, 129,802 unique images; avg. 7.34 images/dialogue, avg. 4.77 images/utterance, avg. 8.20 utterances/dialogue 8.20 DialogCC is a high-quality, diverse multi-modal dialogue dataset constructed via a fully automatic pipeline that uses GPT-4 to infer image-sharing moments in text-only social dialogues and CLIP to align and filter relevant images from Conceptual Captions 3M. It contains substantially more images per dialogue and per utterance than existing multi-modal dialogue datasets, supporting improved generalization in image-sharing dialogue models. Lee et al. 2024
ADEA German Text Text (labeled user utterances, argument graph annotations) Argumentative dialogue on ethical issues concerning future AI applications (medical AI, legal AI, autonomous cars, AI referee) Human-System 378 dialogues, 2,880 user utterances 7.8 user turns per dialogue ADEA is a German argumentative dialogue dataset collected from two user studies in which university students interacted with an argumentative chatbot on four AI ethics topics (medical AI, legal AI, autonomous cars, AI referee). Each of the 2,880 user utterances is annotated using German argument graphs that serve as both the system knowledge base and annotation scheme, labeling argumentative units (well-founded, unfounded) and non-argumentative units (questions, miscellaneous). Hauptmann et al. 2024
MAGID English Multimodal (text and image) Synthetic text dialogues augmented with AI-generated images Open domain Human-Human 58,279 dialogues (53,071 train / 5,208 test); 84,592 images (75,654 train / 8,938 test) 8.53 (train), 11.37 (test) MAGID (Multimodal Augmented Generative Images Dialogues) is a synthetically generated multimodal dialogue dataset created by augmenting text-only dialogues (from DailyDialog, Persona-Chat, and PhotoChat) with diverse, high-quality images produced via a Stable Diffusion XL model, guided by an LLM-based scanner and a quality assurance module. The dataset was generated as a proof-of-concept release accompanying the MAGID automated pipeline framework. Aboutalebi et al. 2024
Chitchat-as-Interference (MultiWOZ + User Backstories) English Text Text (automatically augmented task-oriented dialogues with LLM-generated user backstories and system chitchat reactions) Task-oriented dialogue with chitchat interference (travel, restaurant, hotel, train booking and other MultiWOZ domains) Human-System 3,529 training examples, 458 validation examples, 488 test examples (after filtering); augmented turns average 36.7 tokens (backstory turns) and 20.06 tokens (reaction turns) An automatically augmented version of MultiWOZ 2.2 in which user turns are enriched with LLM-generated backstory chitchat (using few-shot prompting with Llama-2-70B), and corresponding system turns are prepended with supportive chitchat reactions. The dataset is designed to test and train TOD systems on inter-mode user turns that seamlessly blend chitchat and task-oriented requests. Stricker et al. 2024
BlendX English Text Text Multi-intent detection for task-oriented dialogue (airline travel, banking, general/out-of-scope, and voice command domains) Human-System 179,535 total utterances across four sub-datasets: BlendATIS (22,500), BlendSNIPS (55,853), BlendBanking77 (40,420), BlendCLINC150 (60,762) BlendX is a suite of four multi-intent detection datasets (BlendATIS, BlendSNIPS, BlendBanking77, BlendCLINC150) derived from ATIS, SNIPS, Banking77, and CLINC150, featuring more complex and diverse multi-intent utterance patterns than prior MixX datasets. Utterances combining 1–3 intents are constructed via rule-based manual heuristics and ChatGPT-based generative concatenation with a similarity-driven utterance selection strategy, supporting explicit and implicit (omissions, coreferences, gerund phrases) merging patterns. Yoon et al. 2024
LUCID English Text Text (LLM-generated dialogues with semantic labels, intent/slot annotations, and turn-level conversational phenomenon labels) Task-oriented dialogue across 13 domains and 100 intents (e.g. hotel booking, music, exercise logging, restaurant review, reminders, transportation) Human-System (LLM-simulated user and system agents) 4,277 conversations, 92,699 turns, 100 intents, 501 slots, 13 domains 21.7 LUCID (LLM-Generated Utterances for Complex and Interesting Dialogues) is a seed dataset of 4,277 task-oriented dialogues generated by a modular, automated LLM-driven pipeline across 100 intents and 13 domains. Dialogues are annotated with intent and slot labels and include nine explicitly labelled challenging conversational phenomena (e.g., sarcasm, in-turn corrections, overheard conversations, ASR early-end errors), with train, dev, and seen/unseen test splits provided. Stacey et al. 2024
MSDC (Minecraft Structured Dialogue Corpus) English Text (chat) and nonlinguistic game actions (pick and place moves) Discourse-annotated transcripts with elementary discourse units (EDUs), elementary event units (EEUs), and SDRT-style discourse relation labels; game action logs Situated collaborative construction task (Minecraft block building) Human-Human 541 dialogues; 22,552 EDUs; 32,818 EEUs (6,162 squished); 34,574 relation instances; 6,280 multi-parent discourse units 31.1 speaker turns per dialogue (mean); 53.1 discourse units per dialogue (mean) The Minecraft Structured Dialogue Corpus (MSDC) is a discourse-annotated version of the Minecraft Dialogue Corpus (MDC; Narayan-Chen et al., 2019), providing complete situated discourse structures in the style of SDRT (Segmented Discourse Representation Theory) for 541 two-party Architect–Builder dialogues. Structures include both linguistic discourse moves (EDUs) and nonlinguistic builder actions (EEUs), annotated with 16 discourse relation types by three linguists and two NLP experts. Thompson et al. 2024
EmoProgress English Text Text, emotion category annotations, appraisal dimension ratings (10-dimensional, 5-point Likert scale) Cumulative emotion progression analysis; two subsets: dream self-reports and customer service dialogues Human-Human (customer service, Wizard-of-Oz origin); Single-author narratives (dream reports) 149 dreams (890 annotated parts) and 339 customer service dialogues (2,010 annotated parts); 46 conversations (271 parts) and 45 dreams (264 parts) annotated for IAA Dreams: avg. 5.97 parts (SD 1.85); CS dialogues: avg. 5.92 parts (SD 1.7) EmoProgress is a corpus of dream self-reports and customer service dialogues annotated for cumulative emotion progression, in which annotators label emotion categories and appraisal dimensions incrementally (part-by-part) so that each label reflects the experienced emotion up to and including the currently revealed sentence or bi-turn, rather than in isolation. The corpus supports research on how individual textual units contribute to the global emotional trajectory of a discourse. Wemmer et al. 2024
Weights Task Dataset (WTD) — augmented with common ground annotations English Multimodal (speech/text, gesture, physical action, video) Speech transcripts, prosodic features (openSMILE), Gesture-AMR annotations, participant action annotations, collaborative problem-solving (CPS) indicators, common ground annotations (CGA) Collaborative problem-solving (deducing block weights using a balance scale — Weights Task) Multi-party human (triads) 10 groups; 1,822 utterances total; 271 utterances with common ground annotation An augmented version of the Weights Task Dataset (WTD) featuring co-situated triadic problem-solving dialogues annotated with Gesture-AMR (GAMR), participant actions (VoxML), prosodic features, collaborative problem-solving (CPS) indicators, and a new layer of common ground annotations (CGA) covering dialogue moves such as STATEMENT, ACCEPT, DOUBT, OBSERVATION, INFERENCE, QUESTION, and ANSWER. The dataset enables research on multimodal common ground tracking in shared, task-oriented physical environments. Khebour et al. 2024
Multimodal AMR Corpus (Speech and Gesture) English Multimodal (speech/audio and gesture/video) Video recordings, speech transcripts, gesture morphology annotations, speech AMRs, gesture AMRs, multimodal coreference/bridging relations (MS-AMR) Task-based block-building instruction (Human-Human collaborative assembly) Human-Human 21 video segments (~23 minutes total); 662 AMRs (343 speech, 319 gesture); 436 cross-modal relations (388 coreference chains, 28 set-member, 20 part-whole); 1,933 coreference mentions A multilayered annotated corpus of multimodal Abstract Meaning Representation (AMR) built on top of the EGGNOG dataset, covering 21 one-minute video segments of pairs of English-speaking participants in a block-building task. The corpus provides temporally aligned speech and gesture AMRs, gesture morphology annotations, and cross-modal coreference and bridging relations using Multi-sentence AMR, enabling fine-grained analysis of how gesture and natural language semantics interact. Lai et al. 2024
RECIPE4U English, Korean (code-mixed) Text Conversation logs, intent annotations (13 labels), self-rated satisfaction scores (5-point Likert), utterance-level essay edit histories EFL essay writing education (student–ChatGPT dialogues for essay revision) Human-System (EFL university students interacting with ChatGPT) 504 dialogues (97 single-turn, 407 multi-turn); 4,330 utterances (1,913 student, 2,417 ChatGPT); 380,364 total tokens; 16,118 unique tokens; 1,913 utterance-level essay edit history records 3.38 utterances per dialogue RECIPE4U (RECIPE for University) is a task-oriented dialogue dataset collected from a semester-long study with 212 EFL university students in South Korea who conversed with ChatGPT to revise their essays. It includes conversation logs, student utterances annotated with 13 intent labels, self-rated satisfaction scores, and utterance-level essay edit histories, supporting subtasks such as intent detection and satisfaction estimation in educational dialogue systems. Han et al. 2024
SCOUT (Situated Corpus Of Understanding Transactions) English Multimodal (text transcripts, speech, images, LIDAR maps) Manually transcribed and ASR speech, text messages, robot camera images, LIDAR maps, AMR annotations, Dialogue-AMR annotations, Dialogue Structure (TU and Relations) annotations Collaborative robot navigation and exploration (human instructs a remotely-located robot to move and gather environmental information) Human-WoZ (Commander human participant + two Wizard-of-Oz experimenters acting as Dialogue Manager and Robot Navigator) 278 dialogues, 89,056 utterances, 310,095 words, 5,785 images, 30 LIDAR maps, 569 AMR-annotated sentences, 13,663 Dialogue Structure TUs, 69,430 Dialogue Structure Relations 320 utterances per dialogue SCOUT is a multi-modal, human-robot dialogue corpus collected via Wizard-of-Oz experiments across four studies in which human Commanders gave verbal instructions to a remotely-located (physical or simulated) robot to explore and assess its environment. The corpus includes time-aligned transcripts, robot camera images, and LIDAR maps, and is annotated with Abstract Meaning Representation (AMR), Dialogue-AMR, and Dialogue Structure (Transactional Units and Relations). Lukin et al. 2024
MWoZGPT (NeutralGPT, FriendlyGPT, MultistyleGPT) English Text Text, dialogue-act annotations, slot-value annotations Restaurant search (task-oriented dialogue, restaurant domain from MultiWOZ) Human-System (LLM-generated user and system turns, with manual annotation correction) Three datasets of ~1,311 dialogues each (1,180 train + 131 test per style): NeutralGPT, FriendlyGPT (131 test only), and MultistyleGPT NeutralGPT: 6.05 avg. system turns/dialogue; FriendlyGPT: 7.28; MultistyleGPT: 5.32 Multi-style extensions of the MultiWOZ restaurant-domain dataset, generated using GPT-3.5-turbo with style-specific prompts (neutral, friendly, and mixed). Each collection is semantically annotated with dialogue acts and slot-value pairs, with test sets manually corrected, to support research on stylistic variation in task-oriented dialogue systems. Labruna et al. 2024
KCDD Korean Text Text (human-written dialogues with conversation-level crime class labels and utterance-level speaker type annotations) Violence/crime dialogue classification (Serious Threats, Extortion or Blackmail, Harassment in the Workplace, Other Harassment, Clean Dialogue) Human-Human (crowd-worker authored fictional dialogues) 22,249 dialogues, 178,991 utterances, 1,307,678 words; train/dev/test split of 17,799/2,225/2,225 8 The Korean Crime Dialogue Dataset (KCDD) is the first Korean NLP dataset for context-based violence detection, comprising 22,249 crowd-sourced dialogues categorised into four criminal classes aligned with the UN ICCS international standards (Serious Threats, Extortion or Blackmail, Harassment in the Workplace, Other Harassment) plus one Clean Dialogue class. Each dialogue is annotated at the conversation level with a crime class and at the utterance level with speaker roles (perpetrator, victim, normal person). Kim et al. 2024
ProMISe English Text Text (suggested question-answer pairs, user intent labels, user turn choices, dialogue history) Open-domain information-seeking intent resolution via proactive multi-turn suggested question-answering Human-System (human annotators simulate user choices; LLM simulates agent) 1,025 dialogues, 4,453 turns, 17,812 suggested question-answer (SQA) pairs 4.35 ProMISe is a proactive multi-turn dialogue dataset for open-domain information-seeking intent resolution, in which an LLM agent generates sets of suggested question-answer (SQA) pairs at each turn and human annotators (via MTurk) simulate user choices to progressively satisfy a predefined information-seeking intent. Dialogues are grounded in real-world trending queries from Google Trends and generated using web-retrieval-augmented LLMs with chain-of-thought prompting. Butala et al. 2024
DialogStudio English Text Text (unified dialogue transcripts with metadata, external knowledge, dialogue state annotations, intent annotations, and prompts) Multi-domain: open-domain dialogue, task-oriented dialogue, natural language understanding, conversational recommendation, dialogue summarization, knowledge-grounded dialogue Human-Human, Human-System 80+ dialogue datasets unified into a single collection DialogStudio is the largest and most diverse unified collection of publicly available dialogue datasets, aggregating more than 80 datasets spanning open-domain, task-oriented, NLU, conversational recommendation, dialogue summarization, and knowledge-grounded dialogues. All datasets are standardized into a consistent JSON format while preserving original information, and are enriched with domain-aware prompts, external knowledge, dialogue state, and intent annotations to facilitate dialogue research and instruction-aware model training. Zhang et al. 2024
PRODIGy English Text Text (movie script dialogues aligned with speaker profile annotations: MBTI personality type, binary gender, biography sentences, and character dialogue history) Open-domain profile-based dialogue generation (movie script dialogues) Human-Human (fictional movie characters) 20,850 dialogues, 80,604 turns, 339 annotated characters, 8,498 biography sentences 4 (±3.28) PRODIGy (PROfile-based DIalogue Generation) is a dataset of over 20K movie script dialogues sourced from the Cornell Movie Dialogs Corpus, enriched with diverse speaker profile representations including MBTI personality type, binary gender, character biography sentences, and implicit linguistic style captured via dialogue history. It is designed for training and evaluating open-domain dialogue agents that maintain consistent, coherent speaker profiles. Occhipinti et al. 2024
Reddit Conversational EL Dataset English Text Text (Reddit conversation threads with entity mention annotations linked to Fandom knowledge base) Zero-shot conversational entity linking; fan/entertainment domain (Fandom wikis) Human-Human 6,097 conversations, 8,771 threads, 54,252 utterances, 11,228 annotations (train: 5,352 conversations, 8,026 threads, 49,695 utterances, 10,263 annotations; test: 745 conversations, 745 threads, 4,557 utterances, 965 annotations) 6.19 (train), 6.11 (test) A conversational entity linking dataset curated from Reddit discussions on Fandom topics, where entity annotations (mention spans linked to Fandom KB entries) are derived from user-included hyperlinks to the Fandom website. The dataset is designed to evaluate zero-shot EL models in realistic conversational settings with domain-specific, long-tail entities and an unfamiliar knowledge base. Hoveyda et al. 2024
HING-POEM Hinglish (Hindi-English code-mixed) Text Text (code-mixed dialogues), utterance-level politeness labels, politeness causal span annotations, politeness intensity values Mental health and legal counseling of crime victims Human-System (victim and counseling agent) 5,000 dialogues, 129,325 utterances (Train: 2,859 dialogues / 77,806 utterances; Validation: 1,080 dialogues / 25,775 utterances; Test: 1,061 dialogues / 25,744 utterances) ~25 utterances per dialogue (Train: 27.21, Validation: 23.87, Test: 24.26) HING-POEM is a code-mixed Hinglish conversational dataset for mental health and legal counseling of crime victims, derived from the English POEM dataset by converting utterances to Hinglish using LLM-based generation and human verification. Each utterance is annotated with politeness labels (polite, neutral, impolite), politeness causal spans, and ordinal politeness intensity values, supporting the novel Politeness Cause Elicitation and Intensity Tagging (PCEIT) task. Priya et al. 2024
MEDIATOR Chinese Text Automatically annotated chain-of-thought diagnostic thought process reasoning paths for existing medical dialogue turns Medical consultation / clinical diagnosis Human-Human 407K thought processes (122K over MedDG dialogues; 285K over KaMed dialogues) ~4 reasoning steps per thought process (avg. 4.17 steps for MedDG, 4.13 for KaMed) MEDIATOR is a medical dialogue thought process corpus in which each doctor turn from the MedDG and KaMed datasets is annotated with a multi-step chain-of-thought diagnostic reasoning path, generated automatically using GPT-4. Each thought process averages ~4 steps and ~237–240 tokens, capturing the abductive and deductive reasoning a clinician uses before formulating a response. Xu et al. 2024
LEGO-MRTA English Multimodal (text and image) Text (conversations, instruction manuals), Images, Vision Question Answering pairs, XR tool responses LEGO brick assembly training in Mixed Reality (MR) environments Human-System (synthetically generated trainer–trainee dialogues via LLM) 65 instruction manuals, 1,423 conversations, 35,131 utterances, 26,405 context-response pairs, 13,994 instruction steps 24.8 LEGO-MRTA is a multimodal fine-grained assembly dialogue dataset for Mixed Reality training assistants, automatically synthesized using a commercial LLM grounded on 65 LEGO instruction manuals. It comprises 1,423 trainer–trainee conversations with vision-language pairs, XR tool responses (18 functional tools), and vision question answering pairs for LEGO brick assembly tasks. Pei et al. 2024
CAUSE (Counterfactual Augmented User Satisfaction Estimation test collections) English Text Text (counterfactual task-oriented dialogues with human-annotated binary user satisfaction labels and dialogue coherence labels) Task-oriented dialogue (multi-domain: hotel, restaurant, train booking, etc.); user satisfaction estimation Human-System (Wizard-of-Oz) 543 counterfactual dialogues (MultiWOZ CF) + 742 counterfactual dialogues (SGD CF); augmenting original test sets of 851 (MultiWOZ) and 924 (SGD) turn-level samples Counterfactual augmentations of the MultiWOZ and SGD user satisfaction estimation test collections, generated by GPT-4 and curated via human annotation. Each counterfactual sample replaces the last system utterance with one that flips the binary satisfaction label (satisfied ↔ dissatisfied), addressing the severe class imbalance in existing benchmarks and enabling robustness evaluation of user satisfaction estimators in task-oriented dialogue systems. Abolghasemi et al. 2024
StableLLAVA Synthesized Image-Dialogue Dataset English Multimodal (text and image) Synthesized images (via Stable Diffusion) and synthesized dialogues (via ChatGPT) Visual instruction tuning; covers single-image capabilities (recognition, physical attributes, anomaly detection, profession, color, etc.) and multi-image reasoning (similarity, difference, logical relations), as well as interleaved multi-turn dialogues Human-System 38K image-dialogue pairs (single-image, stage 1); 3K multi-image instances (stage 2) A synthesized visual instruction tuning dataset created by pairing ChatGPT-generated dialogues with Stable Diffusion-generated images, covering diverse single-image capabilities and multi-image reasoning tasks. The dataset is designed to reduce domain bias found in benchmark-derived datasets and to support flexible, scalable training of multimodal large language models. Li et al. 2024
SMILECHAT Mandarin Chinese Text Text (multi-turn dialogues generated by rewriting single-turn QA pairs via ChatGPT) Mental health support / psychological counseling Human-System (ChatGPT-rewritten help-seeker and supporter turns) 55,165 dialogues; 1,833,856 utterances (693,756 help-seeker, 1,140,100 supporter) 5.7 turns per dialogue (33.2 utterances per dialogue) SMILECHAT is a large-scale Chinese multi-turn dialogue dataset for mental health support, generated by the SMILE method, which prompts ChatGPT to rewrite publicly available single-turn QA pairs (from PsyQA) into multi-turn counseling conversations between a help-seeker and a supporter. The dataset covers 60 distinct mental health dialogue topics and is designed to be lifelike, diverse, and privacy-preserving. Qiu et al. 2024
FEDI English Text Text (dialogues annotated with implicit user feedback types, generation error types, user emotions, demographic information, slot values, intents, and knowledge documents) Task-oriented and document-grounded dialogue; domains include parcel shipping, SIM card top-up, access control (receptionist), and insurance/financial question answering Human-System (LLM-generated training/validation dialogues; human-human test dialogues collected by computer science students) 8,852 dialogues (1,988 feedback-free including 326 test; 6,864 feedback dialogues in four versions) 7.6 FEDI is the first English task-oriented and document-grounded dialogue dataset annotated with implicit user feedback (generation error and feedback types), user emotions (11 categories), and demographic information (gender, age, occupation, name, language style). Training and validation dialogues are LLM-generated (GPT-3.5); test dialogues were collected via human-human interaction across four service domains. Petrak et al. 2024
STARK English Multimodal (text and image) Text dialogues, synthetic images (generated via diffusion models, image retrieval, and web search), persona profiles (demographic, commonsense, narrative), temporal event sequences Long-term social multi-modal conversation with personalized image-sharing behavior Human-System 93K episodes, ~0.5M sessions, ~0.9M images 10.5 turns per session STARK is a large-scale, automatically constructed long-term multi-modal dialogue dataset featuring personalized image-sharing behavior grounded in rich social personas (demographics, commonsense knowledge, personal narratives) and covering multiple sessions with realistic time intervals. Dialogues are distilled from ChatGPT using the MCU framework with a Plan-and-Execute image aligner that sources images via text-to-image generation, retrieval, and web search. Lee et al. 2024
MLMCID-dataset Multilingual (English, Spanish, Thai) Text Text Multi-label multi-class intent detection in task-oriented dialogue Human-annotated Approx. 20,000+ instances across 10 splits (e.g., Mix-SNIPS: 11,000 train/2,197 dev/2,198 test; Mix-ATIS: 13,161/600/829; FB-EN/ES/TH: 800/100/100 each; HWU64: 780/97/97; BANKING: 1,156/144/144; CLINC: 1,353/169/169; Yahoo: 498/62/162; MPQA: 284/36/136) MLMCID-dataset is a multilingual (English, Spanish, Thai) multi-label multi-class intent detection dataset curated and re-annotated from existing benchmark NLU datasets (SNIPS, ATIS, Facebook, HWU64, BANKING, CLINC, MPQA, Yahoo). Each instance is annotated with multiple intent spans, coarse and fine-grained intent labels, and primary/non-primary intent markings, enabling joint intent span extraction and multi-intent detection in task-oriented dialogue settings. Mullick et al. 2024
CACTUS English Text Text (synthetic multi-turn counseling dialogues) Psychological counseling using Cognitive Behavioral Therapy (CBT) Human-System (LLM-simulated counselor and LLM-simulated client) 31,577 dialogues, 995,512 utterances 16.6 CACTUS (CBT-augmented Counseling Chat Corpus) is a large-scale synthetic multi-turn dialogue dataset simulating realistic psychological counseling interactions grounded in Cognitive Behavioral Therapy (CBT). Dialogues are generated by LLM-simulated counselors and clients with diverse personas and attitudes, with counselors following structured CBT technique planning before each session. Lee et al. 2024
TransferTOD Mandarin Chinese Text Text (dialogues with slot-value annotations) Multi-domain task-oriented information collection across 30 life service scenarios (e.g., hotel, food delivery, courier, sanitation, water delivery) Human-System 5,460 dialogues, 35,965 turns, across 30 domains (27 in-domain + 3 out-of-domain); 188 slots ~6.6 turns per dialogue (35,965 turns / 5,460 dialogues) TransferTOD is a Chinese multi-domain task-oriented dialogue dataset simulating system-driven human-computer information collection conversations across 30 popular life service scenarios. It is constructed via a four-step pipeline (script generation, noise injection, GPT-based diversity augmentation, and expert fluency refinement) and includes 5,460 dialogues with slot-filling annotations, designed to support proactive questioning and robust slot filling. Zhang et al. 2024
ConvPlan English Text Text (natural language conversation plans with targets, user settings, and plan descriptions distilled from existing target-driven conversation corpora) Target-driven conversational recommendation (movie recommendation) Human-Human (plans distilled from the DuRecDial human-to-human dialogues) 12K high-quality plans ConvPlan is a dataset of 12K high-quality natural language conversation plans distilled from target-driven dialogue corpora (DuRecDial) using an LLM-based two-stage framework (EnPL). Each plan includes a user setting, a target item (e.g., a movie), and a free-text plan sketch describing the conversational path toward the target, filtered for quality via entity-consistency scoring. Zheng et al. 2024
MediTOD English Text Transcripts of staged doctor-patient interviews with comprehensive intent, slot, and attribute annotations (CMAS schema), canonicalized to UMLS medical concepts Medical history taking (task-oriented dialogue); covers respiratory and musculoskeletal specialties Human-Human (staged/simulated doctor-patient interactions performed by medical professionals) 213 dialogues, 22,503 utterances (175 train / 20 validation / 18 in-domain test / 20 out-of-domain test) 96.57 utterances per dialogue MediTOD is the first publicly available English task-oriented dialogue dataset for medical history taking, annotated by medical professionals using a novel Comprehensive Medical Attribute Schema (CMAS) that captures slots (e.g., symptoms, medications) together with their attributes (e.g., onset, severity, progression), with medical values canonicalized to UMLS concepts. It supports benchmarking of NLU, policy learning, and NLG subtasks in both supervised and few-shot settings. Saley et al. 2024
PERPDSCD English Text Text (multi-turn dialogues with user profile annotations for gender, age, OCEAN personality traits, politeness levels, and empathy levels) Personalized physical disability support (covering topics such as mobility aids, home modifications, physical therapy, assistive technology, pain management, ADLs, emotional support, employment and education, social interaction, fitness, peer support, parenting with disabilities, and life transitions) Human-System 18,026 dialogues; 403,086 utterances (train: 14,421 dialogues / 313,495 utterances; val: 1,803 / 49,238; test: 1,800 / 40,353) ~22–27 utterances per dialogue (avg. 21.73 train, 27.30 val, 22.41 test) PERPDSCD (Persona-tailored Physical Disability Support Conversational Dataset) is a large-scale, GPT-3.5-generated and human-verified dataset of multi-turn dialogues between users with physical disabilities and a doctor-role system. Each dialogue is annotated with user profile information (gender, age, and OCEAN-model personality traits) and utterance-level labels for politeness (polite/neutral/impolite) and empathy (empathetic/neutral/non-empathetic), spanning 14 disability types and 13 support topics. Mishra et al. 2024
AAC Personal Narrative Conversational Dataset English Text Text (dialogue transcripts, prompt-response pairs) Personalized conversational assistance for Augmentative and Alternative Communication (AAC) users Human-System (AAC user and conversational AI partner) 511 dialogues, 4,053 utterances, 2,023 prompt-response pairs (1,423 train / 200 validation / 400 test) ~4 turns per dialogue (average 7.93 utterances per dialogue) A personalized conversational dataset centred on the life experiences and communication style of a single primary AAC user, constructed by prompting Google Gemini with authored content from the user and then refining the generated dialogues with AAC domain experts. Designed to fine-tune language models for deeply personal and contextually relevant AAC communication support. Pal et al. 2024
The Dining Llamas of Oz English Text Text (automatically generated dialogues between LLM-simulated user and system agents) Restaurant search and reservation (Cambridge domain, based on MultiWOZ) Human-System (Human-Llama phase); System-System (Llama-Llama phase) 1,311 dialogues (1,049 train, 131 validation, 131 test) 6.21 A corpus of 1,311 task-oriented dialogues generated via LLM-LLM (Llama-3 8B) interaction in the restaurant domain, based on the MultiWOZ knowledge base. Dialogues are annotated with KB-Alignment and KB-Grounding metrics to support research on LLM consistency and trustworthiness in task-oriented dialogue systems. Labruna et al. 2024
Food Salt Content Conversational Dataset English Text Text (template-generated dialogues with belief state and action state annotations) Food salt/sodium content inquiry for heart failure patient dietary management Human-System 87,425 dialogues, 525,392 turns 6 A template-based task-oriented conversational dataset designed for food-based salt content inquiries, modelled after MultiWOZ. Dialogues simulate a patient asking about sodium content in food items, with the system posing clarification questions (covering slots such as food, cook, type, animal, part, foodweight, and metric) to identify the precise food item and its salt value, sourced from the USFDC database. Tayal et al. 2024
Synthetic Arabic Medical Dialogues Arabic (Saudi Najdi dialect) Text Synthetic medical dialogue transcripts generated from clinical notes Medical consultation (doctor-patient dialogue generation from clinical notes) Human-System (simulated doctor and patient roles generated by LLMs) 207 dialogues (derived from 207 dialogue-note pairs in ACI-bench) ~50 exchanges per dialogue A synthetic Arabic medical dialogue dataset generated from English clinical notes (ACI-bench) using a multi-agent LLM pipeline (Claude-3-Opus and GPT-4). Dialogues are in the Saudi Najdi dialect and cover patient-physician consultations including chief complaints, medical history, diagnosis, and treatment plans. ALMutairi et al. 2024
DailyPersuasion English Text Text (dialogue sessions annotated with user intents, persuader strategies, and intent-to-strategy reasoning processes) Multi-domain persuasive dialogue (35 domains including science, travel, culture, marketing, history, politics, and more) Human-System (GPT-4-generated dialogues simulating persuader and user roles via third-person storytelling) 76,000 dialogue sessions across 13,000 scenarios and 35 domains, with 229,598 strategies 5.08 DailyPersuasion is the first large-scale multi-domain persuasive dialogue dataset, comprising 76,000 GPT-4-generated dialogue sessions spanning 13,000 scenarios across 35 daily-life domains. Each session is annotated with user intents, persuader strategies, and intent-to-strategy reasoning processes, enabling research on cross-domain persuasive dialogue systems. Jin et al. 2024
HealMe Psychotherapy Dialogue Dataset English Text Text (simulated multi-turn psychotherapy dialogues between AI client and AI therapist) Cognitive reframing psychotherapy (CBT-based) Human-System (AI client simulated by ChatGPT interacting with AI therapist; also small-scale real human client sessions) 1,300 cases (900 train, 100 validation, 300 test), each with 3 dialogue rounds (6 turns per case) 3 rounds (6 turns) per case A multi-turn psychotherapy dialogue dataset constructed by prompting ChatGPT to simulate both client and therapist roles based on (thinking trap, client’s thought) pairs, following a structured three-step cognitive reframing procedure. Dialogues are annotated for empathy, logical coherence, and guidance quality by domain experts and GPT-4. Xiao et al. 2024
LUAS Generated DST Dialogues English Text Text dialogues annotated with dialogue state tracking labels (slot-value pairs) Multi-domain task-oriented dialogue (attraction, hotel, restaurant, taxi, train) Human-System (GPT-4 simulated user and agent) 7,556 dialogues, 102,602 turns 13.57 A synthetically generated task-oriented dialogue dataset created via GPT-4-backed user-agent simulation (LUAS), covering 5 domains (attraction, hotel, restaurant, taxi, train) aligned with the MultiWOZ schema. Dialogues are annotated with DST slot-value labels and verified through consistency checks to reduce hallucination noise. Wang et al. 2024
DialogueMRC Mandarin Chinese Text Text (dialogue scripts with MRC question-answer pair annotations and discourse parsing annotations) Machine reading comprehension over multi-party dialogues (span extraction QA); source material is scripts from the Chinese sitcom “I Love My Family” (我爱我家) Multi-party human 705 dialogues, 24,451 utterance units, 8,305 question-answer pairs (6,654 train / 830 dev / 818 test); 7,877 answerable and 1,425 unanswerable questions ~34.7 utterances per dialogue (24,451 utterances / 705 dialogues) DialogueMRC is the first machine reading comprehension dataset targeting Chinese multi-party dialogues, built from scripts of the 120-episode sitcom “I Love My Family.” It contains 705 dialogue instances with 24,451 utterance units and 8,305 span-extraction QA pairs (including unanswerable questions), annotated via a multi-stage pipeline combining GPT-4 generation and human review, and is designed to challenge models on dynamic conversational understanding and discourse parsing. Jiang et al. 2024
ExTES English Text Text (LLM-generated dialogues with emotional support strategy annotations) Emotional support conversation Human-System (LLM-generated user and assistant turns) 11,177 dialogues, 200,393 utterances, 97,893 annotated strategy instances 18.2 utterances per dialogue ExTES is a large-scale emotional support conversation dataset generated via an iterative LLM-based expansion framework (using ChatGPT as a “counseling teacher”), covering 36 emotional support scenarios and 16 fine-grained response strategies. Each dialogue is annotated with the emotional support strategy used in each assistant turn, and the dataset was quality-checked through human review and toxicity assessment. Zheng et al. 2024
MMC (Multilingual Multiparty Coreference) English, Mandarin Chinese, Farsi Text TV show transcripts and subtitle texts with coreference annotations (gold for English, silver via annotation projection for Chinese and Farsi) Entity coreference resolution in multiparty dialogue (TV sitcom transcripts: Friends and The Big Bang Theory) Multi-party human 1,222 scenes; English: 955 train / 134 dev / 133 test scenes, 22,964 utterances, 79,034 mentions, 33,961 clusters; Chinese and Farsi splits of comparable scale via projection MMC is a large-scale multilingual multiparty coreference resolution dataset built from transcripts and subtitles of two TV sitcoms (Friends and The Big Bang Theory). It provides gold exhaustive entity coreference annotations in English and silver annotations in Chinese and Farsi created via automatic annotation projection, covering over 1,200 scenes and approximately 101 hours of content. Zheng et al. 2023
COD (Cross-lingual Outline-based Dialogue dataset) Arabic, Indonesian, Russian, Kiswahili Text Text (dialogue utterances with intent, slot, and dialogue state annotations) Task-oriented dialogue across 11 domains (Alarm, Flights, Homes, Movies, Music, Media, Banks, Payment, RideSharing, Travel, Weather) Human-Human (Wizard-of-Oz-style outline-guided native speaker dialogue writing) Dev set: 1,138 turns; Test set: 1,352 turns; covering 11 domains across 4 languages COD is a large-scale multilingual task-oriented dialogue dataset in Arabic, Indonesian, Russian, and Kiswahili, created via a novel outline-based annotation process in which domain-specific dialogue schemata are mapped to natural language outlines that guide native-speaker annotators in writing culturally localized dialogues. It supports natural language understanding (intent detection and slot labeling), dialogue state tracking, and end-to-end dialogue evaluation across 11 domains. Majewska et al. 2023
SK-TOD English Text Text (dialogue contexts, manually annotated system responses, customer reviews with aspect and sentiment annotations) Task-oriented dialogue (hotel and restaurant booking) grounded in subjective knowledge (customer reviews) Human-System 19,696 dialogue instances; 143 entities; 1,430 reviews; 8,013 review sentences; train/val/test split: 14,768 / 2,129 / 2,799 ~9.3 utterances per instance SK-TOD is a large-scale, manually annotated dataset for subjective-knowledge-based task-oriented dialogue, built by augmenting MultiWOZ with crowd-sourced customer reviews and subjective user requests in the hotel and restaurant domains. Each instance pairs a dialogue context containing a subjective knowledge-seeking user turn with a human-written system response grounded in multiple customer review snippets annotated with aspect and sentiment information. Zhao et al. 2023
BrainKT French Multimodal (audio, video, EEG, physiological signals) Audio (48kHz), Video (25fps), EEG (BioSemi ActiveTwo, 64 electrodes, 2048Hz), Physiological signals (Empatica E4: BVP, EDA, IBI, HR, skin temperature, accelerometer), transcripts, morpho-syntactic labels, facial landmarks, gaze and head movement annotations, post-experiment questionnaires Common ground instantiation and information exchange; includes a collaborative video game task (Keep Talking and Nobody Explodes) and a free conversation task (moral dilemma + open topic) Human-Human (dyads) 28 dyads (56 participants), ~14 hours total, ~60K words (game task), ~75K words (free conversation) BrainKT is a naturalistic French conversational corpus of 28 dyads (56 participants, ~14 hours) recorded with synchronised audio, video, 64-channel EEG, and Empatica E4 physiological signals, designed to study information exchange and common ground instantiation. Sessions comprised a collaborative video-game task (Keep Talking and Nobody Explodes) followed by a free conversation task (moral dilemma then open topic), and are annotated with transcripts, part-of-speech tags, facial landmarks, gaze/head movements, and conversation themes. Maës et al. 2023
DiaBiz.Kom Polish Text (transcripts of telephone calls) Transcripts, dialogue act annotations (communicative functions, dimensions, functional and dependence relations) Business call-centre interactions (varied business settings) Human-Human 1,100 dialogues; 1,277,965 tokens; 151,520 final annotations for communicative functions DiaBiz.Kom is a Polish dialogue act corpus comprising 1,100 telephone conversation transcripts drawn from the DiaBiz call-centre corpus and annotated according to the ISO 24617-2:2012 standard. Each dialogue is annotated by two independent annotators and a super-annotator, covering communicative functions, dimensions, and functional/dependence relations, and is released under a CC BY-NC-ND 4.0 licence. Hwaszcz et al. 2023
OptiMouse-Quest English Text Synthetic dialogues (LLM-generated), human annotations Goal-oriented information elicitation for linear programming problem formulation Human-System (simulated via dual LLM agents: Question Generation Agent and Question Answering Agent) 476 dialogues, 9,480 turns; 28 dialogues with manual human annotations 20 A synthetic dialogue dataset generated by a dual-agent LLM setup (GPT-4), in which a Question Generation Agent elicits information from a Question Answering Agent to reconstruct linear programming problem descriptions sourced from the NL4Opt dataset. A subset of 28 dialogues has been manually annotated by human evaluators. Abdullin et al. 2023
MAIA-DQE Multilingual (German, Brazilian Portuguese, European Portuguese; agent side in English) Text Text (source and machine-translated turns), emotion annotations (8-class, sentence-level), dialogue quality annotations (sentence-, turn-, and dialogue-level) Customer support Human-Human 612 dialogues, 24,960 sentences, 690,146 tokens 40 sentences per dialogue (average) MAIA-DQE extends the MAIA bilingual customer support corpus (originally released for the WMT22 Chat shared task) with holistic emotion and dialogue quality annotations at sentence, turn, and dialogue levels. Annotations cover 612 dialogues across German and Portuguese (Brazilian and European) customer languages, with an 8-class emotion scheme and multiple quality dimensions including Interaction Quality and Task Success. Mendonça et al. 2023
PersonalityChat English Text Text (synthetic dialogues) Persona- and personality-grounded open-domain social conversation Human-System (ChatGPT-generated single-agent dialogues) 10,907 dialogues 17.3 PersonalityChat is a synthetic open-domain conversational dataset distilled from ChatGPT, conditioned on both persona facts (drawn from PersonaChat) and Big-5 personality trait labels. It provides a parallel, one-to-one counterpart to PersonaChat for studying trait-based personalization of dialogue models. The authors also release PersonaTraits, a companion dataset of ChatGPT-generated Big-5 personality trait speculations for 5,156 PersonaChat personas. Lotfi et al. 2023
RED (Reddit Emotional Distress) English Text Text (Reddit peer support dialogues) Emotional distress peer support / empathetic response generation Human-Human 1,275,486 dialogues, 3,396,476 turns, ~88.3M tokens 2.66 RED is a large-scale dyadic dialogue dataset scraped from 8 distress-related subreddits (e.g., r/depression, r/SuicideWatch) via the Pushshift API, containing approximately 1.3M peer support conversations spanning more than 4,000 distress-related topic clusters. Dialogues are anonymised and filtered for profanity, and the dataset is annotated with emotion and empathetic response intent labels to support the development of empathetic chatbots for distress support. Yeh et al. 2023
BanglaNLG Multi-turn Dialogue Dataset Bangla (Bengali) Text Text (machine-translated training set; human-translated evaluation sets) Open-domain multi-turn dialogue generation (translated from DailyDialog; covers everyday social situations) Human-Human 89,761 dialogues total (76,052 train / 7,069 dev / 6,640 test) A Bangla multi-turn dialogue dataset created by translating the English DailyDialog corpus, with training data translated automatically and evaluation sets translated by expert human translators. It is the first public dialogue generation dataset for Bangla, released as part of the BanglaNLG benchmark. Bhattacharjee et al. 2023
VIRADialogs English Text Text (dialogue transcripts, user feedback, predicted intents, dialog acts, offensive language predictions) COVID-19 vaccine hesitancy information and Q&A Human-System 8,088 dialogues, 28,202 total turns (20,304 free-text turns excluding feedback turns) 3.5 VIRADialogs is a dataset of real-world conversations between users and VIRA, a dialogue system addressing COVID-19 vaccine hesitancy, collected from July 2021 to May 2022. The dataset includes full dialogues, user feedback, predicted intents, dialog acts, and offensive language predictions, and is anonymized; it is intended as a benchmark for intent discovery in a rapidly evolving domain. Gretz et al. 2023
CUDON Mandarin Chinese Text Text (natural language utterances with function expression annotations) Task-oriented dialogue semantic parsing; financial assistant domain spanning 7 domains and 52 intents (e.g., fund search, stock evaluation) Human-System (semi-synthetic: agenda-based simulator with crowd-worker paraphrasing) 9,996 dialogues; 8 train/dev/test splits with varying compound divergence; ~206,531 total data samples across all splits (train sizes ranging from ~132K–196K samples per split) 41 turns per dialogue CUDON (Chinese dialogUe Dataset for compOsitional geNeralization) is a large-scale, semi-synthetic, multi-turn cross-domain task-oriented dialogue dataset in Chinese, designed to evaluate compositional generalization of semantic parsers. It contains ~10,000 dialogues annotated with function expressions and provides 8 train/dev/test splits with varying degrees of train-test distribution divergence (IID, 6 TMCD, and Length splits). Zheng et al. 2023
NewsDialogues Mandarin Chinese Text Text (dialogues with annotations for target topics, dialog acts, and knowledge spans) Proactive news-grounded conversation (hot news topics; both information-seeking and chit-chat scenarios) Human-Human 1,000 dialogues, 14,591 utterances (6,847 user + 7,744 agent); 1,000 news articles 14.59 turns per dialogue NewsDialogues is a human-to-human Chinese dialogue dataset grounded in hot news articles, designed for the Proactive News Grounded Conversation task. It contains 1K conversations with 14.6K utterances and rich annotations including key topics, dialog acts (Chit-chat, Inform, Guide), and knowledge spans, covering both information-seeking and chit-chat scenarios. Li et al. 2023
Multi3NLU++ Multilingual (English, Spanish, Marathi, Turkish, Amharic) Text Text (utterances with intent and slot annotations, manually translated) Intent detection and slot labelling for task-oriented dialogue; Banking and Hotels domains Human-System 3,080 utterances per language (5 languages); 62 intents, 17 slot types Multi3NLU++ is a multilingual, multi-intent, multi-domain NLU dataset for task-oriented dialogue, extending the English NLU++ dataset with expert manual translations into Spanish, Marathi, Turkish, and Amharic across Banking and Hotels domains. It supports multi-label intent detection and slot labelling benchmarking across high-, medium-, and low-resource languages. Moghe et al. 2023
X-RiSAWOZ Multilingual (English, French, Hindi, Korean, code-mixed English-Hindi) Text Text (dialogue utterances, dialogue state annotations, dialogue acts, database results/API calls) Multi-domain task-oriented dialogue (12 domains including TV, hotel, restaurant, tourist attractions, etc.) Human-Human (Wizard-of-Oz) 18,000+ human-verified utterances per language; 11,200 dialogues and 151,982 turns in the underlying RiSAWOZ source; per-language splits: 100 few-shot dialogues (1,318 utterances), 600 validation dialogues (8,116 utterances), 600 test dialogues (9,286 utterances) X-RiSAWOZ is a large-scale, high-quality, end-to-end multilingual task-oriented dialogue benchmark created by translating the Chinese RiSAWOZ dataset into English, French, Hindi, Korean, and code-mixed English-Hindi, using a methodology combining neural machine translation, hybrid entity alignment, and manual post-editing. It covers 12 domains and includes full end-to-end annotations (dialogue state, dialogue acts, API calls, database results) with 18,000+ human-verified utterances per language, supporting both zero-shot and few-shot agent training. Moradshahi et al. 2023
Werewolf Among Us (Persuasion in Social Deduction Games Dataset) English Multimodal (Text, Video, Audio) Dialogue transcripts, video recordings, utterance-level persuasion strategy annotations, game-level voting outcome annotations Persuasion strategy modeling in multi-player social deduction games (One Night Ultimate Werewolf and The Resistance: Avalon) Multi-party human 199 dialogue transcriptions and videos, 26,647 utterance-level persuasion strategy annotations; Ego4D subset: 5,815 utterances (48 games); YouTube subset: 20,832 utterances (151 game clips) A multimodal dataset for modeling persuasion behaviors in naturalistic multi-player social deduction games, comprising 199 video recordings and dialogue transcriptions sourced from the Ego4D Social dataset and YouTube. The dataset includes 26,647 utterance-level annotations across six persuasion strategies and game-level voting outcome annotations, enabling research on persuasion strategy prediction and social deduction outcome modeling. Lai et al. 2023
NatCS English (EN-US) Speech (recorded and transcribed) and Text (self-written spoken-style dialogues) Transcripts, audio recordings (NATCS_SPOKE), written spoken-style dialogues (NATCS_SELF), dialogue act annotations, intent and slot annotations Customer service / support (Banking, Finance, Health, Travel, Insurance) Human-Human 6,934 dialogues total: NATCS_SELF Insurance (954), NATCS_SPOKE Banking (980), Finance (3,000), Health (1,000), Travel (1,000) 59.6–72.1 turns/dialogue depending on domain (Banking: 59.6, Finance: 65.6, Health: 67.0, Travel: 72.1, Insurance: 70.6) NatCS is a multi-domain collection of spoken and spoken-style customer service conversations in English, comprising two sub-collections: NATCS_SPOKE (pairs of participants recorded and transcribed across Banking, Finance, Health, and Travel domains) and NATCS_SELF (self-written spoken-form dialogues for Insurance). A subset is annotated with task-oriented dialogue acts (InformIntent, ElicitSlot, etc.) and open-schema intent/slot labels, designed to better approximate real human-to-human customer support interactions than existing TOD datasets. Gung et al. 2023
CausalDialogue English Text Text (expert-written scripts and crowd-sourced utterances) Open-domain chit-chat dialogue with utterance-level causal structure (branching conversations in a directed acyclic graph) Human-Human 2,322 dialogues, 4,866 branches, 46,109 utterances, 51 speakers; train/validation/test split: 3,457/741/715 dialogues 26.8 utterances per dialogue (overall); 17.0 (Ori.-2S), 51.4 (Multi), 5.6 (Expansion) CausalDialogue is a chit-chat dialogue dataset structured as conversational directed acyclic graphs (DAGs), enabling the study of utterance-level causality (branch-splitting and branch-colliding). It combines expert-written scripts from the role-playing game Fire Emblem: Three Houses with crowd-sourced expansions collected via Amazon Mechanical Turk, and includes rich speaker profiles and situational metadata. Tuan et al. 2023
CLASS Datasets (Scaffolding & Conversational) English Text Synthetic scaffolding data (problems, subproblems, hints, incorrect responses, feedback) and synthetic conversational student-tutor dialogues Intelligent tutoring / introductory college-level biology education Human-System (simulated student–AI Tutorbot, generated by GPT-4) Scaffolding dataset: 648 problems, 2,198 subproblems; Conversational dataset: 648 conversations, ~20K student-tutorbot interactions ~30 turns per conversation (approx., based on 648 conversations and ~20K interactions) Two synthetic datasets created via GPT-4 under the CLASS (Conversational Learning with Analytical Step-by-Step Strategies) framework for training intelligent tutoring systems in college-level biology. The scaffolding dataset contains challenging biology problems with decomposed subproblems, hints, incorrect student responses, and feedback; the conversational dataset contains simulated student–Tutorbot dialogues that apply these scaffolding strategies in natural language interactions. Sonkar et al. 2023
RefGPT-Fact and RefGPT-Code English, Mandarin Chinese Text Text (GPT-4-generated multi-turn dialogues) Factual knowledge (RefGPT-Fact); Programming/code discussion, creation, and bug fixing (RefGPT-Code) Human-System 176K dialogues total: RefGPT-Fact — 100K dialogues (50K English, 50K Chinese); RefGPT-Code — 76K dialogues (37K English, 39K Chinese) ~4 turns (RefGPT-Fact and RefGPT-Code-cr); ~4 turns (RefGPT-Code-ds and RefGPT-Code-bg) RefGPT-Fact and RefGPT-Code are two large multi-turn dialogue datasets generated by GPT-4 using the RefGPT method, which grounds generation in external references (Wikipedia/Baidu Baike and GitHub code repositories, respectively) to minimize hallucination. RefGPT-Fact contains 100K English and Chinese dialogues on factual knowledge; RefGPT-Code contains 76K English and Chinese dialogues covering code discussion, creation, and bug fixing across multiple programming languages. Yang et al. 2023
SEAME-C and ASCEND-C Chinese-English (code-switching) Text ASR transcripts (reference and ASR system output pairs) with error annotations Code-switching ASR error correction Human-System SEAME-C: 297 dialogues, 54,698 sentences, 901,385 tokens; ASCEND-C: 49 dialogues, 10,455 sentences, 131,271 tokens SEAME-C: ~184 sentences/dialogue; ASCEND-C: ~213 sentences/dialogue Two Chinese-English code-switching ASR error correction datasets derived from the SEAME and ASCEND speech corpora, covering bilingual speakers from Singapore, Malaysia, and Hong Kong. Each dataset consists of paired ASR system outputs and manual transcriptions, annotated with four error types (redundant, missing, word selection, word ordering) using an automatic annotator. Wan et al. 2023
KBP (Knowledge Behind Persona) Mandarin Chinese Text Text (dialogues with persona descriptions, persona-related knowledge from knowledge bases, and grounding source labels) Personalized knowledge-grounded open-domain dialogue, with explicit dependency between persona and implicit knowledge Human-Human (single-person setup: one annotator plays both user and system roles) 2,477 dialogues, 24,554 utterances; Train: 1,981 dialogues / 9,821 samples; Valid: 248 dialogues / 1,227 samples; Test: 248 dialogues / 1,229 samples 4.96 KBP (Knowledge Behind Persona) is a Chinese personalized knowledge-grounded dialogue dataset in which system responses are conditioned on persona descriptions and persona-related knowledge retrieved from Chinese knowledge bases (Baike and Ownthink). It is the first dataset to explicitly model the dependency between persona and implicit knowledge, with grounding source labels (NULL, PERSONA, or PERSONA+DOCUMENTS) provided for each response. Wang et al. 2023
Multi-User MultiWOZ English Text Text (dialogue transcripts with dialogue state annotations, system acts, system responses, and ground-truth query rewrites) Task-oriented dialogue (hotel, attraction, restaurant, train, taxi, bus, police, hospital booking/information) Human-System (two users + one agent; multi-party human-system) 16,706 multi-user chats; train: 8,859 chats across 2,284 dialogues; dev: 3,936 chats across 995 dialogues; test: 3,911 chats across 994 dialogues 11 turns per dialogue (2.7–2.8x MultiWOZ 2.2) Multi-User MultiWOZ is an extension of MultiWOZ 2.2 to multi-party task-oriented dialogues, where each user utterance is replaced by a collaborative chat between two users that preserves the original dialogue state and system response. The dataset supports research on multi-user dialogue systems and the novel task of multi-user contextual query rewriting, and captures social dynamics such as slot elicitation, social chatter, and deliberation. Jo et al. 2023
GapChat English Text Text (dialogue transcripts with event timelines, session gap annotations, and event progress labels) Multi-session open-domain chit-chat grounded in simulated life event timelines with explicit time gaps between sessions Human-Human 650 dialogues, 2,650 sessions, 56,254 utterances ≥20 utterances per session (minimum enforced during collection) GapChat is a multi-session dialogue dataset in which the time gap between each session varies (from minutes to a year), and conversations are grounded in procedurally generated timelines of simulated life events and real-world news events. It extends the MSC dataset format with explicit, realistic long-term temporal structure and annotated event progress, enabling research into time-aware long-term dialogue generation. Zhang et al. 2023
Avalon-NLU English Text (chat + structured game state) Text chat transcripts, game state records, hand-annotated persuasion strategy labels, deception strategy labels, player belief annotations Social deduction game (Avalon: The Resistance) — long-horizon multi-party deception and persuasion Multi-party human (6 players per game) 20 games, 2,384 utterances, 30 unique players, 19 unique team compositions ~119 utterances per game A benchmark dataset of 20 complete human-played games of Avalon: The Resistance, comprising 2,384 chat utterances from six-player cooperative-competitive sessions. Each utterance is hand-annotated with persuasion strategies, deception strategies (for evil players), player role beliefs, and ground-truth game state, supporting research on long-horizon multi-party dialogue understanding involving deception and persuasion. Stepputtis et al. 2023
TopDial English Text Text (LLM-synthesized multi-turn dialogues with user profiles, personality descriptors, domain knowledge triples, and dialogue act–topic targets) Personalized target-oriented proactive dialogue; domains include movies, music, food, and point-of-interest (POI) restaurants Human-System (simulated via LLM role-playing agents: user agent, system agent, moderator agent) 18,009 dialogues, 202,734 utterances (train 141,928 / valid 20,310 / test 40,496); 501 unique targets 12.3 utterances per dialogue TopDial is a large-scale, LLM-synthesized dataset for personalized target-oriented proactive dialogue, where each dialogue is grounded in a user profile, Big-5 personality traits, domain knowledge triples, and a predefined ⟨dialogue act, topic⟩ target. It covers four domains (movies, music, food, POIs) and was constructed automatically via a ChatGPT-based role-playing framework involving user, system, and moderator agents. Wang et al. 2023
ScreenEval English Text Text (TV scripts/dialogues, human- and model-generated summaries, human annotations for factual consistency and relevant utterance identification) Factual inconsistency detection in long-form dialogue summarization (TV screenplays) Human-System 52 dialogues, 624 summary sentences (455 model-generated, remainder human-written); avg. 6,073 tokens per dialogue; 168 factually consistent and 58 factually inconsistent sentences annotated 309 utterances per dialogue (average) ScreenEval is a dataset for evaluating factual inconsistency detection in long-form TV-script dialogues. It pairs 52 long TV scripts (averaging 6,073 tokens) with human-, Longformer-, and GPT-4-generated summaries, annotated by crowdworkers for sentence-level factual consistency and relevant supporting utterances. Lattimer et al. 2023
FGAER-dia Mandarin Chinese Text Transcripts (ASR output from spoken dialogues), augmented synthetic dialogues with labeled fine-grained address entities Fine-grained address entity recognition from customer service / e-commerce spoken dialogues Human-Human 9.8K dialogue contexts, 16 fine-grained address entity types; supplementary CUCC-labeled set of 200 real-world dialogues up to 10 turns (max); average character length 167 (max 366) FGAER-dia is a labeled multi-turn spoken dialogue dataset for fine-grained address entity recognition, constructed via an ontology-based data augmentation paradigm (leveraging UR template pairs and ChatGPT) combined with noise injection to simulate real call-center scenarios. It covers 16 hierarchical address entity types (e.g., province, city, district, town, POI) and is accompanied by a small human-labeled real-world subset (CUCC-labeled, 200 dialogues). Han et al. 2023
IEMPATHIZE + TwittEmp with Empathy Intent Annotations English Text Text (utterances with empathy direction labels and expert-annotated empathy intent labels) Empathy detection in online health communities (cancer survivors and Twitter cancer topics) Human (single-turn utterances from online communities) IEMPATHIZE: 5,007 utterances; TwittEmp: 3,000 utterances; both annotated with 8 empathy intent labels (Acknowledging, Consoling, Questioning, Sympathizing, Wishing, Positive, Negative, Neutral) Expert-annotated extensions of the existing IEMPATHIZE (5,007 utterances) and TwittEmp (3,000 utterances) empathy detection datasets, enriched with 8 empathy intent labels manually assigned by 3 experts (Cohen’s kappa 88.4% and 80.0% respectively). The annotations enable joint training of empathy detection and empathy intent recognition tasks. Jiang et al. 2023
ORCHID Mandarin Chinese Speech (transcribed) ASR transcripts with manual post-correction, stance-annotated utterances, stance-specific summaries Competitive debate (multi-domain: education & profession, science & technology, philosophy & ethics, politics & law, culture & society, art & entertainment, economy & business, health & environment) Human-Human 1,218 debates, 476 unique topics, 14,133 annotated utterances, 2,436 stance-specific summaries 11.5 utterances per debate ORCHID (Oral Chinese Debate) is the first Chinese dataset for benchmarking target-independent stance detection and argumentative dialogue summarization. It consists of 1,218 real-world competitive debates in Mandarin Chinese covering 476 unique topics, with 14,133 stance-annotated utterances and 2,436 stance-specific summaries derived from ASR-transcribed and manually corrected debate videos. Zhao et al. 2023
ReSee-WoW and ReSee-DD English Multimodal (text and image) Text (dialogue turns, extracted entities), Images (turn-level and entity-level) Knowledge-grounded conversation (ReSee-WoW); Daily conversation (ReSee-DD) Human-Human ReSee-WoW: 22.3K dialogues, 100.4K utterances; ReSee-DD: 13.1K dialogues, 49.2K utterances ReSee-WoW: avg. 24.19 images/session; ReSee-DD: avg. 13.83 images/session Two automatically constructed multimodal dialogue datasets extended from text-only corpora (Wizard of Wikipedia and DailyDialog) by augmenting each dialogue with fine-grained visual knowledge at two granularities: turn-level images (retrieved from a large image-caption pool of ~826K pairs) and entity-level images (searched online via Qwant and Pixabay for named entities and nouns). The datasets support open-domain multimodal dialogue research where visual knowledge is explicitly split into turn-level and entity-level. Tu et al. 2023
Conversation Chronicles English Text Text (LLM-generated multi-session dialogues with time interval and speaker relationship annotations) Open-domain long-term multi-session conversation with diverse temporal intervals and fine-grained speaker relationships Human-System (LLM-generated two-speaker dialogues) 1M dialogue sessions, 200K episodes (each episode has 5 sessions), 11.7M turns 11.7 turns per session Conversation Chronicles is a large-scale English multi-session dialogue dataset of 200K episodes (1M sessions, 11.7M turns) generated via ChatGPT, incorporating diverse time intervals (a few hours to a couple of years) and 10 fine-grained speaker relationships (e.g., classmates, co-workers, husband and wife). It is designed to support research on long-term open-domain conversational AI with chronological and relational dynamics. Jang et al. 2023
RiSAWOZ-EN / RiSAWOZ-DE English, German Text Text transcripts with dialogue state annotations (slot-value pairs) Multi-domain task-oriented dialogue (attraction, restaurant, hotel, flight, train, weather, movie, TV, computer, car, hospital, courses) Human-WOZ 10,000 dialogues, 134,580 turns (training split; English and German translations of RiSAWOZ) Automatically machine-translated English and German versions of the Chinese RiSAWOZ dataset (RiSAWOZ-EN-auto and RiSAWOZ-DE-auto), created using an improved slot-value alignment method to ensure faithful translation of ontology entities without human post-editing. The datasets span 12 domains and are annotated with dialogue states for task-oriented dialogue state tracking research. Moradshahi et al. 2023
NormDial English, Mandarin Chinese (Bilingual) Text Synthetically generated dialogues with turn-by-turn norm adherence/violation labels and textual explanations Social norm adherence and violation detection across Chinese and American cultures (apology, compliment, condolence, criticism, greeting, leave, persuasion, request, response to compliment, giving thanks) Human-System (LLM-generated dyadic dialogues with human-in-the-loop validation) 4,231 dyadic dialogues, 29,550 conversational turns; 133 Chinese and 134 American expert-verified social norms NormDial is a high-quality bilingual (Chinese and English) synthetic dyadic dialogue dataset for studying social norm adherence and violation in Chinese and American cultural contexts. Dialogues are generated using a human-in-the-loop LLM prompting pipeline and annotated with turn-by-turn labels (Adhered, Violated, Not Relevant) and textual explanations grounded in expert-verified social norms across 10 norm categories. Li et al. 2023
PSYCON English Text Text (synthetic dialogues generated via GPT-J with manual intervention, annotated at dialogue-level and utterance-level) Psychotherapy / mental health support (depression, anxiety, stress, bipolar disorder, disruptive behaviour and dissocial disorders, PTSD, schizophrenia) Human-System (therapist–user, synthetically generated) 1,020 dialogues, 25,071 utterances (train: 816 dialogues / 19,568 utterances; validation: 102 / 2,692; test: 102 / 2,811) ~26 utterances per dialogue (23.98 train, 26.39 validation, 27.56 test) PSYCON is a synthetic conversational dataset for psychotherapy covering seven psychological conditions (depression, anxiety, stress, bipolar disorder, PTSD, disruptive behaviour/dissocial disorders, and schizophrenia). Dialogues are generated via GPT-J with manual quality control and annotated at two levels: dialogue-level (user gender, age, persona, therapist’s psychotherapeutic approach) and utterance-level (user sentiment; therapist politeness and interpersonal behaviour using the IPC model). Mishra et al. 2023
StatCan Dialogue Dataset English, French Text Text (live chat transcripts, data table metadata) Information seeking / table retrieval — users seeking published Statistics Canada data tables via live chat Human-Human 4,468 conversations (3,675 English, 793 French), 19,379 conversation turns, 51,872 messages (English split only); broader release covers 25,397 conversations 4.44 turns per conversation (English split) A collection of conversations sourced from live chats between online visitors of Statistics Canada’s website and StatCan agents, in which users express genuine intents to find published data tables. The dataset supports two tasks: (1) automatic retrieval of relevant tables given an ongoing conversation, and (2) automatic generation of appropriate agent responses at each turn. Lu et al. 2023
MUStARD++ with Gaze Features English Multimodal (text, audio, video, and eye-tracking/gaze) Text transcripts, audio, video, eye-tracking gaze features (fixations, saccades, regressions) Sarcasm detection in conversational settings Human annotators reading TV show dialogues (gaze recorded); underlying content is Human-Human TV dialogue 1,155 gaze-annotated samples (231 dialogue instances × 5 participants); full MUStARD++ base dataset: 1,202 instances 2–13 speaker turns per dialogue context An enrichment of the MUStARD++ multimodal conversational sarcasm dataset with eye-tracking/gaze features collected from 5 human participants over 231 dialogue instances (1,155 participant-instance samples), sourced from TV shows (Friends, The Big Bang Theory, The Golden Girls, Burnistoun, Silicon Valley). The resource includes 25 gaze features per sample (fixation durations, regression counts, saccade metrics, etc.) along with synthetic predicted gaze features for the remaining ~971 instances, enabling multimodal sarcasm detection research. Tiwari et al. 2023
ECPE-D English Text Text (dialogue transcripts with emotion-cause pair annotations and cause type labels) Emotion-Cause Pair Extraction in dialogue Human-Human 1,122 dialogues (1,106 ECPE-D-DD + 16 ECPE-D-IE); 3,370 no-context pairs, 4,161 inter-personal pairs, 2,403 self-contagion pairs (across both splits) ~10 turns (ECPE-D-DD); ~42 turns (ECPE-D-IE) ECPE-D is an English dialogue dataset reconstructed from RECCON (DailyDialog and IEMOCAP subsets) for the Emotion-Cause Pair Extraction (ECPE) task. It provides emotion-cause pair labels per dialogue along with cause type annotations (no-context, inter-personal, self-contagion), and contains substantially more emotion-cause pairs per document than existing news-article ECPE corpora. Jeong et al. 2023
PTCD (Persona-aware Topic-guiding Conversational Dataset) English Text Text (semi-automatically generated dialogues with persona profiles, concept paths, and topic-transition annotations) Persona-aware topic-guided open-domain conversation with targeted concept transitions (e.g., chit-chat to task-oriented domains such as restaurant, travel, shopping, electronics) Human-System (semi-automatic generation via GPT-J with human-in-the-loop quality checks) 2,586 dialogues, 13,746 utterances, 1,843 unique concepts, 1,738 unique concept paths 5.31 PTCD is a semi-automatically constructed English conversational dataset for persona-aware topic-guided dialogue, in which a conversational agent steers discourse toward target concepts (e.g., restaurant, travel, shopping) along ConceptNet-derived concept paths conditioned on the user’s persona profile. Dialogues are generated using few-shot prompting of GPT-J, with human-in-the-loop quality filtering, and are grounded in persona profiles drawn from the PERSONA-CHAT dataset. Ahmad et al. 2023
SafeConv Mandarin Chinese Text Text (multi-turn dialogues with utterance-level safety labels, unsafe span annotations, and safe alternative responses) Conversational safety / open-domain dialogue detoxification Human-Human 160,000 dialogues (133,153 safe responses, 26,847 unsafe responses; 148,271 safe prompts, 11,729 unsafe prompts) SafeConv is a large-scale Chinese multi-turn dialogue dataset for conversational safety research, sourced from Weibo-based corpora (LCCC-base and PchatbotW). Beyond utterance-level binary safety labels, it provides annotated unsafe spans indicating which words contribute to unsafe behavior, and context-relevant safe alternative responses to guide conversations toward safe trajectories. Zhang et al. 2023
VSTAR English Multimodal (video and text) Video clips, dialogue transcripts, scene boundary annotations, topic boundary annotations, episode metadata (genres, keywords, storylines) Video-grounded dialogue understanding and generation; scene segmentation, topic segmentation, and response generation in TV series Human-Human 185K video-grounded dialogue clips, 4.6M turns, 265K scene segments, 499K topic segments; sourced from 395 TV series across 8,159 episodes 25.1 VSTAR (Video-grounded Scene&Topic AwaRe dialogue) is a large-scale video-grounded dialogue dataset built from 395 TV series (8,159 episodes), comprising 185K 90-second multimodal dialogue clips with manually annotated scene and topic boundaries. It supports three benchmarks: video-grounded dialogue scene segmentation, topic segmentation, and response generation, with a focus on situated semantic understanding across scene and topic transitions. Wang et al. 2023
MidMed Mandarin Chinese Text Text (dialogue transcripts with dialogue type annotations and knowledge graph triples) Medical consultation covering mixed dialogue types: task-oriented diagnosis, recommendation, knowledge-grounded dialogue, QA, and chitchat; across four departments (otorhinolaryngology, ophthalmology, skin, digestive system) Human-Human 8,175 dialogues, ~98,000 utterances, 1,887,227 tokens; 229,570 knowledge graph triples 11.79 utterances per dialogue MidMed is a Chinese human-to-human mixed-type medical consultation dialogue corpus covering five dialogue types (task-oriented diagnosis, recommendation, knowledge-grounded dialogue, QA, and chitchat) across four medical departments. Each dialogue contains at least three dialogue types with natural topic transitions, constructed via crowdsourcing on top of real medical dialogues from MedDialog, augmented with a medical knowledge graph of 229,570 triples. Shi et al. 2023
XDailyDialog Multilingual (English, German, Chinese, Italian) Text Text (human-translated dialogues) Open-domain dialogue Human-Human 52K dialogues, 411K utterances (13K dialogues aligned across 4 languages) 7.9 XDailyDialog is the first publicly available multilingual parallel open-domain dialogue corpus, consisting of 13K dialogues professionally translated and aligned across four languages (English, German, Chinese, and Italian), yielding 52K dialogues and 411K utterances in total. It supports monolingual, multilingual, and cross-lingual open-domain dialogue research. Liu et al. 2023
TopiOCQA English Text Text (information-seeking questions, free-form answers, rationale spans) Open-domain conversational question answering with topic switching, based on Wikipedia Human-Human 3,920 conversations, 50,466 turns (QA pairs) 13 turns per conversation TopiOCQA is an open-domain conversational question answering dataset built on Wikipedia, featuring information-seeking dialogues in which the topic (Wikipedia document) may switch across turns. Conversations are collected by pairs of annotators (questioner and answerer) and include free-form answers; on average each conversation spans 13 QA turns and covers 4 distinct topics (documents). Adlakha et al. 2022
CareCall Corpus Korean Text Text (LM-generated dialogues with human filtering, and human-bot dialogues with corrections) Role-specified open-domain dialogue; caring chatbot for senior citizens living alone Human-System ~19,240 dialogues (17,617 filtered + 1,623 feedback); ~184,268 turns (154,903 filtered + 29,365 feedback); 57,920 positive and 22,112 negative example pairs 8.79 (filtered set); 18.09 (feedback set) A Korean open-domain dialogue dataset for a role-specified chatbot designed to have casual conversations with senior citizens living alone. Constructed via in-context few-shot generation using large-scale language models (HyperCLOVA), followed by human filtering and human-bot feedback collection, with positive and negative example pairs annotated for out-of-bounds utterances. Bae et al. 2022
DVD-DST English Text and Video Synthetic dialogues grounded on 3D-rendered videos, with annotated multimodal dialogue states (object attributes and temporal slots) Multimodal dialogue state tracking over video-grounded dialogues featuring 3D visual objects Human-System (synthetically generated) 13,992 dialogues, 139,920 turns, across 13,998 videos 10 DVD-DST is a synthetic benchmark for multimodal dialogue state tracking built on CATER 3D-rendered videos. Each dialogue is grounded on a video containing visually varied objects, and dialogue states include object attribute slots (size, color, material, shape) plus temporal start/end slots, requiring models to update states turn-by-turn as new objects or segments are mentioned. Le et al. 2022
Cleaned DailyDialog and Cleaned OpenSubtitles English Text Text Open-domain dialogue generation Human-Human Cleaned DailyDialog: 60,005–60,138 train / 6,594–6,612 validation / 6,955–6,980 test (single- and multi-turn); Cleaned OpenSubtitles: 979,230–1,002,026 train / 11,982–12,289 validation / 12,152–12,506 test (single- and multi-turn) Deduplicated and re-split versions of the DailyDialog and OpenSubtitles open-domain dialogue benchmarks, produced by identifying and removing identical or near-identical train/test overlapping samples using a bag-of-words overlap ratio method. The cleaned datasets provide a more rigorous evaluation protocol for open-domain dialogue generation research. Wen et al. 2022
DIASER English Text Text (transcripts with unified dialogue act, belief state, and ontology annotations) Multi-domain task-oriented dialogue (restaurant, hotel, and ~19 domains total, including travel, services, and more) Human-Human and Human-System (mixed, drawn from source corpora) ~37,100 dialogues, ~662,800 turns 17.83 DIASER (DIAlog System extendER) is a large-scale, unified task-oriented dialogue dataset created by merging and re-annotating four publicly available corpora — MultiWOZ 2.2, CamRest676, DSTC2, and the Schema-Guided Dialogue Dataset — under a common ontology and annotation schema spanning 19 domains and 166 slots. It is intended to support end-to-end dialogue model training with richer, more diverse annotated data than any single source dataset. Hudeček et al. 2022
CRECIL Mandarin Chinese Text Text (TV scripts/transcripts), character relationship triples, referential relationship annotations Character relationship extraction from multi-party dialogue (Chinese sitcom “I Love My Family”) Multi-party human 679 dialogues, 20,183 turns, 121 character entities, 501 global character relationship triples, 8,282 referential relationship triples, 53,646 dialogue-based character relationship triples, 30 relationship types 29.7 CRECIL is a freely available Chinese multi-party dialogue corpus extracted from the TV scripts of the Chinese sitcom “I Love My Family” (120 episodes, 679 scenes), annotated for dialogue-based character relationship extraction. It provides global character relationship maps, referential relationship annotations, and automatically generated character relationship triples across 30 relationship types (including 13 new Chinese-oriented types), covering 121 character entities. Jiang et al. 2022
Character Speech Corpus (CHAR) / Narrator Speech Corpus (NARR) / Neutral Speech Corpus (NEU) Estonian Speech Audio recordings with aligned transcripts Text-to-speech synthesis training; conversational style voice synthesis Human-System (single professional speaker recording scripted/audiobook material) Three experimental corpora of equal size (99,500 characters each): CHAR – 2,063 sentences; NARR – 867 sentences; NEU – 1,535 sentences. Underlying fiction audiobook sub-corpus (PT): ~8 hours (6h narrator + 2h character speech); neutral sentence corpus: 1,849 sentences (2.47 hours). Three Estonian speech corpora recorded by the same professional male speaker, created for training and comparing conversational-style TTS voices: a Character Speech Corpus (audiobook dialogues), a Narrator Speech Corpus (audiobook narration), and a Neutral Speech Corpus (neutral-style isolated sentences). All three experimental corpora are balanced to 99,500 characters and identical technical parameters (48 kHz, 16-bit, mono). Piits et al. 2022
UgChDial Uyghur Text Text (chat messages), annotated question-response pairs Open domain chat; response space classification Human-Human (two-party and multi-party) 12,911 total turns (7,323 two-party + 4,142 MP-chitchat + 1,446 MP-topic); 86,504 total words; 2,419 annotated question-response pairs Two-party sessions avg. ~293 turns/session (25 sessions); multi-party ~approx. 20 dialogues UgChDial is the first dialogue corpus for Uyghur, collected via a customized open-source chatroom (Rocket.Chat) and comprising both two-party dialogues (25 sessions of 120 minutes each, covering 16 scenarios/topics) and multi-party chitchat and topic-oriented dialogues. All question-response pairs are annotated with a fine-grained response space taxonomy (10 classes) to support research on response space classification in a low-resource language. Yusupujiang et al. 2022
Bazinga! English Multimodal (Speech and Text) Audio, manual transcripts, forced-alignment timestamps, speaker labels, addressee labels, entity linking annotations Multi-party dialogue structuring; tasks include speaker diarization, speaker identification, ASR, punctuation restoration, named entity recognition, entity linking, and addressee detection Multi-party human (scripted TV and movie series) 1,765 episodes (127 gold + 1,638 silver); ~8M tokens (569K gold + 7,584K silver); 400+ hours of speech (~29.4h gold, ~399.4h silver); 64,698 entity linking mentions in gold subset; 73K+ sentences annotated with addressee Bazinga! is a large-scale multimodal dataset of multi-party dialogues drawn from 13 TV series and 3 movie series (16 in total), comprising 400+ hours of speech and 8M+ tokens. A gold-standard subset (~30 hours, 569K tokens, 127 episodes) provides word-level annotations for speaker, addressee, and entity linking, while the larger silver-standard subset supplies transcripts and forced-alignment timestamps to support self- and weakly-supervised learning research. Lerner et al. 2022
ChiSense-12 English Text Transcripts (sense-annotated utterances with verb-object coding) Child-directed speech; word sense disambiguation Human-Human (caregiver–child naturalistic interactions) 15,581 utterances for 12 ambiguous words; sourced from 53 corpora covering 958 target children; 115,272 word tokens, 4,805 word types ChiSense-12 is a large-scale sense-annotated corpus of American and British English child-directed speech, drawn from 53 CHILDES corpora covering 958 children up to 59 months of age. It contains 15,581 sense-tagged utterances for 12 ambiguous words (dominant and subordinate senses), with additional coding of verb instances in which the target ambiguous word appears as a verb object, enabling study of verb-event structure in child word sense disambiguation. Cabiddu et al. 2022
MMChat Mandarin Chinese Multimodal (text and image) Text dialogues, Images, Image captions, Detected object labels Open-domain image-grounded conversation on social media Human-Human 120.84K filtered dialogues, 204.32K images, 314.13K utterances (raw: 32.4M dialogues, 8.41M images); MMChat-hf subset: 19.90K dialogues, 52.66K images, 81.06K utterances 2.59 utterances per dialogue (MMChat); 4.07 utterances per dialogue (MMChat-hf) MMChat is a large-scale Chinese multi-modal dialogue corpus collected from real conversations on social media, comprising 120.84K filtered image-grounded dialogues (from 32.4M raw sessions) paired with images, object labels, and captions. It includes a human-filtered subset (MMChat-hf, 19.90K dialogues) and is designed to study the “sparsity” phenomenon where image-initiated dialogues drift to non-image-related topics. Zheng et al. 2022
Travel Agency Task Dialogue Corpus with Age-Diverse Speakers Japanese Multimodal (video, audio, transcripts) Video (mp4), Audio (m4a, including separate per-speaker channels), manual transcripts, dialogue act annotations, facial action unit annotations, system query/retrieval logs Tourism consultation / travel agency task (operator recommends tourist spots to customer) Human-Human (operator and customer, face-to-face via Zoom video call) 330 dialogues, 111,771 utterances (turns), 246,316 annotated functional segments, ~6,948 minutes (~115+ hours) of recorded dialogue 338.7 turns per dialogue (111,771 turns / 330 dialogues) A large multimodal Japanese dialogue corpus of two-party travel agency consultations between an operator and customers spanning a wide age range (7–72 years old), including minors, adults, and older adults. Dialogues are manually transcribed and annotated with ISO 24617-2 dialogue act tags; video data also supports facial action unit analysis. Inaba et al. 2022
Badalona Corpus Spanish Multimodal (Audio, Video, EEG, Physiological signals) Audio (head-mounted microphone, 44.1kHz/24-bit), Video (frontal cameras), EEG (Emotiv EpocX 14-channel, 128Hz), Physiological signals (Empatica E4 wristband: BVP, EDA, IBI, HR, body temperature, 3-axis accelerometer), Automatic transcriptions and phoneme-level alignments, Facial landmark and gaze annotations Natural dyadic conversation; includes controlled divergent thinking tasks (Alternative Uses Test, Name Invention Task) and free conversation (moral dilemma discussion) Human-Human (dyadic; 5 dyads, 10 participants) 5 dyads × 3 sessions × ~30 min each; approximately 12.5 hours of multimodal data total The Badalona Corpus is the first natural conversational dataset combining audio, video, EEG (Emotiv EpocX), and electro-physiological signals (Empatica E4) recorded simultaneously. Five Spanish-speaking dyads (10 participants) were recorded across three longitudinal sessions (spaced 4 days apart), enabling investigation of interlocutor alignment and convergence over time; the corpus is enriched with automatic transcriptions, phoneme alignments, and facial expression annotations. Blache et al. 2022
EmoInHindi Hindi Text Text (dialogue transcripts with multi-label emotion and intensity annotations) Mental health counselling and legal assistance for crime victims; multi-label emotion and intensity recognition in conversations Human-WoZ 1,814 dialogues, 44,247 utterances, 7,036 unique tokens 24.39 utterances per dialogue EmoInHindi is a large-scale Hindi conversational dataset constructed in Wizard-of-Oz style for multi-label emotion and intensity recognition in dialogues, covering mental health counselling and legal assistance for crime victims. Each utterance is annotated with one or more of 16 emotion labels (including Neutral) and corresponding intensity values (0–3), making it the first multi-label emotion and intensity annotated conversational dataset in Hindi. Singh et al. 2022
INSPIRED English Text Text (crowdsourced natural-language dialogues with SPARQL logical forms, sub-questions, and correction feedback) Interactive semantic parsing / question answering over knowledge bases (KBQA); complex multi-hop questions derived from ComplexWebQuestions Human-System 10,374 dialogues (3,492 train / 3,441 dev / 3,441 test complex questions) Varies; average 1.9 predicted sub-questions and 1.4 edit operations per dialogue INSPIRED (INteractive Semantic ParsIng for CoRREction with Decomposition) is a crowdsourced dialogue dataset for interactive KBQA, derived from the ComplexWebQuestions (CWQ) dataset. Each dialogue has a system agent explaining a predicted SPARQL parse step-by-step in natural language, and crowdworkers providing natural-language feedback to correct individual sub-steps. Mo et al. 2022
HybriDialogue English Text Text (crowdsourced natural language dialogues grounded on Wikipedia tables and text passages) Information-seeking dialogue grounded on hybrid tabular and textual Wikipedia knowledge Human-Human (single Turker simulating both seeker and expert roles) 4,844 dialogues, 21,070 turns (QA pairs) 4.34 HybriDialogue is a crowdsourced information-seeking dialogue dataset in which multi-turn conversations are grounded on both Wikipedia tables and text. Dialogues are constructed by decomposing complex multi-hop questions from OTT-QA into sequences of simpler, natural intermediate question-answer turns, each associated with a structured or unstructured Wikipedia reference (table rows, cells, linked paragraphs, or intro text). Nakamura et al. 2022
DomainCC and DomainReddit English Text Flat text (DomainCC) and dialogic context-response triples (DomainReddit) Task-oriented dialogue domain specialization; five MultiWOZ domains: Taxi, Attraction, Train, Hotel, Restaurant Human-Human (Reddit threads) DomainCC: 200K sentences per domain (5 domains); DomainReddit: Taxi – 120K, Attraction – 157K, Hotel – 229K, Train – 229K, Restaurant – 243K context–response triples DomainCC and DomainReddit are domain-specific pretraining corpora for five MultiWOZ task-oriented dialogue domains (Taxi, Attraction, Train, Hotel, Restaurant), constructed by filtering CCNet and Reddit using TF-IDF-extracted domain ngrams. DomainCC provides flat text for masked language modelling, while DomainReddit provides dialogic context–true response–false response triples drawn from travel-related subreddits for response-selection pretraining. Hung et al. 2022
DMR-FastFood English Text Text transcripts annotated with Dialogue Meaning Representation (DMR) graphs, including conjunction, negation, and cross-turn coreference annotations Fast-food ordering (task-oriented dialogue) Human-System 7,194 dialogues, 70,328 annotated utterances, 16,087 conjunctions, 557 negations, 7,846 coreference references 18.5 DMR-FastFood is a multi-turn task-oriented dialogue dataset in the fast-food ordering domain, derived from the MultiDoGO dataset and annotated with Dialogue Meaning Representation (DMR) — a rooted directed acyclic graph representation capturing rich compositional semantics including conjunction, negation, modification, quantification, and cross-turn coreference. It features substantially more turns per dialogue and richer linguistic semantic annotations than comparable datasets. Hu et al. 2022
KETOD English Text Text (task-oriented dialogues enriched with knowledge-grounded chit-chat, dialogue state annotations, dialogue act annotations, Wikipedia knowledge snippets) Knowledge-enriched task-oriented dialogue across 16 domains (e.g., Restaurant, Music, Hotels, Movies, Events, Weather) Human-System 5,324 dialogues, 52,063 turns (6,302 enriched with chit-chat), 4,639 entities, 33,761 knowledge snippets; split 4,247/545/532 train/dev/test 9.78 KETOD (Knowledge-Enriched Task-Oriented Dialogue) is a dataset built upon the SGD task-oriented dialogue corpus, augmented by human annotators who enrich system responses with knowledge-grounded chit-chat based on Wikipedia entity knowledge retrieved from dialogue states and actions. It covers 16 domains and is designed to support research on integrating task-oriented dialogue with knowledge-grounded open-domain conversation. Chen et al. 2022
JILDA 2.0 Italian Text Text transcripts with Dialogue Act and slot annotations Job offer / job application (task-oriented) Human-Human 745 dialogues, 17,889 utterances, 263,104 tokens 17 JILDA 2.0 is an updated version of the JILDA Italian task-oriented dialogue corpus, consisting of 745 complex human-human dialogues in the job offer domain, annotated with 12 dialogue acts and 14 slot types. The update corrects annotation errors and aligns the resource with the MultiWOZ 2.1 standard, providing a benchmark for Italian NLU (Dialogue Act Recognition and Slot Recognition). Sucameli et al. 2022
BSBT (Blended Skill BotsTalk) English Text Text (automatically generated multi-skill dialogues with skill annotations) Open-domain multi-skill dialogue (blending personality, knowledge, and empathy) Human-System (bot-bot conversations generated by multiple skill-grounded dialogue agents) 300K dialogues, 3M utterances 10 BSBT is a large-scale, machine-sourced multi-skill open-domain dialogue dataset comprising 300K conversations (3M utterances) automatically generated by the BotsTalk framework, in which multiple skill-grounded agents (persona, knowledge, empathy) collaboratively produce dialogues blending skills derived from ConvAI2, Wizard of Wikipedia, and Empathetic Dialogues. Each utterance is annotated with a skill type label and skill distribution. Kim et al. 2022
DIALOCONAN English Text Text (fictitious multi-turn dialogues between a hater and an NGO operator, post-edited by human expert annotators from machine-generated candidates) Hate speech countering via counter narratives; multi-turn dialogical counter narrative generation Human-Human (fictitious: simulated hater vs. NGO operator, written/post-edited by expert annotators) 3,059 dialogues, 16,625 turns, covering 6 hate targets (LGBT+, Migrants, Muslims, Jews, POC, Women) ~5.43 turns per dialogue (16,625 turns / 3,059 dialogues); dialogues have 4, 6, or 8 turns DIALOCONAN (DIALOgical COunter-NArratives collectioN) is the first multi-turn dialogue dataset for hate speech countering, comprising 3,059 fictitious dialogues (16,625 turns) between a simulated hater and an NGO operator across 6 hate targets. It was built via a hybrid human-machine approach combining 19 generation/augmentation strategies with expert post-editing. Bonaldi et al. 2022
ProsocialDialog English Text Text (multi-turn dialogues, rules-of-thumb, dialogue safety labels with free-form rationales) Prosocial response generation; dialogue safety detection in unethical, toxic, biased, and problematic conversational contexts Human-AI (GPT-3 generates problematic utterances; crowdworkers provide prosocial responses) 58,137 dialogues, 331,362 utterances, 160,295 unique rules-of-thumb, 497,043 dialogue safety labels with rationales 5.7 ProsocialDialog is a large-scale multi-turn English dialogue dataset designed to teach conversational agents to respond prosocially to problematic, toxic, biased, or unethical user utterances in accordance with commonsense social norms (rules-of-thumb). Created via a human-AI collaborative framework, each dialogue is annotated with rules-of-thumb grounding prosocial responses and fine-grained dialogue safety labels (Casual, Needs Caution, Needs Intervention) accompanied by free-form rationales. Kim et al. 2022
D4 Mandarin Chinese Text Text transcripts of simulated doctor-patient dialogues, with topic annotations, psychiatrist-authored diagnosis summaries, symptom summaries, and severity labels for depressive episode and suicide risk Depression diagnosis consultation (mental health) Human-Human (crowdsourced workers playing doctor and patient roles, supervised by licensed psychiatrists) 1,339 dialogues; avg. 60.9 utterances per dialogue; avg. 877.6 tokens per dialogue 21.6 D4 is a Chinese dialogue dataset of simulated clinical depression-diagnosis consultations, constructed via a 3-phase procedure grounded in ICD-11 and DSM-5 criteria. Each of the 1,339 conversations is annotated with dialogue topic tags, and accompanied by a professional psychiatrist’s symptom summary and severity ratings for depressive episode and suicide risk. Yao et al. 2022
KPRS (Knowledge Probing using Response Selection) English Text Text (dialogue contexts and contrastive response pairs) Task-oriented dialogue; knowledge probing across restaurant, hotel, attraction, and train domains Human-System 3,055 samples derived from 831 unique dialogues / 1,997 unique dialogue contexts 1.52 samples per dialogue context; 3.65 samples per dialogue KPRS is a contrastive probing benchmark derived from MultiWOZ 2.2 development and test dialogues, covering four task-oriented domains (restaurant, hotel, attraction, train). Each sample pairs a dialogue context with a KB-consistent reference response and a minimally perturbed distractor response, designed to evaluate whether TOD models can access and correctly retrieve domain-specific factual knowledge stored in their parameters. Emelin et al. 2022
META-GUI English Multimodal (text and screenshots/images) Text dialogues, GUI operation traces (screenshots, Android view hierarchies, action sequences) Task-oriented dialogue via mobile GUI operations; domains include weather, calendar, search, taxi, hotel, and restaurant Human-WoZ (annotators acting as both user and system, with GUI traces recorded on real Android devices) 1,125 dialogues, 4,684 turns, 18,337 action-level data points 4.16 META-GUI is a multimodal dataset for training conversational agents to operate real Android mobile apps via GUI interactions, without relying on backend APIs. Each dialogue is paired with GUI operation traces comprising screenshots, Android view hierarchies, and action sequences (click, swipe, input, etc.) across six task domains. Sun et al. 2022
TamilATIS Tamil Text Text (translated and manually annotated utterances with intent and slot labels) Airline travel inquiry / flight information (task-oriented dialogue NLU) Human-System 4,874 utterances; 23 intent labels; 45 slot labels; vocabulary size 1,819 TamilATIS is a task-oriented dialogue dataset for Tamil, derived by automatically translating a modified version of the English ATIS corpus into Tamil using the Google Translate API and then manually annotating slot labels. It contains 4,874 utterances covering airline-related enquiries, annotated with 23 intent classes and 45 slot label types, intended to support NLU research (intent detection and slot filling) in Tamil as a low-resource language. Ramaneswaran et al. 2022
Task2Dial English Text Text (dialogues grounded in recipe documents) Cooking / recipe-following instruction-giving (document-grounded, commonsense-enhanced) Human-Human (Information Giver and Information Follower, via self-dialogue annotation) 353 dialogues (documents/recipes); avg. 18.15 turns per dialogue; avg. 19.79 tokens per turn 18.15 Task2Dial is a dataset of document-grounded, task-based dialogues in the cooking domain, where an Information Giver (IG) provides recipe instructions to an Information Follower (IF), who may ask clarification questions requiring commonsense knowledge not present in the underlying document. The dataset is notable for its lexical richness, longer turns, and the need for commonsense reasoning and sequential planning compared to existing document-grounded dialogue datasets. Strathearn et al. 2022
Wired Explaining Dialogue Corpus English Text (transcripts from video) Transcripts, manually annotated turn-level labels (topic relation, dialogue act, explanation move) Dialogical explanations of science and technology topics (e.g., blockchain, machine learning, black holes) Human-Human (expert explainer and explainees of varying proficiency: child, teenager, undergrad, grad student, colleague) 65 dialogues, 1,550 turns, 51,344 words, covering 13 topics 23.8 A corpus of 65 transcribed English expert-to-layperson explaining dialogues drawn from the Wired video series “5 Levels,” in which a domain expert explains 13 science/technology topics to five explainees of increasing proficiency (child to fellow expert). All 1,550 turns are manually annotated by five independent professionals for topic relation, dialogue act (10 categories), and explanation move (10 categories). Wachsmuth et al. 2022
HLA-Chat++ English Text Text (dialogue transcripts from TV drama scripts, persona tags, sentence-level and document-level persona tag interpretations from TVTropes) Open-domain multi-party personalized dialogue generation Multi-party human (TV drama characters) 823,204 conversations; 239 characters; 30 TV dramas; 4,778 sentence-level knowledge entries; 4,778 document-level knowledge entries Average of 10.7 nodes and 24.6 edges per dialogue graph HLA-Chat++ is a multi-party personalized dialogue dataset constructed from 30 English TV drama scripts, where each character is annotated with TVTropes persona tags enriched with sentence-level (laconic) and document-level (main) text knowledge interpretations. It extends the earlier HLA-Chat dataset by retaining original persona tags and adding structured external text knowledge to address the challenge of incomprehensible persona representations in multi-party dialogue generation. Ju et al. 2022
JDDC 2.1 Mandarin Chinese Multimodal (text and image) Text, Images, Product knowledge bases, Image category annotations, Query rewriting annotations, Discourse parsing annotations, Dialogue summarization annotations E-commerce customer service (multi-task: response generation, query rewriting, discourse parsing, summarization) Human-Human 246,153 dialogue sessions, 3,459,888 utterances, 507,678 images; 2,000 sessions annotated for query rewriting, discourse parsing, and summarization 14.06 JDDC 2.1 is a large-scale multimodal multi-turn Chinese dialogue dataset collected from JD.com, a major Chinese e-commerce platform, covering small home appliances and fashion product categories. It contains 246,153 dialogue sessions with 3,459,888 utterances and 507,678 images, and provides joint annotations for four tasks: multimodal dialogue response generation, multimodal query rewriting, multimodal dialogue discourse parsing, and multimodal dialogue summarization, all over the same dialogue sessions. Zhao et al. 2022
MultilingualDatasets (MulZDG code-switching dialogue datasets) Multilingual (English, Chinese, German, Russian, Spanish, French, Italian) Text Text (code-switching dialogue context-response pairs, automatically constructed via NMT-based utterance translation) Open-domain dialogue generation (daily life conversations and social media dialogue) Human-Human Derived from DailyDialog (11,118 training / 1,000 validation / 1,000 test context-response pairs) and DSTC7 (76,590 training / 17,870 validation / 1,710 test pairs), each expanded into 6 bilingual code-switching variants (English paired with Chinese, German, Russian, Spanish, French, Italian) Multilingual code-switching dialogue datasets constructed from the English DailyDialog and DSTC7 corpora by randomly selecting utterances from dialogue histories and translating them into six target languages (Chinese, German, Russian, Spanish, French, Italian) using NMT systems, producing bilingual code-switching training sets for zero-shot and data-augmentation dialogue generation research. Liu et al. 2022
Reddit Multi-turn Dialogue Corpus English Text Text (automatically generated multi-turn dialogues from Reddit posts and comments) Open domain (covering 10 subreddits: Advice, Books, College, Casual Conversation, Fitness, LetsTalkMusic, Movies, TrueGaming, Writing, TalesFromRetail) Human-Human (reconstructed from Reddit post authors and commenters) 10,098 dialogues, 109,916 utterances, 3,317,807 tokens 10.9 utterances per dialogue A large-scale multi-turn dialogue corpus automatically constructed from Reddit posts and their comments across 10 subreddits using a threading-based algorithm that sequences post sentences and comment segments into coherent two-speaker dialogues. The corpus is intended for pretraining neural dialogue models and was found to be more engaging but slightly less natural than comparable crowdsourced datasets. Huryn et al. 2022
Wizard of Tasks English Text (with multimodal content sharing such as images and structured recipe/article content) Text transcripts, intent labels, teacher action labels, relevance/usefulness annotations, shared multimodal content (step images, step text, ingredients, tools) Conversational Task Assistance: Cooking and Home Improvement (DIY) Human-WOz (asynchronous Wizard-of-Oz crowdsourcing via Amazon Mechanical Turk; one worker as ‘student’, another as ‘teacher’) 549 conversations, 18,077 utterances (272 conversations / 7,908 utterances in Cooking; 277 conversations / 10,169 utterances in DIY) 29.1 turns per conversation (Cooking); 36.7 turns per conversation (DIY) Wizard of Tasks is the first conversational corpus designed for Conversational Task Assistants (CTAs), covering two real-world task domains: Cooking and Home Improvement (DIY). It was crowd-sourced via an asynchronous Wizard-of-Oz setup on Amazon Mechanical Turk, with 549 conversations and 18,077 utterances annotated with student intents, teacher actions, relevance/usefulness labels, and shared multimodal content, supporting tasks such as Intent Classification and Abstractive Question Answering. Choi et al. 2022
DEER English Multimodal (text, audio, video) Video recordings, manual transcripts, audio features, speaker information, utterance-level timestamps Mental health: emotion detection and emotional reasoning detection in doctor-patient conversations Human-Human (doctor-patient dyadic interactions) 30 videos, 3,753 annotated utterances (743 ER utterances) DEER (Detection of Emotion and Emotional Reasoning) is a multimodal mental health conversational corpus of 30 doctor-patient interview videos (20 real, 10 enacted) sourced from YouTube, manually annotated at the utterance level with one of seven emotion classes (Ekman’s six basic emotions plus ‘others’) and binary emotional reasoning (ER) labels, along with speaker information and start/end timestamps. It is the first publicly released multimodal corpus targeting emotional reasoning detection in clinical mental health conversations. Ghosh et al. 2022
CODI-CRAC 2022 Corpus English Text Text transcripts annotated for identity coreference, bridging references, and discourse deixis Multi-domain dialogue: meeting recordings (AMI), fantasy text adventure game dialogues (LIGHT), persuasion conversations (Persuasion for Good), and spontaneous telephone conversations (Switchboard) Human-Human 218 documents, 214,625 tokens, 60,933 markables (29,363 identity/DO anaphors, 6,626 bridging references, 1,583 discourse deixis) The CODI-CRAC 2022 corpus is a large-scale English dialogue dataset newly annotated for identity anaphora, bridging references, and discourse deixis, comprising conversations from four domains (AMI, LIGHT, Persuasion for Good, and Switchboard) annotated using the ARRAU annotation scheme. It is reported to be the largest dataset annotated for anaphoric interpretation in dialogue, and one of the largest for bridging references. Yu et al. 2022
Hindi–English Chat and QnA Translation Corpus English, Hindi Text Parallel text (synthetic machine-translated and gold-standard human-translated sentence pairs) Chat translation (customer service, e-commerce) and Question-Answer (QnA) translation Human-System Three corpora: (1) WMT20 Chat: 550 dialogues, 16,249 sentences (train/val/test); (2) MMD: 1,437 dialogues, 52,535 sentences (train/val/test); (3) QnA: 2.1M QnA pairs (~4.2M sentences, plus 1,000 gold-standard val and 1,000 gold-standard test sentences). Total synthetic sentences: ~68.7K (chat) + 4.19M (QnA); gold-standard: 3,037 sentences (chat) + 2,000 sentences (QnA). WMT20 Chat: ~25.17 turns/dialogue; MMD: ~35.6 turns/dialogue A benchmark English–Hindi parallel corpus for chat and QnA translation, covering service and e-commerce domains. It comprises Hindi translations (synthetic and gold-standard) of the WMT20 Chat dataset (based on Taskmaster-1), the MultiModal Dialogue (MMD) corpus, and a large-scale Flipkart QnA corpus of 2.1M question-answer pairs, with professional human translations for test and validation sets. Gain et al. 2022
GlobalWoZ Multilingual (20 languages including Chinese, Spanish, Indonesian, Arabic, Danish, German, Greek, French, Hebrew, Italian, Japanese, Korean, Dutch, Norwegian, Portuguese, Russian, Swedish, Thai, Turkish, Vietnamese) Text Text (machine-translated and human post-edited dialogue transcripts with dialogue state annotations) Task-oriented dialogue (hotel booking, restaurant search, attraction finding, train booking, taxi); multi-domain Human-System ~566,280 train/dev dialogues and 60,000 test dialogues across 20 languages and 3 use cases (9,438 train/dev + 1,000 test dialogues per language per use case) GlobalWoZ is a large-scale multilingual task-oriented dialogue dataset derived from MultiWoZ 2.2 by translating dialogue templates (via machine translation with professional post-editing for test sets in Chinese, Spanish, and Indonesian) and filling them with locally crawled entities from target-language cities. It covers 20 languages across three novel use cases (foreign-language speaker in foreign-language country, foreign-language speaker in English country, English speaker in foreign-language country) for dialogue state tracking research. Ding et al. 2022
MDMD (Multi-label Dialogue Malevolence Detection) English Text Text (multi-turn dialogue utterances with multi-label malevolence annotations) Malevolence detection in dialogues (negative emotions, inappropriate behavior, unethical content) Human-Human 8,462 utterances (2,098 malevolent, 6,364 non-malevolent); re-annotated validation and test splits from MDRDC (701 and 1,397 utterances respectively); training set from original MDRDC (6,000 dialogues, 10,299 malevolent + 21,081 non-malevolent utterances) MDMD is a multi-label dialogue malevolence detection dataset constructed by re-annotating the validation and test sets of the existing MDRDC dataset via Amazon MTurk, allowing each utterance to be assigned multiple malevolence labels from an 18-category, 3-level taxonomy covering negative emotions, negative psychological behavior, and unethical issues. It is designed to support research on multi-label malevolence detection in multi-turn dialogues. Zhang et al. 2022
MSCTD English, Chinese, German Multimodal (text and image) Text (utterances with translations), Images, Sentiment labels Multimodal chat translation and multimodal dialogue sentiment analysis; movie-based bilingual conversations Human-Human 17,841 bilingual dialogues; 173,241 utterance-image-sentiment quadruplets (142,871 English-Chinese and 30,370 English-German utterance pairs) ~10 turns per dialogue MSCTD is a human-annotated multimodal sentiment chat translation dataset comprising 17,841 bilingual (English–Chinese and English–German) movie-based dialogues, each utterance paired with a scene image and a sentiment label (positive/neutral/negative). It supports benchmarking of multimodal chat translation across four language directions as well as multimodal dialogue sentiment analysis in three languages. Liang et al. 2022
Large-scale In-domain Paired Bilingual Dialogue Dataset (Movie Subtitles) English, Chinese, German Text Text (aligned bilingual movie subtitle dialogues) Chat translation (Neural Chat Translation); movie subtitle dialogues Human-Human En↔Zh: 28,214,769 dialogues, 28,238,877 utterances; En↔De: 18,041,125 dialogues, 18,048,573 utterances 4 Two large-scale in-domain paired bilingual dialogue corpora constructed from aligned movie subtitles for neural chat translation research: an English–Chinese set (~28M dialogues) and an English–German set (~18M dialogues). Each dialogue consists of four consecutive aligned utterances, built using the Vecalign sentence alignment tool and LASER multilingual embeddings. Liang et al. 2022
QAConv English Text Text (conversation transcripts and QA pairs, including human-written and machine-generated questions with span-extractable or unanswerable answers) Question answering on informative conversations (business emails, panel discussions, work/Slack channels) Multi-party human 34,608 QA pairs from 10,259 conversations; split into 27,287 train / 3,660 validation / 3,661 test samples Avg. 568.8 words per dialogue (avg. 2.8 speakers per dialogue) QAConv is a question answering dataset grounded in informative multi-party conversations — business emails (BC3, Enron), panel discussions (Court, Media), and work channels (Slack) — featuring 34,608 QA pairs (human-written and machine-generated) over 10,259 conversations. It supports two evaluation modes: chunk mode (oracle conversational chunk provided) and full mode (retrieval required), and includes both answerable and unanswerable questions. Wu et al. 2022
SalesBot English Text Text (automatically generated dialogues with human annotations for relevance, aggressiveness, transition quality, and implicit intent rankings) Sales dialogues transitioning from open-domain chit-chat to task-oriented conversations (movie finding, attraction finding, music lookup, song playing) Human-System (simulated user and simulated salesperson, with human crowdsourced annotations) 3,916 dialogues (sampled for human evaluation); full dataset larger (unlimited generation possible) 17 (average over sampled dialogues; ranges from 13 to 21 by subset) SalesBot is a large-scale dataset of dialogues that naturally transition from open-domain chit-chat to task-oriented conversations, simulating a salesperson discovering users’ implicit intents and guiding them toward completing tasks such as finding movies, music, or attractions. Dialogues are automatically generated using BlenderBot-based simulators and include detailed human annotations for transition quality, relevance, aggressiveness, and implicit intent. Chiu et al. 2022
WITS Hindi-English code-mixed Multimodal (text, audio, video) Text transcripts, audio, video Sarcasm explanation in dialogue; code-mixed multi-party conversations from the Indian TV show ‘Sarabhai v/s Sarabhai’ Multi-party human 2,240 sarcastic dialogues, 9,080 utterances 4.05 WITS (“Why Is This Sarcastic”) is a multimodal, multi-party, Hindi-English code-mixed dialogue dataset built by extending the MASAC dataset with human-annotated natural language explanations for sarcastic utterances. Each instance includes text transcripts, audio, and video, and is annotated with an explanation identifying the sarcasm source, target, action word, and description. Kumar et al. 2022
Expressed and Experienced Emotions Dialogue Corpus Japanese Text Text (Twitter dialogue transcripts with crowdsourced emotion annotations) Open-domain / emotion-aware dialogue (Twitter conversations) Human-Human 3,828 dialogues, 13,806 utterances 3.61 A Japanese multi-turn dialogue corpus collected from Twitter and annotated via crowdsourcing with two types of emotions per utterance: the emotion expressed by the speaker and the emotion experienced by the listener. Each utterance may carry multiple emotion labels (from Plutchik’s eight basic emotions) at two intensity levels (strong and weak). Ide et al. 2022
PPMD English Multimodal (text and image) Text (dialogues), Images E-commerce assistant (buying/selling electronic gadgets, phones, tablets) Human-Human 1,031 dialogues, 11,602 utterances, 1,861 images 11.25 The Personalized Persuasive Multi-modal Dialogue (PPMD) corpus is a human-curated e-commerce conversational dataset covering goal-unavailability and persuasion scenarios in phone/tablet buying-selling. Each utterance is annotated with intent, slot, sentiment, dialogue act, user persona, image information, and persuasion strategy (6 strategies); the corpus also includes 1,861 product images across 5 visual attribute categories. Tiwari et al. 2022
MRCWOZ English Text Dialogue transcripts with annotated question–answer pairs (slot-based comprehension questions) Task-oriented dialogue comprehension (restaurant, hotel, and taxi booking domains) Human-System 2,409 dialogues; 8,950 train slot–question pairs and 791 test slot–question pairs across three domains 8.92 MRCWOZ is a task-oriented dialogue comprehension dataset derived from MultiWOZ 2.2, covering restaurant, hotel, and taxi domains. It pairs existing MultiWOZ dialogues with manually annotated slot-based comprehension questions (averaging 4.2 questions per dialogue), enabling machine reading comprehension evaluation over task-oriented conversations. Aksu et al. 2021
CCC DA Annotations English Text Transcripts with dialogue act annotations Clinical conversational interviews with Alzheimer’s Disease patients and elderly controls Human-Human 30 conversations, 5,082 utterances A dialogue act annotation layer applied to a subset of 30 conversations from the Carolinas Conversation Collection (CCC), covering 10 Alzheimer’s Disease patients and 10 Non-AD elderly controls. Conversations are manually annotated with a 20-tag DAMSL-derived tagset specifically designed to capture rare dialogue acts indicative of AD, including clarification requests, signal-non-understanding, and question/answer types. Nasreen et al. 2021
HuRDL Corpus English Text (with video recordings of robot actions) Text transcripts, video recordings Collaborative tool-organization task in a virtual spacecraft environment; situated human-robot dialogue for learning Human-Human (participant plays robot role; human confederate plays Commander) 22 dialogues, 1122 participant utterances, 760 questions, 13 hours total duration 51 participant utterances per dialogue (mean) The Human-Robot Dialogue Learning (HuRDL) Corpus is a collection of 22 annotated text-based dialogues from an online interactive virtual environment in which human participants tele-operate a robot to perform a collaborative tool-organization task, asking questions of a human confederate Commander to manage uncertainty about novel objects and procedures. The corpus includes an annotation scheme covering question form (YNQ, AQ, WHQ, Statement) and clarification type categories relevant to situated learning. Gervits et al. 2021
HealthCareMagic and iCliniq Question Summarization Datasets English Text Text (consumer health questions and expert-written summaries extracted from medical dialogues) Consumer health question summarization (medical domain) Human-Human HealthCareMagic: 226,405 pairs (181,122 train / 22,641 dev / 22,642 test); iCliniq: 31,062 pairs (24,851 train / 3,105 dev / 3,106 test) Two medical question summarization datasets extracted from the MedDialog large-scale medical dialogue dataset, sourced from HealthCareMagic.com and iCliniq.com. Each example pairs a patient’s long utterance (consumer health question) with a single-sentence description serving as its summary, supporting training and evaluation of question summarization models in the biomedical domain. Mrini et al. 2021
XPersona Multilingual (Chinese, French, Indonesian, Italian, Korean, Japanese, English) Text Text (dialogues and persona descriptions; training set machine-translated, validation and test sets human-annotated) Personalized open-domain chit-chat Human-Human Validation: ~1,668 dialogues, ~25,992 utterances across 6 languages; Test: ~1,670 dialogues, ~26,090 utterances across 6 languages (plus English from original Persona-Chat); training set ~130K utterances XPersona is a multilingual extension of the Persona-Chat dataset covering six languages beyond English (Chinese, French, Indonesian, Italian, Korean, and Japanese). Training sets are automatically translated using APIs with human-in-the-loop correction, while validation and test sets are fully human-annotated, enabling evaluation of multilingual and cross-lingual personalized dialogue systems. Lin et al. 2021
Contrast Set for Knowledge-seeking Turn Detection English Text Text (dialogue utterances / user queries) Task-oriented dialogue; knowledge-seeking turn detection (hotel, restaurant, attraction domains) Human-System 2,817 samples total (617 knowledge-seeking turns collected from Tripadvisor forums, mixed with 2,200 non-knowledge-seeking turns from the DSTC9 Track 1 test set) A contrast test set for evaluating knowledge-seeking turn detectors in task-oriented dialogue systems, curated by collecting real user questions from Tripadvisor forums, filtering for queries outside the MultiWOZ API schema, and manually paraphrasing them into dialogue utterances. It is designed to assess generalisation beyond the DSTC9 Track 1 benchmark distribution. Jin et al. 2021
ACCENTOR English Text Text (chit-chat augmented task-oriented dialogues with good/bad candidate annotations and justification labels) Task-oriented dialogues augmented with chit-chat (multi-domain: restaurants, hotels, events, ride-sharing, music, etc.) Human-System ACCENTOR-SGD: 22,825 dialogues (228,250 chit-chat candidates annotated, 94,600 good); ACCENTOR-MultiWOZ: 997 dialogues ACCENTOR consists of chit-chat-augmented versions of two popular task-oriented dialogue datasets (Schema-Guided Dialogue and MultiWOZ 2.1), created via a Human↔AI collaborative annotation pipeline using GPT-2 and BlenderBot for candidate generation, automatic filtering, and crowdworker labeling. Each system turn is annotated with good/bad chit-chat candidate add-ons and justification labels (social, useful, inappropriate, misleading). Sun et al. 2021
DialogueMT Test Set Chinese, English Text Parallel text utterances with manual translation and annotation labels (ProDrop, PunDrop, DialTypo) Dialogue machine translation (Chinese-English) Human-Human 300 dialogues, 1,931 sentence pairs, 19,155/15,976 total tokens (Chinese/English) 6.44 A manually annotated Chinese-English benchmark test set for dialogue machine translation, comprising 300 dialogues and 1,931 parallel utterance pairs. Each utterance is annotated for three dialogue-specific translation challenges: pronoun dropping (ProDrop), punctuation dropping (PunDrop), and typos (DialTypo). Wang et al. 2021
Sentimental Douban Conversation Corpus Mandarin Chinese Text Text (sentiment-annotated multi-turn dialogues) Open-domain affective/sentiment-controlled chat Human-Human 500,000 dialogues (1,400 manually annotated + 498,600 automatically annotated); 4,347,200 utterances total (10,712 manual + 4,347,200 automatic, with positive/neutral/negative sentiment labels) 6.69 (training set of base corpus) A sentiment-annotated extension of the Douban Conversation Corpus for affective response research in retrieval-based chatbots. Sentiment polarity labels (positive, neutral, negative) were applied to 1,400 dialogues (10,712 utterances) via human annotation and to 498,600 further dialogues (~4.3M utterances) via an automatic RoBERTa-based classifier. Lu et al. 2021
ForumSum English Text Text (forum conversations and human-written abstractive summaries) Multi-speaker internet forum conversation summarization Human-Human (multi-party) 4,058 dialogues; avg 303.45 input words per conversation; 1.18M total input words 10.13 ForumSum is a diverse, high-quality multi-speaker conversation summarization dataset collected from 281 internet forums, annotated with human-written abstractive summaries via Amazon Mechanical Turk. It features an average of 6.73 speakers per conversation and longer, more descriptive summaries compared to existing chat summarization datasets. Khalman et al. 2021
CEDAR English Text Text (dialogue history, responses, and human-verified commonsense causal explanations) Commonsense reasoning for open-domain dialogue response generation Human-Human 6,000 generated explanations (1,560 human-verified valid) across 1,200 dialogues sampled from four existing datasets CEDAR (CommonSense in DiAlogue Response generation) is an annotation dataset of commonsense causal explanations justifying dialogue responses, collected from 1,200 dialogues drawn from four public dialogue datasets (DailyDialog, EmpatheticDialogues, MuTual, and SocialIQA-prompted dialogues). Each dialogue is annotated with five dimensions of causal explanation (event, emotion, location, possession, attribute), yielding 6,000 generated explanations of which 1,560 were verified as valid by human crowdworkers, along with corrupted versions for probing response generation models’ commonsense reasoning capabilities. Zhou et al. 2021
SGD-S, SGD-M, Multiwoz-T English Text Text (dialogue transcripts with task/intent-based cluster labels) Task-oriented dialogue clustering (travel, flight search, restaurant booking, and other task-oriented domains) Human-Human SGD-S: 3,925 dialogues (29 tasks); SGD-M: 4,722 dialogues (59 tasks); Multiwoz-T: 9,695 dialogues (35 tasks) SGD-S: 15.57 turns avg; SGD-M: 21.68 turns avg; Multiwoz-T: 13.94 turns avg Three task-divided dialogue datasets constructed from the Schema-Guided Dialogue (SGD) and MultiWoz corpora for evaluating Task-Oriented Dialogue Clustering (TODC). Dialogues are grouped into tasks based on matching sets of active intents, yielding single-domain (SGD-S), multi-domain (SGD-M), and MultiWoz-derived (Multiwoz-T) splits with task-level cluster labels. Lv et al. 2021
Reddit Trendings English Text Text (crawled dialogue threads) Open-domain chit-chat about trending topics Human-Human 407 dialogues over 40 trending topics A test set of real-world dialogues crawled from the Reddit Trendings panel in 2021, covering 40 hot topics (e.g., recent news entities) that are largely absent from standard knowledge bases. It is designed to evaluate dialogue generation models on unseen, out-of-KB entities in practical settings. Cui et al. 2021
CI-ToD English Text Text (dialogue history, system responses, knowledge base entries, fine-grained inconsistency labels) Consistency identification in task-oriented dialogue (navigation, weather, calendar scheduling) Human-System 3,190 dialogues (2,553 train / 319 validation / 318 test) 3.693 CI-ToD is a human-annotated dataset for Consistency Identification in Task-oriented Dialogue systems, built on top of the KVRET corpus. Each sample is labeled with a single overall consistency label as well as fine-grained inconsistency source labels (Dialogue History Inconsistency, User Query Inconsistency, and Knowledge Base Inconsistency), enabling models to identify both whether and why a system response is contradictory. Qin et al. 2021
ToDCL English Text Text (dialogue transcripts with intent, dialogue state, and NLG annotations) Task-oriented dialogue across 37 domains (e.g., restaurant, hotel, flight, music, ridesharing), spanning intent recognition, dialogue state tracking, NLG, and end-to-end settings Human-System 40,287 dialogues (31,426 train / 4,043 valid / 4,818 test); 347,885 train + 44,047 valid + 53,641 test input-output pairs in E2E setting; 37 domains, 280 intents 16.23 A continual learning benchmark for task-oriented dialogue systems, assembled by jointly pre-processing four existing datasets (TaskMaster 2019, TaskMaster 2020, MultiWoZ, and Schema-Guided Dialogue) into a unified curriculum of 37 domains. It supports four learning settings—intent recognition, dialogue state tracking, NLG, and end-to-end—to facilitate research on catastrophic forgetting and continual learning in dialogue systems. Madotto et al. 2021
TOD-Dravidian Test Set Kannada, Tamil Text Text, intent and slot annotations (BIO notation) Task-oriented dialogue (intent detection and slot filling) Human-annotated (native speaker graduate students) 600 utterances (300 per language) A manually curated, gold-standard test dataset for task-oriented dialogue (intent detection and slot filling) in two low-resource Dravidian languages, Kannada and Tamil. Utterances were translated from the Facebook Multilingual Task-Oriented Dialog dataset and annotated by native-speaker graduate students with inter-annotator Cohen’s Kappa scores of 0.93 (Kannada) and 0.96 (Tamil). Kanakagiri et al. 2021
Gutenberg Dialogue Dataset Multilingual (English, German, Dutch, Spanish, Italian, Hungarian, Portuguese) Text Text (dialogues extracted from public-domain books) Open domain (fiction/literary dialogue) Human-Human English: 14,773,741 utterances, 2,526,877 dialogues; German: 226,015 utterances, 43,440 dialogues; Dutch: 129,471 utterances, 23,541 dialogues; Spanish: 58,174 utterances, 6,912 dialogues; Italian: 41,388 utterances, 6,664 dialogues; Hungarian: 18,816 utterances, 2,826 dialogues; Portuguese: 16,228 utterances, 2,233 dialogues English: 5.85; German: 5.20; Dutch: 5.50; Spanish: 8.42; Italian: 6.21; Hungarian: 6.66; Portuguese: 7.27 A high-quality open-domain dialogue dataset extracted from public-domain books on Project Gutenberg using an automated pipeline with multiple heuristic filtering steps. The dataset contains 14.8M utterances in English and smaller datasets (20K–226K utterances) in German, Dutch, Spanish, Italian, Hungarian, and Portuguese, offering a better size-quality trade-off than existing corpora such as Opensubtitles. Csaky et al. 2021
FewShotSGD English Text Text (MR-to-Text pairs, meaning representations and natural language utterances) Task-oriented dialogue NLG (Natural Language Generation from meaning representations), covering 16 domains including Restaurants, Hotels, Flights, Calendar, Banks, Weather, Buses, Events, Homes, Media, Movies, Music, Rentalcars, Ridesharing, Services, and Travel Human-authored (derived from the Schema-Guided Dialogue corpus) 16 domains; avg. ~35 training instances and ~5,618 test instances per domain; avg. ~31 delexicalized MRs in training and ~31 in testing per domain FewShotSGD is a few-shot NLG benchmark dataset constructed by applying the same preparation steps as FewShotWOZ to the Schema-Guided Dialogue (SGD) corpus, covering 16 domains. It features fewer training instances per domain (avg. ~35) and a higher novelty of test n-grams compared to FewShotWOZ, making it a more challenging few-shot NLG testbed. Xu et al. 2021
DECODE English Text Text (human-written dialogues with annotated contradictions and supporting evidence utterances; human-bot dialogues with contradiction labels) Open-domain dialogue contradiction detection; consistency evaluation across multiple conversational domains (knowledge, emotion, persona, general chit-chat) Human-Human and Human-System (Human-Bot) Main train: 27,184 dialogues; Main dev: 4,026 dialogues; Main test: 4,216 dialogues; Human-Bot test: 764 dialogues; A2T auxiliary test: 2,079 dialogues; RCT auxiliary test: 2,011 dialogues. 17,713 human-written contradicting dialogues collected in total. DECODE (DialoguE COntradiction DEtection) is a conversational dataset for dialogue contradiction detection, containing human-written dialogues in which one speaker deliberately contradicts prior utterances, along with annotator-verified supporting evidence. It includes both human-human and human-bot dialogue subsets spanning multiple open-domain topics, with balanced contradiction/non-contradiction labels and auxiliary diagnostic test sets. Nie et al. 2021
Snips-NSD and ATIS-NSD English Text Text (utterances with slot labels including novel/unknown slot annotations) Task-oriented dialogue slot filling / novel slot detection Human-System Snips-NSD (15% split): 9,329 train / 700 dev / 700 test utterances; ATIS-NSD: 4,478 train / 500 dev / 893 test utterances (base); multiple splits at 5%, 15%, 30% unknown slot proportions Two benchmark datasets for Novel Slot Detection (NSD) in task-oriented dialogue, derived from Snips and ATIS slot filling datasets. Each dataset re-annotates a portion of slot types (5%, 15%, or 30%) as unknown/out-of-domain novel slots to support research on detecting previously unseen slot types at inference time. Wu et al. 2021
BMELD English, Mandarin Chinese Text Text (bilingual dialogue utterances with manual translations) Bilingual conversational chat translation (English↔Chinese); derived from the MELD emotion dialogue dataset Human-Human 1,418 dialogues (train/valid/test); 13,672 utterances across En⇒Ch and Ch⇒En directions (train: 5,560 En⇒Ch + 4,427 Ch⇒En; valid: 567 + 517; test: 1,466 + 1,135) BMELD (Bilingual MELD) is a bilingual dialogue corpus for English↔Chinese chat translation, constructed by crawling Chinese translations of the MELD dataset and having them manually post-edited by native Chinese speakers according to dialogue history. It simulates bilingual conversations where 50% of speakers are assigned as Chinese speakers, and is released publicly to support research on neural chat translation. Liang et al. 2021
Multi-Modal Dialogue Dataset English Multimodal (text and image) Text dialogues with semantically relevant images replacing selected utterances Open-domain chit-chat / multi-modal dialogue Human-Human 45K dialogues (39,956 train / 2,401 valid / 2,673 test instances); 12,272 unique training images ~13 turns per dialogue (13.01 train, 13.62 valid, 13.59 test) A 45K multi-modal dialogue dataset constructed semi-automatically by replacing semantically relevant sentences in existing text-only dialogue datasets (DailyDialog, Persona-Chat, EmpatheticDialogues) with contextually coherent images sourced from MS-COCO and Flickr 30k, using text-to-image similarity and contextual-similarity-based filtering. The dataset is designed for training and evaluating multi-modal dialogue systems that must jointly understand images and dialogue context. Lee et al. 2021
Document-aligned Japanese-English Conversation Parallel Corpus (BSD+AMI+ON) Japanese, English Text Text (parallel sentence- and document-aligned conversation transcripts with speaker, scene, and ambiguity-type metadata) Business conversations, meetings, broadcast conversation, telephone conversation; Machine Translation Human-Human (multi-party) ~219,000 sentence pairs across three sub-corpora: BSD ~84,800 sentence pairs (42,400 JA→EN + 42,400 EN→JA), AMI 110,483 sentence pairs, ON 28,429 sentence pairs; development set 2,051 sentence pairs; evaluation set 2,120 sentence pairs A document-aligned Japanese-English parallel corpus of multi-party conversations comprising three sub-corpora: an expanded Business Scene Dialogue (BSD) corpus written by professional scenario writers and translated by professional translators, a Japanese translation of the AMI Meeting Corpus, and a Japanese translation of broadcast and telephone conversation subsets of OntoNotes 5.0. The corpus includes speaker information, scene metadata, and a balanced development/evaluation split with annotations of context-dependent linguistic phenomena (zero anaphora, phrase ambiguity) to support document-level machine translation research. Rikters et al. 2020
PhotoChat English Multimodal (text and image) Text (dialogue transcripts), Images Photo sharing in online messaging / open-domain chat Human-Human 12,286 dialogues, 10,917 unique images, 156,099 turns, 988,215 tokens 12.7 turns per dialogue (9.5 when consecutive same-speaker turns are merged) PhotoChat is the first human-human dialogue dataset capturing photo-sharing behavior in online messaging, collected via crowdsourcing. Each of the 12,286 dialogues is paired with a user photo (drawn from Open Images V4) that is shared during the conversation, supporting tasks such as photo-sharing intent prediction and dialogue-based image retrieval. Zang et al. 2021
NeuralWOZ Synthetic Corpus English Text Text (synthetically generated dialogues with dialogue state and active domain annotations) Task-oriented dialogue; multi-domain travel (attraction, hotel, restaurant, taxi, train) Human-System (model-simulated user and system via Collector and Labeler modules) Up to 5,000 dialogues per domain for zero-shot experiments; 1,000 dialogues for full augmentation; e.g., ~38K–46K turns and ~870K–1.1M tokens per domain batch (see Table 7) ~7.5–9.1 turns per dialogue (derived from Table 7 per-domain statistics) A synthetically generated task-oriented dialogue corpus produced by NeuralWOZ, a model-based dialogue simulation framework consisting of a Collector (BART-based dialogue generator) and a Labeler (RoBERTa-based annotation model). Dialogues are grounded in natural-language goal instructions and knowledge base API call results, and are automatically annotated with dialogue states and active domain labels for use in zero-shot and few-shot domain transfer learning for dialogue state tracking. Kim et al. 2021
CrossWOZ Mandarin Chinese Text Text (utterances), structured dialogue states, dialogue acts (user and system sides), user goals Cross-domain task-oriented dialogue for tourism in Beijing (hotel, restaurant, attraction, metro, taxi) Human-Human (Wizard-of-Oz) 6,012 dialogues, 102K utterances, ~1.65M tokens (train+valid+test); training set: 5,012 dialogues, 84,692 turns, 1,376,033 tokens 16.9 CrossWOZ is the first large-scale Chinese cross-domain Wizard-of-Oz task-oriented dialogue dataset, containing 6,012 dialogue sessions and 102K utterances across 5 domains (hotel, restaurant, attraction, metro, taxi). It features rich annotation of dialogue states and dialogue acts on both user and system sides, with approximately 60% of dialogues involving cross-domain user goals that require natural inter-domain transitions. Zhu et al. 2020
COM / ONTO-COM English Text Meaning representation (MR) / utterance pairs Restaurant information and recommendation (task-oriented NLG) Human-System ~77K training MR/utterance pairs; 3,040 test MRs (COM); 3,040 additional test MRs (COM-2) A combined-ontology NLG dataset for the restaurant domain created by merging two existing datasets (NYC and E2E) into a new larger ontology (ONTO-COM). The dataset includes a ~77K balanced training set remapped to the combined ontology and two held-out test sets (COM and COM-2, 3,040 MRs each) whose MRs always combine attributes from both source ontologies, providing combinations never seen in training. Reed et al. 2020
JSL Dialogue Corpus Japanese Sign Language (JSL) Sign Language Video (multimodal: manual signs, mouth movements, non-manual movements, gaze) Video recordings, word gloss annotations, utterance unit annotations, mouth movement tiers, non-manual movement tiers, gaze tiers (annotated in ELAN) Spontaneous dialogue (animation narrative, curry recipe explanation, personal narrative); sign language conversation Human-Human (Deaf signer pairs) 60 dialogues, 120 participants, ~40 hours 52 minutes total recording (15 hours 43 minutes for dialogue tasks); 27,371 annotated word gloss tokens across 85 annotated files A corpus of spontaneous Japanese Sign Language (JSL) dialogues collected from 120 Deaf signers across 7 Japanese prefectures (2012–2016), comprising 60 dyadic dialogues (~40 hours of video). The corpus is annotated in ELAN using a novel multimodal ‘utterance unit’ scheme with tiers for word glosses, mouth movements, non-manual movements, and gaze, developed from a Conversation Analysis perspective to identify interactional boundaries in sign language without reliance on spoken language writing systems. Bono et al. 2020
SMCalFlow English Text Text dialogues with executable dataflow program annotations per turn Calendar events, weather, places, and people (task-oriented, cross-domain) Human-WoZ 41,517 dialogues, 155,923 user turns SMCalFlow is a large-scale English task-oriented dialogue dataset collected via a Wizard-of-Oz process, featuring complex, open-ended conversations about calendar events, weather, places, and people. Each dialogue turn is annotated with an executable dataflow program (including metacomputation operators for reference and revision) representing the agent’s response to the user’s intent. Andreas et al. 2020
ILLC-IER + Low-level Image Editing Dialogues English Text Text (typed natural language utterances, image edit requests, dialogue turns) Natural language image editing (low-level attribute adjustment of localized image regions) Human-System 2,537 ILLC-IERs (32,194 tokens, 1,034 unique tokens); 83 dialogues with 1,359 user utterances (2,753 tokens, 534 unique) 17.4 turns per dialogue Two datasets collected to support low-level natural language image editing research: (1) the ILLC-IER dataset of 2,537 crowd-sourced Imperative Low-Level Complete Image Edit Requests (split 2,055/242/240 train/dev/test), annotated with ACTION, REFER, ATTRIBUTE, and VALUE BIO tags; and (2) 83 task-oriented dialogues with 1,359 user utterances collected via a user study in which AMT workers interacted with a dialogue system to perform low-level image attribute adjustments. Lin et al. 2020
Margarita Dialogue Corpus English Multimodal (text and video) Text transcripts, annotated dialogues, video clips Time-Offset Interaction / conversational avatar QA (personal information and university information domains) Human-Human (interrogators and avatar maker) 892 KB question-answer pairs (431 unique answers), 20 annotated dialogues, 659 dialogue Q-A pairs, 60,860 total words across KB and dialogues 33 (avg. turns per dialogue; range: 22–36 across subsets) The Margarita Dialogue Corpus is a dataset for Time-Offset Interaction Applications (TOIAs), comprising a knowledge base of 892 question-answer pairs (with corresponding answer video clips) and 20 annotated human-human dialogues recorded between random interrogators and an avatar maker. The corpus supports research in unstructured multi-turn dialogue, answer retrieval, and conversational avatar systems, and includes two interaction modes: university information (EDU) and personal conversation (PER). Chierici et al. 2020
Cheese! French Multimodal (audio and video) Audio recordings, video recordings, orthographic transcriptions, automatic annotations (phonemes, syllables, tokens, POS tags), manual annotations of smiling and humor Spontaneous face-to-face dyadic conversation; smiling and conversational humor Human-Human (dyadic, mixed and non-mixed pairs of French native university students) 11 dialogues; 5 interactions fully annotated; 2,130 unique words; 20,201 word occurrences; ~165 minutes total audio/video ~15 minutes per interaction Cheese! is a multimodal corpus of 11 face-to-face French dyadic conversations (~15 minutes each), recorded at the LPL laboratory in 2016. It was designed for cross-cultural comparison of smiling behavior in humorous and non-humorous sequences, and includes high-quality audio/video recordings along with orthographic transcriptions, automatic linguistic annotations (via SPPAS and MarsaTag), and manual annotations of smile intensity (using the Smiling Intensity Scale) and conversational humor. Priego-Valverde et al. 2020
KOMODIS English Text Text (chat transcripts with fact and opinion profile annotations, named entity and sentiment labels) Movie discussions (open-domain chit-chat grounded in movie facts and opinions) Human-Human 7,519 dialogues, 103,500 utterances, 1,487,284 tokens 13.8 KOMODIS (Knowledgeable and Opinionated MOvie DIScussions) is a crowd-sourced dialogue dataset collected via Amazon Mechanical Turk in which each dialogue is grounded in pre-specified IMDb-derived facts and discrete opinion profiles about movies and related entities. Every dialogue is annotated with named entity mentions and sentiment labels derived from validation of participant adherence to their assigned profiles, covering 500 movies across 7,519 conversations. Galetzka et al. 2020
French Medical Conversations Corpus (LabForSIMS2) French Text Transcripts of medical consultation dialogues, annotated with question/response category tags Medical consultation / virtual patient (surgical emergency for abdominal pain) Human-System (medical interns interacting with a Virtual Standardized Patient) Single-turn dataset: 1 consultation, 5,402 sentences, 2,733 vocabulary items; Context QA dataset: 41 dialogues, 1,818 sentences, 812 vocabulary items ~44 turns per dialogue (1,818 sentences across 41 dialogues) A French annotated corpus of doctor–patient medical consultation dialogues built for virtual patient dialogue systems. It comprises a single-turn QA dataset and a context QA dataset (41 end-to-end dialogues collected from medical interns interacting with a virtual standardized patient), annotated with seven semantic categories covering aim of consultation, personal data, medical history, symptoms, lifestyle, treatments, and other. Laleye et al. 2020
PACO French Multimodal (Audio and Video) Audio, Video, Speech transcripts (Enriched Orthographic Transcription), IPU segmentation, smile intensity annotations Spontaneous face-to-face dyadic conversation; study of common ground, topic transitions, and smile behavior Human-Human (dyadic; strangers meeting for the first time) 15 dialogues, ~5 hours of conversational data, 30 participants PACO is a French audio-video corpus of 15 face-to-face dyadic interactions (~20 min each, totalling ~5 hours) between participants who did not know each other, designed to study the impact of personal common ground on conversational organization and smile behavior during topic transitions. It replicates the protocol of the “Cheese!” corpus (friends condition) and includes speech transcriptions, IPU segmentation, and semi-automatic smile intensity annotations using the Smiling Intensity Scale. Amoyal et al. 2020
Hindi Customer Care Conversational Dataset (courteousH) English, Hindi Text Text (Twitter conversations with generic and courteous/polite response pairs, annotated as informative, courteous, or hybrid) Customer care / customer support (complaint handling, suggestions) Human-Human (customers and customer care agents on Twitter) English: 200,300 conversations, 256,014 utterances; Hindi: 65,094 conversations, 97,488 utterances (combined train/valid/test splits) A large-scale multilingual conversational dataset (English and Hindi) collected from Twitter, comprising real interactions between customers and customer care agents, with each utterance annotated as informative, courteous, or hybrid, and paired generic and polite/courteous response versions. The Hindi portion is a newly created resource; the English portion is based on prior work (Golchha et al., 2019). Firdaus et al. 2020
AIA-BDE Portuguese Text Text (FAQ questions, answers, and question variations/paraphrases) FAQ retrieval and question answering; economic activities and entrepreneurship services in Portugal Human-authored (original FAQs from Balcão do Empreendedor); variations created manually by human volunteers or automatically via Google Translate 380 FAQs with 380 VG1 variations, 380 VG2 variations, and 936 manual (VUC) variations; 1,696 total question variants AIA-BDE is a corpus of 380 domain-oriented FAQs in Portuguese drawn from the Portuguese Entrepreneur’s Desk (Balcão do Empreendedor), covering three service domains (RJACSR, AL, PE), each paired with question variations created either automatically via round-trip Google Translate (VG1, VG2; 380 each) or manually by native-speaker volunteers (VUC; 936 variations). It is intended as a benchmark for FAQ retrieval, automatic question answering, task-oriented dialogue systems, and natural language inference in interrogative contexts. Gonçalo Oliveira et al. 2020
Dicta-Sign-LSF-v2 French Sign Language (LSF) Video (RGB) Video recordings, framewise annotations (lexical and non-lexical manual units), preprocessed body/face/hand pose features Spontaneous dialogue on the theme of travel; Sign Language Processing (recognition of lexical and non-lexical structures) Human-Human (face-to-face dialogue pairs) 94 videos, 16 signers, 11 hours (~1,007,593 frames), ~35,000 manual units Dicta-Sign-LSF-v2 is a remake of the French Sign Language portion of the multilingual Dicta-Sign corpus, comprising 11 hours of video dialogue from 16 signers across 8 pairs performing 9 loosely constrained tasks on the theme of travel. It provides cleaned lexical and non-lexical (depicting signs, pointing signs, fragment buoys, fingerspelling, numbering) manual-unit annotations totalling ~35,000 units, along with preprocessed 3D body, face, and hand pose features. Belissen et al. 2020
E-Commerce Customer Service Dialogue Dataset (CHG) Mandarin Chinese Text Text (customer-seller dialogue transcripts) Customer service / E-commerce (clothing domain) Human-Human 60,000 multi-turn dialogues 9 utterances per dialogue (avg. 27 characters per utterance) A real-world multi-turn customer service dialogue dataset collected from a top Chinese online shopping platform in the clothing domain. Each dialogue is paired with 1–5 historical dialogues from the same seller, partitioned into 80/10/10 train/validation/test splits. Zhang et al. 2020
MDMMD (Multi-domain Multi-modal Dialogue) English Multimodal (text and image) Text, Images, Aspect category and aspect term annotations Task-oriented dialogue across three domains: restaurants, electronics, and furniture Human-WoZ (domain experts as system agents, crowd workers as customer agents) 121,023 dialogues total (99,813 train / 11,081 valid / 10,129 test); 208,609 utterances (train) / 121,718 (valid) / 87,436 (test); vocabulary size 45,453 ~20.7 (train: 20.9, valid: 19.6, test: 20.7) MDMMD is a large-scale multi-domain, multi-modal task-oriented dialogue dataset containing dyadic conversations across restaurant, electronics, and furniture domains, with both textual utterances and images. Each utterance is annotated with aspect categories and aspect terms to support aspect-guided response generation research. Firdaus et al. 2020
VFD (Visually-grounded First-person Dialogue Dataset) Japanese Multimodal (text and image) Text (utterances and verbal/non-verbal responses), First-person images, Eye-gaze location annotations Visually-grounded first-person dialogue; task-oriented and non-task-oriented Human-System 308,793 verbal dialogues, 81,867 non-verbal dialogues; based on 34,775 first-person images 1 (single-turn: one human utterance + one agent verbal and/or non-verbal response) VFD is a large-scale Japanese multimodal dialogue dataset in which human utterances and agent verbal and non-verbal responses are manually annotated for first-person images drawn from the GazeFollow dataset, supplemented with eye-gaze location annotations. It supports research on visually-grounded dialogue understanding and response generation, including both verbal replies and non-verbal (action) responses. Kamezawa et al. 2020
doc2dial English Text Text (dialogue utterances annotated with dialogue acts and document grounding spans, plus associated HTML and plain-text documents) Goal-oriented information-seeking dialogue grounded in government service documents (ssa.gov, va.gov, dmv.gov, cdc.gov) Human-Human (crowdsourced agent and user roles) 4,470 dialogues, ~69,820 turns, grounded in 458 documents across four domains 14 doc2dial is a goal-oriented, document-grounded dialogue dataset in which conversations between an assisting agent and a user are grounded in public government service web documents. Dialogues are constructed via a pipeline that generates dialogue flows from document structure and discourse relations, which are then converted into natural utterances by crowdworkers; each turn is annotated with a dialogue act and a reference span in the grounding document. Feng et al. 2020
CIMA English (with Italian target language content) Text Text (dialogue utterances with action type labels) Foreign language tutoring (English speakers learning Italian vocabulary and prepositional phrases) Human-Human (crowdworkers role-playing as students and tutors) 741 exercises (350 Shape + 391 Prepositional Phrase); 5,850 tutor responses (2,970 Shape + 2,880 Prepositional Phrase) Shape: 3.09 turns/exercise; Prepositional Phrase: 3.65 turns/exercise CIMA (Conversational Instruction with Multi-responses and Actions) is a large open-access collection of tutoring dialogues in two sub-datasets (Shape and Prepositional Phrase) collected asynchronously via crowdworkers role-playing as students and tutors. It is notable for providing multiple (three) distinct tutor responses per student conversational turn and dialogue-level action type labels for both student and tutor utterances, supporting training of next-utterance generation models conditioned on tutoring action strategies. Stasaski et al. 2020
KdConv Mandarin Chinese Text Text (dialogue utterances with sentence-level knowledge graph annotations) Knowledge-driven open-domain conversation across three domains: film, music, and travel Human-Human 4.5K dialogues, 86K utterances (85,596); 1,500 dialogues per domain; split 8:1:1 into train/dev/test 19.0 KdConv is a Chinese multi-domain knowledge-driven conversation dataset grounding multi-turn dialogues to domain-specific knowledge graphs across film, music, and travel domains. Each utterance is annotated at the sentence level with the knowledge triples it draws upon, enabling research on knowledge planning, knowledge grounding, and domain adaptation in open-domain conversational systems. Zhou et al. 2020
Non-Conversational Text Corpus Chinese Text Text (forum comments, idioms/quotes/proverbs, book snippets) Open-domain dialogue generation augmentation; non-conversational text covering diverse daily-life topics 1,040,135 utterances (781,847 forum comments; 51,948 idioms/quotes; 206,340 book snippets) A large-scale Chinese non-conversational text corpus collected from three sources — Zhihu forum comments, idioms/famous quotes/proverbs, and highlighted book snippets from WeChat Read — intended to augment open-domain dialogue generation with more diverse and topically broad content. Utterances are filtered for offensive language and constrained to 10–30 words in length. Su et al. 2020
WikiHow Intent Detection Dataset English, Spanish, Thai Text Text (goal-step pairs from wikiHow articles formatted as 4-choice multiple-choice intent detection examples) Intent detection pretraining across broad open domains (instructional/how-to content) Human-System 107,298 English examples, 64,803 Spanish examples, 6,342 Thai examples A multilingual pretraining dataset derived from the wikiHow instructional website, in which each example pairs a wikiHow step (approximating a user utterance) with its corresponding article goal (approximating an intent) in a 4-choice multiple-choice format. Covers English, Spanish, and Thai and is intended to enable robust zero- and few-shot intent detection across diverse domains. Zhang et al. 2020
Repository of Conversational Datasets (Reddit, OpenSubtitles, AmazonQA) English (OpenSubtitles also available in 62 languages) Text Text (conversational context–response pairs stored as TensorFlow records) Open domain conversational response selection; three sub-domains: social media (Reddit), movie/TV subtitles (OpenSubtitles), and e-commerce Q&A (AmazonQA) Human-Human Reddit: 727M examples (654M train, 72.6M test); OpenSubtitles: 316.9M examples (283.7M train, 33.2M test); AmazonQA: 3.7M examples (3.3M train, 373K test); total: ~1.04 billion context–response pairs A public repository of three large conversational datasets (Reddit, OpenSubtitles, AmazonQA) comprising hundreds of millions of context–response pair examples, together with reproducible preprocessing scripts and a standardised 1-of-100 accuracy evaluation framework for conversational response selection models. Henderson et al. 2019
ViGGO English Text Structured meaning representations (MRs) and crowdsourced reference utterances Data-to-text natural language generation; video game domain; open-domain conversation Human-System 6,900 MR-utterance pairs, 2,253 unique MRs, 3 references per MR ViGGO is a parallel data-to-text NLG corpus in the video game domain, comprising 6,900 crowdsourced and manually cleaned MR–utterance pairs covering 9 conversational dialogue act types and 14 slot types across more than 100 video game titles. It is designed to support open-domain dialogue systems with more conversational and linguistically diverse utterances than prior task-oriented NLG datasets. Juraska et al. 2019
DREAM English Text Text (dialogue transcripts with multiple-choice reading comprehension questions) Dialogue-based reading comprehension (multiple-choice QA over multi-turn multi-party dialogues from English language examinations) Human-Human (multi-party dialogues, avg. 2.0 speakers per dialogue) 6,444 dialogues, 10,197 multiple-choice questions, 30,183 turns 4.7 DREAM is the first dialogue-based multiple-choice reading comprehension dataset, collected from English as a Foreign Language examinations designed by human experts to assess Chinese learners of English. It contains 10,197 three-way multiple-choice questions for 6,444 multi-turn multi-party dialogues, with 85% of questions requiring multi-sentence reasoning and 34% requiring commonsense knowledge. Sun et al. 2019
TreeNLG Weather Dataset English Text Text (user queries, tree-structured meaning representations, natural language responses, span-level response annotations) Weather domain NLG; generating natural language responses to weather-related user queries from tree-structured semantic representations Human-System (crowdworker-annotated responses to system-generated MRs) 33,493 examples (25,390 training, 3,121 test); vocabulary size 1,485 40.6 tokens average response length (not turn count; single-turn NLG pairs) A task-oriented NLG dataset for the weather domain in which each example comprises a user query, synthetic user context (datetime and location), a tree-structured meaning representation (MR) encoding discourse relations (JOIN, CONTRAST, JUSTIFY) and dialog acts, a natural language response, and a complete tree-structured span annotation of the response. The dataset was collected via crowdsourcing and quality-filtered, yielding 33,493 examples ranging from simple single-act MRs to complex nested structures of depth and width up to 4. Balakrishnan et al. 2019
IDS Benchmark Dataset Chinese (English translation provided) Text Text (rule-generated dialogues) Customer service (product query, purchase, delivery, after-sales, emotional utterances) Human-System 5 sub-datasets; each with 20,000 training, 5,000 validation, and 5,000 test dialogues (125,000 dialogues total) 9.8–12.4 utterances per dialogue (varies by sub-dataset) A rule-generated Chinese task-oriented dialogue benchmark consisting of five incremental sub-datasets (SubD1–SubD5) in a customer service domain, designed to simulate unanticipated user needs at deployment time. Each subsequent sub-dataset covers a superset of dialogue scenarios (e.g., product queries, after-sales service, emotional utterances), enabling evaluation of dialogue systems’ robustness to unconsidered user actions. Wang et al. 2019
PhotoBook English Text (chat messages) Text (chat utterances, image labelling actions, timestamps, participant IDs, self-reported collaboration scores) Visually-grounded collaborative image identification; referring expression generation and resolution Human-Human 2,506 dialogues (games); 164,615 utterances; 130,322 actions; 11,805 unique tokens; 18,321 reference chains; 44,669 dialogue segments approximately 65 utterances per dialogue (164,615 utterances / 2,506 games) The PhotoBook dataset is a large-scale collection of 2,506 visually-grounded, task-oriented human-human dialogues in English, collected via Amazon Mechanical Turk. Two participants play a collaborative five-round game in which they identify shared images from MS COCO by chatting, producing rich reference chains that track how referring expressions are established and refined over shared dialogue history. Haber et al. 2019
BANKING English Text Text (FAQ question-answer pairs) E-banking customer support FAQ (e.g., card activation, closing account, refund request) Human-System 11,880 question-answer pairs total (10,395 train, 1,485 test); 77 intent categories, 10 paraphrases per question A FAQ-style dataset for the e-banking domain comprising question-answer pairs divided into 77 unique intent categories (e.g., “card activation”, “closing account”, “refund request”). Each question has 10 paraphrases mapping to the same answer, and the dataset is split into training (70%), validation (20%), and test (10%) portions. Henderson et al. 2019
IRC Disentanglement Corpus English Text Text (IRC chat logs manually annotated with reply-structure graphs) Conversation disentanglement; technical support (Ubuntu and Linux IRC channels) Multi-party human 77,563 messages (74,963 from #Ubuntu IRC, 2,600 from #Linux IRC); sampled across 173 time points (2004–2018); splits: ~47,500 train (uniform) + ~18,963 train (1-hr spans) + 1,000 train (agreement subset) + 2,500 dev + 5,000 test + 2,600 out-of-domain A large-scale manually annotated corpus of IRC chat messages from the #Ubuntu and #Linux channels, in which each message is labeled with reply-to relations forming conversation-structure graphs suitable for conversation disentanglement research. At 77,563 messages it is 16 times larger than all previously released disentanglement datasets combined, and is the first to include context windows and adjudicated annotations for development and test sets. Kummerfeld et al. 2019
CYCCD (Courteously Yours Customer Care Dataset) English Text Text (tweets; paired generic and courteous agent responses with conversation history) Customer care / complaint and support interactions on Twitter Human-Human 200,300 conversations (train: 140,203; valid: 20,032; test: 40,065); 256,914 utterances total CYCCD is a large conversational dataset of real Twitter interactions between customers and professional customer care agents, providing paired “generic” (neutral/informative-only) and “courteous” forms of agent responses alongside conversation history. It was constructed by manually annotating and filtering courteous expressions from actual customer care tweets, with annotator agreement (Kappa ~80%). Golchha et al. 2019
Korean Online Counseling Dialogue Corpus Korean Text Text (counselor-client chat dialogues with utterance-level category annotations) Text-based online psychotherapy / counseling (cognitive behavioral therapy) Human-Human (professional counselor and client) 1,448 total dialogues; 100 labelled dialogues; 21,100 annotated triples (train/valid/test: 14,679/3,166/3,165) ~163 counselor utterances and ~239 client utterances per session (labelled dialogues) A Korean text-based online counseling corpus collected from the Trost platform, comprising 1,448 anonymised counselor-client dialogues, of which 100 are annotated at the utterance level with five CBT-grounded client utterance categories (Factual Information, Anecdotal Experience, Appealing Problem, Psychological Change, Counseling Process) by professional counselors. Park et al. 2019
Clothes & Makeup CS Dialogue Datasets Mandarin Chinese Text Text (dialogue transcripts with dialogue-level satisfaction ratings and utterance-level sentiment labels) E-commerce customer service (clothing and makeup domains); service satisfaction analysis Human-Human 13,540 dialogues total (Clothes: 10,000 dialogues, 123,242 utterances; Makeup: 3,540 dialogues, 46,255 utterances) ~26 (Clothes: 25.99, Makeup: 26.67) Two Chinese multi-turn customer service dialogue datasets collected from a top E-commerce platform (Taobao), covering the Clothes and Makeup domains. Each dialogue is annotated with a three-class service satisfaction label (well satisfied, met, unsatisfied) derived from 1–5 star customer ratings, and all customer and server utterances are annotated with three-class sentiment labels (positive, neutral, negative). Song et al. 2019
DyKgChat Multilingual (Mandarin Chinese and English) Text Text (dialogue scripts paired with dynamic knowledge graphs including entities and relation triplets) Knowledge-grounded conversation generation; TV series dialogue (Chinese palace drama and English sitcom) Human-Human 4,339 dialogues total (1,247 HGZHZ + 3,092 Friends); 74,921 turns total (17,164 HGZHZ + 57,757 Friends); 1,301,560 tokens total (462,647 HGZHZ + 838,913 Friends) 13.76 (HGZHZ), 18.68 (Friends) DyKgChat is a TV series conversation corpus comprising a Chinese palace drama (Hou Gong Zhen Huan Zhuan) and an English sitcom (Friends), each paired with manually constructed dynamic knowledge graphs containing character relation triplets. It is designed to benchmark dialogue generation models on their ability to zero-shot adapt to updated, unseen knowledge graphs. Tuan et al. 2019
GoRecDial English Text Text (dialogue utterances, movie recommendations, accept/reject decisions, engagingness ratings) Goal-oriented movie recommendation Human-Human 9,125 dialogues, 170,904 utterances (81,260 conversation turns per abstract) 23.0 GoRecDial is a large-scale goal-driven recommendation dialogue dataset in which pairs of Amazon Mechanical Turk workers play a cooperative game: an “expert” must identify and recommend the correct movie (grounded in real MovieLens user preferences) to a “seeker” through natural language conversation, while avoiding incorrect candidates. The dataset includes dialogue utterances, recommendation and accept/reject actions, justifications, and engagingness ratings. Kang et al. 2019
CoSQL English Text Text (user utterances, SQL queries, system responses, dialogue act annotations) Cross-domain database querying via natural language (Text-to-SQL conversational interfaces) Human-WOZ (crowd worker as DB user, SQL expert as wizard) 3,007 dialogues, 31,148 turns, 10,000+ annotated SQL queries, spanning 200 databases across 138 domains 10.36 (training set); ~80% of dialogues have 8 or more turns CoSQL is a large-scale cross-domain Wizard-of-Oz conversational text-to-SQL corpus for building general-purpose database querying dialogue systems. It contains 3,007 dialogues over 200 complex databases spanning 138 domains, with annotated SQL queries, dialogue acts, and natural language system responses, supporting three tasks: SQL-grounded dialogue state tracking, response generation from query results, and user dialogue act prediction. Yu et al. 2019
CamRest676-GECOR English Text Text transcripts with ellipsis and co-reference annotations Restaurant search (task-oriented dialogue) Human-Human 676 dialogues, 2,744 user utterances, 1,174 ellipsis versions, 1,209 co-reference versions An annotated dataset for ellipsis and co-reference resolution in multi-turn task-oriented dialogue, constructed by manually annotating the public CamRest676 restaurant-domain dataset. Each user utterance is labelled and supplemented with pragmatically complete versions resolving ellipsis and/or co-reference, enabling both standalone resolution model training and multi-task learning with end-to-end dialogue systems. Quan et al. 2019
Indonesian Conversational SRL Dataset Indonesian Text Text (chat logs annotated with semantic role labels and entity recognition tags) Semantic Role Labeling and Entity Recognition in conversational chat logs between a virtual friend bot and human users Human-System 6,057 sentences containing predicates; 7,330 role instances (AGENT: 2,843; PATIENT: 3,040; BENEFACTOR: 293; GREET: 572; LOCATION: 183; TIME: 399) A low-resource Indonesian conversational corpus of human–chatbot chat logs annotated with semantic roles (a PropBank-derived tagset augmented with a GREET role) and entity recognition labels (PERSON, LOCATION, ORGANIZATION, MISC). Private information is anonymised; annotation was performed by three linguistically trained annotators. Ikhwantri et al. 2018
Stylistic Variation NLG Corpus English Text Text (meaning representations paired with stylistically varied natural language utterances) Restaurant information (task-oriented NLG, restaurant domain) Human-System 88,855 training utterances (3,784 unique MRs × 5 personalities, ~17,771 references per personality); 1,390 test utterances (278 unique MRs × 5 personalities) A large parallel corpus of over 88,000 restaurant-domain utterances synthesized by the PERSONAGE statistical generator, in which the same meaning representations (drawn from the E2E Generation Challenge) are realized in multiple stylistically distinct variants corresponding to five Big Five personality traits (agreeable, disagreeable, conscientious, unconscientious, extravert). The corpus provides total control over both semantic content and stylistic variation, enabling systematic study of style-content disentanglement in neural NLG. Oraby et al. 2018
Sentence Planning for NLG Corpus English Text Text (meaning representation / dialogue act pairs with reference utterances exhibiting sentence planning operations) Task-oriented dialogue response generation (restaurant information); sentence scoping, distributive aggregation, and discourse contrast Human-System ~204,955 utterance/MR pairs (automatically generated via PERSONAGE); supplemented with crowdsourced E2E data; training sets ranging from ~3K to ~64K instances per experiment A systematically constructed corpus of meaning representation (MR) / utterance pairs designed to test neural NLG models on sentence planning operations including sentence scoping, distributive aggregation, and discourse contrast. It combines automatically generated data from the PERSONAGE stylistic generator with crowdsourced E2E data, yielding over 200K utterance/MR pairs with controlled sentence planning phenomena. Reed et al. 2018
Medical Dialogue Dataset for Automatic Diagnosis Mandarin Chinese Text Text (patient self-reports and doctor-patient conversational transcripts with annotated symptoms) Medical diagnosis / symptom collection (pediatric diseases: infantile diarrhea, children’s functional dyspepsia, upper respiratory infection, children’s bronchitis) Human-Human 710 dialogues (user goals): 200 infantile diarrhea, 150 children functional dyspepsia, 160 upper respiratory infection, 200 children’s bronchitis; 144 unique symptoms identified (67 kept with frequency ≥ 10) A Chinese medical dialogue dataset collected from a pediatric online healthcare community, comprising patient self-reports and doctor-patient conversations annotated with explicit and implicit symptoms (BIO tagging + SNOMED CT normalization) across four pediatric disease categories. The dataset is designed to support task-oriented dialogue systems for automatic diagnosis. Liu et al. 2018
Samsung QA English Text Text (question-answer pairs crawled from web pages) Consumer electronics product question answering (Samsung products) Human-System 183,616 QA pairs (163,616 train, 10,000 validation, 10,000 test) A consumer electronics domain question-answer dataset crawled from Samsung Electronics’ official website and crowd QA websites, containing user questions paired with answers from certified company users. Questions cover six top-level product categories (mobile, office, photo, tv/video, accessories, home appliance), with answers averaging ~173 tokens across ~6 sentence groups. Yoon et al. 2018
QC3 (Qatar Computing Conversational Corpus) English Text Text (forum posts annotated with speech act labels) Speech act recognition in asynchronous forum conversations Human-Human 47 conversations, avg. 13.32 comments and 33.28 sentences per conversation 13.32 comments per conversation QC3 is a forum conversation corpus collected from the Qatar Living community Q&A site, annotated at the sentence level with five speech act types (Statement, Question, Suggestion, Response, Polite) using a standard tagset derived from MRDA. Two native English speakers annotated each conversation, with disagreements resolved by a third annotator. Joty et al. 2018
LSDSCC English Text Text (query-response pairs extracted from online forum threads) Open-domain conversational response generation; movie discussion domain Human-Human 738,095 single-turn dialogues; 346,543 multi-turn conversations; 300-query multi-reference test set (each query with ~15 references) 1 (single-turn focus; multi-turn subset also available) LSDSCC is a large-scale domain-specific conversational corpus of high-quality query-response pairs crawled from the Reddit movie discussion board, with thorough preprocessing and cleansing. It includes a 300-query multi-reference test set with human-annotated, group-aware diverse responses, and is accompanied by diversity-oriented evaluation metrics (MaxBLEU, MDS, PDS) for benchmarking neural response generation models. Xu et al. 2018
CraigslistBargain English Text Text (chat transcripts with annotated dialogue acts) Price negotiation over real Craigslist items (housing, furniture, cars, bikes, phones, electronics) Human-Human 6,682 dialogues 9.2 CraigslistBargain is a human-human negotiation dialogue dataset collected via Amazon Mechanical Turk, in which a buyer and a seller negotiate the price of real items scraped from Craigslist across six categories. Compared to prior negotiation datasets, it features longer dialogues, richer vocabulary, and diverse negotiation phenomena such as embellishment, side offers, cheap talk, and appeals to sympathy. He et al. 2018
Twitch-FIFA English Multimodal (video and text chat) Video (live broadcast soccer game footage), Text (live user chat transcripts) Video-grounded dialogue; live soccer game chat (FIFA-18 on Twitch.tv) Multi-party human 15,083 instances total (10,150 train, 2,153 val, 2,780 test); 49 FIFA-18 game videos; ~85.7 total hours of video A many-speaker, video-context dialogue dataset built from live-broadcast FIFA-18 soccer game videos and concurrent user chat streams on Twitch.tv. Each instance consists of a 20-second video clip with its associated chat context and a target response drawn from the immediately following 10-second window, enabling visually-grounded, multi-party dialogue research. Pasunuru et al. 2018
Stanford Multi-turn Multi-domain Dialogue LU Annotation English Text Semantic frame annotations (slot-span alignments) over existing dialogue transcripts Task-oriented dialogue language understanding; Navigation, Scheduling, and Weather domains Human-System Approx. 2,604 annotated utterances across three domains (500 training + 337/212/271 test utterances per domain; 321/201/262 dev utterances per domain) A slot-filling annotation layer added to the Stanford Multi-turn, Multi-domain Dialogue Dataset (Eric and Manning, 2017), assigning semantic slot types to corresponding word spans in utterances across three domains (navigation, scheduling, weather). Annotated by two annotators per dialogue with reported inter-annotator Kappa values of 0.68–0.92. Hou et al. 2018
E-commerce Dialogue Corpus (ECD) Mandarin Chinese Text Text (tokenized conversation transcripts) E-commerce customer service (commodity consultation, logistics, recommendation, negotiation, chitchat) Human-Human 1M training, 10K validation, 10K test context-response pairs 5.51 (train), 5.48 (valid), 5.64 (test) A large-scale Chinese e-commerce dialogue corpus collected from real customer-service conversations on Taobao, covering over 5 conversation types (commodity consultation, logistics, recommendation, negotiation, and chitchat) across more than 20 commodity categories. It is the first publicly released e-commerce dataset for multi-turn dialogue research, with 1M training, 10K validation, and 10K test context-response pairs. Zhang et al. 2018
Deal or No Deal Negotiation Dataset English Text Text (dialogue transcripts, item pool descriptions, agent value functions, output decisions) Multi-issue bargaining / negotiation (dividing books, hats, and balls between two agents) Human-Human 5,808 dialogues, 2,236 unique scenarios, 252 held-out test scenarios (526 test dialogues) 6.6 A large dataset of human-human natural language negotiations collected via Amazon Mechanical Turk, in which two agents with different, private reward functions must agree on how to divide a pool of items (books, hats, balls) through multi-turn dialogue. Each dialogue is paired with the agents’ input goal specifications and their output division decisions. Lewis et al. 2017
PentoRef English, German Speech (audio recordings with transcriptions), Visual (virtual and real-world scenes) Audio recordings, transcripts, referring expression annotations, visual scene representations (logical and perceptual features), dialogue act tags, disfluency annotations Task-oriented puzzle game (Pentomino): object reference, referring expression generation and resolution Human-Human and Human-Wizard-of-Oz More than 20,000 utterances (approx. 216,343 tokens across sub-corpora); 8 sub-corpora PentoRef is a multilingual (English and German) corpus of task-oriented spoken dialogues in a Pentomino puzzle-playing domain, collected across multiple systematically manipulated experimental settings varying interactivity, visual access, and verbal channel. The corpus is fully transcribed and annotated with referring expressions mapped to objects in corresponding visual scenes, providing a rich resource for research on spoken referring expression generation and resolution. Zarrieß et al. 2016
AIMU English Text Transcripts with actionable item intent/action annotations and argument annotations, plus CDSSM vector embeddings Meeting understanding; actionable item detection (calendar, reminders, communication, device settings, search) in multi-party meetings Multi-party human 22 meetings, 21,035 utterances, 318 utterances annotated with actionable items, 10 intent/action types AIMU is an extended annotation layer on top of 22 meetings from the ICSI meeting corpus, where participant utterances are labelled with actionable intents (10 types across calendar, reminders, communication, device, and search domains) and associated slot arguments, together with CDSSM vector representations, to support automated meeting assistant research. Chen et al. 2016
STAC English Text Text (online chat transcripts), discourse structure annotations (SDRT), dialogue act annotations, game event logs Multi-party negotiation / strategic conversation (trading in the board game Settlers of Catan) Multi-party human 1,081 dialogues, 9,160 turns, 10,678 EDUs, 10,513 relation instances, 1,284 CDUs STAC is a corpus of multi-party online chat dialogues collected from an online version of the board game Settlers of Catan, annotated for discourse structure in the style of SDRT and for dialogue acts (offers, counter-offers, acceptances, refusals, etc.). It is the first corpus to provide full discourse structures for multi-party dialogues, featuring interleaved threads, creative language, and interactions between linguistic and extra-linguistic (game event) contexts. Asher et al. 2016
METALOGUE Multi-Issue Bargaining Corpus English Speech Audio recordings, ASR transcripts, manual transcriptions, dialogue act annotations (ISO 24617-2 extended with negotiation moves), rhetorical and dependence relation annotations Multi-issue bargaining / negotiation (anti-smoking regulation scenario) Human-Human 50 dialogues, ~4,000 speaking turns, 8 hours total duration ~80 turns per dialogue The METALOGUE corpus consists of 50 spoken human-human multi-issue bargaining dialogues (totalling ~8 hours and ~4,000 turns) collected from 16 participants negotiating over anti-smoking regulations. Dialogues are annotated with dialogue acts following an extended ISO 24617-2 scheme that adds negotiation-specific moves (e.g., OfferValue, CounterOfferValue, Deal), plus functional dependence, feedback dependence, and rhetorical relations. Petukhova et al. 2016
Negochat Corpus English Text Text (natural language utterances annotated with formal semantic intent labels) Negotiation (job-candidate domain: bilateral multi-issue closed negotiation) Human-Wizard (WOZ: human turkers as employer, automated agent backed by wizard NLU as candidate) 105 dialogues, 1484 human utterances, 2140 agent utterances (3624 total utterances) The Negochat Corpus is the first publicly available annotated natural language human-agent negotiation dialogue corpus, collected via Amazon Mechanical Turk using a Wizard-of-Oz approach in a job-candidate negotiation domain. Each human utterance is annotated with formal semantic intent labels (Offer, Accept, Reject, Query, Greet, Quit) by two independent annotators, achieving a Krippendorff’s α inter-annotator agreement of 0.95. Konovalov et al. 2016
DUEL German, French, Mandarin Chinese Speech, Video, Body tracking (multimodal face-to-face) Audio, Video, Body tracking (Kinect 2 skeleton data), Transcripts with disfluency/laughter/exclamation annotations Loosely task-directed face-to-face dialogue (Dream Apartment, Film Script, Border Control role-play tasks) Human-Human (dyads, 10 pairs per language) 24 hours (30 dyads across 3 languages, ~45 minutes per dyad) DUEL (Disfluency, Exclamations and Laughter in Dialogue) is a 24-hour multilingual, multimodal corpus of natural face-to-face dyadic dialogue in German, French, and Mandarin Chinese, recorded with audio, video, and Kinect body-tracking data. The corpus is transcribed and annotated for disfluency, laughter, and exclamations using a unified, cross-linguistically consistent annotation scheme, making it a unique resource for cross-linguistic spontaneous dialogue research. Hough et al. 2016
Artwalk Corpus English Speech (mobile phone / Skype calls), Transcripts Transcripts, Audio (phone/Skype recordings), Photos of target artworks, GPS coordinates, Post-experiment questionnaire responses, Weather and session metadata Pedestrian navigation and referential communication (public art identification in a real-world outdoor setting) Human-Human (Director on campus via Skype + Follower on mobile phone downtown; 24 friend pairs and 24 stranger pairs) 48 dialogues (24 friend pairs, 24 stranger pairs) ~40 minutes per dialogue (range: 24–55 minutes) The Artwalk Corpus consists of 48 mobile phone conversations between pairs of friends and strangers performing a naturalistic referential communication task: a Director on a university campus gives verbal instructions via Skype to a Follower walking downtown Santa Cruz to identify and photograph public artworks. The corpus is designed to study entrainment, referring expression coordination, wayfinding dialogue, and the effect of friendship on dialogue in real-world, out-of-lab conditions. Liu et al. 2016
HCA Conversation Corpus English Speech, Gesture (deictic) Audio recordings, transcripts, deictic gesture annotations Interactive storytelling / Hans Christian Andersen narrative system Human-System 5 sub-corpora, ~57 hours of interaction The Hans Christian Andersen (HCA) Conversation Corpus consists of five sub-corpora comprising approximately 57 hours of transcribed and annotated English spoken and deictic gesture interactions, recorded primarily with children between 2002 and 2005 as part of the development and evaluation of two consecutive research prototypes for an interactive storytelling system. Andersen et al. 2006