MUStARD: Multimodal Sarcasm Detection Dataset
Description
Description
The MUStARD dataset is a multimodal video corpus for research in automated sarcasm discovery. The dataset is compiled from popular TV shows including Friends, The Golden Girls, The Big Bang Theory, and Sarcasmaholics Anonymous. MUStARD consists of audiovisual utterances annotated with sarcasm labels. Each utterance is accompanied by its context, which provides additional information on the scenario where the utterance occurs.
Please cite the following if you use the data:
Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. 2019. Towards Multimodal Sarcasm Detection (An Obviously Perfect Paper). In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4619–4629, Florence, Italy. Association for Computational Linguistics.
Variables
| Name | Description |
|---|---|
| utterance | The text of the target utterance to classify. |
| speaker | Speaker of the target utterance. |
| context | List of utterances (in chronological order) preceding the target utterance. |
| context_speakers | Respective speakers of the context utterances. |
| sarcasm | Binary label for sarcasm tag. |
Details
| Resource type | Open dataset |
| Title | MUStARD: Multimodal Sarcasm Detection Dataset |
| Creators |
|
| Publisher | Association for Computational Linguistics |
| Year of publication | 2019 |
| Research fields | Psychology Business Administration Other |
| Formats | JSON format (.json) |
| External resource | https://github.com/soujanyaporia/MUStARD#mustard-multimodal-sarcasm-detection-dataset |
Additional details
Related works
- Is referenced by
- Journal article: 10.18653/v1/2020.acl-main.401 (DOI)
- Journal article: 10.3390/electronics12030666 (DOI)