Published 2019 | Version v4
Open dataset

MUStARD: Multimodal Sarcasm Detection Dataset

Description

Description

The MUStARD dataset is a multimodal video corpus for research in automated sarcasm discovery. The dataset is compiled from popular TV shows including Friends, The Golden Girls, The Big Bang Theory, and Sarcasmaholics Anonymous. MUStARD consists of audiovisual utterances annotated with sarcasm labels. Each utterance is accompanied by its context, which provides additional information on the scenario where the utterance occurs.

Please cite the following if you use the data:

Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. 2019. Towards Multimodal Sarcasm Detection (An Obviously Perfect Paper). In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4619–4629, Florence, Italy. Association for Computational Linguistics.

Variables

Name Description
utterance The text of the target utterance to classify.
speaker Speaker of the target utterance.
context List of utterances (in chronological order) preceding the target utterance.
context_speakers Respective speakers of the context utterances.
sarcasm Binary label for sarcasm tag.

Details

Resource type Open dataset
Title MUStARD: Multimodal Sarcasm Detection Dataset
Creators
  • Castro, Santiago ORCID icon
  • Hazarika, Devamanyu ORCID icon
  • Pérez-Rosas, Verónica ORCID icon
  • Zimmermann, Roger ORCID icon
  • Mihalcea, Rada ORCID icon
  • Poria, Soujanya ORCID icon
  • Publisher Association for Computational Linguistics
    Year of publication 2019
    Research fields Psychology Business Administration Other
    Formats JSON format (.json)
    External resource https://github.com/soujanyaporia/MUStARD#mustard-multimodal-sarcasm-detection-dataset

    Additional details

    Related works

    Is referenced by
    Journal article: 10.18653/v1/2020.acl-main.401 (DOI)
    Journal article: 10.3390/electronics12030666 (DOI)