RST Signalling Corpus: A Corpus of Signals of Coherence Relations

Peer reviewed: 
Yes, item is peer reviewed.
Scholarly level: 
Final version published as: 

Das, D., and Taboada, M. (2018). RST Signalling Corpus: A corpus of signals of coherence relations. Language Resources and Evaluation 52: 149. 10.1007/s10579-017-9383-x

Date created: 
RST Signalling Corpus
RST Discourse Treebank
Coherence relations
Rhetorical Structure Theory
Discourse markers

We present the RST Signalling Corpus (Das et al. in RST signalling corpus, LDC2015T10., a corpus annotated for signals of coherence relations. The corpus is developed over the RST Discourse Treebank (Carlson et al. in RST Discourse Treebank, LDC2002T07. which is annotated for coherence relations. In the RST Signalling Corpus, these relations are further annotated with signalling information. The corpus includes annotation not only for discourse markers which are considered to be the most typical (or sometimes the only type of) signals in discourse, but also for a wide array of other signals such as reference, lexical, semantic, syntactic, graphical and genre features as potential indicators of coherence relations. We describe the research underlying the development of the corpus and the annotation process, and provide details of the corpus. We also present the results of an inter-annotator agreement study, illustrating the validity and reproducibility of the annotation. The corpus is available through the Linguistic Data Consortium, and can be used to investigate the psycholinguistic mechanisms behind the interpretation of relations through signalling, and also to develop discourse-specific computational systems such as discourse parsing applications.

Document type: 
Rights remain with the authors.
Natural Sciences and Engineering Research Council of Canada (NSERC)