The SpeakUnique Corpus of Heard and Imagined Natural Speech (CHINS)
Overview
The Corpus of Heard and Imagined Natural Speech (CHINS) is a single-participant electroencephalography (EEG) corpus comprising over 22 hours of time-aligned natural speech data, with approximately 11 hours of heard speech and 11 hours of imagined speech. CHINS is intended to support research on the neural representation of natural speech, with particular emphasis on the relationship between auditory speech perception and imagined speech (inner speech).
The dataset is released under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) licence. It is made available for non-commercial purposes, including academic research and teaching, with appropriate attribution. Users may process the data for their own research; the licensing conditions governing redistribution and commercial use are described in the accompanying LICENSE file.
Publication
S. Wellington, O. Watts, D. Coyle, and B. Metcalfe, “Shared phone-level neural representations of auditory perception and ‘inner voice’ production: One-to-one mapping using a single-subject EEG corpus of heard and imagined natural speech,” in Proc. Interspeech 2026, 2026, pp. 4864–4869, doi: 10.21437/Interspeech.2026-2683.
Hosting and Access
The data is hosted at https://zenodo.org/records/22974624. Access is available on request: please contact [email protected] with details of your affiliation and description of how you will use the data to request access and download permissions.
