Linguistic Corpora at the HZSK Repository
EXMARaLDA Demo Corpus 1.0
A selection of short audio and video recordings in various languages to be used for instruction or demonstration of the EXMARaLDA system.
Language: German, English, French, Spanish, Turkish, Polish, Vietnamese, Swedish, Norwegian, Italian, Russian, Afrikaans, Portuguese
License: HZSK-PUB (public)
The Spoken Wikipedia Corpora
The Spoken Wikipedia project unites volunteer readers of Wikipedia articles. Hundreds of spoken articles in multiple languages are available to users who are – for one reason or another – unable or unwilling to consume the written version of the article. Our resource, the Spoken Wikipedia Corpus, consolidates the Spoken Wikipediae, adding text segmentation, normalization, time-alignment and further annotations, making it accessible for research and fostering new ways of interacting with the material.
Language: English, German, Dutch
License: Creative Commons Attribution-ShareAlike 4.0 International (public)
Covert translation: Business Communication (old)
Translation corpora of original texts with translations and comparable texts from the genre external business communication
Language: German, English
License: HZSK-ACA (academic)
Covert translation: Business Communication (new)
Translation corpora of original texts with translations and comparable texts from the genre external business communication.
Language: German, English
License: HZSK-ACA (academic)
Community Interpreting Database Pilot Corpus (ComInDat)
Audio and video recordings of various types of community interpreted discourse (doctor-patient communication, simulated doctor-patient communication, courtroom communication) in German (simulated and authentic doctor-patient communication) and US (courtroom communication) institutions with varying community languages. Video recordings only exist for the simulated communication. For the authentic interpreted doctor-patient communication, no audio files will be made available.
Language: German, English, Spanish, Turkish, Polish, Portuguese, Romanian, Russian, Haitian
License: HZSK-RES (restricted)
Hamburg Corpus of Old Swedish with Syntactic Annotations (HaCOSSA)
Religious and secular prose, law texts, non-fiction literature (geographical, theological, historic, natural science), diploma.
Language: English, German, Latin, Old Swedish, Swedish
License: FID-AKA (restricted)