Lokal zugängliche Daten

English reference corpora

Das DH-Lab stellt die hier aufgeführten English reference corpora nach Prof. Christian Mairs Eintritt in den Ruhestand zur Verfügung.
Mitglieder der Universität Freiburg können Zugang zu diesen Daten erhalten und müssen dafür versichern, dass sie die Korpora nicht weitergeben.

Bei Interesse schicken Sie bitte eine Nachricht von Ihrer Universitätsemailadresse aus an dh-lab[at]mail.uni-freiburg.de mit den folgenden Informationen:

  • welches Korpus/welche Korpora benötigt werden
  • wofür sie verwendet werden – z.B. eine Kurzbeschreibung des Projekts oder Seminars
  • eine Bestätigung, dass verstanden wurde, dass es nicht erlaubt ist, die Daten an andere weiterzugeben

Available corpora

Brown family corpora

American EnglishBritish EnglishTime Period
B-BrownB-LOB1930s
BrownLOB (Lancaster-Oslo/Bergen) Corpus1960s
Frown (Freiburg-Brown corpus)FLOB1990s
CrownCLOB2009
  • family of corpora that can be used to study differences between the decades or to compare British and American English
  • each contain one million words of written English divided into 2,000 word samples from various genres (e.g. press reportage, different kinds of fiction, government documents)

Wellington Corpus of Written New Zealand English

  • one million words of written New Zealand English
  • 2,000 word excerpts of various texts
  • mirrors the structure of the Brown Corpus
  • writings published between 1986 and 1990
  • click here for the manual

Helsinki Corpus of English Texts

  • about 1.6 million words
  • divided into three main periods: Old, Middle and Early Modern English, each being subdivided into 100-year subperiods
  • covers a range of genres, regional varieties and sociolinguistic variables (e.g. gender, age, education, social class)
  • additional ‘satellite’ corpora of Early Scots and Early American English
  • click here for the manual

ICE (International Corpus of English)

  • a range of million-word corpora of different varieties of English, representing native and official-language national varieties of English
  • aims at achieving comparability between the different corpora; each one is constructed from 500 (300 spoken and 200 written language) 2,000 word samples to produce a 1,000,000 word corpus for each variety of English covered
  • available at the lab: Australia, Canada, East Africa, Great Britain, Hong Kong, India, Ireland, Jamaica, New Zealand, Philippines, Singapore, Sri Lanka, USA
  • some ICE corpora may also be downloaded here

LLC (London-Lund Corpus)

  • contains 100 texts of 500,000 words of spoken British English
  • various genres (e.g. spontaneous dialogues, radio broadcasts)
  • compiled 1975-1981 and 1985-1988
  • click here for the manual