Summary of the paper

Title Wikicorpus: A Word-Sense Disambiguated Multilingual Wikipedia Corpus
Authors Samuel Reese, Gemma Boleda, Montse Cuadros, Lluís Padró and German Rigau
Abstract This article presents a new freely available trilingual corpus (Catalan,Spanish, English) that contains large portions of the Wikipedia and has beenautomatically enriched with linguistic information. To our knowledge, this isthe largest such corpus that is freely available to the community: In itspresent version, it contains over 750 million words.The corpora have been annotated with lemma and part of speech information usingthe open source library FreeLing. Also,they have been sense annotated with the state of the art Word SenseDisambiguation algorithm UKB. As UKB assigns WordNet senses, andWordNet has been aligned across languages via the InterLingual Index, this sortof annotation opens the way tomassive explorations in lexical semantics that were not possible before.We present a first attempt at creating a trilingual lexical resource from thesense-tagged Wikipedia corpora, namely, WikiNet.Moreover, we present two by-products of the project that are of use for the NLPcommunity: An open source Java-based parser for Wikipedia pages developed forthe construction of the corpus, and the integration of the WSD algorithm UKB inFreeLing.
Language Word Sense Disambiguation
Topics Corpus (creation, annotation, etc.), Lexicon, lexical database, Word Sense Disambiguation
Full paper Wikicorpus: A Word-Sense Disambiguated Multilingual Wikipedia Corpus
Bibtex @InProceedings{REESE10.222,
  author = {Samuel Reese, Gemma Boleda, Montse Cuadros, Lluís Padró and German Rigau},
  title = {Wikicorpus: A Word-Sense Disambiguated Multilingual Wikipedia Corpus},
  booktitle = {Proceedings of the Seventh conference on International Language Resources and Evaluation (LREC'10)},
  year = {2010},
  month = {may},
  date = {19-21},
  address = {Valletta, Malta},
  editor = {Nicoletta Calzolari (Conference Chair), Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odjik, Stelios Piperidis, Mike Rosner, Daniel Tapias},
  publisher = {European Language Resources Association (ELRA)},
  isbn = {2-9517408-6-7},
  language = {english}
 }
Powered by ELDA © 2010 ELDA/ELRA