mirror of
https://github.com/awesomedata/awesome-public-datasets.git
synced 2024-04-18 07:30:58 +08:00
Update README.rst
added Personae and CSI corpus to Natural Language
This commit is contained in:
parent
908654d7e8
commit
59a5dc490b
|
@ -385,6 +385,7 @@ Natural Language
|
||||||
----------------
|
----------------
|
||||||
|
|
||||||
* `Blogger Corpus <http://u.cs.biu.ac.il/~koppel/BlogCorpus.htm>`_
|
* `Blogger Corpus <http://u.cs.biu.ac.il/~koppel/BlogCorpus.htm>`_
|
||||||
|
* `CLiPS Stylometry Investigation Corpus <http://www.clips.uantwerpen.be/datasets/csi-corpus>`_
|
||||||
* `ClueWeb09 FACC <http://lemurproject.org/clueweb09/FACC1/>`_
|
* `ClueWeb09 FACC <http://lemurproject.org/clueweb09/FACC1/>`_
|
||||||
* `ClueWeb12 FACC <http://lemurproject.org/clueweb12/FACC1/>`_
|
* `ClueWeb12 FACC <http://lemurproject.org/clueweb12/FACC1/>`_
|
||||||
* `DBpedia - 4.58M things with 583M facts <http://wiki.dbpedia.org/Datasets>`_
|
* `DBpedia - 4.58M things with 583M facts <http://wiki.dbpedia.org/Datasets>`_
|
||||||
|
@ -396,6 +397,7 @@ Natural Language
|
||||||
* `Hansards text chunks of Canadian Parliament <http://www.isi.edu/natural-language/download/hansard/>`_
|
* `Hansards text chunks of Canadian Parliament <http://www.isi.edu/natural-language/download/hansard/>`_
|
||||||
* `Machine Comprehension Test (MCTest) of text from Microsoft Research <http://research.microsoft.com/en-us/um/redmond/projects/mctest/index.html>`_
|
* `Machine Comprehension Test (MCTest) of text from Microsoft Research <http://research.microsoft.com/en-us/um/redmond/projects/mctest/index.html>`_
|
||||||
* `Machine Translation of European languages <http://statmt.org/wmt11/translation-task.html#download>`_
|
* `Machine Translation of European languages <http://statmt.org/wmt11/translation-task.html#download>`_
|
||||||
|
* `Personae Corpus <http://www.clips.uantwerpen.be/datasets/personae-corpus>`_
|
||||||
* `SaudiNewsNet Collection of Saudi Newspaper Articles (Arabic, 30K articles) <https://github.com/ParallelMazen/SaudiNewsNet>`_
|
* `SaudiNewsNet Collection of Saudi Newspaper Articles (Arabic, 30K articles) <https://github.com/ParallelMazen/SaudiNewsNet>`_
|
||||||
* `SMS Spam Collection in English <http://www.dt.fee.unicamp.br/~tiago/smsspamcollection/>`_
|
* `SMS Spam Collection in English <http://www.dt.fee.unicamp.br/~tiago/smsspamcollection/>`_
|
||||||
* `USENET postings corpus of 2005~2011 <http://www.psych.ualberta.ca/~westburylab/downloads/usenetcorpus.download.html>`_
|
* `USENET postings corpus of 2005~2011 <http://www.psych.ualberta.ca/~westburylab/downloads/usenetcorpus.download.html>`_
|
||||||
|
|
Loading…
Reference in New Issue
Block a user