Civil Rights
Movements, leaders, victories and the continuing fight for equality.
Explore the people, places, events, achievements, struggles and stories that shaped our journey.
Movements, leaders, victories and the continuing fight for equality.
Innovation, patents, science, technology and world-changing contributions.
Pioneers, champions, Negro Leagues, records, activism and excellence.
Meet the people whose lives, choices and achievements shaped the journey.
Black towns, communities, institutions and places where history happened.
Moments that changed communities, movements, institutions and the nation.
Mansa Musa was the ruler of the Mali Empire in West Africa. Details recorded here should be sourced; unknown information is left blank.
MORE →Reflects the personal views, recollections, and perspective of the author, Mike Davis.
This is a personal recollection on the Move fire on May 13, 1985
In linguistics and natural language processing, a corpus (pl.: corpora) or text corpus is a dataset, consisting of natively digital and older, digitalized, language resources, either annotated or unannotated. Annotated, they have been used in corpus linguistics for statistical hypothesis testing, checking occurrences or validating linguistic rules within a specific language territory.
A corpus may contain texts in a single language (monolingual corpus) or text data in multiple languages (multilingual corpus). In order to make the corpora more useful for doing linguistic research, they are often subjected to a process known as annotation. An example of annotating a corpus is part-of-speech tagging, or POS-tagging, in which information about each word's part of speech (verb, noun, adjective, etc.) is added to the corpus in the form of tags. Another example is indicating the lemma (base) form of each word. When the language of the corpus is not a working language of the researchers who use it, interlinear glossing is used to make the annotation bilingual.[citation needed]
Some corpora have further structured levels of analysis applied. In particular, smaller corpora may be fully parsed. Such corpora are usually called Treebanks or Parsed Corpora. The difficulty of ensuring that the entire corpus is completely and consistently annotated means that these corpora are usually smaller, containing around one to three million words. Other levels of linguistic structured analysis are possible, including annotations for morphology, semantics and pragmatics.[citation needed]
Corpora are the main knowledge base in corpus linguistics.[citation needed] Other notable areas of application include:
Source: Wikipedia. Article content is retrieved live through the MediaWiki API.
In linguistics and natural language processing, a corpus (pl.: corpora) or text corpus is a dataset, consisting of natively digital and older, digitalized, language resources, either annotated or unannotated. Annotated, they have been used in corpus linguistics for statistical hypothesis testing, checking occurrences or validating linguistic rules within a specific language territory.
The Electronic Text Corpus of Sumerian Literature (ETCSL) is an online digital library of texts and translations of Sumerian literature that was created by a now-completed project based at the Oriental Institute of the University of Oxford. This project's website contains "Sumerian text, English prose translation and bibliographical information" for "over 400 literary works composed in the Sumerian language in ancient Mesopotamia (modern Iraq) during the late third and early second millennia BCE." It is both browsable and searchable and includes transliterations, composite texts, a bibliography of Sumerian literature and a guide to spelling conventions for proper nouns and literary forms. The purpose of the project was to make Sumerian literature accessible to those wishing to read or study it, and make it known to a wider public. The project was founded by Jeremy Black in 1997 and is based at the Oriental Institute of the University of Oxford. It was funded by the University along with the Leverhulme Trust and the Arts and Humanities Research Board. Various other bodies have been involved in the project including All Souls College, Oxford, the British Academy, the Hungarian Scientific Research Fund (OTKA) and the Hungarian Academy of Sciences. Contributors to the project have included Graham Cunningham, Eleanor Robson, Gábor Zólyomi, Miguel Civil, Bendt Alster, Joachim Krecher and Piotr Michałowski. Other libraries from the University of Chicago and the University of Pennsylvania now usually follow the ETCSL in regards to abbreviations. Funding for the project ended and it was closed in 2006, but the web site remains available.
The Neo-Assyrian Text Corpus Project is an international scholarly project aimed at collecting and publishing ancient Assyrian texts of the Neo-Assyrian Empire and studies based on them. Its headquarters are in Helsinki in Finland.
The AsoSoft text corpus is the first large-scale Kurdish text corpus, collected and processed by the AsoSoft research and development group. It contains 458,000 documents (188 million tokens) that are collected from sources such as websites, news agencies, books, and magazines. The corpus is partially tagged by topic, so it can be used for topic identification tasks. Also, it is applicable for extracting language model and computational lexicon information. Part of the corpus (75 million tokens) is available online for non-commercial use. The corpus uses the TEI format.
Before the 1921 destruction of Tulsa’s Greenwood District, Black residents had created a remarkable center of business and community life. The district included stores, professional offices, entertainment venues and homes owned by Black citizens. Understanding Greenwood means learning what was built—not only what was burned.
MORE →The Greenwood District.