q2K BHP
Black History Portal
THE BHP WIRE —
HIDDEN TRUTHS
What's New!
THE JOURNEY THROUGH TIME

Explore Black History

Explore the people, places, events, achievements, struggles and stories that shaped our journey.

✊🏾

Civil Rights

Movements, leaders, victories and the continuing fight for equality.

⚙️

Black Inventors

Innovation, patents, science, technology and world-changing contributions.

🏆

Sports

Pioneers, champions, Negro Leagues, records, activism and excellence.

♟️

People

Meet the people whose lives, choices and achievements shaped the journey.

📍

Places

Black towns, communities, institutions and places where history happened.

📜

Events

Moments that changed communities, movements, institutions and the nation.

Enter a person, place, event, or topic.
MY'STORY

The MOVE Fire

This is a personal recollection on the Move fire on May 13, 1985 Philadelphia police fired thousands of rounds at the MOVE house, city officials approved dropping an explosive device on the roof, the resulting fire was allowed to burn, 11 people—including five children—died, and 61 homes were destroyed. Philadelphia City Council later called it a “brutal attack carried out by the City of Philadelphia on its own citizens” and acknowledged that no individual faced criminal consequences for the bombing. One timeline correction worth preserving for the BHP record: the major previous MOVE-police confrontation was August 8, 1978, about seven years before the bombing, not a year or two earlier. Officer James Ramp was killed, other police and firefighters were wounded, nine MOVE members were later convicted, and television cameras recorded police beating Delbert Africa during his arrest. The 1985 MOVE Commission later specifically criticized city planners for failing to adequately use lessons from that 1978 confrontation. And that actually strengthens the point you’re making: 1985 did not happen without precedent or institutional memory. There had already been a deadly confrontation with MOVE, years of conflict, negotiations and police involvement before Osage Avenue.

MORE →
BLACK FACTS
The Truths They Never Taught You...

Ruler of the Mali Empire in the 14th century

Mansa Musa was the ruler of the Mali Empire in West Africa. Details recorded here should be sourced; unknown information is left blank.

MORE →
BHP gathered finds from its connected research sources. Showing the 4 strongest Black History matches.
← BACK TO RESULTS
Wikipedia

Text corpus

In linguistics and natural language processing, a corpus (pl.: corpora) or text corpus is a dataset, consisting of natively digital and older, digitalized, language resources, either annotated or unannotated. Annotated, they have been used in corpus linguistics for statistical hypothesis testing, checking occurrences or validating linguistic rules within a specific language territory.

Overview

[edit]

A corpus may contain texts in a single language (monolingual corpus) or text data in multiple languages (multilingual corpus). In order to make the corpora more useful for doing linguistic research, they are often subjected to a process known as annotation. An example of annotating a corpus is part-of-speech tagging, or POS-tagging, in which information about each word's part of speech (verb, noun, adjective, etc.) is added to the corpus in the form of tags. Another example is indicating the lemma (base) form of each word. When the language of the corpus is not a working language of the researchers who use it, interlinear glossing is used to make the annotation bilingual.[citation needed]

Some corpora have further structured levels of analysis applied. In particular, smaller corpora may be fully parsed. Such corpora are usually called Treebanks or Parsed Corpora. The difficulty of ensuring that the entire corpus is completely and consistently annotated means that these corpora are usually smaller, containing around one to three million words. Other levels of linguistic structured analysis are possible, including annotations for morphology, semantics and pragmatics.[citation needed]

Applications

[edit]

Corpora are the main knowledge base in corpus linguistics.[citation needed] Other notable areas of application include:

  • Machine translation
    • Multilingual corpora that have been specially formatted for side-by-side comparison are called aligned parallel corpora. There are two main types of parallel corpora which contain texts in two languages. In a translation corpus, the texts in one language are translations of texts in the other language. In a comparable corpus, the texts are of the same kind and cover the same content, but they are not translations of each other.[2] To exploit a parallel text, some kind of text alignment identifying equivalent text segments (phrases or sentences) is a prerequisite for analysis. Machine translation algorithms for translating between two languages are often trained using parallel fragments comprising a first-language corpus and a second-language corpus, which is an element-for-element translation of the first-language corpus.[3]
  • Philologies
    • Text corpora are also used in the study of historical documents, for example in attempts to decipher ancient scripts, or in Biblical scholarship. Some archaeological corpora can be of such short duration that they provide a snapshot in time. One of the shortest corpora in time may be the 15–30 year Amarna letters texts (1350 BC). The corpus of an ancient city, (for example the "Kültepe Texts" of Turkey), may go through a series of corpora, determined by their find site dates.

Some notable text corpora

[edit]

See also

[edit]

References

[edit]
  1. ^ Yoon, H., & Hirvela, A. (2004). ESL Student Attitudes toward Corpus Use in L2 Writing. Journal of Second Language Writing, 13(4), 257–283. Retrieved 21 March 2012.
  2. ^ Wołk, K.; Marasek, K. (7 April 2014). "Real-Time Statistical Speech Translation". New Perspectives in Information Systems and Technologies, Volume 1. Advances in Intelligent Systems and Computing. Vol. 275. Springer. pp. 107–114. arXiv:1509.09090. doi:10.1007/978-3-319-05951-8_11. ISBN 978-3-319-05950-1. ISSN 2194-5357. S2CID 15361632.
  3. ^ Wolk, Krzysztof; Marasek, Krzysztof (2015). "Tuned and GPU-accelerated parallel data mining from comparable corpora". In Král, Pavel; Matoušek, Václav (eds.). Text, Speech, and Dialogue – 18th International Conference, TSD 2015, Plzeň, Czech Republic, September 14–17, 2015, Proceedings. Lecture Notes in Computer Science. Vol. 9302. Springer. pp. 32–40. arXiv:1509.08639. doi:10.1007/978-3-319-24033-6_4. ISBN 978-3-319-24032-9.
[edit]


Source: Wikipedia. Article content is retrieved live through the MediaWiki API.

No preview image
Wikipedia

Text corpus

In linguistics and natural language processing, a corpus (pl.: corpora) or text corpus is a dataset, consisting of natively digital and older, digitalized, language resources, either annotated or unannotated. Annotated, they have been used in corpus linguistics for statistical hypothesis testing, checking occurrences or validating linguistic rules within a specific language territory.

MORE →
Wikipedia

Electronic Text Corpus of Sumerian Literature

The Electronic Text Corpus of Sumerian Literature (ETCSL) is an online digital library of texts and translations of Sumerian literature that was created by a now-completed project based at the Oriental Institute of the University of Oxford. This project's website contains "Sumerian text, English prose translation and bibliographical information" for "over 400 literary works composed in the Sumerian language in ancient Mesopotamia (modern Iraq) during the late third and early second millennia BCE." It is both browsable and searchable and includes transliterations, composite texts, a bibliography of Sumerian literature and a guide to spelling conventions for proper nouns and literary forms. The purpose of the project was to make Sumerian literature accessible to those wishing to read or study it, and make it known to a wider public. The project was founded by Jeremy Black in 1997 and is based at the Oriental Institute of the University of Oxford. It was funded by the University along with the Leverhulme Trust and the Arts and Humanities Research Board. Various other bodies have been involved in the project including All Souls College, Oxford, the British Academy, the Hungarian Scientific Research Fund (OTKA) and the Hungarian Academy of Sciences. Contributors to the project have included Graham Cunningham, Eleanor Robson, Gábor Zólyomi, Miguel Civil, Bendt Alster, Joachim Krecher and Piotr Michałowski. Other libraries from the University of Chicago and the University of Pennsylvania now usually follow the ETCSL in regards to abbreviations. Funding for the project ended and it was closed in 2006, but the web site remains available.

MORE →
No preview image
Wikipedia

Neo-Assyrian Text Corpus Project

The Neo-Assyrian Text Corpus Project is an international scholarly project aimed at collecting and publishing ancient Assyrian texts of the Neo-Assyrian Empire and studies based on them. Its headquarters are in Helsinki in Finland.

MORE →
No preview image
Wikipedia

AsoSoft text corpus

The AsoSoft text corpus is the first large-scale Kurdish text corpus, collected and processed by the AsoSoft research and development group. It contains 458,000 documents (188 million tokens) that are collected from sources such as websites, news agencies, books, and magazines. The corpus is partially tagged by topic, so it can be used for topic identification tasks. Also, it is applicable for extracting language model and computational lexicon information. Part of the corpus (75 million tokens) is available online for non-commercial use. The corpus uses the TEI format.

MORE →
TOPIC OF THE DAY

Greenwood / Black Wall Street

Before the 1921 destruction of Tulsa’s Greenwood District, Black residents had created a remarkable center of business and community life. The district included stores, professional offices, entertainment venues and homes owned by Black citizens. Understanding Greenwood means learning what was built—not only what was burned.

MORE →
TRIVIA QUESTION OF THE DAY

What prosperous Tulsa district became widely known as “Black Wall Street”?

The Greenwood District.