Linguistic Data Consortium
The leading non-profit consortium that creates, shares and preserves high quality language resources
09/21/2026
MATERIAL Lithuanian-English Language Pack was developed by Appen for the IARPA MATERIAL program and contains 64 hours of Lithuanian conversational telephone speech, transcripts, English translations, annotations and queries. Calls were made using different telephones from a variety of environments. Transcripts cover 100% of the speech files, 6% of which were translated into English. This release also includes domain annotations, English queries and their relevance annotations. The MATERIAL program focused on underserved languages with the ultimate goal of building cross language information retrieval systems to find speech and text content using English search queries. https://catalog.ldc.upenn.edu/LDC2026S12
09/18/2026
CALLHOME Mandarin Chinese Lexicon Second Edition was developed by LDC and contains 44,404 Mandarin Chinese words with morphological, phonological and frequency information. This second edition updates file formats, directory structure and documentation. The first edition is available as CALLHOME Mandarin Chinese Lexicon (LDC96L15). The words in the lexicon were derived from transcripts representing unscripted telephone conversations between native Mandarin Chinese speakers contained in CALLHOME Mandarin Chinese Second Edition (LDC2026S11) and from Chinese news text. This release also includes a pronunciation dictionary derived from the lexicon in CMUdict format. https://catalog.ldc.upenn.edu/LDC2026L06
09/17/2026
CALLHOME Mandarin Chinese Second Edition was developed by LDC and contains 38 hours of speech from 120 unscripted telephone conversations between native Mandarin Chinese speakers. This publication is a re-release of the original CALLHOME Mandarin Chinese collection, combining CALLHOME Mandarin Chinese Speech (LDC96S34) and CALLHOME Mandarin Chinese Transcripts (LDC96T16), with additional transcription and updated directory structure, file formats, and documentation.
Participants spoke on topics of their choice in a single telephone call lasting up to 30 minutes. Calls were manually audited for gender, language, recording quality, channel characteristics, dialect, and accent. For this second edition, all audio was converted from SPHERE files to FLAC format, and the original training/development/evaluation partitioning was removed.
In addition to the original transcripts published in CALLHOME Mandarin Chinese Transcripts, this release has updated transcripts addressing normalization of annotation formats, standardization of speaker-produced and background noises, application of foreign-language marking, whitespace cleanup, and corrections and consistency fixes. https://catalog.ldc.upenn.edu/LDC2026S11
09/16/2026
Check out the September newsletter for info on LDC’s three new publications – CALLHOME Mandarin Chinese Speech Second Edition, CALLHOME Mandarin Chinese Lexicon Second Edition and MATERIAL Lithuanian-English Language Pack http://ldc-upenn.blogspot.com/
08/21/2026
The LORELEI series continues with LORELEI Arabic Representative Language Pack, a collection of comprehensive resources for HLT development -- monolingual text, Arabic-English parallel text, entity annotation, noun phrase chunking, semantic and situation frame annotation, and related software tools – developed by LDC for the DARPA LORELEI program. The LORELEI program was concerned with building human language technology for low resource languages in the context of emergent situations. Data was collected from discussion forum, news, reference, social network, and weblogs. https://catalog.ldc.upenn.edu/LDC2026T08
08/20/2026
Multi-Language Conversational Telephone Speech 2014 – Slavic Group is comprised of 20 hours (95 recordings) of Polish and Russian telephone speech collected by LDC to support research and technology evaluation in automatic language identification. Portions of these recordings were used in the NIST 2015 and 2017 language recognition evaluations. The collection focused on language pair discrimination for 20 languages/dialects, some of which could be considered mutually intelligible or closely related https://catalog.ldc.upenn.edu/LDC2026S10
08/19/2026
Parsed Early English Books Online - Text Creation Partnership (EEBO-TCP), developed by LDC, is a part-of-speech tagged and syntactically parsed version of the EEBO-TCP collection of Early Modern English texts. The corpus consists of 59,433 texts dating primarily from 1600-1700, with a smaller number of texts from earlier and later periods. It contains 48.6 million parsed sentences (trees) comprising more than 1.5 billion tokens. The parses were produced automatically using a parser trained on the Penn Helsinki Parsed Corpus of Early Modern English, part of the Penn Parsed Corpora of Historical English (LDC2020T16). The automatically generated parses were not manually reviewed.
This release includes the CorpusSearch 2 program and associated documentation. This tool allows users to search the data for syntactic structure, word sequences and words. An alternative version of CorpusSearch 2 for use on very large corpora is also included in this release. https://catalog.ldc.upenn.edu/LDC2026T09
08/18/2026
LDC’s August newsletter has the latest on the fall data scholarship application deadline and three new releases – Parsed EEBO-TCP, Multi-Language Conversational Telephone Speech 2014 – Slavic Group and LORELEI Arabic Representative Language Pack http://ldc-upenn.blogspot.com/
07/21/2026
CALLHOME American English Lexicon (PRONLEX) Second Edition was developed by LDC and contains 90,988 English words with citation-form pronunciations. This second edition updates file formats, directory structure and documentation. The first edition is available as CALLHOME American English Lexicon (PRONLEX) (LDC97L20). The words in the lexicon were derived from Wall Street Journal text used in the continuous speech recognition publication series CSR-1 WSJ0 Complete (LDC93S6A), transcripts from the Switchboard telephone collection (LDC97S62), and transcripts representing unscripted telephone conversations between native American English speakers contained in CALLHOME American English Second Edition (LDC2026S08). PRONLEX transcription is a phonemic transcription system that provides standard American English pronunciations. This release also includes a pronunciation dictionary derived from the lexicon in CMUdict format. https://catalog.ldc.upenn.edu/LDC2026L05
07/20/2026
CALLHOME American English Second Edition was developed by LDC and contains 56 hours of speech from 120 telephone conversations between native American English speakers. It is a re-release of the original CALLHOME American English collection, combining CALLHOME American English Speech LDC97S42 and CALLHOME American English Transcripts LDC97T14, with additional transcription and updated directory structure, file formats, and documentation. Participants spoke on topics of their choice in a single telephone call lasting up to 30 minutes. Calls were manually audited for gender, language, recording quality, channel characteristics, dialect, and region. For this second edition, all audio was converted from SPHERE files to FLAC format, and the original training/development/evaluation partitioning was removed.
In addition to the original transcripts published in CALLHOME American English Transcripts, this release has updated transcripts addressing normalization of annotation formats, standardization of speaker-produced and background noises, application of foreign-language marking, whitespace cleanup, and corrections and consistency fixes. https://catalog.ldc.upenn.edu/LDC2026S08
Click here to claim your Sponsored Listing.
Category
Contact the organization
Telephone
Website
Address
3600 Market Street, Ste 810
Philadelphia, PA
19104
Alerts
Be the first to know and let us send you an email when Linguistic Data Consortium posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.