Starting from a TCP Transcription
One possible starting point for a LEMDO semi-diplomatic transcription of a printed
playbook is the text created by the Text Creation Partnership (TCP), if one exists. An EEBO-TCP text offers a transcription of the copy microfilmed
for Early English Books (EEB) and subsequently digitized for Early English Books Online (EEBO). EEBO-TCP transcriptions are encoded in TEI. The EEBO-TCP customization of
TEI is similar to LEMDO’s customization, which means we can run a conversion on the
EEBO-TCP file to give you or an RA a convenient—if imperfect—starting point.
TCP is an imperfect starting point for various reasons:
Even so, having much of the transcription and tagging in place does save time.
TCP does not transcribe certain features that LEMDO does transcribe (e.g., forme works).
TCP’s transcriptions leave gaps and occasionally mistranscribe.
TCP transcriptions are made from digitized black-and-white microfilms and are thus
two removes from the physical book.
The copy filmed for EEB is not necessarily either the best copy of the printed playbook
or the copy that we will take as the basis for the LEMDO transcription. We try to
obtain a high-resolution colour scan of a copy so that we can embed images in the
transcription, and we align the transcription with the copy for which we have scans.
Normally, the LEMDO RAs remediate (i.e., correct) the converted TCP transcriptions
for editors.1 The editor’s job consists of two tasks:
Identify the correct TCP transcription to take as a starting point. This task needs
to happen before LEMDO runs the conversion.
Peer review LEMDO’s work and submit any corrections to LEMDO. This task happens after
the LEMDO RAs have completed their remediation.
How to find an EEBO-TCP GitHub File ID
Once you have identified the STC or Wing number of the publication for which you want
a semi-diplomatic transcription prepared,2 you will need to find the ID of the XML file of the corresponding TCP transcription.
The TCP file ID is an alphanumeric string that begins with an upper case letter. Example:
A68727.
The TCP’s XML files are stored on GitHub, along with a master list in the form of
a CSV file with the file extension .csv. CSV stands for “comma separated values.”
The direct link to this master list of TCP texts is as follows (link valid as of August
21, 2026): https://raw.githubusercontent.com/textcreationpartnership/Texts/master/TCP.csv. If the link does not work, visit https://github.com and type “EEBO texts” into the search bar. One of the top results will be textcreationpartnership/Texts. Navigate to this page and click TCP.csv from the list of files. Click “View raw” to open the file in your browser. Even though
the file is plain text with no formatting, it is very large and may take a few minutes
to fully load in your browser. Be patient and resist the urge to refresh the page.
Each TCP XML file is represented by a row in the master list. There are ten fields
in each row, each field wrapped in double quotation marks. The fields are separated
from each other by a comma. The field headings are:
"TCP","EEBO","VID","STC","Status","Author","Date","Title","Terms","Pages"The first ("TCP") and fourth ("STC") fields are the ones that matter, although the other fields are helpful for confirming that you have found the right publication.
When the TCP.csv file has fully loaded, use ctrl+f to search the file for the STC or Wing number. Be sure to include STC or Wing in your search string. (The ESTC number is also listed in the STC field, if you prefer
to search by ESTC number.) It will take time for your browser to perform a search
on the file. Example: Q1 Merchant of Venice is STC 22296. If you search the CSV file for STC 22296, you will find this entry:
"A68727","99846609","11589","STC 22296; ESTC S111215","Free","Shakespeare, William, 1564-1616.","1600","The most excellent historie of the merchant of Venice VVith the extreame crueltie of Shylocke the Iewe towards the sayd merchant, in cutting a iust pound of his flesh: and the obtayning of Portia by the choyse of three chests. As it hath beene diuers times acted by the Lord Chamberlaine his Seruants. Written by William Shakespeare.; Merchant of Venice","",70
The STC or Wing (and ESTC) number for which you searched is captured in the fourth
field. The first field contains the TCP ID that LEMDO needs. For Q1 Merchant, the ID is A68727. Send this number to the LEMDO team at lemdo@uvic.ca.
The TCP Editorial Declaration
Every TCP transcription will have information about the EEBO-TCP transcription and
encoding process contained in the file’s
<editorialDecl>
element. LEMDO’s RAs will move EEBO-TCP’s
<editorialDecl>
into
<xenoData>
so that we can use the
<editorialDecl>
to describe our remediation process. Once the file is fully remediated to comply
with LEMDO’s practices, we normally delete the entire
<xenoData>
element and use the
<sourceDesc>
element to capture the fact that we took the TCP transcription as our starting point.
For reference, we include a typical
<editorialDecl>
below.
<editorialDecl>
<p>EEBO-TCP is a partnership between the Universities of Michigan and Oxford and the publisher ProQuest to create accurately transcribed and encoded texts based on the image sets published by ProQuest via their Early English Books Online (EEBO) database (http://eebo.chadwyck.com). The general aim of EEBO-TCP is to encode one copy (usually the first edition) of every monographic English-language title published between 1473 and 1700 available in EEBO.</p>
<p>EEBO-TCP aimed to produce large quantities of textual data within the usual project restraints of time and funding, and therefore chose to create diplomatic transcriptions (as opposed to critical editions) with light-touch, mainly structural encoding based on the Text Encoding Initiative (http://www.tei-c.org).</p>
<p>The EEBO-TCP project was divided into two phases. The 25,368 texts created during Phase 1 of the project (2000-2009), initially available only to institutions that contributed to their creation, were released into the public domain on 1 January 2015. The approximately 40,000 texts produced during Phase 2 (2009- ), of which 34,963 had been released as of 2020, originally similarly restricted, were similarly freed from all restrictions on 1 August 2020. As of that date anyone is free to take and use these texts for any purpose (modify them, annotate them, distribute them, etc.). But we do respectfully request that due credit and attribution be given to their original source.</p>
<p>Users should be aware of the process of creating the TCP texts, and therefore of any assumptions that can be made about the data.</p>
<p>Text selection was based on the New Cambridge Bibliography of English Literature (NCBEL). If an author (or for an anonymous work, the title) appears in NCBEL, then their works are eligible for inclusion. Selection was intended to range over a wide variety of subject areas, to reflect the true nature of the print record of the period. In general, first editions of a works in English were prioritized, although there are a number of works in other languages, notably Latin and Welsh, included and sometimes a second or later edition of a work was chosen if there was a compelling reason to do so.</p>
<p>Image sets were sent to external keying companies for transcription and basic encoding. Quality assurance was then carried out by editorial teams in Oxford and Michigan. 5% (or 5 pages, whichever is the greater) of each text was proofread for accuracy and those which did not meet QA standards were returned to the keyers to be redone. After proofreading, the encoding was enhanced and/or corrected and characters marked as illegible were corrected where possible up to a limit of 100 instances per text. Any remaining illegibles were encoded as <gap>s. Understanding these processes should make clear that, while the overall quality of TCP data is very good, some errors will remain and some readable characters will be marked as illegible. Users should bear in mind that in all likelihood such instances will never have been looked at by a TCP editor.</p>
<p>The texts were encoded and linked to page images in accordance with level 4 of the TEI in Libraries guidelines.</p>
<p>Copies of the texts have been issued variously as SGML (TCP schema; ASCII text with mnemonic sdata character entities); displayable XML (TCP schema; characters represented either as UTF-8 Unicode or text strings within braces); or lossless XML (TEI P5, characters represented either as UTF-8 Unicode or TEI g elements).</p>
<p>Keying and markup guidelines are available at the <ref target="http://www.textcreationpartnership.org/docs/.">Text Creation Partnership web site</ref>.</p>
</editorialDecl>
<p>EEBO-TCP is a partnership between the Universities of Michigan and Oxford and the publisher ProQuest to create accurately transcribed and encoded texts based on the image sets published by ProQuest via their Early English Books Online (EEBO) database (http://eebo.chadwyck.com). The general aim of EEBO-TCP is to encode one copy (usually the first edition) of every monographic English-language title published between 1473 and 1700 available in EEBO.</p>
<p>EEBO-TCP aimed to produce large quantities of textual data within the usual project restraints of time and funding, and therefore chose to create diplomatic transcriptions (as opposed to critical editions) with light-touch, mainly structural encoding based on the Text Encoding Initiative (http://www.tei-c.org).</p>
<p>The EEBO-TCP project was divided into two phases. The 25,368 texts created during Phase 1 of the project (2000-2009), initially available only to institutions that contributed to their creation, were released into the public domain on 1 January 2015. The approximately 40,000 texts produced during Phase 2 (2009- ), of which 34,963 had been released as of 2020, originally similarly restricted, were similarly freed from all restrictions on 1 August 2020. As of that date anyone is free to take and use these texts for any purpose (modify them, annotate them, distribute them, etc.). But we do respectfully request that due credit and attribution be given to their original source.</p>
<p>Users should be aware of the process of creating the TCP texts, and therefore of any assumptions that can be made about the data.</p>
<p>Text selection was based on the New Cambridge Bibliography of English Literature (NCBEL). If an author (or for an anonymous work, the title) appears in NCBEL, then their works are eligible for inclusion. Selection was intended to range over a wide variety of subject areas, to reflect the true nature of the print record of the period. In general, first editions of a works in English were prioritized, although there are a number of works in other languages, notably Latin and Welsh, included and sometimes a second or later edition of a work was chosen if there was a compelling reason to do so.</p>
<p>Image sets were sent to external keying companies for transcription and basic encoding. Quality assurance was then carried out by editorial teams in Oxford and Michigan. 5% (or 5 pages, whichever is the greater) of each text was proofread for accuracy and those which did not meet QA standards were returned to the keyers to be redone. After proofreading, the encoding was enhanced and/or corrected and characters marked as illegible were corrected where possible up to a limit of 100 instances per text. Any remaining illegibles were encoded as <gap>s. Understanding these processes should make clear that, while the overall quality of TCP data is very good, some errors will remain and some readable characters will be marked as illegible. Users should bear in mind that in all likelihood such instances will never have been looked at by a TCP editor.</p>
<p>The texts were encoded and linked to page images in accordance with level 4 of the TEI in Libraries guidelines.</p>
<p>Copies of the texts have been issued variously as SGML (TCP schema; ASCII text with mnemonic sdata character entities); displayable XML (TCP schema; characters represented either as UTF-8 Unicode or text strings within braces); or lossless XML (TEI P5, characters represented either as UTF-8 Unicode or TEI g elements).</p>
<p>Keying and markup guidelines are available at the <ref target="http://www.textcreationpartnership.org/docs/.">Text Creation Partnership web site</ref>.</p>
</editorialDecl>
Further Reading about EEBO and EEBO-TCP
A Text Creation Partnership Companion.https://www.textpartnership.net/index.html.
Gadd, Ian.
The Use and Misuse of Early English Books Online.Literature Compass 6 (2009): 680–692. DOI 10.1111/j.1741-4113.2009.00632.x.
Gavin, Michael.
How To Think About EEBO.Textual Cultures: Text, Contexts, Interpretation 11.1–2 (2019): 70–105. DOI 10.14434/textual.v11i1-2.23570.
Gavin, Michael.
EEBO and Us.Textual Cultures: Text, Contexts, Interpretation 14.1 (2021): 270–278. DOI 10.14434/tc.v14i1.32860.
Herman, Peter C.
EEBO and Me: An Autobiographical Response to Michael Gavin,Textual Cultures: Text, Contexts, Interpretation 13.1 (2020): 207–216. DOI 10.14434/textual.v13i1.30078.How to Think About EEBO.
Kichuk, Diana.
Metamorphosis: Remediation in Early English Books Online (EEBO).Literary and Linguistic Computing 22.3 (2007): 291–303. DOI 10.1093/llc/fqm018.
Mäkelä, Eetu, James Misson, Devani Singh, and Mikko Tolonen.
Opening the black box of EEBO.Digital Scholarship in the Humanities 41.1 (2026): 236–254. DOI 10.1093/llc/fqaf086.
Maurer, Margaret C.
Facsimiles and Transcription: EEBO-TCP and Narratives of Textual Production.Journal of Early Modern Studies 14 (2025): 17–31. DOI 10.36253/jems-2279-7149-16516.
Quiring, Ana.
Fingerprints of British Book History: A Feminist Labor History of EEBO.Digital Humanities Quarterly 18.1 (2024). DOI 10.63744/4c9n5tkrpqeq.
Notes
1.LEMDO can commit to doing this work as long as Janelle has grant funding, donated
funds, or a revenue trickle to hire RAs. If you would like to hire an RA at your institution
to do the remediation (thereby making LEMDO funds go further), LEMDO will gladly train
your RA.↑
2.Use Greg’s Bibliography of the English Printed Drama to the Restoration in conjunction with Farmer and Lesser’s Database of Early English Playbooks to ensure that you pick the right publication (especially in those few cases where
there are multiple publications in the same year) and that you have the correct STC
number.↑
Prosopography
Illya
Illya has a BA in English and Sociocultural Anthropology and an MA in English. Prior
to joining the HCMC, he was a PhD candidate in English and Book History at the University
of Toronto and worked on Records of Early English Drama and on the Modernist Archives Publishing Project. His work at the HCMC focuses on creating web-based applications for research projects
led by members of the faculty of Humanities at the University of Victoria. This involves
creating schemas for new and existing datasets, writing XSLT and build files to transform
datasets into structured TEI and HTML formats, implementing staticSearch, and ensuring
that new projects are Endings Principles compliant.
Janelle Jenstad
Janelle Jenstad is a Professor of English at the University of Victoria, Director
of The Map of Early Modern London, and Director of Linked Early Modern Drama Online. With Jennifer Roberts-Smith and Mark Beatrice Kaethler, she co-edited Shakespeare’s Language in Digital Media: Old Words, New Tools (Routledge). She has edited John Stow’s A Survey of London (1598 text) for MoEML and is currently editing The Merchant of Venice (with Stephen Wittek) and Heywood’s 2 If You Know Not Me You Know Nobody for DRE. Her articles have appeared in Digital Humanities Quarterly, Elizabethan Theatre, Early Modern Literary Studies, Shakespeare Bulletin, Renaissance and Reformation, and The Journal of Medieval and Early Modern Studies. She contributed chapters to Approaches to Teaching Othello (MLA); Teaching Early Modern Literature from the Archives (MLA); Institutional Culture in Early Modern England (Brill); Shakespeare, Language, and the Stage (Arden); Performing Maternity in Early Modern England (Ashgate); New Directions in the Geohumanities (Routledge); Early Modern Studies and the Digital Turn (Iter); Placing Names: Enriching and Integrating Gazetteers (Indiana); Making Things and Drawing Boundaries (Minnesota); Rethinking Shakespeare Source Study: Audiences, Authors, and Digital Technologies (Routledge); and Civic Performance: Pageantry and Entertainments in Early Modern London (Routledge). For more details, see janellejenstad.com.
Joey Takeda
Joey Takeda is LEMDO’s Consulting Programmer and Designer, a role he assumed in 2020
after three years as the Lead Developer on LEMDO.
Mahayla Galliford
Project Manager, 2025-present; Assistant Project Manager, 2024-2025; Research Assistant,
2021-present. Mahayla Galliford (she/her) graduated from the University of Victoria
with a BA (honours with distinction) in 2024, and an MA English in 2026. Mahayla’s
undergraduate research explored early modern stage directions and civic water pageantry.
Her SSHRC-funded MA thesis project focuses on transcribing, editing, and encoding
early modern girls’ manuscripts, specifically Lady Rachel Fane’s May Masque in collaboration with LEMDO.
Martin Holmes
Martin Holmes has worked as a developer in the UVic’s Humanities Computing and Media
Centre for over two decades, and has been involved with dozens of Digital Humanities
projects. He has served on the TEI Technical Council and as Managing Editor of the
Journal of the TEI. He took over from Joey Takeda as lead developer on LEMDO in 2020.
He is a collaborator on the SSHRC Partnership Grant led by Janelle Jenstad.
Navarra Houldin
Training and Documentation Lead 2025–present. LEMDO project manager 2022–2025. Textual
remediator 2021–present. Navarra Houldin (they/them) completed their BA with a major
in history and minor in Spanish at the University of Victoria in 2022. Their primary
research was on gender and sexuality in early modern Europe and Latin America. They
are continuing their education through an MA program in Gender and Social Justice
Studies at the University of Alberta where they will specialize in Digital Humanities.
Samuel Seaberg
Samuel Seaberg, a University of Victoria English undergrad, enjoys riding his bike.
During the summer of 2025, he began working with LEMDO as a recipient of the Valerie
Kuehne Undergraduate Research Award (VKURA). Unfortunately, due to his summer being
spent primarily in working to establish an edition of Thomas Heywood’s If You Know Not Me, You Know Nobody, Part 2 and consequently working out how to represent multi-text works in a digital space,
his bike has suffered severely of sheltered seclusion from the sun. Note: Samuel now
works for LEMDO as the Assistant Project Manager, much to his bike’s chagrin.
Tracey El Hajj
Junior Programmer 2019–2020. Research Associate 2020–2021. Tracey received her PhD
from the Department of English at the University of Victoria in the field of Science
and Technology Studies. Her research focuses on the algorhythmics of networked communications. She was a 2019–2020 President’s Fellow in Research-Enriched
Teaching at UVic, where she taught an advanced course on
Artificial Intelligence and Everyday Life.Tracey was also a member of the Map of Early Modern London team, between 2018 and 2021. Between 2020 and 2021, she was a fellow in residence at the Praxis Studio for Comparative Media Studies, where she investigated the relationships between artificial intelligence, creativity, health, and justice. As of July 2021, Tracey has moved into the alt-ac world for a term position, while also teaching in the English Department at the University of Victoria.
Bibliography
DEEP: Database of Early
English Playbooks. Ed. Alan B.
Farmer and Zachary Lesser.
2007. https://deepplaybooks.org/.
Gadd, Ian.
The Use and Misuse of Early English Books Online.Literature Compass 6 (2009): 680–692. DOI 10.1111/j.1741-4113.2009.00632.x.
Greg, W.W.
Dramatic Documents from the Elizabethan Playhouses: Stage Plots, Actors’ Parts, Prompt
Books. 2 vols. Clarendon Press, 1931.
Kichuk, Diana.
Metamorphosis: Remediation in Early English Books Online (EEBO).Literary and Linguistic Computing 22.3 (2007): 291–303. DOI 10.1093/llc/fqm018.
Orgography
LEMDO Team (LEMD1)
The LEMDO Team is based at the University of Victoria and normally comprises the project
director, the lead developer, project manager, junior developers(s), remediators,
encoders, and remediating editors.
Metadata
| Authority title | Starting from a TCP Transcription |
| Type of text | Documentation |
| Publisher | University of Victoria on the Linked Early Modern Drama Online Platform |
| Series | Linked Early Modern Drama Online |
| Source |
TEI Customization created by Martin Holmes, Joey Takeda, and Janelle Jenstad; documentation written by members of the LEMDO Team
|
| Editorial declaration | n/a |
| Edition | Released with Linked Early Modern Drama Online 1.0 |
| Encoding description | Encoded in TEI P5 according to the LEMDO Customization and Encoding Guidelines |
| Document status | prgGenerated |
| Funder(s) | Social Sciences and Humanities Research Council of Canada |
| License/availability |
This file is licensed under a CC BY-NC_ND 4.0 license, which means that it is freely downloadable without permission under the following
conditions: (1) credit must be given to the author and LEMDO in any subsequent use
of the files and/or data; (2) the content cannot be adapted or repurposed (except
in quotations for the purposes of academic review and citation); and (3) commercial
uses are not permitted without the knowledge and consent of the editor and LEMDO.
This license allows for pedagogical use of the documentation in the classroom.
|