Starting from a TCP Transcription

One possible starting point for a LEMDO semi-diplomatic transcription of a printed playbook is the text created by the Text Creation Partnership (TCP), if one exists. An EEBO-TCP text offers a transcription of the copy microfilmed for Early English Books (EEB) and subsequently digitized for Early English Books Online (EEBO). EEBO-TCP transcriptions are encoded in TEI. The EEBO-TCP customization of TEI is similar to LEMDO’s customization, which means we can run a conversion on the EEBO-TCP file to give you or an RA a convenient—if imperfect—starting point.
TCP is an imperfect starting point for various reasons:
TCP does not transcribe certain features that LEMDO does transcribe (e.g., forme works).
TCP’s transcriptions leave gaps and occasionally mistranscribe.
TCP transcriptions are made from digitized black-and-white microfilms and are thus two removes from the physical book.
The copy filmed for EEB is not necessarily either the best copy of the printed playbook or the copy that we will take as the basis for the LEMDO transcription. We try to obtain a high-resolution colour scan of a copy so that we can embed images in the transcription, and we align the transcription with the copy for which we have scans.
Even so, having much of the transcription and tagging in place does save time.
Normally, the LEMDO RAs remediate (i.e., correct) the converted TCP transcriptions for editors.1 The editor’s job consists of two tasks:
Identify the correct TCP transcription to take as a starting point. This task needs to happen before LEMDO runs the conversion.
Peer review LEMDO’s work and submit any corrections to LEMDO. This task happens after the LEMDO RAs have completed their remediation.

How to find an EEBO-TCP GitHub File ID

Once you have identified the STC or Wing number of the publication for which you want a semi-diplomatic transcription prepared,2 you will need to find the ID of the XML file of the corresponding TCP transcription. The TCP file ID is an alphanumeric string that begins with an upper case letter. Example: A68727.
The TCP’s XML files are stored on GitHub, along with a master list in the form of a CSV file with the file extension .csv. CSV stands for “comma separated values.”
The direct link to this master list of TCP texts is as follows (link valid as of August 21, 2026): https://raw.githubusercontent.com/textcreationpartnership/Texts/master/TCP.csv. If the link does not work, visit https://github.com and type “EEBO texts” into the search bar. One of the top results will be textcreationpartnership/Texts. Navigate to this page and click TCP.csv from the list of files. Click “View raw” to open the file in your browser. Even though the file is plain text with no formatting, it is very large and may take a few minutes to fully load in your browser. Be patient and resist the urge to refresh the page.
Each TCP XML file is represented by a row in the master list. There are ten fields in each row, each field wrapped in double quotation marks. The fields are separated from each other by a comma. The field headings are:
"TCP","EEBO","VID","STC","Status","Author","Date","Title","Terms","Pages"
The first ("TCP") and fourth ("STC") fields are the ones that matter, although the other fields are helpful for confirming that you have found the right publication.
When the TCP.csv file has fully loaded, use ctrl+f to search the file for the STC or Wing number. Be sure to include STC or Wing in your search string. (The ESTC number is also listed in the STC field, if you prefer to search by ESTC number.) It will take time for your browser to perform a search on the file. Example: Q1 Merchant of Venice is STC 22296. If you search the CSV file for STC 22296, you will find this entry:
"A68727","99846609","11589","STC 22296; ESTC S111215","Free","Shakespeare, William, 1564-1616.","1600","The most excellent historie of the merchant of Venice VVith the extreame crueltie of Shylocke the Iewe towards the sayd merchant, in cutting a iust pound of his flesh: and the obtayning of Portia by the choyse of three chests. As it hath beene diuers times acted by the Lord Chamberlaine his Seruants. Written by William Shakespeare.; Merchant of Venice","",70
The STC or Wing (and ESTC) number for which you searched is captured in the fourth field. The first field contains the TCP ID that LEMDO needs. For Q1 Merchant, the ID is A68727. Send this number to the LEMDO team at lemdo@uvic.ca.

The TCP Editorial Declaration

Every TCP transcription will have information about the EEBO-TCP transcription and encoding process contained in the file’s <editorialDecl> element. LEMDO’s RAs will move EEBO-TCP’s <editorialDecl> into <xenoData> so that we can use the <editorialDecl> to describe our remediation process. Once the file is fully remediated to comply with LEMDO’s practices, we normally delete the entire <xenoData> element and use the <sourceDesc> element to capture the fact that we took the TCP transcription as our starting point. For reference, we include a typical <editorialDecl> below.
<editorialDecl>
  <p>EEBO-TCP is a partnership between the Universities of Michigan and Oxford and the publisher ProQuest to create accurately transcribed and encoded texts based on the image sets published by ProQuest via their Early English Books Online (EEBO) database (http://eebo.chadwyck.com). The general aim of EEBO-TCP is to encode one copy (usually the first edition) of every monographic English-language title published between 1473 and 1700 available in EEBO.</p>
  <p>EEBO-TCP aimed to produce large quantities of textual data within the usual project restraints of time and funding, and therefore chose to create diplomatic transcriptions (as opposed to critical editions) with light-touch, mainly structural encoding based on the Text Encoding Initiative (http://www.tei-c.org).</p>
  <p>The EEBO-TCP project was divided into two phases. The 25,368 texts created during Phase 1 of the project (2000-2009), initially available only to institutions that contributed to their creation, were released into the public domain on 1 January 2015. The approximately 40,000 texts produced during Phase 2 (2009- ), of which 34,963 had been released as of 2020, originally similarly restricted, were similarly freed from all restrictions on 1 August 2020. As of that date anyone is free to take and use these texts for any purpose (modify them, annotate them, distribute them, etc.). But we do respectfully request that due credit and attribution be given to their original source.</p>
  <p>Users should be aware of the process of creating the TCP texts, and therefore of any assumptions that can be made about the data.</p>
  <p>Text selection was based on the New Cambridge Bibliography of English Literature (NCBEL). If an author (or for an anonymous work, the title) appears in NCBEL, then their works are eligible for inclusion. Selection was intended to range over a wide variety of subject areas, to reflect the true nature of the print record of the period. In general, first editions of a works in English were prioritized, although there are a number of works in other languages, notably Latin and Welsh, included and sometimes a second or later edition of a work was chosen if there was a compelling reason to do so.</p>
  <p>Image sets were sent to external keying companies for transcription and basic encoding. Quality assurance was then carried out by editorial teams in Oxford and Michigan. 5% (or 5 pages, whichever is the greater) of each text was proofread for accuracy and those which did not meet QA standards were returned to the keyers to be redone. After proofreading, the encoding was enhanced and/or corrected and characters marked as illegible were corrected where possible up to a limit of 100 instances per text. Any remaining illegibles were encoded as <gap>s. Understanding these processes should make clear that, while the overall quality of TCP data is very good, some errors will remain and some readable characters will be marked as illegible. Users should bear in mind that in all likelihood such instances will never have been looked at by a TCP editor.</p>
  <p>The texts were encoded and linked to page images in accordance with level 4 of the TEI in Libraries guidelines.</p>
  <p>Copies of the texts have been issued variously as SGML (TCP schema; ASCII text with mnemonic sdata character entities); displayable XML (TCP schema; characters represented either as UTF-8 Unicode or text strings within braces); or lossless XML (TEI P5, characters represented either as UTF-8 Unicode or TEI g elements).</p>
  <p>Keying and markup guidelines are available at the <ref target="http://www.textcreationpartnership.org/docs/.">Text Creation Partnership web site</ref>.</p>
</editorialDecl>

Further Reading about EEBO and EEBO-TCP

A Text Creation Partnership Companion. https://www.textpartnership.net/index.html.
Gadd, Ian. The Use and Misuse of Early English Books Online. Literature Compass 6 (2009): 680–692. DOI 10.1111/j.1741-4113.2009.00632.x.
Gavin, Michael. How To Think About EEBO. Textual Cultures: Text, Contexts, Interpretation 11.1–2 (2019): 70–105. DOI 10.14434/textual.v11i1-2.23570.
Gavin, Michael. EEBO and Us. Textual Cultures: Text, Contexts, Interpretation 14.1 (2021): 270–278. DOI 10.14434/tc.v14i1.32860.
Herman, Peter C. EEBO and Me: An Autobiographical Response to Michael Gavin, How to Think About EEBO. Textual Cultures: Text, Contexts, Interpretation 13.1 (2020): 207–216. DOI 10.14434/textual.v13i1.30078.
Kichuk, Diana. Metamorphosis: Remediation in Early English Books Online (EEBO). Literary and Linguistic Computing 22.3 (2007): 291–303. DOI 10.1093/llc/fqm018.
Mäkelä, Eetu, James Misson, Devani Singh, and Mikko Tolonen. Opening the black box of EEBO. Digital Scholarship in the Humanities 41.1 (2026): 236–254. DOI 10.1093/llc/fqaf086.
Maurer, Margaret C. Facsimiles and Transcription: EEBO-TCP and Narratives of Textual Production. Journal of Early Modern Studies 14 (2025): 17–31. DOI 10.36253/jems-2279-7149-16516.
Quiring, Ana. Fingerprints of British Book History: A Feminist Labor History of EEBO. Digital Humanities Quarterly 18.1 (2024). DOI 10.63744/4c9n5tkrpqeq.

Notes

1.LEMDO can commit to doing this work as long as Janelle has grant funding, donated funds, or a revenue trickle to hire RAs. If you would like to hire an RA at your institution to do the remediation (thereby making LEMDO funds go further), LEMDO will gladly train your RA.↑
2.Use Greg’s Bibliography of the English Printed Drama to the Restoration in conjunction with Farmer and Lesser’s Database of Early English Playbooks to ensure that you pick the right publication (especially in those few cases where there are multiple publications in the same year) and that you have the correct STC number.↑

Prosopography

Illya

Illya has a BA in English and Sociocultural Anthropology and an MA in English. Prior to joining the HCMC, he was a PhD candidate in English and Book History at the University of Toronto and worked on Records of Early English Drama and on the Modernist Archives Publishing Project. His work at the HCMC focuses on creating web-based applications for research projects led by members of the faculty of Humanities at the University of Victoria. This involves creating schemas for new and existing datasets, writing XSLT and build files to transform datasets into structured TEI and HTML formats, implementing staticSearch, and ensuring that new projects are Endings Principles compliant.

Janelle Jenstad

Janelle Jenstad is a Professor of English at the University of Victoria, Director of The Map of Early Modern London, and Director of Linked Early Modern Drama Online. With Jennifer Roberts-Smith and Mark Beatrice Kaethler, she co-edited Shakespeare’s Language in Digital Media: Old Words, New Tools (Routledge). She has edited John Stow’s A Survey of London (1598 text) for MoEML and is currently editing The Merchant of Venice (with Stephen Wittek) and Heywood’s 2 If You Know Not Me You Know Nobody for DRE. Her articles have appeared in Digital Humanities Quarterly, Elizabethan Theatre, Early Modern Literary Studies, Shakespeare Bulletin, Renaissance and Reformation, and The Journal of Medieval and Early Modern Studies. She contributed chapters to Approaches to Teaching Othello (MLA); Teaching Early Modern Literature from the Archives (MLA); Institutional Culture in Early Modern England (Brill); Shakespeare, Language, and the Stage (Arden); Performing Maternity in Early Modern England (Ashgate); New Directions in the Geohumanities (Routledge); Early Modern Studies and the Digital Turn (Iter); Placing Names: Enriching and Integrating Gazetteers (Indiana); Making Things and Drawing Boundaries (Minnesota); Rethinking Shakespeare Source Study: Audiences, Authors, and Digital Technologies (Routledge); and Civic Performance: Pageantry and Entertainments in Early Modern London (Routledge). For more details, see janellejenstad.com.

Joey Takeda

Joey Takeda is LEMDO’s Consulting Programmer and Designer, a role he assumed in 2020 after three years as the Lead Developer on LEMDO.

Mahayla Galliford

Project Manager, 2025-present; Assistant Project Manager, 2024-2025; Research Assistant, 2021-present. Mahayla Galliford (she/her) graduated from the University of Victoria with a BA (honours with distinction) in 2024, and an MA English in 2026. Mahayla’s undergraduate research explored early modern stage directions and civic water pageantry. Her SSHRC-funded MA thesis project focuses on transcribing, editing, and encoding early modern girls’ manuscripts, specifically Lady Rachel Fane’s May Masque in collaboration with LEMDO.

Martin Holmes

Martin Holmes has worked as a developer in the UVic’s Humanities Computing and Media Centre for over two decades, and has been involved with dozens of Digital Humanities projects. He has served on the TEI Technical Council and as Managing Editor of the Journal of the TEI. He took over from Joey Takeda as lead developer on LEMDO in 2020. He is a collaborator on the SSHRC Partnership Grant led by Janelle Jenstad.

Navarra Houldin

Training and Documentation Lead 2025–present. LEMDO project manager 2022–2025. Textual remediator 2021–present. Navarra Houldin (they/them) completed their BA with a major in history and minor in Spanish at the University of Victoria in 2022. Their primary research was on gender and sexuality in early modern Europe and Latin America. They are continuing their education through an MA program in Gender and Social Justice Studies at the University of Alberta where they will specialize in Digital Humanities.

Samuel Seaberg

Samuel Seaberg, a University of Victoria English undergrad, enjoys riding his bike. During the summer of 2025, he began working with LEMDO as a recipient of the Valerie Kuehne Undergraduate Research Award (VKURA). Unfortunately, due to his summer being spent primarily in working to establish an edition of Thomas Heywood’s If You Know Not Me, You Know Nobody, Part 2 and consequently working out how to represent multi-text works in a digital space, his bike has suffered severely of sheltered seclusion from the sun. Note: Samuel now works for LEMDO as the Assistant Project Manager, much to his bike’s chagrin.

Tracey El Hajj

Junior Programmer 2019–2020. Research Associate 2020–2021. Tracey received her PhD from the Department of English at the University of Victoria in the field of Science and Technology Studies. Her research focuses on the algorhythmics of networked communications. She was a 2019–2020 President’s Fellow in Research-Enriched Teaching at UVic, where she taught an advanced course on Artificial Intelligence and Everyday Life. Tracey was also a member of the Map of Early Modern London team, between 2018 and 2021. Between 2020 and 2021, she was a fellow in residence at the Praxis Studio for Comparative Media Studies, where she investigated the relationships between artificial intelligence, creativity, health, and justice. As of July 2021, Tracey has moved into the alt-ac world for a term position, while also teaching in the English Department at the University of Victoria.

Bibliography

DEEP: Database of Early English Playbooks. Ed. Alan B. Farmer and Zachary Lesser. 2007. https://deepplaybooks.org/.
Gadd, Ian. The Use and Misuse of Early English Books Online. Literature Compass 6 (2009): 680–692. DOI 10.1111/j.1741-4113.2009.00632.x.
Greg, W.W. Dramatic Documents from the Elizabethan Playhouses: Stage Plots, Actors’ Parts, Prompt Books. 2 vols. Clarendon Press, 1931.
Kichuk, Diana. Metamorphosis: Remediation in Early English Books Online (EEBO). Literary and Linguistic Computing 22.3 (2007): 291–303. DOI 10.1093/llc/fqm018.

Orgography

LEMDO Team (LEMD1)

The LEMDO Team is based at the University of Victoria and normally comprises the project director, the lead developer, project manager, junior developers(s), remediators, encoders, and remediating editors.

Metadata