A Substack post by Ranke de Vries criticising the use of AI for reading medieval Irish manuscripts was going around my circles today [link]. It provoked a lot of reactions, partly because people in digital humanities felt attacked, and partly because the Substack post draws conclusions based on ill-informed premises. The Bluesky post it spread through didn't really help:
Substack post by Ranke de Vries which not only shows that AI won't be taking over Medieval Irish Studies any time soon (thankfully), but also makes an eloquent case in favour of sitting with the difficulty of manuscript transcription. #MedievalSky rankedevries.substack.com/p/ai-and-med...
De Vries' complaints about quality are generally fair. Irish script is poorly represented in the training data, and the current status of handwritten text recognition (HTR) results in a lot of errors for this material. However, what the original post clearly shows is a gap in understanding between digital and traditional humanities scholars.
First of all, all AI is not alike. It is through our universal acceptance of media and marketing that we consider "AI" as synonym for generative AI, or Large Language Models (LLMs), which is not what Transkribus1 or Kraken2 models do. These are models based on machine learning that are trained for a single, specific task: the recognition of layout or text. They provide the same output for the same image every time, and their errors are also visible: wrong letters, broken words, obvious gibberish and we can measure any given model's error rate3 on our own material before deciding to trust it. That use is very different from LLMs which can give you three different answers for the very same question, and when they transcribe, they tend to produce fluent, plausible text that silently "corrects" the scribe, skips a line or invents a word. Thus scholars training HTR models on and for their specific material is a completely different matter than Oxford letting OpenAI train its models on Bodleian collections. On that matter, many of us in the digital humanities share the concerns.
A more serious misunderstanding is how to evaluate the output of HTR models. I don't know about a single colleague in DH who would say the performance of machine-read transcription is comparable to a human palaeographer publishing a text edition. An edition requires meticulous transcription rules, knowledge of the context and expert editorial judgement. By itself, an HTR model does not have these capabilities, as it was taught to recognise that certain pattern in the image should be converted to a certain unicode character. Although it can learn character context implicitly, it still misses the understanding of the context that humans have to identify hard-to-read words and thus produce wrong case ending, an impossible verb form or plain gibberish.
That is one of the reasons why HTR scholars, just like traditional palaeographers, have discussions about transcription rules.4 Whether to expand the abbreviations or keep them as they are, whether we should aim for language accuracy when creating the ground truth5 or keep mistakes in the text, who the target group using the model is and therefore how easy-to-read the output should be. Discussion about abbreviations, for example, does not just revolve around whether to represent the abbreviations in the text as they are or if, and if so, how, to expand them but also in what form: should it be a precomposed ā or combining macron or a MUFI character?6 De Vries' complaint that the CoMMA sample doesn't expand anything is the very result of this discussion. CoMMA deliberately transcribes letters as represented in the text without expanding abbreviations, following the CATMuS transcription guidelines.7 Her specific examples like respoi for respon, are real misreadings, though.
The Substack post is right that Transkribus has trouble identifying large initials. Its models sometimes truly do struggle with that. Especially if they span over multiple lines or are coloured they are just invisible for the model because there is not enough of them in the training data. However, testing the Irish material in itself does not say much. More seriously, de Vries did not try to create a free account and familiarise herself with the tool, its functions and best practices. Instead, she used the Quick Transcription function on the website to prove her point that the a model not trained on Irish hands was not able to recognize Irish language and script features well.
Unfortunately, another misconception in the article, namely that providing several images would lead to the training of an "AI program", again speaks about the author's unfamiliarity with the technology. A single transcribed page that was used for her test won't produce a useful model. That would take much more material.
The Point on Transkribus' business model deserves answer too. Ground truth creation and curation is expert labour and it's fair to ask who profits from it. Models trained on your own transcriptions stay private unless you choose to share them, however there is a clause that allows them to use your data to improve their models.8 Lucky for us, Transkribus is not the only option available here and we can use fully open-source tools like Kraken9 and eScriptorium,10 where nothing leaves our own server.
So yes, the technology is not flawless and it produces error-ridden output. So why we should bother ourselves with it, right? There is a debate regarding CoMMA11 and whether their approach to abbreviations is good or bad, and I'm not going to join that discussion. However, what this database shows is exactly the reason why HTR can be beneficial. Imagine you are researching a niche text that might be preserved in hundreds of manuscripts. You have two options – either there is a catalogue, you go through it and search for all occurrences of your text, and then you look for the digitalised manuscript and confirm that the text exists in it. Or, if there is no catalogue, you can go through the manuscript page by page by yourself and possibly find the text or something new. However, you can also consult the CoMMA website (or a similar project) and try the full text search and maybe discover an occurrence of the text that was not documented in any catalogue, or something related to it which you would otherwise missed. You might argue that there is no fun in full-text search and that you decided to study manuscripts precisely because you can go through the catalogues and visit the archives, and that is okay. The goal of CoMMA is (I hope) to not to rob you of this experience, but to give you an additional option, to allow you to discover additional resources that you would otherwise miss, or to help you find the manuscript from which you took the note and you forgot the shelf-mark.
Another very good use case for error-ridden transcriptions can be for named entity recognition (NER) like people, places and organisations, though there is a caveat: this is where the transcription models fail the most, because rare words are rare in training data, too. Still, even partial results can feed a network chart or a map and lead to questions you wouldn't have thought to ask.
So as you can see, the use of HTR as technology is much broader than just transcribing text for editions. I would even argue that many of the people working with this technology would agree with the opinion that editing a single text manually is much faster and the only way how to accurately represent the text. The technology is not meant to replace a scholarly edition nor produce it for you, understanding palaeography and being able to interpret the content of the text is still a requirement.
Most importantly for all of us, the article shows how poorly our two groups communicate. A scholar took Transkribus' "100 languages" marketing at face value, which is what marketing is for, and concluded that the whole technology is useless. Digital humanities scholars, in turn, are upset about the tone and her dismissal of a technology that was not properly explained in the first place. The explaining part is on us.