TY - JOUR A1 - Reul, Christian A1 - Christ, Dennis A1 - Hartelt, Alexander A1 - Balbach, Nico A1 - Wehner, Maximilian A1 - Springmann, Uwe A1 - Wick, Christoph A1 - Grundig, Christine A1 - Büttner, Andreas A1 - Puppe, Frank T1 - OCR4all—An open-source tool providing a (semi-)automatic OCR workflow for historical printings JF - Applied Sciences N2 - Optical Character Recognition (OCR) on historical printings is a challenging task mainly due to the complexity of the layout and the highly variant typography. Nevertheless, in the last few years, great progress has been made in the area of historical OCR, resulting in several powerful open-source tools for preprocessing, layout analysis and segmentation, character recognition, and post-processing. The drawback of these tools often is their limited applicability by non-technical users like humanist scholars and in particular the combined use of several tools in a workflow. In this paper, we present an open-source OCR software called OCR4all, which combines state-of-the-art OCR components and continuous model training into a comprehensive workflow. While a variety of materials can already be processed fully automatically, books with more complex layouts require manual intervention by the users. This is mostly due to the fact that the required ground truth for training stronger mixed models (for segmentation, as well as text recognition) is not available, yet, neither in the desired quantity nor quality. To deal with this issue in the short run, OCR4all offers a comfortable GUI that allows error corrections not only in the final output, but already in early stages to minimize error propagations. In the long run, this constant manual correction produces large quantities of valuable, high quality training material, which can be used to improve fully automatic approaches. Further on, extensive configuration capabilities are provided to set the degree of automation of the workflow and to make adaptations to the carefully selected default parameters for specific printings, if necessary. During experiments, the fully automated application on 19th Century novels showed that OCR4all can considerably outperform the commercial state-of-the-art tool ABBYY Finereader on moderate layouts if suitably pretrained mixed OCR models are available. Furthermore, on very complex early printed books, even users with minimal or no experience were able to capture the text with manageable effort and great quality, achieving excellent Character Error Rates (CERs) below 0.5%. The architecture of OCR4all allows the easy integration (or substitution) of newly developed tools for its main components by standardized interfaces like PageXML, thus aiming at continual higher automation for historical printings. KW - optical character recognition KW - document analysis KW - historical printings Y1 - 2019 U6 - http://nbn-resolving.de/urn/resolver.pl?urn:nbn:de:bvb:20-opus-193103 SN - 2076-3417 VL - 9 IS - 22 ER - TY - JOUR A1 - Christ, Andreas A1 - Härtl, Patrick A1 - Kloster, Patrick A1 - Bode, Matthias A1 - Leisegang, Markus T1 - Influence of band structure on ballistic transport revealed by molecular nanoprobe JF - Physical Review Research N2 - In this study we characterize the tautomerization of HPc on Cu(111) as a charge-carrier-induced reversible one-electron process. An analysis of the bias-dependent tautomerization rate finds an energy threshold that corresponds to the energy of the N-H stretching mode. By using the tautomerization of the molecule as a detector for charge carrier transport in the so-called molecular nanoprobe (MONA) technique, we provide evidence for an inhomogeneous coupling between the fourfold-symmetric molecule and sixfold-symmetric surface. We conclude the study by comparing the energy dependence of charge carrier transport on the Cu(111) to the Ag(111) surface. While the MONA technique is limited to the detection of hot-electron transport for Ag(111), our data reveal that the lower onset energy of the Cu surface state also allows for the detection of hot-hole transport. The influence of surface and bulk transport on the MONA technique is discussed. KW - tautomerization KW - HPc KW - Cu(111) Y1 - 2022 U6 - http://nbn-resolving.de/urn/resolver.pl?urn:nbn:de:bvb:20-opus-300855 VL - 4 IS - 4 ER -