Orbis Pictus: Zpřístupnění netextových dat z digitálních knihoven

Účel - Projekt "Orbis Pictus - oživení knihy pro kulturní a kreativní odvětví" si klade za cíl zpřístupnit netextový obsah českých digitálních knihoven, který je ve srovnání s textovými daty obtížně dosažitelný a neprohledatelný. Tento článek přináší přehled plánovaných výstupů projektu s důrazem na klíčové výsledky dosažené v prvních dvou letech. Metody - Zpřístupnění netextových objektů v digitalizovaných dokumentech lze rozdělit na tři úlohy: detekci, popis a vyhledání. Identifikaci, lokalizaci a kategorizaci objektů zajistí nástroj AnnoPage, který umožní extrakci popisů objektů a jejich uložení ve standardizovaném formátu. V dalších fázích projektu naváže na AnnoPage nástroj PeopleGator, který identifikuje osoby na fotografiích či kresbách a umožní propojení dokumentů s vyobrazením stejné osoby a vytvoření databáze identifikovaných osob. Projekt bude zakončen softwarovým řešením integrujícím všechny vyvinuté nástroje. Výsledky - V prvních dvou letech projektu byla vytvořena metodika pro zpracování obrazových dokumentů. Ta popisuje způsob detekce netextových objektů, jejich rozdělení do 25 kategorií a zápis informací pomocí mezinárodních standardů, čímž pokládá základ pro nástroj AnnoPage. K detekci objektů je využíván detektor trénovaný na vlastní datové sadě. Detekované objekty jsou popsány pomocí vektorových reprezentací a textových popisů. Originalita/hodnota - Výstupy projektu budou integrovány do České digitální knihovny, což umožní využívání vyvinutých nástrojů širokému spektru knihoven, které platforma agreguje. Orbis Pictus je unikátní projekt v oblasti digital humanities díky rozsáhlému shromáždění netextových dat. Výsledky najdou uplatnění nejen v identifikaci objektů a metadat, ale i ve výzkumu a kulturním a kreativním průmyslu, kde mohou zpřístupněné objekty sloužit jako inspirace pro marketing, vzdělávání, gamifikaci nebo umělou inteligenci.
Purpose - The project "Book Revival for Cultural and Creative Sectors" aims to make the non-textual content of Czech digital libraries easily available, since it is now difficult to access and search compared to textual data. This article provides an overview of the planned outputs of the project, with an emphasis on the key results achieved in the first two years. Method - Accessing non-textual objects in digitized documents can be divided into three tasks: detection, description and retrieval. The identification, localization and categorization of objects will be provided by AnnoPage. This tool will allow extracting object descriptions and storing them in a standardized format. In the next phases of the project, AnnoPage will be followed by PeopleGator, which identifies people in photographs or drawings and allows linking documents depicting the same person and creating a database of identified people. At the project's conclusion, a software solution integrating all the developed tools will be provided. Results - In the first two years of the project, a methodology for processing image documents was developed. This methodology describes how to detect non-text objects, classify them into 25 categories and store this information using international standards, thus laying the foundation for the AnnoPage tool. A detector trained on a custom dataset is used to detect the objects. Detected objects are described using vector representations and textual descriptions. Originality/value - The outputs of the project will be integrated into the Czech Digital Library, which will enable a wide range of libraries aggregated by the platform to use the developed tools. Orbis Pictus is a unique project in the field of digital humanities due to its extensive collection of non-textual data. The results will find applications not only in object and metadata identification, but also in research and the cultural and creative industries, where the detected objects can serve as inspiration for marketing, education, gamification or artificial intelligence.

Keywords

digital libraries , machine learning , image recognition , image retrieval , creative industries

Citation

ITlib. 2024, vol. 2024, issue 2, p. 22-31.
https://doi.org/10.52036/1335793X.2024.2.22-31

Document type

Peer-reviewed

Document version

Published version

Language of document

cs

DOI

10.52036/1335793X.2024.2.22-31

URI

http://hdl.handle.net/11012/252167

Collections

Ústav počítačové grafiky a multimédií

Creative Commons license

Except where otherwised noted, this item's license is described as Creative Commons Attribution 4.0 International

Citace PRO

Full item page

Orbis Pictus: Zpřístupnění netextových dat z digitálních knihoven

Files

Date

Authors

Advisor

Referee

Mark

Journal Title

Journal ISSN

Volume Title

Publisher

ORCID

Altmetrics

Abstract

Description

Keywords

Citation

Document type

Document version

Date of access to the full text

Language of document

Study field

Comittee

Date of acceptance

Defence

Result of defence

DOI

URI

Collections

Endorsement

Review

Supplemented By

Referenced By

Creative Commons license

Citace PRO