Chinese and Tibetan bilingual ancient book knowledge-based method and system

By constructing a multimodal large model that supports Chinese and Tibetan bilingualism, the problem of insufficient support for complex layouts and low-resource languages ​​in ancient Chinese and Tibetan bilingual books is solved, efficient text recognition and entity relationship extraction are achieved, and multimodal knowledge graph is formed to support interdisciplinary research.

CN120047953AActive Publication Date: 2025-05-27INSTITUTE OF ETHNOLOGY & ANTHROPOLOGY CHINESE ACADEMY OF SOCIAL SCIENCES

Patent Information

Application Number
CN202510108191.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively deal with ancient Chinese and Tibetan books on bilingual books, especially when complex layouts, diverse fonts and low-resource languages ​​are insufficient, resulting in greater difficulty in recognition accuracy and in-depth analysis.

Method used

By introducing Tibetan language expanded vocabulary list and supporting pre-training tasks, a multimodal large model that supports both Chinese and Tibetan languages ​​is built, so that it has the ability to process and understand text across languages. This model is used for layout analysis, text recognition, and entity relationship extraction, and uniformly maps information to multimodal knowledge graphs.

Benefits of technology

It has achieved accurate distinction between pictures, Chinese and Tibetan areas in complex ancient books, improved the accuracy of text recognition and entity relationship extraction, significantly lowered the threshold for low-resource language ancient books processing, and formed a multimodal knowledge graph that supports interdisciplinary research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005256010830000021
    Figure BDA0005256010830000021
  • Figure BDA0005256010830000031
    Figure BDA0005256010830000031
  • Figure BDA0005256010830000041
    Figure BDA0005256010830000041
Patent Text Reader

Abstract

The invention relates to the technical field of ancient book processing, and particularly discloses a Chinese and Tibetan bilingual ancient book knowledge-based method and system, and the method comprises the steps: constructing a multi-modal large model which supports the Chinese and Tibetan at the same time through introducing a Tibetan word expansion list and a matched pre-training task, and enabling the multi-modal large model to have cross-language text processing and image understanding capabilities; performing layout analysis by using a multi-modal large model, and automatically identifying and distinguishing a picture region, a Chinese text region and a Tibetan text region in the ancient book; executing entity and relation extraction on the recognized Chinese text and Tibetan text respectively, and extracting core elements and interrelation thereof; and uniformly mapping cross-language and cross-modal text and image information into a queriable knowledge graph to form semantic and associated description of ancient book contents. According to the method, the identification problem caused by Chinese and Tibetan bilingual mixed arrangement is effectively solved, the accuracy and efficiency of automatic analysis and knowledge extraction are greatly improved, and inheritance and propagation of academic resources are promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ancient book processing, and specifically to a method and system for knowledge-based processing of Chinese-Tibetan bilingual ancient books. Background Art

[0002] With the rise of digital humanities research and cross-language information processing technologies, the need for digitalization and knowledge-based processing of ancient books has become increasingly prominent. Especially for Chinese-Tibetan bilingual ancient books, they contain extremely rich academic resources in aspects such as history, culture, religion, and medicine. However, these documents usually have problems such as complex layouts, diverse characters (including Chinese characters, Tibetan letters, as well as charts, illustrations, etc.), difficult-to-unify typesetting methods, and damaged fonts. In order to comprehensively and efficiently explore and utilize this precious heritage, there is an urgent need for a new solution that integrates modern artificial intelligence, multi-modal large models, and cross-language processing. Based on this, the present invention is committed to researching and developing a method and system for constructing a Chinese-Tibetan bilingual ancient book knowledge graph based on multi-modal large models, with the expectation of generating important value at the levels of academic research, cultural protection, and industrial applications.

[0003] Currently, the mainstream OCR technology mainly targets modern typesetting or single-language text scenarios. Facing the complex layouts in Chinese-Tibetan bilingual ancient books with interlaced rows and columns, mixed vertical and horizontal typesetting, and inserted illustrations, the recognition accuracy and block segmentation effect are poor, making it difficult to meet the requirements of subsequent in-depth processing; most existing multi-modal or multi-language large models are pre-trained for high-resource languages such as English and Chinese, with insufficient support for low-resource languages like Tibetan, and lacking processing strategies for the specialized typesetting and diverse fonts of ancient books, resulting in greater difficulties in cross-language and cross-modal in-depth analysis; in addition, the ancient book digitalization process often stays at simple text retrieval or image management, unable to automatically extract entities, concepts, and relationships in the literature, and even more difficult to form a multi-modal knowledge graph for in-depth interdisciplinary research. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for knowledge-based processing of Chinese-Tibetan bilingual ancient books to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] A method for knowledge-based processing of Chinese-Tibetan bilingual ancient books, the method comprising:

[0007] S1. By introducing a Tibetan word expansion table and a supporting pre-training task, construct a multi-modal large model that supports both Chinese and Tibetan, enabling it to have cross-language text processing and image understanding capabilities;

[0008] S2. Use a multi-modal large model for layout analysis, automatically identify and distinguish the picture areas, Chinese text areas, and Tibetan text areas in ancient books, and separate the illustrations or photos to the picture server for subsequent separate processing and display;

[0009] S3. Perform entity and relationship extraction on the identified Chinese text and Tibetan text respectively, and extract the core elements and their interconnections. Among them, the core elements include people, place names, time, medicinal materials, systems, etc.;

[0010] S4. Uniformly map the cross-language and cross-modal text and image information to a queryable knowledge base or knowledge graph to form a semantic and associated description of the content of ancient books.

[0011] As a further technical solution of the present invention, the method further includes data preparation:

[0012] Collect a large amount of Tibetan language materials, clean, segment or sub-letter segment them, and combine them with Chinese language materials for unified multi-modal pre-training; for the scanning of ancient book images, collect diverse images containing different fonts, degrees of mutilation, and typesetting methods to enable the model to maintain robustness in complex scenarios.

[0013] As a further technical solution of the present invention, the steps of constructing a multi-modal large model that supports both Chinese and Tibetan and has cross-language text processing and image understanding capabilities by introducing a Tibetan word expansion table and a supporting pre-training task include:

[0014] Tibetan word list expansion: Add Tibetan syllables or common subsets of Tibetan words to the original word list of the large model, and randomly or incrementally initialize the corresponding embedding vectors;

[0015] Large model pre-training: Combine the language model and the image-text matching task to fine-tune or incrementally train "ancient book scenarios + Chinese-Tibetan bilingual";

[0016] Supervised fine-tuning: Add a translation task head to the model, and perform Tibetan-Chinese translation or bilingual alignment when necessary to solve the mapping problem of the same entity in different languages. To improve the accuracy of cross-language translation and alignment, the following loss function examples can be introduced:

[0017]

[0018] Among them, represents the cross-entropy loss of Chinese-Tibetan translation, can be the cross-language vector alignment loss, and λ 1 and λ 2 are used to balance the relative weights of the translation and alignment tasks;

[0019] To address the problem of scarce data in low-resource languages, a small amount of real optical character recognition (OCR) corpus in Tibetan can be combined with automatically synthesized Tibetan data to improve the model coverage. An example of the loss function is as follows:

[0020]

[0021] Among them, represents the real low-resource Tibetan OCR corpus set, represents the automatically synthesized Tibetan data set, is the training loss. The parameter β ∈ [0, 1] is used to balance the proportion of real data and synthetic data in actual training, ensuring both the adaptation to real scenarios and the expansion of the diversity of Tibetan word forms, fonts, etc., and improving the coverage and robustness of the model in low-resource environments.

[0022] As a further technical solution of the present invention, the steps of using a multi-modal large model for layout analysis, automatically identifying and distinguishing picture areas, Chinese text areas, and Tibetan text areas in ancient books, and separating illustrations or photos to a picture server for subsequent separate processing and display include:

[0023] Graphic and text input: Input an ancient book image of a certain page into the multi-modal large model;

[0024] Chunking result: The multi-modal large model outputs corresponding segmentation masks or bounding boxes to distinguish picture areas, Chinese text areas, and Tibetan text areas;

[0025] Picture storage: Crop the picture area separately and save it to the server, and the picture can be referenced in subsequent knowledge graph construction.

[0026] As a further technical solution of the present invention, the steps of respectively performing entity and relationship extraction on the identified Chinese text and Tibetan text, and extracting core elements and their interconnections include:

[0027] 1), Chinese text: Obtain Chinese characters through OCR and use natural language processing (NLP) tools to identify entities {E h} and relationships {R h};

[0028] 2), Tibetan text: Use a large model to recognize Tibetan glyphs to obtain a text sequence, then perform word segmentation or sub-syllable decomposition, and extract entities {E t} and relationships {R t}. If bilingual merging is required, the cosine similarity of similar entity vectors can be judged. The formula is as follows:

[0029]

[0030] When the similarity is greater than the threshold θ, they can be regarded as the same concept.

[0031] As a further technical solution of the present invention, the step of uniformly mapping cross - language and cross - modal text and image information into a queryable knowledge graph to form a semantic and associative description of the content of ancient books includes:

[0032] Multi - modal data integration: Using a graph database, store nodes in the graph and connect them to each other through relationships. The nodes include person names (Person), place names (Location), events (Event), illustration resources (ImageResource), etc., and the relationships include belonging to, occurring in, being located in, etc.

[0033] Cross - language entity linking: If the entities represented by Tibetan and Chinese texts are the same, merge them or establish a synonym link in the graph to facilitate cross - language retrieval by users, as shown in the following formula:

[0034]

[0035] Among them, vt and vh respectively represent the embedding vectors of Tibetan and Chinese entities, sim(·) is the cosine similarity function, and θ is the similarity threshold for determining whether to merge or link entities; when the similarity is greater than θ, it indicates that the Tibetan entity and the Chinese entity have a very high correspondence relationship and can be regarded as the same object, so as to perform unified processing or cross - language synonym linking in the knowledge graph.

[0036] Visualization and interface: Provide a visualization page or API to support researchers in searching, reasoning, and statistical analysis in the graph.

[0037] Another object of the present invention is to provide a knowledge - based system for Chinese - Tibetan bilingual ancient books, and the system includes:

[0038] A multi - modal large - model construction module for constructing a multi - modal large model that supports both Chinese and Tibetan by introducing a Tibetan word expansion table and a supporting pre - training task, enabling it to have cross - language text processing and image understanding capabilities.

[0039] A layout analysis and block recognition module for using the multi - modal large model for layout analysis, automatically identifying and distinguishing the picture area, Chinese text area, and Tibetan text area in ancient books, and separating illustrations or photos to the picture server for subsequent separate processing and display.

[0040] A text recognition and entity relationship extraction module for respectively performing entity and relationship extraction on the recognized Chinese text and Tibetan text, extracting core elements and their interrelationships. Among them, the core elements include characters, place names, time, medicinal materials, systems, etc.

[0041] The multimodal knowledge graph construction module is used to uniformly map cross - language and cross - modal text and image information into a queryable knowledge base or knowledge graph, forming a semantic and associated description of the content of ancient books.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] By constructing a multimodal large model that supports Chinese - Tibetan bilingual, the present invention can accurately distinguish picture, Chinese and Tibetan regions in complex ancient book layouts, and efficiently complete text recognition and entity relationship extraction, thus significantly reducing the processing threshold for ancient books in low - resource languages. By jointly mapping the recognition results and picture information into a multimodal knowledge graph, the multi - dimensional content such as culture, history, and medicine contained in ancient books can be comprehensively displayed, realizing in - depth support for cross - language and cross - disciplinary research.

[0044] Compared with traditional methods, the present invention not only effectively solves the recognition problems brought by Chinese - Tibetan bilingual mixed layout, but also takes into account the particularities such as complex ancient book layout, diverse fonts, and interspersed illustrations, greatly improving the accuracy and efficiency of automatic parsing and knowledge extraction. In addition, through the two - way alignment and translation training of Chinese - Tibetan texts, researchers can freely query and compare the literature content in the two languages, promoting the inheritance and dissemination of academic resources, and providing a broader application prospect for digital humanities, museum informatization, and cultural exchanges. Through the above innovative design, the present invention demonstrates a feasible and practical new method in the fields of ancient book digitization, knowledge - based, and cross - language information processing, bringing significant improvement space to related scientific research, cultural heritage protection, and education work. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.

[0046] Figure 1 It is a flowchart of a method for knowledge - based processing of Chinese - Tibetan bilingual ancient books. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] In order to make the technical problems, technical solutions, and beneficial effects to be solved by the present invention more clearly understood, the following further details the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0048] The present invention aims to utilize a multi-modal large model to perform graphic layout recognition, bilingual text parsing, and knowledge extraction on Chinese-Tibetan bilingual ancient books, and finally form a multi-modal knowledge graph to support interdisciplinary research. Its process includes: First, based on the existing Chinese multi-modal model, synchronous support for Chinese-Tibetan bilingual texts is achieved through "Tibetan word expansion table + language model pre-training + supervised translation task"; then, for the complex layout of ancient books, the picture, Chinese, and Tibetan regions are segmented, and the image part is stored separately for subsequent associated display; subsequently, OCR and entity relationship extraction are respectively performed on each text region, and Chinese-Tibetan translation or alignment is assisted when necessary; then, the identified cross-language entities and relationships are mapped to the graph structure (such as RDF or graph database), and multi-modal information such as illustrations, geographical coordinates, and time axes are linked; finally, this knowledge graph can be applied to research in fields such as culture, medicine, history, and literature, and can also be used for more in-depth analysis and display in digital humanities and museum systems.

[0049] Please refer to Figure 1 , an embodiment of the present invention provides a method for knowledge-based Chinese-Tibetan bilingual ancient books, and the method includes:

[0050] S1. Collect a large amount of Tibetan language materials, clean, segment, or sub-letter segment them, and uniformly use them for multi-modal pre-training in combination with Chinese language materials; for the scanning of ancient book images, collect diverse images including different fonts, degrees of mutilation, and typesetting methods, so that the model can maintain robustness in complex scenarios.

[0051] S2. By introducing a Tibetan word expansion table and a supporting pre-training task, construct a multi-modal large model that supports both Chinese and Tibetan, enabling it to have cross-language text processing and image understanding capabilities;

[0052] Expansion of the Tibetan word list: Add Tibetan syllables or a subset of common Tibetan words to the word list of the original large model, and randomly or incrementally initialize the corresponding embedding vectors;

[0053] Pre-training of the large model: Combine the language model (Chinese + Tibetan) and the graphic-text matching tasks (detection of text blocks in images, alignment of images and texts, etc.) to perform fine-tuning or incremental training on "ancient book scenarios + Chinese-Tibetan bilingual";

[0054] Supervised fine-tuning: Add a translation task head to the model, and perform mutual translation or bilingual alignment between Tibetan and Chinese when necessary to solve the mapping problem of the same entity in different languages. To improve the accuracy of cross-language translation and alignment, the following example of a loss function can be introduced:

[0055]

[0056] Among them, represents the cross-entropy loss of Chinese-Tibetan translation, can be a cross - language vector alignment loss (such as the negative logarithm of the vector cosine similarity), and λ 1 and λ 2 are used to balance the relative weights of the translation and alignment tasks;

[0057] To address the problem of scarce data for low - resource languages, a small amount of real OCR corpus can be combined with automatically synthesized Tibetan data to improve the model coverage. An example of the loss function is as follows:

[0058]

[0059] where, represents the real low - resource Tibetan OCR corpus set, represents the automatically synthesized Tibetan data set, is the training loss (such as OCR or language model loss). The parameter β ∈ [0, 1] is used to balance the proportion of real data and synthetic data in actual training, ensuring both the adaptation to real scenarios and the expansion of the diversity of Tibetan word forms, fonts, etc., and improving the model coverage and robustness in low - resource environments.

[0060] S3. Use a multi - modal large model for layout analysis, automatically identify and distinguish the picture area, Chinese text area, and Tibetan text area in ancient books, and separate the illustrations or photos to the picture server for subsequent separate processing and display;

[0061] Graphic input: Input the image of a certain page of an ancient book into the multi - modal large model;

[0062] Chunking result: The multi - modal large model outputs the corresponding segmentation mask or position box to distinguish the picture area, Chinese text area, and Tibetan text area;

[0063] Picture storage: Crop and save the picture area separately to the server, and the picture can be referenced in the subsequent knowledge graph construction.

[0064] S4. Perform entity and relationship extraction on the identified Chinese text and Tibetan text respectively, and extract the core elements and their inter - relationships. Among them, the core elements include people, place names, time, medicinal materials, systems, etc.;

[0065] 1). Chinese text: Obtain Chinese characters through OCR and use NLP tools (word segmentation, NER, relationship extraction) to identify entities {E h} and relationships {R h};

[0066] 2). Tibetan text: Use a large model to recognize the Tibetan word form to obtain the text sequence, then perform word segmentation or sub - syllable decomposition, and extract entities {E t} and relationships {R t}, if bilingual merging is required, the cosine similarity of similar entity vectors can be judged, and the formula is as follows:

[0067]

[0068] When the similarity is greater than the threshold θ, it can be regarded as the same concept.

[0069] S5. Uniformly map cross - language and cross - modal text and image information into a queryable knowledge base or knowledge graph to form a semantic and associated description of the content of ancient books.

[0070] Multi - modal data integration: Using a graph database, store nodes in the graph and connect them through relationships. The nodes include person names (Person), place names (Location), events (Event), illustration resources (ImageResource), etc., and the relationships include belonging to, occurring in, being located in, etc.;

[0071] Cross - language entity linking: If the entities represented by Tibetan and Chinese texts are the same, they are merged in the graph or a synonym link is established to facilitate cross - language retrieval by users, as shown in the following formula:

[0072]

[0073] Among them, vt and vh respectively represent the embedding vectors of Tibetan and Chinese entities (which can be obtained by a multi - modal large - model or other cross - language models), sim(·) is the cosine similarity function, and θ is the similarity threshold for determining whether to merge or link entities; when the similarity is greater than θ, it indicates that the Tibetan entity and the Chinese entity have a very high correspondence relationship and can be regarded as the same object, so as to be uniformly processed or cross - language synonym - linked in the knowledge graph;

[0074] Visualization and interface: Provide a visualization page or API to support researchers in searching, reasoning, and statistical analysis in the graph.

[0075] This graph can not only support in - depth research in fields such as culture, medicine, and history, but also provide accurate data support for applications such as digital humanities, museum informatization, and education popularization. By integrating low - resource language processing, complex layout analysis, and entity relationship extraction, the present invention realizes the efficient knowledge - based conversion of Chinese - Tibetan bilingual ancient books, significantly improving the deficiencies of existing methods in recognition accuracy, cross - language compatibility, and semantic structure mining, and is of great significance for promoting the digitization of ancient books and academic research.

[0076] Another object of the embodiment of the present invention is to provide a knowledge - based system for Chinese - Tibetan bilingual ancient books, and the system includes:

[0077] The multi-modal large model construction module is used to construct a multi-modal large model that supports both Chinese and Tibetan by introducing a Tibetan word expansion table and a supporting pre-training task, enabling it to have cross-language text processing and image understanding capabilities;

[0078] The layout analysis and block recognition module is used to perform layout analysis using the multi-modal large model, automatically identify and distinguish the picture area, Chinese text area, and Tibetan text area in ancient books, and separate the illustrations or photos to the picture server for subsequent separate processing and display;

[0079] The text recognition and entity relationship extraction module is used to perform entity and relationship extraction on the recognized Chinese text and Tibetan text respectively, and extract the core elements and their interconnections. Among them, the core elements include people, place names, time, medicinal materials, systems, etc.;

[0080] The multi-modal knowledge graph construction module is used to uniformly map cross-language and cross-modal text and image information into a queryable knowledge base or knowledge graph, forming a semantic and associated description of the content of ancient books.

[0081] It should be noted that in this article, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article, or system. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or system including that element.

[0082] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A method for converting Chinese and Tibetan ancient books into knowledge, characterized in that: The method comprises: S1. By introducing the Tibetan vocabulary expansion table and supporting pre-training tasks, a multimodal large model that supports both Chinese and Tibetan is constructed, enabling it to have cross-language text processing and image understanding capabilities; S2. Use a multimodal large model to perform layout analysis, automatically identify and distinguish image areas, Chinese text areas, and Tibetan text areas in ancient books, and separate illustrations or photos to an image server for subsequent separate processing and display; S3, performing entity and relationship extraction on the identified Chinese text and Tibetan text respectively, extracting the core elements and their mutual connections, wherein the core elements include people, place names, time, medicinal materials and systems; S4. Unify and map cross-language and cross-modal text and image information into a searchable knowledge graph to form a semantic and associative description of the content of ancient books.

2. The method for converting Chinese-Tibetan bilingual ancient books into knowledge according to claim 1 is characterized in that: The method also includes data preparation: A large amount of Tibetan corpus is collected, cleaned, word-segmented or sub-letter-segmented, and then combined with Chinese corpus for multimodal pre-training. For ancient book image scanning, diverse images with different fonts, degrees of damage, and typesetting methods are collected so that the model can remain robust in complex scenarios.

3. The method for converting Chinese-Tibetan bilingual ancient books into knowledge according to claim 1 is characterized in that: The steps of introducing the Tibetan extended vocabulary and supporting pre-training tasks to construct a multimodal large model that supports both Chinese and Tibetan languages, so that it has cross-language text processing and image understanding capabilities, include: Tibetan vocabulary expansion: Add Tibetan syllables or a subset of common Tibetan words to the vocabulary of the original large model, and perform random or incremental initialization on the corresponding embedding vectors; Large model pre-training: Combine language models with image-text matching tasks to perform fine-tuning or incremental training on "ancient book scenes + Chinese-Tibetan bilingual"; Supervised fine-tuning: Add a translation task head to the model to translate Tibetan and Chinese or perform bilingual alignment to solve the mapping problem of the same entity in different languages. To improve the accuracy of cross-language translation and alignment, introduce the following loss function: in, represents the cross entropy loss of Chinese-Tibetan translation, It can be the cross-language vector alignment loss, while λ1 and λ2 are used to balance the relative weights of translation and alignment tasks; In order to solve the problem of lack of low-resource language data, the real text recognition corpus is combined with the automatically synthesized Tibetan data to improve the model coverage. The loss function is as follows: in, represents a real low-resource Tibetan OCR corpus, represents the automatically synthesized Tibetan dataset, is the training loss, and the parameter β∈[0, 1] is used to balance the proportion of real data and synthetic data in actual training.

4. The method for converting Chinese-Tibetan bilingual ancient books into knowledge according to claim 1 is characterized in that: The steps of using a multimodal large model to perform layout analysis, automatically identifying and distinguishing image areas, Chinese text areas, and Tibetan text areas in ancient books, and separating illustrations or photos to an image server for subsequent separate processing and display include: Image and text input: input a multimodal large model into an image of a certain page of ancient books; Segmentation results: The multimodal large model outputs the corresponding segmentation mask or location box to distinguish the image area, Chinese text area, and Tibetan text area; Image storage: Crop the image area separately and save it to the server. The image can be referenced in subsequent knowledge graph construction.

5. The method for converting Chinese-Tibetan bilingual ancient books into knowledge according to claim 1 is characterized in that: The steps of performing entity and relationship extraction on the recognized Chinese text and Tibetan text respectively to extract core elements and their mutual relations include: 1) Chinese text: Obtain Chinese characters through OCR and use natural language processing tools to identify entities {E h } and the relation {R h }; 2) Tibetan text: Use a large model to recognize Tibetan characters, obtain text sequences, perform word segmentation or subsyllable decomposition, and extract entities {E t } and the relationship {R t }, if bilingual merging is required, the cosine similarity is determined for similar entity vectors. The formula is as follows: When the similarity is greater than the threshold θ, they can be considered as the same concept.

6. The method for converting Chinese-Tibetan bilingual ancient books into knowledge according to claim 1 is characterized in that: The steps of uniformly mapping cross-language and cross-modal text and image information into a queryable knowledge graph to form a semantic and associative description of the contents of ancient books include: Multimodal data integration: Using a graph database, nodes are stored in a graph and connected to each other through relationships. The nodes include names of people, places, events, and illustrations, and the relationships include belonging to, occurring in, and located in; Cross-language entity linking: If the entities represented by Tibetan and Chinese texts are the same, they are merged or synonyms are established in the graph to facilitate cross-language search by users, as shown in the following formula: Where vt and vh represent the embedding vectors of Tibetan and Chinese entities respectively, sim(·) is the cosine similarity function, and θ is the similarity threshold used to determine whether to merge or link entities. When the similarity is greater than θ, it means that the Tibetan entity and the Chinese entity have a very high correspondence and can be regarded as the same object, so they can be uniformly processed or linked across languages ​​in the knowledge graph. Visualization and interface: Provide visualization pages or APIs to support researchers in searching, reasoning, and statistical analysis in graphs.

7. A knowledge system for Chinese-Tibetan bilingual ancient books, characterized by: The system comprises: The multimodal large model construction module is used to build a multimodal large model that supports both Chinese and Tibetan by introducing a Tibetan vocabulary expansion table and supporting pre-training tasks, so that it has cross-language text processing and image understanding capabilities; The layout analysis and block recognition module is used to perform layout analysis using a multimodal large model, automatically identify and distinguish image areas, Chinese text areas, and Tibetan text areas in ancient books, and separate illustrations or photos to an image server for subsequent separate processing and display; The text recognition and entity relationship extraction module is used to perform entity and relationship extraction on the recognized Chinese text and Tibetan text respectively, and extract the core elements and their mutual relations, wherein the core elements include people, place names, time, medicinal materials and systems; The multimodal knowledge graph construction module is used to uniformly map cross-language and cross-modal text and image information into a queryable knowledge graph, forming a semantic and associative description of the content of ancient books.

Citation Information

Patent Citations

  • Tibetan language entity knowledge information extraction method

    CN104133848A

  • Entity relationship extracting method of Zang language

    CN104809176A

  • Method and system for constructing Tibetan Andor dialect speech synthesis corpus

    CN116798404A

  • Chinese and Tibetan language multi-mode image-text processing method and processing system

    CN118709147A

Cited By

  • Method and system for constructing and interacting Yao language corpus based on artificial intelligence

    CN121833926A