Multi-modal ancient poetry knowledge graph construction method for mutual conversion of ancient poetry and image

By constructing a multimodal ancient poetry knowledge graph, combining natural language processing and computer vision technology, the text of ancient poetry and image information is integrated, and the problem of deviations in the understanding of ancient poetry and insufficient multimodal associations in the existing technology is solved, and high-accurate semantic analysis and image generation of ancient poetry is achieved.

CN120179828APending Publication Date: 2025-06-20BEIJING UNIV OF TECH

Patent Information

Application Number
CN202510095342.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When processing the content of ancient poems, the existing Chinese knowledge graphs have deviations in understanding due to differences in the meanings of ancient and modern languages, resulting in the inconsistent image generation results that are inconsistent with the context of ancient poems and lack multimodal correlation with the real world, making it difficult to accurately support the semantic analysis and image generation tasks of ancient poems and content.

Method used

The multimodal ancient poetry knowledge graph construction method is used to analyze ancient poetry texts through natural language processing technology, and computer vision technology is used to understand image information related to ancient poetry, organically integrate text and image information to build a multimodal knowledge graph. Specific steps include crawling ancient poetry data, cleaning and counting word frequency, obtaining the interpretation of candidate words, calculating the similarity between interpretations, establishing entity relationships, and using multimodal models for text-to-picture mapping.

Benefits of technology

The accurate correspondence between ancient poems and images is achieved, the semantic analysis accuracy and image generation quality of the knowledge graph are improved, and the multimodal application of ancient poems is better supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179828A_ABST
    Figure CN120179828A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal ancient poetry knowledge graph construction method for mutual conversion of ancient poetries and images, which comprises the following steps of: taking text interpretation as a text entity of a knowledge graph by utilizing ancient poetry information in search, annotation information in Chinese poetry and translation information in the ancient poetry; the method comprises the following steps: establishing a semantic relation between mutually associated entities to obtain an ancient poetry knowledge graph, and generating an image for an ancient poetry text by using a multi-modal model to obtain a multi-modal ancient poetry knowledge graph. In a text-to-image conversion task, compared with a traditional ancient poetry knowledge graph, only explanation of the ancient poetry is reserved, the one-to-many relation between the ancient poetry and the explanation is ignored, and a corresponding relation can be better found from candidate ancient poetry, so that an image with more accurate corresponding relation and better quality is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of natural language processing and computer vision. Specifically, it relates to a method for constructing a multimodal ancient Chinese poetry knowledge graph based on text and image and its application. This method uses natural language processing technology to analyze ancient Chinese poetry texts, and at the same time uses computer vision technology to understand images related to ancient Chinese poetry, and organically integrates text and image information to construct a multimodal knowledge graph. Background Art

[0002] As a large-scale semantic network structure with entities and concepts as nodes and semantic relationships as edges, knowledge graphs have been widely used in text-to-image conversion tasks in recent years. However, in the scenario of mutual conversion between ancient Chinese poetry and images, the construction of existing Chinese knowledge graphs is usually based on modern Chinese. When dealing with the content of ancient Chinese poetry, such knowledge graphs based on modern semantics often have understanding deviations due to the differences in the meanings of ancient and modern languages, which in turn leads to the inconsistency between the generated image results and the context of ancient Chinese poetry. For example, in the sentence "A heroic man in his old age still cherishes high aspirations", the word "heroic man" in the ancient context refers to "a man with ambitions and lofty aspirations", while in modern Chinese it is understood as "a person who sacrifices for a just cause". This semantic deviation makes it difficult for existing knowledge graphs to accurately support the semantic analysis and image generation tasks of ancient Chinese poetry content.

[0003] Existing Chinese knowledge graphs are mostly presented in pure text form and lack multimodal associations with the real world. Especially in the process of corresponding ancient Chinese poetry texts with images, the unimodal knowledge expression cannot effectively enhance the model's understanding of semantics. For example, for the entity "pine tree", it is difficult to accurately establish its association with the image of a pine tree in the real world only through text description. By introducing image information, the semantic analysis ability of the knowledge graph for entities can be significantly enhanced, helping the model to generate more context-compliant images in the text-to-image conversion task. At the same time, for the processing of large-scale image-text pair datasets in common corpora, there is also an urgent need for an efficient evaluation method to accurately measure the semantic correspondence relationship between images and texts.

[0004] Chinese culture has a long history and is profound. Ancient poetry, as an important part of it, expresses rich images and emotions in concise words. Many words in ancient poetry have a high degree of semantic polysemy and may express completely different meanings in different contexts. For example, the sentence "the white day relies on the mountain and ends" can be divided into four words: "white day", "rely", "mountain" and "end". Among them, "white day" may mean "daytime", "sun" or "time"; "rely" may mean "rely on", "according to", "obey" or "intimate"; "mountain" may refer to "the towering part formed by the ground", "an object shaped like a mountain" or other symbolic meanings; "end" may mean "finished", "extreme" or "all". Accurately identifying the specific semantics of these polysemous words and establishing a correspondence with objects or scenes in the real world is one of the core challenges in realizing semantic parsing and image generation of ancient poetry. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention aims to provide a method for constructing a multimodal knowledge graph, proposes a technical solution that effectively solves the difficulty of matching text and images in public corpora, and applies it to the field of ancient poetry to construct a knowledge graph dedicated to the mutual conversion between ancient poetry and images. The graph has a wide coverage and high semantic analysis accuracy, providing technical support for the multimodal application of ancient poetry.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for constructing a multimodal ancient poetry knowledge graph for converting ancient poetry and images, comprising the following steps:

[0008] Step 1. Use Python language to write a web crawler algorithm, take the Souyun website as the crawling target, crawl all the poems, lyrics, songs, etc. included in Souyun from the Pre-Qin Dynasty to the modern times, take the ancient poetry website as the target, crawl the corresponding explanations of ancient poems, take the Guoxue Huicui website as the target, crawl the author information and poetry information, and crawl the dynasty-related information from Baidu Encyclopedia.

[0009] Step 2: Clean the crawled ancient poems, and then count the word frequencies of the cleaned ancient poems. Words that appear more than 5 times will be regarded as text candidate words for the multimodal ancient poetry knowledge graph.

[0010] Step 3: Use Python to write a web crawler algorithm, take Han Dian as the crawling target, crawl the explanations of all existing candidate words, clean, sentence and format all the explanations to obtain the multi-group between ancient poems and explanations, and use the explanations of ancient poems to form the text entities of the multimodal ancient poems knowledge graph.

[0011] Step 4: Calculate the similarity between different ancient poetry interpretations and establish connections between related entities

[0012] Step 5: Use the existing multimodal model to convert the text into the corresponding image and use the multimodal model to ensure that the text and the image correspond to each other to obtain a multimodal knowledge graph

[0013] Furthermore, the specific process of step 2 is as follows:

[0014] Step 2.1: Remove all punctuation marks and other useless special characters from ancient poems.

[0015] Step 2.2: Continuously split the ancient poems according to their length, count the number of occurrences of ancient poems of different lengths, and select those that appear more than 5 times as candidate words.

[0016] Furthermore, the specific process of step 3 is as follows:

[0017] Step 3.1: Use Python to write a web crawler algorithm, take Han Dian as the crawling target, crawl the explanations of all candidate words, and count all candidate words with existing explanations.

[0018] Step 3.2: Remove useless special characters in the explanation of ancient poems

[0019] Step 3. Divide the cleaned ancient poetry multi-word group into two copies S and P, and delete the parts that supplement the explanation, such as the source and examples, by writing a Python algorithm to construct a knowledge graph G, in which the parts that appear in P are used as the explanation of ancient poetry, and the parts that appear in S but not in P are used as supplements to ancient poetry.

[0020] Furthermore, the specific process of step 4 is as follows:

[0021] Step 4.1 uses the word embedding algorithm to embed all entities into corresponding word vectors. By calculating the cosine similarity between different word vectors, "synonyms" are used as the unified relationship between entities. For entities whose similarity exceeds the threshold, the "entity-synonym-entity" ternary knowledge is established.

[0022] Furthermore, the specific process of step 5 is as follows:

[0023] Step 5.1: Use a multimodal model to map text to images. For entities with "synonym" relationships, only one entity needs to be selected to use a multimodal model to map text to images, and a multimodal ancient poetry knowledge graph is obtained.

[0024] Step 5.2: The existing text-image calculation method calculates the CLIP score between the text-image pair as the final evaluation index. The present invention converts the regression problem into a classification problem by setting interference items to match the correct answer, thereby efficiently realizing image-text alignment.

[0025] The present invention utilizes the ancient poetry information in Soyun, the annotation information in Han Dian, and the translation information in ancient poetry, takes the text interpretation as the text entity of the knowledge graph, establishes semantic connections between interrelated entities, obtains the ancient poetry knowledge graph, uses a multimodal model to generate images for the ancient poetry text, and obtains a multimodal ancient poetry knowledge graph. In the text-to-image conversion task, compared with the traditional ancient poetry knowledge graph that only retains the interpretation of ancient poetry and ignores the one-to-many relationship between ancient poetry and interpretation, we can better find the corresponding relationship from the candidate ancient poetry, thereby generating images with more accurate corresponding relationships and better quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 : The construction process of multimodal ancient poetry knowledge graph.

[0027] Figure 2 : Examples of ancient poetry from Souyun (left) and examples of ancient poetry interpretation from Handian (right).

[0028] Figure 3 :Chinese character statistical method.

[0029] Figure 4 : The format of the information after processing in the present invention.

[0030] Figure 5 : An example of the knowledge graph constructed by the present invention. DETAILED DESCRIPTION

[0031] The following examples further explain the specific contents of the present invention in detail.

[0032] like Figure 1 The method for constructing a multimodal ancient poetry knowledge graph for converting ancient poetry and images comprises the following steps:

[0033] Step 1: Obtain candidate words for ancient poetry

[0034] Souyun.com uses computer technology to deeply mine and integrate relevant humanities data, develop numerous poetry tools, and disseminate poetry and literary resources through network technology. After years of hard work, Souyun.com has developed into a complete scale with complete functions. With its expertise in poetry and powerful computer technology, Souyun.com has been able to stand at the forefront of the digital humanities field and has a wide influence in the global poetry community. Therefore, it was decided to crawl data for Souyun. The web crawler algorithm was written in Python language, with dynasties as the distinction. Each time a dynasty was crawled and saved separately, all poems, lyrics, and songs included in Souyun from the pre-Qin period to modern times were crawled, the sentence explanations in ancient poems and essays were crawled, and the author information, poetry information, and dynasty information in Guoxuehui were crawled.

[0035] The ancient poems from various dynasties that were crawled were summarized, cleaned, special characters were removed, format problems were fixed, and word frequency statistics were performed on the cleaned ancient poems. Words that appeared more than 5 times were used as text candidate words for the multimodal ancient poetry knowledge graph. Figure 1 As shown in the figure, when counting word frequency, the UTF-8 encoding of Chinese characters is combined with the abstract tree in the data structure. Chinese characters occupy 2 to 4 bytes in the UTF-8 encoding. For each Chinese character, each byte after the UTF-8 encoding is used as a layer of branch nodes. The last leaf node not only saves the number of times the Chinese character appears, but also stores a pointer to the next leaf node. Finally, the number of leaf nodes can be obtained by traversing all leaf nodes.

[0036] Step 2: Obtain candidate word explanations

[0037] HanDian was founded in 2004. It is a free online dictionary with a huge capacity of characters, words, phrases, idioms and other Chinese language forms. The number of word explanations contained in HanDian and the richness of word explanations are unmatched by other platforms, so HanDian was chosen as the platform for crawling keyword explanations. The web crawler algorithm was written in Python language, with HanDian as the crawling target, crawling all candidate word explanations. If the candidate word explanation exists, it will be retained, and if it does not exist, it will be discarded.

[0038] The keywords of the crawled explanations have problems such as dirty data, format issues, reference issues, and redundant information. For dirty data such as line breaks, whitespace characters, and tags for modifying text, write scripts in the Python language to uniformly clean the dirty data; for format issues, correct them in batches by writing Python scripts; the reference issue refers to the fact that the explanations of many words with similar meanings are given by referring to another word rather than directly stating the meaning. For example, the explanation of "sacrificial feast" is: the same as "sacrificial offering", and the explanation of "sacrificial offering" needs to be referred to "sacrificial feast". This situation exists in multiple keywords such as "also written as", "also called", "the same as", "see", etc. For this situation, it is necessary to first determine whether the referred word contains a meaning or continues to refer to other words until a word with a meaning is found and referred to the target word. If the synonym does not exist in the knowledge graph, record it and then perform secondary crawling on it. For redundant information, we use a comparison method. Keep one copy unchanged, and perform batch processing on the other copy using the Python language to delete the irrelevant information. Then write a Python script for batch comparison. First, the comparison interval is the length of the poem explanation to remove the repeatedly appearing words, and then the comparison interval is reduced by one to remove the repeated words. Repeat until the matching interval reaches 3 and stop to prevent the loss of semantics caused by the accidental appearance of shorter words. Take all the remaining words as supplementary explanations for the current poem. For example, "mountain seedlings" has a meaning of "the newly grown plants and trees on the mountain. Metaphor for mediocre talents in hereditary high positions", and the supplementary explanation is "quoted from Zuo Si's 'Ode to History' II in the Jin Dynasty: 'The lush pines at the bottom of the ravine, the scattered mountain seedlings... The aristocrats step into high positions, while the outstanding talents sink to lower positions.'"

[0039] Step 3: Construct entities and relationships

[0040] Take the word explanations processed in step 2 as the text entities of the knowledge graph, a total of 250,000. Use the sentosa / ZNV-Embedding model to embed all text entities into word vectors. The sentosa / ZNV-Embedding model has excellent performance in calculating Chinese text similarity and ranks first in the Chinese text similarity model leaderboard of mteb (June 2024). Therefore, choosing sentosa / ZNV-Embedding for text embedding can obtain relatively accurate results. For the word vectors of all text entities, calculate their cosine similarities pairwise. Add the relationship "synonym" between two entities whose cosine similarity exceeds the specified threshold, thus obtaining the text knowledge graph. When performing text embedding, all text entities will be embedded into a vector with a length of 4096. The total storage of all text entity vectors is 15GB. In addition, since it is necessary to calculate the similarity between any two entities, the number of calculations is 250,000 * 250,000 / 2, which is also difficult to calculate in a short time. These two items have a huge overhead on the system. Therefore, the sharding idea in the operating system is adopted to shard the vectors after text entity embedding. Each 2000 explanations form a file, and the 15GB file is split into 126 files. When calculating text similarity, only read the required text into memory, greatly reducing the system overhead. When calculating similarity, adopt the idea of bubble sort. Calculate the vectors of the i-th file and the k-th file (k = i, i + 1, i + 2... 126), calculate the vectors in the (i + 1)-th file and the k-th file (k = i, i + 1, i + 2... 126), until the vectors of the 126th file calculate the cosine similarity with themselves. Since the calculation of the i-th file and the k-th file (k = i, i + 1, i + 2... 126) does not affect the calculation of the (i - 1)-th file, if the computing power is sufficient, a separate thread can be opened for each file, and multi-threaded simultaneous calculation can greatly reduce the calculation efficiency.

[0041] Step 4: Construct a multimodal knowledge graph

[0042] For all text entities, the multi-modal model Taiyi-Diffusion-XL (Taiyi-XL) is used for text-to-image generation. Taiyi-XL is a model in the Fengshenbang project, which was officially announced and launched by Xiangyang Shen at the IDEA conference. Because the basic models, especially language models, are currently dominated by the English community. Text-to-image models such as Google's Imagen, OpenAI's DALL-E 3, and StabilityAI's Stable Diffusion have led a new wave of AIGC and digital art creation. However, the performance of Chinese text-to-image models based on SD v1.5, such as Taiyi-Diffusion-v0.1 and Alt-Diffusion, is still mediocre. Many AI painting platforms in China only support English or rely on translation tools for Chinese-to-English translation. Currently, the open-source text-to-image models mainly support English, with limited bilingual support. Based on these developments, Taiyi-XL focuses on enhancing Chinese text-to-image generation while retaining English comprehension ability. It achieves efficient vocabulary expansion by integrating the most commonly used Chinese characters into the tokenizer and embedding layer of CLIP, and also adds an extension of absolute position encoding. In addition, Taiyi-X enriches text prompts through large vision-language models, obtaining better image captions and higher visual quality. These enhancements are then applied to downstream text-to-image models. Empirical results show that the developed CLIP model performs well in bilingual image-text retrieval. In addition, the bilingual image generation ability of Taiyi-Diffusion-XL exceeds that of previous models, especially for Chinese applications. Therefore, I choose Taiyi-X as the text-to-image model to generate images. The obtained images are corresponded with the text entities to obtain a multi-modal ancient Chinese poetry knowledge graph.

[0043] Due to the instability of the image generation results, although the images are generated based on text, the correspondence between the generated images and the text may not be strong. The present invention abandons using the CLIP score as an evaluation metric because some ancient Chinese poetry images are too abstract to find a suitable evaluation metric. The present invention uses multiple texts to match 10 images generated from the target text. If all ten images exceed the threshold and the percentage of exceeding also exceeds the threshold, the present invention considers this to be an appropriate text-image pair.

[0044] Combining the above-mentioned ancient Chinese poetry-related knowledge, author-related knowledge, dynasty information, image-text and image-image pairs will be used as nodes of the multi-modal ancient Chinese poetry knowledge graph, thus constructing a multi-modal ancient Chinese poetry knowledge graph.

Claims

1. A method for constructing a multimodal ancient poetry knowledge graph for converting ancient poetry and images, characterized in that; The following steps are involved: Step 1: Use Python to write a web crawler algorithm, take the website as the crawling target, crawl all the poems, lyrics, and songs from the pre-Qin period to the modern times included in the website, take the ancient poetry website as the target, crawl the corresponding explanations of ancient poems, take the Chinese Studies Collection website as the target, crawl the author information and poetry information, and crawl the dynasty related information from Baidu Encyclopedia; Step 2: Clean the crawled ancient poems, and then count the word frequencies of the cleaned ancient poems, and use the words that appear more than 5 times as text candidate words for the multimodal ancient poems knowledge graph; Step 3: Use Python to write a web crawler algorithm, take Han Dian as the crawling target, crawl the explanations of all candidate words, clean, sentence and format all the explanations to obtain the multi-tuple between ancient poems and explanations, and use the explanations of ancient poems to form the text entities of the multimodal ancient poems knowledge graph; Step 4: Calculate the similarity between different ancient poetry interpretations and establish connections between related entities; Step 5: Use the existing multimodal model to convert text into corresponding images and correspond them to obtain a multimodal knowledge graph.

2. The method for constructing a multimodal ancient poetry knowledge graph for converting ancient poetry and images according to claim 1 is characterized in that: The specific process of step 2 is: Step 2.1, remove all punctuation marks and useless special characters in ancient poems; Step 2.2: Continuously split the ancient poems according to their length, count the number of occurrences of ancient poems of different lengths, and select those that appear more than 5 times as candidate words.

3. The multimodal knowledge graph construction method for converting ancient poetry and images according to claim 1 is characterized in that: The specific process of step 3 is as follows: Step 3.1, use Python language to write a web crawler algorithm, take Han Dian as the crawling target, crawl all the explanations of candidate words, and count all the candidate words with existing explanations; Step 3.2, remove useless special characters in the explanation of ancient poems; Step 3. Make two copies of the cleaned ancient poetry multi-word group S and P, and delete the parts that supplement the explanation, such as the source and examples, by writing a Python algorithm to construct a knowledge graph G, in which the parts that appear in P are used as explanations of ancient poetry, and the parts that appear in S but not in P are used as supplements to ancient poetry.

4. The method for constructing a multimodal ancient poetry knowledge graph for converting ancient poetry and images according to claim 1 is characterized in that: The specific process of step 4 is as follows: Step 4.1: Use the word embedding algorithm to embed all entities into corresponding word vectors. By calculating the cosine similarity between different word vectors, "synonyms" are used as the unified relationship between entities. For entities whose similarity exceeds the threshold, the "entity-synonym-entity" ternary knowledge is established.

5. The method for constructing a multimodal ancient poetry knowledge graph for converting ancient poetry and images according to claim 1, characterized in that: The specific process of step 5 is as follows: Step 5.1: Use the multimodal model to map text to images. For entities with "synonym" relationships, only one entity needs to be selected to perform the multimodal model mapping from text to images. Step 5.2: The existing text-image calculation method calculates the CLIP score between the text-image pair as the final evaluation index. The present invention converts the regression problem into a classification problem by setting interference items to match the correct answer, thereby efficiently realizing image-text alignment.

6. The method for constructing a multimodal ancient poetry knowledge graph for converting ancient poetry and images according to claim 1, characterized in that: The specific process of step 6 is as follows: The comprehensive ancient poetry-related knowledge, author-related knowledge, dynasty information, image text and image image pairs will be used as nodes of the multimodal ancient poetry knowledge graph, thus constructing a multimodal ancient poetry knowledge graph.

Citation Information

Patent Citations

  • Knowledge graph construction method for sentiment analysis of ancient poetry

    CN118035469A

  • Mask segmentation map guided Chinese landscape painting generation model construction method

    CN118262195A

  • Formalization of a natural language

    US20120101803A1

Cited By

  • Document knowledge management method and system based on text retrieval enhancement generation

    CN120407749A

  • Image-text matching method, electronic equipment and storage medium

    CN121579719A