Multi-modal knowledge graph construction method and system based on image-text joint analysis
By employing a combined text-image parsing method, this study utilizes a web crawler system and the U-Net image segmentation model to process multimodal data, constructing a multimodal knowledge graph. This approach resolves the inconsistency in multimodal data representation, achieving efficient semantic understanding and unified representation, and is suitable for intelligent question answering and scientific research-assisted analysis.
Patent Information
- Application Number
- CN202511198394.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2026-01-27
AI Technical Summary
Existing multimodal data is represented inconsistently on different platforms, making it difficult to achieve accurate semantic alignment and efficient cross-analysis. Traditional database systems lack the ability to express complex relationships and cannot support unified representation and comprehensive reasoning of multimodal data.
By using a graph-text joint parsing method, multimodal data is collected using a web crawler system. The U-Net image segmentation model and OCR technology are combined to perform graph-text alignment, construct multimodal knowledge graph entity nodes, and generate a unified structure through graph structure fusion and cross-modal alignment techniques.
It enhances the structured processing and semantic understanding capabilities of multi-source heterogeneous data, and realizes unified representation and comprehensive reasoning of multimodal data, making it suitable for intelligent question answering and scientific research auxiliary analysis.
Smart Images

Figure CN121413718A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal information processing and knowledge graph construction, and more specifically, to a method and system for constructing a multimodal knowledge graph based on joint text-graph parsing. Background Technology
[0002] With the rapid development of information technology, a large amount of scientific research data is showing a trend of multimodal and heterogeneous origins. Especially in interdisciplinary fields such as biology, geology, and environment, research objects often involve multiple types of information, including text, images, and structured records. This data is usually scattered across multiple databases, academic platforms, literature, and image resources, resulting in significant information fragmentation and inconsistent expression methods, posing considerable challenges to researchers in information retrieval, integration, and analysis.
[0003] Taking data on certain protist phyla as an example, relevant information is widely available in open databases, academic encyclopedias, scientific literature, and image libraries, covering classification characteristics, geographical distribution, morphological descriptions, ecological evolutionary relationships, and correlations with environmental parameters. However, due to differences in naming conventions, storage structures, and update frequencies among various data sources, the representation of the same entity often varies across different platforms, making accurate semantic alignment and efficient cross-analysis difficult.
[0004] Currently, traditional data management methods mainly rely on relational databases, using structured storage to meet some query needs. However, when faced with semantically complex and modally diverse scientific research data, traditional tabular database systems often lack sufficient expressive power and struggle to support the modeling of complex relationships between entities, especially in expressing evolutionary paths, ecological associations, or contextual semantics.
[0005] In recent years, knowledge graphs have gradually become an effective tool for handling complex knowledge organization and relational analysis. By constructing a graph structure of entities and relationships through nodes and edges, they can integrate heterogeneous data, enhance semantic connectivity, and improve knowledge representation capabilities. However, most mainstream knowledge graph construction methods are currently based on text-based triples, which have limited coverage of modal information such as images, hindering the unified representation and comprehensive reasoning of multimodal data. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for constructing a multimodal knowledge graph based on graph-text joint parsing, which can improve the structured processing capability, semantic understanding capability and knowledge expression efficiency of multi-source heterogeneous data.
[0007] This invention provides a method for constructing a multimodal knowledge graph based on joint text-graph parsing, comprising the following steps: S1: Obtain literature data using a web crawler system based on heterogeneous data sources; process the literature data to obtain text data and image data; S2: Based on the text data and image data, perform multimodal knowledge completion to obtain multimodal data pairs; S3: Based on the multimodal data pairs, perform multimodal cleaning and completion to obtain knowledge graph entity nodes; S4: Based on the entity nodes of the knowledge graph, a multimodal knowledge graph structure is obtained by using graph structure fusion and cross-modal alignment methods.
[0008] This invention also provides a multimodal knowledge graph construction system based on graph-text joint parsing, the system comprising the following modules: The multi-source data acquisition module is configured to: obtain literature data from heterogeneous data sources using a web crawler system; and process the literature data to obtain text data and image data. The multimodal knowledge completion module is configured to: perform multimodal knowledge completion based on the text data and image data to obtain multimodal data pairs; The multimodal knowledge graph normalization module is configured to perform multimodal cleaning and completion based on the multimodal data pairs to obtain knowledge graph entity nodes; The multimodal knowledge graph structure module is configured to: obtain the multimodal knowledge graph structure based on the entity nodes of the knowledge graph using graph structure fusion and cross-modal alignment methods.
[0009] The method and system for constructing a multimodal knowledge graph based on joint text-graph parsing provided by this invention have the following beneficial effects: This invention addresses the challenges of acquiring, managing, and unifying the representation of multimodal data from complex sources with inconsistent representation methods. First, an optimized crawler module efficiently collects text and image data from multiple heterogeneous sources, improving the completeness and timeliness of raw data acquisition. Second, an image segmentation model based on the U-Net architecture extracts structural and semantic regions from images, and OCR technology is used to align and complement text and image content, effectively alleviating the inconsistency in text and image information representation. Building upon this, a constructed domain dictionary and rule base are used to complete multimodal entity recognition, attribute standardization, and semantic structure fusion. Finally, a unified multimodal knowledge graph is constructed based on entity, attribute, and relation triples, and a cross-modal alignment mechanism is introduced to optimize the graph structure and semantic consistency, achieving graph structure fusion and cross-modal alignment.
[0010] This invention can improve the structured processing capability, semantic understanding capability and knowledge expression efficiency of multi-source heterogeneous data, and provides an efficient and universal solution for the organization, management and intelligent application of scientific research corpora. It is applicable to various application scenarios such as intelligent question answering, bioinformatics modeling and scientific research auxiliary analysis. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of the multimodal knowledge graph construction method based on graph-text joint parsing provided by the present invention; Figure 2 This is a schematic diagram illustrating the implementation steps of the multimodal knowledge graph construction method based on graph-text joint parsing provided by the present invention; Figure 3 This invention provides a multimodal knowledge graph construction method based on graph-text joint parsing, and its multimodal knowledge graph construction model structure diagram. Figure 4 The multimodal knowledge graph construction method based on graph-text joint parsing provided by this invention relates to the data format of foraminifera-related websites in the geosciences field. Figure 5 This is an image segmentation model architecture diagram of the multimodal knowledge graph construction method based on graph-text joint parsing provided by the present invention; Figure 6 This is a schematic diagram of foraminifera image segmentation based on U-Net, illustrating the multimodal knowledge graph construction method based on joint text and image parsing provided by this invention. Detailed Implementation
[0012] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0013] Figure 1 This diagram illustrates the multimodal knowledge graph construction method based on graph-text joint parsing according to this embodiment. In this embodiment, the multimodal knowledge graph construction method based on graph-text joint parsing includes the following steps: S1: Obtain literature data using a web crawler system based on heterogeneous data sources; process the literature data to obtain text data and image data; In one exemplary embodiment, the heterogeneous data source includes public databases, encyclopedia platforms, and scientific research literature resources; As an exemplary embodiment, the web crawler system has the ability to focus on domain data crawling, dynamic IP switching and anti-crawling mechanism bypass function, access frequency control and stability guarantee, URL deduplication and link management mechanism, and supports the identification of naming rules for multi-source heterogeneous data and high-quality data output; In one exemplary embodiment, the web crawler system constructs a dynamic proxy IP pool. By setting access times, it periodically changes the IP address used to access websites to bypass anti-crawling mechanisms. At the same time, it sets random delays between requests to reduce the access frequency, thereby avoiding triggering IP restrictions. For each target website, a proxy IP is selected from the proxy pool, and web page content is crawled from each target website via an HTTP GET request. If the HTTP GET request is blocked, it switches to a new proxy IP and retryes the HTTP GET request until the target web page content, i.e., the literature data, is obtained. In one exemplary embodiment, the processing of the document data includes: converting PDF format documents in the document data into HTML format pages, and extracting text data and image data from the HTML format pages; S2: Based on the text data and image data, perform multimodal knowledge completion to obtain multimodal data pairs; In one exemplary embodiment, step S2 specifically includes: S21: The image data is segmented using an image segmentation model to obtain the target sub-image; In one exemplary embodiment, the image segmentation model is U-Net; As an exemplary embodiment, the image segmentation model includes an encoder, a decoder, and a skip connection module, wherein: the encoder is used to extract the contextual semantic features of the image, and uses multiple 3×3 convolutional layers and 2×2 max pooling layers to achieve image downsampling; the decoder uses 2×2 transposed convolution to upsample the feature map; the skip connection module connects the high-resolution features of each layer in the encoder with the corresponding layers in the decoder to enhance the accuracy of the image segmentation boundary; In one exemplary embodiment, the target sub-map is a single fossil sample; S22: Use optical character recognition (OCR) to identify the marking information of the target sub-image and extract the image number label; S23: Match the illustrated number label with the corresponding species information in the text data to obtain the associated image and text; S24: Based on the associated images and text, extract the species names using a dictionary matching method to obtain multimodal data pairs; As an exemplary embodiment, in step S2, each individual fossil is precisely cut out from the acquired fossil images and matched with its corresponding text; the input is the acquired text information and image data; the output is each individual fossil and its corresponding text matching information; specifically, an image segmentation model based on the U-Net architecture is applied to the acquired image data to extract structural regions and semantic content in the image, and the text annotation in the image is obtained by combining the optical character recognition (OCR) method to achieve image-text alignment and multimodal completion; specifically including the following steps: (1) For the image content, perform image segmentation operation to divide the image containing multiple fossil samples into a single labeled fossil image; In this embodiment, the fossil image extracted from the literature illustration is input into the U-Net image segmentation model for image segmentation to obtain a sub-image containing a single fossil sample; (2) The optical character recognition (OCR) method is used to identify the marking information around each sub-image and extract the image number label; specifically, the letters or numbers used for marking in each sub-image are located and identified; (3) Match the identified illustration number with the species information in the corresponding illustration text description to realize the association between the image and the text; that is, align the identified number with the corresponding text description in the HTML page structure to realize the association between the image and the text. (4) Extract species names using dictionary matching method to generate consistent multimodal data pairs containing image and text information; the dictionary matching method identifies and standardizes the species field in the text based on a species name dictionary; S3: Based on the multimodal data pairs, perform multimodal cleaning and completion to obtain knowledge graph entity nodes; As an exemplary embodiment, in step S3, a multimodal cleaning and completion module containing a domain-specific terminology dictionary and an entity recognition rule base is constructed to identify, standardize, resolve synonyms, and fuse structures of entity information in text and images. Specifically, the multimodal cleaning and completion module containing a domain-specific terminology dictionary and an entity recognition rule base can identify candidate entities from text and images, and perform attribute standardization, synonym resolution, and missing information completion on the identified entity information. This includes name unification, unit conversion, type labeling, and context- and semantic-based ambiguity elimination and synonym merging, ultimately outputting knowledge graph entity nodes with standardized structures and consistent semantics, providing a foundation for subsequent knowledge graph relationship construction. It should be noted that in step S3, the input is the original foraminifera multimodal data (including inconsistency identifiers and formats) from multiple sources after matching; the output is a foraminifera multimodal knowledge graph dataset that is standardized in terms of entities, attributes, and relationships and can be directly imported; the specific steps are as follows: construct a multimodal cleaning and completion module containing a terminology dictionary for a specific domain and an entity recognition rule base, and perform attribute standardization, synonym resolution, and missing information completion on the identified entity information, specifically including name unification, unit conversion, type labeling, and context- and semantic ambiguity elimination and synonym merging, and finally output knowledge graph entity nodes with standardized structure and consistent semantics; S4: Based on the entity nodes of the knowledge graph, a multimodal knowledge graph structure is obtained by using graph structure fusion and cross-modal alignment methods; As an exemplary embodiment, in step S4, based on the normalized entity types, attribute relationships and semantic boundaries, a multimodal knowledge graph that integrates text entities and image entities and their relationships is constructed, and a unified graph structure is generated through graph structure fusion and cross-modal alignment techniques; It should be noted that the input to step S4 is the normalized triplet extracted in step S3, and its output is a knowledge graph.
[0014] This embodiment provides a multimodal knowledge graph construction system based on graph-text joint parsing. The system includes the following modules: a multi-source data acquisition module, configured to: obtain literature data from heterogeneous data sources using a web crawler system; process the literature data to obtain text data and image data; a multimodal knowledge completion module, configured to: perform multimodal knowledge completion based on the text data and image data to obtain multimodal data pairs; a multimodal knowledge graph normalization module, configured to: perform multimodal cleaning and completion based on the multimodal data pairs to obtain knowledge graph entity nodes; and a multimodal knowledge graph structure module, configured to: obtain the multimodal knowledge graph structure based on the knowledge graph entity nodes using graph structure fusion and cross-modal alignment methods.
[0015] Specifically, the aforementioned heterogeneous data sources include public databases, encyclopedia platforms, and scientific research literature resources.
[0016] Specifically, the above-mentioned document data processing includes: converting PDF format documents in the document data into HTML format pages, and extracting text data and image data from the HTML format pages.
[0017] Specifically, the multimodal knowledge completion module is configured as follows: segmenting the image data using an image segmentation model to obtain a target sub-image; recognizing the marking information of the target sub-image using an optical character recognition method to extract the illustration number label; matching the illustration number label with the corresponding species information in the text data to obtain the associated image and text; and extracting the species name using a dictionary matching method based on the associated image and text to obtain a multimodal data pair.
[0018] Specifically, the image segmentation model mentioned above is U-Net.
[0019] In some embodiments, the above-described method for constructing a multimodal knowledge graph based on joint text-graph parsing can also be implemented in the following ways.
[0020] The flowchart of the multimodal knowledge graph construction method based on graph-text joint parsing in this embodiment is as follows: Figure 2 As shown, the structure diagram of the constructed multimodal knowledge graph model is as follows: Figure 3 As shown, the specific steps include: Step 1: Use the optimized web crawler system to collect text information and image data from multiple heterogeneous data sources, including public databases, encyclopedia platforms, and scientific literature resources. In this embodiment, in order to achieve multimodal data acquisition and structured organization of specific marine microorganisms (such as foraminifera), a web crawler with target focusing capabilities is optimized.
[0021] Heterogeneous data sources such as the World Register of Marine Species (WoRMS), Foraminifera.eu, and some open-source literature platforms are utilized. Figure 4 As shown, Figure 4 This embodiment uses a data format specific to the geosciences field related to foraminifera, which automatically captures and organizes multimodal information, including classification structure, morphological characteristics, image data, and species nomenclature.
[0022] To improve crawling efficiency and effectively circumvent anti-crawler mechanisms, the system constructs a dynamic proxy IP pool: , By periodically changing the IP address of the website accessed by setting access times, anti-scraping mechanisms can be bypassed. At the same time, random delays are set between requests to reduce the access frequency, thereby avoiding triggering IP restrictions. For each target website... Select a proxy IP from the proxy pool. :
[0023] Use the selected proxy IP From each target website via HTTP GET request crawl web page content :
[0024] If the request is blocked (e.g., with an HTTP status code of 403), the algorithm will rotate to a new proxy IP. And retry the request:
[0025]
[0026] Subsequently, the system performs structured information extraction on the target webpage. First, it identifies key entities and their attributes, distinguishing between triples used for graph construction and unstructured fields used only for retrieval assistance. Combining webpage layout parsing, it locates the position of each attribute within the HTML structure and processes them accordingly. Within each webpage, the system uses XPath or regular expressions to extract relevant data. :
[0027] Because the extraction results from different platforms differ in naming rules and hierarchical granularity, and unstructured data lacks a consistent information layout, it is difficult to directly extract text data and image information from such data sources. Therefore, a literature processing module is added before multimodal information completion. This module specifically includes the following steps: (1) Use PDF to HTML technology to convert the original PDF document into an HTML page in order to extract the page structure and layout information; (2) Extract image information and main text content from the HTML page; (3) Perform image segmentation operation on the image content to divide the image containing multiple fossil samples into single-labeled fossil images; (4) Use OCR recognition technology to identify the labeled numbers in the image; (5) Align the identified number with the corresponding text description in the HTML body to achieve image-text association annotation.
[0028] Step 2: Apply an image segmentation model based on the U-Net architecture to the collected image data to extract structural regions and semantic content in the image, and combine it with optical character recognition (OCR) to obtain text annotations in the image, thereby achieving image-text alignment and multimodal completion; For the collected multimodal image data, this embodiment employs an image segmentation model based on the U-Net architecture to accurately segment the images, solving problems such as multiple sample overlays and mixed annotations in fossil images from literature. U-Net extracts image contextual features through an encoder, restores spatial resolution through a decoder, and fuses high-resolution information using skip connections to achieve efficient segmentation of blurred target regions.
[0029] The U-Net image segmentation model includes an encoder, a decoder, and skip connection modules. The U-Net model architecture diagram is shown below. Figure 5 As shown, the encoder is used to extract the contextual semantic features of the image, and uses multiple 3×3 convolutional layers and 2×2 max pooling layers to achieve image downsampling; the decoder uses 2×2 transposed convolution to upsample the feature map; the skip connection module connects the high-resolution features of each layer in the encoder with the corresponding layers in the decoder to enhance the accuracy of image segmentation boundaries.
[0030] Subsequently, OCR technology is used to identify the marker numbers in each segmented sub-image, achieving precise alignment between the image and the corresponding text annotations, such as... Figure 6 As shown, Figure 6 This is an exemplary embodiment of the present invention, specifically a U-Net-based foraminifera image segmentation. Species information is extracted from the text using a dictionary matching method. This algorithm identifies fossil image labels and corresponding species information from OCR-extracted image caption text using regular expressions. First, a regular expression pattern is defined that matches a numerical identifier followed by a species name consisting of two words. Then, this pattern is used to search for all matches in the text. For each match, the label and species name are extracted and stored as key-value pairs in a dictionary. Finally, a mapping dictionary containing all fossil labels and their corresponding species names is returned, realizing the correspondence between image tags and text descriptions. This ultimately achieves multimodal completion of image and text information, providing support for the subsequent construction of a structured multimodal knowledge graph.
[0031] Step 3: Construct a multimodal cleaning and completion module that includes a domain-specific terminology dictionary and an entity recognition rule base to identify, standardize, resolve synonyms, and fuse structures in text and images; In this embodiment, the multimodal cleaning and completion module further includes three key sub-processes: entity normalization, attribute standardization, and relationship definition, to support unified modeling of heterogeneous source data in the knowledge graph.
[0032] First, the system performs standardized identification and encoding of entities extracted from images and text based on a pre-built domain terminology dictionary. Entity categories, including but not limited to species, organs, geological units, and geographical regions, can be normalized and labeled using authoritative external databases (such as scientific taxonomy catalogs, geographic information systems, or professional knowledge bases). For example, classification information is broken down and organized into entities by identifying hierarchical structures such as phylum, class, and species, while geographical information is standardized by region, depth, or coordinate range. For locally unique identifiers (such as certain numeric IDs) in some databases that are incompatible with external knowledge systems, unified mapping or discarding is used to ensure consistency and robustness of multi-source data alignment.
[0033] Secondly, during attribute normalization, the system converts entity feature values (such as environmental parameters, distribution range, morphological features, etc.) into a standard format according to predefined attribute mapping rules. The units, terminology, and naming conventions used follow industry-standard practices; for example, temperature is expressed in degrees Celsius, depth in meters, and names use Latin scientific names or industry-standard terms. Boolean categories are labeled using standard Boolean values. This process significantly improves the comparability of attributes in the map and enhances the system's automated processing capabilities.
[0034] Finally, the relation normalization module performs structural abstraction and standardized modeling of semantic relationships between entities. The system pre-sets basic relation templates such as "belongs to", "distributed in", "lived in", and "is a", and automatically identifies the connection methods between entities and forms reasonable knowledge links by combining contextual information. All relations are bound to standard semantic tags, supporting logical calculations in subsequent graph reasoning, question answering systems, or decision support.
[0035] In typical application scenarios, this module is applicable to data processing tasks in multiple fields, including paleontology, biodiversity, and ecological surveys. For example, when processing images of a certain type of microfossil and related literature, this module can construct a structured knowledge graph containing multi-dimensional information such as taxonomic hierarchy, paleogeographic distribution, and environmental conditions, enabling unified representation and integrated association of entities from different sources.
[0036] Step 4: Based on the normalized entity types, attribute relationships, and semantic boundaries, construct a multimodal knowledge graph that integrates text entities and image entities and their relationships, and generate a unified graph structure through graph structure fusion and cross-modal alignment techniques.
[0037] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for constructing a multimodal knowledge graph based on joint text-graph parsing, characterized in that, Includes the following steps: S1: Obtain literature data using a web crawler system based on heterogeneous data sources; process the literature data to obtain text data and image data; S2: Based on the text data and image data, perform multimodal knowledge completion to obtain multimodal data pairs; S3: Based on the multimodal data pairs, perform multimodal cleaning and completion to obtain knowledge graph entity nodes; S4: Based on the entity nodes of the knowledge graph, a multimodal knowledge graph structure is obtained by using graph structure fusion and cross-modal alignment methods.
2. The method for constructing a multimodal knowledge graph based on joint text-graph parsing according to claim 1, characterized in that, The heterogeneous data sources include public databases, encyclopedia platforms, and scientific research literature resources.
3. The method for constructing a multimodal knowledge graph based on joint text-graph parsing according to claim 1, characterized in that, The processing of the document data includes: converting PDF documents in the document data into HTML format pages, and extracting text data and image data from the HTML format pages.
4. The method for constructing a multimodal knowledge graph based on joint text-graph parsing according to claim 1, characterized in that, Step S2 specifically includes: S21: The image data is segmented using an image segmentation model to obtain the target sub-image; S22: Use optical character recognition (OCR) to identify the marking information of the target sub-image and extract the image number label; S23: Match the illustrated number label with the corresponding species information in the text data to obtain the associated image and text; S24: Based on the associated images and text, extract the species names using a dictionary matching method to obtain multimodal data pairs.
5. The method for constructing a multimodal knowledge graph based on joint text-graph parsing according to claim 4, characterized in that, The image segmentation model is U-Net.
6. A multimodal knowledge graph construction system based on graph-text joint parsing, characterized in that, The system includes the following modules: The multi-source data acquisition module is configured to: obtain literature data from heterogeneous data sources using a web crawler system; and process the literature data to obtain text data and image data. The multimodal knowledge completion module is configured to: perform multimodal knowledge completion based on the text data and image data to obtain multimodal data pairs; The multimodal knowledge graph normalization module is configured to perform multimodal cleaning and completion based on the multimodal data pairs to obtain knowledge graph entity nodes; The multimodal knowledge graph structure module is configured to: obtain the multimodal knowledge graph structure based on the entity nodes of the knowledge graph using graph structure fusion and cross-modal alignment methods.
7. The multimodal knowledge graph construction system based on graph-text joint parsing according to claim 6, characterized in that, The heterogeneous data sources include public databases, encyclopedia platforms, and scientific research literature resources.
8. The multimodal knowledge graph construction system based on graph-text joint parsing according to claim 6, characterized in that, The processing of the document data includes: converting PDF documents in the document data into HTML format pages, and extracting text data and image data from the HTML format pages.
9. The multimodal knowledge graph construction system based on graph-text joint parsing according to claim 6, characterized in that, The specific configuration of the multimodal knowledge completion module is as follows: The image data is segmented using an image segmentation model to obtain the target sub-image; The marking information of the target sub-image is identified using optical character recognition methods, and the image number is extracted and labeled. The illustrated number is matched with the corresponding species information in the text data to obtain the associated image and text; Based on the associated images and text, the species names are extracted using a dictionary matching method to obtain multimodal data pairs.
10. The multimodal knowledge graph construction system based on graph-text joint parsing according to claim 9, characterized in that, The image segmentation model is U-Net.