Cross-modal data retrieval method, system and equipment based on multi-modal knowledge graph

By preprocessing and decoupling features of text and image modal data, a path-constrained contrastive learning loss function and a multi-scale path-aware rejection loss function are constructed to optimize the path encoding features of entities in multimodal knowledge graphs. This resolves the contradiction between cross-modal semantic alignment and graph structure preservation, and improves the accuracy of text and image retrieval.

CN120804352AActive Publication Date: 2025-10-17UNICOM WOYUEDU TECH CULTURE CO LTD

Patent Information

Application Number
CN202511324517.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-10-17
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing multimodal knowledge graph construction technologies suffer from alignment-structure mutual exclusion bottlenecks in cross-modal comparative learning, leading to a decrease in the accuracy of text and image retrieval.

Method used

By preprocessing text and image modal data, extracting visual and structural feature vectors, decoupling features, and constructing path constraint contrastive learning loss function and multi-scale path-aware rejection loss function, the path encoding features of entities in multimodal knowledge graphs are optimized to achieve cross-modal data retrieval.

Benefits of technology

It improves the accuracy of image and text retrieval, ensures the semantic fidelity of text and images, enhances the semantic alignment and path semantic distinguishability of multimodal data, and resolves the contradiction between cross-modal semantic alignment and graph structure preservation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804352A_ABST
    Figure CN120804352A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal data retrieval method, system and equipment based on a multi-modal knowledge graph, and the method comprises the steps: carrying out the feature decoupling based on a visual feature vector and a structural feature vector, and determining a target visual projection vector; constructing a path constraint contrast learning loss function based on the target visual projection vector and the text feature vector; constructing a multi-modal knowledge graph, and mining effective paths among entities in the multi-modal knowledge graph; coding the effective path into a path coding feature, and constructing a multi-scale path perception rejection loss function based on the path coding feature; joint optimization is carried out on the path constraint contrast learning loss function and the multi-scale path perception rejection loss function, and a target path coding feature subset is determined; fusing the text feature vector, the target visual projection vector and the target path coding feature subset to obtain a joint vector; and performing cross-modal data retrieval based on the joint vector to obtain a retrieval result. The image-text retrieval method and device can improve the accuracy of image-text retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cross-modal retrieval, in particular to a cross-modal data retrieval method, system and device based on a multi-modal knowledge graph. BACKGROUND

[0002] Multi-modal knowledge graph vectorization is a core topic in the field of artificial intelligence and knowledge graph technology, and its goal is to model heterogeneous modal data (such as text, images, and videos, etc.) and their complex relationships in a unified vector space, supporting cross-modal semantic understanding and intelligent reasoning. Current multi-modal knowledge graph construction techniques mainly revolve around two paradigms: independent embedding followed by fusion (such as the MCEAS model), which uses dedicated encoders (such as BERT for text and ResNet for images) to generate single-modal vectors, and then fuses them into entity vectors through an attention mechanism. Graph structure expansion methods (such as the M3-Agent framework), which directly aggregate multi-modal neighborhood information based on graph neural networks (such as GAT / R-GCN), preserving local graph structure.

[0003] Although the above two methods have made progress in heterogeneous data fusion, there is still a bottleneck of alignment-structure incompatibility, that is, when cross-modal contrast learning forces text / image vectors to be close, it will dilute key topological relationships, and graph structure injection will pollute visual features, resulting in a decrease in the accuracy of image-text retrieval. SUMMARY

[0004] The present application aims to provide a cross-modal data retrieval method, system and device based on a multi-modal knowledge graph, which can improve the accuracy of image-text retrieval.

[0005] In a first aspect, the present application provides a cross-modal data retrieval method based on a multi-modal knowledge graph, comprising: preprocessing the obtained text modal data and image modal data to obtain standardized text modal data and standardized image modal data; extracting a text feature vector from the standardized text modal data, and extracting a visual feature vector and a structure feature vector from the standardized image modal data; decoupling features based on the visual feature vector and the structure feature vector to determine a target visual projection vector; constructing a path constraint contrastive learning loss function based on the target visual projection vector and the text feature vector; constructing a multi-modal knowledge graph between text and image, and mining effective paths between entities in the multi-modal knowledge graph; encoding the effective paths into path encoding features, and constructing a multi-scale path perception repulsion loss function based on the path encoding features; jointly optimize the path constraint contrastive learning loss function and the multi-scale path-aware repulsion loss function to determine a target path encoding feature subset; fuse the text feature vector, the target visual projection vector, and the target path encoding feature subset to obtain a joint vector; perform cross-modal data retrieval based on the joint vector to obtain a retrieval result.

[0006] Compared with the prior art, the first aspect of the present application has the following beneficial effects: The method pre-processes the obtained text modal data and image modal data to obtain standardized text modal data and standardized image modal data; extracts a text feature vector from the standardized text modal data, and extracts a visual feature vector and a structural feature vector from the standardized image modal data; performs feature decoupling based on the visual feature vector and the structural feature vector to determine a target visual projection vector; constructs a path constraint contrastive learning loss function based on the target visual projection vector and the text feature vector; constructs a multi-modal knowledge graph between text and image, and mines effective paths between entities in the multi-modal knowledge graph; encodes the effective paths into path encoding features, and constructs a multi-scale path-aware repulsion loss function based on the path encoding features; jointly optimizes the path constraint contrastive learning loss function and the multi-scale path-aware repulsion loss function to determine a target path encoding feature subset; fuses the text feature vector, the target visual projection vector, and the target path encoding feature subset to obtain a joint vector; and performs cross-modal data retrieval based on the joint vector to obtain a retrieval result. In this way, by feature decoupling, the semantic fidelity of text and image can be ensured; by path constraint contrastive learning loss, the similarity distribution of multi-modal features is optimized to align the semantics of multi-modal data; by multi-scale path-aware repulsion loss, the path feature distance of positive sample entity pairs with semantic correlation is forced to be smaller than the path feature distance of positive sample and negative sample pairs, thereby strengthening the discriminability of path semantics; by jointly optimizing the path constraint contrastive learning loss function and the multi-scale path-aware repulsion loss function, the optimal path encoding feature subset can be determined, and thus fusing the text feature vector, the target visual projection vector, and the target path encoding feature subset, and performing cross-modal data retrieval based on the joint vector, can solve the inherent contradiction between cross-modal semantic alignment and graph structure preservation in the multi-modal knowledge graph, and can improve the accuracy of image-text retrieval.

[0007] In some embodiments, the feature decoupling based on the visual feature vector and the structural feature vector to determine a target visual projection vector comprises: mapping the visual feature vector and the structural feature vector to an orthogonal subspace by an orthogonal projection operation to obtain an initial visual projection vector and an initial structural projection vector; constructing a mutual information constraint loss function for feature decoupling, and determining whether the initial visual projection vector and the initial structure projection vector satisfy the mutual information constraint loss function; If the initial visual projection vector and the initial structure projection vector do not satisfy the mutual information constraint loss function, updating the gradient of the mutual information constraint loss function through back propagation to optimize the visual feature vector and the structure feature vector extracted in the forward propagation process; Based on the optimized visual feature vector and the structure feature vector, an orthogonal projection operation is used to obtain an optimized initial visual projection vector and an optimized initial structure projection vector, and after satisfying the mutual information constraint loss function, the optimized initial visual projection vector is taken as a target visual projection vector.

[0008] In some embodiments, the path constraint contrastive learning loss function is constructed based on the target visual projection vector and the text feature vector, including: ; Wherein, represents the path constraint contrastive learning loss function, represents a set of positive sample pairs satisfying the path constraint, represents a text feature vector, represents a target visual projection vector, represents a cosine similarity function of a feature pair, represents a temperature coefficient, represents the number of negative samples, represents a text entity, represents an image entity.

[0009] In some embodiments, the effective path between entities in the multi-modal knowledge graph is mined, including: a pre-set semantic constraint condition; adopting depth-first search and breadth-first search to mine potential connection paths between entities in the multi-modal knowledge graph; mining effective paths between entities satisfying the semantic constraint condition from the potential connection paths.

[0010] In some embodiments, the multi-scale path-aware repulsion loss function is constructed based on the path encoding feature, including: ; Wherein, represents a multi-scale path-aware repulsion loss function, represents a sampling weight of a path with different hops, represents an anchor entity after path encoding features after jumping the relation path, representing a positive sample entity after path encoding features after jumping the relation path, representing a negative sample entity after path encoding features after jumping the relation path, representing a boundary margin.

[0011] In some embodiments, the fusing the text feature vector, the target visual projection vector and the target path encoding feature subset to obtain a joint vector comprises: respectively performing layer normalization on the text feature vector, the target visual projection vector and the target path encoding feature subset to obtain a layer-normalized text feature vector, a layer-normalized target visual projection vector and a layer-normalized target path encoding feature subset; splicing the layer-normalized text feature vector, the layer-normalized target visual projection vector and the layer-normalized target path encoding feature subset to obtain a joint vector.

[0012] In some embodiments, the performing cross-modal data retrieval based on the joint vector to obtain a retrieval result comprises: storing the joint vector generated by each piece of data into a vector database having a same space vector representation, and establishing an approximate nearest neighbor index of the vector database; encoding the obtained to-be-retrieved modal data into a query vector having a same space vector representation; performing cross-modal data retrieval on the query vector in the vector database having the approximate nearest neighbor index by using an approximate nearest neighbor search to obtain a retrieval result.

[0013] In a second aspect, the embodiments of the present application further provide a cross-modal data retrieval system based on a multi-modal knowledge graph, and the system comprises: a data processing unit configured to pre-process obtained text modal data and image modal data to obtain standardized text modal data and standardized image modal data; a feature extraction unit configured to extract a text feature vector from the standardized text modal data, and extract a visual feature vector and a structural feature vector from the standardized image modal data; a feature decoupling unit configured to perform feature decoupling based on the visual feature vector and the structural feature vector to determine a target visual projection vector; a first construction unit configured to construct a path constraint contrastive learning loss function based on the target visual projection vector and the text feature vector; a path mining unit configured to construct a multi-modal knowledge graph between texts and images, and mine effective paths between entities in the multi-modal knowledge graph; a second construction unit configured to encode the effective paths into path encoding features, and construct a multi-scale path-aware repulsion loss function based on the path encoding features; a joint optimization unit configured to jointly optimize the path-constrained contrastive learning loss function and the multi-scale path-aware repulsion loss function, and determine a target path encoding feature subset; a feature fusion unit configured to fuse the text feature vector, the target visual projection vector, and the target path encoding feature subset to obtain a joint vector; a data retrieval unit configured to perform cross-modal data retrieval based on the joint vector to obtain a retrieval result.

[0014] In a third aspect, an electronic device is provided, including at least one control processor and a memory in communication connection with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to perform the cross-modal data retrieval method based on a multi-modal knowledge graph as described above.

[0015] In a fourth aspect, a computer-readable storage medium is provided, which stores computer-executable instructions for causing a computer to perform the cross-modal data retrieval method based on a multi-modal knowledge graph as described above.

[0016] It can be understood that the beneficial effects of the second aspect to the fourth aspect compared with the related art are the same as the beneficial effects of the first aspect compared with the related art, and reference can be made to the related description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings, in which: Figure 1 is a flowchart of an embodiment of the cross-modal data retrieval method based on a multi-modal knowledge graph provided by the present application; Figure 2 is a flowchart of the overall process in the best embodiment of the cross-modal data retrieval method based on a multi-modal knowledge graph provided by the present application; Figure 3 is a structural diagram of an embodiment of the cross-modal data retrieval system based on a multi-modal knowledge graph provided by the present application; Figure 4 is a structural schematic diagram of an embodiment of an electronic device provided by the present application. DETAILED DESCRIPTION

[0018] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the accompanying drawings are exemplary and are only used to explain the present application and cannot be understood as a limitation of the present application.

[0019] In the description of the present application, if the first, second, etc. are described, it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features or the order of the indicated technical features.

[0020] In the description of the present application, it should be understood that the orientation description, such as the orientation or position relationship indicated by up, down, etc. is based on the orientation or position relationship shown in the drawings, only for the purpose of facilitating the description of the present application and simplifying the description, and is not intended to indicate or imply that the device or element indicated must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.

[0021] In the description of the present application, it should be noted that, unless otherwise explicitly limited, the words such as setting, installing, connecting, etc. should be broadly understood, and those skilled in the art can reasonably determine the specific meaning of the above words in the present application in combination with the specific content of the technical solution.

[0022] Although the existing method has made progress in heterogeneous data fusion, there is still a bottleneck of alignment-structure exclusivity, that is, when cross-modal contrast learning forces text / image vectors to approach, it will dilute the key topological relationship, and graph structure injection will pollute the visual features, resulting in a decrease in the accuracy of image-text retrieval.

[0023] To solve the problem of low accuracy of image-text retrieval in the prior art, the present application provides a cross-modal data retrieval method, system and device based on a multi-modal knowledge graph.

[0024] Reference Figure 1 The flowchart of the cross-modal data retrieval method based on the multi-modal knowledge graph provided by the embodiments of the present application is shown. The cross-modal data retrieval method based on the multi-modal knowledge graph is applied to an electronic device, which can be a server or a mobile terminal, etc. As shown in Figure 1 The cross-modal data retrieval method based on the multi-modal knowledge graph can include the following steps: Step S101, pre-processing the obtained text modal data and image modal data to obtain standardized text modal data and standardized image modal data; Step S102, extracting a text feature vector from the canonical text modality data, and extracting a visual feature vector and a structure feature vector from the canonical image modality data; Step S103, performing feature decoupling based on the visual feature vector and the structure feature vector to determine a target visual projection vector; Step S104, constructing a path constraint contrastive learning loss function based on the target visual projection vector and the text feature vector; Step S105, constructing a multi-modal knowledge graph between the text and the image, and mining effective paths between entities in the multi-modal knowledge graph; Step S106, encoding the effective paths into path encoding features, and constructing a multi-scale path perception repulsion loss function based on the path encoding features; Step S107, jointly optimizing the path constraint contrastive learning loss function and the multi-scale path perception repulsion loss function to determine a target path encoding feature subset; Step S108, fusing the text feature vector, the target visual projection vector, and the target path encoding feature subset to obtain a joint vector; Step S109, performing cross-modal data retrieval based on the joint vector to obtain a retrieval result.

[0025] In the embodiment, the obtained text modal data and image modal data are preprocessed to obtain standardized text modal data and standardized image modal data; a text feature vector is extracted from the standardized text modal data, and a visual feature vector and a structure feature vector are extracted from the standardized image modal data; feature decoupling is performed based on the visual feature vector and the structure feature vector to determine a target visual projection vector; a path constraint contrast learning loss function is constructed based on the target visual projection vector and the text feature vector; a multi-modal knowledge graph between text and image is constructed, and an effective path between entities in the multi-modal knowledge graph is mined; the effective path is encoded into a path encoding feature, and a multi-scale path perception repulsion loss function is constructed based on the path encoding feature; the path constraint contrast learning loss function and the multi-scale path perception repulsion loss function are jointly optimized to determine a target path encoding feature subset; the text feature vector, the target visual projection vector, and the target path encoding feature subset are fused to obtain a joint vector; cross-modal data retrieval is performed based on the joint vector to obtain a retrieval result. In this way, through feature decoupling, the semantic fidelity of text and image can be ensured; through the path constraint contrast learning loss, the similarity distribution of multi-modal features is optimized to enable semantic alignment of multi-modal data; through the multi-scale path perception repulsion loss, the path feature distance of a positive sample entity pair with semantic correlation is forced to be smaller than the path feature distance of a positive sample and a negative sample pair, thereby strengthening the discriminability of path semantics; through joint optimization of the path constraint contrast learning loss function and the multi-scale path perception repulsion loss function, an optimal path encoding feature subset can be determined, and thus the text feature vector, the target visual projection vector, and the target path encoding feature subset are fused, and cross-modal data retrieval is performed based on the joint vector, which can solve the inherent contradiction between cross-modal semantic alignment and graph structure preservation in the multi-modal knowledge graph, and can improve the accuracy of image-text retrieval.

[0026] The above extracting a text feature vector from the standardized text modal data, and extracting a visual feature vector and a structure feature vector from the standardized image modal data can be extracting the text feature vector from the standardized text modal data using a BERT_base model, extracting the visual feature vector from the standardized image modal data using a ResNet-50 model, and extracting the structure feature vector from the standardized image modal data using a graph neural network. It should be noted that the BERT_base model, the ResNet-50 model, and the graph neural network in the embodiment are all models known to those skilled in the art, and the embodiment will not be described in detail.

[0027] The above encoding the effective path as a path encoding feature can be mapping each relationship in the effective path to an embedding vector, fusing the forward and backward sequence semantics of the path (such as the logical order of "place of birth → dynasty to which it belongs") through a bidirectional long short-term memory network (BiLSTM), and generating a unified vector (i.e., the path encoding feature) capable of representing the semantics of the path.

[0028] The above joint optimization of the path constraint contrast learning loss function and the multi-scale path-aware repulsion loss function can be joint optimization of the path constraint contrast learning loss function and the multi-scale path-aware repulsion loss function by using a projected gradient descent and back propagation.

[0029] In some embodiments, based on the visual feature vector and the structural feature vector, feature decoupling is performed to determine a target visual projection vector, including: The visual feature vector and the structural feature vector are mapped to an orthogonal subspace by using an orthogonal projection operation to obtain an initial visual projection vector and an initial structural projection vector; An mutual information constraint loss function for feature decoupling is constructed, and it is determined whether the initial visual projection vector and the initial structural projection vector satisfy the mutual information constraint loss function; If the initial visual projection vector and the initial structural projection vector do not satisfy the mutual information constraint loss function, the gradient of the mutual information constraint loss function is updated by back propagation to optimize the visual feature vector and the structural feature vector extracted in the forward propagation process; Based on the optimized visual feature vector and the structural feature vector, the orthogonal projection operation is used to obtain an optimized initial visual projection vector and an optimized initial structural projection vector, and until the mutual information constraint loss function is satisfied, the optimized initial visual projection vector is taken as the target visual projection vector.

[0030] In the present embodiment, by constructing the mutual information constraint loss function for feature decoupling, effective decoupling of the structural feature and the visual feature in the orthogonal space can be achieved, and by eliminating the coupling between the structural feature and the visual feature, the semantic fidelity of the text and the image can be guaranteed.

[0031] The above orthogonal projection operation can refer to the projection of the image space U and the null space W to each other orthogonal subspace. It should be noted that the orthogonal projection operation in the present embodiment is an operation known to those skilled in the art, and the present embodiment will not be described in detail.

[0032] In some embodiments, based on the target visual projection vector and the text feature vector, a path constraint contrast learning loss function is constructed, including: ; Wherein, represents the path constraint contrast learning loss function, a set of positive sample pairs satisfying the path constraint, a text feature vector, a target visual projection vector, a cosine similarity function of the feature pairs, a temperature coefficient, a negative sample number, a text entity, an image entity.

[0033] In the embodiment, the similarity distribution of the multi-modal features is optimized by the constructed path constraint contrast learning loss function, so as to perform semantic alignment on the multi-modal data.

[0034] In some embodiments, the effective path between entities in the multi-modal knowledge graph is mined, including: a preset semantic constraint condition; adopting deep-first search and breadth-first search to mine the potential connection path between entities in the multi-modal knowledge graph; mining the effective path between entities satisfying the semantic constraint condition from the potential connection path.

[0035] In the embodiment, the effective path between entities satisfying the semantic constraint condition is mined from the potential connection path, which can guarantee semantic correlation, filter the relationships irrelevant to the target field, and exclude some invalid entity type chains, thereby screening the effective path and laying a good data foundation for improving data retrieval in the later stage.

[0036] In some embodiments, a multi-scale path perception repulsion loss function is constructed based on path encoding features, including: ; wherein, the multi-scale path perception repulsion loss function, a sampling weight of a path with different hop numbers, an anchor entity the path encoding feature after hop relationship path, the path encoding feature after hop relationship path, the path encoding feature after hop relationship path, the path encoding feature after hop relationship path, a boundary margin.

[0037] In the embodiment, the multi-scale path perception rejection loss is used to force the path feature distance of the semantic associated positive sample entity pair to be smaller than the path feature distance of the positive sample and the negative sample pair, so as to enhance the distinguishability of the path semantics.

[0038] In some embodiments, the text feature vector, the target visual projection vector and the target path encoding feature subset are fused to obtain a joint vector, including: The text feature vector, the target visual projection vector and the target path encoding feature subset are respectively layer normalized to obtain a layer normalized text feature vector, a layer normalized target visual projection vector and a layer normalized target path encoding feature subset; The layer normalized text feature vector, the layer normalized target visual projection vector and the layer normalized target path encoding feature subset are spliced to obtain a joint vector.

[0039] In the embodiment, the layer normalized text feature vector, the layer normalized target visual projection vector and the layer normalized target path encoding feature subset are spliced to obtain a joint vector, and all modal data (text, image) are mapped to vector representations in the same space, thereby realizing deep semantic alignment and laying a good data foundation for improving data retrieval later.

[0040] In some embodiments, cross-modal data retrieval is performed based on the joint vector to obtain a retrieval result, including: The joint vector generated by each piece of data is stored in a vector database with the same space vector representation, and an approximate nearest neighbor index of the vector database is established; The obtained to-be-retrieved modal data is encoded into a query vector with the same space vector representation; In the vector database with the approximate nearest neighbor index, an approximate nearest neighbor search is used to perform cross-modal data retrieval on the query vector to obtain a retrieval result.

[0041] In the embodiment, the joint vector generated by each piece of data is stored in a vector database with the same space vector representation, and in the vector database with the approximate nearest neighbor index, an approximate nearest neighbor search is used to perform cross-modal data retrieval on the query vector, thereby improving the accuracy of image-text retrieval.

[0042] For the convenience of those skilled in the art, a set of best embodiments is provided as follows: Multi-modal knowledge graph vectorization is a core issue in the intersection of artificial intelligence and knowledge graph technology, which aims to model heterogeneous modal data (e.g., text, image, and video) and their complex relationships in a unified vector space, supporting cross-modal semantic understanding and intelligent reasoning. Current multi-modal knowledge graph construction techniques mainly revolve around two paradigms: independent embedding followed by fusion (such as the MCEAS model), which uses dedicated encoders (such as BERT for text and ResNet for images) to generate single-modal vectors, and then fuses them into entity vectors through an attention mechanism. Graph structure expansion methods (such as the M3-Agent framework), which directly aggregate multi-modal neighborhood information based on graph neural networks (such as GAT / R-GCN), preserving local graph structure.

[0043] Although the above two methods have made progress in heterogeneous data fusion, there is still a bottleneck of alignment-structure incompatibility, that is, cross-modal contrastive learning forces text / image vectors to be close, which dilutes key topological relationships, while graph structure injection pollutes visual features. And cross-modal alignment needs to minimize the distance between text and image vectors, while structure preservation needs to maximize the distance between entity vectors on different paths, which conflicts in the unified space. Therefore, existing methods can lead to a decrease in the accuracy of graph-text retrieval.

[0044] To this end, in view of the core contradiction between cross-modal semantic alignment destroying graph structure and graph structure injection distorting visual features in multi-modal knowledge graph construction, the embodiment aims to design a topologically aware multi-modal dynamic joint encoding framework, which realizes accurate cross-modal semantic understanding and improves the accuracy of graph-text retrieval through a path constraint alignment mechanism with spatial decoupling and an intent-aware feature reorganization technique. Referring to Figure 2 The method of the embodiment specifically includes the following contents: The system includes two main steps: multi-modal data preprocessing and feature decoupling, and dual-space joint modeling.

[0045] 1. Multi-modal data preprocessing and feature decoupling.

[0046] This step is characterized by decoupling structural features and visual features, ensuring the semantic fidelity of text and image, and laying a good data foundation for subsequent dual-space modeling.

[0047] (1) Multi-modal data normalization processing: First, data preprocessing is performed to remove HTML / XML tags and functionally irrelevant noise characters (such as random code symbols), but special symbols with semantic value (such as #, & and * special symbols to avoid destroying entity structure, for example, "AT&T" or "C#") are retained; Pointer Network is used to extract text core information paragraphs (i.e. text modal data), and TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is introduced to realize the refined extraction of key semantics, and standardized text modal data is obtained. The image (i.e. image modal data) is scaled to 224x224 using bilinear interpolation, and the standardized image modal data is obtained.

[0048] (2) Feature independent extraction: BERT_base model is used to extract text feature vector from standardized text modal data. Image feature encoding is realized based on ResNet-50 model, which first passes the input preprocessed image (i.e. standardized image modal data) through layer1 to layer4 in ResNet-50 model for forward propagation, and finally outputs a high-dimensional feature map from the last convolutional layer. Subsequently, global average pooling operation is performed on the feature map, thereby aggregating the two-dimensional feature matrix of each channel into a single scalar value, and finally generating a fixed-dimensional vector representation . To balance the general feature extraction capability and domain adaptability of the model, a hierarchical parameter freezing strategy is adopted: the pre-training parameters of lower level layer1 to layer3 in ResNet-50 are fixed to retain the ability to extract general low-level features (such as texture and edge); at the same time, the parameters of higher level layer4 are allowed to be fine-tuned during training to learn high-level semantic features highly related to the specific domain, thereby enhancing the adaptability of the model to the target task; then the topological association information of entities is obtained through knowledge extraction, and graph neural network (such as GAT or R-GCN) is used for encoding, which first maps the original node information and associated relationships of entities into initial vectors, and then aggregates the structural semantic information of 1-2 hop neighborhood of entities to generate structure feature vectors that can represent the topological properties of entities.

[0049] (3) Feature orthogonalization decoupling: The original structure feature vector and the visual feature vector are respectively mapped to orthogonal subspaces through orthogonal projection operation to obtain projection vectors (initial structure projection vector) and (initial visual projection vector), so that they satisfy the orthogonality constraint (i.e. mutual information constraint loss) in spatial dimension. Mutual information constraint loss function to control the statistical correlation between the features: ; where, is the loss weight coefficient, denotes and the Jensen-Shannon divergence for measuring the difference between two probability distributions. The calculation of the Jensen-Shannon divergence is based on the Kullback-Leibler divergence, so it can be defined as: ; where, and are the probability distributions of and respectively, is the average distribution of the two, is the KL divergence. This loss function ensures that the orthogonal feature mutual information after projection is lower than the threshold of 0.25 by penalizing the case where the JSD divergence is greater than the threshold, ultimately achieving effective decoupling of the structural features and visual features in the orthogonal subspace, obtaining and (i.e., the target visual projection vector).

[0050] 2. Joint modeling in dual spaces.

[0051] This step aims to solve the inherent contradiction between cross-modal semantic alignment and graph structure preservation in multi-modal knowledge graphs.

[0052] (1) Aligning subspace modeling: With the relationship paths between entities in the knowledge graph (KG) as constraints, filter the positive sample pairs with semantic correlation, and enhance the feature discriminability through an adaptive temperature coefficient. Construct the path-constrained contrastive learning loss to optimize the similarity distribution of multi-modal features: ; where, is the cosine similarity function of the feature pair; is the temperature coefficient for controlling the sharpness of the similarity distribution; is the number of negative samples; is the set of positive sample pairs that meet the path constraints, defined by the relationship path constraints of entities in the knowledge graph: ; that is, if and only if the text entity and the image entity exist in the knowledge graph with a relationship path of no more than 5 hops, The positive sample pairs are determined to ensure the semantic relevance of the positive sample pairs. To enhance the discrimination of features in contrast learning, a temperature coefficient is designed With the similarity value , the strategy is dynamically adjusted, and the formula is: ; When the similarity variance increases (the feature discrimination decreases), it is decreased, thereby enhancing the discrimination of features in contrast learning; on the contrary, when the similarity variance decreases (the feature discrimination increases), it is increased to avoid overfitting.

[0053] (2) Structural subspace modeling: first, from unstructured multi-modal data such as text, image, etc., atomic facts (such as "Li Bai - birthplace - Shuiye City" and "Shuiye City - dynasty to which it belongs - Tang Dynasty") are extracted through entity extraction, relation extraction, etc. Form a symbolic triple in the form of "entity-relation-entity"; then, based on the initial "entity-relation" graph composed of these triples, the potential connection paths between entities are mined by means of depth-first search (DFS) and breadth-first search (BFS) graph traversal algorithms, and combined with semantic constraint conditions (such as path length ≤ 5 hops to ensure semantic relevance, filtering "height" and "friends" and other irrelevant relationships to the target domain, excluding "person-relation-number-relation-color" and other invalid entity type chains), effective paths are selected; then enter the semantic abstraction stage, first map each relationship in the effective path to an embedding vector, and generate a unified vector that can represent the semantic of the path by fusing the forward and backward sequence semantics of the path (such as the logical order of "place of birth → dynasty to which it belongs") through a bidirectional long short-term memory network (BiLSTM), and then relying on the domain ontology and pre-set logical rules, abstracting specific paths into general logical patterns (such as "Li Bai - birthplace - Shuiye City - dynasty to which it belongs - Tang Dynasty" and "Du Fu - birthplace - Gong County - dynasty to which it belongs - Tang Dynasty" are abstracted into "person - birthplace - place - dynasty to which it belongs - dynasty"), while extracting the implicit relationships in the path (such as abstracting the indirect association of "person → dynasty to which it belongs" from the above paths); finally, integrate these abstracted structured logical knowledge with the original symbolic triples to organize a knowledge graph structure that has both topological connectivity and semantic logic. Then inject it into the vector space to enhance the distinguishability of entity paths in the knowledge graph, thereby reducing the loss of graph topology information. Construct multi-scale path perception repulsion loss to impose repulsion constraints on relationship path entity features greater than two, forcing the path feature distance of the positive sample entity pair associated with semantics to be less than that of the positive and negative sample pairs The path feature distance of the path characteristics, so as to strengthen the distinguishability of the path semantics: ; Wherein, represents the anchor entity After the path encoding characteristics of the jump relationship path, represents the positive sample entity After the path encoding characteristics of the jump relationship path, represents the negative sample entity After the path encoding characteristics of the jump relationship path, is the sampling weight of different hop paths; is the boundary margin, which is used to control the strictness of the exclusion constraint. Since the multi-hop relationship path of the entity in the knowledge graph has the sequence semantic attribute, BiLSTM is used to encode the path. For the relationship path of length k of entity e (composed of k relationships ), each relationship is first mapped to an embedding vector , and then the embedding vectors are fused through BiLSTM to finally output the encoding feature of the path , the formula is: ; In order to cover relationship paths of different lengths and balance their influence, different sampling weights and boundary margins are set for 2-5 hop paths, and the specific parameters are shown in Table 1: Table 1 sets parameters for 2-5 hop paths

[0054] The sampling weight decreases with the increase of path length (short path semantics is more direct, and contributes more significantly to distinguishability); the boundary margin increases with the increase of path length (long path semantics is more complex, and needs more strict exclusion constraint to ensure feature distinguishability).

[0055] If only relying on relationship embedding to encode the path will lose the semantic information of the entity type (such as the type difference between “person” and “organization” will affect the path semantics). Therefore, the semantic expression of the relationship path is enhanced by fusing the entity type embedding, and the relationship The final embedding of the connection source entity src and the target entity dst is calculated as: ;​​​ where, , is the learned weight matrix, is the original embedding of the relation , , are the type embeddings of the source entity src and the target entity dst respectively. The entity type information is fused by vector concatenation operation to make the relation embedding more accurately reflect the semantic properties of the path.

[0056] In the multi-loss joint optimization scenario, and the gradients of the two types of losses may have serious direction conflicts, i.e. , leading to unstable training. To this end, the projected gradient descent (PGD) is used to coordinate the gradients, and the updated gradient is: ; where, is a dynamic parameter, , t is the training round, which gradually decays from 0.7 with the training process to achieve strong coordination of the gradient direction in the early training stage and weak intervention in the later stage to ensure convergence and balance the gradient update of the two types of losses.

[0057] It should be noted that the projected gradient descent of the present embodiment not only optimizes the parameters in the two types of loss functions and , but also can automatically adjust the parameters based on the global optimization of the projected gradient descent, i.e. the parameters in the mutual information constraint loss, ResNet-50 model, graph neural network and bidirectional long short-term memory network (BiLSTM) can also be adjusted by back propagation combined with the projected gradient descent. For example, when the projected feature (i.e. the projected vector and ) does not satisfy the mutual information constraint loss, it will not simply perform a projection operation again, but will perform a global optimization of the parameters through back propagation and the projected gradient descent mechanism. The gradient calculated by the mutual information constraint loss function will update the front-end feature encoder (i.e. the ResNet-50 model and the graph neural network) and the learnable projection matrix parameters (i.e. the parameters in the orthogonal projection operation) at the same time, guiding them to learn how to generate original features (i.e. extracted visual feature vectors and structural feature vectors) with lower coupling degree and find a better projection direction, so as to guide the projection vectors and obtained after the orthogonal projection operation to be more optimal, and make the projection vectors satisfy the mutual information constraint loss.

[0058] 3. Dual space fusion output and cross-modal retrieval.

[0059] joint vector It is constructed by segmented concatenation of text features, visual features, optimal path structure features, and layer normalization. Its expression is as follows: ; in, Represents the target path encoding feature subset, that is, the optimal path structure feature set.

[0060] Finally, generate the corresponding joint vector for each data , all joint vectors are stored in the efficient vector database Milvus, and an approximate nearest neighbor (ANN) index is established.

[0061] Through the above steps, all modal data (text and images) are mapped into vector representations within the same space, achieving deep semantic alignment. In practical applications, by encoding user queries (text, image, or mixed input) into query vectors in the same space and utilizing approximate nearest neighbor search techniques, rapid retrieval can be performed directly within this vector space (i.e., the vector database). This effectively and accurately returns the most relevant cross-modal search results, improving the efficiency and precision of cross-modal retrieval.

[0062] Compared with the existing technology, the technical solution of this embodiment has the following advantages: The algorithm optimizes mutually exclusive objectives (cross-modal alignment and graph structure preservation) within independent subspaces to improve retrieval accuracy. Path-constrained contrastive learning ensures that cross-modal semantic alignment does not compromise valuable topological associations. A multi-scale path-aware exclusion loss is employed to explicitly encode 2- to 5-hop relational paths. BiLSTM aggregates path semantics and increases the vector distance between terminal entities in divergent paths. This gives the model deep reasoning capabilities and enables it to effectively distinguish locally similar but globally distinct entities, thereby improving multimodal retrieval accuracy.

[0063] Reference Figure 3 , the embodiment of the present application also provides a cross-modal data retrieval system based on a multimodal knowledge graph, the system comprising: The data processing unit 100 is used to pre-process the acquired text modality data and image modality data to obtain standard text modality data and standard image modality data; A feature extraction unit 200 is configured to extract text feature vectors from the standard text modality data, and to extract visual feature vectors and structural feature vectors from the standard image modality data; A feature decoupling unit 300 is configured to perform feature decoupling based on the visual feature vector and the structural feature vector to determine a target visual projection vector; A first construction unit 400 is configured to construct a path-constrained contrastive learning loss function based on the target visual projection vector and the text feature vector; The path mining unit 500 is configured to construct a multi-modal knowledge graph between texts and images, and mine effective paths between entities in the multi-modal knowledge graph. The second construction unit 600 is configured to encode the effective paths into path encoding features, and construct a multi-scale path perception repulsion loss function based on the path encoding features. The joint optimization unit 700 is configured to jointly optimize the path constraint contrast learning loss function and the multi-scale path perception repulsion loss function, and determine a target path encoding feature subset. The feature fusion unit 800 is configured to fuse the text feature vector, the target visual projection vector, and the target path encoding feature subset to obtain a joint vector. The data retrieval unit 900 is configured to perform cross-modal data retrieval based on the joint vector to obtain a retrieval result.

[0064] It should be noted that, since the cross-modal data retrieval system based on a multi-modal knowledge graph in the embodiment and the cross-modal data retrieval method based on a multi-modal knowledge graph described above are based on the same inventive concept, the corresponding content in the method embodiment is also applicable to the system embodiment, and will not be described in detail here.

[0065] Referring to Figure 4 The embodiment of the present application further provides an electronic device, and the electronic device comprises: at least one memory; at least one processor; at least one program; The program is stored in the memory, and the processor executes the at least one program to implement the cross-modal data retrieval method based on a multi-modal knowledge graph described above.

[0066] The electronic device can be any intelligent terminal, including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.

[0067] The electronic device of the embodiment of the present application will be described in detail below.

[0068] The processor 1600 can be implemented in a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute a related program to implement the technical solutions provided by the embodiments of the present application. The memory 1700 can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present disclosure are implemented by software or firmware, the related program codes are stored in the memory 1700 and are called and executed by the processor 1600 to implement the multi-modal knowledge graph-based cross-modal data retrieval method provided by the embodiments of the present disclosure.

[0069] The input / output interface 1800 is configured to realize information input and output. The communication interface 1900 is configured to realize the communication interaction between the device and other devices. The communication can be realized by a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.). The bus 2000 is configured to transmit information between various components (for example, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900) of the device. The processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are connected to each other through the bus 2000 to realize the communication connection between them in the device.

[0070] The present disclosure also provides a storage medium, which is a computer readable storage medium, and stores computer executable instructions for causing a computer to execute the multi-modal knowledge graph-based cross-modal data retrieval method.

[0071] The memory is a non-transitory computer readable storage medium, which can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0072] The embodiments described in the present disclosure are used to more clearly illustrate the technical solutions of the present disclosure, and do not constitute a limitation on the technical solutions provided by the present disclosure. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the present disclosure are also applicable to similar technical problems.

[0073] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and can include more or fewer steps than the figures, or combine certain steps, or different steps.

[0074] The apparatus embodiments described above are merely illustrative, and units described as separate components can or can not be physically separate, i.e., can be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.

[0075] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the function modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.

[0076] The terms "first", "second", "third", "fourth" and the like used in the description of the present application and the above-described drawings (if any) are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0077] It should be understood that in the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the association between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0078] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiments described above are merely illustrative, for example, the division of units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.

[0079] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or can be distributed to a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0080] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0081] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various program storage media. The embodiments of the present application are described in detail above in combination with the drawings, but the present application is not limited to the above embodiments, and various changes can be made within the knowledge range of those skilled in the art without departing from the purpose of the present application.

[0082] The embodiments of the present application are described in detail above in combination with the drawings, but the present application is not limited to the above embodiments, and various changes can be made within the knowledge range of those skilled in the art without departing from the purpose of the present application.

Claims

1. A cross-modal data retrieval method based on multimodal knowledge graph, characterized in that: The method comprises: Preprocessing the acquired text modal data and image modal data to obtain standardized text modal data and standardized image modal data; Extracting a text feature vector from the canonical text modality data, and extracting a visual feature vector and a structural feature vector from the canonical image modality data; Perform feature decoupling based on the visual feature vector and the structural feature vector to determine a target visual projection vector; Constructing a path-constrained contrastive learning loss function based on the target visual projection vector and the text feature vector; Constructing a multimodal knowledge graph between text and images, and mining effective paths between entities in the multimodal knowledge graph; Encoding the valid path into a path encoding feature, and constructing a multi-scale path-aware rejection loss function based on the path encoding feature; Jointly optimizing the path-constrained contrastive learning loss function and the multi-scale path-aware rejection loss function to determine a target path encoding feature subset; Fusing the text feature vector, the target visual projection vector, and the target path encoding feature subset to obtain a joint vector; Cross-modal data retrieval is performed based on the joint vector to obtain a retrieval result.

2. The cross-modal data retrieval method based on a multimodal knowledge graph according to claim 1 is characterized in that: The performing feature decoupling based on the visual feature vector and the structural feature vector to determine the target visual projection vector includes: Mapping the visual feature vector and the structural feature vector to an orthogonal subspace using an orthogonal projection operation to obtain an initial visual projection vector and an initial structural projection vector; Constructing a mutual information constraint loss function for feature decoupling, and determining whether the initial visual projection vector and the initial structural projection vector satisfy the mutual information constraint loss function; If the initial visual projection vector and the initial structural projection vector do not satisfy the mutual information constraint loss function, updating the gradient of the mutual information constraint loss function through backpropagation to optimize the visual feature vector and the structural feature vector extracted during the forward propagation process; Based on the optimized visual feature vector and structural feature vector, an orthogonal projection operation is used to obtain an optimized initial visual projection vector and an optimized initial structural projection vector. After the mutual information constraint loss function is satisfied, the optimized initial visual projection vector is used as the target visual projection vector.

3. The cross-modal data retrieval method based on a multimodal knowledge graph according to claim 1 is characterized in that: The constructing of a path-constrained contrastive learning loss function based on the target visual projection vector and the text feature vector includes: ; in, represents the path-constrained contrastive learning loss function, represents the set of positive sample pairs that satisfy the path constraint, represents the text feature vector, represents the target visual projection vector, represents the cosine similarity function of the feature pair, represents the temperature coefficient, represents the number of negative samples, Represents a text entity, Represents an image entity.

4. The cross-modal data retrieval method based on a multimodal knowledge graph according to claim 1 is characterized in that: The mining of effective paths between entities in the multimodal knowledge graph includes: Preset semantic constraints; Using depth-first search and breadth-first search to mine potential connection paths between entities in the multimodal knowledge graph; Valid paths between entities that meet the semantic constraint conditions are mined from the potential connection paths.

5. The cross-modal data retrieval method based on multimodal knowledge graph according to claim 1 is characterized in that: The constructing of a multi-scale path-aware rejection loss function based on the path encoding feature includes: ; in, represents the multi-scale path-aware rejection loss function, represents the sampling weight of different hop paths, Represents an anchor entity go through Path encoding features after skipping relation paths, Represents a positive sample entity go through Path encoding features after skipping relation paths, Represents negative sample entities go through Path encoding features after skipping relation paths, Represents the boundary margin.

6. The cross-modal data retrieval method based on a multimodal knowledge graph according to claim 1 is characterized in that: The fusing of the text feature vector, the target visual projection vector, and the target path encoding feature subset to obtain a joint vector includes: performing layer-normalization on the text feature vector, the target visual projection vector, and the target path encoding feature subset, respectively, to obtain a layer-normalized text feature vector, a layer-normalized target visual projection vector, and a layer-normalized target path encoding feature subset; The layer-normalized text feature vector, the layer-normalized target visual projection vector, and the layer-normalized target path encoding feature subset are concatenated to obtain a joint vector.

7. The cross-modal data retrieval method based on multimodal knowledge graph according to claim 1 is characterized in that: The cross-modal data retrieval based on the joint vector is performed to obtain a retrieval result, including: storing the joint vector generated from each piece of data into a vector database having the same spatial vector representation, and establishing an approximate nearest neighbor index of the vector database; Encoding the acquired modality data to be retrieved into a query vector represented by the same space vector; In a vector database with an approximate nearest neighbor index, an approximate nearest neighbor search is used to perform cross-modal data retrieval on the query vector to obtain a retrieval result.

8. A cross-modal data retrieval system based on a multimodal knowledge graph, characterized in that: The system comprises: A data processing unit, configured to pre-process the acquired text modality data and image modality data to obtain standardized text modality data and standardized image modality data; A feature extraction unit, configured to extract a text feature vector from the standard text modality data, and to extract a visual feature vector and a structural feature vector from the standard image modality data; a feature decoupling unit, configured to perform feature decoupling based on the visual feature vector and the structural feature vector to determine a target visual projection vector; A first construction unit is configured to construct a path-constrained contrastive learning loss function based on the target visual projection vector and the text feature vector; A path mining unit, configured to construct a multimodal knowledge graph between text and images, and to mine valid paths between entities in the multimodal knowledge graph; A second construction unit is configured to encode the valid path into a path encoding feature, and construct a multi-scale path-aware rejection loss function based on the path encoding feature; a joint optimization unit, configured to jointly optimize the path-constrained contrastive learning loss function and the multi-scale path-aware rejection loss function to determine a target path encoding feature subset; a feature fusion unit, configured to fuse the text feature vector, the target visual projection vector, and the target path encoding feature subset to obtain a joint vector; A data retrieval unit is used to perform cross-modal data retrieval based on the joint vector to obtain a retrieval result.

9. An electronic device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the cross-modal data retrieval method based on a multimodal knowledge graph as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the cross-modal data retrieval method based on a multimodal knowledge graph as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent visualization and text association method for multi-modal knowledge graph

    CN119441281A

  • Knowledge graph-combined brilliance large-scale model culture knowledge generation and retrieval method

    CN120256644A

  • Method and apparatus for completing knowledge graph, electronic device, and computer-readable medium

    WO2024120385A1

Cited By

  • Road disease retrieval and diagnosis method based on multi-modal knowledge graph

    CN121809626A