Method for constructing multi-modal knowledge graph in cross-media retrieval
Patent Information
- Application Number
- CN202511108141.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-08-08
AI Technical Summary
然而,构建多模态知识图谱并非易事,它面临着诸多技术挑战
[0017] As can be seen from the above technical solution, this application provides a method and apparatus for constructing a multimodal knowledge graph in cross-media retrieval. The extracted multimodal features are mapped to a multimodal feature space through a linear transformation layer. In the multimodal feature space, intramodal attention weights of each modal feature are calculated based on a self-attention mechanism, and cross-modal attention weights of different modal features are calculated based on a cross-attention mechanism. The two are fused to obtain a final fusion weight, which is then used to weight each modal feature and input into a graph attention network to obtain a multimodal fusion graph structure. Semantic parsing of each modal feature is performed and matched with a preset multimodal semantic knowledge base to determine potential semantic associations. The multimodal fusion graph structure is semantically aligned based on a preset graph matching algorithm and potential semantic associations to obtain a multimodal knowledge graph. Cross-modal data retrieval is then performed based on the multimodal knowledge graph, thereby improving the efficiency and accuracy of cross-media multimodal data retrieval.
Smart Images

Figure CN120611774B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, specifically to a method for constructing a multimodal knowledge graph in cross-media retrieval. Background Technology
[0002] In today's era of rapid information technology development, cross-media data such as text, images, and audio are experiencing explosive growth, containing a wealth of information and knowledge. However, due to the unique feature representations and semantic structures of various media modalities, traditional single-modality-based retrieval methods face severe challenges and struggle to meet the comprehensiveness and accuracy requirements of cross-media retrieval.
[0003] Multimodal knowledge graphs, as an emerging semantic representation structure, offer new solutions for cross-media retrieval. However, constructing a multimodal knowledge graph is no easy task, facing numerous technical challenges. First, effectively integrating features from different modalities is a key issue. Data from different modalities have different representations and feature dimensions; unifying and integrating these features to enable them to jointly support cross-media retrieval is a problem that urgently needs to be solved. Second, the semantic gap between modalities is another challenge that needs to be overcome in constructing a multimodal knowledge graph. Data from different modalities may have semantic differences and gaps; narrowing these differences and ensuring the semantic accuracy of the knowledge graph in cross-media retrieval is another important technical challenge.
[0004] In summary, there is an urgent need for a method to construct multimodal knowledge graphs in cross-media retrieval to improve the efficiency and accuracy of cross-media retrieval of multimodal data. Summary of the Invention
[0005] To address the problems in the prior art, this application provides a method and apparatus for constructing a multimodal knowledge graph in cross-media retrieval, which can improve the efficiency and accuracy of cross-media retrieval of multimodal data.
[0006] To solve at least one of the above problems, this application provides the following technical solution: Firstly, this application provides a method for constructing a multimodal knowledge graph in cross-media retrieval, including: Feature extraction is performed on the multimodal features respectively, and the modal features obtained after feature extraction are mapped to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space; In the multimodal feature space, self-attention calculation is performed on each modal feature according to the self-attention mechanism to determine the corresponding intramodal attention weight. Correlation calculation is performed on different modal features according to the cross-attention mechanism to determine the corresponding cross-modal attention weight. The intramodal attention weight and the cross-modal attention weight are fused, and the modal features are weighted based on the final fused weight to determine the corresponding multimodal fusion graph structure. Semantic parsing is performed on each modal feature, and the semantic features after semantic parsing are matched with a preset multimodal semantic knowledge base to determine the potential semantic associations between each modal feature. The multimodal fusion graph structure is semantically aligned according to a preset graph matching algorithm and the potential semantic associations to determine the corresponding multimodal knowledge graph. Cross-modal data retrieval is performed based on the multimodal knowledge graph.
[0007] Furthermore, the feature extraction of the multimodal features includes: Semantic extraction is performed on text modal data based on a pre-trained language model to determine the corresponding text modal features, wherein the text modal features include vocabulary, grammatical structure, and global semantic features; The image modal data is subjected to multi-layer convolution and pooling operations based on a preset convolutional neural network to determine the corresponding image modal features, wherein the image modal features include local image features and global image features; Acoustic features are extracted from audio modal data according to a preset time-frequency transformation algorithm, and semantic features are extracted from the acoustic features obtained by the acoustic feature extraction according to a preset temporal modeling network to determine the corresponding audio modal features.
[0008] Further, the step of mapping the modal features obtained after feature extraction to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space includes: The text modal features, image modal features, and audio modal features are normalized, and the normalized modal features are projected onto a shared feature space of the same dimension according to the learnable weight matrix. In the shared feature space, activation functions are applied to each modal feature after projection to determine the corresponding multimodal feature space.
[0009] Further, the step of performing self-attention calculation on each modal feature based on the self-attention mechanism to determine the corresponding intra-modal attention weights includes: Matrix calculations are performed on each modal feature using trainable parameters to determine the query vector, key vector, and value vector corresponding to each modal feature. Self-attention is calculated for feature elements within the same modality to determine the corresponding same-modality attention score; The corresponding intramodal attention weights are determined by weighting the value vector of the modality based on the intramodal attention scores.
[0010] Furthermore, the step of performing correlation calculations on features of different modalities based on the cross-attention mechanism to determine the corresponding cross-modal attention weights includes: Different modalities are divided into target modalities and source modalities. The query vector of the target modalities is interacted with the key vector of the source modalities to determine the corresponding cross-modal attention scores. The cross-modal attention weights of the target modality are determined by weighting the value vector of the source modality based on the cross-modal attention scores.
[0011] Furthermore, the semantic parsing of the modal features includes: Named entity recognition and relation extraction are performed on text modal features to determine the corresponding structured semantic representation; The image modal features are analyzed based on the preset visual semantic parsing model to determine the corresponding visual entities and spatial relationships; Audio modal features are analyzed based on acoustic event detection and speech semantic understanding to determine the corresponding high-level semantic labels.
[0012] Further, the step of semantically aligning the multimodal fusion graph structure according to a preset graph matching algorithm and the latent semantic association to determine the corresponding multimodal knowledge graph includes: Based on the potential semantic associations, the semantic similarity between entities of different modalities in the multimodal fusion graph structure is calculated to determine the corresponding cross-modal association weight matrix; The cross-modal association weight matrix is solved for cross-modal entity alignment according to the preset graph matching algorithm to determine the corresponding optimal alignment path; The multimodal fusion graph structure is semantically aligned according to the optimal alignment path to determine the corresponding multimodal knowledge graph.
[0013] Secondly, this application provides a multimodal knowledge graph construction apparatus for cross-media retrieval, comprising: The multimodal feature extraction module is used to extract features from multimodal features respectively, and to map the modal features obtained after feature extraction to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space. The multimodal fusion graph construction module is used to perform self-attention calculation on each modal feature in the multimodal feature space according to the self-attention mechanism to determine the corresponding intramodal attention weights, perform association calculation on different modal features according to the cross-attention mechanism to determine the corresponding cross-modal attention weights, fuse the intramodal attention weights and the cross-modal attention weights, and weight each modal feature based on the final weight after fusion to determine the corresponding multimodal fusion graph structure. The multimodal knowledge graph determination module is used to perform semantic parsing on the modal features, match the semantic features after semantic parsing with a preset multimodal semantic knowledge base, determine the potential semantic associations between the modal features, perform semantic alignment on the multimodal fusion graph structure according to a preset graph matching algorithm and the potential semantic associations, determine the corresponding multimodal knowledge graph, and perform cross-modal data retrieval based on the multimodal knowledge graph.
[0014] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the multimodal knowledge graph construction method in cross-media retrieval.
[0015] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multimodal knowledge graph construction method in cross-media retrieval.
[0016] Fifthly, this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the multimodal knowledge graph construction method in cross-media retrieval.
[0017] As can be seen from the above technical solution, this application provides a method and apparatus for constructing a multimodal knowledge graph in cross-media retrieval. The extracted multimodal features are mapped to a multimodal feature space through a linear transformation layer. In the multimodal feature space, intramodal attention weights of each modal feature are calculated based on a self-attention mechanism, and cross-modal attention weights of different modal features are calculated based on a cross-attention mechanism. The two are fused to obtain a final fusion weight, which is then used to weight each modal feature and input into a graph attention network to obtain a multimodal fusion graph structure. Semantic parsing of each modal feature is performed and matched with a preset multimodal semantic knowledge base to determine potential semantic associations. The multimodal fusion graph structure is semantically aligned based on a preset graph matching algorithm and potential semantic associations to obtain a multimodal knowledge graph. Cross-modal data retrieval is then performed based on the multimodal knowledge graph, thereby improving the efficiency and accuracy of cross-media multimodal data retrieval. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is one of the flowcharts illustrating the multimodal knowledge graph construction method in cross-media retrieval according to an embodiment of this application; Figure 2 This is a structural diagram of the multimodal knowledge graph construction device in cross-media retrieval according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of the electronic device in the embodiments of this application.
[0020] Figure label: Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] The acquisition, storage, use, and processing of data in this application all comply with relevant regulations.
[0023] The semantic gap between different data modalities is a challenge that needs to be overcome in constructing a multimodal knowledge graph. This application provides a method and apparatus for constructing a multimodal knowledge graph in cross-media retrieval. The method involves mapping extracted multimodal features to a multimodal feature space through a linear transformation layer. In the multimodal feature space, intramodal attention weights for each modal feature are calculated using a self-attention mechanism, and cross-modal attention weights for different modal features are calculated using a cross-attention mechanism. The two are fused to obtain a final fused weight, which is then applied to each modal feature and input into a graph attention network to obtain a multimodal fused graph structure. Semantic parsing of each modal feature is performed and matched with a pre-defined multimodal semantic knowledge base to determine potential semantic associations. The multimodal fused graph structure is then semantically aligned based on a pre-defined graph matching algorithm and the potential semantic associations to obtain a multimodal knowledge graph. Cross-modal data retrieval is then performed based on the multimodal knowledge graph, thereby improving the efficiency and accuracy of cross-media multimodal data retrieval.
[0024] To improve the efficiency and accuracy of cross-media retrieval of multimodal data, this application provides an embodiment of a multimodal knowledge graph construction method in cross-media retrieval, see [link to embodiment]. Figure 1 The method for constructing a multimodal knowledge graph in cross-media retrieval specifically includes the following: Step S101: Perform feature extraction on the multimodal features respectively, and map the modal features obtained after feature extraction to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space; Optionally, in this embodiment, in order to construct a multimodal knowledge graph, we first perform multimodal feature extraction and map all features to the same vector space so that feature alignment can be performed subsequently.
[0025] Specifically, for text modalities, text preprocessing is first performed, including lexical analysis, syntactic analysis, and part-of-speech tagging, to clean and standardize the text data; a pre-trained word vector model (such as BERT) is used to convert the text into vector representation, while extracting key semantic information from the text, such as entities and relationships.
[0026] Specifically, for image modalities, convolutional neural networks (CNNs) are used to extract features from images, such as using network structures like ResNet, to extract local and global features of the image.
[0027] Specifically, for audio modalities, audio feature extraction algorithms (such as Mel frequency cepstral coefficients (MFCC)) are used to obtain the spectral features of the audio, and then deep learning models (such as Long Short-Term Memory networks (LSTM)) are used to further extract the semantic features of the audio.
[0028] Next, a fusion framework is constructed to map the extracted text, image, and audio features to a unified feature space.
[0029] Specifically, for the feature vector of each modality, a weighted summation operation is performed using a linear transformation layer, where the weight parameters are learned through the training process. Through the mapping of the linear transformation layer, features originally belonging to different modalities and having different dimensions and semantic spaces are integrated into a unified feature space. In this space, features from different modalities can be represented and compared in a unified way, providing a foundation for subsequent multimodal fusion operations.
[0030] Constructing a unified feature space is one of the key steps in multimodal fusion. Only when features from different modalities reside in the same space can complementary information in multimodal data be effectively mined, thereby improving the performance of multimodal tasks such as multimodal classification and retrieval.
[0031] Step S102: In the multimodal feature space, self-attention calculation is performed on each modal feature according to the self-attention mechanism to determine the corresponding intramodal attention weight. Correlation calculation is performed on different modal features according to the cross-attention mechanism to determine the corresponding cross-modal attention weight. The intramodal attention weight and the cross-modal attention weight are fused, and the modal features are weighted based on the final weight after fusion to determine the corresponding multimodal fusion graph structure. Optionally, in this embodiment, an attention mechanism is used to dynamically allocate weights according to the importance of different modal features, so that important modal features occupy a larger proportion in the fusion process.
[0032] Optionally, in this embodiment, a self-attention mechanism is adopted for each modal feature, the core of which is to calculate the similarity of each feature in the unified feature set. Specifically, for each feature vector in the modal feature set, it is used as a query vector, and all feature vectors in the entire modal feature set are used as key vectors and value vectors.
[0033] An attention score matrix is obtained by calculating the degree of similarity between the query vector and the key vector (after scaling using a dot product operation with a scaling factor equal to the reciprocal of the square root of the key vector's dimension). This score matrix represents the degree of association between each feature vector and other feature vectors. For example, in a text modality, the feature vector of a word is compared with the feature vectors of other words in the same text to determine the word's importance within the overall text context.
[0034] Then, the attention score matrix is normalized using the softmax function, ensuring that the sum of the attention scores for each feature vector is 1, thus obtaining the intra-modal attention weights. These weights reflect the importance of each feature vector relative to others within the modality, highlighting important feature vectors while suppressing less important ones.
[0035] Optionally, in this embodiment, a cross-attention mechanism is used between different modal features, the core of which is to calculate the similarity between different feature vectors. For example, in the text-image modality, if a word in the text has a high cross-modal attention weight with a certain object region feature in the image, then it means that the word may be describing that object.
[0036] Cross-attention mechanisms can establish semantic associations between different modalities. It enables models to understand the correspondences between data from different modalities. For example, in image-text retrieval tasks, through cross-attention, the model can determine which parts of the image match the content of the text description, thus retrieving images that more accurately match the text description.
[0037] Optionally, in this embodiment, the calculated intra-modal attention weights and cross-modal attention weights are fused. The weight coefficients are adjusted according to the actual task and data. For example, in some multimodal tasks where text is the main information, the weight coefficient of the intra-modal attention weight may be set relatively high, while in some tasks that emphasize the complementarity between modalities, the weight coefficient of the cross-modal attention weight may be increased.
[0038] Optionally, in this embodiment, each modal feature is weighted according to the final fusion weight. Each modal feature vector is multiplied by the corresponding final fusion weight, so that the resulting weighted feature vector better conforms to the overall semantic structure of the multimodal data.
[0039] The weighted feature vectors are input into a graph attention network. The weighted modal features are treated as node features in the graph, and the connections between nodes are determined by the associations between modalities, such as the association between words in text and regions in an image. The graph attention network aggregates information from neighboring nodes by calculating attention coefficients between nodes, thereby updating the features of each node. For example, for a text node in a text-image fusion graph structure, it updates its own features based on the features of the image nodes connected to it and the attention coefficients between them, enabling the text features to better integrate with image information.
[0040] The resulting updated set of node features forms a graph structure for the fusion of text, images, and audio multimodal data. This structure can more accurately represent the entity relationships and structural information between multimodal data, providing an effective data representation for subsequent multimodal tasks.
[0041] Step S103: Perform semantic parsing on each modal feature, match the semantic features after semantic parsing with a preset multimodal semantic knowledge base, determine the potential semantic associations between each modal feature, perform semantic alignment on the multimodal fusion graph structure according to a preset graph matching algorithm and the potential semantic associations, determine the corresponding multimodal knowledge graph, and perform cross-modal data retrieval based on the multimodal knowledge graph.
[0042] Optionally, in this embodiment, we delve into the intrinsic relationships between the features of each modality at the semantic level, and update the multimodal fusion graph structure based on this.
[0043] Optionally, in this embodiment, the core objective is to first transform the feature vectors of each modality, after previous feature extraction and fusion processing, into feature representations with clear semantic meaning.
[0044] Specifically, for text modal features, semantic parsing uses natural language processing technology to abstract text words, sentences, etc. into semantic vectors, thereby obtaining semantic entities and relationships between entities.
[0045] Specifically, for image modal features, semantic parsing of images is performed by combining computer vision techniques such as object detection, image segmentation, and scene understanding to form semantic features of the image. For example, in a city street scene image, the semantic features obtained after semantic parsing may include elements such as "roads," "buildings," "pedestrians," and "vehicles" and their relative positions, as well as the semantic concept that the overall scene is a "city street."
[0046] Specifically, in audio modal features, speech recognition technology first converts speech into text, then performs semantic analysis on the text to extract the core semantic information conveyed by the speech, such as the intent of the instruction and the topic of the dialogue.
[0047] Optionally, in this embodiment, after semantic parsing is completed, the obtained modal semantic features will be matched with a preset multimodal semantic knowledge base.
[0048] Specifically, the matching process is achieved by calculating the similarity between semantic features and semantic entries in the knowledge base.
[0049] For text semantic features, cosine similarity is used to compare the similarity between the text semantic vector and the text topic vector in the knowledge base; For image semantic features, techniques such as the bag-of-words model and similarity measurement networks in deep learning are used to match the semantics of image scenes or object categories in the knowledge base; For audio semantic features, the similarity between the audio features and the audio category features in the knowledge base is used for judgment.
[0050] Specifically, based on the semantic matching results, we further explore the potential semantic relationships between features of different modalities. In multimodal scenarios, different modalities often describe the same thing or scene from different perspectives, and there are rich semantic relationships between them. For example, in a multimodal dataset containing product images, product description text, and product promotional audio, the description of the product's "high-end and sophisticated" appearance in the text is related to the exquisite appearance of the product shown in the image, and the semantic concepts such as "superior quality" promoted in the audio.
[0051] By analyzing the relationships between the matched semantic entries in the knowledge base, we can identify the parts of each modality feature that have similar or related semantic concepts.
[0052] Optionally, in this embodiment, after determining the potential semantic associations between the modal features, the multimodal fusion graph structure is then semantically aligned using a preset graph matching algorithm.
[0053] The core function of graph matching algorithms is to adjust the connection methods of nodes and edges in the fused graph structure based on semantic association information, so that the graph structure can more accurately reflect the semantic consistency between various modal features.
[0054] Specifically, the algorithm compares the semantic association strength between feature nodes of different modalities. For nodes with strong semantic association, it strengthens the weight of the connection edges between them, and may even add new connection edges when necessary. For nodes with weak or irrelevant semantic association, it may reduce the weight of the connection edges or delete the connection edges.
[0055] For example, in a multimodal dataset containing news report text, news scene photos, and news anchor audio, if the text and images semantically refer to the topic of "major sporting events," and the audio also mentions this topic, the graph matching algorithm will strengthen the connections between text, image, and audio nodes, ensuring that they are closely linked in the fused graph structure, forming a semantically consistent subgraph structure around the topic of "major sporting events."
[0056] We integrate and optimize the nodes and edges in the fusion graph structure according to semantic hierarchy and semantic relationships. After semantic alignment, the multimodal fusion graph structure is finally transformed into a multimodal knowledge graph. Nodes with the same or similar semantic concepts are merged or categorized to avoid duplication and redundancy. For parts with complex semantic relationships, we add intermediate semantic nodes or redefine the semantic types of edges to make them express semantic associations more clearly and accurately.
[0057] In cross-modal retrieval scenarios, users can input a query in one modality (such as an image), and the system can quickly locate other modal data (such as related text descriptions, audio introductions, etc.) that are semantically related to the image based on the knowledge graph. This greatly improves the accuracy and efficiency of retrieval, enabling users to obtain information from multimodal data more comprehensively and in-depth.
[0058] This example demonstrates how this embodiment constructs a multimodal knowledge graph through multimodal feature extraction, fusion, and semantic similarity matching, thereby improving the efficiency and accuracy of cross-media retrieval of multimodal data.
[0059] As described above, the multimodal knowledge graph construction method for cross-media retrieval provided in this application can map the extracted multimodal features to a multimodal feature space through a linear transformation layer. In the multimodal feature space, the intramodal attention weights of each modal feature are calculated according to the self-attention mechanism, and the cross-modal attention weights of different modal features are calculated according to the cross-attention mechanism. The two are fused to obtain the final fusion weight, which is then used to weight each modal feature and input into a graph attention network to obtain a multimodal fusion graph structure. Semantic parsing is performed on each modal feature and matched with a preset multimodal semantic knowledge base to determine potential semantic associations. The multimodal fusion graph structure is semantically aligned according to a preset graph matching algorithm and potential semantic associations to obtain a multimodal knowledge graph. Cross-modal data retrieval is then performed based on the multimodal knowledge graph, thereby improving the efficiency and accuracy of cross-media multimodal data retrieval.
[0060] In one embodiment of the multimodal knowledge graph construction method in cross-media retrieval of this application, see [link to relevant documentation]. Figure 2 It can also specifically include the following: Step S201: Extract semantics from the text modal data based on the pre-trained language model to determine the corresponding text modal features, wherein the text modal features include vocabulary, grammatical structure and global semantic features; Step S202: Perform multi-layer convolution and pooling operations on the image modal data according to the preset convolutional neural network to determine the corresponding image modal features, wherein the image modal features include local image features and global image features; Step S203: Extract acoustic features from audio modal data according to a preset time-frequency transformation algorithm, and extract semantic features from the acoustic features obtained by the acoustic feature extraction according to a preset time-series modeling network to determine the corresponding audio modal features.
[0061] Optionally, in this embodiment, a pre-trained language model is used to process text modal data.
[0062] Specifically, pre-trained language models are trained on a large amount of text corpus.
[0063] First, the language model focuses on lexical features, including information such as the individual words appearing in the text and their parts of speech.
[0064] Secondly, the model analyzes grammatical structural features, namely the combinational relationships between words and the structural form of sentences.
[0065] Finally, the model also extracts global semantic features, that is, it understands the overall semantic information of the entire text segment (such as a paragraph or an article).
[0066] This step yields text modal features that comprehensively reflect the semantic information of the text, providing a rich semantic foundation for subsequent multimodal fusion and other operations.
[0067] Optionally, in this embodiment, a preset convolutional neural network is used to extract image modal features.
[0068] Specifically, in the feature extraction process, multiple convolution operations are first performed. In the shallow convolutional layers, local features such as image edges and textures are mainly extracted. As the convolutional layers deepen, the model gradually extracts more abstract and higher-level features.
[0069] Next, pooling will be performed. The purpose of pooling is to downsample the convolutional feature map, reducing its spatial size while retaining important feature information.
[0070] Specifically, after multiple convolution and pooling operations, the resulting image modal features include both rich local image features, such as edges and textures, and global image features that can reflect the overall semantics of the image, such as object shapes and spatial layouts.
[0071] Optionally, in this embodiment, in terms of audio modal data processing, a preset time-frequency transformation algorithm is first used to extract acoustic features from the audio data to extract the acoustic features of the audio.
[0072] Specifically, after obtaining the acoustic features, a pre-defined temporal modeling network is used to extract semantic features from these acoustic features. Audio signals are sequential data that change over time; therefore, the temporal modeling network processes the temporal relationships in the audio data to identify the text content corresponding to the speech, thereby extracting the semantic features of the speech. These semantic features enable audio modal data to express its semantic information in a way that more closely resembles human understanding of language and sound.
[0073] In step S203, this embodiment extracts features from text, image, and audio data respectively, determines the features of each modality, and lays a solid data foundation for the subsequent construction of the knowledge graph.
[0074] In one embodiment of the multimodal knowledge graph construction method in cross-media retrieval of this application, the method may further include the following: Step S301: Normalize the text modal features, the image modal features, and the audio modal features, and project each modal feature after normalization to a shared feature space of the same dimension according to the learnable weight matrix; Step S302: In the shared feature space, apply activation functions to each modal feature after projection to determine the corresponding multimodal feature space.
[0075] Optionally, in this embodiment, normalization processing is performed on the text modal features, image modal features, and audio modal features respectively, so that the numerical range of different modal features is adjusted to a relatively uniform range.
[0076] Specifically, data from different modalities may have different scales and dimensions in their raw state. If normalization is not performed, subsequent processing may be affected by the differences in numerical magnitude, causing features of some modalities to dominate the calculation process while features of other modalities are ignored. For example, image modal features may contain large values such as pixel values, while text modal features may contain relatively small values after encoding.
[0077] Optionally, in this embodiment, the normalized modal features are projected onto a shared feature space of the same dimension based on the learnable weight matrix.
[0078] Specifically, a learnable weight matrix is a parameterized method that learns the importance and correlation between features of different modalities. During the projection process, the learnable weight matrix is automatically adjusted based on the training data to find the optimal way to map features of different modalities into a shared space.
[0079] Optionally, in this embodiment, in the shared feature space, activation functions are applied to each modal feature after projection to determine the corresponding multimodal feature space.
[0080] Specifically, the role of activation functions is to introduce non-linearity into the model, enabling it to learn more complex feature representations and mapping relationships. For example, the semantic relationship between text and images is not a simple linear one; activation functions can uncover these complex semantic connections. Features processed by activation functions can better reflect the inherent characteristics of each modality in the shared space, thus forming a richer and more effective multimodal feature space.
[0081] Through step S302, this embodiment successfully constructed a multimodal feature space, laying the foundation for subsequent feature processing in this feature space.
[0082] In one embodiment of the multimodal knowledge graph construction method in cross-media retrieval of this application, the method may further include the following: Step S401: Perform matrix calculations on each modal feature using trainable parameters to determine the query vector, key vector, and value vector corresponding to each modal feature; Step S402: Perform self-attention calculation on feature elements within the same modality to determine the corresponding same-modality attention score; Step S403: Weight the value vector of the modality according to the same-modal attention score to determine the corresponding intra-modal attention weight.
[0083] Optionally, in this embodiment, during the multimodal feature processing, it is first necessary to vectorize each modal feature to facilitate subsequent calculation of the attention mechanism. Specifically, matrix calculations are performed on each modal feature using trainable parameters to determine the query vector, key vector, and value vector corresponding to each modal feature.
[0084] Specifically, for the feature matrix of each mode (Where m represents the modality type, n is the number of features, and d is the feature dimension) Calculate the query matrix Qm, the key matrix Km, and the value matrix Vm respectively.
[0085] Query Matrix , These are trainable parameters.
[0086] The query vector represents the "query intent" of the current feature when searching for related information. It is obtained by performing a linear transformation on the original modality features, and the parameters of the linear transformation are trainable. Key matrix , These are trainable parameters.
[0087] The key vector represents the "key information" contained in the feature elements, used for matching with the query vector. Similarly, the key vector is obtained by performing a linear transformation on the original modality features; Value matrix , These are trainable parameters.
[0088] The value vector represents the actual "value" or "content" of the feature elements. It is weighted and summed according to attention weights to generate the final feature representation. The value vector is also obtained from the original modality features through a linear transformation.
[0089] Optionally, in this embodiment, after determining the query vector, key vector, and value vector, self-attention calculation is then performed on the feature elements within the same modality to determine the corresponding same-modality attention score.
[0090] Specifically, calculate the self-attention score matrix. It is used to characterize the dependencies between feature elements within the same modality, allowing features within the same modality to influence and complement each other.
[0091] Specifically, output intra-modal enhancement features The value vector of each feature element is multiplied by its corresponding intramodal attention weight, and then the weighted vectors of all feature elements are summed to obtain a comprehensive feature representation of the modality.
[0092] Through step S403, this embodiment successfully realizes the self-attention calculation of each modal feature and the determination of intramodal attention weights, generating a comprehensive feature representation that can better represent each modal feature, providing an effective intramodal feature foundation for subsequent cross-modal interaction and fusion.
[0093] In one embodiment of the multimodal knowledge graph construction method in cross-media retrieval of this application, the method may further include the following: Step S501: Divide different modalities into target modalities and source modalities, and interact the query vector of the target modality with the key vector of the source modality to determine the corresponding cross-modal attention score; Step S502: Weight the value vector of the source modality according to the cross-modal attention score to determine the cross-modal attention weight of the corresponding target modality.
[0094] Optionally, in this embodiment, in order to achieve effective fusion and interaction between different modalities, these modalities are divided into "target modalities". and "source mode" The target modality is the primary modality from which we want to acquire information, while the source modality is the modality used to assist the target modality in acquiring information.
[0095] The query matrix Q of the target modality m Key matrix with source mode Perform interaction and calculate the cross-modal attention score matrix. ; Based on cross-modal attention scores, the value matrix of the source modal is aggregated. Generate cross-modal contextual features:
[0096] Repeat the above steps for all source modes to obtain multiple sets of cross-modal context features for the target mode m.
[0097] Modal attention weights reflect the most relevant and useful information that the target modality acquires from the source modality. They not only retain important information from the source modality but also optimize for the needs of the target modality. This weighting mechanism makes multimodal data fusion more efficient and targeted, better meeting the information needs of the target modality.
[0098] Through step S502, this embodiment successfully determined the cross-modal attention weights, enabling the target modality to clearly express its need for source modal information, thereby enhancing the ability to understand and utilize data from different modalities.
[0099] In one embodiment of the multimodal knowledge graph construction method in cross-media retrieval of this application, the method may further include the following: Step S601: Perform named entity recognition and relation extraction on the text modal features to determine the corresponding structured semantic representation; Step S602: Analyze the image modal features according to the preset visual semantic parsing model to determine the corresponding visual entities and spatial relationships; Step S603: Based on acoustic event detection and speech semantic understanding, analyze the audio modal features to determine the corresponding high-level semantic labels.
[0100] Optionally, in this embodiment, named entity recognition and relation extraction are performed on the text modal features. Through these two steps, key information in the text can be presented in a structured semantic representation. For example, sentence information can be represented as structured data, which facilitates subsequent processing and application.
[0101] Optionally, in this embodiment, the visual semantic parsing model is trained based on a large amount of image data, and it is able to understand various semantic information in the image.
[0102] First, the model extracts features from the image. These features can be low-level visual features such as color, texture, and shape, or higher-level semantic features. Using these features, the model can identify visual entities in the image. Visual entities can be various objects in the image, such as people, animals, vehicles, and buildings. For example, in an image of a city street, the model can identify different visual entities such as pedestrians, cars, streetlights, and buildings. This entity recognition is achieved through the model learning and matching image features. During training, the model learns the feature patterns of different objects and can find matching regions in new images, thus confirming the presence of these objects.
[0103] Beyond recognizing visual entities, the model also determines the spatial relationships between these entities. Spatial relationships refer to the positional relationships between different objects in an image, such as "to the left of," "above," or "around." For example, in an image of a family living room, the model can identify objects such as a sofa, coffee table, and television, and determine that the coffee table is in front of the sofa, and the television is opposite the sofa, etc. These spatial relationships are determined by analyzing information such as the coordinate positions and relative sizes of objects in the image. Through this visual semantic parsing, the content of an image can be expressed in a semantic form, making the image not merely a collection of pixels, but visual information with clear semantic meaning. This plays an important role in image understanding and image retrieval.
[0104] Optionally, in this embodiment, the processing of audio data involves two aspects: acoustic event detection and speech semantic understanding.
[0105] Acoustic event detection primarily involves identifying various sound events in audio. This is achieved by analyzing the characteristics of the audio signal to extract time-domain features (such as energy and zero-crossing rate) and frequency-domain features (such as spectrum, Mel-frequency cepstral coefficients, and MFCC), and then using these features to determine whether a specific acoustic event exists in the audio.
[0106] Speech semantic understanding, on the other hand, processes the speech portion of audio. When audio contains human speech, the goal of speech semantic understanding is to understand the semantic meaning of what the speaker is saying. This requires first performing speech recognition to convert the speech signal into text, and then using natural language processing techniques to perform semantic analysis on the text.
[0107] Through step S603, this embodiment successfully achieved semantic parsing of multimodal features, laying the foundation for subsequent knowledge graph updates based on semantic associations.
[0108] In one embodiment of the multimodal knowledge graph construction method in cross-media retrieval of this application, the method may further include the following: Step S701: Calculate the semantic similarity between entities of different modalities in the multimodal fusion graph structure based on the latent semantic association, and determine the corresponding cross-modal association weight matrix; Step S702: Solve the cross-modal entity alignment problem of the cross-modal association weight matrix according to the preset graph matching algorithm, and determine the corresponding optimal alignment path; Step S703: Perform semantic alignment on the multimodal fusion graph structure according to the optimal alignment path to determine the corresponding multimodal knowledge graph.
[0109] Optionally, in this embodiment, based on latent semantic association, semantic similarity is calculated for each pair of different modal entities in the multimodal fusion graph structure, and a cross-modal association weight matrix is constructed based on the semantic similarity values. The rows and columns of the matrix correspond to entities of different modalities, and each element in the matrix is the semantic similarity value between the corresponding pair of different modal entities. The matrix intuitively reflects the association strength between entities of different modalities in the multimodal fusion graph structure, providing a precise quantitative basis for subsequent alignment operations.
[0110] Optionally, in this embodiment, after obtaining the cross-modal association weight matrix, the next task is to process this matrix using a preset graph matching algorithm to determine the optimal cross-modal entity alignment path.
[0111] In practice, we treat entities of different modalities as two sets in a bipartite graph, use the elements of the association weight matrix as the edge weights, and iteratively search for augmenting paths using the Hungarian algorithm to gradually construct the optimal matching scheme. This allows us to select the optimal alignment path from numerous possible cross-modal entity alignment schemes.
[0112] For example, in a multimodal data scenario containing text descriptions and image content, this step ensures that the objects mentioned in the text are matched with the corresponding object entities in the images in the best way, laying a solid foundation for the subsequent construction of an accurate multimodal knowledge graph.
[0113] Optionally, in this embodiment, the different modal entities in the multimodal fusion graph structure are reorganized and adjusted according to the optimal alignment path. Guided by the optimal alignment path, different modal entities with high semantic relevance are aligned and connected in the graph structure to form a semantically coherent and structurally clear multimodal knowledge graph.
[0114] For example, if the optimal alignment path indicates that the word "car" in the text is associated with the shape of a "car" in the image and the audio clip of "car horn sound", then in a multimodal knowledge graph, these three entities of different modalities will be closely connected together to form a complete semantic unit.
[0115] Through step S702, this embodiment successfully obtained a multimodal knowledge graph, improving the efficiency and accuracy of cross-media retrieval of multimodal data.
[0116] To improve the efficiency and accuracy of cross-media retrieval of multimodal data, this application provides an embodiment of a multimodal knowledge graph construction apparatus for cross-media retrieval, which implements all or part of the multimodal knowledge graph construction method in the aforementioned cross-media retrieval. See [link to embodiment]. Figure 2 The multimodal knowledge graph construction device in cross-media retrieval specifically includes the following components: The multimodal feature extraction module 10 is used to extract features from the multimodal features respectively, and to map the modal features obtained after the feature extraction to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space. The multimodal fusion graph construction module 20 is used to perform self-attention calculation on each modal feature in the multimodal feature space according to the self-attention mechanism to determine the corresponding intramodal attention weight, perform association calculation on different modal features according to the cross-attention mechanism to determine the corresponding cross-modal attention weight, fuse the intramodal attention weight and the cross-modal attention weight, and weight each modal feature based on the final weight after fusion to determine the corresponding multimodal fusion graph structure; The multimodal knowledge graph determination module 30 is used to perform semantic parsing on the modal features, match the semantic features after semantic parsing with a preset multimodal semantic knowledge base, determine the potential semantic associations between the modal features, perform semantic alignment on the multimodal fusion graph structure according to a preset graph matching algorithm and the potential semantic associations, determine the corresponding multimodal knowledge graph, and perform cross-modal data retrieval based on the multimodal knowledge graph.
[0117] As described above, the multimodal knowledge graph construction device for cross-media retrieval provided in this application embodiment can map the extracted multimodal features to a multimodal feature space through a linear transformation layer. In the multimodal feature space, the intramodal attention weights of each modal feature are calculated according to the self-attention mechanism, and the cross-modal attention weights of different modal features are calculated according to the cross-attention mechanism. The two are fused to obtain the final fusion weight, which is then used to weight each modal feature and input into a graph attention network to obtain a multimodal fusion graph structure. The modal features are semantically parsed and matched with a preset multimodal semantic knowledge base to determine potential semantic associations. The multimodal fusion graph structure is semantically aligned according to a preset graph matching algorithm and potential semantic associations to obtain a multimodal knowledge graph. Cross-modal data retrieval is then performed based on the multimodal knowledge graph, thereby improving the efficiency and accuracy of cross-media multimodal data retrieval.
[0118] From a hardware perspective, in order to improve the efficiency and accuracy of cross-media retrieval of multimodal data, this application provides an embodiment of an electronic device for implementing all or part of the multimodal knowledge graph construction method in the cross-media retrieval, wherein the electronic device specifically includes the following: The system comprises a processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to realize information transmission between the multimodal knowledge graph construction method in cross-media retrieval and core business systems, user terminals, and related databases and other related devices; the logic controller can be a desktop computer, tablet computer, or mobile terminal, etc., and this embodiment is not limited to these. In this embodiment, the logic controller can be implemented with reference to the embodiments of the multimodal knowledge graph construction method in cross-media retrieval in the present embodiment, and the contents of the embodiments are incorporated herein, and repeated parts will not be described again.
[0119] It is understood that the user terminal may include smartphones, tablet computers, network set-top boxes, portable computers, desktop computers, personal digital assistants (PDAs), in-vehicle devices, smart wearable devices, etc. Among these, the smart wearable devices may include smart glasses, smartwatches, smart bracelets, etc.
[0120] In practical applications, parts of the multimodal knowledge graph construction method in cross-media retrieval can be executed on the electronic device side as described above, or all operations can be completed on the client device. The choice can be made based on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations on this. If all operations are completed on the client device, the client device may further include a processor.
[0121] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.
[0122] Figure 3 This is a schematic block diagram illustrating the system configuration of the electronic device 9600 according to an embodiment of this application. Figure 3 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that... Figure 3 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.
[0123] In one embodiment, the multimodal knowledge graph construction method functionality in cross-media retrieval can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following controls: Step S101: Perform feature extraction on the multimodal features respectively, and map the modal features obtained after feature extraction to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space; Step S102: In the multimodal feature space, self-attention calculation is performed on each modal feature according to the self-attention mechanism to determine the corresponding intramodal attention weight. Correlation calculation is performed on different modal features according to the cross-attention mechanism to determine the corresponding cross-modal attention weight. The intramodal attention weight and the cross-modal attention weight are fused, and the modal features are weighted based on the final weight after fusion to determine the corresponding multimodal fusion graph structure. Step S103: Perform semantic parsing on each modal feature, match the semantic features after semantic parsing with a preset multimodal semantic knowledge base, determine the potential semantic associations between each modal feature, perform semantic alignment on the multimodal fusion graph structure according to a preset graph matching algorithm and the potential semantic associations, determine the corresponding multimodal knowledge graph, and perform cross-modal data retrieval based on the multimodal knowledge graph.
[0124] As described above, the electronic device provided in this application's embodiments maps the extracted multimodal features to a multimodal feature space through a linear transformation layer. In the multimodal feature space, intramodal attention weights of each modal feature are calculated based on a self-attention mechanism, and cross-modal attention weights of different modal features are calculated based on a cross-attention mechanism. The two are fused to obtain the final fusion weight, which is then used to weight each modal feature and input into a graph attention network to obtain a multimodal fusion graph structure. Semantic parsing is performed on each modal feature and matched with a preset multimodal semantic knowledge base to determine potential semantic associations. The multimodal fusion graph structure is semantically aligned based on a preset graph matching algorithm and potential semantic associations to obtain a multimodal knowledge graph. Cross-modal data retrieval is performed based on the multimodal knowledge graph, thereby improving the efficiency and accuracy of multimodal data cross-media retrieval.
[0125] In another implementation, the multimodal knowledge graph construction method in cross-media retrieval can be configured separately from the central processing unit 9100. For example, the multimodal knowledge graph construction method in cross-media retrieval can be configured as a chip connected to the central processing unit 9100, and the function of the multimodal knowledge graph construction method in cross-media retrieval can be realized through the control of the central processing unit.
[0126] like Figure 3 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily need to include these components. Figure 3 All components shown; in addition, the electronic device 9600 may also include Figure 3 For components not shown, please refer to existing technologies.
[0127] like Figure 3 As shown, the central processing unit 9100, sometimes also referred to as a controller or operating control, may include a microprocessor or other processor device and / or logic device, which receives inputs and controls the operation of various components of the electronic device 9600.
[0128] The memory 9140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information, and also store a program for executing that information. The central processing unit 9100 may execute the program stored in the memory 9140 to perform information storage or processing, etc.
[0129] Input unit 9120 provides input to central processing unit 9100. Input unit 9120 may be, for example, a keypad or touch input device. Power supply 9170 provides power to electronic device 9600. Display 9160 displays images and text. Display may be, for example, an LCD display, but is not limited thereto.
[0130] The memory 9140 can be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), a SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROMs. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 via the central processing unit 9100.
[0131] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for the electronic device for communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0132] The communication module 9110 is a transmitter / receiver that sends and receives signals via the antenna 9111. The communication module 9110 is coupled to the central processing unit 9100 to provide input signals and receive output signals, which is the same as in a conventional mobile communication terminal.
[0133] Based on different communication technologies, multiple communication modules 9110 can be configured in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication module 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby realizing typical telecommunications functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 9130 is also coupled to a central processing unit 9100, enabling on-device recording via the microphone 9132 and on-device playback of stored sound via the speaker 9131.
[0134] Embodiments of this application also provide a computer-readable storage medium capable of implementing all steps of the multimodal knowledge graph construction method in cross-media retrieval where the execution subject is a server or client, as described in the above embodiments. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all steps of the multimodal knowledge graph construction method in cross-media retrieval where the execution subject is a server or client, as described in the above embodiments. For example, when the processor executes the computer program, it implements the following steps: Step S101: Perform feature extraction on the multimodal features respectively, and map the modal features obtained after feature extraction to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space; Step S102: In the multimodal feature space, self-attention calculation is performed on each modal feature according to the self-attention mechanism to determine the corresponding intramodal attention weight. Correlation calculation is performed on different modal features according to the cross-attention mechanism to determine the corresponding cross-modal attention weight. The intramodal attention weight and the cross-modal attention weight are fused, and the modal features are weighted based on the final weight after fusion to determine the corresponding multimodal fusion graph structure. Step S103: Perform semantic parsing on each modal feature, match the semantic features after semantic parsing with a preset multimodal semantic knowledge base, determine the potential semantic associations between each modal feature, perform semantic alignment on the multimodal fusion graph structure according to a preset graph matching algorithm and the potential semantic associations, determine the corresponding multimodal knowledge graph, and perform cross-modal data retrieval based on the multimodal knowledge graph.
[0135] As described above, the computer-readable storage medium provided in this application's embodiments maps the extracted multimodal features to a multimodal feature space through a linear transformation layer. In the multimodal feature space, intramodal attention weights of each modal feature are calculated based on a self-attention mechanism, and cross-modal attention weights of different modal features are calculated based on a cross-attention mechanism. The two are fused to obtain the final fusion weight, which is then used to weight each modal feature and input into a graph attention network to obtain a multimodal fusion graph structure. Semantic parsing is performed on each modal feature and matched with a preset multimodal semantic knowledge base to determine potential semantic associations. The multimodal fusion graph structure is semantically aligned based on a preset graph matching algorithm and potential semantic associations to obtain a multimodal knowledge graph. Cross-modal data retrieval is performed based on the multimodal knowledge graph, thereby improving the efficiency and accuracy of multimodal data cross-media retrieval.
[0136] Embodiments of this application also provide a computer program product capable of implementing all steps in the multimodal knowledge graph construction method for cross-media retrieval where the execution subject is a server or client, as described in the above embodiments. When executed by a processor, this computer program / instruction implements the steps of the multimodal knowledge graph construction method for cross-media retrieval. For example, the computer program / instruction implements the following steps: Step S101: Perform feature extraction on the multimodal features respectively, and map the modal features obtained after feature extraction to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space; Step S102: In the multimodal feature space, self-attention calculation is performed on each modal feature according to the self-attention mechanism to determine the corresponding intramodal attention weight. Correlation calculation is performed on different modal features according to the cross-attention mechanism to determine the corresponding cross-modal attention weight. The intramodal attention weight and the cross-modal attention weight are fused, and the modal features are weighted based on the final weight after fusion to determine the corresponding multimodal fusion graph structure. Step S103: Perform semantic parsing on each modal feature, match the semantic features after semantic parsing with a preset multimodal semantic knowledge base, determine the potential semantic associations between each modal feature, perform semantic alignment on the multimodal fusion graph structure according to a preset graph matching algorithm and the potential semantic associations, determine the corresponding multimodal knowledge graph, and perform cross-modal data retrieval based on the multimodal knowledge graph.
[0137] As described above, the computer program product provided in this application's embodiments maps the extracted multimodal features to a multimodal feature space through a linear transformation layer. In the multimodal feature space, the intramodal attention weights of each modal feature are calculated based on a self-attention mechanism, and the cross-modal attention weights of different modal features are calculated based on a cross-attention mechanism. The two are fused to obtain the final fusion weight, which is then used to weight each modal feature and input into a graph attention network to obtain a multimodal fusion graph structure. Semantic parsing is performed on each modal feature and matched with a preset multimodal semantic knowledge base to determine potential semantic associations. The multimodal fusion graph structure is semantically aligned based on a preset graph matching algorithm and potential semantic associations to obtain a multimodal knowledge graph. Cross-modal data retrieval is performed based on the multimodal knowledge graph, thereby improving the efficiency and accuracy of multimodal data cross-media retrieval.
[0138] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0140] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0141] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0142] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A method for constructing a multimodal knowledge graph in cross-media retrieval, characterized in that, The method includes: Feature extraction is performed on the multimodal features separately, and the modal features obtained after feature extraction are mapped to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space. The modal features obtained after feature extraction include text modal features, image modal features, and audio modal features. In the multimodal feature space, self-attention is calculated for each modal feature according to the self-attention mechanism to determine the corresponding intramodal attention weights. Correlation calculation is performed for different modal features according to the cross-attention mechanism to determine the corresponding cross-modal attention weights. The intramodal attention weights and the cross-modal attention weights are fused, and the modal features are weighted based on the final weights after fusion. The weighted modal features are then input into the graph attention network to determine the corresponding multimodal fusion graph structure. Semantic parsing is performed on each modal feature. The semantically parsed features are then matched with a preset multimodal semantic knowledge base to determine the potential semantic associations between the modal features. The multimodal fusion graph structure is then semantically aligned according to a preset graph matching algorithm and the potential semantic associations. The step of semantically aligning the multimodal fusion graph structure according to the preset graph matching algorithm and the potential semantic associations includes: calculating the semantic similarity between different modal entities in the multimodal fusion graph structure based on the potential semantic associations to determine the corresponding cross-modal association weight matrix; solving for cross-modal entity alignment in the cross-modal association weight matrix using the preset graph matching algorithm to determine the corresponding optimal alignment path; performing semantic alignment on the multimodal fusion graph structure according to the optimal alignment path; and finally determining the corresponding multimodal knowledge graph. Cross-modal data retrieval is then performed based on the multimodal knowledge graph.
2. The method for constructing a multimodal knowledge graph in cross-media retrieval according to claim 1, characterized in that, The feature extraction for the multimodal features includes: Semantic extraction is performed on text modal data based on a pre-trained language model to determine the corresponding text modal features, wherein the text modal features include vocabulary, grammatical structure, and global semantic features; The image modal data is subjected to multi-layer convolution and pooling operations based on a preset convolutional neural network to determine the corresponding image modal features, wherein the image modal features include local image features and global image features; Acoustic features are extracted from audio modal data according to a preset time-frequency transformation algorithm, and semantic features are extracted from the acoustic features obtained by the acoustic feature extraction according to a preset temporal modeling network to determine the corresponding audio modal features.
3. The method for constructing a multimodal knowledge graph in cross-media retrieval according to claim 2, characterized in that, The step of mapping the modal features obtained after feature extraction to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space includes: The text modal features, image modal features, and audio modal features are normalized, and the normalized modal features are projected onto a shared feature space of the same dimension according to the learnable weight matrix. In the shared feature space, activation functions are applied to each modal feature after projection to determine the corresponding multimodal feature space.
4. The method for constructing a multimodal knowledge graph in cross-media retrieval according to claim 1, characterized in that, The step of performing self-attention calculation on each modal feature according to the self-attention mechanism to determine the corresponding intra-modal attention weights includes: Matrix calculations are performed on each modal feature using trainable parameters to determine the query vector, key vector, and value vector corresponding to each modal feature. Self-attention is calculated for feature elements within the same modality to determine the corresponding same-modality attention score; The corresponding intramodal attention weights are determined by weighting the value vector of the modality based on the intramodal attention scores.
5. The method for constructing a multimodal knowledge graph in cross-media retrieval according to claim 4, characterized in that, The step of performing correlation calculations on features of different modalities based on the cross-attention mechanism to determine the corresponding cross-modal attention weights includes: Different modalities are divided into target modalities and source modalities. The query vector of the target modalities is interacted with the key vector of the source modalities to determine the corresponding cross-modal attention scores. The cross-modal attention weights of the target modality are determined by weighting the value vector of the source modality based on the cross-modal attention scores.
6. The method for constructing a multimodal knowledge graph in cross-media retrieval according to claim 2, characterized in that, The semantic parsing of the modal features includes: Named entity recognition and relation extraction are performed on text modal features to determine the corresponding structured semantic representation; The image modal features are analyzed based on the preset visual semantic parsing model to determine the corresponding visual entities and spatial relationships; Audio modal features are analyzed based on acoustic event detection and speech semantic understanding to determine the corresponding high-level semantic labels.
7. A multimodal knowledge graph construction device for cross-media retrieval, characterized in that, The device includes: The multimodal feature extraction module is used to extract features from multimodal features respectively, and to map the modal features obtained after feature extraction to a unified feature space through a linear transformation layer to determine the corresponding multimodal feature space. The multimodal fusion graph construction module is used to perform self-attention calculation on each modal feature in the multimodal feature space according to the self-attention mechanism to determine the corresponding intramodal attention weights, perform association calculation on different modal features according to the cross-attention mechanism to determine the corresponding cross-modal attention weights, fuse the intramodal attention weights and the cross-modal attention weights, and weight each modal feature based on the final weight after fusion, and input the weighted modal features into the graph attention network to determine the corresponding multimodal fusion graph structure; A multimodal knowledge graph determination module is used to perform semantic parsing on the modal features, match the semantically parsed features with a preset multimodal semantic knowledge base to determine the potential semantic associations between the modal features, and perform semantic alignment on the multimodal fusion graph structure according to a preset graph matching algorithm and the potential semantic associations. The step of semantically aligning the multimodal fusion graph structure according to the preset graph matching algorithm and the potential semantic associations includes: calculating the semantic similarity between different modal entities in the multimodal fusion graph structure according to the potential semantic associations to determine the corresponding cross-modal association weight matrix; solving the cross-modal entity alignment problem in the cross-modal association weight matrix according to the preset graph matching algorithm to determine the corresponding optimal alignment path; performing semantic alignment on the multimodal fusion graph structure according to the optimal alignment path, and finally determining the corresponding multimodal knowledge graph; and performing cross-modal data retrieval based on the multimodal knowledge graph.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the multimodal knowledge graph construction method in cross-media retrieval as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method for constructing a multimodal knowledge graph in cross-media retrieval as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent search method and system based on multi-source heterogeneous data
CN116049454A
Multi-modal image automatic labeling system and method
CN120411667A