Knowledge fusion retrieval method and system based on multi-modal large model
By optimizing node representation through cross-modal encoders and multi-head self-attention mechanisms, the problems of inaccurate alignment and information loss in multimodal data are solved, achieving efficient multimodal information fusion and knowledge graph construction, and improving the accuracy and efficiency of query retrieval.
Patent Information
- Application Number
- CN202510739839.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-10-31
AI Technical Summary
Existing multimodal information processing methods are insufficient in terms of the accuracy of cross-modal alignment and information fusion, making it difficult to effectively integrate the complex relationships between multiple modalities such as images, text, and speech. In particular, when processing complex multimodal data, the alignment is inaccurate and information loss is severe.
A cross-modal encoder is used to construct node vectors for cross-modal alignment. The node representation is optimized by combining a multi-head self-attention mechanism and a variational hybrid expert network. The node representation is updated by an attention weighting mechanism, and a knowledge graph is constructed. The BERT model is used to understand the user query content and form query records for storage.
It improves the accuracy of node alignment and the effect of multimodal information fusion, enhances the accuracy and efficiency of knowledge graph construction and query retrieval, and strengthens the representation ability and query accuracy of cross-modal information.
Smart Images

Figure CN120873246A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graph retrieval technology, and in particular to a knowledge fusion retrieval method and system based on a multimodal large model. Background Technology
[0002] With the rapid development of artificial intelligence and big data technologies, multimodal information processing has become one of the important research directions in the field of artificial intelligence. Multimodal information refers to data with different forms of expression obtained from multiple different data sources, such as text, images, and speech. In recent years, with the continuous progress of deep learning, especially large models and neural networks, multimodal artificial intelligence technologies have achieved remarkable results in fields such as natural language processing, computer vision, and speech recognition. However, existing multimodal information processing still has shortcomings in terms of the accuracy of cross-modal alignment and information fusion. When processing more complex multimodal data, it still faces problems such as inaccurate alignment and serious information loss, making it difficult to effectively integrate the complex relationships between multiple modalities such as images, text, and speech. Summary of the Invention
[0003] In view of the aforementioned existing problems, the present invention is proposed.
[0004] Therefore, this invention provides a knowledge fusion retrieval method and system based on a multimodal large model, which solves the problems of inaccurate alignment and severe information loss that existing technologies still face when processing more complex multimodal data, and the difficulty in effectively integrating the complex relationships between multiple modalities such as images, text and speech.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: Firstly, this invention provides a knowledge fusion retrieval method based on a multimodal large model, which includes: Multimodal information is acquired and text features, image features and speech features are extracted respectively. Cross-modal encoders are used to perform cross-modal alignment to construct node vectors and form a knowledge graph. Construct a multi-relation embedding representation between nodes, and calculate the multi-relation awareness message between nodes to update the node representation through an attention weighting mechanism. Simultaneously, use a variational hybrid expert network to approximate the potential distribution of node vectors and fuse them to obtain a fused node representation. Concatenate the updated node representation and the fused node representation to obtain an optimized node representation. A multi-head self-attention mechanism is used to further optimize the node representation and output the final node representation. Based on the user query, the query results are matched, and the query results are displayed and stored as query records. The multimodal information includes text information, image information, and voice information.
[0006] As a preferred embodiment of the knowledge fusion retrieval method based on a multimodal large model described in this invention, the step of using a cross-modal encoder to perform cross-modal alignment and construct node vectors to form a knowledge graph refers to using a multilayer perceptron to construct cross-modal encoders for text features, image features, and speech features respectively. , as well as The text features, image features, and speech features are converted into text embeddings through a cross-modal encoder. Image embedding and voice embedding ; The same number of text features, image features, and speech features are obtained from the training data as a training set and labeled. A contrastive loss function is constructed and the cross-modal encoder is trained by gradient descent. The extracted text features, image features, and speech features are respectively input into the trained cross-modal encoder to obtain the text embedding. Image embedding and voice embedding And each embedding is treated as a node vector. Nodes are formed, the cosine similarity of node vectors is calculated, and node pairs with similarity greater than a set threshold are connected to form connecting edges and a knowledge graph is constructed.
[0007] As a preferred embodiment of the knowledge fusion retrieval method based on a multimodal large model described in this invention, the following steps are included: constructing a multi-relation embedding representation between nodes, calculating multi-relation-aware messages between nodes and updating node representations through an attention weighting mechanism, and simultaneously using a variational hybrid expert network to approximate the latent distribution of node vectors and fusing them to obtain a fused node representation. This refers to constructing a multi-relation embedding representation between nodes through multi-relation embedding. ; Multi-relation embedding representation is used to compute multi-relation-aware messages between nodes. ; Calculate the attention weights between nodes that have connecting edges. ; The node representation is updated using attention-weighted scaling based on the attention weights between nodes. ; A variational hybrid expert network is constructed, including an expert network, a variational information bottleneck, and a hybrid expert mechanism. All node vectors are input into the expert network to calculate the mean vector and standard deviation vector of the node vectors, and variational inference is used to approximate the latent distribution Z. Assign activation weights G to each expert network; The fusion node representation H is obtained by weighted summation of the potential distributions of each expert network output by activating weights.
[0008] As a preferred embodiment of the knowledge fusion retrieval method based on a multimodal large model described in this invention, the following steps are taken: The optimized node representation is further optimized using a multi-head self-attention mechanism to output the final node representation. Based on the user query and matching query results, the optimized node representation is mapped to a query vector, key vector, and value vector through a linear transformation. The multi-head self-attention mechanism obtains the attention weights between the optimized node representations through a softmax operation, and the value vector is weighted and summed using the attention weights to obtain the weighted representation of each attention head. The weighted representations of each attention head are then concatenated and mapped using a convolutional neural network to obtain the final node representation. The BERT model is used to convert user query content into an embedding vector X. The cosine similarity between the embedding vector X and the final node representation of each node is calculated, and the node with the highest similarity is selected as the final user query result.
[0009] As a preferred embodiment of the knowledge fusion retrieval method based on multimodal large model described in this invention, wherein: obtaining multimodal information and extracting text features, image features and speech features respectively refers to preprocessing the text information, image information and speech information in the multimodal information respectively after obtaining the multimodal information; The text information is mapped into text features through the BERT model; The image information is used to extract image features through a convolutional neural network; The speech features are extracted using MFCC and deep neural networks.
[0010] As a preferred embodiment of the knowledge fusion retrieval method based on a multimodal large model described in this invention, the step of displaying the query results refers to displaying the node content in the user's query results to the user, collecting user feedback and evaluation, and notifying staff of the user feedback and evaluation as a query evaluation.
[0011] As a preferred embodiment of the knowledge fusion retrieval method based on a multimodal large model described in this invention, the step of forming and storing query records refers to forming query records from user query content and query results and storing them, and binding and storing the knowledge graph and query records synchronously.
[0012] Secondly, this invention provides a knowledge fusion retrieval system based on a multimodal large model, including: The knowledge graph construction module is used to collect multimodal information and extract features, and uses a cross-modal encoder to construct node vectors to form a knowledge graph; The node optimization module is used to optimize the node vectors using an attention mechanism and a variational hybrid expert network, and then further optimizes them using a multi-head attention mechanism to obtain the final node representation. The query storage module is used to match user query results with the knowledge graph, display the query results, and finally store the query records.
[0013] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the knowledge fusion retrieval method based on a multimodal large model as described in the first aspect of the present invention.
[0014] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the knowledge fusion retrieval method based on a multimodal large model as described in the first aspect of the present invention.
[0015] The beneficial effects of this invention are as follows: By collecting multimodal information and extracting features, this invention uses a cross-modal encoder to perform cross-modal alignment to construct node vectors, which effectively improves the accuracy of node alignment. Furthermore, by combining a self-attention mechanism with a variational hybrid expert network to optimize node vectors, the effect of multimodal information fusion optimization is further enhanced, thereby improving the accuracy and efficiency of knowledge graph construction and query retrieval. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of the knowledge fusion retrieval method based on a multimodal large model in Example 1.
[0018] Figure 2 This is a structural diagram of the knowledge fusion retrieval system based on a multimodal large model in Example 1. Detailed Implementation
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0022] Example 1, referring to Figure 1 and Figure 2 This is the first embodiment of the present invention, which provides a knowledge fusion retrieval method based on a multimodal large model, including the following steps: S1. Obtain multimodal information and extract text features, image features and speech features respectively. Use a cross-modal encoder to perform cross-modal alignment to construct node vectors and form a knowledge graph. The multimodal information includes text information, image information, and voice information; Specifically, acquiring multimodal information and extracting text features, image features, and speech features separately refers to preprocessing the text information, image information, and speech information in the multimodal information separately after acquiring the multimodal information; The text information is mapped into text features through the BERT model; As a pre-trained language model, BERT can understand the context of text and overcome the surface grammar matching problem in traditional text processing methods. It can capture the polysemy and subtle differences of words through contextual information, providing a more accurate text representation for subsequent knowledge reasoning. The image information is used to extract image features through a convolutional neural network; CNNs can automatically learn key image features from raw image data, avoiding the difficulty of manually designing features. Their convolutional layers can effectively extract low-level features such as edges and textures, while pooling layers further enhance the spatial invariance of the image, making the image representation more robust. The speech features are extracted using MFCC and deep neural networks.
[0023] MFCC, as a classic feature in speech processing, can effectively extract the spectral information of speech signals, reflecting the tone quality and timbre characteristics of speech. By further processing these features through deep neural networks, deeper speech features can be extracted, improving the speech representation capability.
[0024] Furthermore, cross-modal encoders are used for cross-modal alignment to construct node vectors and form a knowledge graph. This involves using multilayer perceptrons to construct cross-modal encoders for text features, image features, and speech features respectively. , as well as The text features, image features, and speech features are converted into text embeddings through a cross-modal encoder. Image embedding and voice embedding : ; in For text features, For image features, For speech features; Multilayer perceptron (MLP) is used to align features from different modalities. The MLP maps input text, image, and speech features to a common embedding space through nonlinear transformation, enabling information from different modalities to be compared and fused within this space. The same number of text features, image features, and speech features are obtained from the training data and labeled. A contrastive loss function is constructed, and the cross-modal encoder is trained using gradient descent. for: ; Where m and n are any two of the text features, image features, and speech features. and For the i-th pair of feature embeddings, including text embeddings Image embedding and voice embedding Obtained through annotation. For the embedding of the j-th feature, This is the sensitivity coefficient; The contrastive loss function optimizes the cross-modal encoder by maximizing the cosine similarity of positive samples (matching text, image, and speech pairs) and minimizing the similarity between other mismatched sample pairs. By maximizing the similarity of matching samples and minimizing the similarity of non-matching samples, accurate cross-modal alignment is achieved. Through contrastive learning, the representations of text, images, and speech can be optimized into a common semantic space, making it easier for subsequent tasks (such as knowledge graph construction and query retrieval) to find relevant nodes in high-dimensional space. By using a contrastive loss function, the model can significantly improve the alignment accuracy between modalities, thereby reducing information loss in traditional methods and avoiding semantic conflicts between modalities. The extracted text features, image features, and speech features are respectively input into the trained cross-modal encoder to obtain the text embedding. Image embedding and voice embedding And each embedding is treated as a node vector. Nodes are formed, each containing a node vector and corresponding node modality information. The cosine similarity of the node vectors is calculated, and node pairs with similarities greater than a set threshold are connected to form connection edges and a knowledge graph is constructed.
[0025] Multilayer perceptrons can effectively align different modalities through nonlinear transformations, avoiding the misalignment or information loss between modalities in traditional methods. Through an optimized contrastive loss function, the model can significantly improve the alignment between text, image, and speech modalities during training, thereby enhancing the accuracy of the entire knowledge graph construction and retrieval task. The representation of different modal features (text, image, and speech) in the same semantic space can help subsequent knowledge reasoning and retrieval to be more accurate. Through multimodal alignment and accurate knowledge graph construction, queries can be efficiently matched based on multimodal features, improving query accuracy.
[0026] S2. Construct a multi-relation embedding representation between nodes, and calculate the multi-relation awareness message between nodes. Update the node representation through an attention weighting mechanism, and simultaneously use a variational hybrid expert network to approximate the potential distribution of the node vectors and fuse them to obtain a fused node representation. Concatenate the updated node representation and the fused node representation to obtain the optimized node representation. Specifically, a multi-relation embedding representation between nodes is constructed, and the node representation is updated through an attention weighting mechanism by calculating the multi-relation-aware messages between nodes. Simultaneously, a variational hybrid expert network is used to approximate the latent distribution of node vectors and fuse them to obtain a fused node representation. This multi-relation embedding representation between nodes is constructed through multi-relation embedding. : ; in It is a 2-layer MLP used to embed relation and node type into a unified relation embedding space. and Let i and j be the node vectors, respectively. The edge connecting node i and node j is represented using one-hot encoding; Multi-relation embedding representation is used to compute multi-relation-aware messages between nodes. : ; in For linear transformation, node representation and relation embedding are combined to generate messages; Calculate the attention weights between nodes that have connecting edges. : ; in Let i be the query vector for node i. Let be the key vector of node j, and D be the dimension of the embedding space. Let Q be the index of the neighboring node of node i, and Q is the transpose operation; The node representation is updated using attention-weighted scaling based on the attention weights between nodes. : ; in Let k be the node representation of the k-th layer, initially set to , It is a 2-layer MLP used to transform the weighted message and update the node representation; By using attention weighting, the model can prioritize the transmission of important information when updating node representations, thereby achieving more accurate knowledge reasoning. Especially when dealing with large-scale, multimodal data, the attention mechanism can help the model effectively filter out the most valuable features, avoiding the interference of redundant or irrelevant information on the final result. A variational hybrid expert network is constructed, comprising expert networks, a variational information bottleneck, and a hybrid expert mechanism. The expert networks are feedforward neural networks (FFNNs), with each expert network corresponding to a subnetwork responsible for extracting specific features from the input data. The variational information bottleneck is used to optimize information propagation, ensuring that each modal input is mapped to a latent low-dimensional space through variational inference, while retaining sufficient information to support downstream tasks. The hybrid expert mechanism determines which experts are activated through a routing mechanism, thereby dynamically adjusting the computational load so that each input has a suitable subnetwork for processing. The mean and standard deviation vectors of all node vectors are calculated by inputting them into the expert networks, and the latent distribution Z is approximated using variational inference. ; in Given all the input node vectors, For the parameters of the feedforward neural network, It is the mean vector. The standard deviation vector, The potential distribution Z is a Gaussian distribution; During training, the approximation of the latent distribution is performed through variational inference, ensuring that the distribution of node representations in the latent space is close to the true distribution. Furthermore, each expert network makes a different response to the input data. In this way, the model can efficiently process different modalities and complex data, thereby improving the expressive power of the multimodal knowledge graph. An activation weight G is assigned to each expert network through a dynamic routing mechanism: ; in and For training parameters; The introduction of dynamic routing mechanism enables each expert network to flexibly select when processing different inputs, improving computational efficiency and model adaptability. By dynamically allocating activation weights, the model can select the most suitable expert for different modal inputs, thereby avoiding the impact of irrelevant information on model performance. The fusion node representation H is obtained by weighted summation of the latent distributions of each expert network output by activating weights: ; in Let i be the potential distribution of the output of the i-th expert network. Let K be the activation weight of the i-th expert network, and K be the total number of expert networks.
[0027] Multi-relation embedding representations between nodes enable the alignment and expression of relationships between different modalities (such as text, image, speech, etc.) in the same semantic space. By using multilayer perceptron (MLP) for mapping, a tighter and more consistent embedding representation is created for each node. This not only accurately captures the semantic features of each node, but also ensures efficient learning of relationships between nodes through weighted combination of relation embeddings. Through the computation and transmission of multi-relation-aware messages, nodes can propagate and update information with the support of multilayer graph neural networks, forming a structured and interconnected knowledge graph. This effectively aligns nodes of different modalities and optimizes node representations through information flow, thereby improving the accuracy and complexity of graph construction.
[0028] S3. The multi-head self-attention mechanism is used to further optimize the node representation and output the final node representation. Based on the user query and the matching query results, the query results are displayed and the query records are stored. Specifically, a multi-head self-attention mechanism is used to further optimize the node representation and output the final node representation. Based on the user query and matching query results, the optimized node representation is mapped to query vector, key vector, and value vector through linear transformation. The multi-head self-attention mechanism is used to obtain the attention weights between the optimized node representations through softmax operation, and the attention weights are used to perform a weighted summation on the value vector to obtain the weighted representation of each attention head. The weighted representations of each attention head are concatenated and then mapped using a convolutional neural network to obtain the final node representation. The BERT model is used to convert user query content into an embedding vector X. The cosine similarity between the embedding vector X and the final node representation of each node is calculated, and the node with the highest similarity is selected as the final user query result.
[0029] Multi-head self-attention mechanisms focus on the most relevant information between nodes, rather than relying on global information, thereby enhancing node performance in tasks. Nodes dynamically adjust their weights based on their relevance to the query during information flow, making information transmission more flexible and particularly suitable for complex cross-modal retrieval tasks. The introduction of multi-head attention mechanisms helps improve the accuracy of node representations in query matching, especially in multimodal data, avoiding feature loss and information redundancy problems that may occur in traditional methods. CNNs can extract deeper feature information from the outputs of multiple attention heads, thereby improving the expressive power of node representations. CNNs can not only effectively fuse the outputs of multi-head self-attention mechanisms, but also perform deeper combination and understanding of information from multiple modalities, improving the representation ability of cross-modal information. Through convolution operations, CNNs can compress information dimensions, reduce unnecessary computation, and improve the training and inference speed of the model. By accurately understanding user queries, BERT can help generate query results that better meet user needs, improving user experience. BERT can understand the contextual information in queries and maintain a high level of understanding in multi-level and complex queries.
[0030] Furthermore, displaying the query results refers to showing the node content in the user's query results to the user, collecting user feedback and evaluations, and notifying staff of the user feedback and evaluations, so that staff can optimize the query process based on the evaluations.
[0031] Furthermore, storing query records refers to storing user query content and query results as query records, and binding and synchronizing the knowledge graph with the query records.
[0032] This embodiment also provides a knowledge fusion retrieval system based on a multimodal large model, including: The knowledge graph construction module is used to collect multimodal information and extract features, and uses a cross-modal encoder to construct node vectors to form a knowledge graph; The node optimization module is used to optimize the node vectors using an attention mechanism and a variational hybrid expert network, and then further optimizes them using a multi-head attention mechanism to obtain the final node representation. The query storage module is used to match user query results with the knowledge graph, display the query results, and finally store the query records.
[0033] This embodiment also provides a computer device applicable to the knowledge fusion retrieval method based on a multimodal large model, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the knowledge fusion retrieval method based on a multimodal large model as proposed in the above embodiment.
[0034] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0035] This embodiment also provides a storage medium on which a computer program is stored. When executed by a processor, the program implements the knowledge fusion retrieval method based on a multimodal large model as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0036] In summary, this invention effectively improves the accuracy of node alignment by collecting multimodal information and extracting features, and constructing node vectors through cross-modal alignment using a cross-modal encoder. Furthermore, by optimizing node vectors through a self-attention mechanism combined with a variational hybrid expert network, the effect of multimodal information fusion optimization is further enhanced, thereby improving the accuracy and efficiency of knowledge graph construction and query retrieval.
[0037] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A knowledge fusion retrieval method based on a multimodal large model, characterized by: include, Multimodal information is acquired and text features, image features and speech features are extracted respectively. Cross-modal encoders are used to perform cross-modal alignment to construct node vectors and form a knowledge graph. Construct a multi-relation embedding representation between nodes, and calculate the multi-relation awareness message between nodes to update the node representation through an attention weighting mechanism. Simultaneously, use a variational hybrid expert network to approximate the potential distribution of node vectors and fuse them to obtain a fused node representation. Concatenate the updated node representation and the fused node representation to obtain an optimized node representation. A multi-head self-attention mechanism is used to further optimize the node representation and output the final node representation. Based on the user query, the query results are matched, and the query results are displayed and stored as query records. The multimodal information includes text information, image information, and voice information.
2. The knowledge fusion retrieval method based on a multimodal large model as described in claim 1, characterized in that: The method of using a cross-modal encoder for cross-modal alignment to construct node vectors and form a knowledge graph refers to using a multilayer perceptron to construct cross-modal encoders for text features, image features, and speech features respectively. , as well as The text features, image features, and speech features are converted into text embeddings through a cross-modal encoder. Image embedding and voice embedding ; The same number of text features, image features, and speech features are obtained from the training data as a training set and labeled. A contrastive loss function is constructed and the cross-modal encoder is trained by gradient descent. The extracted text features, image features, and speech features are respectively input into the trained cross-modal encoder to obtain the text embedding. Image embedding and voice embedding And each embedding is treated as a node vector. Nodes are formed, the cosine similarity of node vectors is calculated, and node pairs with similarity greater than a set threshold are connected to form connecting edges and a knowledge graph is constructed.
3. The knowledge fusion retrieval method based on a multimodal large model as described in claim 2, characterized in that: The process involves constructing a multi-relation embedding representation between nodes, calculating multi-relation-aware messages between nodes, updating node representations through an attention-weighted mechanism, and simultaneously using a variational hybrid expert network to approximate the latent distribution of node vectors and fusing them to obtain a fused node representation. This process refers to constructing a multi-relation embedding representation between nodes through multi-relation embedding. ; Multi-relation embedding representation is used to compute multi-relation-aware messages between nodes. ; Calculate the attention weights between nodes that have connecting edges. ; The node representation is updated using attention-weighted scaling based on the attention weights between nodes. ; A variational hybrid expert network is constructed, including an expert network, a variational information bottleneck, and a hybrid expert mechanism. All node vectors are input into the expert network to calculate the mean vector and standard deviation vector of the node vectors, and variational inference is used to approximate the latent distribution Z. Assign activation weights G to each expert network; The fusion node representation H is obtained by weighted summation of the potential distributions of each expert network output by activating weights.
4. The knowledge fusion retrieval method based on a multimodal large model as described in claim 3, characterized in that: The process involves using a multi-head self-attention mechanism to further optimize the node representation and output the final node representation. Based on the user query and matching query results, the optimized node representation is mapped to a query vector, key vector, and value vector through a linear transformation. The multi-head self-attention mechanism uses a softmax operation to obtain the attention weights between the optimized node representations and then uses the attention weights to perform a weighted summation on the value vectors to obtain the weighted representation of each attention head. Finally, the weighted representations of each attention head are concatenated and mapped using a convolutional neural network to obtain the final node representation. The BERT model is used to convert user query content into an embedding vector X. The cosine similarity between the embedding vector X and the final node representation of each node is calculated, and the node with the highest similarity is selected as the final user query result.
5. The knowledge fusion retrieval method based on a multimodal large model as described in claim 4, characterized in that: The process of acquiring multimodal information and extracting text features, image features, and speech features refers to preprocessing the text information, image information, and speech information in the multimodal information after acquiring it. The text information is mapped into text features through the BERT model; The image information is used to extract image features through a convolutional neural network; The speech features are extracted using MFCC and deep neural networks.
6. The knowledge fusion retrieval method based on a multimodal large model as described in claim 5, characterized in that: Displaying the query results refers to showing the node content in the user's query results to the user, collecting user feedback and evaluation, and using the user feedback and evaluation as a query evaluation notification to the staff.
7. The knowledge fusion retrieval method based on a multimodal large model as described in claim 6, characterized in that: The process of forming and storing query records refers to storing user query content and query results as query records, and binding and storing the knowledge graph and query records synchronously.
8. A knowledge fusion retrieval system based on a multimodal large model, based on the knowledge fusion retrieval method based on a multimodal large model as described in any one of claims 1 to 7, characterized in that: include, The knowledge graph construction module is used to collect multimodal information and extract features, and uses a cross-modal encoder to construct node vectors to form a knowledge graph; The node optimization module is used to optimize the node vectors using an attention mechanism and a variational hybrid expert network, and then further optimizes them using a multi-head attention mechanism to obtain the final node representation. The query storage module is used to match user query results with the knowledge graph, display the query results, and finally store the query records.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the knowledge fusion retrieval method based on a multimodal large model as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the knowledge fusion retrieval method based on a multimodal large model as described in any one of claims 1 to 7.
Citation Information
Cited By
Property information cross-modal retrieval method and system based on knowledge graph
CN121350283A
Method and system for cross-modal retrieval of property information based on knowledge graph
CN121350283B