A Generative Dialogue Method and System Based on Multimodal Knowledge Enhancement
Through multimodal fusion and external knowledge enhancement methods, the problem of insufficient utilization of image information and knowledge in the existing dialogue system is solved, and more reasonable dialogue reply generation is achieved, improving the application effect of the system.
Patent Information
- Application Number
- CN202310303970.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-27
AI Technical Summary
The existing dialogue system fails to effectively integrate image information and external knowledge in real scenarios, resulting in unreasonable dialogue responses and ignore the important role of knowledge in response generation.
Through multimodal fusion method and external knowledge enhancement, the attention mechanism of visual semantic features and global text features is enhanced, and dialogue replies are generated in combination with BART encoder and decoder to achieve preliminary fusion of images and text, and attention weight allocation is used to use knowledge feature vectors.
More reasonable dialogue replies were generated, improving the scope of application of dialogue systems in real scenarios and the effectiveness of response generation.
Smart Images

Figure CN116450787B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing dialogue generation, and particularly relates to a generative dialogue method and system based on multi-modal knowledge enhancement. Background Art
[0002] Dialogue is one of the main ways for humans to communicate and conduct daily activities. Through dialogue, we can book train tickets for travel, discuss work with colleagues, and chat with family and friends to communicate feelings.
[0003] In natural language processing tasks, how to develop an intelligent dialogue system that can communicate with humans naturally is a long-term research goal and also a complex task.
[0004] With the development of artificial intelligence technology, the application of dialogue systems in current society is becoming more and more common. Many intelligent assistants (such as Alibaba Xiaomi, Microsoft Xiaoice, etc.) can give corresponding responses according to user input, which not only meets the needs of users but also saves a large amount of manpower and resources.
[0005] However, most of such dialogue systems are only applied in a single modality (i.e., the text modality), ignoring some important image information during the dialogue process. For example, when a user is consulting about shopping, sometimes they will use images to describe their needs instead. Therefore, there is certain research value in task-oriented multi-modal dialogue systems.
[0006] Although some existing works have proposed using state-of-the-art multi-modal models to generate multi-modal dialogues, they mainly focus on question-and-answer style dialogue generation based on a single image, rather than each image in each dialogue turn. This learning paradigm limits the application scope of multi-modal dialogue generation models in real-world scenarios, where dialogues occur in a visually changing context over time.
[0007] At the same time, due to the development of the knowledge graph field and the improvement of publicly available encyclopedic knowledge bases on the Internet, in recent years, some researchers have also tried to generate responses by introducing information provided in the knowledge graph to promote the progress of the dialogue. However, most of the research stops at the early-stage integration of knowledge and does not directly act on response generation. Based on this research status, the current challenges mainly include the following two aspects:
[0008] (1) In real-world scenarios, dialogues do not only revolve around a single image, and the relevance of each dialogue to the image is also different. That is to say, the influence degree of the information in the image on the dialogue content is different, and it is necessary to explore a more effective multi-modal fusion method to generate more reasonable responses;
[0009] (2) In task-oriented dialogue systems, most of the research using external knowledge bases only performs early-stage fusion and does not explicitly use knowledge to generate responses. Such an approach ignores the important role of knowledge in response generation. Therefore, it is necessary to explore the important role of external knowledge in response generation. Summary of the Invention
[0010] The present invention provides a generative dialogue method based on multi-modal knowledge enhancement, which can achieve more reasonable dialogue responses through multi-modal fusion and by incorporating external knowledge.
[0011] A generative dialogue method based on multi-modal knowledge enhancement, comprising:
[0012] Obtain multi-modal context, where the multi-modal context includes text context and image context. Obtain a knowledge sequence that matches the text context, respectively perform vector representations on the knowledge sequence and the text context to obtain a knowledge feature vector and a text feature vector, and at the same time perform feature extraction on the image context to obtain visual semantic features; fuse the knowledge feature vector and the text feature vector and then obtain global text features through a text encoder;
[0013] Use the global text features to perform attention weighting on the visual semantic features to obtain enhanced visual semantic features, use the visual semantic features to perform attention weighting on the global text features to obtain enhanced global text features, splice the enhanced visual semantic features and the enhanced global text features, pass the splicing result through a fully connected layer to obtain multi-modal context features, and use the knowledge feature vector to perform attention weighting on the multi-modal context features to obtain enhanced multi-modal context features;
[0014] Use a BART encoder to encode the enhanced multi-modal context features to obtain high-order features, use a BART decoder to decode the high-order features in an autoregressive manner to obtain hidden state vectors, convert the dimension of the hidden state vectors to the vocabulary dimension of the BART pre-trained language model, normalize the converted-dimension hidden state vectors and then map them to the probability space to obtain the prediction probabilities of words in the vocabulary, and generate a dialogue based on the prediction probabilities of the words.
[0015] Further, the use of the global text features to perform attention weighting on the visual semantic features to obtain enhanced visual semantic features includes:
[0016] Based on the similarity between the word vectors of each visual semantic feature and all the word vectors of the global text features, obtain visual semantic feature attention weights, and assign corresponding visual semantic feature attention weights to the word vectors of each visual semantic feature to obtain enhanced visual semantic features.
[0017] Further, the visual semantic feature attention weight is obtained based on the similarity between the word vectors of each visual semantic feature and all the word vectors of the global text feature, including:
[0018] The word vectors of each visual semantic feature are respectively matched with the word vectors of the global text feature by using the scaled dot product method to obtain a first similarity vector set, and the first similarity vector set is summed and then normalized to obtain the visual semantic feature attention weight of the corresponding word vector of the visual semantic feature.
[0019] Further, the obtaining of the knowledge sequence matching the text context includes:
[0020] Use an NLP tool to perform named entity recognition on the text context to obtain an entity name sequence, obtain a knowledge base, which includes multiple knowledge entities, each knowledge entity includes multiple knowledge attributes, compare the cosine similarity between the word vectors of the knowledge entities in the entity name sequence and the word vectors of the knowledge entities in the knowledge base, retain the knowledge entities with the top K similarities, and integrate the knowledge attributes corresponding to the knowledge entities with the top K similarities to obtain a knowledge sequence.
[0021] Further, the knowledge sequence and the text context are respectively vectorized through an embedding layer to obtain a knowledge feature vector and a text feature vector; a ViT-B / 32 module pre-trained based on the CLIP model is used to extract the features of the image context to obtain visual semantic features.
[0022] Further, the knowledge feature vector and the text feature vector are fused and then encoded to obtain the global text feature T t as:
[0023] T t = Text_Encoder(E h + E k ) = {e t1 , e t2 , … e ti …, e tI}
[0024] where, Text_Encoder is a text encoder, e ti is the i-th word vector in the global text feature, I is the number of word vectors in the global text feature, E h is the text feature vector, and E k is the knowledge feature vector.
[0025] Further, a BART encoder is used to encode the enhanced multimodal context features to obtain high-order features. The BART encoder includes multiple encoding layers, and each encoding layer includes a multi-head self-attention unit and a feed-forward network unit. The BART decoder includes decoding layers corresponding to the number of encoding layers, and each decoding layer includes a masked multi-head self-attention unit, a multi-head encoder-decoder attention unit, and a feed-forward network unit.
[0026] A generative dialogue system based on multimodal knowledge enhancement, comprising:
[0027] A feature processing module for obtaining a knowledge sequence that matches the text context in the multimodal context;
[0028] A feature extraction module for respectively performing vector representations on the knowledge sequence and the text context in the multimodal context to obtain a knowledge feature vector and a text feature vector, and simultaneously extracting features from the image context in the multimodal context to obtain visual semantic features;
[0029] A feature fusion module for fusing the knowledge feature vector and the text feature vector and then obtaining global text features through a text encoder; using the global text features to perform attention weighting on the visual semantic features to obtain enhanced visual semantic features, using the visual semantic features to perform attention weighting on the global text features to obtain enhanced global text features, concatenating the enhanced visual semantic features and the enhanced global text features, passing the concatenated result through a fully connected layer to obtain multimodal context features, and using the knowledge feature vector to perform attention weighting on the multimodal context features to obtain enhanced multimodal context features;
[0030] A response generation module for using a BART encoder to encode the enhanced multimodal context features to obtain high-order features, using a BART decoder to decode the high-order features in an autoregressive manner to obtain hidden state vectors, converting the dimension of the hidden state vectors to the vocabulary dimension of the BART pre-trained language model, normalizing the hidden state vectors with the converted dimension and then mapping them to the probability space to obtain the predicted probabilities of the words in the vocabulary, and generating a dialogue based on the predicted probabilities of the words.
[0031] Compared with the prior art, the beneficial effects of the present invention are:
[0032] The present invention enhances the global text features based on visual semantic features using an attention mechanism, enhances the visual semantic features based on the global text features using an attention mechanism, fuses the enhanced results to obtain multi-modal context features that initially fuse images and text, and then assigns attention weights to the multi-modal context features based on knowledge feature vectors to obtain enhanced multi-modal context features fused with external knowledge. The method provided by the present invention fuses external knowledge, corresponding semantic features, and text features to obtain a relatively reasonable dialogue response. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a system block diagram of a generative dialogue method based on multi-modal knowledge enhancement provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following takes the publicly available multi-modal dialogue dataset MMConv as an example, and in combination with the accompanying drawings, the present invention is further described in detail. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0035] As Figure 1 shown, the present invention provides a generative dialogue method based on multi-modal knowledge enhancement, including:
[0036] In a specific embodiment of the present invention, multi-modal contexts are obtained, and multiple multi-modal contexts are used to construct a dataset. The dataset contains a total of 5,106 dialogue segments, with a total of 39,759 dialogue turns. After processing according to the dialogue turns, a total of 34,653 dialogue data are obtained.
[0037] The multi-modal contexts provided by the specific embodiment of the present invention include text contexts and image contexts. An NLP tool is used to perform named entity recognition on the text contexts to obtain an entity name sequence. The NLP tool is Hanlp or jieba, and a knowledge base is obtained. The knowledge base includes multiple knowledge entities, and each knowledge entity includes multiple knowledge attributes. The cosine similarity between the word vectors of the knowledge entities in the entity name sequence and the word vectors of the knowledge entities in the knowledge base is compared, and the knowledge entities with the top K similarities are retained, and the knowledge attributes corresponding to the knowledge entities with the top K similarities are integrated to obtain a knowledge sequence. The knowledge attributes include name, field, category, etc.
[0038] In a specific embodiment, the cosine similarity calculation formula is:
[0039]
[0040] Among them, a is the word vector of the knowledge entity in the entity name sequence, and b is the word vector of the knowledge entity in the knowledge base;
[0041] In the specific embodiment of the present invention, the knowledge sequence and the text context are respectively vectorized through an Embedding Layer to obtain the knowledge feature vector E k ={e k1 , e k2 , … e kg …, e kG} and the text feature vector E h ={e h1 , e h2 , … e hf …, e hF}, where e ke is the word vector of the e-th in the knowledge feature vector, E is the number of word vectors in the knowledge feature vector, e hf is the word vector of the f-th in the text feature vector, and F is the number of word vectors in the text feature vector. The ViT-B / 32 (Vision Transformer-Base) module pre-trained based on the CLIP model (Contrastive Language-Image Pre-training, a pre-training model for contrasting text-image pairs) is used to extract the features of the image context to obtain the visual semantic feature T v ={e v1 , e v2 , … e vm …, e vM}, where e vm is the word vector of the m-th in the visual semantic feature, and M is the number of word vectors in the visual semantic feature.
[0042] In the specific embodiment of the present invention, the knowledge feature vector and the text feature vector are fused and then the global text feature T t is obtained through a text encoder:
[0043] T t =Text_Encoder(E h +E k )={e t1 , e t2 , … e ti …, e tI}
[0044] Among them, Text_Encoder is the text encoder, e ti is the word vector of the i-th in the global text feature, I is the number of word vectors in the global text feature, E h is the text feature vector, and E kIt is a knowledge feature vector.
[0045] In a specific embodiment of the present invention, the visual semantic feature attention weight of the word vector corresponding to the visual semantics is obtained based on the similarity between the word vector of each visual semantic feature and all word vectors of the global text feature.
[0046] In a specific embodiment, the semantic feature attention weight α of the m-th word vector in the visual semantic feature t→vm is:
[0047]
[0048]
[0049] where Similarity(.) is the modality similarity. The word vector of each visual semantic feature and the word vector of the global text feature are respectively subjected to similarity matching using the scaled dot product method to obtain the first similarity vector set, and after summing and normalizing the first similarity vector set, the visual semantic feature attention weight of the word vector corresponding to the visual semantics is obtained.
[0050] In a specific embodiment of the present invention, the corresponding semantic feature attention weight is assigned to the word vector of each visual semantic feature to obtain the enhanced visual semantic feature.
[0051] In a specific embodiment, the corresponding semantic feature attention weight is assigned to the word vector of the m-th visual semantic feature to obtain the enhanced visual semantic feature is:
[0052]
[0053] where d is the dimension index of the visual semantic feature attention weight α of the m-th word vector in the visual semantic feature, D is the number of dimensions, and the m-th word vector in the visual semantic feature corresponds to the dimension of the corresponding visual semantic feature attention weight α. Each word vector in the visual semantic feature is enhanced to obtain the enhanced visual semantic feature. t→vm of, and the m-th word vector in the visual semantic feature corresponds to the dimension of the corresponding visual semantic feature attention weight α. Each word vector in the visual semantic feature is enhanced to obtain the enhanced visual semantic feature. t→vm Each word vector in the visual semantic feature is enhanced to obtain the enhanced visual semantic feature.
[0054] In an embodiment of the present invention, the global text feature is subjected to attention weighting using the visual semantic feature to obtain the enhanced global text feature, and the method for obtaining the enhanced global text feature is the same as the method for obtaining the enhanced visual semantic feature.
[0055] In a specific embodiment, the semantic feature attention weight α of the i-th word vector in the global text feature v→ti is:
[0056]
[0057]
[0058] Among them, Similarity(.) is the modal similarity. The word vectors of each global text feature and the word vectors of the visual semantic features are respectively subjected to similarity matching by using the scaled dot product method to obtain the third similarity vector set. After summing and normalizing the third similarity vector set, the attention weight of the global text feature corresponding to the word vector of the global text feature is obtained.
[0059] In a specific embodiment of the present invention, the attention weight of the global text feature corresponding to the word vector of each global text feature is assigned to obtain the enhanced global text feature.
[0060] In a specific embodiment, the attention weight of the global text feature corresponding to the word vector of the i-th global text feature is assigned to obtain the enhanced global text feature as:
[0061]
[0062] where r is the dimension index of the attention weight α of the global text feature corresponding to the i-th word vector in the global text feature, R is the number of dimensions, and the i-th word vector in the global text feature corresponds to the dimension of the corresponding attention weight α of the global text feature. Each word vector of the global text feature in the global text feature is enhanced to obtain the enhanced visual semantic feature. v→ti of the global text feature. v→ti The present invention embodiment splices the enhanced visual semantic feature and the enhanced global text feature, and passes the splicing result through a fully connected layer to obtain the multi-modal context feature T as:
[0063] T = concat(T
[0064] , T v ′ , T t ′ ) = {e p1 , e p2 , … e pq , …, e pQ}
[0065] where T v ′ is the enhanced visual semantic feature, T t ′ is the enhanced global text feature, and e pqis the word vector of the q-th in the multi-modal context features, and Q is the number of word vectors in the multi-modal context features. To utilize knowledge information in a more refined manner, in addition to the global enhancement described above, the knowledge feature vector is used to perform attention weighting on the multi-modal context features to obtain enhanced multi-modal context features. The method steps for obtaining the enhanced multi-modal context features are the same as those for obtaining the enhanced global text features.
[0066] In a specific embodiment, the multi-modal context feature attention weight α of the q-th word vector in the visual semantic features k→pq is:
[0067]
[0068]
[0069] where Similarity(.) is the modality similarity. The word vectors of each multi-modal context feature and the word vectors of the knowledge feature vector are respectively subjected to similarity matching using the scaled dot product method to obtain the second similarity vector set. After summing and normalizing the second similarity vector set, the multi-modal context feature attention weight corresponding to the word vector of the multi-modal context feature is obtained.
[0070] In the specific embodiment of the present invention, the corresponding multi-modal context feature attention weight is assigned to the word vector of each multi-modal context feature to obtain the enhanced multi-modal context features.
[0071] In a specific embodiment, the corresponding multi-modal context feature attention weight is assigned to the word vector of the q-th multi-modal context feature to obtain the enhanced visual multi-modal context features is:
[0072]
[0073] where c is the dimension index of the multi-modal context feature attention weight α of the q-th word vector in the multi-modal context features, C is the number of dimensions, and the q-th word vector in the multi-modal context features corresponds to the dimension of the corresponding multi-modal context feature attention weight α. k→pq The word vectors of each multi-modal context feature are enhanced to obtain the enhanced multi-modal context features. k→pq
[0074] In the specific embodiments of the present invention, a BART encoder is used to learn high-order information of multimodal features. The encoder consists of 6 layers of Encoders, and each layer contains two sub-layers: a multi-head self-attention layer (MSA) and a feed-forward network (FFN). After each sub-layer, residual connection and normalization (LN) operations are used to enhance the generalization of the model. At the same time, we use a BART decoder for autoregressive generation. The model consists of 6 layers of Decoders, and each layer contains three sub-layers: a masked multi-head self-attention layer (MMSA), a multi-head encoder-decoder attention layer (MEA), and a feed-forward network (FFN).
[0075] The output layer is a fully connected layer, which converts the vector dimension from the dimension size of the hidden state vector output by the last layer of the decoder to the vocabulary size through a linear transformation, and then maps the output to the probability space through the softmax function to obtain the prediction probability of the words in the vocabulary. It is:
[0076]
[0077] Among them, represents the hidden state vector output by the last layer of the decoder, W y and b y are parameters to be learned.
[0078] The model is trained using the cross-entropy loss function, which is used to measure the difference between the true distribution of the data labels and the predicted probability distribution. Let y t represent the true response sequence, and let represent the model prediction sequence. Then the calculation formula of the loss function is as follows:
[0079]
[0080] The embodiments of the present invention also provide a generative dialogue system based on multimodal knowledge enhancement, which is characterized by including:
[0081] A feature processing module for obtaining a knowledge sequence that matches the text context in the multimodal context;
[0082] A feature extraction module for respectively vectorizing the knowledge sequence and the text context in the multimodal context to obtain a knowledge feature vector and a text feature vector, and at the same time extracting features from the image context in the multimodal context to obtain a visual semantic feature;
[0083] A feature fusion module for fusing the knowledge feature vector and the text feature vector, and then obtaining the global text feature through a text encoder; using the global text feature to perform attention weighting on the visual semantic feature to obtain an enhanced visual semantic feature, using the visual semantic feature to perform attention weighting on the global text feature to obtain an enhanced global text feature, splicing the enhanced visual semantic feature and the enhanced global text feature through a feature bidirectional refinement module, passing the splicing result through a fully connected layer to obtain a multi-modal context feature, and using the knowledge feature vector to perform attention weighting on the multi-modal context feature to obtain an enhanced multi-modal context feature;
[0084] A response generation module for encoding the enhanced multi-modal context feature using a BART encoder to obtain a high-order feature, decoding the high-order feature in an autoregressive manner using a BART decoder to obtain a hidden state vector, and converting the dimension of the hidden state vector to the vocabulary dimension, where the vocabulary dimension is: the vocabulary dimension of the BART pre-trained language model, which is pre-trained on a large-scale corpus, and the corresponding vocabulary is all the words that appear in the pre-trained corpus.
[0085] Mapping the hidden state vector with the converted dimension to the probability space through an output layer to obtain the prediction probability of the words in the vocabulary, and generating a dialogue based on the prediction probability of the words.
[0086] For the dialogue generation system based on multi-modal context enhancement, there are the following evaluation metrics:
[0087] 1) PPL (Perplexity): It is used to evaluate whether the generated response conforms to the language rules in the dialogue system. The smaller the value, the more it conforms to the legal probability of the language model. The calculation formula is as follows:
[0088]
[0089] Among them, P(S) is the sentence score, N represents the number of words in the sentence, and loss is the cross-entropy loss.
[0090] 2) BLEU (Bilingual Evaluation Understudy): BLEU is a commonly used metric for evaluating the relevance of the responses generated by the dialogue system. BLEU-N counts the number of occurrences of n-gram phrases in the generated response and the real response in the entire training corpus. BLEU is used to calculate the co-occurrence word frequency of two sentences. The larger the value, the more consistent the generated response is with the real response. The calculation formula is as follows:
[0091]
[0092] Among them, N represents the phrase size, P nis the accuracy of n-gram phrases in the entire dataset, ω represents the weight of each n-gram, and BP is the penalty factor for short length.
[0093] In this embodiment, we use the BART pre-trained language model as the baseline and conduct comparative experiments on the MMConv dataset. Among them, Baseline only uses the text context as the input, With Feature introduces image input on the basis of the baseline and fuses it with the text features, and With Knowledge Enhancement introduces external knowledge for feature enhancement on the premise of text and image input. The model is trained for a total of 100 epochs, and the model with the best performance on the validation set is selected for testing. The experimental results are shown in Table 1. It can be seen from the table that after introducing visual features, the BLEU index has improved, and the PPL has also increased slightly. This shows that multi-modal input can generate responses that are more in line with real responses than single-text input without affecting the fluency of the conversation. Although the PPL of the model with external knowledge enhancement has decreased slightly, our knowledge enhancement model has obtained the best performance in terms of the BLEU index, which proves the effectiveness of introducing external knowledge to improve the system response.
[0094] Table 1. Comparison of experimental results before and after adding visual features on the MMConv dataset, BLEU-N(%)
[0095]
[0096] The above has described the present invention in detail. The above specific embodiments or examples are only examples revealing the technical content of the present invention. The present invention is not limited to the above specific examples, and should not be interpreted narrowly. Various changes can be made within the spirit and scope of the claims of the present invention for implementation.
Claims
1. A generative dialogue method based on multi-modal knowledge enhancement, characterized in that, Including: Obtain a multi-modal context, where the multi-modal context includes a text context and an image context. Obtain a knowledge sequence that matches the text context. Respectively perform vector representations on the knowledge sequence and the text context to obtain a knowledge feature vector and a text feature vector, and simultaneously perform feature extraction on the image context to obtain a visual semantic feature. After fusing the knowledge feature vector and the text feature vector, obtain a global text feature through a text encoder; Use the global text feature to perform attention weighting on the visual semantic feature to obtain an enhanced visual semantic feature, use the visual semantic feature to perform attention weighting on the global text feature to obtain an enhanced global text feature, concatenate the enhanced visual semantic feature and the enhanced global text feature, pass the concatenated result through a fully connected layer to obtain a multi-modal context feature, and use the knowledge feature vector to perform attention weighting on the multi-modal context feature to obtain an enhanced multi-modal context feature; Use a BART encoder to encode the enhanced multi-modal context feature to obtain a high-order feature, use a BART decoder to decode the high-order feature in an autoregressive manner to obtain a hidden state vector, convert the dimension of the hidden state vector to the vocabulary dimension of the BART pre-trained language model, normalize the hidden state vector with the converted dimension and map it to a probability space to obtain the prediction probability of the words in the vocabulary, and generate a dialogue based on the prediction probability of the words.
2. The generative dialogue method based on multi-modal knowledge enhancement according to claim 1, wherein The using the global text feature to perform attention weighting on the visual semantic feature to obtain an enhanced visual semantic feature includes: Based on the similarity between the word vector of each visual semantic feature and all the word vectors of the global text feature, obtain the attention weight of the visual semantic feature, and assign the corresponding attention weight of the visual semantic feature to the word vector of each visual semantic feature to obtain an enhanced visual semantic feature.
3. The generative dialogue method based on multi-modal knowledge enhancement according to claim 2, wherein, The based on the similarity between the word vector of each visual semantic feature and all the word vectors of the global text feature to obtain the attention weight of the visual semantic feature includes: Adopt a scaled dot product method to perform similarity matching between the word vector of each visual semantic feature and the word vectors of the global text feature respectively to obtain a first similarity vector set, sum and normalize the first similarity vector set to obtain the attention weight of the visual semantic feature corresponding to the word vector of the visual semantic feature.
4. The generative dialogue method based on multi-modal knowledge enhancement according to claim 1, wherein The obtaining a knowledge sequence that matches the text context includes: Use an NLP tool to perform named entity recognition on the text context to obtain an entity name sequence, obtain a knowledge base, where the knowledge base includes multiple knowledge entities, and each knowledge entity includes multiple knowledge attributes. Compare the cosine similarity between the word vectors of the knowledge entities in the entity name sequence and the word vectors of the knowledge entities in the knowledge base, retain the knowledge entities with the top K similarities, and integrate the knowledge attributes corresponding to the knowledge entities with the top K similarities to obtain a knowledge sequence.
5. The generative dialogue method based on multi-modal knowledge enhancement according to claim 1, wherein Respectively perform vector representations on the knowledge sequence and the text context through an embedding layer to obtain a knowledge feature vector and a text feature vector; use a ViT-B / 32 module pre-trained based on the CLIP model to extract the features of the image context to obtain a visual semantic feature.
6. The generative dialogue method based on multi-modal knowledge enhancement according to claim 1, wherein Fuse the knowledge feature vector and the text feature vector and then encode them to obtain the global text feature T t It is: T t = Text_Encoder(E h + E k ) = {e t1 , e t2 , … e ti …, e tI} Among them, Text_Encoder is the text encoder, e ti is the word vector of the i-th in the global text feature, I is the number of word vectors in the global text feature, E h is the text feature vector, E k is the knowledge feature vector.
7. The generative dialogue method based on multi-modal knowledge enhancement according to claim 1, wherein The enhanced multimodal context features are encoded by a BART encoder to obtain high-order features. The BART encoder includes multiple encoding layers, and each encoding layer includes a multi-head self-attention unit and a feed-forward network unit. The BART decoder includes decoding layers corresponding to the number of encoding layers, and each decoding layer includes a masked multi-head self-attention unit, a multi-head encoder-decoder attention unit, and a feed-forward network unit.
8. A generative dialogue system based on multi-modal knowledge enhancement, characterized in that, Comprising: A feature processing module for obtaining a knowledge sequence that matches the text context in the multimodal context; A feature extraction module for respectively vectorizing the knowledge sequence and the text context in the multimodal context to obtain a knowledge feature vector and a text feature vector, and simultaneously extracting features from the image context in the multimodal context to obtain visual semantic features; A feature fusion module for fusing the knowledge feature vector and the text feature vector and then obtaining global text features through a text encoder; using the global text features to perform attention weighting on the visual semantic features to obtain enhanced visual semantic features, using the visual semantic features to perform attention weighting on the global text features to obtain enhanced global text features, concatenating the enhanced visual semantic features and the enhanced global text features, obtaining multimodal context features through a fully connected layer on the concatenated result, and using the knowledge feature vector to perform attention weighting on the multimodal context features to obtain enhanced multimodal context features; A response generation module for encoding the enhanced multimodal context features by a BART encoder to obtain high-order features, decoding the high-order features in an autoregressive manner by a BART decoder to obtain hidden state vectors, converting the dimension of the hidden state vectors to the vocabulary dimension of the BART pre-trained language model, normalizing the hidden state vectors with the converted dimension and then mapping them to a probability space to obtain the prediction probabilities of words in the vocabulary, and generating a dialogue based on the prediction probabilities of words.
Citation Information
Patent Citations
Knowledge-introduced generative session system
CN112084314A
Visual question and answer oriented method of context awareness based on multi-modal interaction
CN114970517A