Multi-modal Dialogue Q&A Generation Method Based on Multi-relation Graph Model
By introducing multi-relational graph model and graph convolutional neural network into the multi-modal dialogue system, the problem of existing systems only considering scene serialization information is solved, and a higher quality multi-modal dialogue question-and-answer generation is achieved, which significantly improves the performance of the model.
Patent Information
- Application Number
- CN202211451009.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-11-18
AI Technical Summary
The existing multimodal dialogue system only considers scene serialization information, resulting in a general model effect and is unable to effectively capture the multi-grained semantic complementary relationship between different modes.
A multimodal dialogue question-and-answer generation method based on multi-relational graph model is proposed. By constructing a multi-relational graph model, the entity correspondence between video clips and syntactic and semantic relationships in dialogue are modeled, and multimodal feature representation and response generation are carried out in combination with graph convolutional neural network and GPT-2 model.
Through the combination of multi-relational graph model and graph convolutional neural network, complex dependencies and semantic information between multimodals can be more effectively captured, significantly improving the quality and performance of generated answers, especially in the CIDEr index, which outperforms existing models.
Smart Images

Figure CN115712709B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of dialogue question - answering, and particularly relates to a multi - modal dialogue question - answering generation method. Background Art
[0002] Current research in the field of dialogue question - answering systems mainly falls into two major branches: text and multi - modal. There are mainly two major difficulties in text dialogue question - answering tasks: answer generation requires dialogue context reasoning and lacks large - scale dialogue datasets. Since pre - trained language models (LMs) have learned rich semantic information from other text data, can perform a certain degree of reasoning, and effectively make up for the problem of insufficient dialogue data volume, enabling the system to still achieve good results in a low - resource background. Therefore, introducing pre - trained language models into dialogue question - answering tasks can deepen the system's understanding of text, reason based on historical dialogue turns to process the current user's question, and improve the quality of generated answers. In ISCA2020, Whang et al. applied pre - trained language models in open - domain dialogue to select candidate answers, where the output of the pre - trained language model (such as the [CLS] token in BERT) was used as the context representation for each dialogue context and candidate answer pair. In WNGT2019, Budzianowski et al. assumed that the true dialogue state could be obtained, and merged the inputs into a single sequence to generate responses for task - oriented dialogue. Since the dialogue state and database state can be regarded as raw text inputs, pre - trained language models can be used to fine - tune the system. In ICASSP2020, Lai et al. introduced the GPT - 2 model, using the output representation of the model to predict slot values, and then tracking the dialogue state.
[0003] In ACL2020, hierarchical pointer networks have also been widely applied in text dialogue systems. In ICLR2019, Wu et al. merged a global encoder and a local decoder to achieve sharing of external knowledge in a task - oriented dialogue setting. In NAACL2019, Reddy et al. designed a multi - level storage framework for task - oriented dialogue. In ACL2019, Tian et al. explored how to extract valuable information during training and built a memory - priming architecture based on this. In addition, multi - task learning has also been proven to optimize the performance of natural language answers. In the work of Chen et al. in ACL2019, working memory was introduced into this task. By fully interacting with two long - term memories, it can capture tuples of dialogue history and knowledge base to generate high - quality answers. In EMNLP2019, Lin et al. also applied heterogeneous memory networks to the research in this field in view of their ability to simultaneously utilize dialogue context, user questions, and knowledge base information.
[0004] Multi-modal dialogue question answering has opened up a new pattern for realizing powerful dialogue systems. Current research mainly focuses on how to bridge the gap between language and vision by leveraging the multi-granularity complementary information between static images and text. For example, in CVPR2017, the Visual Dialog task proposed by Das et al. provides pictures and related multi-turn dialogues, requiring the model to correctly answer questions related to the questioner in natural language based on the given image and dialogue history. Although this task is of great significance in promoting the development of multi-modal dialogue question answering systems, dialogue based on static images has certain inherent limitations. It largely restricts the dynamic perception ability of the question answering system to spatio-temporal changes, making it unable to reasonably handle many applications that require understanding the context of a specific scene to make reasonable inferences. Therefore, to enhance the spatio-temporal intelligence of the question answering system, a new task - Audio-Visual Scene-Aware Dialog (AVSD) is introduced. It can be regarded as a general form of visual dialogue, that is, visual dialogue based on consecutive picture frames and audio information, with a broader application prospect compared to visual dialogue. However, existing solutions mainly use independent encoders to encode different modalities separately, and then use the attention mechanism to fuse their representations and generate response sentences. This late fusion solution only considers the sequential features of the scene and dialogue, and ignores the multi-granularity semantic complementary relationship between different modalities, resulting in the unsatisfactory performance of existing models. Therefore, exploring the semantic information representation and modality fusion method of the dialogue scene is of great significance for realizing a higher-performance multi-modal dialogue question answering system.
[0005] Meanwhile, the focus of related research in the multi-modal dialogue field is shifting towards how to fully integrate multi-source heterogeneous information, including images, audio, video, and text, etc. Compared with text dialogue question answering, the multi-modal dialogue question answering task additionally introduces audio-visual features related to the dialogue, thus requiring the solution of fine-grained interaction problems between different modalities.
[0006] The attention mechanism is the mainstream research method in this field, which can narrow the gap between visual and language modality representations. In CVPR2018, the CoAtt model designed by Wu et al. contains an encoder with a sequential co-attention mechanism, enabling each input feature to be jointly provided by the other two features in a sequential manner. The ReDAN model proposed by Gan et al. in ACL2019 and the DMRM model proposed by Chen et al. in AAAI2020 answer a series of questions related to images through multi-step reasoning based on the dual attention mechanism. In ECCV2020, the LTMI model designed by Nguyen et al. uses the multi-head attention mechanism to focus on the interaction relationship between modalities.
[0007] The pre-trained language model architecture based on the Transformer architecture implemented with attention mechanism also performs well in learning cross-modal representations for visual-text natural language processing tasks. In the image captioning task, Li et al. in AAAI 2020 constructed a BERT-based architecture to improve text and visual representations. In NIPS 2019, Lu et al. used a similar method to deal with the visual question answering task. The difference from the former is that when processing multi-modal inputs, the visual and text representations are separated rather than merged into an entire sequence. In IJCNLP 2019, Alberti et al. noticed the important role of early fusion or late fusion methods in enriching cross-modal representations. In ICCV 2019, Sun et al. proposed the VideoBERT model, which uses the BERT model to generate video captions and abandons the method of representing video frames with visual features. Instead, it transforms the frame-level features into visual tokens as the original input of the model.
[0008] Recent research has also explored higher-level semantic representations of pictures or dialogue history, especially the way of modeling pictures or dialogue history based on graph structures. In AAAI 2020, the DualVD model proposed by Jiang et al. describes the features of pictures in detail from both visual and semantic perspectives. Specifically, the visual graph model helps extract surface information including entities and relationships, and the semantic graph model promotes the transformation of the dialogue question answering system from global to local visual semantic understanding. In CVPR 2020, the CAG model designed by Guo et al. establishes graph nodes with entity-related visual representations and history-related context representations, and updates the corresponding edge weights with an adaptive Top-K information transfer mechanism to establish a visual-semantic related dynamic graph for subsequent reasoning. In ACL 2021, the GoG model proposed by Chen et al. takes into account that there are also interactions between different relationships, so it models the current question dependency graph based on dialogue history and the object (region) relationship graph based on the current question. Summary of the Invention
[0009] The purpose of the present invention is to solve the problem that the existing multi-modal dialogue systems only consider scene serialization information, resulting in general performance of existing models, and further propose a multi-modal dialogue question answering generation method based on a multi-relationship graph model.
[0010] A multi-modal dialogue question answering generation method based on a multi-relationship graph model includes the following steps:
[0011] S1. Use a sliding window of fixed size to serially segment the video into multiple video clips. For each clip, obtain the color feature and optical flow feature as well as the audio feature
[0012] The color feature Optical flow features and audio features are concatenated to obtain and then the location information and modality information are added to obtain the sequence representation V of each video segment t ; The expression is:
[0013]
[0014]
[0015] where the location information uses numbers to represent the order in which each video segment appears, and the modality information uses the identifier [video] to uniformly identify video features, which are converted into vectors of fixed dimensions during actual calculations;
[0016] S2. For the audio-visual scene representation V=(V 1 ,V 2 ,...,V m ), V 1 ,V 2 ,...,V m is the sequence representation of each video segment. Each video segment is regarded as a vertex to construct a video graph based on the fully connected relationship where ε V is the set of directed dependency edges. For each directed dependency edge (V i ,V j ,l ij ), l ij represents the dependency relationship from V i to V j and is set to 1;
[0017] The video graph is input into the graph convolutional neural network, and the video hidden layer sequence G V is output;
[0018] S3. The video hidden layer sequence G V and the original video sequence representation V are input into the linear layer to obtain the fused representation of the video and used as part of the input of the subsequent multi-layer TRANSFORMER model based on the GPT-2 architecture;
[0019] S4. Based on the audio-visual scene title C and the dialogue history H, the corresponding word vector representations C feature and H feature are obtained; The title word vector C feature and the dialogue history word vector H feature are concatenated, and then the location information T is addedpos and modal information T mod obtain the text sequence representation T; the expression is:
[0020] T feature =[C feature ,H feature ,
[0021] T = T feature +T mod +T pos ,
[0022] wherein, for the position information T pos numbers are used to refer to the order of appearance of words in the title and each Q&A pair, and for T pos the title of the audio-visual scene is uniformly identified by the identifier [cap], the questioner is identified by the identifier [usr1], and the answerer is identified by the identifier [usr2]. During actual calculation, they are respectively converted into vectors of a fixed dimension;
[0023] S5. Regarding each word vector in the text sequence representation T obtained in S4 as a vertex, construct a graph structure based on sentence-level dependency relationships and / or a graph structure based on the co-reference relationship of the complete dialogue
[0024] Then, for the graph structure based on sentence-level dependency relationships and / or the graph structure based on the co-reference relationship of the complete dialogue input them into a graph convolutional neural network respectively to obtain their corresponding text hidden layer sequences;
[0025] S6. Input the text hidden layer sequences corresponding to the graph structure based on sentence-level dependency relationships and / or the graph structure based on the co-reference relationship of the complete dialogue and the original text sequence representation T into a linear layer to obtain the fused representation of the text and use it as part of the input for the multi-layer TRANSFORMER model based on the GPT-2 architecture;
[0026] S7. Concatenate and to obtain the enhanced multi-modal input, and generate an answer based on the enhanced multi-modal input using the multi-layer Transformer model based on the GPT-2 architecture.
[0027] Furthermore, the process of regarding each word vector in the text sequence representation T obtained in S4 as a vertex to construct a graph structure based on sentence-level dependency relationships in S5 and obtaining the corresponding text hidden layer sequence includes the following steps:
[0028] First, use the GPT2 Tokenizer to obtain the word vector representation corresponding to each word. Use the StanfordCoreNLP text parsing tool to analyze the syntactic dependency relationship of the sentence. Treat each word vector as a vertex and model the graph structure based on the syntactic dependency relationship. Then input it into the graph convolutional neural network to output the text hidden layer sequence G D 。
[0029] Or,
[0030] In S5, each word vector in the text sequence representation T obtained in S4 is regarded as a vertex, and a graph structure based on the coreference relationship of the complete conversation is constructed And the process of obtaining their respective corresponding text hidden layer sequences includes the following steps:
[0031] First, use the GPT2 Tokenizer to obtain the word vector representation corresponding to each word. Use the StanfordCoreNLP text parsing tool to analyze the coreference relationship of the sentence. Treat each word vector as a vertex and model the graph structure based on the coreference relationship of the sentence. Then input it into the graph convolutional neural network to output the text hidden layer sequence G C 。
[0032] Or,
[0033] In S5, each word vector in the text sequence representation T obtained in S4 is regarded as a vertex, and a graph structure based on the sentence-level dependency relationship and a graph structure based on the coreference relationship of the complete conversation are constructed and a graph structure based on the coreference relationship of the complete conversation And the process of obtaining their respective corresponding text hidden layer sequences includes the following steps:
[0034] First, use the GPT2 Tokenizer to obtain the word vector representation corresponding to each word. Use the StanfordCoreNLP text parsing tool to analyze the syntactic dependency relationship and the coreference relationship of the sentence respectively. Treat each word vector as a vertex and construct a graph structure based on the sentence-level dependency relationship and a graph structure based on the coreference relationship of the complete conversation Then input the two graph structures into the graph convolutional neural network respectively to output the text hidden layer sequences G D and G C 。
[0035] Furthermore, the expression calculated by each layer of the graph convolutional neural network in the graph convolutional neural network described in S5 is:
[0036]
[0037] where f(H v (l+1) ,A v)Indicates each layer of graph convolution, for or A d respectively represent their corresponding adjacency matrices, is the corresponding degree matrix, adding the identity matrix I d to obtain Correspondingly, the corresponding degree matrix is obtained for the convenience of normalization operation, l is the number of layers of the graph convolutional neural network, is the hidden state of the l-th layer of the graph convolutional neural network,
[0038] Furthermore, the expression calculated by each layer of the graph convolutional neural network in the graph convolutional neural network described in S2 is:
[0039]
[0040] where f(H v (l+1) , A v ) represents each layer of graph convolution; A v is 's adjacency matrix, i and j respectively represent 's i-th and j-th nodes, is the value of the i-th row and j-th column of A v ; is 's degree matrix, is the value of the i-th row and i-th column of D v ; to enable the model to consider the node's own representation, add the identity matrix I v to obtain Correspondingly, the corresponding degree matrix is obtained for the convenience of normalization operation; l is the number of layers of the graph convolutional neural network, is the hidden state of the l-th layer of the graph convolutional neural network, V is the original video sequence representation, is the trainable weight.
[0041] Furthermore, when processing with the multi-layer Transformer model based on the GPT-2 architecture, the in S3 also needs to be input into the linear fully connected layer, and the output result is projected into the same vector space as the in S6, and the two are concatenated to obtain a complete multi-modal input representation, and then it is input into the pre-trained language model GPT-2.
[0042] Preferably, the multi-layer Transformer model based on the GPT-2 architecture in S7 is stacked by 12 layers of Transformer decoder modules with masked multi-head attention mechanisms.
[0043] Furthermore, the multi-layer Transformer model based on the GPT-2 architecture is trained using the negative log-likelihood loss function during the training process, and the training process includes the following steps:
[0044] Generate an answer based on the audio-visual feature V, title C, conversation history H <n and the current question Q n Generate an answer By minimizing the negative log-likelihood loss function, maximize the likelihood that the output next word is the corresponding word in the source sequence:
[0045]
[0046] where represents the first j-1 words of the answer R n θ refers to the trainable model parameters, the set (V, C, H, Q) is sampled from the entire training set D, and E (V,C,H,Q,R)~D represents the expectation.
[0047] Or
[0048] The multi-layer Transformer model based on the GPT-2 architecture is jointly trained based on the answer prediction task RPT, title prediction task CPT of audio-visual, and audio-visual-text matching task VTMT during the training process, and the training process includes the following steps:
[0049] The RPT part aims to generate an answer based on the audio-visual feature V, title C, conversation history H <n and the current question Q n Generate an answer By minimizing the negative log-likelihood loss function, maximize the likelihood that the next word output by the model is the corresponding word in the source sequence:
[0050]
[0051] where represents the first j-1 words of the answer R n θ refers to the trainable model parameters, the set (V, C, H, Q) is sampled from the entire training set D, and E (V,C,H,Q,R)~D represents the expectation;
[0052] The CPT part is similar to the RPT part. For the given audio-visual feature V, generate the title C = {c 1 , c 2 ,..., c L} by minimizing the negative log-likelihood loss function, and the loss function is as follows:
[0053]
[0054] Among them, c <i represents the first i - 1 words of title C;
[0055] The VTMT part aims to determine whether the given audio - visual feature V matches the given text features. The given text features include title C, dialogue history H <n , current question Q n and generated answer R n ; Select a certain proportion of training data, randomly replace the corresponding original input with incorrect audio - visual features, and obtain the probability of matching or not by passing the final output of the hidden state of the GPT2 module through a linear fully - connected layer, and then use binary cross - entropy to calculate the loss function:
[0056]
[0057] Among them, X=(V, C, H, Q, R), and Y is the label indicating whether the audio - visual feature and the text feature match.
[0058] Beneficial effects:
[0059] A multi - modal dialogue question - answering generation method based on a multi - relational graph model according to the present invention constructs a multi - relational graph model according to the characteristics of different modalities to enrich the multi - modal feature representation. By modeling the corresponding relationships of entities in continuous video segments and the syntactic and semantic relationships implicit in continuous dialogues, the understanding of the scene and dialogue by the system is further deepened, improving the deficiency of existing methods that only consider temporal order and word order encoding, and further improving the quality of the generated answer.
[0060] The present invention aims to improve the situation where the current pre - trained model architecture only uses the serialized information of videos or texts to obtain the embedding representation corresponding to each video segment or word, resulting in unsatisfactory generated answers for multi - modal dialogue question - answering systems. An additional graph convolutional neural network is added to enable the model to have the ability to effectively encode the information of the multi - relational graph structure. Specifically, the model introduces a graph convolutional neural network, calculates the data distribution of the adjacent node set of the current node, assigns corresponding weights to the current node, and after stacking multiple layers of graph convolutional neural networks, the model has the ability to infer nodes that are multiple jumps away from the current node, so as to capture the interaction information between the current node and distant nodes, and further represent the syntactic or semantic relationships between all words within a framework.
[0061] The invention can effectively improve various indicators of the system's generated responses, including BLEU, METEOR, ROUGE-L, CIDEr, etc. On the dataset related to audio-visual scene perception dialogue, the invention compares the system-generated answers and the manually annotated answers with multiple indicators. The experimental results show that the answers generated by this method conform to the basic habits of human expression, are basically matched with the manually annotated results, and are superior to the effects of all existing models in this task. Among them, the improvement in the CIDEr index, which represents the naturalness of sentences, is particularly obvious, with an average increase of 1% compared to the most advanced baseline results, which fully demonstrates its effectiveness and superiority. Description of the Drawings
[0062] Figure 1 It is a flowchart of the present invention;
[0063] Figure 2 It is the overall model architecture of the present invention;
[0064] Figure 3 It is a block diagram of the basic unit of the pre-trained language model;
[0065] Figure 4 It is an example of constructing a graph model based on dependency relationships;
[0066] Figure 5 It is an example of constructing a graph model based on co-reference relationships. Detailed Implementation Modes
[0067] Detailed Implementation Mode 1: Combining Figure 1 to illustrate this implementation mode,
[0068] This implementation mode is a multi-modal dialogue question and answer generation method based on a multi-relationship graph model, including the following steps:
[0069] Step 1: Use a sliding window of a fixed size to serially divide the video into multiple video segments. For each segment, use the I3D model to obtain the color features of the segment and optical flow features Use the VGGish model to obtain the audio features of the segment The color features optical flow features and audio features are concatenated to obtain Add position information and modality information to obtain the sequence representation V of each video segment t ; The expression is:
[0070]
[0071]
[0072] Among them, the position information uses numbers to represent the order in which each video segment appears, and the modality information uses the identifier [video] to uniformly identify video features, which are converted into vectors of fixed dimensions during actual calculation. Figure 2 In they are represented as V1, V2, V3, V4, V5, and their corresponding are all marked as [video].
[0073] Step 2: Given the audiovisual scene representation V = (V 1 , V 2 ,..., V m ), where V 1 , V 2 ,..., V m are the sequence representations of each video segment. Each video segment is regarded as a vertex, and a video graph based on the fully connected relationship is constructed. Among them ε V is the set of directed dependency edges. For each directed dependency edge (V i , V j , l ij ), l ij represents the dependency relationship from V i to V j , and is set to 1.
[0074] The video graph is input into the graph convolutional neural network, and the video hidden layer sequence G V is output; the calculation expression for each layer of the graph convolutional neural network is:
[0075]
[0076] Among them, f(H v (l+1) , A v ) represents each layer of graph convolution; A v is 's adjacency matrix, where i and j respectively represent 's i-th and j-th nodes, is the value of the i-th row and j-th column of A v ; is 's degree matrix, is the value of the i-th row and i-th column of D v ; To enable the model to consider the node's own representation, the identity matrix I v is added to obtain Correspondingly, the corresponding degree matrix To facilitate normalization operation; l is the number of graph convolutional neural network layers, is the hidden state of the l-th layer graph convolutional neural network, V is the original video sequence representation, are trainable weights.
[0077] Step 3: Transform the video hidden layer sequence G V And the original video sequence representation V is input into the linear layer to obtain the fusion representation of the video And as part of the input of the subsequent GPT2 model; the expression is:
[0078]
[0079] Among them, W M , W V is the training weight.
[0080] Step 4: Use GPT2 Tokenizer to obtain the word vector representation C of the audio-visual scene title C and the dialogue history H based on Word Pieces feature and H feature . The title word vector C feature and the conversation history word vector H feature Splice them together and add location information T pos and modal information T mod Get the text sequence representation T; the expression is:
[0081] T feature =[C feature ,H feature ],
[0082] T=T feature +T mod +T pos ,
[0083] Among them, the location information T pos Numbers are used to refer to the order in which words appear in the title and each question-answer pair. pos The identifier [cap] is used to uniformly identify the audio-visual scene title, the identifier [usr1] identifies the questioner, and the identifier [usr2] identifies the answerer. In actual calculation, they are converted into vectors of fixed dimensions. Figure 2 A specific example is given. The Feature layer uses GPT2 Tokenizer to encode the text "[cap]a woman…[eos][usr1]is the woman…[eos]" into T feature , according to the source of each sentence, the text "[cap]a woman…[eos]" corresponding to the T modLabeled as [cap], the T corresponding to the text "[usr1]is the woman…" mod Labeled as [usr1], the T corresponding to the texts "[usr2]yes she is…" and "[usr2]nothing much…[eos]" mod Labeled as [usr2].
[0084] Step 5: Treat each word vector as a vertex, and construct a text graph based on the sentence-level dependency relationship according to the Stanford CoreNLP text parsing tool Figure 4 A specific example is given. For the current text "does the womaneat or drink anything", first use the GPT2 Tokenizer to obtain the word vector representation corresponding to each word, use the Stanford CoreNLP text parsing tool to analyze the syntactic dependency relationship of the sentence (the connection lines between word vectors in the figure), treat each word vector as a vertex, then the graph structure can be modeled according to this dependency relationship, and input this graph into the graph convolutional neural network to output the text hidden layer sequence G D ; The expression is:
[0085]
[0086] where A d is the adjacency matrix of is the degree matrix of. To enable the model to consider the node's own representation, add the identity matrix I d to obtain Correspondingly, obtain the corresponding degree matrix for easy normalization operation. l is the number of layers of the graph convolutional neural network, is the hidden state of the l-th layer of the graph convolutional neural network,
[0087] Step 6: Input the text hidden layer sequence G D and the original text sequence representation T into the linear layer to obtain the fused representation of the text and use it as part of the input of the subsequent GPT2 model; The expression is:
[0088]
[0089] where, W N , W D are trainable weights.
[0090] Step 7: Construct the GPT2 model for multi-modal dialogue question answering (see Figure 2 andFigure 3 )。 Figure 2 shows the overall architecture of the model, which is a multi-layer Transformer model based on the GPT-2 architecture. This model is stacked by 12 Transformer decoder modules with masked multi-head attention mechanisms. To enable the model to have the ability to fuse multi-modal features and generate reasonable answers at the same time, certain modifications are made to the GPT-2 model for generation tasks to make it more in line with the needs of multi-modal dialogue question-answering tasks. Specifically, the model takes the result of step three and inputs it into a linear fully connected layer, and projects the output result into the same vector space as the result of step six , concatenates the two to obtain a complete multi-modal input representation, and then inputs it into the pre-trained language model GPT-2.
[0091] Figure 3 shows the specific architecture of each Transformer decoder module in the GPT-2 model. This module is mainly composed of a masked multi-head attention mechanism and a feed-forward neural network. Among them, the masked multi-head attention mechanism Masked Self-Attention can detect fine-grained long-term dependencies within and between the inputs of each modality, including spatio-temporal relationships of video objects, co-reference relationships in the dialogue history, reference relationships between local video features and text vocabulary, etc., to generate reasonable answers based on audiovisual features and in line with the user's question.
[0092] Step eight: Concatenate and to obtain an enhanced multi-modal input representation as the complete input of the GPT2 model for multi-modal dialogue question-answering constructed in step seven. Design a loss function based on this model for joint training to obtain a trained model, and then generate reasonable answers for a given audiovisual scene and user question.
[0093] During training, the negative log-likelihood loss function is used to enable the model to have the ability to predict answers based on audio-visual, caption, and dialogue history features. Formally, the model is based on audio-visual features V, caption C, dialogue history H <n and the current question Q n to generate an answer By minimizing the negative log-likelihood loss function, the probability that the output next word is the corresponding word in the source sequence is maximized:
[0094]
[0095] where, represents the first j - 1 words of the answer R n , θ refers to the trainable model parameters, the set (V, C, H, Q) is sampled from the entire training set D, E(V,C,H,Q,R)~D Indicates expectation. Specific Embodiment 2:
[0097] This embodiment is a multi-modal dialogue question and answer generation method based on a multi-relational graph model. The difference between this embodiment and the first specific embodiment is as follows:
[0098] In step 5, each word vector in the text sequence representation T obtained in step 4 is regarded as a vertex, and a text graph based on the co-reference relationship of the complete dialogue is constructed according to the Stanford CoreNLP text parsing tool Figure 5 A specific example is given. For the current text "a woman…a fridge…the woman…she…it", first use GPT2Tokenizer to obtain the word vector representation corresponding to each word, and use the Stanford CoreNLP text parsing tool to analyze the co-reference relationship of the sentence, that is, there is a co-reference relationship between "a woman", "the woman", and "she" in the context, and there is a co-reference relationship between "afridge" and "it". Regarding each word vector as a vertex, the graph structure can be modeled according to this co-reference relationship, that is, an edge is established between "a woman", "the woman", and "she", and an edge is also established between "a fridge" and "it". Input this graph into the graph convolutional neural network, and output the text hidden layer sequence G C , and the specific calculation formula is the same as the expression in step 5 of the first specific embodiment;
[0099] In step 6, the text hidden layer sequences G obtained in steps 4 and 5 C and the original text sequence representation T are input into the linear layer to obtain the fusion representation of the text and used as part of the input of the subsequent GPT2 model; the expression is:
[0100]
[0101] Among them, is a trainable weight.
[0102] Other steps and parameters are the same as those in the first specific embodiment. Specific Embodiment 3:
[0104] This embodiment is a multi-modal dialogue question and answer generation method based on a multi-relational graph model. The difference between this embodiment and the first or second specific embodiment is as follows:
[0105] In step five, each word vector in the text sequence representation T obtained in step four is regarded as a vertex, and a text graph based on sentence-level dependency relationships and a text graph based on co-reference relationships of the entire conversation are constructed according to the Stanford CoreNLP text parsing tool. and a text graph based on co-reference relationships of the entire conversation The two graphs are respectively input into a graph convolutional neural network, and a text hidden layer sequence G is output. D and G C , and the specific calculation formula is the same as the expression in step five of the specific implementation manner one.
[0106] In step six, the text hidden layer sequences G D and G C obtained in step four and step five and the original text sequence representation T are input into a linear layer to obtain a fused representation of the text and used as part of the input for the subsequent GPT2 model; the expression is:
[0107]
[0108] where and are trainable weights.
[0109] Other steps and parameters are the same as those in the specific implementation manner one or two. Specific implementation manner four:
[0111] This implementation manner is a multi-modal dialogue question-answering generation method based on a multi-relational graph model. The difference between this implementation manner and one of the specific implementation manners one to three is that:
[0112] In step eight, to promote the fusion of different modal information, three tasks are introduced for fine-tuning during model training, including a response prediction task (RPT) based on audio-visual, title, and dialogue history features, a caption prediction task (CPT) based on audio-visual, and a video-text matching task (VTMT). The first three specific implementation manners only use one loss function and are in the form of single-task learning. Specific implementation manner four designs three loss functions and uses multi-task learning to enhance the model's ability to understand different modal information.
[0113] The RPT part aims to generate an answer <n based on the audio-visual feature V, title C, dialogue history H n and the current question Q By minimizing the negative log-likelihood loss function, the probability that the next word output by the model is the corresponding word in the source sequence is maximized:
[0114]
[0115] Among them, represents the first j - 1 words of the response R, θ refers to the trainable model parameters, the set (V, C, H, Q) is sampled from the entire training set D, and E n represents the expectation. (V,C,H,Q,R)~D represents the expectation.
[0116] The CPT part is similar to the RPT part. For the given audio - visual feature V, the caption C = {c 1 , c 2 ,..., c L} is generated by minimizing the negative log - likelihood loss function. The loss function is as follows:
[0117]
[0118] Among them, c <i represents the first i - 1 words of the caption C.
[0119] The VTMT part aims to determine whether the given audio - visual feature V and the given text features (including the caption C, the dialogue history H <n , the current question Q n and the generated response R n ) match, so as to successfully apply it to the dialogue domain task by fine - tuning the pre - trained language model. Specifically, this task selects about 15% of the training data, randomly replaces the corresponding original input with incorrect audio - visual features, and obtains the probability of matching or not by passing the final output of the hidden state of the GPT2 module through a linear fully - connected layer, and then calculates the loss function using binary cross - entropy to enhance the system's understanding of the scenario. The calculation formula is as follows:
[0120]
[0121] Among them, X = (V, C, H, Q, R), and Y is the label representing whether the audio - visual feature and the text feature match.
[0122] Other steps and parameters are the same as those in any one of the specific embodiments one to three.
[0123] The following embodiments are used to verify the beneficial effects of the present invention:
[0124] Embodiment 1:
[0125] The dataset selects the audio-visual scene perception dialogue dataset of the 7th Dialogue System Technology Challenge (DSTC7) released by Hori et al. in ICASSP2019 for system performance evaluation. To ensure the fairness and reasonableness of measuring the performance differences between different models, the dataset partitioning method is consistent with the task settings in the challenge. The size and partitioning method of this dataset are shown in Table 1.
[0126] Table 1 Overview of the DSTC7-AVSD Dataset
[0127]
[0128] The evaluation metrics are selected from the commonly used metrics in natural language generation tasks, including BLEU, METEOR, ROUGE-L, and CIDEr, etc. These metrics can calculate the semantic similarity and language fluency between the predicted answer and the true answer from different perspectives, so as to scientifically reflect the system performance.
[0129] The experimental parameters are set as shown in Table 2. Specifically, during the encoding process, the learning rate of the Adam optimizer is set to 6.25e-5, the maximum number of dialogue history rounds involved is 3, the hidden state of the Transformer module is 768, and the batch size is 8. During the decoding process, the beam search algorithm is adopted, the beam width is set to 5, the maximum sentence length is 20, and the length penalty is 0.3.
[0130] Table 2 Experimental Parameter Settings
[0131]
[0132] Table 3 compares the differences between the baseline model and the generation results of the present invention on DSTC7-AVSD. It can be seen that in this example, the "television" mentioned by the questioner does not appear in the title and dialogue history. Therefore, the system needs to combine audio-visual information and perform simple reasoning to correctly answer this question. At this time, the answer of the baseline model to this question fully shows that the baseline model does not fully understand the question direction of the questioner, lacks reasoning ability, and cannot give the correct answer to questions where specific information cannot be found in the title, abstract, or dialogue history, and even gives an irrelevant answer.
[0133] Compared with the baseline model, the present invention can fully interact the information of the given video and text, thereby capturing the complex hidden dependencies between different modality inputs, extracting richer feature representations, and generating high-quality and natural answers based on reasoning.
[0134] Table 3 DSTC7-AVSD Example Generated by the VGPT Model
[0135]
[0136]
[0137] To objectively and comprehensively verify the effectiveness of the present invention, for the DSTC7-AVSD dataset, it is compared with relevant baseline methods. The specific results are shown in Table 4, where the optimal results for each metric are shown in bold:
[0138] (1) The Naive Fusion model proposed by Hori et al. in ICASSP2019 is a multi-modal baseline method provided by the DSTC7 organizers. It uses an LSTM model with question orientation to extract video and audio features respectively, and uses a hierarchical LSTM to encode the dialogue history. Finally, all modalities are combined by a projection matrix to generate answers.
[0139] (2) The Hierarchical Attention (HA) model proposed by Sanabria et al. in AAAI2019 introduces transfer learning for the video summarization task to obtain more visual details and won the first place in the DSTC7-AVSD challenge.
[0140] (3) The Multimodal Transformer Networks (MTN) proposed by Le et al. in ACL2019 is the highest-level system before the DSTC8-AVSD challenge. It uses an auto-encoding module based on Transformer to focus on visual features with questions as the guidance.
[0141] (4) The Universal Multimodal Transformer (UMT) proposed by Li et al. in TASLP2021 is the most advanced dialogue question-answering system for this task at present. It introduces the pre-trained GPT-2 model and uses multi-task learning to learn the fusion representation of the audio-visual scene.
[0142] Table 4 Objective evaluation results based on the DSTC7-AVSD dataset
[0143]
[0144] The experimental results show that the present invention using the third specific implementation mode is superior to the existing methods in almost all automated metrics of the DSTC7-AVSD test set, and on the BLEU-2, BLEU-2, BLEU-2 and CIDEr metrics, it has an average improvement of 1% compared to the current state-of-the-art model UMT in this task. This indicates that by introducing the multi-relational graph structure encoding, the dialogue system can generate higher-quality answers and significantly improve the performance of the model. Thanks to the structural characteristics of the graph convolutional neural network, various syntactic and semantic information between all words can be presented within one framework. Compared with the Multilayer Perceptron (MLP), when calculating the representation of the current node, the information of its neighbor nodes can be comprehensively considered, and for distant connected points, the "receptive field" can be expanded by stacking multiple layers of graph convolutional neural networks.
[0145] Example Two:
[0146] The dataset selects the audio-visual scene perception dialogue dataset of the 8th Dialog System Technology Challenge (DSTC8) published by Kim et al. in TASLP2021 for system performance evaluation. To ensure the fairness and rationality of measuring the performance differences between different models, the dataset division method is consistent with the task settings in the challenge.
[0147] The size and division method of this dataset are shown in Table 5.
[0148] Table 5 Overview of the DSTC8-AVSD Dataset
[0149]
[0150] The experimental parameter settings are the same as those in Table 2.
[0151] To objectively and comprehensively verify the effectiveness of the present invention, it is compared with relevant baseline methods for the DSTC8-AVSD dataset. The specific results are shown in Table 6, where the optimal results for each metric are presented in bold:
[0152] (1) The Multi-step Joint-Modality Attention Network (JMAN) proposed by Chu et al. in arXiv2020 designed a model architecture based on the recurrent neural network, applied the multi-step attention mechanism, and considered the representations of both vision and text in each inference process to better integrate the information of the two different modalities.
[0153] (2) Compared with the traditional Transformer architecture, the Multimodal Semantic Transformer Network (MSTN) proposed by Lee et al. in arXiv2020 additionally designs an attention-based word embedding layer, enabling the model to more comprehensively consider word meanings during the generation phase.
[0154] Table 6 Objective evaluation results based on the DSTC8-AVSD dataset
[0155]
[0156]
[0157] The experimental results show that the present invention, using the first specific implementation manner, outperforms the existing models in almost all automated metrics of the DSTC8-AVSD test set. Among them, the improvement in the CIDEr metric, which characterizes the naturalness of sentences, is particularly significant, with an increase of 0.012 (1.240 vs. 1.252). This indicates that local dependency relationships and global co-reference relationships can reflect the functional similarity of texts from different perspectives to encode text information, thereby improving the performance of existing models.
[0158] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention. However, these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A multi-modal dialogue question-answering generation method based on a multi-relational graph model, characterized in that, it includes the following steps: S1. Serialize and segment the video into multiple video clips using a sliding window of fixed size. For each clip, obtain the color features of the clip and optical flow features as well as audio features The color features The optical flow features and the audio features are concatenated to obtain Then, the location information and the modality information are added to obtain the sequence representation V of each video segment t ; The expression is: Among them, the position information uses numbers to represent the order in which each video segment appears. The modality information uses the identifier [video] to uniformly identify video features, which are converted into vectors of a fixed dimension during actual calculation; S2. For the audiovisual scene representation \(V=(V 1 ,V 2 ,\cdots,V m ), V 1 ,V 2 ,\cdots,V m is the sequence representation of each video segment. Regarding each video segment as a vertex, a video graph based on the fully connected relationship is constructed where is the set of directed dependency edges. For each directed dependency edge \((V i ,V j ,l ij ), l ij represents the dependency relationship from \(V i to \(V j and is set to 1; Input the video graph into the graph convolutional neural network and output the video hidden layer sequence G V ; S3. Input the video hidden layer sequence G V and the original video sequence representation V into a linear layer to obtain the fused representation of the video and use it as part of the input for the subsequent multi-layer TRANSFORMER model based on the GPT-2 architecture; S4. Obtain the corresponding word vector representations C feature and H feature based on the audiovisual scene title C and the dialogue history H; concatenate the title word vector C feature and the dialogue history word vector H feature , then add the position information T pos and the modality information T mod to obtain the text sequence representation T; the expression is: T feature = [C feature , H feature , T = T feature + T mod + T pos , Among them, the position information T pos uses numbers to refer to the order of the appearance of words in the title and each Q&A pair. T pos uses the identifier [cap] to uniformly identify the title of the audio-visual scene, the identifier [usr1] to identify the questioner, and the identifier [usr2] to identify the answerer. During actual calculation, they are respectively converted into vectors with a fixed dimension; S5. Treat each word vector in the text sequence representation T obtained in S4 as a vertex, and construct a graph structure based on sentence-level dependency relationships and / or a graph structure based on coreference relationships of the entire conversation Then, the graph structure of the sentence-level dependency relationship and / or the graph structure based on the coreference relationship of the complete conversation are respectively input into the graph convolutional neural network to obtain their respective corresponding text hidden layer sequences; S6. Take the graph structure of the sentence-level dependency relationship and / or the graph structure based on the coreference relationship of the complete conversation The corresponding text hidden layer sequence and the original text sequence are input into the linear layer to obtain the fused representation of the text And use it as part of the input for the subsequent multi-layer TRANSFORMER model based on the GPT-2 architecture; S7. Concatenate and to obtain enhanced multimodal input, and generate an answer for the enhanced multimodal input based on the multi-layer Transformer model of the GPT-2 architecture.
2. The multi-modal dialogue question-answering generation method based on a multi-relational graph model according to claim 1, characterized in that, In S5, each word vector in the text sequence representation T obtained in S4 is regarded as a vertex to construct a graph structure based on sentence-level dependency relationships The process of obtaining the corresponding text hidden layer sequence includes the following steps: First, use the GPT2 Tokenizer to obtain the word vector representation corresponding to each word. Use the Stanford CoreNLP text parsing tool to analyze the syntactic dependency relationship of the sentence. Treat each word vector as a vertex and model the graph structure based on the syntactic dependency relationship. Then, input it into the graph convolutional neural network to output the text hidden layer sequence G D 。 3. The multi-modal dialogue question-answering generation method based on a multi-relational graph model according to claim 1, characterized in that, In S5, each word vector in the text sequence representation T obtained in S4 is regarded as a vertex, and a graph structure based on the coreference relationship of the complete conversation is constructed. The process of obtaining their respective corresponding text hidden layer sequences includes the following steps: First, use the GPT2 Tokenizer to obtain the word vector representation corresponding to each word. Use the Stanford CoreNLP text parsing tool to analyze the coreference relationship of the sentence. Treat each word vector as a vertex and model the graph structure based on the coreference relationship of the sentence. Then input it into the graph convolutional neural network to output the text hidden layer sequence G C 。 4. The multi-modal dialogue question-answering generation method based on a multi-relational graph model according to claim 1, characterized in that, In S5, each word vector in the text sequence representation T obtained in S4 is regarded as a vertex, and a graph structure based on sentence-level dependency relationships and a graph structure based on full-dialogue co-reference relationships are constructed. and a graph structure based on full-dialogue co-reference relationships The process of obtaining the respective corresponding text hidden layer sequences includes the following steps: First, use the GPT2 Tokenizer to obtain the word vector representation corresponding to each word. Use the Stanford CoreNLP text parsing tool to analyze the syntactic dependency relationship and the co-reference relationship of the sentence respectively. Treat each word vector as a vertex, and construct a graph structure based on the sentence-level dependency relationship and a graph structure based on the co-reference relationship of the complete dialogue Then, input the two graph structures into the graph convolutional neural network respectively, and output the text hidden layer sequences G D and G C .
5. The multi-modal dialogue question-answering generation method based on a multi-relational graph model according to claim 1, 2, 3 or 4, characterized in that, The expression calculated by each layer of the graph convolutional neural network in the graph convolutional neural network in S5 is: Among them, f(H v (l+1) , A v ) represents the graph convolution of each layer. For or A d respectively represent their corresponding adjacency matrices. is the corresponding degree matrix. Adding the identity matrix I d obtains Correspondingly, the corresponding degree matrix is obtained for the convenience of the normalization operation. l is the number of layers of the graph convolutional neural network. is the hidden state of the l-th layer of the graph convolutional neural network.
6. The multi-modal dialogue question-answering generation method based on a multi-relational graph model according to claim 5, characterized in that, The expression calculated by each layer of the graph convolutional neural network in the graph convolutional neural network in S2 is: Among them, f(H v (l+1) , A v ) represents graph convolution for each layer; A v is 's adjacency matrix, where i and j respectively represent 's i-th and j-th nodes, is the value of the i-th row and j-th column of A v ; is 's degree matrix, is the value of the i-th row and i-th column of D v ; To enable the model to consider the node's own representation, the identity matrix I v is added to obtain Correspondingly, the corresponding degree matrix is obtained for easy normalization operation; l is the number of layers of the graph convolutional neural network, is the hidden state of the l-th layer of the graph convolutional neural network, V is the representation of the original video sequence, is the trainable weight.
7. The multi-modal dialogue question-answering generation method based on a multi-relational graph model according to claim 6, characterized in that, When processing with a multi-layer Transformer model based on the GPT-2 architecture, it is also input into a linear fully connected layer, and the output result is projected into the same vector space as in S6. The two are concatenated to obtain a complete multi-modal input representation, and then it is input into the pre-trained language model GPT-2.
8. The multi-modal dialogue question-answering generation method based on a multi-relational graph model according to claim 7, characterized in that, The multi-layer Transformer model based on the GPT-2 architecture in S7 is stacked by 12 layers of Transformer decoder modules with masked multi-head attention mechanisms.
9. The multi-modal dialogue question-answering generation method based on a multi-relational graph model according to claim 8, characterized in that, The multi-layer Transformer model based on the GPT-2 architecture is trained using a negative log-likelihood loss function during the training process, and the training process includes the following steps: Based on the audio-visual features V, the title C, the dialogue history H <n and the current question Q n Generate an answer By minimizing the negative log-likelihood loss function, maximize the likelihood that the next word output is the corresponding word in the source sequence: Among them, represents the first j - 1 words of the answer R n , θ refers to the trainable model parameters, the set (V, C, H, Q) is sampled from the entire training set D, and E (V,C,H,Q,R)-D represents the expectation.
10. The multi-modal dialogue question-answering generation method based on a multi-relational graph model according to claim 8, characterized in that, The multi-layer Transformer model based on the GPT-2 architecture is jointly trained based on the response prediction task RPT of audio-visual, title and dialogue history features, the caption prediction task CPT of audio-visual, and the audio-visual-text matching task VTMT during the training process, and the training process includes the following steps: The RPT part aims to generate answers based on the audio-visual features V, the title C, the dialogue history H <n and the current question Q n by minimizing the negative log-likelihood loss function to maximize the likelihood that the next word output by the model is the corresponding word in the source sequence: Among them, represents the first j - 1 words of the answer R n , θ refers to the trainable model parameters, the set (V, C, H, Q) is sampled from the entire training set D, and E (V,C,H,Q,R)~D represents the expectation; The CPT part is similar to the RPT part. For a given audio-visual feature V, a caption C = {c 1 , c 2 ,..., c L} is generated by minimizing the negative log-likelihood loss function. The loss function is shown as follows: Among them, c <i represents the first i - 1 words of the title C; The VTMT part aims to determine whether the given audio-visual feature V matches the given text features, where the given text features include the title C, the dialogue history H <n , the current question Q n and the generated answer R n ; Select a certain proportion of the training data, randomly replace the corresponding original input with incorrect audio-visual features, and obtain the probability of matching or not by passing the final output of the hidden state of the GPT2 module through a linear fully connected layer, and then use binary cross-entropy to calculate the loss function: where X = (V, C, H, Q, R), and Y is a label representing whether the audio-visual feature and the text feature match.
Citation Information
Patent Citations
Visual dialogue generation method based on context perceptual map neural network
CN110609891A
Method and apparatus for automatically generating inference questions and answers
WO2021184311A1