A Video Question Answering Method Based on Descriptive Text
By converting video features into description text and supplementing them with the knowledge base, the problem of inaccurate video Q&A answers in the prior art is solved, and higher accuracy is achieved.
Patent Information
- Application Number
- CN202111439356.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-30
AI Technical Summary
Existing video question-and-answer methods cannot effectively express the information related to the problem in video through high-level neural networks, resulting in inaccurate answer generation.
By converting video features into description text and using the knowledge base to supplement the description content, combining question text for fusion reasoning, video description text is generated and classified to predict answers.
Improve the accuracy of video Q&A, and explore relevant information to a greater extent through single-modal information fusion, avoiding the problem of insufficient information retention in cross-modal fusion.
Smart Images

Figure CN114387537B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of natural language processing and computer vision, and particularly relates to a video question answering method based on descriptive text. Background Art
[0002] Video question answering generally refers to providing a computer with a video and several questions related to the video content, enabling the computer to understand the video content and answer these questions in natural language. In recent years, it has attracted the attention of researchers and is an important multi-modal understanding task with many application scenarios in reality, such as improving the human-computer interaction ability of robots, helping visually impaired people obtain video information, and quickly obtaining video information, etc.
[0003] With the rapid development of neural networks, in recent years, question answering systems based on deep learning have emerged. Early video question answering systems focused on separately processing visual information and text information, and fusing features at the final stage of the model to infer answers. The conference proceedings of CVPR in 2017, TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering, pages 1359-1367 disclose a method of separately extracting the visual information of the video and the question features of the question, and finally fusing the two to infer the answer. This method uses C3D and ResNet neural networks to extract the visual features of the video in the time and space dimensions respectively and fuse them, uses the LSTM network to extract the text features of the question, then encodes the fused visual features and text features, and finally decodes to obtain the answer. After that, many researchers have devoted themselves to improving the question system based on deep learning. A more representative improvement is the use of the attention mechanism and the memory network. However, most current methods fuse the two different modalities of visual information and text information, resulting in insufficient utilization of the corresponding information. Specifically, the existing video question answering methods cannot effectively express the information related to the question in the video through a high-level neural network, so they cannot fully utilize the video information when answering questions, making the generated answers inaccurate. Summary of the Invention
[0004] Aiming at the deficiencies of the existing methods, the present invention proposes a video question answering method based on descriptive text, considering describing the video features through text, supplementing the content of the descriptive text using a knowledge base, then fusing and reasoning the relevant descriptions of the video and the question text, mining the potential association between the question and the description, finally classifying to obtain the predicted answer, and judging the semantic relationship between the predicted answer and the expected answer through the knowledge base.
[0005] Technical solution: To achieve the object of the present invention, the technical solution adopted by the present invention is: A video question answering method based on descriptive text, comprising the following steps:
[0006] Step 1, generate a corresponding prior knowledge base G according to all questions and answers;
[0007] Step 2, extract the frame features and motion features of the video through a convolutional neural network, and represent them as two groups of feature real-value vectors Vf and Vm with fixed dimensions;
[0008] Step 3, extract the common sense feature real-value vector Vc of the video frame through a common sense feature extraction network;
[0009] Step 4, perform attention calculation on the video feature real-value vectors Vf, Vm and the common sense feature real-value vector Vc to obtain the final feature vector representation Φ(V) of the video;
[0010] Step 5, generate a video description text Dw through the description network for the video feature representation Φ(V) and the question text;
[0011] Step 6, extract keywords from the video description text Dw, query the corresponding knowledge in the prior knowledge base G, generate a supplementary description Gw according to the query result, and merge the video description text Dw and the supplementary description Gw to obtain the final description text Vd;
[0012] Step 7, extract features from the video description text Vd and the question text, process the two through a question answering processing network, input the processing result into a classifier to obtain a predicted answer, and combine the prior knowledge base G to judge the accuracy of the prediction, and complete the video question answering based on the descriptive text.
[0013] Preferably, the specific steps of Step 1 are:
[0014] Step 1-1, replace the interrogative words in the question text according to the answer to obtain a corresponding declarative sentence;
[0015] Step 1-2, extract keywords from the question statement, and generate a triple <W, R, W> according to the relationship between words, where W represents a word and R represents the relationship between words;
[0016] Step 1-3, according to the relationship triple <W, R, W>, use W as the node of the graph and R as the edge to construct the prior knowledge base G.
[0017] Preferably, the specific steps of Step 2 are:
[0018] Step 2-1, extract n frames of images at equal intervals according to the number of frames included in the video;
[0019] Step 2-2: According to the extracted n frames of images, use the convolutional neural network ResNet to obtain the frame feature vector representation Vf = {Vf1, Vf2,..., Vf n}, and use the frame feature vector representation Vf of the video as the frame feature real-value vector corresponding to the video, where Vf n represents the frame feature corresponding to the nth video frame;
[0020] Step 2-3: Divide the video into t video segments at equal intervals according to the number of frames contained in the video;
[0021] Step 2-4: According to the divided t video segments, use the convolutional neural network ResNext to obtain the motion feature vector representation Vm = {Vm1, Vm2,..., Vm t}, and use the motion feature vector representation Vm of the video as the motion feature real-value vector corresponding to the video, where Vm t represents the motion feature corresponding to the tth video segment.
[0022] Preferably, the specific steps of Step 3 are as follows:
[0023] Step 3-1: Extract K frames of images of the video by means of sparse sampling;
[0024] Step 3-2: Input the extracted K frames of images into the VC-RNN network pre-trained on the COCO dataset for common sense feature extraction to obtain the corresponding common sense feature vector representation Vc = {Vc1, Vc2,..., Vc k}, and use the common sense feature vector representation Vc as the common sense feature real-value vector corresponding to the video, where Vc k represents the common sense feature corresponding to the kth video frame.
[0025] Preferably, the specific steps of Step 4 are as follows:
[0026] Step 4-1: Use the attention mechanism to calculate the attention of the common sense feature real-value vector Vc to the motion feature real-value vector Vm to obtain the motion feature Φ c (V m ); The expression of Φ c (V m ) is as follows:
[0027]
[0028] Among them, Vf i represents the frame feature corresponding to the ith video frame, represents the attention weight of the ith frame feature to the jth common sense feature calculated by the attention mechanism at each time step t, and n represents the number of common sense features;
[0029] Step 4-2: Using the attention mechanism, calculate the attention of the common sense feature real-value vector Vc to the frame feature real-value vector Vf to obtain the frame feature Φ c (V f ); The expression of Φ c (V f ) is as follows:
[0030]
[0031] where Vm i represents the motion feature corresponding to the i-th video segment, represents the attention weight of the i-th motion feature to the j-th common sense feature calculated by the attention mechanism at each time step t, and n represents the number of common sense features;
[0032] Step 4-3: Concatenate the common sense feature Vc, the motion feature Φ c (V m ), and the frame feature Φ c (V f ) to obtain the final feature representation vector Φ(V) = {v1, v2, …… v t}, where v t represents the video feature representation obtained by the t-th concatenation.
[0033] Preferably, the specific steps of Step 5 are as follows:
[0034] Step 5-1: Process the input question in the form of a word sequence according to the question word text;
[0035] Step 5-2: Use the pre-trained BERT model to convert the question word sequence into a set of word vectors Q = {q1, q2, ……, q n} with a fixed dimension, where q n represents the word vector representation corresponding to the n-th word;
[0036] Step 5-3: At the first time step t, generate the start identifier <sos>And initialize the LSTM network;
[0037] Step 5-4: Calculate based on the relevance of words to obtain the phrase representation matrix P t =[p 1,t ……p j,t , where p j,t represents the phrase representation formed by the j-th to the t-th words;
[0038] Step 5-5: Calculate the relevant video feature weights based on the phrase representation matrix P t to obtain the final video semantic feature representation Φ atten (V) (t) ;
[0039] Step 5-6: Input the video semantic feature Φ atten (V) (t) into the LSTM network to generate the video description word w at this time step;
[0040] Step 5-7: Loop through steps 5-4 to 5-6 to generate video description words until the end flag is generated <eos>The generation of the video description is completed, and finally the corresponding description text Dw = {w1, w2, ……, w n} is obtained, where n represents the number of words.
[0041] In step 5-4, the calculation process of P t is as follows:
[0042]
[0043] A t = ψ p (W t ) = [α 1,t , ……, α j,t
[0044]
[0045] P t = [p 1,t , ……, p j,t
[0046] At any time step t, W t represents the matrix composed of the word vector representations of the description words generated in the previous t - 1 steps, where E wt-1 represents the word vector representation obtained by word2vec for the (t - 1)-th word, ψ p represents the calculation of word attention, A t represents the word attention weight matrix, α j,t represents the attention weight of the 1st to the (t - 1)-th words for generating the phrase p j,t , P t is the phrase representation matrix, and p j,t represents the phrase representation composed of the j-th to the t-th words.
[0047] In step 5-5, the calculation process of Φ atten (V) is as follows:
[0048]
[0049]
[0050]
[0051]
[0052]
[0053] At any time step t, β i,j,t represents the phrase pair p i,t and the video feature v j The relevant score, where σ represents the tanh activation function, U s and s represent the trainable parameter matrix, b s represents the trainable bias, p i,t represents the phrase representation formed by the i-th word to the t-th word, v j represents the j-th video feature representation; represents the video feature representation calculated based on the relevant score, N represents the total number of video features; s i,t represents concatenating p i,t and and using the concatenation result as the word group feature of the phrase pair p i,t ; γ i,t represents the relevant score between the current word group feature and the historical feature, where σ represents the tanh activation function, U d and d represent the trainable parameter matrix, b d represents the trainable bias, h t-1 represents the hidden layer representation at the previous time step t - 1 of the LSTM network; Φ atten (V)(t) is the final feature representation obtained based on the relevant scores γ i,t of all word groups, M represents the number of word groups.
[0054] Preferably, the specific steps of step 6 are as follows:
[0055] Step 6 - 1, extract keywords from the video description text Dw through the jieba segmentation library to obtain the corresponding keyword list Kw = {kw1, kw2,..., kw m}, where kw m represents the m-th keyword, m represents the number of keywords;
[0056] Step 6 - 2, use each word in the keyword list as a node to query associated knowledge, and select the relationship with the highest association weight to generate the corresponding description information Dk = {Dk1, Dk2,..., Dk i}, where Dk i represents the description information corresponding to the i-th keyword;
[0057] Step 6 - 3, merge the video description text Dw and the description information Dk of all keywords as the final video description text Vd.
[0058] Preferably, the specific steps of step 7 are as follows:
[0059] Step 7 - 1, process the question text and the video description text Dw in the form of a word sequence;
[0060] Step 7-2: Use the pre-trained BERT model to convert the question text sequence into a set of word vectors Q = {q1, q2, ……, q n}, where q n represents the word vector representation corresponding to the nth word of the question;
[0061] Step 7-3: Use the pre-trained BERT model to convert the video description text sequence into a set of word vectors D = {d1, d2, ……, d m}, where d m represents the word vector representation corresponding to the mth word of the video description;
[0062] Step 7-4: Input the video description word vector set D and the question word vector set Q into the encoding layer of the question-answering processing network, and use the pre-trained BERT model to extract the sentence-level features of both and obtain the feature-encoded representations Td and Tq;
[0063] Step 7-5: Input the feature-encoded representations Td and Tq into the bidirectional attention layer of the question-answering processing network to obtain the attention matrix X of question-description and the attention matrix Y of description-question;
[0064] Step 7-6: Input the feature-encoded representation Td, the row a of matrix X, and the row b of matrix Y into the model encoding layer of the question-answering processing network, and obtain the final encoded representation M through the pooling layer processing;
[0065] Step 7-7: Input the encoded representation M into the classifier for classification to obtain the initially predicted answer p of the model, and perform knowledge base search based on the predicted answer and the expected answer to judge the correctness of the final generated result, and complete the video question-answering based on the description text.
[0066] In Step 7-5, the expressions and related calculations of matrix X and matrix Y are as follows:
[0067] S i,j = W0[tq i , td j , tq i ⊙td j
[0068]
[0069] X = softmax(S, axis = row) · Td T
[0070] Y = X · (Td T ) -1 ·softmax(S, axis=column) T ·Tq T
[0071] wherein, tq i and td j are respectively the i-th word vector of the encoded question and the j-th word vector of the description. ⊙ represents element-wise multiplication, W0 is a trainable parameter matrix, and LayerNorm represents a regularization layer.
[0072] Step 7-7 includes the following steps:
[0073] Step 7-7-1: Use the predicted answer as a node in the knowledge base G to query its associated information, and use the directly associated t words as its associated word group R pw ={pw1, pw2, ……, pw t}, where pw t represents the t-th word associated with the predicted answer queried from the knowledge base;
[0074] Step 7-7-2: Use the standard answer as a node in the knowledge base G to query its associated information, and use the directly associated t words as its associated word group R rw ={rw1, rw2, ……, rw t}, where rw t represents the t-th word associated with the standard answer queried from the knowledge base;
[0075] Step 7-7-3: Respectively form sentences with the associated word groups R pw and R rw and input them into the BERT model for semantic similarity calculation. If the semantic similarity is greater than the set threshold, the answer prediction is correct; otherwise, the answer prediction fails.
[0076] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:
[0077] The present invention describes the visual features of the video through text, and at the same time uses the prior knowledge base to further expand the knowledge of the description text, avoiding the problem of insufficient retention of cross-modal fusion video semantic information. During the answer reasoning process, relevant information can be mined to a greater extent through the interaction of the text single modality, thereby improving the effect of video question answering. Description of the Drawings
[0078] Figure 1 is a schematic flow chart of the method of the present invention;
[0079] Figure 2 is a schematic flow chart in this embodiment. Detailed Embodiments
[0080] The specific implementation manners of the present invention will be described below to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manners. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0081] As Figure 1 and Figure 2 shown, the present invention provides a video question answering method based on descriptive text, and its implementation method is as follows:
[0082] Step 1, generate a corresponding prior knowledge base G according to all questions and answers, and its implementation method is as follows:
[0083] Step 1-1, replace the interrogative words such as what, who, where in the question text according to the answer to obtain corresponding declarative sentences;
[0084] Step 1-2, extract keywords from the question statement sentence, and generate a triple <W, R, W> according to the relationship between words, where W represents a word and R represents the relationship between words;
[0085] Step 1-3, according to the relationship triple <W, R, W>, use W as the node of the graph and R as the edge to construct the prior knowledge base G.
[0086] Step 2, extract the frame features and motion features of the video through a convolutional neural network, and represent them as two groups of fixed-dimensional feature real-valued vectors Vf and Vm, and its implementation method is as follows:
[0087] Step 2-1, extract n frames of images at equal intervals according to the number of frames contained in the video;
[0088] Step 2-2, according to the n frames of images extracted, use the convolutional neural network ResNet to obtain the frame feature vector representation Vf = {Vf1, Vf2,..., Vf n}, and use the frame feature vector representation Vf of the video as the frame feature real-valued vector corresponding to the video, where Vf n represents the frame feature corresponding to the nth video frame;
[0089] Step 2-3, divide the video into t video segments at equal intervals according to the number of frames contained in the video;
[0090] Step 2-4, according to the t video segments divided, use the convolutional neural network ResNext to obtain the motion feature vector representation Vm = {Vm1, Vm2,..., Vm t}, and use the motion feature vector representation Vm of the video as the real-valued motion feature vector corresponding to the video, where Vm t represents the motion feature corresponding to the t-th video segment.
[0091] Step 3: Extract the real-valued common sense feature vector Vc of the video frame through the common sense feature extraction network. The implementation method is as follows:
[0092] Step 3-1: Extract K frame images of the video by sparse sampling;
[0093] Step 3-2: Input the extracted K frame images into the VC-RNN network pre-trained on the COCO dataset for common sense feature extraction, and obtain the corresponding common sense feature vector representation Vc = {Vc1, Vc2,..., Vc k}, and use the common sense feature vector representation Vc as the real-valued common sense feature vector corresponding to the video, where Vc k represents the common sense feature corresponding to the k-th video frame.
[0094] Step 4: Calculate the attention of the video feature real-valued vectors Vf, Vm, and the common sense feature real-valued vector Vc to obtain the final feature vector representation Φ(V) of the video. The implementation method is as follows:
[0095] Step 4-1: Use the attention mechanism to calculate the attention of the common sense feature real-valued vector Vc to the motion feature real-valued vector Vm, and obtain the motion feature concerned by the common sense feature Φ c (V m );
[0096] In this embodiment, by using the attention mechanism, the real-valued motion feature vector Vm of the video is processed to calculate the weight of the motion feature to the common sense feature at any time t Then, the motion feature vector related to the common sense feature at any time t is obtained by weighting
[0097]
[0098] This makes the motion feature contain richer common sense information, which is beneficial to improving the diversity of the generated content of the description text.
[0099] Step 4-2: Use the attention mechanism to calculate the attention of the common sense feature real-valued vector Vc to the frame feature real-valued vector Vf, and obtain the frame feature concerned by the common sense feature Φ c (V f );
[0100] In this embodiment, by using the attention mechanism, the real-valued frame feature vector Vf of the video is processed to calculate the weight of the frame feature to the common sense feature at any time t Then, a weighted motion feature vector related to the common sense feature at any time t is obtained.
[0101]
[0102] This makes the frame features contain richer common sense information and also helps to improve the diversity of the generated content of the descriptive text.
[0103] Step 4-3: Concatenate the common sense feature Vc, the motion feature Φ c (V m ) that the common sense feature focuses on, and the frame feature Φ c (V f ) that the common sense feature focuses on to obtain the final feature representation vector Φ(V) = {v1, v2, ……, v t}, where v t represents the video feature representation obtained by the t-th concatenation.
[0104] Step 5: Generate the video description text Dw through the description network using the video feature representation Φ(V) and the question text. The implementation method is as follows:
[0105] Step 5-1: Process the input question in the form of a word sequence according to the question word text.
[0106] Step 5-2: Use the pre-trained BERT model to convert the question word sequence into a set of word vectors Q = {q1, q2, ……, q n}, where q n represents the word vector representation corresponding to the n-th word.
[0107] Step 5-3: At the first time step t, generate the start identifier <sos>And initialize the LSTM network;
[0108] Step 5-4: Calculate according to the relevance of words to obtain the phrase representation matrix P t =[p 1,t ……p j,t , where p j,t represents the phrase representation formed by the j-th to the t-th words;
[0109] In this embodiment, considering that the meanings of different word combinations are different, calculate the relevance between words according to the attention between words to obtain the phrase representation matrix P t =[p 1,t ……p j,t and the attention weight matrix A t =[α 1,t ……α j,t . This method can mine the semantic relationship between problem words and helps the model understand the problem semantics.
[0110] Step 5-5: Calculate the relevant video feature weights according to the phrase representation matrix P t to obtain the final video semantic feature representation Φ atten (V) (t) ;
[0111] In this embodiment, by calculating the relevance between different phrases and video features:
[0112]
[0113] And comprehensively select the most relevant video features according to the relevance:
[0114]
[0115] Furthermore, obtain different semantic feature group representations:
[0116]
[0117] Then calculate the relevance between different semantic feature groups and historical features:
[0118]
[0119] Finally, combine all semantic feature groups to obtain the final video semantic feature representation:
[0120]
[0121] This is of great help to the information selection and positioning of video descriptions, enabling the content of video descriptions to capture the core information of both the video and the problem.
[0122] In the above formula, at any time step t, β i,j,t represents the correlation score between the phrase pair p i,t and the video feature v j , where σ represents the tanh activation function. U s and O s represent trainable parameter matrices, b s represents the trainable bias, p i,t represents the phrase representation formed by the i-th word to the t-th word, and v j represents the j-th video feature representation; represents the video feature representation calculated based on the correlation score, and N represents the total number of video features; s i,t represents concatenating p i,t and , and using the concatenation result as the word group feature of the phrase pair p i,t ; γ i,t represents the correlation score between the current word group feature and the historical feature, where σ represents the tanh activation function. U d and O d represent trainable parameter matrices, b d represents the trainable bias, h t-1 represents the hidden layer representation of the previous time step t - 1 of the LSTM network; Φ atten (V)(t) is the final feature representation obtained based on the correlation scores γ i,t of all word groups, and M represents the number of word groups.
[0123] Step 5 - 6: Input the video semantic feature Φ atten (V)(t) into the LSTM network to generate the video description word w at this time step;
[0124] Step 5 - 7: Repeat Steps 5 - 4 to 5 - 6 to generate video description words until the end flag <eos>To complete the generation of the video description, finally obtain the corresponding description text Dw = {w1, w2, ……, w n}, where n represents the number of words.
[0125] Step 6, generate the video description text Dw through the description network using the video feature representation Φ(V) and the question text. The implementation method is as follows:
[0126] Step 6-1, extract keywords from the video description text Dw through the jieba segmentation library to obtain the corresponding keyword list Kw = {kw1, kw2, ……, kw m}, where kw m represents the m-th keyword, and m represents the number of keywords;
[0127] Step 6-2, use each word in the keyword list as a node to query associated knowledge, and select the i relationships with the highest association weights to generate the corresponding description information Dk = {Dk1, Dk2, ……, Dk i}, where Dk i represents the description information corresponding to the i-th keyword;
[0128] Step 6-3, merge the video description text Dw and the description information Dk of all keywords as the final video description text Vd.
[0129] Step 7, extract features from the video description text Vd and the question text, use the question-answering processing network to process both, input the processing results into the classifier to obtain the predicted answer, and combine the knowledge base information to judge the accuracy of the prediction to complete the video question-answering based on the description text. The implementation method is as follows:
[0130] Step 7-1, process the question text and the video description text Dw in the form of a word sequence;
[0131] Step 7-2, use the pre-trained BERT model to convert the question text sequence into a fixed-dimensional word vector set Q = {q1, q2, ……, q n}, where q n represents the word vector representation corresponding to the n-th word of the question;
[0132] Step 7-3, use the pre-trained BERT model to convert the video description text sequence into a fixed-dimensional word vector set D = {d1, d2, ……, d m}, where d m represents the word vector representation corresponding to the n-th word of the video description.
[0133] Step 7-4: Input the video description word vector set D and the question word vector set Q into the encoding layer of the question-answering processing network. Use the pre-trained BERT model to extract the sentence-level features of both and obtain the feature-encoded representations Td and Tq;
[0134] Step 7-5: Input the feature-encoded representations Td and Tq into the bidirectional attention layer of the question-answering processing network to obtain the attention matrix X of question-description and the attention matrix Y of description-question;
[0135] Step 7-6: Input the feature-encoded representation Td, the row a of matrix X, and the row b of matrix Y into the model encoding layer of the question-answering processing network, and obtain the final encoded representation M after passing through the pooling layer;
[0136] Step 7-7: Input the encoded representation M into the classifier for classification to obtain the initially predicted answer p of the model, and perform a knowledge base search based on the predicted answer and the expected answer to judge the correctness of the final generated result, thus completing the video question-answering based on the descriptive text.
[0137] This method represents the visual information of the video through the corresponding descriptive text. During the video description generation process, the attention weights are calculated for the word groups of the question, making the generated description more relevant to the question. In addition, the prior knowledge base is used to supplement the descriptive content, so that the descriptive text contains richer information. During the answer reasoning process, the advantages of single-modal fusion can be fully utilized to avoid the problem of insufficient retention of video features during the multi-modal fusion process. Therefore, the accuracy of the question-answering result can be greatly improved.
[0138] In this embodiment, in order to make the generated result more in line with the reality, the semantic relevance between the predicted answer and the standard answer is selected for comparison, and this semantic relevance is expanded through the knowledge base, which helps to improve the matching ability of the question answer. As long as the semantic of the question answer is similar enough, it is considered a correct prediction, and different forms are acceptable, which also conforms to the actual life.
[0139] The following further describes the present invention.
[0140] This embodiment was tested on the datasets MSVD-QA and MSRVTT-QA. The MSVD-QA dataset is a small dataset with 5K questions, which are derived from 1,970 video clips and can be divided into 5 types: what, who, how, when, and where. The MSRVTT-QA dataset contains 243K questions, which are derived from 10K video clips. The types of questions are similar to those in the MSVD-QA dataset, but its videos contain more complex scenarios and the average duration of the videos is also longer. Since all questions are open-ended answers, all 5 types of questions in the answers are evaluated by accuracy. As shown in Table 1, Table 1 shows the comparison of the effects of this method and existing methods on the MSVD-QA dataset. It can be seen from the table that the effect of this method is better than that of the existing ST-VQA (Spatio-Temporal Attention Network), Co-Mem (Collaborative Attention Memory Network), and HME (Heterogeneous Memory Enhanced Network) methods.
[0141] Table 1
[0142] What Who How When Where All ST-VQA 0.181 0.500 0.838 0.724 0.286 0.313 Co-Mem 0.196 0.487 0.816 0.741 0.317 0.317 HME 0.224 0.501 0.730 0.707 0.429 0.337 This method 0.237 0.513 0.746 0.682 0.413 0.372
[0143] As shown in Table 2, Table 2 shows the comparison of the effects of this method and existing methods on the MSRVTT-QA dataset. It can be seen from the table that the effect of this method is generally better than that of other methods.
[0144] Table 2
[0145] What Who How When Where All ST-VQA 0.245 0.412 0.780 0.765 0.349 0.309 Co-Mem 0.239 0.425 0.741 0.690 0.429 0.32 HME 0.265 0.436 0.824 0.760 0.286 0.33 This method 0.266 0.445 0.824 0.731 0.322 0.364
[0146] In summary, the present invention introduces a video description mechanism and a prior knowledge base. The information contained in the video is no longer represented by the intermediate state of a high-level neural network, but by corresponding descriptive texts. Furthermore, some techniques of text question answering, such as knowledge bases, are applied to video question answering to improve the effect of video question answering. At the same time, multi-level interaction processing is performed on the video description text and the question description to improve the ability of text semantic matching, and thus better predict the answer. The experiment was carried out on two datasets, MSVD-QA and MSRVTT-QA. The results show that the accuracy of the answers of the present invention is improved compared with other existing methods.< / eos> < / sos> < / eos> < / sos>
Claims
1. A video question answering method based on descriptive text, characterized in that: The method includes the following steps: Step 1, generate a corresponding prior knowledge base G according to all questions and answers; Step 2, extract the frame features and motion features of the video through a convolutional neural network, and represent them as two sets of feature real-value vectors Vf and Vm with fixed dimensions; Step 3, extract the common sense feature real-value vector Vc of the video frame through a common sense feature extraction network; Step 4, perform attention calculation on the video feature real-value vectors Vf, Vm and the common sense feature real-value vector Vc to obtain the final feature vector representation Φ(V) of the video; Step 5, generate a video description text Dw through the description network using the video feature representation Φ(V) and the question text; Step 5-1, process the input question in the form of a word sequence according to the question word text; Step 5-2: Use the pre-trained BERT model to convert the problem word sequence into a set of word vectors Q = {q1, q2, ……, q n}, where q n represents the word vector representation corresponding to the nth word; Step 5-3, at the first time step t, generate a start identifier <sos>And initialize the LSTM network; < / sos> Step 5-4, calculate according to the relevance of words to obtain the phrase representation matrix P t = [p 1,t ……p j,t , where p j,t represents the phrase representation formed by the j-th to the t-th words; Step 5-5, calculate the relevant video feature weights according to the phrase representation matrix P t , and obtain the final video semantic feature representation Φ atten (V) (t) ; Step 5-6, input the video semantic feature Φ atten (V)(t) into the LSTM network to generate the video description word w at this time step; Steps 5-7, loop through Steps 5-4 to 5-6 to generate video description words until the end flag is generated <eos>The generation of the video description is completed, and finally the corresponding description text Dw = {w1, w2, ……, w n} is obtained, where n represents the number of words;< / eos> Step 6, extract keywords from the video description text Dw, query the corresponding knowledge in the prior knowledge base G, generate a supplementary description Gw according to the query result, and merge the video description text Dw and the supplementary description Gw to obtain the final description text Vd; Step 7, extract features from the video description text Vd and the question text, process the two using a question-answering processing network, input the processing result into a classifier to obtain a predicted answer, and combine with the prior knowledge base G to judge the accuracy of the prediction, completing video question-answering based on the description text.
2. The video question answering method based on descriptive text according to claim 1, characterized in that: The said Step 1 includes the following steps: Step 1-1, replace the interrogative words in the question text according to the answer to obtain a corresponding declarative sentence; Step 1-2, extract keywords from the question statement sentence, and generate a triple <W, R, W> according to the relationship between words, where W represents a word and R represents the relationship between words; Step 1-3, according to the relationship triple <W, R, W>, use W as the node of the graph and R as the edge to construct the prior knowledge base G.
3. The video question answering method based on descriptive text according to claim 1, wherein: The said Step 2 includes the following steps: Step 2-1, extract n frame images at equal intervals according to the number of frames contained in the video; Step 2-2: According to the extracted n frames of images, use the convolutional neural network ResNet to obtain the frame feature vector representation Vf = {Vf1, Vf2, ……, Vf n}, and use the frame feature vector representation Vf of the video as the frame feature real-value vector corresponding to the video, where Vf n represents the frame feature corresponding to the nth video frame; Step 2-3, divide the video into t video segments at equal intervals according to the number of frames contained in the video; Step 2-4, according to the t video segments divided, use the convolutional neural network ResNext to obtain the motion feature vector representation Vm = {Vm1, Vm2, ……, Vm t}, and use the motion feature vector representation Vm of the video as the motion feature real-value vector corresponding to the video, where Vm t represents the motion feature corresponding to the t-th video segment.
4. The video question answering method based on descriptive text according to claim 1, characterized in that: The said Step 3 includes the following steps: Step 3-1, extract K frame images of the video through a sparse sampling method; Step 3-2: Input the extracted K-frame images into the VC-RNN network pre-trained on the COCO dataset for common sense feature extraction to obtain the corresponding common sense feature vector representation Vc = {Vc1, Vc2, ……, Vc k}, and use the common sense feature vector representation Vc as the real value vector of the common sense features corresponding to the video, where Vc k represents the common sense features corresponding to the k-th video frame.
5. The video question answering method based on descriptive text according to claim 1, characterized in that: The said Step 4 includes the following steps: Step 4-1: Using the attention mechanism, calculate the attention of the common sense feature real-value vector Vc to the motion feature real-value vector Vm to obtain the motion feature Φ concerned by the common sense feature c (V m );The expression of Φ c (V m ) is as follows: Among them, Vf i represents the frame feature corresponding to the i-th video frame, represents the attention weight of the i-th frame feature to the j-th common sense feature calculated by the attention mechanism at each time step t, and n represents the number of common sense features; Step 4-2: Using the attention mechanism, calculate the attention of the frame feature real-value vector Vf by the common sense feature real-value vector Vc to obtain the frame feature Φ concerned by the common sense feature c (V f ); The expression of Φ c (V f ) is as follows: Among them, Vm i represents the motion feature corresponding to the i-th video segment, represents the attention weight of the i-th motion feature to the j-th common sense feature calculated by the attention mechanism at each time step t, and n represents the number of common sense features; Step 4-3: Concatenate the common sense feature Vc and the motion feature Φ c (V m ) that the common sense feature focuses on, and the frame feature Φ c (Vf) that the common sense feature focuses on to obtain the final feature representation vector Φ(V) = {v1, v2, …… v t}, where v t represents the video feature representation obtained from the t-th concatenation.
6. The video question answering method based on descriptive text according to claim 1, wherein: In step 5-4, P t is calculated as follows: W t = [E w1 , ……, E wt-1 A t = ψ p (W t ) = [α 1,t , ……, α j,t P t = [p 1,t , ……, p j,t At any time step t, W t represents a matrix composed of the word vector representations of the descriptive words generated in the previous t - 1 steps, where E wt-1 represents the word vector representation obtained by word2vec for the (t - 1)-th word, ψ p represents the calculation of word attention, A t represents the word attention weight matrix, α j,t represents the attention weights of the 1st to the (t - 1)-th words for generating the phrase p j,t , P t is the phrase representation matrix, p j,t represents the phrase representation formed by the j-th to the t-th words.
7. The video question answering method based on descriptive text according to claim 1, wherein: In the said step 5-5, Φ atten (V) (t) is calculated as follows: At any time step t, β i,j,t represents the relevance score of the phrase pair p i,t and the video feature v j , where σ represents the tanh activation function, U s , O s represent trainable parameter matrices, b s represents the trainable bias, p i,t represents the phrase representation formed by the i-th word to the t-th word, v j represents the j-th video feature representation; represents the video feature representation calculated based on the relevance score, and N represents the total number of video features; s i,t Indicates that the concatenation result of p i,t and is used as the word group feature of phrase pair p i,t ; γ i,t represents the relevance score of the current word group features and historical features, where σ represents the tanh activation function, U d 、O d represent the trainable parameter matrices, b d represents the trainable bias, h t-1 represents the hidden layer representation at the previous time step t - 1 of the LSTM network; Φ atten (V) (t) is the final feature representation obtained according to the relevance scores γ of all word groups i,t and M represents the number of word groups.
8. The video question answering method based on descriptive text according to any one of claims 1-7, characterized in that: The said Step 6 includes the following steps: Step 6-1, extract keywords from the video description text Dw through the jieba word segmentation library to obtain the corresponding keyword list Kw = {kw1, kw2, ……, kw m}, where kw m represents the m-th keyword, and m represents the number of keywords; Step 6-2: Take each word in the keyword list as a node, query the associated knowledge, and select the relationship with the highest association weight to generate the corresponding description information Dk = {Dk1, Dk2, ……, Dk i}, where Dk i represents the description information corresponding to the i-th keyword; Step 6-3, merge the video description text Dw and the description information Dk of all keywords as the final video description text Vd.
9. The video question answering method based on descriptive text according to any one of claims 1-7, characterized in that: The said Step 7 includes the following steps: Step 7-1, process the question text and the video description text Dw in the form of a word sequence; Step 7-2, use the pre-trained BERT model to convert the problem text sequence into a set of word vectors Q = {q1, q2, ……, q n}, where q n represents the word vector representation corresponding to the nth word of the problem; Step 7-3, use the pre-trained BERT model to transform the video description text sequence into a set of word vectors D = {d1, d2, ……, d m}, where d m represents the word vector representation corresponding to the m-th word of the video description; Step 7-4, input the video description word vector set D and the question word vector set Q into the encoding layer of the question-answering processing network, and extract the sentence-level features of the two through a pre-trained BERT model to obtain the feature-encoded representations Td and Tq; Step 7-5, input the feature-encoded representations Td and Tq into the bidirectional attention layer of the question-answering processing network to obtain the attention matrix X of question-description and the attention matrix Y of description-question; Step 7-6: Input the representation Td after feature encoding, row a of matrix X, and row b of matrix Y into the model encoding layer of the question-answering processing network, and obtain the final encoded representation M through the processing of the pooling layer; Step 7-7: Input the encoded representation M into the classifier for classification to obtain the initially predicted answer p of the model, and perform knowledge base search based on the predicted answer and the expected answer to judge the correctness of the final generated result, thus completing the video question-answering based on the descriptive text.
Citation Information
Patent Citations
A combined video description method based on multi-modal features and multi-layer attention mechanism
CN109344288A
Method for solving video question and answer tasks needing common knowledge by using question-knowledge guided progressive space-time attention network
CN110704601A