Keyword-aware multimodal attention video question answering method and system

By adopting a multimodal attention processing method based on keyword perception in the video Q&A technology, the problem of difficulty in extracting core feature information in the prior art is solved, and a higher video Q&A accuracy and more effective multimodal feature fusion are achieved.

CN113902964BActive Publication Date: 2025-05-23SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111053387.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-09
Publication Date
2025-05-23
Estimated Expiration
2041-09-09

AI Technical Summary

Technical Problem

Existing video Q&A technology is difficult to accurately extract core feature information, which is affected by invalid information, resulting in low accuracy of Q&A.

Method used

The multimodal attention video question-and-answer method is adopted based on keyword perception. Through multimodal feature extraction and keyword extraction algorithm, video frames, subtitle text and question text information are filtered and processed, and the soft attention mechanism, self-attention mechanism and two-way attention mechanism are used to feature association and fusion, and predicted answers are output.

Benefits of technology

The accuracy of video Q&A is significantly improved, and by combining keyword characteristics and multimodal information, video features are more accurately extracted and fused, reducing the impact of redundant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113902964B_ABST
    Figure CN113902964B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal attention video question-answering method and system based on keyword perception. It includes: using multimodal feature extraction and pre-training model KeyBert keyword extraction algorithm to extract various multimodal features of the input video; using the keyword-aware multimodal attention algorithm to process the extracted multimodal features, and output the multimodal features after effective association and fusion; passing the fused multimodal features through a multi-layer perceptron MLP to output the predicted answer. The present invention also discloses a multimodal attention video question-answering computer device and a computer-readable storage medium based on keyword perception. When extracting video features, the present invention combines more implicit keyword features to extract richer video features; when fusing features, it combines the self-attention mechanism to capture the temporal nature of features, and uses the bidirectional attention mechanism to emphasize the information related to each other between modalities, so as to more effectively fuse multimodal features and significantly improve the accuracy of video question-answering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a keyword-aware multimodal attention video question-answering method, a keyword-aware multimodal attention video question-answering system, a computer device, and a computer-readable storage medium. Background Art

[0002] In recent years, with the development of artificial intelligence technology, video question-answering technology has emerged. Video question-answering technology can quickly and effectively predict the corresponding answers based on the content of the video according to the questions asked, thereby helping users quickly understand the video content, obtain the desired video information, and reduce the time people spend screening information in lengthy videos. Traditional visual question-answering technology mainly targets single static images, while videos are composed of a large number of video frames. Videos semantically contain visual, textual, and audio information, and have the characteristics of unstructured, multimodal, temporal, and spatial. Therefore, video question-answering technology must process more input data, and requires specific methods to extract visual content and textual content and effectively fuse them.

[0003] Currently, most video question-answering technology models directly use all video information to answer questions, which makes it difficult to accurately extract core and effective feature information. They are usually affected by invalid and redundant information and have the disadvantage of low video question-answering accuracy, making them difficult to be widely used.

[0004] One of the current existing technologies, the patent "A video question-answering system and method based on action-relational network", uses the results of the temporal action detection network to assist in the encoding of video features, emphasizing the action factors of the video, and then the action probability distribution and the initial video features are input into the encoder of the neural network together to learn the video features so that the final video features can contain action information. Finally, the output video features and question features are input into a multi-head relational converter network, and the final result is output through this network for video question answering. The disadvantage of this technology is that it does not emphasize the interrelated parts of the multimodal features and does not consider the subtitle mode of the video.

[0005] The second existing technology, the patent "An artificial intelligence video question-answering method", first obtains visual features and text features; then extracts visual features, performs multimodal fusion of visual features and semantic features to obtain fusion features; finally, generates answers based on fusion features and semantic features. The disadvantage of this technology is that the feature fusion method used is relatively simple and does not pay much attention to the relevant information between multimodal features.

[0006] The third existing technology, patent "A method and system for improving the accuracy of video question answering based on a multimodal fusion model", inputs the video question answering questions into the trained multimodal fusion model to obtain the answers to the questions; according to the characteristics of the questions, different target entity instances are focused on for different questions to improve the accuracy of the model's answer selection. The disadvantage of this technology is that although the relevant content between the modalities is associated, the implicit feature information is not taken into account, and the key information is not further associated with the keyword features. Summary of the invention

[0007] The purpose of the present invention is to overcome the shortcomings of existing methods and propose a multimodal attention video question-answering method, system, device and storage medium based on keyword perception. The main problem solved by the present invention is that in video question-answering, the answers related to the questions only appear in some sentences or words in the video, while most of the methods in the prior art directly use global video information, resulting in low question-answering efficiency and more redundant information. That is, how to screen and process the input video frames, subtitle texts and question text information through multimodal feature extraction algorithms and keyword extraction algorithms, so as to more accurately output the predicted answer.

[0008] In order to solve the above problems, the present invention proposes a keyword-aware multimodal attention video question answering method, which includes:

[0009] Input video frames, subtitle text and question text information, and use multimodal feature extraction and keyword extraction algorithms to extract multimodal features of the input video;

[0010] Using a keyword-aware multimodal attention algorithm, the multimodal features of the video are processed, and after effective association and fusion, the fused multimodal features are output;

[0011] The multi-layer perceptron MLP is used to process the fused multimodal features and output a predicted answer.

[0012] Preferably, the multimodal features of the input video are extracted by using a multimodal feature extraction and keyword extraction algorithm, specifically:

[0013] Using a convolutional network C3D to extract action labels of the video frames, using an object detection algorithm Yolo to extract visual labels of the video frames, and combining the action labels and the visual labels into a visual label set;

[0014] Integrate the visual label set, question text, and subtitle text into a long sentence, extract keywords using the pre-trained model KeyBert, and output the extracted keyword set;

[0015] Using a pre-trained model BERT and a bidirectional neural network LSTM encoder, the visual label set, the question text, the subtitle text, and the keyword set are processed to obtain an encoding of the text features;

[0016] Input the video frame into the neural network ResNet, directly extract the visual features of the picture corresponding to the video frame, and input it into the LSTM to obtain the visual feature representation;

[0017] The text feature and the visual feature are combined to obtain a multimodal feature.

[0018] Preferably, the keyword-aware multimodal attention algorithm processes the multimodal features of the video, and after effective association and fusion, outputs the fused multimodal features, specifically:

[0019] Using a soft attention mechanism, the keyword feature and the subtitle text feature in the multimodal feature are associated to select the subtitle text that is more relevant to the keyword feature, and the two features are combined into a keyword subtitle text feature;

[0020] Similarly, the keyword feature and the question text feature in the multimodal feature are associated, the question text that is more relevant to the keyword feature is screened out, and the two features are combined into a key question text feature;

[0021] Applying a self-attention mechanism to the multimodal features, key subtitle text features, and key question text features respectively, enhancing the temporal nature of the features, and outputting feature representations of each modality respectively;

[0022] A bidirectional attention mechanism is applied between each modal feature to associate relevant information in different modal features to improve the effect of feature fusion.

[0023] Preferably, after processing the fused multimodal features using MLP, the predicted answer is output, specifically:

[0024] A two-layer MLP is defined as a classifier, and the structure of the classifier is as follows:

[0025] FC(2048)-ReLU-FC(n)

[0026] Among them, FC is the fully connected layer of the neural network, 2048 is the number of neurons; ReLU is the activation function of the neural network, and n is the output dimension of the fully connected layer, which is determined by the number of candidate answers;

[0027] After MLP, the predicted scores for each candidate answer are output as follows:

[0028]

[0029] in, is the fused multimodal feature, x is the prediction score of each candidate answer, x=x 1 ,x 2 ,…,x n ;

[0030] Normalizing the prediction scores using a softmax function to obtain the prediction probability of each candidate answer;

[0031] Use the argmax function to select the maximum predicted probability among all the candidate answers, as follows:

[0032] y = argmax(softmax(x))

[0033] Wherein, y is the maximum value of the predicted probability;

[0034] During training, the cross entropy loss function is used to measure the gap between the model output and the actual output. The specific formula is as follows:

[0035]

[0036] Among them, x is a sample, probability distribution p is the expected output of the true answer, and probability distribution q represents the actual output; the closer the two probability distributions are, the smaller the value of the loss function H(p,q), and the closer the actual output when predicting the answer is to the expected output of the true answer; conversely, the farther the two probability distributions are, the larger the value of the loss function H(p,q), and the more the actual output when predicting the answer deviates from the expected output of the true answer.

[0037] Accordingly, the present invention also provides a keyword-aware multimodal attention video question-answering method and system, comprising:

[0038] A multimodal feature extraction unit, used to extract multimodal features of an input video;

[0039] A keyword and subtitle text feature fusion unit, used to associate the extracted keyword feature with the subtitle text feature, filter out the subtitle text that is more relevant to the keyword feature, and combine the two features into one keyword and subtitle text feature;

[0040] A key question text feature fusion unit is used to associate the extracted keyword features with the question text features, filter out question texts that are more relevant to the keyword features, and combine the two features into one key question text feature;

[0041] A multimodal feature fusion unit, used to apply a self-attention mechanism to the multimodal features, key subtitle text features and key question text features respectively, to enhance the temporal nature of the features, and to output feature representations of each modality respectively;

[0042] The answer prediction unit is used to process the fused multimodal features and output a predicted answer.

[0043] Correspondingly, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the steps of the above-mentioned video question-and-answer method.

[0044] Correspondingly, the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned video question-and-answer method are implemented.

[0045] The implementation of the present invention has the following beneficial effects:

[0046] In the feature extraction of video question and answer, the present invention combines the more implicit feature of keywords, extracts richer video features, and significantly improves the accuracy of video question and answer. In the feature fusion of video question and answer, the soft attention mechanism is applied to the information between the associated keyword set and the subtitle text, as well as the associated keyword set and the question text, combined with the self-attention mechanism to capture the temporal nature of the features, and the bidirectional attention mechanism is applied to emphasize the information related to each other between the modalities, so as to more effectively fuse the multimodal features. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is an overall flow chart of a keyword-aware multi-modal attention video question answering method according to an embodiment of the present invention;

[0048] Figure 2 is a flow chart of a multimodal feature representation part of an embodiment of the present invention;

[0049] Figure 3 is a flow chart of a keyword extraction part of an embodiment of the present invention;

[0050] Figure 4 is a multimodal attention flow chart of keyword perception in an embodiment of the present invention;

[0051] Figure 5 is a flow chart of a question answer prediction part of an embodiment of the present invention;

[0052] Figure 6 It is a structural diagram of a keyword-aware multimodal attention video question-answering system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] Figure 1 is an overall flow chart of the keyword-aware multi-modal attention video question answering method according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0055] S1, in this embodiment of the present invention, a video and a question to be answered (including a video frame, a subtitle text, and a question text) are input, and a multimodal feature extraction algorithm and a keyword extraction algorithm are used to extract multimodal features of the input video;

[0056] S2, using a keyword-aware multimodal attention algorithm to process the multimodal features of the video, and after effective association and fusion, output the fused multimodal features;

[0057] S3, using MLP, the fused multimodal features (set as ) is processed and the predicted answer is output.

[0058] Step S1, such as Figure 2 As shown, the details are as follows:

[0059] S1-1, extracting an action label of the video frame using C3D, extracting a visual label of the video frame using Yolo, and combining the action label and the visual label into a visual label set;

[0060] S1-2, such as Figure 3 As shown, the visual label set, question text and subtitle text are integrated into a long sentence, and then the long sentence is input into the KeyBert pre-training model for keyword extraction, and a set of extracted keywords is output;

[0061] S1-3, inputting the visual label set, question text, subtitle text and keyword set into BERT and bidirectional LSTM encoders for processing respectively to obtain the encoding of the text features;

[0062] S1-4, input the video frame into ResNet, directly extract the visual features of the picture corresponding to the video frame, and input it into LSTM to obtain the visual feature representation to ensure that the visual signal is not lost.

[0063] Step S2, such as Figure 4 As shown, the details are as follows:

[0064] S2-1, using a soft attention mechanism, associating keyword features and subtitle text features in the multimodal features, screening out subtitle texts that are more relevant to the keyword features, and combining the two features into a keyword subtitle text feature;

[0065] Similarly, the keyword feature and the question text feature in the multimodal feature are associated, the question text that is more relevant to the keyword feature is screened out, and the two features are combined into a key question text feature;

[0066] S2-2, applying a self-attention mechanism to the multimodal features, key subtitle text features, and key question text features, respectively, to enhance the temporal nature of the features, and outputting feature representations of each modality respectively;

[0067] S2-3, applying a bidirectional attention mechanism between the modal features in pairs to associate relevant information in different modal features to improve the effect of feature fusion.

[0068] Step S3, such as Figure 5 As shown, the details are as follows:

[0069] S3-1, the video question answering task in the embodiment of the present invention is regarded as a multi-classification task, and a two-layer MLP is defined as a classifier. The structure of the classifier is:

[0070] FC(2048)-ReLU-FC(n).

[0071] Among them, FC is the fully connected layer of the neural network, 2048 is the number of neurons; ReLU is the activation function of the neural network, and n is the output dimension of the fully connected layer, which is determined by the number of candidate answers;

[0072] Preferably, after the fused multimodal features are passed through MLP, the prediction score for each candidate answer is output, as follows:

[0073]

[0074] in, is the fused multimodal feature, x is the prediction score of each candidate answer, x=x 1 ,x 2 ,…,x n ;

[0075] S3-2, using the softmax function to normalize the prediction score to obtain the prediction probability of each candidate answer, and then using the argmax function to select the maximum value of the prediction probability among all candidate answers, as follows:

[0076] y = argmax(softmax(x))

[0077] Wherein, y is the maximum value of the predicted probability;

[0078] S3-3, the predicted candidate answers in the embodiment of the present invention are regarded as a multi-classification problem in a neural network. During training, a cross entropy loss function is used to measure the gap between the output of the model and the actual output. The specific formula is as follows:

[0079]

[0080] Among them, x is a sample, probability distribution p is the expected output (true answer), and probability distribution q represents the actual output; the closer the two probability distributions are, the smaller the value of the loss function H(p,q), and the closer the actual output when predicting the answer is to the expected output of the true answer; conversely, the farther the two probability distributions are, the larger the value of the loss function H(p,q), and the more the actual output when predicting the answer deviates from the expected output of the true answer.

[0081] Accordingly, the present invention also provides a keyword-aware multimodal attention video question answering method and system, such as Figure 6 As shown, including:

[0082] The multimodal feature extraction unit 1 is used to extract multimodal features of an input video.

[0083] Specifically, C3D is used to extract action labels of the video frames, Yolo is used to extract visual labels of the video frames, and the action labels and visual labels are combined into a visual label set. In this embodiment, visual labels are such as "standing person", "blue shirt", "gray door", etc., and action labels are such as "running", "walking", "picking up", etc.; the visual label set, question text and subtitle text are integrated into a long sentence, and the KeyBert pre-training model is used to extract keywords, and the extracted keyword set is output; BERT and bidirectional LSTM encoders are used to process the visual label set, question text, subtitle text and keyword set to obtain the encoding of the text features; the video frame is input into ResNet, the visual features of the picture corresponding to the video frame are directly extracted, and the visual features are input into LSTM to obtain the visual feature representation; the text features and the visual features are combined to obtain multimodal features.

[0084] The keyword and subtitle text feature fusion unit 2 is used to associate the extracted keyword features with the subtitle text features, filter out the subtitle text that is more relevant to the keyword features, and combine the two features into one keyword and subtitle text feature.

[0085] Specifically, the soft attention mechanism is used to associate the keyword features and the subtitle text features in the multimodal features, screen out the subtitle text that is more relevant to the keyword features, and combine the two features into one keyword subtitle text feature.

[0086] The key question text feature fusion unit 3 is used to associate the extracted keyword features with the question text features, filter out the question text that is more relevant to the keyword features, and combine the two features into one key question text feature.

[0087] Specifically, the soft attention mechanism is used to associate the keyword features and the question text features in the multimodal features, filter out the question text that is more relevant to the keyword features, and combine the two features into a key question text feature.

[0088] The multimodal feature fusion unit 4 is used to apply the self-attention mechanism to the multimodal features, key subtitle text features and key question text features respectively, enhance the temporal nature of the features, and output the feature representation of each modality respectively.

[0089] Specifically, a self-attention mechanism is applied to the multimodal features, key subtitle text features and key question text features respectively to enhance the temporal nature of the features, and the feature representations of each modality are output respectively; a bidirectional attention mechanism is applied between each modal feature to associate relevant information in different modal features to improve the effect of feature fusion.

[0090] The answer prediction unit 5 is used to process the fused multimodal features and output a predicted answer.

[0091] Specifically, the fused multimodal features are processed using MLP, and a predicted answer is output.

[0092] Therefore, in the feature extraction of video question and answer, the present invention combines the more implicit feature of keywords, extracts richer video features, and significantly improves the accuracy of video question and answer; in the feature fusion of video question and answer, the soft attention mechanism is applied to the information between the associated keyword set and the subtitle text, as well as the associated keyword set and the question text, combined with the self-attention mechanism to capture the temporal nature of the features, and the bidirectional attention mechanism is applied to emphasize the information related to each other between the modalities, thereby more effectively fusing multimodal features.

[0093] Accordingly, the present invention further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned video question-answering method when executing the computer program. At the same time, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and the steps of the above-mentioned video question-answering method are implemented when the computer program is executed by the processor.

[0094] The above is a detailed introduction to the keyword-aware multimodal attention video question-answering method, system, device and storage medium provided in the embodiments of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A keyword-aware multimodal attention video question answering method, It is characterized in that The method comprises: Input video frames, subtitle text and question text information, and use multimodal feature extraction algorithm and keyword extraction algorithm to extract multimodal features of the input video; Using a keyword-aware multimodal attention algorithm, the multimodal features of the video are processed, and after effective association and fusion, the fused multimodal features are output; Using a multi-layer perceptron MLP, the fused multimodal features are processed and a predicted answer is output; The input video frame, subtitle text and question text information are extracted using a multimodal feature extraction algorithm and a keyword extraction algorithm to extract multimodal features of the input video, specifically: Extracting action labels of the video frames using a convolutional network C3D, extracting visual labels of the video frames using an object detection algorithm Yolo, and combining the action labels and visual labels into a visual label set; Integrate the visual label set, question text and subtitle text into a long sentence, extract keywords using the pre-trained model KeyBert, and output the extracted keyword set; Using the pre-trained model BERT and the bidirectional neural network LSTM encoder, the visual label set, the question text, the subtitle text and the keyword set are processed to obtain the encoding of the text features; Input the video frame into the neural network ResNet, directly extract the visual features of the picture corresponding to the video frame, and input the bidirectional LSTM to obtain the visual feature representation; Combining the text feature and the visual feature to obtain a multimodal feature; The keyword-aware multimodal attention algorithm processes the multimodal features of the video, and after effective association and fusion, outputs the fused multimodal features, specifically: By using a soft attention mechanism, the keyword feature and the subtitle text feature in the multimodal feature are associated, the subtitle text that is more relevant to the keyword feature is selected, and the two features are combined into a keyword subtitle text feature; Similarly, the keyword feature and the question text feature in the multimodal feature are associated, the question text that is more relevant to the keyword feature is screened out, and the two features are combined into a key question text feature; Applying a self-attention mechanism to the multimodal features, key subtitle text features, and key question text features respectively, enhancing the temporal nature of the features, and outputting feature representations of each modality respectively; A bidirectional attention mechanism is applied between each modal feature to associate relevant information in different modal features to improve the effect of feature fusion.

2. The keyword-aware multimodal attention video question answering method according to claim 1, It is characterized in that Using MLP, after processing the fused multimodal features, the predicted answer is output, specifically: A two-layer MLP is defined as a classifier, and the structure of the classifier is as follows: FC(2048)-ReLU-FC(n) Among them, FC is the fully connected layer of the neural network, 2048 is the number of neurons; ReLU is the activation function of the neural network, and n is the output dimension of the fully connected layer, which is determined by the number of candidate answers; After MLP, the predicted scores for each candidate answer are output as follows: in, is the fused multimodal feature, x is the prediction score of each candidate answer, x=x 1 ,x 2 ,…,x n ; Normalizing the prediction scores using a softmax function to obtain the prediction probability of each candidate answer; Use the argmax function to select the maximum predicted probability among all the candidate answers, as follows: y = srgmax(softmax(x)) Wherein, y is the maximum value of the predicted probability; During training, the cross entropy loss function is used to measure the gap between the model output and the actual output. The specific formula is as follows: Among them, x is a sample, probability distribution p is the expected output of the true answer, and probability distribution q is the actual output; the closer the two probability distributions are, the smaller the value of the loss function H(p,q), and the closer the actual output when predicting the answer is to the expected output of the true answer; conversely, the farther the two probability distributions are, the larger the value of the loss function H(p,q), and the more the actual output when predicting the answer deviates from the expected output of the true answer.

3. A keyword-aware multimodal attention video question answering system, It is characterized in that The system comprises: A multimodal feature extraction unit, used to extract multimodal features of an input video; A keyword subtitle text feature fusion unit, used to associate the keyword feature and the subtitle text feature in the multimodal feature, filter out the subtitle text that is more relevant to the keyword feature, and combine the two features into one keyword subtitle text feature; A key question text feature fusion unit is used to associate the keyword feature and the question text feature in the multimodal feature, filter out the question text that is more relevant to the keyword feature, and combine the two features into a key question text feature; A multimodal feature fusion unit, used to apply a self-attention mechanism to the multimodal features, key subtitle text features and key question text features respectively, to enhance the temporal nature of the features, and to output feature representations of each modality respectively; to apply a bidirectional attention mechanism to each modality feature in pairs, and finally to output the fused multimodal features; An answer prediction unit, used to process the fused multimodal features and output a predicted answer; Among them, the multimodal feature extraction unit needs to use C3D to extract the action label of the video frame, use Yolo to extract the visual label of the video frame, and combine the action label and the visual label into a visual label set; integrate the visual label set, question text and subtitle text into a long sentence, use KeyBert to extract keywords, and output the extracted keyword set; use BERT and bidirectional LSTM encoder to process the visual label set, question text, subtitle text and keyword set to obtain the encoding of text features; input the video frame into ResNet, directly extract the visual features of the picture corresponding to the video frame, and input LSTM to obtain the visual feature representation; combine the text features and the visual features to obtain multimodal features.

4. The keyword-aware multimodal attention video question answering system according to claim 3, It is characterized in that The keyword subtitle text feature fusion unit needs to use a soft attention mechanism to associate the keyword features and subtitle text features in the multimodal features, screen out subtitle texts that are more relevant to the keyword features, and combine the two features into one keyword subtitle text feature.

5. The keyword-aware multimodal attention video question answering system according to claim 3, It is characterized in that The key question text feature fusion unit needs to use a soft attention mechanism to associate the keyword features and question text features in the multimodal features, screen out question texts that are more relevant to the keyword features, and combine the two features into one key question text feature.

6. The keyword-aware multimodal attention video question answering system according to claim 3, It is characterized in that The multimodal feature fusion unit needs to apply a self-attention mechanism to the multimodal features, key subtitle text features and key question text features respectively to enhance the temporal nature of the features and output feature representations of each modality respectively; apply a bidirectional attention mechanism to each of the modal features to associate relevant information in different modal features to improve the effect of feature fusion.

7. The keyword-aware multimodal attention video question answering system according to claim 3, It is characterized in that The answer prediction unit needs to use MLP to process the fused multimodal features and output the predicted answer.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 2 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 2 are implemented.

Citation Information

Patent Citations

  • Dynamic attention-based super-network method for fusing answer accuracy of visual questions and answers

    CN112818889A

  • Video information processing method and device based on video information processing model

    CN112861580A