Video Q&A Method, Device, System and Storage Medium
By splicing and modal fusion processing of video and text feature vectors, and using self-attention and mutual attention mechanisms, the problem of insufficient multimodal information processing capabilities in the existing technology is solved, and a deeper semantic understanding and more accurate answer prediction are achieved.
Patent Information
- Application Number
- CN202211043431.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-08-29
AI Technical Summary
Existing video Q&A technology lacks semantic understanding ability and is difficult to effectively process multimodal information, especially in complex scenarios involving video, language and environment information.
Cross-modal information is obtained by splicing the video feature vectors with the text feature vectors and inputting the pre-trained model for self-attention mechanism learning. Then, the processed eigenvectors are input into the modal fusion model, and deep semantic interaction is performed using the mutual attention mechanism, and finally the correct candidate answer is predicted through the decoding layer.
It has improved the semantic representation ability of the video Q&A system, enhanced the model's inference effect in multimodal scenarios, and is suitable for video retrieval, intelligent Q&A system, assisted driving system and other fields.
Smart Images

Figure CN115391511B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to, but are not limited to, the field of natural language processing technologies, and particularly refer to a video question answering method, apparatus, system, and storage medium. Background Art
[0002] In the current era of mobile Internet and big data, video data on the network has shown explosive growth. As an increasingly rich information carrier, understanding the semantics of videos is a technology for many video intelligent applications, and it has important research significance and practical application value. Video Question Answering (Video QA) is a task of inferring the correct answer from a candidate set given a video clip and a question. With the progress of computer vision and natural language processing, the wide application of video question answering in aspects such as video retrieval, intelligent question answering systems, assisted driving systems, and autonomous driving has received more and more attention. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.
[0004] The embodiments of the present disclosure provide a video question answering method, including:
[0005] Extracting a video feature vector for the input video, and extracting text feature vectors for the question text and candidate answer texts, where the question text is used to describe the question, and the candidate answer texts are used to provide multiple candidate answers; concatenating the video feature vector and the text feature vectors to obtain a concatenated feature vector, and inputting the concatenated feature vector into a first pre-trained model. The first pre-trained model learns the cross-modal information between the video feature vector and the text feature vectors through a self-attention mechanism to obtain an encoded second concatenated feature vector;
[0006] Dividing the second concatenated feature vector into a second video feature vector and a second text feature vector; inputting the second video feature vector and the second text feature vector into a modality fusion model. The modality fusion model processes the second video feature vector and the second text feature vector through a mutual-attention mechanism to obtain a video expression and a text expression, and respectively performs pooling and fusion on the video expression and the text expression to obtain a fusion feature vector;
[0007] Inputting the fusion feature vector into a decoding layer to predict the correct candidate answer.
[0008] In an exemplary embodiment, the extracting a video feature vector for the input video includes:
[0009] Extracting frames from the input video at a preset speed, and using a second pre-trained model to extract video feature vectors for the extracted frames.
[0010] In an exemplary embodiment, extracting text feature vectors for the problem text and the candidate answer text includes:
[0011] Generating a sequence string according to the problem text and the candidate answer text, where the sequence string includes multiple sequences, and each word or character in the problem text and the candidate answer text corresponds to one or more sequences;
[0012] Inputting the sequence string into the first pre-trained model to obtain text feature vectors.
[0013] In an exemplary embodiment, before the method, it further includes:
[0014] Constructing and initializing the first pre-trained model;
[0015] Pre-training the first pre-trained model through multiple self-supervised tasks. The multiple self-supervised tasks include a label classification task, a masked language model task, and a masked frame model task. The label classification task is used for multi-label classification of videos. The masked language model task is used for randomly masking text and predicting the masked words. The masked frame model task is used for randomly masking video frames and predicting the masked frames;
[0016] Calculating the loss of the first pre-trained model through the weighted sum of losses of the multiple self-supervised tasks.
[0017] In an exemplary embodiment, calculating the losses of the label classification task and the masked language model task based on binary cross-entropy, and calculating the loss of the masked frame model task based on noise contrast estimation.
[0018] In an exemplary embodiment, the first pre-trained model is a 24-layer deep Transformer encoder cascaded neural network, with a hidden layer dimension of 1024 and 16 attention heads. The parameters of the first pre-trained model are initialized by the bidirectional encoder representations from Transformers, BERT.
[0019] In an exemplary embodiment, processing the second video feature vector and the second text feature vector through a cross-attention mechanism includes:
[0020] Using the second video feature vector as the query vector, and using the second text feature vector as the key vector and the value vector to perform multi-head attention;
[0021] Using the second text feature vector as the query vector, and using the second video feature vector as the key vector and the value vector to perform multi-head attention.
[0022] In an exemplary embodiment, before the method, the following steps are further included:
[0023] Receiving a voice input from a user;
[0024] Converting the voice input into the question text through speech recognition.
[0025] In an exemplary embodiment, before the method, the following steps are further included:
[0026] Obtaining the question text;
[0027] Generating the candidate answer text corresponding to the question text according to the question text.
[0028] In an exemplary embodiment, the generating the candidate answer text corresponding to the question text according to the question text includes:
[0029] Querying triples matching the question text from a common sense knowledge graph through keyword matching or an attention mechanism model;
[0030] Generating the candidate answer text corresponding to the question text according to the matched triples.
[0031] In an exemplary embodiment, the method further includes:
[0032] Processing the video feature vector and / or the text feature vector so that when the video feature vector is concatenated with the text feature vector, the dimensions of the video feature vector and the text feature vector are the same.
[0033] An embodiment of the present disclosure further provides a video question answering device, including a memory; and a processor coupled to the memory, the processor being configured to execute the steps of the video question answering method as described in any embodiment of the present disclosure based on instructions stored in the memory.
[0034] An embodiment of the present disclosure further provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the video question answering method as described in any embodiment of the present disclosure.
[0035] An embodiment of the present disclosure further provides a video question answering system, including a video question answering device, a monitoring system, a speech recognition device, a voice input device, and a knowledge base, where:
[0036] The monitoring system is configured to obtain one or more monitoring videos, process the monitoring videos according to instruction texts, and output the monitoring videos to the video question answering device;
[0037] The voice input device is configured to receive voice input and output it to the voice recognition device;
[0038] The voice recognition device is configured to convert the voice input into an instruction text or a question text through voice recognition, input the instruction text into the monitoring system, and input the question text into the video question-answering device;
[0039] The knowledge base is configured to store a common sense knowledge graph;
[0040] The video question-answering device is configured to receive a question text and a monitoring video, generate candidate answer texts according to the question text, where the question text is used to describe a question, and the candidate answer texts are used to provide multiple candidate answers; it is also configured to extract a video feature vector from the received monitoring video, extract text feature vectors for the question text and the candidate answer texts, splice the video feature vector and the text feature vectors to obtain a spliced feature vector, input the spliced feature vector into a first pre-trained model, and the first pre-trained model learns cross-modal information between the video feature vector and the text feature vectors through a self-attention mechanism to obtain an encoded second spliced feature vector; divide the second spliced feature vector into a second video feature vector and a second text feature vector; input the second video feature vector and the second text feature vector into a modality fusion model, and the modality fusion model processes the second video feature vector and the second text feature vector using a mutual attention mechanism to obtain a video expression and a text expression, and respectively perform pooling and fusion on the video expression and the text expression to obtain a fusion feature vector; input the fusion feature vector into a decoding layer to predict the correct candidate answer.
[0041] Other aspects will be apparent after reading the accompanying drawings and the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings are used to provide an understanding of the technical solutions of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation to the technical solutions of the present disclosure.
[0043] Figure 1 It is a schematic flow chart of a video question-answering method according to an exemplary embodiment of the present disclosure;
[0044] Figure 2 It is a schematic structural diagram of a vision-language cross-modal model created according to an embodiment of the present disclosure;
[0045] Figure 3 It is a schematic structural diagram of a modality fusion model created according to an exemplary embodiment of the present disclosure;
[0046] Figure 4Schematic diagram of the structure of a common sense knowledge graph for traffic accidents according to an exemplary embodiment of the present disclosure;
[0047] Figure 5 Schematic diagram of the structure of a video question answering device according to an exemplary embodiment of the present disclosure;
[0048] Figure 6 Schematic diagram of the structure of a video question answering system according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0049] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The implementation manners can be implemented in multiple different forms. It is easy for those of ordinary skill in the art to understand the fact that the manners and contents can be transformed into one or more forms without departing from the purpose and scope of the present disclosure. Therefore, the present disclosure should not be construed as being limited only to the contents described in the following implementation manners. Without conflict, the embodiments and features in the embodiments of the present disclosure can be combined arbitrarily with each other.
[0050] In the drawings, sometimes for clarity, the sizes of one or more constituent elements, the thicknesses of layers, or regions are exaggerated. Therefore, one aspect of the present disclosure is not necessarily limited to this size, and the shapes and sizes of multiple components in the drawings do not reflect the true proportions. In addition, the drawings schematically show ideal examples, and one aspect of the present disclosure is not limited to the shapes or values shown in the drawings, etc.
[0051] The ordinal numbers such as "first", "second", "third", etc. in the present disclosure are set to avoid confusion of constituent elements, rather than to limit in terms of quantity. The "multiple" in the present disclosure means two or more quantities.
[0052] Computer vision (CV) and natural language processing (NLP) are two major branches of artificial intelligence, which focus on simulating human intelligence in vision and language. In the past decade, deep learning has greatly promoted the development of unimodal learning in these two fields. Some visual question answering technologies encode the input image through an image encoder, encode the input question through a question encoder, perform dot product combination on the encoded image and question features, and then predict the probability of candidate answers through a fully connected layer. However, these video question answering technologies only consider unilateral modal information and lack semantic understanding ability. And real-world problems often involve multiple modalities. For example, an autonomous driving vehicle should be able to process human commands (language), traffic signals (vision), and road conditions (vision and sound).
[0053] As Figure 1 shown, the embodiments of the present disclosure provide a video question answering method, including the following steps:
[0054] Step 101: Extract video feature vectors for the input video, and extract text feature vectors for the question text and candidate answer texts, where the question text is used to describe the question and the candidate answer texts are used to provide multiple candidate answers; Concatenate the video feature vectors and the text feature vectors to obtain concatenated feature vectors, and input the concatenated feature vectors into a first pre-trained model. The first pre-trained model, through the self-attention mechanism, learns the cross-modal information between the video feature vectors and the text feature vectors to obtain the encoded second concatenated feature vectors;
[0055] Step 102: Divide the second concatenated feature vectors into second video feature vectors and second text feature vectors; Input the second video feature vectors and the second text feature vectors into a modality fusion model. The modality fusion model, through the mutual attention mechanism, processes the second video feature vectors and the second text feature vectors to obtain a video expression and a text expression, and respectively performs pooling and fusion on the video expression and the text expression to obtain fused feature vectors;
[0056] Step 103: Input the fused feature vectors into a decoding layer to predict the correct candidate answer.
[0057] In the video question answering method of the embodiments of the present disclosure, by concatenating the video feature vectors and the text feature vectors and then inputting them into the first pre-trained model, the visual-language cross-modal information is used to enhance the semantic representation ability of video question answering, and the mutual attention mechanism is used to perform deeper semantic interaction on the video and text option pairs, which is expected to perform a deeper semantic understanding of the video content, thereby improving the model inference effect. The video question answering method of the embodiments of the present disclosure can be applied to many application fields such as video retrieval, intelligent question answering systems, assisted driving systems, and autonomous driving.
[0058] In some exemplary embodiments, before step 101, the method further includes:
[0059] Construct and initialize the first pre-trained model;
[0060] Pre-train the first pre-trained model through multiple self-supervised tasks. The multiple self-supervised tasks include a label classification task, a masked language model task, and a masked frame model task. The label classification task is used to perform multi-label classification on the video, the masked language model task is used to randomly mask the text and predict the masked words, and the masked frame model task is used to randomly mask the video frames and predict the masked frames;
[0061] Calculate the loss of the first pre-trained model through the weighted sum of the losses of the self-supervised tasks.
[0062] Figure 2 Schematic diagram of the structure of the visual language cross-modal model created for the embodiments of the present disclosure. As Figure 2As shown, the visual language cross-modal model includes a cascaded first pre-training model, a modality fusion model, and a decoding layer.
[0063] In some exemplary embodiments, the first pre-training model can be a cascaded neural network of 24-layer deep Transformer encoders, with a hidden layer dimension of 1024 and 16 attention heads. The parameters pre-trained by BERT are used to initialize the first pre-training model.
[0064] The Transformer model consists of an encoder and a decoder. The encoder and the decoder each include multiple network blocks. Each network block of the encoder consists of a multi-head attention sub-layer and a feed-forward neural network sub-layer. The structure of the decoder is similar to that of the encoder, except that each network block of the decoder has an additional multi-head attention layer.
[0065] Bidirectional Encoder Representations from Transformers (BERT) is a successful application of Transformer. It utilizes the Transformer encoder and introduces bidirectional masking technology, allowing each language token to attend to other tokens bidirectionally.
[0066] The first pre-training model of the embodiments of the present disclosure adopts a cascaded structure of multiple Transformer encoders. Exemplarily, the first pre-training model can be a cascaded neural network of 24-layer deep Transformer encoders, with a hidden layer dimension of 1024 and 16 attention heads. At this time, the network structure of the first pre-training model is consistent with BERT Large (a natural language pre-training model proposed by Google). Therefore, the pre-training results of natural language pre-training models such as BERT can be directly used to initialize the first pre-training model of the embodiments of the present disclosure.
[0067] In some exemplary embodiments, extracting video feature vectors for the input video includes:
[0068] Frame the input video at a preset speed, and use the second pre-training model to extract video feature vectors from the framed frames.
[0069] The input for video question answering is a video clip, a question, and a set of candidate answers. Different from the processing of images, a continuous video stream can be understood as a set of rapidly played pictures, where each picture is defined as a frame. For video information, we extract frames from the input video at a preset speed (exemplarily, at a speed of 1 frame per second (fps)), and each video is extracted up to N frames at most (exemplarily, N can be 30). We use a pre-trained model EfficientNetB3 in the field of computer vision (CV) to extract video feature vectors (Visual Features). Optionally, in the embodiments of the present disclosure, pre-trained models such as MobileNet or ResNet101 can also be used to extract video feature vectors from the input video.
[0070] In some exemplary embodiments, for the question text and the candidate answer text, extracting text feature vectors includes:
[0071] Generating a sequence string according to the question text and the candidate answer text, the sequence string includes multiple sequences, and each word or character in the question text and the candidate answer text corresponds to one or more sequences;
[0072] Inputting the sequence string into a first pre-trained model to obtain text feature vectors.
[0073] In the embodiments of the present disclosure, for the question text and the candidate answer text, we first perform sequence list marking on them with a vocabulary of size n_vocab, convert each word or character into one or more sequences, and all words and / or characters in the question text and the candidate answer text form a sequence string. Among them, the sequence can be a numerical ID. Exemplarily, the numerical ID can be between 0 and n_vocab - 1. For example, converting "What is the current road type? Trail, Highway, Road, Street" into a sequence string like "2560 2013 1300 100 567...". Inputting this sequence string into the first pre-trained model (here the first pre-trained model can be the initialized first pre-trained model), we obtain text feature vectors of n_word * 1024, where n_word is the number of sequences in the sequence string, and 1024 is the dimension of each text feature vector.
[0074] In the embodiments of the present disclosure, when the video feature vector and the text feature vector are concatenated and input into the first pre-trained model, the input form of the concatenated feature vector can be [CLS]Video_frame[SEP]Question-Answer[SEP], where Video_frame is the video feature vector, Question-Answer is the text feature vector, the [CLS] flag is the first flag of the concatenated feature vector, and the [SEP] flag is used to separate the video feature vector and the text feature vector. Moreover, position embedding vectors and segment embedding vectors identical to those of BERT are added to the input of the first pre-trained model. Among them, the position embedding vectors are used to specify positions in the sequence, and the segment embedding vectors are used to specify the frame segments to which they belong.
[0075] In some exemplary embodiments, before concatenating the video feature vector and the text feature vector, the method further includes: processing the video feature vector and / or the text feature vector so that when the video feature vector and the text feature vector are concatenated, the dimensions of the video feature vector and the text feature vector are the same.
[0076] In the embodiments of the present disclosure, when the video feature vector extracted by the pre-trained model EfficientNetB3 is a 1536-dimensional feature vector, the 1536-dimensional feature vector is reduced in dimension through a fully connected layer, and the feature dimension is reduced to 1024 dimensions, which is the same as the dimension of the text feature vector.
[0077] In the embodiments of the present disclosure, the training of the vision-language cross-modal model includes two stages: the pre-training stage of the first pre-trained model and the fine-tuning stage of the vision-language cross-modal model.
[0078] In the pre-training stage of the first pre-trained model, we use three tasks, namely Tag classify (TC), Mask language model (MLM), and Mask frame model (MFM), to pre-train the first pre-trained model. Among them, the tag classification task is used for multi-label classification of videos, the mask language model task is used for randomly masking text and predicting the masked words, and the mask frame model task is used for randomly masking video frames and predicting the masked frames.
[0079] Exemplarily, in the tag classification task, the top 100 tags with higher occurrence frequencies can be used for the multi-label classification task. Here, the tags are manually annotated video tags. For example, whether an accident occurs, the cause of the accident, the location where the accident occurs, the vehicle type, the weather condition, the time information, etc. The predicted tag of the tag is obtained by connecting the [CLS] corresponding vector of the last layer of Bert to the fully connected layer, and the binary cross-entropy loss (Binary Cross Entropy loss, BCEloss) is calculated with the true tag.
[0080] In the masked language model task, 15% of the random text is masked to predict the masked text. In the multi-modal scenario, combining the information of the video to predict the masked word can effectively fuse multi-modal information.
[0081] In the masked frame model task, 15% of the random video frames are masked, and the masked video frames are filled with all-zero vectors. Since the video feature vector is a continuous real-valued vector without word segmentation (token), it is difficult to perform a classification task similar to the masked language model task. Since it is desired that the predicted frame of the mask is as similar as possible to the masked frame within the range of all frames in the entire batch, therefore, the embodiment of the present disclosure calculates the loss of the masked frame model task based on noise contrastive estimation (NCE, Noise Contrastive Estimation) to maximize the mutual information between the mask frame and the predicted frame.
[0082] Multi-task joint training is adopted. The loss of the total pre-training task uses the weighted sum of the losses of the above three pre-training tasks, Loss = Loss(TC)*a + Loss(mlm)*b + Loss(mfm)*c, where a, b, and c are the weights of the losses of the three pre-training tasks respectively.
[0083] The video question answering method of the embodiment of the present disclosure uses the first pre-training model to represent the information of the video and text modalities. The video feature vector corresponding to the video and the question text are concatenated with the text feature vector corresponding to the candidate answer text and input into the first pre-training model to obtain the encoded second concatenated feature vector. According to the flag bit in the second concatenated feature vector, the second concatenated feature vector can be divided into a second video feature vector and a second text feature vector.
[0084] Suppose V = [v1, v2, …, vm], Q = [q1, q2, …, q3], and A = [a1, a2, …, a3] are the video, question, and answer respectively, where vi, qi, and ai are the corresponding sequences. Let VLP(·) denote the VLP model (i.e., the first pre-trained model). Then the encoded representation E obtained by inputting [CLS]V[SEP]Q + A[SEP] into the VLP model can be expressed as:
[0085]
[0086] Among them, E = [e1, e2, …, e(m + n + k)]. By dividing the encoded video frames and question option vectors into two parts, the second video feature vector and the second text feature vector can be obtained. and
[0087] Next, connect the pre-trained first pre-trained model with the modality fusion model and the decoding layer to fine-tune the overall visual language cross-modal model. The processing process of the first pre-trained model is the same as the previous process and will not be elaborated here. Input the second video feature vector and the second text feature vector into the modality fusion model. The modality fusion model processes the second video feature vector and the second text feature vector through the mutual attention mechanism to obtain the video expression and the text expression, and respectively performs pooling and fusion on the video expression and the text expression to obtain the fused feature vector.
[0088] Figure 3 It is a schematic structural diagram of the modality fusion model created for the embodiments of the present disclosure. As Figure 3 shown, in the modality fusion model, the query vector (Query) comes from one modality, while the key vector (Key) and the value vector (Value) come from another modality. The video question answering method of the embodiments of the present disclosure further aligns and fuses language and images semantically by introducing the mutual attention mechanism.
[0089] In some exemplary embodiments, processing the second video feature vector and the second text feature vector through the mutual attention mechanism includes:
[0090] Taking the second video feature vector as the query vector, and taking the second text feature vector as the key vector and the value vector, to perform multi-head attention (Multi Head Attention);
[0091] Taking the second text feature vector as the query vector, and taking the second video feature vector as the key vector and the value vector, to perform multi-head attention.
[0092] In the embodiments of the present disclosure, the modality fusion model uses the second video feature vector as the Query, the second text feature vector as the Key and Value, and performs multi-head attention. At the same time, it uses the second text feature vector as the Query, the second video feature vector as the Key and Value, and performs multi-head attention. Such swapped attention is one layer. In some exemplary embodiments, multiple layers can be cascaded. Then, the video expression representing the video frame sequence is pooled, and the text expression representing the question option sequence is pooled. Then, the two pooled vectors are fused to obtain a fused feature vector. Then, the fused feature vector is input into the decoding layer to predict the probability that each candidate is the correct answer candidate.
[0093] The above process is expressed by the formula as:
[0094]
[0095]
[0096] MHA(E V ,E QA ,E QA )=concat(head 1 ,...head h );
[0097] MHA 1 =MHA(E V ,E QA ,E QA );
[0098] MHA 2 =MHA(E QA ,E V ,E V );
[0099] DUMA(E V ,E QA )=Fuse(MHA 1 ,MHA 2 );
[0100] Among them, The parameter matrix, MHA(·) is the multi-head attention mechanism, DUMA(·) is the bidirectional multi-head co-attention network, and Fuse(·) represents average pooling and fusion of the output results of DUMA. Exemplarily, fusion can be performed by concatenation.
[0101] For each <V, Q, A i > triple, the output after passing through the DUMA network is:
[0102]
[0103] The loss function of the network is as follows:
[0104]
[0105] Among them, W T is the parameter matrix, and s is the number of candidate answers.
[0106] In some exemplary embodiments, the method further includes:
[0107] Receiving a voice input from the user;
[0108] Converting the voice input into question text through speech recognition.
[0109] In some exemplary embodiments, the method further includes:
[0110] Receiving a voice input from the user;
[0111] Converting the voice input into a voice command and / or question text through speech recognition.
[0112] Exemplarily, the voice command can be "Help me switch to the surveillance video of XXXX Road", etc. The system performs corresponding response operations according to the received voice command. For example, it switches the display screen to the surveillance video of XXXX Road. Exemplarily, the question text can be "Is there an accident on the current road section?", "What is the cause of the accident?", etc. The system generates candidate answer texts according to the received question text, and then predicts the correct candidate answer through the video question-answering method of this disclosure embodiment.
[0113] In some exemplary embodiments, the method further includes:
[0114] Announcing the predicted correct candidate answer in voice form through speech synthesis.
[0115] In the embodiments of the present disclosure, the predicted correct candidate answer can be directly displayed on the display screen, or the predicted correct candidate answer can be announced in voice form.
[0116] In some exemplary embodiments, the method further includes:
[0117] Obtaining the question text;
[0118] Generating candidate answer texts corresponding to the question text according to the question text.
[0119] In some exemplary embodiments, generating candidate answer texts corresponding to the question text according to the question text includes:
[0120] Query triples matching the question text from the knowledge graph through keyword matching or an attention mechanism model;
[0121] Generate candidate answer texts corresponding to the question text based on the matching triples.
[0122] With the explosive growth of knowledge, the concept of knowledge graph is applied more and more widely. A knowledge graph can use appropriate knowledge representation methods through the relationship layer to mine the connections between data, making knowledge more easily circulated and collaboratively processed among computers. A knowledge graph represents the relationships between entities in a graph data structure. Compared with pure text information, the graph representation method is more understandable and acceptable. A knowledge graph is composed of relationship edges of "entity-relationship-entity" or "entity-attribute-attribute value", focusing on representing the relationships between entities or between entities and attributes. Relationship search based on knowledge graphs is widely used in entity matching and question-and-answer systems in various industries. It can process and store known knowledge and can quickly perform knowledge matching and answer search.
[0123] The basic composition of a knowledge graph is a triple (S, P, O). Among them, S and O are nodes in the knowledge graph, representing entities. S specifically represents the subject, and O specifically represents the object. P is the edge connecting the two entities (S and O) in the knowledge graph, representing the relationship between the two entities. For example, in a traffic accident common sense knowledge graph, a triple can be (rear-end accident, accident type, traffic accident), etc.
[0124] Exemplarily, the video question-and-answer method of the embodiments of the present disclosure can be used in a video surveillance system. For example, in a multi-screen surveillance system for intelligent transportation, the large screen can be scheduled and interactively questioned through voice control, and the surveillance of an accident can be switched to the main screen for accident analysis. For example, the corresponding interaction instructions can be "Help me switch to the surveillance video of XXXX Road"; "Is there an accident on the current road section?"; "What is the cause of the accident?", etc. In the specific implementation process, it will involve traffic accident-related common sense knowledge, which can be realized by introducing a knowledge base. For example, a traffic accident common sense knowledge graph as shown in Figure 4 can be constructed in the knowledge base, and triples matching the question text are queried from the traffic accident common sense knowledge graph through keyword matching or an attention mechanism, as the common sense knowledge (i.e., candidate answers) corresponding to the displayed semantic information. For example, according to the Figure 4 traffic accident common sense knowledge graph, for the user's question "What is the cause of the accident?", candidate answers are automatically generated: weather, road, driver, poor vehicle condition.
[0125] An embodiment of the present disclosure further provides a video question answering device, including a memory; and a processor coupled to the memory, the processor being configured to execute the steps of the video question answering method as described in any embodiment of the present disclosure based on instructions stored in the memory.
[0126] As Figure 5 shown, in one example, the video question answering device may include: a processor 510, a memory 520, and a bus system 530. Among them, the processor 510 and the memory 520 are connected through the bus system 530. The memory 520 is used to store instructions, and the processor 510 is used to execute the instructions stored in the memory 520 to extract a video feature vector for the input video, and extract text feature vectors for the question text and the candidate answer text, where the question text is used to describe the question, and the candidate answer text is used to provide multiple candidate answers; splice the video feature vector and the text feature vector to obtain a spliced feature vector, input the spliced feature vector into a first pre-trained model, and the first pre-trained model learns the cross-modal information between the video feature vector and the text feature vector through a self-attention mechanism to obtain an encoded second spliced feature vector; divide the second spliced feature vector into a second video feature vector and a second text feature vector; input the second video feature vector and the second text feature vector into a modality fusion model, and the modality fusion model processes the second video feature vector and the second text feature vector through a mutual-attention mechanism to obtain a video expression and a text expression, and perform pooling and fusion on the video expression and the text expression respectively to obtain a fusion feature vector; input the fusion feature vector into a decoding layer to predict the correct candidate answer.
[0127] It should be understood that the processor 510 may be a central processing unit (CPU), and the processor 510 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0128] The memory 520 may include a read-only memory and a random access memory, and provide instructions and data to the processor 510. A part of the memory 520 may also include a non-volatile random access memory. For example, the memory 520 may also store information about the device type.
[0129] In addition to including a data bus, the bus system 530 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, in Figure 5All kinds of buses are labeled as bus system 530.
[0130] In the implementation process, the processing performed by the processing device can be completed by the integrated logic circuit of the hardware in the processor 510 or the instructions in the form of software. That is, the method steps of the embodiments of the present disclosure can be embodied as being executed and completed by the hardware processor, or by a combination of the hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 520, and the processor 510 reads the information in the memory 520 and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0131] As Figure 6 shown, the embodiments of the present disclosure further provide a video question-answering system, including the video question-answering device described in any embodiment of the present disclosure, and further including a monitoring system, a speech recognition device, a speech input device, and a knowledge base, wherein:
[0132] The monitoring system is configured to obtain one or more monitoring videos, process the monitoring videos according to the instruction text, and output the monitoring videos to the video question-answering device;
[0133] The speech input device is configured to receive speech input and output it to the speech recognition device;
[0134] The speech recognition device is configured to convert the speech input into an instruction text or a question text through speech recognition, input the instruction text into the monitoring system, and input the question text into the video question-answering device;
[0135] The knowledge base is configured to store a knowledge graph;
[0136] A video question answering device is configured to receive a question text and a surveillance video, generate candidate answer texts according to the question text, where the question text is used to describe a question, and the candidate answer texts are used to provide multiple candidate answers; it is also configured to extract a video feature vector from the received surveillance video, extract text feature vectors for the question text and the candidate answer texts, splice the video feature vector and the text feature vectors to obtain a spliced feature vector, input the spliced feature vector into a first pre-trained model, and the first pre-trained model learns cross-modal information between the video feature vector and the text feature vectors through a self-attention mechanism to obtain an encoded second spliced feature vector; divide the second spliced feature vector into a second video feature vector and a second text feature vector; input the second video feature vector and the second text feature vector into a modality fusion model, and the modality fusion model uses a mutual attention mechanism to process the second video feature vector and the second text feature vector to obtain a video expression and a text expression, and respectively perform pooling and fusion on the video expression and the text expression to obtain a fusion feature vector; input the fusion feature vector into a decoding layer to predict the correct candidate answer.
[0137] In some exemplary embodiments, the video question answering system further includes a voice synthesis output device, where:
[0138] The voice synthesis output device is used to broadcast the predicted correct candidate answer in a voice form through voice synthesis.
[0139] The video question answering system of the embodiments of the present disclosure is composed of Figure 6 the several modules shown. Among them, a voice instruction or question is given by the user side through a voice input device, the voice recognition device converts natural language into text, the voice synthesis output device broadcasts the answer fed back by the system in a voice form, and the video question answering device performs semantic understanding and multi-modal interaction reasoning according to the question or instruction passed in by the user in combination with the monitoring system and the knowledge base to give a corresponding answer.
[0140] The embodiments of the present disclosure also provide a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the video question answering method as described in any embodiment of the present disclosure.
[0141] In some possible implementation manners, each aspect of the video question answering method provided in the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the video question answering method according to various exemplary embodiments of the present application described above in this specification. For example, the computer device can execute the video question answering method recorded in the embodiments of the present application.
[0142] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0143] The drawings in the present disclosure only relate to the structures involved in the present disclosure, and other structures can refer to the general design. Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0144] Those of ordinary skill in the art should understand that the technical solutions of the present disclosure can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present disclosure, and all should be covered within the scope of the claims of the present disclosure.
Claims
1. A video question answering method, characterized in that, comprising: extracting a video feature vector for the input video, and extracting text feature vectors for the question text and candidate answer texts, wherein the question text is used to describe the question, and the candidate answer texts are used to provide multiple candidate answers; concatenating the video feature vector and the text feature vectors to obtain a concatenated feature vector, and inputting the concatenated feature vector into a first pre-trained model, the first pre-trained model learning cross-modal information between the video feature vector and the text feature vectors through a self-attention mechanism to obtain an encoded second concatenated feature vector; dividing the second concatenated feature vector into a second video feature vector and a second text feature vector; inputting the second video feature vector and the second text feature vector into a modality fusion model, the modality fusion model processing the second video feature vector and the second text feature vector through a mutual-attention mechanism to obtain a video expression and a text expression, and respectively performing pooling and fusion on the video expression and the text expression to obtain a fusion feature vector; inputting the fusion feature vector into a decoding layer to predict the correct candidate answer.
2. The video question answering method according to claim 1, characterized in that, the extracting a video feature vector for the input video includes: extracting frames from the input video at a preset speed, and using a second pre-trained model to extract video feature vectors for the extracted frames.
3. The video question answering method according to claim 1, characterized in that, the extracting text feature vectors for the question text and candidate answer texts includes: generating a sequence string according to the question text and candidate answer texts, the sequence string including multiple sequences, and each word or character in the question text and candidate answer texts corresponding to one or more sequences; inputting the sequence string into the first pre-trained model to obtain text feature vectors.
4. The video question answering method according to claim 1, characterized in that, before the method, it further includes: constructing and initializing the first pre-trained model; pre-training the first pre-trained model through multiple self-supervised tasks, the multiple self-supervised tasks including a label classification task, a masked language model task, and a masked frame model task, the label classification task being used for multi-label classification of videos, the masked language model task being used for randomly masking texts and predicting masked words, and the masked frame model task being used for randomly masking video frames and predicting masked frames; calculating the loss of the first pre-trained model through the weighted sum of losses of the multiple self-supervised tasks.
5. The video question answering method according to claim 4, characterized in that, calculating the losses of the label classification task and the masked language model task based on binary cross-entropy, and calculating the loss of the masked frame model task based on noise contrast estimation.
6. The video question answering method according to claim 1, characterized in that, The first pre-trained model is a 24-layer deep Transformer encoder cascaded neural network with a hidden layer dimension of 1024 and 16 attention heads, and the parameters pre-trained by the Bidirectional Encoder Representations from Transformers (BERT) are used to initialize the first pre-trained model.
7. The video question answering method according to claim 1, wherein, processing the second video feature vector and the second text feature vector through a mutual attention mechanism includes: using the second video feature vector as a query vector, and using the second text feature vector as a key vector and a value vector to perform multi-head attention; using the second text feature vector as a query vector, and using the second video feature vector as a key vector and a value vector to perform multi-head attention.
8. The video question answering method according to claim 1, wherein, before the method, it further includes: receiving a voice input from a user; converting the voice input into the question text through speech recognition.
9. The video question answering method according to claim 1, wherein, before the method, it further includes: obtaining the question text; generating the candidate answer text corresponding to the question text according to the question text.
10. The video question answering method according to claim 9, wherein, generating the candidate answer text corresponding to the question text according to the question text includes: querying triples matching the question text from a common sense knowledge graph through keyword matching or an attention mechanism model; generating the candidate answer text corresponding to the question text according to the matched triples.
11. The video question answering method according to claim 1, wherein, the method further includes: processing the video feature vector and / or the text feature vector so that when the video feature vector and the text feature vector are concatenated, the dimensions of the video feature vector and the text feature vector are the same.
12. A video question answering device, wherein, it includes a memory; and a processor coupled to the memory, and the processor is configured to execute the steps of the video question answering method according to any one of claims 1 to 11 based on instructions stored in the memory.
13. A storage medium, wherein, a computer program is stored thereon, and when the program is executed by a processor, it implements the video question answering method according to any one of claims 1 to 11.
14. A video question answering system, wherein, it includes a video question answering device, a monitoring system, a speech recognition device, a voice input device, and a knowledge base, wherein: the monitoring system is configured to obtain one or more monitoring videos, process the monitoring videos according to instruction texts, and output the monitoring videos to the video question answering device; the voice input device is configured to receive a voice input and output it to the speech recognition device; The voice recognition device is configured to convert a voice input into an instruction text or a question text through voice recognition, input the instruction text into the monitoring system, and input the question text into the video question answering device; The knowledge base is configured to store a common sense knowledge graph; The video question answering device is configured to receive a question text and a monitoring video, generate candidate answer texts according to the question text, where the question text is used to describe a question, and the candidate answer texts are used to provide multiple candidate answers; it is also configured to extract a video feature vector from the received monitoring video, extract text feature vectors for the question text and the candidate answer texts, splice the video feature vector and the text feature vectors to obtain a spliced feature vector, input the spliced feature vector into a first pre-trained model, and the first pre-trained model learns cross-modal information between the video feature vector and the text feature vectors through a self-attention mechanism to obtain an encoded second spliced feature vector; divide the second spliced feature vector into a second video feature vector and a second text feature vector; input the second video feature vector and the second text feature vector into a modality fusion model, and the modality fusion model uses a mutual attention mechanism to process the second video feature vector and the second text feature vector to obtain a video expression and a text expression, respectively pool and fuse the video expression and the text expression to obtain a fused feature vector; input the fused feature vector into a decoding layer to predict the correct candidate answer.
Citation Information
Patent Citations
Medical image question-answering method and system based on deep learning
CN111984772A
Video question answering method and system for end-to-end training based on sparse sampling
CN113807222A
Cited By
Video generation based on 3D window attention
US12726685B2