Video question and answer positioning method and system for Gaussian-guided cross-modal learning

By introducing dynamic Gaussian distributed time positioning, time backtracking mechanism and cross-time causal comparison loss in the video Q&A model, the existing video Q&A model has been solved, and more efficient video Q&A positioning and semantic correlation improvement are achieved.

CN120144693AActive Publication Date: 2025-06-13SUN YAT SEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510008837.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-06-13
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The existing video question-and-answer model has insufficient positioning performance in real scenes, which is difficult to provide a reliable visual basis, and the model operation is high and its efficiency is limited.

Method used

A Gaussian-guided cross-modal learning video Q&A positioning method is proposed. Through dynamic Gaussian distribution time positioning, time backtracking mechanism and cross-time causal comparison loss, the semantic correlation of time segment positioning is improved and the complexity of model operation is reduced.

Benefits of technology

It improves the semantic correlation and practical application efficiency of video Q&A positioning, reduces the complexity of model operations, and is suitable for real-time application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144693A_ABST
    Figure CN120144693A_ABST
Patent Text Reader

Abstract

The invention discloses a video question and answer positioning method and system for Gaussian guided cross-modal learning. Comprising the steps of performing feature extraction on a video question and answer public data set to obtain video frame features and question features; inputting the video frame features into time Transform to obtain global time sequence features and attention weights; constructing time weight distribution by using the global time sequence features and the problem features, and weighting the video frame features to obtain weighted visual features; constructing a time backtracking mechanism by using the attention weight and the time weight distribution to obtain a time backtracking feature; performing feature fusion on the time backtracking feature and the question feature to obtain a multi-modal fusion feature and a prediction answer; cross-time causal comparison loss is introduced for learning, and a trained model is obtained; and the user inputs a to-be-processed video and a question into the trained model, and outputs a predicted answer to the question and a video positioning result. According to the method, the semantic correlation of positioning can be improved, the model operation complexity is reduced, and the positioning performance of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision, natural language processing, and video question answering, and particularly relates to a video question answering localization method and system for Gaussian-guided cross-modal learning. Background Art

[0002] In recent years, video question answering, as one of the core tasks in the field of multi-modal research, has received extensive attention. The goal of a video question answering model is to generate the correct answer by analyzing video content and question text. This task combines the technical advantages of computer vision and natural language processing and has broad application prospects in fields such as human-computer interaction, intelligent transportation, and educational assistance. With the rapid development of deep learning, video question answering models have shown significant improvement in answering question accuracy. However, the performance of these models in real scenarios still has critical deficiencies.

[0003] Most video question answering models adopt an end-to-end training framework, integrating video features, text features, and the answer prediction process into an overall model. Although this method improves the efficiency of the question answering task, since the model directly outputs the answer, its reasoning process is often a "black box", making it difficult to interpret the results, resulting in the model being unable to provide reliable visual evidence and limiting its usability in real scenarios. Thus, the video question answering localization task emerged. This task not only requires the model to generate the correct answer but also must locate the time segment or visual content related to the question in the video.

[0004] Currently, there are still relatively few models that can complete video question answering localization. Existing methods include post hoc attention analysis, multi-level similarity modeling, etc. However, time modeling and causal logic in videos cannot be solved solely by simple text similarity or frame-level visual information, and it is also difficult for a redesigned complex model architecture to efficiently complete the localization of time segments while maintaining the original question answering ability.

[0005] One of the current prior arts is a method for weakly supervised video localization proposed in the paper "Can I Trust Your Answer?Visually Grounded Video Question Answering". This method uses a temporal Transformer to extract the temporal features of the video, models a Gaussian distribution based on its output and generates temporal weights, thereby weighting the video frames to highlight the time segments related to the question, and combines the weighted visual features and question features for multimodal fusion to predict the answer. However, the disadvantage of this method is that the localization effect of this method depends on the output features of the temporal Transformer, can only generate the weight distribution of a single time period, cannot handle complex causal problems that rely on the inference results of multiple time segments, and only uses visual features for localization without combining question features for guidance, resulting in relatively weak semantic relevance.

[0006] Another of the current prior arts is a method based on a gated state-space multimodal Transformer proposed in the paper "Encoding and Controlling Global Semantics for Long-form Video Question Answering". This method introduces a state-space layer in the multimodal Transformer, performs convolutional operations on the video frame features using a series of state matrices to generate the global semantic representation of the time series, and dynamically adjusts the influence of the global semantics on the visual representation through a gating mechanism, alleviating the problem of information loss caused by frame and local feature selection. In addition, this method proposes a cross-modal combination consistency C3 objective to optimize the alignment of visual and language features. However, the disadvantage of this method is that the addition of the state-space layer and the gating mechanism increases the complexity of the model and the demand for computing resources, resulting in limited efficiency in practical applications, and the state-space layer enhances the global semantics, which is not equivalent to the improvement of causal reasoning ability, but only enhances the feature expression of each frame and is not specifically optimized for time-dependent relationships. Summary of the Invention

[0007] The object of the present invention is to overcome the deficiencies of the existing methods and propose a video question answering localization method and system based on Gaussian-guided cross-modal learning. The main problems solved by the present invention are how to improve the semantic relevance of time segment localization, reduce the computational complexity of the model, and improve the localization performance and practical application efficiency of the model.

[0008] To solve the above problems, the present invention proposes a video question answering localization method based on Gaussian-guided cross-modal learning, and the method includes:

[0009] Extract features from the video data, question text data, and answer text data in the video question answering public dataset to obtain video frame features f(v) and question features f(q);

[0010] Input the video frame features f(v) into a temporal Transformer to obtain global temporal features and attention weights A ij ;

[0011] Utilize the global temporal features and the question features f(q) to obtain the center point μ and range σ of the dynamic Gaussian distribution, and construct a normalized temporal weight distribution P(t) to weight the video frame features f(v) to obtain weighted visual features

[0012] Utilize the attention weights A ij and the temporal weight distribution P(t) to construct a temporal backtracking mechanism to obtain temporal backtracking features

[0013] Input the temporal backtracking features and the question features f(q) for feature fusion to obtain multimodal fusion features f cls and predicted answers

[0014] Input the video frame features f(v), the temporal weight distribution P(t), and the temporal backtracking features Introduce cross-temporal causal contrast loss for learning to obtain a trained model;

[0015] The user inputs the video to be processed and the question into the trained model, and the predicted answer to the question and the video localization result corresponding to the answer are output.

[0016] Preferably, the feature extraction is specifically as follows:

[0017] First, sample the video data at a fixed frame rate, record the total number of final sampled frames as a fixed number T, then sample uniformly in fixed quantities to obtain T video frames, adjust the size of the video frames to a unified size, and normalize the pixel values of the video frames to obtain preprocessed sampled video data;

[0018] Input the preprocessed sampled video data into the visual encoder of the pre-trained CLIP model for feature extraction to obtain video frame features f(v), and input the question text data into the text encoder of the pre-trained CLIP model for encoding to obtain question features f(q).

[0019] Preferably, the video frame feature f(v) is input into the temporal Transformer to obtain the global temporal feature and the attention matrix A ij , specifically as follows:

[0020] After inputting the video frame feature f(v) and embedding the position information, an improved self-attention mechanism is used to model the temporal dependence relationship between frames, followed by residual connection and normalization, and then passing through a feed-forward neural network to obtain the global temporal feature and the attention weight A ij ;

[0021] The video frame feature f(v) consists of several temporal frame features , and the temporal frame feature includes a query vector Q t , a key vector K t , and a value vector V t ;

[0022] The calculation formulas for the query vector Q t and the key vector K t are respectively:

[0023]

[0024] where Cov1D represents a one-dimensional convolution operation, and the convolution kernel size k represents covering the information of the adjacent k - 1 temporal frames centered on frame t;

[0025] The video frame feature f(v) is mapped to the vector spaces of the query vector Q t and the key vector K t using convolution to obtain the value vector V t ;

[0026] The attention weight A ij represents the local dependence degree of frame i on frame j, and its calculation formula is:

[0027]

[0028] where d is a scaling factor for the feature dimension to stabilize the attention value, Q i is the query vector of frame i, is the transpose of the key vector of frame j.

[0029] Preferably, using the global temporal feature and the problem feature f(q), the center point μ and the range σ of the dynamic Gaussian distribution are obtained, and a normalized temporal weight distribution P(t) is constructed to weight the video frame feature f(v) to obtain the weighted visual feature Specifically:

[0030] The calculation formulas for the center point μ and the range σ of the dynamic Gaussian distribution are respectively:

[0031]

[0032] Wherein, and are respectively the weight matrices of the global temporal feature , and are respectively the weight matrices of the weighted problem feature f qa , and the weighted problem feature f qa is the result after f(q) is weighted, and the weight is and will be randomly initialized in the model initialization stage and updated during the gradient descent process. The problem feature f(q) is a correction term, and w μ and w σ control the correction intensity;

[0033] Construct the time weight distribution p(t) using the center point μ and the range σ of the dynamic Gaussian distribution. The specific calculation formula is:

[0034]

[0035] Where T is the total number of sampled frames;

[0036] Normalize the time weight distribution p(t) to obtain the normalized time weight distribution P(t). The specific calculation formula is:

[0037]

[0038] Weight the video frame feature f(v) with the normalized time weight distribution P(t) to obtain the weighted visual feature

[0039] Preferably, construct a time backtracking mechanism using the attention weight A ij and the time weight distribution P(t) to obtain the time backtracking feature Specifically:

[0040] Perform a dot product on the attention weight A ij and the value vector V t to generate the causal feature

[0041] Weight the causal feature using the time weight distribution P(t) to obtain the time backtracking feature

[0042] Preferably, the time-backtracking feature and the problem feature f(q) are fused to obtain a multimodal fusion feature f cls and the predicted answer Specifically:

[0043] The time-backtracking feature and the problem feature f(q) are concatenated, and the fusion feature f is obtained through a fully connected layer cls , and the fusion feature f cls is input into the classification layer, and the probability distribution of the answer Softmax(W cls f cls +b cls ) is generated through an activation function, where W cls and b cls are the parameters of the classification layer, and the candidate answer with the highest probability is selected as the predicted answer for the video question answering task

[0044] Preferably, the video frame feature f(v), the time weight distribution P(t) and the time-backtracking feature are input Cross-time causal contrast loss is introduced for learning to obtain a trained model. Specifically:

[0045] Positive and negative samples for contrastive learning are constructed, that is, the time frame features in the video frame feature f(v) are dot-product weighted using the time weight distribution P(t) to obtain positive samples and negative samples are randomly selected from the frames with a lower time weight distribution P(t)

[0046] Using the positive samples and the negative samples A cross-time causal contrast loss function is constructed The specific formula is as follows:

[0047]

[0048] where sim represents the cosine similarity between feature vectors, t is the positive sample index, and t ′ is the negative sample index.

[0049] Correspondingly, the present invention also provides a video question answering localization system for Gaussian-guided cross-modal learning, including:

[0050] An initialization unit for extracting features from video data, question text data, and answer text data in a video question answering public dataset to obtain a video frame feature f(v) and a question feature f(q);

[0051] The model construction and training unit is used to input the video frame feature f(v) into the temporal Transformer to obtain the global temporal feature and the attention weight A ij ; Using the global temporal feature and the question feature f(q), obtain the center point μ and range σ of the dynamic Gaussian distribution, and construct the normalized temporal weight distribution P(t) to weight the video frame feature f(v) to obtain the weighted visual feature Using the attention weight A ij and the temporal weight distribution P(t) to construct a temporal backtracking mechanism to obtain the temporal backtracking feature The temporal backtracking feature and the question feature f(q) are fused to obtain the multi-modal fusion feature f cls and the predicted answer Input the video frame feature f(v), the temporal weight distribution P(t) and the temporal backtracking feature Introduce cross-temporal causal contrast loss for learning to obtain the trained model;

[0052] The model application unit is used for the user to input the video to be processed and the question into the trained model, and output the predicted answer to the question and the video localization result corresponding to the answer.

[0053] Implementing the present invention has the following beneficial effects:

[0054] The dynamic Gaussian distribution time localization method, the temporal backtracking mechanism and the cross-temporal causal contrast loss proposed by the present invention have high modularity and plug-and-play characteristics, and can be easily integrated into existing video question answering models or multi-modal tasks to provide time localization enhancement and causal reasoning ability improvement for these models. In addition, the dynamic Gaussian distribution, the temporal backtracking mechanism and the causal contrast loss in the present invention only depend on simple linear mapping and weighting operations of temporal features during calculation, and the generation and normalization processes of the Gaussian distribution have low overhead, and can adapt to the computing power of most devices. Compared with the end-to-end model that requires complex parameter learning, the module calculation complexity of this method is significantly reduced, and it is suitable for real-time application scenarios. Description of the Drawings

[0055] Figure 1 is a flowchart of a video question answering localization method for Gaussian-guided cross-modal learning according to an embodiment of the present invention;

[0056] Figure 2 is a schematic diagram of the temporal Transformer architecture in an embodiment of the present invention;

[0057] Figure 3 It is a structural diagram of a video question - answering localization system for Gaussian - guided cross - modal learning according to an embodiment of the present invention. Specific implementation manners

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0059] Figure 1 It is a flowchart of a video question - answering localization method for Gaussian - guided cross - modal learning according to an embodiment of the present invention. As Figure 1 shown, the method includes:

[0060] S1. Extract features from the video data, question text data, and answer text data in the video question - answering public dataset to obtain video frame features f(v) and question features f(q);

[0061] S2. Input the video frame features f(v) into a temporal Transformer to obtain global temporal features and attention weights A ij ;

[0062] S3. Use the global temporal features and the question features f(q) to obtain the center point μ and range σ of the dynamic Gaussian distribution, and construct a normalized temporal weight distribution P(t) to weight the video frame features f(v) to obtain weighted visual features

[0063] S4. Use the attention weights A ij and the temporal weight distribution P(t) to construct a temporal backtracking mechanism to obtain temporal backtracking features

[0064] S5. Fuse the temporal backtracking features and the question features f(q) to obtain multi - modal fusion features f cls and predicted answers

[0065] S6. Input the video frame features f(v), the temporal weight distribution P(t), and the temporal backtracking features introduce cross - time causal contrast loss for learning to obtain a trained model;

[0066] In S7, the user inputs the video to be processed and the question into the trained model, and the predicted answer to the question and the video location result corresponding to the answer are output.

[0067] Step S1 is as follows:

[0068] The video question-answering public dataset is general, such as NExT-GQA, MultiHop-EgoQA, AGQA-Decomp, etc.;

[0069] In S1-1, the video data is first sampled at a fixed frame rate. In this embodiment, it is sampled at a frame rate of 6fs, so that the frame rates of different input videos are the same. The total number of finally sampled frames is recorded as a fixed number T, and then uniformly sampled at a fixed number to obtain T video frames. The size of the video frames is adjusted to a unified size. In this embodiment, it is 224×224, and the pixel values of the video frames are normalized to obtain the preprocessed sampled video data;

[0070] In S1-2, the preprocessed sampled video data is input into the visual encoder of the pre-trained CLIP model for feature extraction. The visual encoder of the CLIP model adopts the Vision Transformer (ViT) architecture. Each frame of the image is divided into image patches of a fixed size, and an embedding representation is generated through linear mapping. Feature modeling is performed through multiple layers of Transformer in combination with positional embeddings to obtain the video frame feature f(v). The question text data is input into the text encoder of the pre-trained CLIP model for encoding. The text encoder of the CLIP model adopts the Transformer architecture. The question q is segmented into sub-words through a tokenizer to generate an embedding representation, and the relationship between words is modeled through multiple layers of Transformer in combination with positional encoding to obtain the question feature f(q). This feature represents the semantic information of the question and can be aligned with the visual feature in the shared embedding space.

[0071] In step S2, the architecture of the temporal Transformer is as Figure 2 shown, and the steps are as follows:

[0072] After inputting the video frame feature f(v) and embedding the position information, an improved self-attention mechanism is used to model the temporal dependence relationship between frames, and then residual connection and normalization are performed, and through a feed-forward neural network, the global temporal feature and the attention weight A ij ;

[0073] Among them, the video frame feature f(v) is composed of several temporal frame features which include query vectors Q in the temporal frame features t, key vector K t , value vector V t , query vector Q t , key vector K t , value vector V t not only comes from itself, but also integrates its local context features through convolution operations;

[0074] The query vector Q t and the key vector K t are calculated as follows:

[0075]

[0076] where Cov1D represents a one-dimensional convolution operation, and the convolution kernel size k represents covering the information of the adjacent k - 1 time frames centered on the frame t;

[0077] Map the video frame feature f(v) to the vector spaces of the query vector Q t and the key vector K t to obtain the value vector V t ;

[0078] The attention weight A ij represents the local dependence degree of frame i on frame j, and its calculation formula is:

[0079]

[0080] where d is a scaling factor of the feature dimension, used to stabilize the attention value, Q i is the query vector of frame i, is the transpose of the key vector of frame j.

[0081] Step S3 is as follows:

[0082] S3 - 1, the calculation formulas for the center point μ and the range σ of the dynamic Gaussian distribution are respectively:

[0083]

[0084] where, according to the definition of the Gaussian distribution, μ represents the video moment frame most relevant to the problem, and σ determines the expansion width of the distribution, representing the coverage range of relevant frames; and are respectively the weight matrices of the global temporal feature , and are respectively the weight matrices of the weighted problem feature f qa , the weighted problem feature f qa is the result after f(q) is weighted, and the weight is and will be randomly initialized during the model initialization phase and updated during the gradient descent process. The problem feature f(q) is a correction term used to improve semantic relevance while ensuring the dominant role of video features, w μ and w σ control the correction intensity;

[0085] S3-2. Construct the temporal weight distribution p(t) using the center point μ and range σ of the dynamic Gaussian distribution. The specific calculation formula is:

[0086]

[0087] where T is the total number of sampled frames;

[0088] The closer the temporal frame is to the center point μ, the greater the weight and the higher the correlation;

[0089] S3-3. Normalize the temporal weight distribution p(t) to facilitate subsequent weighted operations and obtain the normalized temporal weight distribution P(t). The specific calculation formula is:

[0090]

[0091] S3-4. Weight the video frame feature f(v) with the normalized temporal weight distribution P(t) to obtain the weighted visual feature

[0092] Step S4 is as follows:

[0093] Model the explicit dependency relationship between temporal frames through the attention weight A ij and combine it with the temporal weight distribution P(t) located by the Gaussian distribution, enabling the model to capture the causal chain before and after events in the video and solve the complex problem of video question answering localization that requires causal reasoning;

[0094] S4-1. Take the dot product of the attention weight A ij and the value vector V t to generate the causal feature

[0095] S4-2. Weight the causal feature with the temporal weight distribution P(t) to integrate the semantics of key temporal frames with their causal dependencies and obtain the temporal backtracking feature

[0096] Step S5 is as follows:

[0097] The temporal backtracking feature Concatenate with the problem feature f(q), and obtain the fused feature f through a fully connected layer cls , and input the fused feature f cls into the classification layer, and generate the probability distribution of the answer through the activation function Softmax(W cls f cls +b cls ), where W cls and b cls are the parameters of the classification layer, and select the candidate answer with the highest probability as the predicted answer for the video question answering task

[0098] Step S6 is as follows:

[0099] S6-1, Construct positive and negative samples for contrast learning, that is, use the time weight distribution P(t) to perform dot product weighting on the time frame features in the video frame feature f(v) to obtain positive samples indicating time segments highly relevant to the question semantics, and randomly select negative samples from the frames with a lower time weight distribution P(t) indicating time segments irrelevant to the question. The random sampling strategy of negative samples ensures the diversity of contrast learning while strengthening the model's ability to distinguish time segments;

[0100] S6-2, Use the positive sample and the negative sample to construct a cross-time causal contrast loss function The specific formula is as follows:

[0101]

[0102] where sim represents the cosine similarity between feature vectors, t is the positive sample index, and t′ is the negative sample index.

[0103] Step S7 is as follows:

[0104] The predicted answer uses the accuracy rate index, and the frame index of the video localization result is the range [μ - λσ, μ + λσ] calculated using the center point μ and range σ of the dynamic Gaussian distribution. The video localization result uses the IoU index.

[0105] Correspondingly, the present invention also provides a video question answering and localization system for Gaussian-guided cross-modal learning, as shown in Figure 3 and includes:

[0106] Initialization unit 1, used to extract features from the video data, question text data, and answer text data in the video question answering public dataset to obtain the video frame feature f(v) and the question feature f(q);

[0107] Specifically, the video data is first sampled at a fixed frame rate, and the total number of finally sampled frames is recorded as a fixed number T. Then, uniform sampling is performed in fixed quantities to obtain T video frames. The size of the video frames is adjusted to a unified size, and the pixel values of the video frames are normalized to obtain preprocessed sampled video data;

[0108] The preprocessed sampled video data is input into the visual encoder of the pre-trained CLIP model for feature extraction to obtain video frame features f(v). The question text data is input into the text encoder of the pre-trained CLIP model for encoding to obtain question features f(q).

[0109] The model construction and training unit 2 is used to input the video frame features f(v) into the temporal Transformer to obtain global temporal features and attention weights A ij ; Using the global temporal features and the question features f(q), the center point μ and range σ of the dynamic Gaussian distribution are obtained, and a normalized temporal weight distribution P(t) is constructed to weight the video frame features f(v) to obtain weighted visual features Using the attention weights A ij and the temporal weight distribution P(t) to construct a temporal backtracking mechanism to obtain temporal backtracking features The temporal backtracking features and the question features f(q) are feature fused to obtain multi-modal fusion features f cls and predicted answers Input the video frame features f(v), the temporal weight distribution P(t) and the temporal backtracking features Introduce cross-temporal causal contrast loss for learning to obtain a trained model;

[0110] Specifically, when inputting the video frame features f(v), after embedding the position information, an improved self-attention mechanism is used to model the temporal dependence relationship between frames, and then residual connection and normalization are performed, and passed through a feed-forward neural network to obtain global temporal features and attention weights A ij ;

[0111] The video frame features f(v) are composed of several temporal frame features The temporal frame features include query vectors Q t 、key vectors K t 、value vectors V t ;

[0112] The query vectors Q t and key vectors Kt The calculation formulas are as follows:

[0113]

[0114] Among them, Cov1D represents a one-dimensional convolution operation, and the convolution kernel size k represents covering the information of the adjacent k - 1 time frames centered on frame t;

[0115] The video frame feature f(v) is mapped to the query vector Q using convolution t and the key vector K t in the vector space to obtain the value vector V t ;

[0116] The attention weight A ij represents the local dependence degree of frame i on frame j, and its calculation formula is:

[0117]

[0118] where d is the scaling factor of the feature dimension, used to stabilize the attention value, Q i is the query vector of frame i, is the transpose of the key vector of frame j.

[0119] The calculation formulas for the center point μ and the range σ of the dynamic Gaussian distribution are respectively:

[0120]

[0121] Among them, and are respectively the weight matrices of the global temporal feature , and are respectively the weight matrices of the weighted problem feature f qa , and the weighted problem feature f qa is the result after f(q) is weighted, and the weight is and will be randomly initialized in the model initialization stage and updated during the gradient descent process. The problem feature f(q) is a correction term, and w μ and w σ control the correction intensity;

[0122] Construct the time weight distribution p(t) using the center point μ and the range σ of the dynamic Gaussian distribution. The specific calculation formula is:

[0123]

[0124] where T is the total number of sampled frames;

[0125] Normalize the time weight distribution p(t) to obtain the normalized time weight distribution P(t). The specific calculation formula is as follows:

[0126]

[0127] Weight the video frame feature f(v) with the normalized time weight distribution P(t) to obtain the weighted visual feature

[0128] Take the attention weight A ij and the value vector V t Perform a dot product to generate the causal feature

[0129] Use the time weight distribution P(t) to weight the causal feature to obtain the time backtracking feature

[0130] Take the time backtracking feature and the question feature f(q) and splice them. Obtain the fused feature f through the fully connected layer cls , and input the fused feature f cls into the classification layer. Generate the probability distribution of the answer through the activation function Softmax(W cls f cls +b cls ), where W cls and b cls are the parameters of the classification layer. Select the candidate answer with the highest probability as the predicted answer for the video question answering task

[0131] Construct positive and negative samples for contrastive learning, that is, use the time weight distribution P(t) to perform dot product weighting on the time frame features in the video frame feature f(v) to obtain the positive sample and randomly select negative samples from the frames with a lower time weight distribution P(t)

[0132] Use the positive sample and the negative sample to construct a cross-time causal contrast loss function The specific formula is as follows:

[0133]

[0134] Among them, sim represents the cosine similarity between feature vectors, t is the positive sample index, and t′ is the negative sample index.

[0135] A model application unit 3, which is used for a user to input a video to be processed and a question into the trained model, and output a predicted answer to the question and a video localization result corresponding to the answer.

[0136] Therefore, the present invention proposes a video question answering and localization method for Gaussian-guided cross-modal learning. The dynamic Gaussian distribution time localization method, the time backtracking mechanism, and the cross-time causal contrast loss proposed by the present invention have high modularity and plug-and-play characteristics, and can be easily integrated into existing video question answering models or multi-modal tasks to provide time localization enhancement and causal reasoning ability improvement for these models. In addition, the dynamic Gaussian distribution, the time backtracking mechanism, and the causal contrast loss in the present invention only rely on simple linear mapping and weighting operations of time features during calculation, and the generation and normalization processes of the Gaussian distribution have low overhead, and can adapt to the computing power of most devices. Compared with the end-to-end model that requires complex parameter learning, the module calculation complexity of this method is significantly reduced, which is suitable for real-time application scenarios.

[0137] The above has introduced in detail a video question answering and localization method and system for Gaussian-guided cross-modal learning provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A video question answering localization method based on Gaussian guided cross-modal learning, characterized in that: The method comprises: Perform feature extraction on the video data, question text data, and answer text data in the video question answering public dataset to obtain video frame features f(v) and question features f(q); Input the video frame feature f(v) into the temporal Transformer to obtain the global temporal feature and attention weight A ij ; Using the global timing characteristics And the problem feature f(q), get the center point μ and range σ of the dynamic Gaussian distribution, and construct a normalized time weight distribution P(t), weight the video frame feature f(v) to get the weighted visual feature Using the attention weight A ij The time weight distribution P(t) is used to construct a time backtracking mechanism to obtain the time backtracking feature The time-tracing feature The feature fusion is performed with the problem feature f(q) to obtain the multimodal fusion feature f cls and predict the answer Input the video frame feature f(v), the time weight distribution P(t) and the time traceback feature Introduce cross-temporal causal contrast loss for learning to obtain a trained model; The user inputs the video to be processed and the question into the trained model, which outputs the predicted answer to the question and the video positioning result corresponding to the answer.

2. The video question answering positioning method of Gaussian guided cross-modal learning as claimed in claim 1, characterized in that: The feature extraction is specifically as follows: The video data is first sampled at a fixed frame rate, and the final total number of sampled frames is recorded as a fixed number T, and then the fixed number is uniformly sampled to obtain T frames of video frames, the size of the video frames is adjusted to a uniform size, and the pixel values ​​of the video frames are normalized to obtain preprocessed sampled video data; The preprocessed sampled video data is input into the pre-trained CLIP model visual encoder for feature extraction to obtain video frame features f(v), and the question text data is input into the pre-trained CLIP model text encoder for encoding to obtain question features f(q).

3. The video question answering positioning method of Gaussian guided cross-modal learning as claimed in claim 1, characterized in that: The video frame feature f(v) is input into the temporal Transformer to obtain the global temporal feature and the attention matrix A ij , specifically: Input the video frame feature f(v), embed the position information, use the improved self-attention mechanism to model the temporal dependency between frames, perform residual connection and normalization, and pass through the feedforward neural network to obtain the global temporal feature and attention weight A ij ; The video frame feature f(v) is composed of several time frame features Composition, time frame characteristics The query vector Q t , key vector K t , value vector V t ; The query vector Q t and the key vector K t The calculation formulas are: Among them, Conv1D represents a one-dimensional convolution operation, and the convolution kernel size k represents the coverage of information centered on frame t and containing adjacent k-1 time frames; Map the video frame features f(v) to the query vector Q using convolution t and the key vector K t The vector space of the value vector V t ; The attention weight A ij It represents the local dependence of frame i on frame j, and its calculation formula is: Where d is the scaling factor of the feature dimension, used to stabilize the attention value, Q i is the query vector for frame i, is the transpose of the key vector for frame j.

4. The video question answering positioning method of Gaussian guided cross-modal learning as claimed in claim 1, characterized in that: The use of the global timing characteristics And the problem feature f(q), get the center point μ and range σ of the dynamic Gaussian distribution, and construct a normalized time weight distribution P(t), weight the video frame feature f(v) to get the weighted visual feature Specifically: The calculation formulas for the center point μ and range σ of the dynamic Gaussian distribution are: in, and The global timing characteristics are The weight matrix of and They are weighted problem features f qa The weight matrix of the weighted problem feature f qa is the weighted result of f(q), and the weight is and It will be randomly initialized in the model initialization stage and updated during the gradient descent process. The problem feature f(q) is the correction term, w μ and w σ Control the intensity of the correction; The time weight distribution p(t) is constructed using the center point μ and range σ of the dynamic Gaussian distribution. The specific calculation formula is: Where T is the total number of sampling frames; The time weight distribution p(t) is normalized to obtain the normalized time weight distribution P(t). The specific calculation formula is: The video frame feature f(v) is weighted by the normalized time weight distribution P(t) to obtain a weighted visual feature 5. The video question answering positioning method of Gaussian guided cross-modal learning as claimed in claim 3, characterized in that: The use of the attention weight A ij The time weight distribution P(t) is used to construct a time backtracking mechanism to obtain the time backtracking feature Specifically: The attention weight A ij and the value vector V t Perform dot product to generate causal features The causal feature is analyzed using the time weight distribution P(t). Weighted to get the time traceback feature 6. The video question answering positioning method of Gaussian guided cross-modal learning as claimed in claim 1, characterized in that: The time back feature The feature fusion is performed with the problem feature f(q) to obtain the multimodal fusion feature f cls and predict the answer Specifically: The time-tracing feature It is concatenated with the problem feature f(q) and the fusion feature f is obtained through the fully connected layer. cls , the fusion feature f cls Input classification layer, generate the probability distribution of answer through activation function Softmax(W cls f cls +b cls ), where W cls and b cls is the parameter of the classification layer, and the candidate answer with the highest probability is selected as the predicted answer for the video question answering task 7. The video question answering positioning method of Gaussian guided cross-modal learning as claimed in claim 1, characterized in that: The input video frame feature f(v), the time weight distribution P(t) and the time traceback feature Introduce cross-temporal causal contrast loss for learning and obtain the trained model, specifically: Construct positive and negative samples for contrastive learning, that is, use the time weight distribution P(t) to compare the time frame features in the video frame features f(v) Perform dot product weighting to obtain positive samples And randomly select negative samples from the frames with lower temporal weight distribution P(t) Using the positive sample And the negative samples Constructing a cross-temporal causal contrast loss function The specific formula is as follows: Among them, sim represents the cosine similarity between feature vectors, t is the positive sample index, and t′ is the negative sample index.

8. A Gaussian guided cross-modal learning video question answering positioning system, characterized in that: The system comprises: An initialization unit, used to extract features from video data, question text data, and answer text data in a public video question-answering dataset to obtain video frame features f(v) and question features f(q); Model building and training unit, used to input the video frame feature f(v) into the time Transformer to obtain the global temporal feature and attention weight A ij ; Using the global timing characteristics And the problem feature f(q), get the center point μ and range σ of the dynamic Gaussian distribution, and construct a normalized time weight distribution P(t), weight the video frame feature f(v) to get the weighted visual feature Using the attention weight A ij The time weight distribution P(t) is used to construct a time backtracking mechanism to obtain the time backtracking feature The time-tracing feature The feature fusion is performed with the problem feature f(q) to obtain the multimodal fusion feature f cls and predict the answer Input the video frame feature f(v), the time weight distribution P(t) and the time traceback feature Introduce cross-temporal causal contrast loss for learning to obtain a trained model; The model application unit is used for the user to input the video to be processed and the question into the trained model, and output the predicted answer to the question and the video positioning result corresponding to the answer.

9. A video question answering positioning system for Gaussian guided cross-modal learning as claimed in claim 8, characterized in that: The initialization unit is used to extract features from the video data, question text data, and answer text data in the video question and answer public data set to obtain video frame features f(v) and question features f(q), specifically: The video data is first sampled at a fixed frame rate, and the final total number of sampled frames is recorded as a fixed number T, and then the fixed number is uniformly sampled to obtain T frames of video frames, the size of the video frames is adjusted to a uniform size, and the pixel values ​​of the video frames are normalized to obtain preprocessed sampled video data; The preprocessed sampled video data is input into the pre-trained CLIP model visual encoder for feature extraction to obtain video frame features f(v), and the question text data is input into the pre-trained CLIP model text encoder for encoding to obtain question features f(q).

10. A Gaussian guided cross-modal learning video question answering positioning system as claimed in claim 8, characterized in that: The model building and training unit is used to input the video frame feature f(v) into the time Transformer to obtain the global temporal feature and attention weight A ij ; Using the global timing characteristics And the problem feature f(q), get the center point μ and range σ of the dynamic Gaussian distribution, and construct a normalized time weight distribution P(t), weight the video frame feature f(v) to get the weighted visual feature Using the attention weight A ij The time weight distribution P(t) is used to construct a time backtracking mechanism to obtain the time backtracking feature The time-tracing feature The feature fusion is performed with the problem feature f(q) to obtain the multimodal fusion feature f cls and predict the answer Input the video frame feature f(v), the time weight distribution P(t) and the time traceback feature Introduce cross-temporal causal contrast loss for learning and obtain the trained model, specifically: Input the video frame feature f(v), embed the position information, use the improved self-attention mechanism to model the temporal dependency between frames, perform residual connection and normalization, and pass through the feedforward neural network to obtain the global temporal feature and attention weight A ij ; The video frame feature f(v) is composed of several time frame features Composition, time frame characteristics The query vector Q t , key vector K t , value vector V t ; The query vector Q t and the key vector K t The calculation formulas are: Among them, Cov1D represents a one-dimensional convolution operation, and the convolution kernel size k represents the coverage of information centered on frame t and containing adjacent k-1 time frames; Map the video frame features f(v) to the query vector Q using convolution t and the key vector K t The vector space of the value vector V t ; The attention weight A ij It represents the local dependence of frame i on frame j, and its calculation formula is: Where d is the scaling factor of the feature dimension, used to stabilize the attention value, Q i is the query vector for frame i, is the transpose of the key vector for frame j; The calculation formulas for the center point μ and range σ of the dynamic Gaussian distribution are: in, and The global timing characteristics are The weight matrix of and They are weighted problem features f qa The weight matrix of the weighted problem feature f qa is the weighted result of f(q), and the weight is and It will be randomly initialized in the model initialization stage and updated during the gradient descent process. The problem feature f(q) is the correction term, w μ and w σ Control the intensity of the correction; The time weight distribution p(t) is constructed using the center point μ and range σ of the dynamic Gaussian distribution. The specific calculation formula is: Where T is the total number of sampling frames; The time weight distribution p(t) is normalized to obtain the normalized time weight distribution P(t). The specific calculation formula is: The normalized time weight distribution P(t) is used to weight the video frame feature f(v) to obtain a weighted visual feature The attention weight A ij and the value vector V t Perform dot product to generate causal features The causal feature is analyzed using the time weight distribution P(t). Weighted to get the time traceback feature The time-tracing feature It is concatenated with the problem feature f(q) and the fusion feature f is obtained through the fully connected layer. cls , the fusion feature f cls Input classification layer, generate the probability distribution of answer through activation function Softmax(W cls f cls +b cls ), where W cls and b cls is the parameter of the classification layer, and the candidate answer with the highest probability is selected as the predicted answer for the video question answering task Construct positive and negative samples for contrastive learning, that is, use the time weight distribution P(t) to compare the time frame features in the video frame features f(v) Perform dot product weighting to obtain positive samples And randomly select negative samples from the frames with lower temporal weight distribution P(t) Using the positive sample And the negative samples Constructing a cross-temporal causal contrast loss function The specific formula is as follows: Among them, sim represents the cosine similarity between feature vectors, t is the positive sample index, and t′ is the negative sample index.

Citation Information

Patent Citations

  • Semantic reconstruction video description method based on time sequence Gaussian mixture cavity convolution

    CN113420179A

  • Context-aware progressive attention video question-answering method and system

    CN114625849A

  • Semantic alignment video question and answer method

    CN115618061A

  • Visual question answering method and apparatus, electronic device and storage medium

    WO2024164616A1