Video question answering method based on aggregation-pruning sampling
By introducing an aggregation-pruning sampler into the video Q&A system, the noise visual marks in long videos are adaptively eliminated, and the problem of low answer accuracy in long video Q&A is solved, achieving higher accuracy and reasoning capabilities.
Patent Information
- Application Number
- CN202510306264.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-03
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
When performing long video Q&A, the answers of existing video Q&A systems are low in accuracy, mainly due to the redundancy of visual information caused by space-time noise.
A video question-and-answer sampling method is proposed, and the visual markers of time and space noise are adaptively eliminated by the aggregation-pruning sampler (APSam), thereby improving the semantic hierarchy and diversified feature granularity of visual information.
It effectively reduces unnecessary interference, improves the accuracy and reasoning ability of long video Q&A, focuses on aggregating similar markers related to the question, and enhances the answer accuracy of the Q&A system.
Smart Images

Figure CN120234446A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video question answering, and particularly relates to a video question answering method based on aggregation-pruning sampling. Background Art
[0002] Video question answering is a well-studied fundamental vision-language task. The basic elements of a video question answering system lie in being able to easily perceive, remember, and understand multi-modal information in daily life to assist humans in tasks such as finding specific objects, recalling past scenes, and analyzing the relationships between multiple events. Currently, methods for video question answering mainly rely on attention mechanisms, graph structures, and pre-trained models. Attention-based methods fuse question and video features into answer representations by calculating cross-modal attention. For example, the attention flow framework proposed by Gao et al. in 2019 captures high-level interactions between language and vision to improve question answering performance. Graph-based methods usually extract video scene graphs (subject-relation-object) and encode the graph structure to obtain information interactions in the video. Cherian et al. proposed a (2.5+1)D scene graph representation in 2021 to better capture spatio-temporal information flows in the video. Pre-trained model-based methods obtain cross-modal knowledge through pre-training on a large amount of video-text data. Lei et al. proposed a general framework in 2021 to achieve end-to-end video question answering learning by using short video clips. However, these methods mainly focus on short video question answering based on sparse frame sampling, lacking sufficient visual information and being difficult to apply to long video question answering.
[0003] Compared with short videos, long videos have a longer duration, which mainly brings two challenges: (1) Temporal noise: Long videos contain transitions of multiple events, so many video segments irrelevant to the problem introduce temporal noise; (2) Spatial noise: Long videos involve numerous visual concepts, including a large number of useless objects and backgrounds, which brings spatial noise. To adapt to more real-world scenarios, some research focuses on long video modeling. These methods mainly adopt dense frame sampling and establish long-term spatio-temporal dependencies. Although increasing the number of input frames can effectively increase useful information, it also introduces a large amount of additional noise, which may reduce the model performance. Therefore, recent research (Gao et al., 2023; Yu et al., 2023; Lee, Kang and Kim, 2023; Wang et al., 2024a; Li et al., 2023c) has considered reducing noise by filtering redundant visual information. For example, Gao et al. proposed a spatio-temporal Transformer in 2023 by iteratively selecting useful frames and image regions. Lee et al. proposed a deformable attention mechanism in 2023 for learning inherent temporal semantics to sample key frames and enhance compositional reasoning. In the same year, Yu et al. used a single image-language model to simultaneously handle temporal key frame localization and question answering in videos. However, these methods only focus on a limited token feature granularity and reduce noise by sampling a fixed number of visual tokens. In contrast, the present invention proposes an "Aggregation-Pruning Sampler" (APSam) to diversify the feature granularity and adaptively prune temporal and spatial noise visual tokens under the condition of each question, thereby reducing redundant interference while maintaining the integrity of visual cues at different scales.
[0004] The aggregation and adaptive pruning mechanism was initially used for model acceleration, where input tokens were adaptively pruned based on different samples to reduce the computational load. Ren et al. proposed aggregating similar frames and patches in each frame in 2023 to reduce visual tokens and thus accelerate video coding. However, they ignored the influence of text input during the aggregation process, which led to relevant visual tokens being ignored or wrongly merged. Wang et al. proposed integrating lightweight networks into the pre-trained backbone in 2024 to make pruning decisions on redundant image tokens within each layer.
[0005] Existing research has not considered how adaptive pruning works in the video field, especially in long video question answering. Different questions require attention to different levels of feature granularity and different numbers of visual cues. This results in low answer accuracy in existing video question answering systems when performing long video question answering. Summary of the Invention
[0006] The object of the present invention is to solve the problem that the accuracy of the answers in the existing video question - answering system is low when performing long - video question - answering. A video question - answering method based on aggregation - pruning sampling is provided, including:
[0007] Step 1: Construct a video question - answering model based on aggregation - pruning sampling; obtain a trained video question - answering model based on aggregation - pruning sampling;
[0008] Step 2: Input the video to be questioned and the question to be answered into the trained video question - answering model based on aggregation - pruning sampling to obtain an answer result;
[0009] In the above - mentioned Step 1, the process of constructing a video question - answering model based on aggregation - pruning sampling and obtaining a trained video question - answering model based on aggregation - pruning sampling is as follows:
[0010] 1. Obtain a video question - answering training set;
[0011] The video question - answering training set includes: the original video (the first input of the video question - answering system), the text question (the second input of the video question - answering system), and the answer label ((the training label of the video question - answering system));
[0012] 2. Construct a video question - answering model based on aggregation - pruning sampling; train the video question - answering model based on aggregation - pruning sampling according to the video question - answering training set to obtain a trained video question - answering model based on aggregation - pruning sampling;
[0013] The aggregation - pruning sampler is used to obtain a visual representation according to the original video and the text question;
[0014] The answerer is used to obtain an answer result according to the original video, the text question, and the visual representation;
[0015] The aggregation - pruning sampler includes: a text encoder module, a video encoder module, a conditional token aggregator module, and a conditional token pruner;
[0016] The answerer is a retrieval - based answerer or a generation - based answerer;
[0017] The text encoder module is used to extract text features from the text question;
[0018] The video encoder module is used to extract visual features from the original video;
[0019] The conditional token aggregator module is used to fuse patch tokens in the visual features to obtain an initial visual representation;
[0020] The conditional token pruner module is used to remove noise in the visual representation to obtain a visual representation;
[0021] The conditional marker pruning module includes: a temporal noise pruning sub-module and a spatial noise pruning sub-module;
[0022] The temporal noise pruning sub-module is used to eliminate visual noise in the visual representation;
[0023] The spatial noise pruning sub-module is used to eliminate spatial noise in the visual representation;
[0024] Furthermore, the responder module of the present invention can be a retrieval-based responder or a generation-based responder. The innovation of the present invention lies in the preprocessing before inputting the responder, so the type and version of the responder are not strictly required.
[0025] In a real-world scenario, different questioning perspectives require different granularity features. For example, some questions require a macroscopic perspective (i.e., viewing multiple patches to represent an overall visual concept), while others require a microscopic perspective (i.e., focusing on specific patches representing details). However, visual information, as a continuous representation, lacks high-level semantic concepts of multiple granularities (such as objects, actions, and relationships).
[0026] Therefore, the conditional marker aggregator module (CTA module) is used to
[0027] can enhance the semantic level of visual information and diversify the feature granularity by aggregating similar question-related markers..
[0028] The beneficial effects of the present invention are:
[0029] The present invention is to solve the problem of visual information redundancy caused by spatio-temporal noise in long video question answering. Therefore, an "aggregation-pruning sampler (APSam)" is proposed, which effectively improves the accuracy and reasoning ability of long video question answering, focuses on aggregating similar markers related to questions to diversify the feature granularity, and adaptively prunes visual noise under each question condition, improving the question answering accuracy. Brief Description of the Drawings
[0030] Figure 1 is a schematic diagram of the processing flow of APSam of the present invention;
[0031] Figure 2 is a schematic diagram of the CTA process of the present invention. Detailed Description of the Invention
[0032] Detailed Description of the Invention 1: In combination with Figure 1 - Figure 2 to illustrate the present invention,
[0033] Step 1: Construct a video question answering model based on aggregation-pruning sampling; Obtain a trained video question answering model based on aggregation-pruning sampling;
[0034] Step 2: Input the video to be questioned and the question to be answered into the trained video question answering model based on aggregation-pruning sampling to obtain an answer result; in Step 1, a video question answering model based on aggregation-pruning sampling is constructed; and a trained video question answering model based on aggregation-pruning sampling is obtained; the specific process is as follows:
[0035] 1. Obtain a video question answering training set;
[0036] The video question answering training set includes: an original video (input of the video question answering system), a text question (input of the video question answering system), and an answer identifier ((training label of the video question answering system));
[0037] 2. Construct a video question answering model based on aggregation-pruning sampling; train the video question answering model based on aggregation-pruning sampling according to the video question answering training set to obtain a trained video question answering model based on aggregation-pruning sampling.
[0038] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that
[0039] In step 2, the construction of the video question answering model based on aggregation-pruning sampling includes: an aggregation-pruning sampler and an answerer;
[0040] The aggregation-pruning sampler is used to obtain a visual representation according to the original video and the text question;
[0041] The answerer is used to obtain an answer result according to the original video, the text question, and the visual representation;
[0042] The aggregation-pruning sampler includes: a text encoder module, a video encoder module, a conditional token aggregator module, and a conditional token pruner;
[0043] The answerer is a retrieval-based answerer or a generation-based answerer;
[0044] The text encoder module is used to extract text features from the text question;
[0045] The video encoder module is used to extract visual features from the original video;
[0046] The conditional token aggregator module is used to fuse patch tokens in the visual features to obtain an initial visual representation;
[0047] The conditional token pruner module is used to remove noise in the visual representation to obtain a visual representation;
[0048] The conditional token pruner module includes: a temporal noise pruning sub-module and a spatial noise pruning sub-module;
[0049] The time noise pruning sub-module is used to eliminate visual noise in the visual representation;
[0050] The spatial noise pruning sub-module is used to eliminate spatial noise in the visual representation;
[0051] Furthermore, the answerer module of the present invention can be a retrieval-based answerer or a generation-based answerer. The innovation of the present invention lies in the preprocessing before inputting the answerer, so the type and version of the answerer are not strictly required.
[0052] In a real-world scenario, different questioning perspectives require different granularities of features. For example, some questions require a macroscopic perspective (i.e., viewing multiple patches to represent an overall visual concept), while others require a microscopic perspective (i.e., focusing on specific patches representing details individually). However, visual information, as a continuous representation, lacks high-level semantic concepts at multiple granularities (such as objects, actions, and relationships).
[0053] Therefore, the conditional token aggregator module (CTA module) is used to
[0054] can enhance the semantic level of visual information and diversify the feature granularity by aggregating similar question-related tokens.
[0055] Other steps and parameters are the same as those in the first specific implementation manner.
[0056] Specific implementation manner three: The difference between this implementation manner and the first specific implementation manner is that
[0057] The process of training the video question answering model based on aggregation-pruning sampling according to the video question answering training set to obtain the trained video question answering model based on aggregation-pruning sampling is as follows:
[0058] S1: Input the original video and text questions in the training set into the aggregation-pruning sampler to obtain the first visual representation, the second visual representation, and the text features;
[0059] S2: When the answerer is a retrieval-based answerer,
[0060] Input the first visual representation, the second visual representation, and the text features into the retrieval-based answerer to obtain the answer result;
[0061] When the answerer is a generation-based answerer,
[0062] Input the second visual representation and the text question into the generation-based answerer to obtain the answer result;
[0063] S3: Train the video question answering model based on aggregation-pruning sampling according to the input and output of the model, and obtain the trained video question answering model based on aggregation-pruning sampling.
[0064] Other steps and parameters are the same as those in one of Embodiments 1 to 2.
[0065] Embodiment 4: The difference between this embodiment and Embodiments 1 to 4 is that
[0066] In S1, the original video and text questions in the training set are input into the aggregation-pruning sampler to obtain the first visual representation, the second visual representation, and the text features. The specific process is as follows:
[0067] S1.1: Input the original video V into the video encoder module for encoding to obtain the first visual feature F.
[0068] The original video is represented as V = {f 1 , f 2 , …, f m}, where f m represents the m-th frame of the original video V.
[0069] The first visual feature is represented as where m is the number of frames, n is the number of patch tokens per frame, d is the dimension of the tokens, represents the dimension of the visual feature;
[0070] The visual feature of the m-th frame is represented as where represents the feature of the n-th patch token, d represents the dimension, is the [CLS] embedding token of the frame f m .
[0071] The video encoder module is a Vision-Language Pre-training Model, which is a well-known model in the art.
[0072] The patch token refers to dividing an image of a frame into small blocks of a fixed size, and each small block is called a "patch token".
[0073] The [CLS] embedding token is a special token used for classification tasks in Transformer models such as BERT, and is usually inserted at the beginning of the input sequence.
[0074] Its function is to generate a global text representation. Through the processing of the Transformer, this representation carries the context information of the input sequence and can be used for subsequent classification or other tasks. Many BERT-based variants (such as RoBERTa, ALBERT, etc.) inherit this approach, which is a well-known practice in the field;
[0075] S1.2: Input the text question q into the text encoder module for encoding processing to obtain the text feature F q ;
[0076] The text question is represented as q,
[0077] The text feature is represented as F q ={q c , w1, w2, …, w k}, where q c is the [CLS] embedding token of the question q, and w k is the text feature of the k-th word of the question q;
[0078] The text encoder module is a Vision-Language Pre-training Model, which is a well-known model in the field;
[0079] The text question q is a text question asked in English. If the question is in other languages, it needs to be translated into English first;
[0080] S1.3: Input the first visual feature F obtained in S1.1 and the text feature F obtained in S1.2 q into the conditional token aggregator module for processing to obtain the second visual feature F';
[0081] S1.4: Input the second visual feature F' obtained in S1.3 and the text feature F obtained in S1.2 q into the conditional token pruner for processing to obtain the first visual representation F' t and the second visual representation F' s ; Other steps and parameters are the same as those in any one of the specific embodiments one to three.
[0082] Specific embodiment five: The difference between this embodiment and the specific embodiments one to four is that,
[0083] In S1.3, input the first visual feature F obtained in S1.1 and the text feature F obtained in S1.2 q into the conditional token aggregator module for processing to obtain the second visual feature F'; The specific process is as follows:
[0084] S1.3.1: The text feature F obtained according to S1.2 q Perform an aggregation connection process on the patch tags of each frame of visual feature in the first visual feature F to obtain the visual feature after the aggregation connection process for each frame;
[0085] S1.3.2: Combine the visual features after the aggregation connection process for each frame into a second visual feature F';
[0086] The second visual feature is represented as
[0087] In a real - world scenario, different questioning perspectives require features at different granularities. For example, some questions require a macroscopic perspective (i.e., viewing multiple patches to represent an overall visual concept), while others require a microscopic perspective (i.e., focusing specifically on a particular patch representing details). However, as a continuous representation, visual information lacks high - level semantic concepts at multiple granularities (such as objects, actions, and relationships).
[0088] Therefore, the conditional tag aggregator module (CTA module) is used to
[0089] Can enhance the semantic level of visual information and diversify the feature granularity by aggregating similar question - related tags.
[0090] Other steps and parameters are the same as those in any one of the specific embodiments one to four.
[0091] Specific embodiment six: The difference between this embodiment and the specific embodiments one to five is that
[0092] The text feature F obtained according to S1.2 in the said S1.3.1 q Perform an aggregation connection process on the patch tags in the visual feature of the m - th frame To obtain the visual feature after the aggregation connection process for the m - th frame The specific process is as follows:
[0093] S1.3.1.1: Calculate the first similarity between all the patch tags in the visual feature of the m - th frame And the text feature F q ;
[0094] S1.3.1.2: Sort all the patch tags in the visual feature of the m - th frame According to the similarity; the higher the similarity, the higher the ranking;
[0095] S1.3.1.3: Divide all the patch tags in the visual feature of the m - th frame Into set A and set B according to the sorting order, where the ranking of the patch tags in set A is odd, and the ranking of the patch tags in set A is even;
[0096] S1.3.1.4: Based on all the patch marks of set A and all the patch marks of set B, form patch mark pairs (A i , B j ); i, j ∈ {1, 2, …, n / 2};
[0097] S1.3.1.5: Calculate the second similarity of all patch mark pairs; Select the R patch mark pairs with the highest second similarity for aggregation to obtain the aggregated patch marks;
[0098] Then connect the aggregated patch marks and the unaggregated patch marks into a sequence as the visual feature after the m-th frame aggregation connection process Expressed by the formula as:
[0099]
[0100] where h a and represent the indices of the aggregated patch mark pairs, h u represents the index of the unaggregated patch mark pair, represents the highest R in the second similarity sorting, represents the R-th to the last in the second similarity sorting, avg represents average aggregation, represents a pair of patch marks aggregated according to the index h a , represents the unaggregated patch marks in set A, represents the unaggregated patch marks in set B;
[0101] . In this way, the conditional tag aggregator CTA aggregates the problem-related tags into high-level semantic concepts, making the feature granularity more diverse;
[0102] In deep learning and computer vision, similarity calculation is usually used to evaluate the similarity between two feature vectors. The first similarity and the second similarity in the present invention are the same similarity metrics, and the similarity metrics can be well-known similarity calculation methods such as cosine similarity or Euclidean distance
[0103] ; Other steps and parameters are the same as those in any one of the specific embodiments one to five.
[0104] Specific embodiment seven: The difference between this embodiment and the specific embodiments one to six is that
[0105] In S1.4, the second visual feature F' obtained in S1.3 and the text feature F obtained in S1.2 qInput conditional label pruner processing, get the first visual representation F' t and a second visual representation F' s ; The specific process is:
[0106] S1.4.1: The second visual feature F' obtained from S1.3 and the text feature F obtained from S1.2 q Get the first visual representation F' t ;
[0107] b is the retained frame;
[0108] S1.4.2: According to the first visual representation F' t , the second visual feature F' and the text feature F obtained in S1.2 q Get the second visual representation F' s ;
[0109] The second visual feature F' obtained in S1.4.1 according to S1.3 and the text feature F obtained in S1.2 q Get the first visual representation F' t ;
[0110] S1.4.1.1: Concatenate the [CLS] embeddings of all frames in the second visual feature F' into a temporal representation F t , expressed as
[0111] S1.4.1.2: Based on the text feature F q Q c and time representation F t Calculate the relevance score S t , expressed as:
[0112] S t =Sigmoid(W q (q c )W p (F t ) T )
[0113] Among them, W q and W p It represents the fully connected layer processing, which is used to map the input features to a new feature space. The core of the projection layer is a weight matrix, which converts the input features into output features through matrix multiplication;
[0114] Sigmoid represents the activation function, (F t ) T F t Transpose; relevance score St Indicates the correlation strength between text features and time representation, with a value range of [0,1].
[0115] S1.4.1.3: According to the correlation score S t Construct a probability distribution for each frame, which is expressed by the formula:
[0116]
[0117] where represents the positive correlation score, represents the score of negative correlation;
[0118] Construct a dynamic time pruning matrix E according to the probability distribution constructed for each frame t , which is expressed by the formula:
[0119]
[0120] Gumbel-Softmax (Jang, Gu, and Poole 2017) is a differentiable function for dynamic end-to-end optimization; it is a function processing well-known to those skilled in the art;
[0121] S1.4.1.4: According to the dynamic time pruning matrix E t Perform dynamic time pruning on the time representation F t to obtain the first visual representation F' t , which is expressed by the formula:
[0122] F′ t = E t × F t
[0123] In S1.4.2, according to the first visual representation F' t , the second visual feature F', and the text feature F obtained from S1.2 q obtain the second visual representation F' s ; the specific process is as follows:
[0124] S1.4.2.1: According to the first visual representation F' t Determine the frames for dynamic time pruning, and remove the visual features of the frames for dynamic time pruning in the second visual feature F' to obtain the second visual feature F after dynamic pruning et ,
[0125]
[0126] S1.4.2.2: Connect the patch labels of all frames in the second visual feature F after dynamic time pruning et to obtain the spatial representation F S , Spatial representation of the b-th frame It is expressed by the formula as follows:
[0127]
[0128] Calculate the second visual feature F after dynamic time pruning et Dynamic spatial pruning matrices of all frames in Denote the dynamic spatial pruning matrix of the b-th frame; The calculation formula of is:
[0129]
[0130] Where H q and H p Are processed by fully connected layers; it maps text features and spatial representations to a new feature space for calculating correlation scores
[0131] Obtain the dynamically restricted receptive field E through the patch probability distribution : s :
[0132] S1.4.2.3: According to the dynamic spatial pruning matrix E of each frame s Perform dynamic spatial pruning on the spatial representation F S To obtain the second visual representation F' s ;
[0133] Among them, according to the dynamic spatial pruning matrix of the b-th frame Perform dynamic spatial pruning on the spatial representation of the b-th frame To obtain the second visual representation of the b-th frame The specific process is as follows:
[0134]
[0135] Some methods (Yu et al., 2023; Gao et al., 2023) use static Top-K strategies to reduce noise by sampling a fixed number of frames or patches. However, in real-world scenarios, the number of visual cues required varies with the complexity of the problem. Therefore, we propose the CTP module to adaptively prune temporal and spatial noise according to problem conditions
[0136] Other steps and parameters are the same as those in any one of the specific embodiments one to six.
[0137] Specific embodiment eight: The difference between this embodiment and the specific embodiments one to seven is that
[0138] The retrieval-based answerer in S2 includes a Transformer network and a feed-forward network;
[0139] The first visual representation F' t , the second visual representation F' s and the text feature F q are input into the retrieval-based answerer to obtain an answer result; which is expressed by the formula:
[0140] Q = K = V = [(F q + T q ); (F′ t + T t ); (F′ s + T s )]
[0141]
[0142] Ans = argmax(F ans , (F can ) T )
[0143] where W q , W k , W v are processed by fully connected layers, used to map the input feature vectors to the query, key, and value spaces of the Transformer architecture; F ans represents the answer feature, and F can represents the candidate answer feature;
[0144] The candidate answer feature is set artificially in the answerer. In the retrieval-based answerer, there is usually a candidate answer set, and each candidate answer has a corresponding feature representation.
[0145] Finally, argmax selects the maximum similarity, Ans represents the answer result, FFN represents the feed-forward network processing, Q represents the query vector, K represents the key vector, d k represents the dimension of the key vector;, V represents the value vector, T q represents the question embedding, T t represents the frame embedding, T s represents the patch token embedding;
[0146] are artificially set embedding vectors, used to distinguish different types of input features (question, frame, patch token). They are learnable parameters and are gradually adjusted through optimization during model training.
[0147] The Transformer network and the feed-forward network are well-known network models in the field.
[0148] The generated answerer in S2 includes a Q-Former network and an LLM network.
[0149] The second visual representation F' s and the text question q are input into the generated answerer to obtain an answer result. The specific process is as follows:
[0150] The second visual representation F' s is input into the Q-Former network to obtain a video-text representation A.
[0151] The video-text representation A and the text question q are input into the LLM network to obtain an answer result.
[0152] The Q-Former network is a vision-language bridge (Li et al., 2023a)
[0153] The LLM network is a large language model (LLM, Chiang et al., 2023), which is a well-known network model in the field. Other steps and parameters are the same as those in any one of the first to seventh specific embodiments.
[0154] Specific embodiment nine: The difference between this embodiment and the first to eighth specific embodiments is that
[0155] in S3, according to the input and output of the video question-answering model based on aggregation-pruning sampling, the video question-answering model based on aggregation-pruning sampling is trained to obtain a trained video question-answering model based on aggregation-pruning sampling. The specific process is as follows:
[0156] When the answerer is a retrieval-based answerer, the loss function for training the video question-answering model based on aggregation-pruning sampling is the cross-entropy loss;
[0157] The cross-entropy loss is a well-known loss function in the field;
[0158] When the answerer is a generation-based answerer, the loss function for training the video question-answering model based on aggregation-pruning sampling is the language modeling loss;
[0159] The language modeling loss is a well-known loss function in the field. Other steps and parameters are the same as those in any one of the first to eighth specific embodiments.
[0160] Specific embodiment ten: The difference between this embodiment and the first to ninth specific embodiments is that
[0161] the video to be questioned and answered in step two is a long video; the question to be answered refers to the text question for the video to be questioned and answered;
[0162] The long video refers to a video that exceeds 15s.
[0163] The text problem is a text problem of natural language questions;
[0164] For example, "Was there any dispute among people in the entrance area of the hall this afternoon?"
[0165] "Was there anyone staying in the designated area after the door was closed last night?"
[0166] "Were there any abnormal personnel activities before the fire alarm was detected?"
[0167] Furthermore, the working scenario of the present invention is to use long - video question - answering technology in an intelligent security monitoring system to assist in analyzing and understanding complex multi - scenario and multi - event monitoring video content. The following is a detailed description of this scenario:
[0168] In the monitoring centers of crowded places such as large shopping malls and subway stations, the system generates a large amount of long - duration video data every day, which contains various different scenarios and events, such as customer flow, emergency events, and daily behaviors. Traditional monitoring video analysis requires staff to manually flip through long video records, which is not only time - consuming and laborious but also may overlook important details due to visual fatigue.
[0169] The intelligent security monitoring system will utilize the aggregation and pruning mechanisms in the present invention to extract multi - granularity visual features related to the questions in the video and automatically prune irrelevant scene noises, thereby accurately locating relevant events and scene details and generating accurate answers. This can not only quickly respond to specific requirements but also provide a powerful multi - level semantic understanding ability in complex video streams, facilitating security personnel to quickly obtain key information, improving the security response speed and accuracy. Such an intelligent security monitoring system can also achieve early warning of potential dangerous behaviors, such as detecting high - frequency crowd gatherings or abnormal movement paths, etc., which helps to take corresponding measures in a timely manner and reduce security risks;
[0170] Other steps and parameters are the same as those in any one of the first to ninth specific embodiments.
[0171] Conduct simulation analysis in combination with the first to tenth specific embodiments
[0172] This invention selects three long video question-answering benchmark datasets, AGQAv2, NExT-QA, and STAR, for evaluation. AGQAv2 (Grunde-McLaughlin et al., 2022) is an open-ended visual compositional reasoning dataset that contains over 9.7K videos and 2.27M question-answer pairs. Compared with the original version of AGQA, it has a more balanced distribution and less language bias (Grunde-McLaughlin, Krishna, and Agrawala, 2021). NExT-QA (Xiao et al., 2021) is a multiple-choice question-answering dataset designed specifically for causal reasoning, temporal reasoning, and common scenario understanding. It contains over 5.4K videos and 52K questions. STAR (Wu et al., 2021) is a multiple-choice question-answering dataset for evaluating the situational reasoning ability of models. This dataset contains 22K videos related to human actions or interactions and 60K questions.
[0173] APSam-R (retrieval-based answerer) and APSam-G (generation-based answerer) are compared with multiple state-of-the-art retrieval-based and generation-based video question-answering models on three datasets. The results in Tables 1, 2, and 3 show that APSam has achieved significant performance improvements. Notably, APSam-R uses the same text and video features as MIST-CLIP (Gao et al., 2023) and ATP (Buch et al., 2022), but performs better (+2.1% on AGQAv2, +1.5% on NExT-QA, +1.7% on STAR). APSam-G is initialized based on VideoChat2 (Li et al., 2024), but also achieves significant performance improvements on all datasets (+1.8% on AGQAv2, +2.1% on NExT-QA, +1.4% on STAR). By observing the detailed experimental results of the three datasets, it can be found that APSam-R and APSam-G both show significant improvements in temporal questions (such as sorting questions in AGQAv2, temporal questions in NExT-QA, and sequential questions in STAR) and reasoning questions (such as causal questions in NExT-QA, feasibility questions in STAR). We believe that these questions involve more multi-granularity concepts and thus require a deeper understanding. This highlights the excellent reasoning ability of APSam in complex situations.
[0174] Table 1 Comparison of the accuracies of different retrieval-based and generation-based SOTA models on the AGQAv2 dataset
[0175]
[0176] Table 2 Comparison of the accuracies of different retrieval - based and generation - based SOTA models on the NExT - QA dataset
[0177]
[0178] Table 3 Comparison of the accuracies of different retrieval - based and generation - based SOTA models on the STAR dataset
[0179]
[0180] Replacement tests for some key steps of the present invention
[0181] Here, the present invention observes the performance changes by replacing the implementation methods of some key steps. The optimization contribution results of each key step are shown in Table 4.
[0182] Effect of the Conditional Token Aggregator (CTA): The CTA diversifies the feature granularity by enhancing the visual semantics related to the question. The results in the third row of Table 4 show that when removing the CTA, the accuracy of APSam - G drops significantly. Specifically, this phenomenon is particularly significant in the temporal questions of NExT - QA involving high - level semantics (such as objects and actions) related to multi - granularity diversification, indicating that the CTA effectively improves the diversity of granularity.
[0183] Effect of the Conditional Token Pruner (CTP): To study the effect of adaptive pruning, we removed the CTP from APSam - G. The results in the second row of Table 4 demonstrate the effectiveness of using the CTP to prune irrelevant noise, significantly improving the accuracy.
[0184] Effect of visual noise: When removing both the CTA and the CTP simultaneously, the answerer is completely filled with all visual tokens without considering visual noise. The results in the first row of Table 4 show a significant performance drop, indicating the effectiveness of APSam. In addition, the retention ratio shows that more than half of the tokens (1 - 0.41 = 0.59) are noise, which reveals the necessity of denoising.
[0185] Table 4 Ablation experiments for the Conditional Token Aggregator (CTA) and Conditional Token Pruner (CTP) of APSam
[0186]
[0187] The above is only a description of the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art, within the scope of the technical solution of the present invention, can make some changes or modifications to equivalent embodiments by using the disclosed technical content, but as long as it does not depart from the technical solution content of the present invention, according to the technical essence of the present invention, within the spirit and principles of the present invention, any simple modification, equivalent replacement and improvement of the above embodiments still fall within the protection scope of the technical solution of the present invention.
Claims
1. A video question answering method based on aggregation-pruning sampling, characterized in that: include: Step 1: Build a video question answering model based on aggregation-pruning sampling; obtain a trained video question answering model based on aggregation-pruning sampling; Step 2: Input the video to be answered and the question to be answered into the trained video question answering model based on aggregation-pruning sampling to obtain the answer result; In the step 1, a video question answering model based on aggregation-pruning sampling is constructed; a trained video question answering model based on aggregation-pruning sampling is obtained; the specific process is:
1. Obtain the video question-answering training set; The video question answering training set includes: original video, text questions and answer labels; 2. Build a video question answering model based on aggregation-pruning sampling; The video question answering model based on aggregation-pruning sampling is trained according to the video question answering training set to obtain a trained video question answering model based on aggregation-pruning sampling.
2. The video question answering method based on aggregation-pruning sampling according to claim 1, characterized in that: The second method of constructing a video question answering model based on aggregation-pruning sampling includes: an aggregation-pruning sampler and an answerer; The aggregation-pruning sampler includes: a text encoder module, a video encoder module, a conditional tag aggregator module and a conditional tag pruner; The responder is a retrieval-based responder or a generation-based responder.
3. The video question answering method based on aggregation-pruning sampling according to claim 2 is characterized in that: The specific process of training the video question answering model based on aggregation-pruning sampling according to the video question answering training set to obtain the trained video question answering model based on aggregation-pruning sampling is as follows: S1: The original video and text questions in the training set are input into the aggregation-pruning sampler to obtain the first visual representation, the second visual representation and the text features; S2: When the responder is a retrieval-based responder, Inputting the first visual representation, the second visual representation and the text feature into a retrieval-based answerer to obtain an answer result; When the responder is a generation-based responder, Inputting the second visual representation and the text question into the generated answerer to obtain an answer result; S3: According to the input and output of the video question answering model based on aggregation-pruning sampling, the video question answering model based on aggregation-pruning sampling is trained to obtain a trained video question answering model based on aggregation-pruning sampling.
4. The video question answering method based on aggregation-pruning sampling according to claim 3 is characterized in that: In S1, the original video and text questions in the training set are input into the aggregation-pruning sampler to obtain the first visual representation, the second visual representation and the text features; the specific process is: S1.1: Input the original video V into the video encoder module for encoding to obtain the first visual feature F; The original video is represented as V = {f 1 ,f 2 ,…,f m }, where f m represents the mth frame of the original video V; The first visual feature is expressed as Where m is the number of frames, n is the number of patch labels per frame, and d is the dimension of the label. Dimensions representing visual features; The visual feature of the mth frame is represented as in, represents the features of the nth patch mark, d represents the dimension, is the frame f m The [CLS] embed tag; S1.2: Input the text question q into the text encoder module for encoding to obtain the text feature F q ; The text question is denoted as q, The text feature is represented as F q = {q c ,w1,w2,…,w k }, where q c is the [CLS] embedding tag for question q, w k is the text feature of the kth word in question q; S1.3: Combine the first visual feature F obtained in S1.1 and the text feature F obtained in S1.2 q The input conditional tag aggregator module processes and obtains the second visual feature F'; S1.4: Combine the second visual feature F' obtained in S1.3 with the text feature F obtained in S1.2 q Input conditional tag pruner processing, get the first visual representation F' t and a second visual representation F' s .
5. The video question answering method based on aggregation-pruning sampling according to claim 4 is characterized in that: In S1.3, the first visual feature F obtained in S1.1 and the text feature F obtained in S1.2 are combined. q The input conditional tag aggregator module processes and obtains the second visual feature F'; The specific process is: S1.3.1: Text feature F obtained according to S1.2 q Performing aggregation connection processing on the patch marks of each frame of visual features in the first visual feature F to obtain the visual features of each frame after aggregation connection processing; S1.3.2: Combining the visual features of each frame after aggregation and connection processing into a second visual feature f'; The second visual feature is expressed as 6. The video question answering method based on aggregation-pruning sampling according to claim 5, characterized in that: The text feature F obtained in S1.3.1 according to S1.2 q Visual features of the mth frame The patch marks in the image are aggregated and connected to obtain the visual features of the mth frame after the aggregation connection processing. The specific process is: S1.3.1.1: Calculate visual features of the mth frame All patch labels and text features in F q The first similarity of S1.3.1.2: Based on the similarity, the visual features of the mth frame Sort all patch tags in the ; the higher the similarity, the higher the ranking; S1.3.1.3: Sort the visual features of the mth frame into All patch marks in are divided into set A and set B, where the ranking of patch marks in set A is odd and the ranking of patch marks in set A is even; S1.3.1.4: Based on all patch labels of set A and all patch labels of set B, form a patch label pair (A i ,B j ); i, j ∈ {1, 2, …, n / 2}; S1.3.1.5: Calculate the second similarity of all patch label pairs; select R patch label pairs with the highest second similarity for aggregation to obtain aggregated patch labels; Then the aggregated patch marks are concatenated with the unaggregated patch marks into a sequence as the visual features after the aggregation connection processing of the mth frame The formula is: Among them, h a and the index of the aggregated patch label pair, h u represents the index of the unaggregated patch-marker pair, represents the highest R in the second similarity sorting, represents the Rth to the last one in the second similarity sorting, avg represents the average aggregation, According to the index h a A pair of patches are aggregated, represents the unaggregated patch markers in set A, Represents the patch markers that are not aggregated in set B.
7. The video question answering method based on aggregation-pruning sampling according to claim 6, characterized in that: In S1.4, the second visual feature F' obtained in S1.3 and the text feature F obtained in S1.2 are combined. q Input conditional tag pruner processing, get the first visual representation F' t and a second visual representation F' s ; The specific process is: S1.4.1: The second visual feature F' obtained from S1.3 and the text feature F obtained from S1.2 q Get the first visual representation F' t ; b<m; b is the frame after being retained; S1.4.2: According to the first visual representation F' t , the second visual feature F' and the text feature F obtained in S1.2 q Get the second visual representation F' s ; The second visual feature F' obtained in S1.4.1 according to S1.3 and the text feature F obtained in S1.2 q Get the first visual representation F' t ; S1.4.1.1: Concatenate the [CLS] embeddings of all frames in the second visual feature F' into a temporal representation F t , expressed as S1.4.1.2: Based on the text feature F q Q c and time representation F t Calculate the relevance score S t , expressed as: Among them, W q and W p It represents the fully connected layer processing. Sigmoid represents the activation function, F t Transpose; S1.4.1.3: According to the relevance score S t Construct a probability distribution for each frame, expressed as: in represents the positive correlation score, denoting a score of negative correlation; Construct the dynamic time pruning matrix E according to the probability distribution of each frame t , expressed as: Gumbel-Softmax is a differentiable function that is dynamically optimized end-to-end; S1.4.1.4: Pruning matrix E according to dynamic time t F for time t Perform dynamic time pruning to obtain the first visual representation F' t , expressed as: F′ t =E t ×F t S1.4.2 according to the first visual representation F' t , the second visual feature F' and the text feature F obtained in S1.2 q Get the second visual representation F' s ; The specific process is: S1.4.2.1: According to the first visual representation F' t Determine the frame of dynamic time inter-pruning, remove the visual features of the frame of dynamic time inter-pruning in the second visual feature F', and obtain the second visual feature F after dynamic time inter-pruning et , S1.4.2.2: The second visual feature F after dynamic time pruning et The patch labels of all frames in are concatenated into a spatial representation F S , Spatial representation of frame b The formula is: Calculate the second visual feature F after dynamic time pruning et The dynamic spatial pruning matrix for all frames in represents the dynamic spatial pruning matrix of the b-th frame; The calculation formula is: Among them, H q and H p Processing for the fully connected layer: S1.4.2.3: Pruning matrix E according to the dynamic space of each frame s For space representation F S Perform dynamic spatial pruning to obtain the second visual representation F' s ; Among them, according to the dynamic space pruning matrix of the b-th frame Spatial representation of frame b Perform dynamic spatial pruning to process the second visual representation of the bth frame The specific process is:
8. The video question answering method based on aggregation-pruning sampling according to claim 7, characterized in that: The retrieval-based answerer in S2 includes a Transformer network and a feedforward network; The first visual representation F' t , second visual representation F' s and text features F q Input the retrieval-based answerer and get the answer result; it can be expressed as: Q=K=V=[(F q +T q );(F′ t +T t );(F′ s +T s )] Among them, W q ,W k ,W v It is processed by the fully connected layer; F ans Indicates the answer feature, F can Represents the characteristics of candidate answers; Finally, argmax selects the maximum similarity, Ans represents the answer result, FFN represents feedforward network processing, Q represents the query vector, K represents the key vector, and d k represents the dimension of the key vector; V represents the value vector, T q Represents the question embedding, T t represents frame embedding, T s represents patch tag embedding; The generation-based answerer in S2 includes a Q-Former network and an LLM network. The second visual representation F' s The text question q is input based on the generated answerer to get the answer result; the specific process is: The second visual representation F' s Input the Q-Former network to obtain the video text representation A; Input the video text representation A and the text question q into the LLM network to get the answer result.
9. The video question answering method based on aggregation-pruning sampling according to claim 8, characterized in that: In S3, the video question answering model based on aggregation-pruning sampling is trained according to the input and output of the video question answering model based on aggregation-pruning sampling to obtain the trained video question answering model based on aggregation-pruning sampling. The specific process is as follows: When the answerer is a retrieval-based answerer, the loss function for training the video question answering model based on aggregation-pruning sampling is the cross-entropy loss; When the answerer is a generation-based answerer, the loss function for training the video question answering model based on aggregation-pruning sampling is the language modeling loss.
10. The video question answering method based on aggregation-pruning sampling according to claim 9, characterized in that: The video to be answered in step 2 is a long video; the question to be answered refers to a text question for the video to be answered; The long video refers to a video that is longer than 15 seconds; The text question is a text question asked in natural language.