A Temporal Enhancement Method and System for Video Question Answering
By defining five timing dimensions and corresponding task designs in timing video Q&A, constructing multi-dimensional timing instruction data and fine-tuning, the problems of insufficient timing relationship modeling and incomplete evaluation benchmarks in the existing technology are solved, and the objectivity and accuracy of timing Q&A capabilities of the video Q&A model are improved and the evaluation objective and accuracy of the evaluation.
Patent Information
- Application Number
- CN202510156668.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The prior art ignores the timing relationship between video frames in timing video Q&A, resulting in poor performance of the model in tasks of high timing perception and understanding. At the same time, the existing evaluation benchmarks have insufficient timing dimension coverage and data distribution, resulting in inaccurate model evaluation.
A time-enhanced video Q&A method is proposed. By defining five timing dimensions (time dynamics, timing reasoning, duration, time positioning and time sequence) and corresponding task design, multi-dimensional timing instruction data is constructed, and through multi-task timing instruction fine-tuning and multi-dimensional timing question and answer evaluation, the timing question and answer capabilities of the video Q&A model are enhanced, and an objective and unbiased evaluation benchmark is provided.
It effectively improves the multi-dimensional timing question and answer capabilities of the video Q&A model, provides a comprehensive and unbiased evaluation benchmark, and can more accurately reflect the model's performance in timing question and answer.
Smart Images

Figure CN119670896B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of temporal video question answering, and particularly to a temporal enhancement video question answering method and system. Background Art
[0002] Video question answering aims to correctly answer questions regarding video content. Different from image question answering that processes static images, video data contains complex temporal dynamic characteristics. Therefore, the perception and understanding of video content in the temporal dimension pose a key technical challenge for video question answering.
[0003] To improve the performance of video question answering models in temporal question answering, adapter-based methods aim to design a new type of adapter to significantly reduce the number of video tokens. Such methods selectively downsample visual tokens in video frames by introducing spatio-temporal convolution or attention mechanisms in the adapter, thereby increasing the number of video frames that can be processed. Meanwhile, by introducing strategies such as long-context training and multi-graph training, training strategy-based methods have significantly improved the ability of video question answering models to process long videos. In addition, data-based methods construct temporally sensitive datasets and enhance the model's perception and understanding ability of temporal dynamics by fine-tuning the video question answering model on these datasets.
[0004] In addition, to comprehensively evaluate the temporal question answering ability of video question answering models, it is urgent to establish a robust and unbiased evaluation benchmark. Early evaluation benchmarks mainly focused on the static visual perception and understanding ability of the models, while ignoring the importance of the dynamic temporal dimension. To make up for this deficiency, a number of evaluation benchmarks specifically for temporal question answering have emerged recently. These benchmarks introduce diverse temporal tasks such as action recognition and temporal prediction, and cover a wider range of evaluation scenarios.
[0005] Although the above research work has made certain progress in the field of temporal video question answering, there are still the following problems:
[0006] (1) Ignoring the explicit modeling of the temporal relationship between video frames. Simply increasing the number of input video frames cannot fundamentally improve the model's perception and understanding ability of temporal dynamics. Existing research has shown that existing video question answering models perform poorly in tasks that require high temporal perception and understanding. These models rely more on prior knowledge or statistical biases during the pre-training process to complete temporal question answering tasks, rather than truly understanding the temporal dependencies between video frames.
[0007] (2) Insufficient coverage in the temporal dimension. Existing benchmarks fail to comprehensively cover all key temporal dimensions in temporal question-answering tasks. Although current benchmarks have designed a variety of task forms, in terms of the temporal question-answering capabilities required to actually complete these tasks, most tasks only involve two to three temporal dimensions. In addition, there are generally statistical biases in the data distribution of existing benchmarks, enabling video question-answering models to rely on statistical priors for evaluation, thus unable to accurately reflect the true performance of the models in temporal question-answering. Summary of the Invention
[0008] To address the deficiencies of the prior art, the present invention provides a method for temporally enhanced video question-answering, which, while enhancing the multi-dimensional temporal question-answering capabilities of video question-answering models, provides an objective and unbiased evaluation benchmark for video temporal question-answering.
[0009] The present invention also provides a temporally enhanced video question-answering system.
[0010] To achieve the above objectives, the present invention adopts the following technical solutions:
[0011] In the first aspect, the present invention provides a method for temporally enhanced video question-answering.
[0012] A method for temporally enhanced video question-answering includes:
[0013] Step 1: Construction of multi-dimensional temporal instruction data;
[0014] Identify and define five temporal dimensions, and establish a data collection and screening process to ensure the completeness of data preparation for each temporal dimension;
[0015] Among them, data collection and screening include the following specific steps:
[0016] Design specialized tasks for each temporal dimension;
[0017] Collect an open-source video dataset with rich temporal dynamics, exclude videos with ineligible durations, and automatically generate questions and answers based on temporal annotations, where the answers include open-ended answers and multiple-choice answers;
[0018] Sample the generated questions and answers;
[0019] Step 2: Multi-task temporal instruction fine-tuning;
[0020] Construct multiple temporal auxiliary tasks, and use the collected data to fine-tune the video question-answering model to enhance the model's temporal question-answering capabilities;
[0021] Among them, constructing multiple temporal auxiliary tasks includes:
[0022] Collect video instruction fine-tuning data, filter out the data containing temporal Q&A using an open-source pre-trained video Q&A model, and construct auxiliary tasks based on the remaining data;
[0023] Construct a video frame index prediction task, randomly select a frame from the original video frame sequence and place it at the front of the sequence, and let the video Q&A model predict the original position index of this video frame;
[0024] Construct an indicative video Q&A task, splice multiple videos, and indicate the target video in natural language form, and let the video Q&A model answer questions based on the target video;
[0025] Step 3: Multi-dimensional temporal Q&A evaluation;
[0026] For five temporal dimensions, additionally collect and construct an evaluation dataset to evaluate the temporal Q&A ability of the video Q&A model.
[0027] According to the preference of the present invention, clarify and define five temporal dimensions, including: classify the video temporal Q&A ability into five dimensions: temporal dynamics, temporal reasoning, duration, temporal localization, and temporal order.
[0028] According to the preference of the present invention, design dedicated tasks for each temporal dimension, including:
[0029] For temporal dynamics, the task form is designed to answer the movement direction of a specified object in the video; among them, the specified object is described in natural language; the movement direction is represented by direction nouns, including four categories: "upper left", "lower left", "upper right", and "lower right";
[0030] For temporal reasoning, the task form is designed to answer the possible next action based on the existing content in the video, where the action is represented by a gerund phrase;
[0031] For duration, the task form is designed to answer the duration of a specified action behavior, where the action behavior is specified in natural language form; the duration has three categories: long, medium, and short, which are respectively expressed as "longer than two-thirds of the video", "between one-third and two-thirds of the video", and "shorter than one-third of the video";
[0032] For temporal localization, the task form is designed to answer the occurrence time of a specified action behavior, where the action behavior is specified in natural language form, and the occurrence time is expressed as three categories: "at the front of the video", "in the middle of the video", and "at the back of the video";
[0033] For temporal order, the task form is designed to answer the correct order of occurrence of three specified actions, where each action is represented by a gerund phrase.
[0034] Preferably according to the present invention, an open-source video dataset with rich temporal dynamics is collected, videos with durations not meeting the requirements are excluded, and questions and answers are automatically generated based on temporal annotations, including:
[0035] Collect videos and temporal annotations from a dataset containing rich temporal dynamics; exclude videos with durations not meeting the requirements;
[0036] Generate open-ended question-and-answer pairs. Manually construct several open-ended question samples, and use ChatGPT to generate several different open-ended questions for each temporal dimension based on the constructed open-ended question samples and task descriptions, and generate open-ended answers based on the video temporal annotations.
[0037] Generate multiple-choice question-and-answer pairs. Manually construct several multiple-choice question samples, and use ChatGPT to generate several different multiple-choice questions for each temporal dimension based on the constructed multiple-choice question samples and task descriptions, and generate one correct option and three incorrect options based on the video temporal annotations.
[0038] Preferably according to the present invention, sample the generated questions and answers; including:
[0039] For the three temporal dimensions of time dynamics, duration, and time localization, which have fixed answer options, balance the quantities between different options through sampling to keep the quantities between options consistent;
[0040] For the two temporal dimensions of temporal reasoning and time order, which have no fixed answer options, take the action category with the smallest quantity as the standard, and balance the quantities between all action categories through sampling to avoid long-tailed distributions;
[0041] Balance the quantities of open-ended question-and-answer pairs and multiple-choice question-and-answer pairs through sampling.
[0042] Preferably according to the present invention, collect video instruction fine-tuning data and filter out the data containing temporal question-and-answer using an open-source pre-trained video question-and-answer model; including:
[0043] Construct the input instructions for the pre-trained video question-and-answer model;
[0044] Input the constructed instructions and videos into the pre-trained video question-and-answer model. For each instruction, there is a corresponding video; according to the responses of the pre-trained video question-and-answer model, filter out the video question-and-answer pairs that do not contain temporal content.
[0045] Preferably according to the present invention, construct a video frame index prediction task; including:
[0046] For the video question-and-answer pairs that do not contain temporal content, randomly extract one frame from their original video frame sequence and place it at the front of the video frame sequence;
[0047] Construct instructions. Use ChatGPT to generate 10 instructions based on manually constructed examples, and randomly sample one instruction for each question-answer pair.
[0048] Construct answers.
[0049] Preferably according to the present invention, construct an indicative video question-answer task; including:
[0050] For video question-answer pairs without temporal content, randomly select a video from the dataset and randomly splice it in front of or behind the original video.
[0051] Construct instructions. Use ChatGPT to generate 10 instructions based on manually constructed examples, and randomly sample one instruction for each question-answer pair.
[0052] Directly use the answer of the original video as the answer of the spliced video.
[0053] Preferably according to the present invention, fine-tune the video question-answer model with three parts of data: the mixed-constructed multi-dimensional temporal instruction data, the constructed multi-task temporal instructions, and the remaining video instruction fine-tuning data, i.e., video question-answer pairs containing temporal content; including:
[0054] Input the instruction text into the encoding layer to obtain text semantic features ;
[0055] Input the video frame sequence into the visual encoder to obtain visual features , and use a visual adapter to map the visual features into the text semantic space to obtain visual semantic features aligned with the text semantic features ;
[0056] The text semantic features and the visual semantic features are spliced and then input into the pre-trained large-scale language model, and the model output is supervised and fine-tuned using cross-entropy loss.
[0057] Preferably according to the present invention, for five temporal dimensions, additionally collect and construct an evaluation dataset to evaluate the temporal question-answer ability of the video question-answer model; including:
[0058] Construct a multi-dimensional temporal question-answer evaluation dataset; including:
[0059] Adopt steps similar to the above to construct the evaluation dataset.
[0060] Using three open-source pre-trained video question-answering models, Qwen2-VL, LLaVA-OneVision, and MiniCPM-V2.6, as filters, input the randomly sampled single-frame video content and the constructed question-answer pairs. Determine whether the filters can answer the questions correctly based on the answers output by the filters. If at least two filters can answer the questions correctly, it is determined that there is a single-frame deviation in the video question-answer pair;
[0061] Delete the question-answer pairs with single-frame deviation from the dataset, and rebalance the number between different answer candidates by sampling to obtain the final evaluation dataset;
[0062] Input the videos and question-answer pairs in the constructed evaluation dataset into the video question-answering model to obtain answers, and compare whether the answers output by the video question-answering model are the same as the true answers;
[0063] Statistically calculate the proportion of the number of questions correctly answered by the video question-answering model to the total number of questions.
[0064] In a second aspect, the present invention provides a video question-answering system with temporal enhancement.
[0065] A video question-answering system with temporal enhancement includes:
[0066] A multi-dimensional temporal instruction data construction unit, configured to: clarify and define five temporal dimensions, establish a data collection and screening process, and ensure the completeness of data preparation for each temporal dimension;
[0067] A multi-task temporal instruction fine-tuning unit, configured to: construct multiple temporal auxiliary tasks, and use the collected data to fine-tune the video question-answering model to enhance the temporal question-answering ability of the model;
[0068] A multi-dimensional temporal question-answering evaluation unit, configured to: for the five temporal dimensions, additionally collect and construct an evaluation dataset to evaluate the temporal question-answering ability of the video question-answering model.
[0069] Compared with the prior art, the beneficial effects of the present invention are:
[0070] 1. The present invention innovatively proposes a video question-answering method and system with temporal enhancement, including multi-dimensional temporal instruction data construction, multi-task temporal instruction fine-tuning, and multi-dimensional temporal question-answering evaluation, which makes up for the deficiencies of the prior art in temporal question-answering.
[0071] 2. The multi-dimensional temporal instruction data construction part in the present invention summarizes the five temporal dimensions involved in temporal question-answering, and designs specific data collection and construction methods for each dimension, providing a data basis for enhancing the temporal question-answering ability of the video question-answering model.
[0072] 3. In the present invention, the multi-task temporal instruction fine-tuning part integrates multiple temporal auxiliary tasks that do not require additional supervision information, breaks through the limitation of data capacity, and can effectively improve the temporal question-answering ability of the video question-answering model in multiple dimensions.
[0073] 4. In the present invention, the multi-dimensional temporal question-answering evaluation part collects a large amount of video data and constructs unbiased question-answer pairs, providing an objective and comprehensive evaluation benchmark for evaluating the temporal question-answering ability of the video question-answering model. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0075] Figure 1 is a flowchart of the multi-dimensional temporal instruction data construction method of the present invention;
[0076] Figure 2 is a schematic diagram of the multi-task temporal instruction fine-tuning method of the present invention;
[0077] Figure 3 is a schematic diagram of the video question-answering model of the present invention;
[0078] Figure 4 is a schematic diagram of the multi-dimensional temporal question-answering evaluation method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0080] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0081] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0082] Embodiment 1
[0083] A temporal-enhanced video question-answering method, comprising:
[0084] Step 1: Multi-dimensional temporal instruction data construction; as Figure 1 shown:
[0085] Define and clarify five temporal dimensions, establish a data collection and screening process, and ensure the completeness of the data preparation work for each temporal dimension;
[0086] Among them, the data collection and screening include the following specific steps:
[0087] Design dedicated tasks for each temporal dimension; these tasks rely on the temporal Q&A capabilities of the corresponding dimension to be completed;
[0088] Collect an open-source video dataset with rich temporal dynamics, exclude videos with durations not meeting the requirements, and automatically generate questions and answers based on temporal annotations, where the answers include open-ended responses and multiple-choice responses;
[0089] Sample the generated questions and answers; to balance the data distribution and eliminate the long-tail distribution;
[0090] Step 2: Multi-task temporal instruction fine-tuning;
[0091] Construct multiple temporal auxiliary tasks, and use the collected data to fine-tune the video Q&A model to enhance the model's temporal Q&A capabilities;
[0092] Among them, constructing multiple temporal auxiliary tasks includes:
[0093] Collect video instruction fine-tuning data, filter the data containing temporal Q&A using an open-source pre-trained video Q&A model, and construct auxiliary tasks based on the remaining data;
[0094] Construct a video frame index prediction task, randomly select a frame from the original video frame sequence and place it at the front of the sequence, and let the video Q&A model predict the original position index of this video frame;
[0095] Construct an indicative video Q&A task, splice multiple videos, and indicate the target video in natural language form, and let the video Q&A model answer questions based on the target video;
[0096] Step 3: Multi-dimensional temporal Q&A evaluation;
[0097] For five temporal dimensions, additionally collect and construct an evaluation dataset to evaluate the temporal Q&A capabilities of the video Q&A model.
[0098] Example 2
[0099] According to a temporal-enhanced video Q&A method described in Example 1, the difference is:
[0100] Clarify and define five temporal dimensions, including: Summarize the video temporal Q&A capabilities into five dimensions: temporal dynamics, temporal reasoning, duration, temporal localization, and temporal order.
[0101] Design dedicated tasks for each temporal dimension, including:
[0102] The questions in the temporal dynamic dimension focus on the dynamic changes of the content in the video over time. For temporal dynamics, the task form is designed to answer the movement direction of a specified object in the video; among them, the specified object is described in natural language; such as "the red car", and the movement direction is represented by direction nouns, including four categories: "upper left", "lower left", "upper right", and "lower right".
[0103] The questions in the temporal reasoning dimension focus on the events that will occur in the future. For temporal reasoning, the task form is designed to answer the possible next actions based on the existing content in the video, among which the actions are represented by gerund phrases; such as "picking up the basketball".
[0104] The questions in the duration dimension focus on the duration of the event. For the duration, the task form is designed to answer the duration of the specified action behavior, among which the action behavior is specified in natural language form; such as "a boy crossed the road", and the duration is divided into three categories: long, medium, and short, which are respectively expressed as "longer than two-thirds of the video", "between one-third and two-thirds of the video", and "shorter than one-third of the video".
[0105] The questions in the time location dimension focus on the time point when a specific event occurs in the video. For time location, the task form is designed to answer the occurrence time of the specified action behavior, among which the action behavior is specified in natural language form, and the occurrence time is expressed as three categories: "at the front of the video", "in the middle of the video", and "at the back of the video".
[0106] The questions in the time sequence dimension focus on the order of occurrence of events. For the time sequence, the task form is designed to answer the correct order of occurrence of three specified actions, among which each action is represented by a gerund phrase. Such as "picking up the box, opening the box, taking out the toy".
[0107] Collect an open-source video dataset with rich temporal dynamics, exclude videos with durations that do not meet the requirements, and automatically generate questions and answers based on temporal annotations, including:
[0108] Collect videos and temporal annotations from datasets containing rich temporal dynamics such as ActivityNet Caption, VidOR, Ego4d Goal-Step, and Charades; exclude videos with durations not meeting the requirements; ActivityNet Caption, a large video caption dataset, consists of videos from various scenarios, each with a corresponding natural language description that usually involves actions, events, or scenes shown in the video; VidOR, composed of multiple video segments, with dynamically changing objects in each video, and provides detailed annotation information for each video, including object categories, locations, timestamps, etc.; Ego4d Goal-Step, a first-person immersive video dataset that mainly focuses on first-person perspective videos captured by wearable cameras (such as smart glasses); Charades, a multi-label dataset for video action recognition that particularly focuses on human activity recognition in daily life.
[0109] Generate open-ended question-answer pairs. Manually construct a small number of open-ended question samples. Use ChatGPT to generate a number of (10) different open-ended questions for each temporal dimension based on the constructed open-ended question samples and task descriptions, and generate open-ended answers based on the video temporal annotations; e.g., "Question: Can you determine the direction in which the <child> moves relative to the video frame? Answer: Lower left". ChatGPT, an artificial intelligence question-answering system based on a large language model, specifically designed to process and generate natural language text;
[0110] Generate multiple-choice question-answer pairs. Manually construct a small number of multiple-choice question samples. Use ChatGPT to generate a number of (10) different multiple-choice questions for each temporal dimension based on the constructed multiple-choice question samples and task descriptions, and generate one correct option and three incorrect options based on the video temporal annotations. E.g., "Question: Where does the part <She increases the volume of her hair by combing it> appear in the video? Options: (a) In the middle of the video (b) At the beginning of the video (c) At the end of the video. Answer: (a)".
[0111] Sample the generated questions and answers; to balance the data distribution and eliminate long-tailed distributions and single-frame biases. Include:
[0112] For the three temporal dimensions of time dynamics, duration, and time localization, which have fixed answer options, balance the quantities between different options through sampling to keep the quantities between options consistent; for example, if they are 450, 670, 890 respectively, all sample them to 450.
[0113] For the two temporal dimensions of temporal reasoning and chronological order, which have no fixed answer options, the smallest number of action categories is used as the standard, and sampling is used to balance the quantities among all action categories to avoid long-tailed distributions;
[0114] Specifically, if the frequencies of some answer options in the dataset are low, an oversampling method is adopted, that is, by duplicating existing question-and-answer pairs to increase the number of low-frequency answer options; while if the frequencies of some answer options are high, then through an undersampling method, some question-and-answer pairs of high-frequency answer options are randomly deleted to achieve the purpose of balancing the quantities of each option.
[0115] Balance the quantities of open-ended question-and-answer pairs and multiple-choice question-and-answer pairs through sampling.
[0116] Construct multiple temporal auxiliary tasks, and use the collected data to fine-tune the video question-answering model to enhance the model's temporal question-answering ability, as Figure 2 shown.
[0117] Collect a video instruction fine-tuning dataset. For example, Video-ChatGPT is a multi-modal dataset that combines video and question-answering, aiming to enable the model to interact with users in natural language based on watching video content and conduct more complex video question-answering; use an open-source pre-trained video question-answering model, such as Qwen2.5-72B, which is an open-source large-scale language model, and "72B" refers to the number of parameters of the model being 7.2 billion. Filter the data in the video instruction fine-tuning dataset, including:
[0118] Construct the input instructions for the pre-trained video question-answering model; the form is "Analyze the following conversation between a human and an assistant to determine whether the conversation reflects an understanding of the video content in the temporal dimension. If so, please reply <Yes>. If not, please reply <No>. The conversation content is as follows: {Human: [Question], Assistant: [Answer]}", where "[Question]" and "[Answer]" are placeholders and need to be replaced with the corresponding question-and-answer content of the video in the instruction fine-tuning dataset.
[0119] Input the constructed instructions and videos into the pre-trained video question-answering model. Each instruction has a corresponding video; according to the reply of the pre-trained video question-answering model, filter out the video question-and-answer pairs that do not contain temporal content.
[0120] Construct a video frame index prediction task, including:
[0121] For the video question-and-answer pairs that do not contain temporal content, randomly select one frame from their original video frame sequence and place it at the front of the video frame sequence;
[0122] Construct instructions. Use ChatGPT to generate 10 instructions based on manually constructed examples. For each question-answer pair, randomly sample one instruction. The final instruction form is like "The first frame in the sequence is a randomly sampled frame from the video, and the remaining video frames are arranged in the original order. Please first predict the position of the first frame in the original video frame sequence, and then perform the following task. Question: [Question]". Here, "[Question]" is a placeholder that needs to be replaced with the question content corresponding to the video.
[0123] Construct answers. The form is "(2, 3). [Answer]". Here, "(2, 3)" represents the original position index of the sampled frame, and "[Answer]" is a placeholder that needs to be replaced with the answer content corresponding to the video.
[0124] Construct video question-answering tasks; including:
[0125] For video question-answer pairs without temporal content, randomly select a video from the dataset, i.e., Video-ChatGPT, and randomly splice it in front of or behind the original video;
[0126] Construct instructions. Use ChatGPT to generate 10 instructions based on manually constructed examples. For each question-answer pair, randomly sample one instruction. The final instruction form is like "This is the splicing of two videos on the timeline. Please perform the following task according to the [index]th video. Question: [Question]". Here, "[index]" is a placeholder for the position index of the original video, and "[Question]" is a placeholder corresponding to the original video content.
[0127] Directly use the answer of the original video as the answer for the spliced video.
[0128] As Figure 3 shown, fine-tune the video question-answering model with three parts of data: the mixed-constructed multi-dimensional temporal instruction data, the constructed multi-task temporal instructions, and the remaining video instruction fine-tuning data, i.e., video question-answer pairs containing temporal content; including:
[0129] Input the instruction text into the encoding layer to obtain text semantic features ;
[0130] Input the video frame sequence into the visual encoder to obtain visual features , and use the visual adapter to map the visual features into the text semantic space to obtain visual semantic features aligned with the text semantic features ;
[0131] Concatenate the text semantic features and the visual semantic features and input them into the pre-trained large-scale language model, and use cross-entropy loss to perform supervised fine-tuning on the model output.
[0132] For five temporal dimensions, additional evaluation datasets are collected and constructed to evaluate the temporal question - answering ability of the video question - answering model; as Figure 4 shown, including:
[0133] Construct a multi - dimensional temporal question - answering evaluation dataset; including:
[0134] Use steps similar to the above to construct the evaluation dataset; the difference is that only multiple - choice question - answer pairs are constructed, and different data sources are selected to verify the out - of - domain generalization ability.
[0135] Use three open - source pre - trained video question - answering models, Qwen2 - VL, LLaVA - OneVision, and MiniCPM - V2.6, as filters. Input randomly sampled single - frame video content and the constructed question - answer pairs. Determine whether the filter can answer the question correctly according to the answers output by the filter. If at least two filters can answer the question correctly, it is determined that there is a single - frame bias in the video question - answer pair;
[0136] Delete the question - answer pairs with single - frame bias from the dataset, and re - balance the number between different answer candidates by sampling to obtain the final evaluation dataset;
[0137] Input the videos and question - answer pairs in the constructed evaluation dataset into the video question - answering model to get answers, and compare whether the answers output by the video question - answering model are the same as the true answers;
[0138] Statistically calculate the proportion of the number of questions correctly answered by the video question - answering model to the total number of questions.
[0139] Example 3
[0140] A temporally enhanced video question - answering system, including:
[0141] A multi - dimensional temporal instruction data construction unit, configured to: clarify and define five temporal dimensions, establish a data collection and screening process, and ensure the completeness of data preparation for each temporal dimension;
[0142] A multi - task temporal instruction fine - tuning unit, configured to: construct multiple temporal auxiliary tasks, and use the collected data to fine - tune the video question - answering model to enhance the temporal question - answering ability of the model;
[0143] A multi - dimensional temporal question - answering evaluation unit, configured to: for five temporal dimensions, additionally collect and construct an evaluation dataset to evaluate the temporal question - answering ability of the video question - answering model.
[0144] It can be understood that the above-mentioned units can be separately or all combined into one or several other units to form, or some of them can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In actual applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the video behavior segment candidate set generation system may also include other units. In actual applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.
Claims
1. A time-enhanced video question-answering method, characterized in that: include: Step 1: Multi-dimensional sequential instruction data construction; Clarify and define the five time series dimensions, establish a data collection and screening process, and ensure that the data preparation for each time series dimension is complete; Among them, data collection and screening include the following specific steps: Design specialized tasks for each temporal dimension; Collect open-source video datasets with rich temporal dynamics, exclude videos that do not meet the duration requirements, and automatically generate questions and answers based on temporal annotations, where the answers include open-ended answers and multiple-choice answers; Sampling generated questions and answers; Step 2: Fine-tune multi-task timing instructions; Construct multiple time-series auxiliary tasks and use the collected data to fine-tune the video question-answering model to enhance the model's time-series question-answering capabilities; Among them, multiple time series auxiliary tasks are constructed, including: Collect video instruction fine-tuning data, use the open source pre-trained video question-answering model to filter the data containing time-series questions and answers, and construct auxiliary tasks based on the remaining data; Construct a video frame index prediction task, randomly extract a frame from the original video frame sequence and place it at the front of the sequence, and use the video question answering model to predict the original position index of the video frame; Construct an instruction video question-answering task, splice multiple videos, and indicate the target video in natural language. The video question-answering model will answer questions based on the target video. Step 3: Multi-dimensional sequential question answering evaluation; For the five time series dimensions, we collected and constructed additional evaluation datasets to evaluate the time series question answering capabilities of the video question answering model. Clarify and define five temporal dimensions, including: summarizing video temporal question-answering capabilities into five dimensions: temporal dynamics, temporal reasoning, duration, temporal positioning, and temporal order.
2. The method for video question answering with time sequence enhancement according to claim 1, characterized in that: Design specialized tasks for each time series dimension, including: In view of temporal dynamics, the task form is designed to answer the movement direction of a specified object in the video; the specified object is described in natural language; the movement direction is represented by a direction noun, including four categories: "upper left", "lower left", "upper right" and "lower right"; For temporal reasoning, the task format is designed to answer the possible actions that may occur next based on the existing content of the video, where the actions are represented by gerund phrases; Regarding duration, the task format is designed to answer the duration of a specified action behavior, where the action behavior is specified in natural language; the duration is divided into three categories: long, medium, and short, which are respectively expressed as "longer than two-thirds of the video", "between one-third and two-thirds of the video", and "shorter than one-third of the video"; For time positioning, the task format is designed to answer the occurrence time of a specified action behavior, where the action behavior is specified in natural language form, and the occurrence time is expressed in three categories: "located in the front of the video", "located in the middle of the video", and "located in the back of the video"; For the time sequence, the task format is designed to answer the correct order in which three specified actions occur, where each action is represented by a gerund phrase.
3. The method for video question answering with time sequence enhancement according to claim 1, characterized in that: Collect open-source video datasets with rich temporal dynamics, exclude videos that do not meet the duration requirements, and automatically generate questions and answers based on temporal annotations, including: Collect videos and temporal annotations from datasets containing rich temporal dynamics; exclude videos whose duration does not meet the requirements; Generate open-ended question-answer pairs, manually construct several open-ended question samples, use ChatGPT to generate several different open-ended questions for each temporal dimension based on the constructed open-ended question samples and task description, and generate open-ended answers based on the video temporal annotations; Generate multiple-choice question-answer pairs, manually construct several multiple-choice question samples, use ChatGPT to generate several different multiple-choice questions for each temporal dimension based on the constructed multiple-choice question samples and task description, and generate one correct option and three incorrect options based on the video temporal annotation.
4. The time-enhanced video question-answering method according to claim 1, characterized in that: Sample generated questions and answers; including: For the three temporal dimensions of time dynamics, duration, and time location, which have fixed answer options, the number of different options is balanced through sampling to keep the number of options consistent; For the two temporal dimensions of temporal reasoning and time order, which have no fixed answer options, the action category with the least number is used, and the number of action categories is balanced through sampling to avoid long-tail distribution; The number of open-ended question-answer pairs and multiple-choice question-answer pairs was balanced through sampling.
5. The time-order enhanced video question-answering method according to claim 1, characterized in that: Collect video instruction fine-tuning data and use the open source pre-trained video question answering model to filter the data containing time-series question answering; including: Construct input instructions for the pre-trained video question answering model; The constructed input instructions and videos are input into the pre-trained video question-answering model, and each instruction has a corresponding video. According to the response of the pre-trained video question-answering model, the video question-answer pairs that do not contain temporal content are filtered out.
6. The time-order enhanced video question-answering method according to claim 1, characterized in that: Construct a video frame index prediction task; including: For video question-answer pairs that do not contain temporal content, a frame is randomly selected from the original video frame sequence and placed at the front of the video frame sequence; Construct instructions, use ChatGPT to generate 10 instructions based on manually constructed examples, and randomly sample one of the instructions for each question-answer pair; Construct the answer.
7. The time-order enhanced video question-answering method according to claim 1, characterized in that: Construct instruction video question answering task; including: For video question-answer pairs that do not contain temporal content, a video is randomly extracted from the dataset and randomly spliced to the front or back of the original video; Construct instructions, use ChatGPT to generate 10 instructions based on manually constructed examples, and randomly sample one of the instructions for each question-answer pair; Directly use the answers from the original video as the answers for the spliced video.
8. A time-order enhanced video question-answering method according to any one of claims 1 to 7, characterized in that: The three parts of data, namely, the video question-answering pair data containing the time sequence content, are mixed with the constructed multi-dimensional time sequence instruction data, the constructed multi-task time sequence instruction data, and the remaining video instruction fine-tuning data, to fine-tune the video question-answering model; include: Input the instruction text into the encoding layer to obtain the text semantic features ; Input the video frame sequence into the visual encoder to obtain visual features , and use the visual adapter to map the visual features into the text semantic space to obtain the visual semantic features aligned with the text semantic features ; Semantic features of text and visual semantic features After concatenation, the pre-trained large-scale language model is input and the model output is supervised and fine-tuned using cross entropy loss; For the five time series dimensions, additional evaluation datasets are collected and constructed to evaluate the time series question answering capabilities of the video question answering model. These include: Construct a multi-dimensional time series question answering evaluation dataset; including: Use similar steps as above to construct the evaluation dataset; Use three open source pre-trained video question-answering models, Qwen2-VL, LLaVA-OneVision, and MiniCPM-V2.6, as filters, input randomly sampled single-frame video content and constructed question-answer pairs, and judge whether the filters can correctly answer the questions based on the answers output by the filters. If at least two filters can correctly answer the questions, it is determined that the video question-answer pair has a single-frame deviation; Remove question-answer pairs with single-frame deviations from the dataset and rebalance the number of different answer candidates through sampling to obtain the final evaluation dataset; Input the video and question-answer pairs in the constructed evaluation dataset into the video question-answering model and obtain the answers. Compare whether the answers output by the video question-answering model are the same as the real answers. Count the proportion of questions correctly answered by the video question-answering model to the total number of questions.
9. A time-order enhanced video question-answering system, used to implement a time-order enhanced video question-answering method according to any one of claims 1 to 8, characterized in that: include: The multi-dimensional time series instruction data construction unit is configured to: clarify and define five time series dimensions, establish data collection and screening processes, and ensure that data preparation for each time series dimension is complete; A multi-task timing instruction fine-tuning unit is configured to: construct multiple timing auxiliary tasks and use the collected data to fine-tune the video question answering model to enhance the timing question answering capability of the model; The multi-dimensional temporal question answering evaluation unit is configured to collect and construct additional evaluation datasets for five temporal dimensions to evaluate the temporal question answering capabilities of the video question answering model.
Citation Information
Patent Citations
Zero-sample video question-answering method for guiding hybrid experts based on time sequence information
CN117612049A
Question and answer method and system based on multi-modal self-adaptive retrieval type enhanced large model
CN117648429A