Video Understanding Large Model Optimization and Evaluation Method, System, Device and Storage Medium

By designing a new connector structure and building a global timing understanding benchmark dataset, the shortcomings of the video understanding big model in terms of global timing understanding capabilities are solved, and a more accurate evaluation of the fine-grained understanding ability of the video understanding big model is achieved.

CN119888581BActive Publication Date: 2025-06-24UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510349413.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-24
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Existing video understanding models have shortcomings in global timing understanding capabilities, and existing evaluation benchmarks focus on video-level and local event-level understanding, making it difficult to measure the model's fine-grained understanding of the entire video.

Method used

A new connector structure is designed, including a spatiotemporal downsampler, a local bidirectional Mamba structure and a linear layer to improve global timing understanding capabilities. At the same time, a semi-automated data generation pipeline and global timing understanding benchmark data set were built to better evaluate the fine-grained understanding ability of video understanding large models.

Benefits of technology

Through the new connector structure and global timing understanding of the benchmark dataset, the global timing understanding ability of video understanding large models is significantly improved, and more accurate evaluation methods are provided, making up for the shortcomings of existing benchmarks in fine-grained understanding evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888581B_ABST
    Figure CN119888581B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device, and storage medium for optimizing and evaluating a large video understanding model. They are still corresponding one-to-one solutions. In the solutions: A new connector structure is designed to enhance the global temporal understanding ability, which consists of a spatio-temporal downsampler, a local bidirectional Mamba structure, and a linear layer. The spatio-temporal downsampler can reduce the token storage overhead; at the same time, the local bidirectional Mamba structure, on the one hand, makes up for the problem of limited receptive field, and on the other hand, it can model both intra-frame features and inter-frame features; in addition, the training of this connector is low-cost, and a three-stage progressive training strategy is used to combat catastrophic forgetting; and, a semi-automated data generation pipeline is also constructed and global temporal understanding data is proposed based on this pipeline to fill the evaluation gap in the existing benchmark field in this ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video understanding large model optimization, and particularly relates to a method, system, device and storage medium for optimizing and evaluating a video understanding large model. Background Art

[0002] The video understanding large model is a key research direction in the current video understanding field. It mainly consists of three parts: a visual encoder, a connector, and a large language model. The core idea of this design is to map video frames to the language feature space to achieve extensive understanding by leveraging the capabilities of the large language model. After a video is input, the model first uniformly samples a fixed number of frames. Then, the visual encoder extracts the features of these frames. To adapt to different sampling strategies, the visual encoder is often an image encoder, that is, the visual branch of CLIP (Contrastive Language–Image Pre-training). Then, the connector maps all the visual features to the language feature space. In addition, the connector also plays the role of modeling the temporal relationship between frames. Common connector structures are divided into three types, namely Q-Former (Transformer structure, Transformer is a neural network based on the attention mechanism), multi-layer perceptron, and convolutional structure. Q-Former uses a small number of query tokens to capture the temporal information and visual information in the video. Its advantage is that the receptive field is global and the number of mapped tokens is small, while its disadvantage is that it cannot maintain the token order. The multi-layer perceptron, on the contrary, maintains the token order but incurs a serious token overhead. The convolutional structure is a compromise between the two. It usually forms a sandwich shape with 2 layers of two-dimensional convolution and 1 layer of three-dimensional convolution. In addition, Q-Former and the convolutional structure also require an additional linear layer to assist in feature mapping. The large language model directly adopts the Vicuna series or the Mistral series. Among them, Vicuna is an open-source large language model based on LLaMA (Large Language Model), and Mistral is an efficient open-source large language model developed by Mistral AI.

[0003] In order to evaluate the capabilities of the video understanding large model, benchmark datasets have also received strong attention. Existing datasets mainly evaluate general video understanding capabilities, that is, within a benchmark, different types of evaluations are included. Divided by the type of understanding, there are actions, objects, positions, scenes, quantities, attributes, postures, roles, cognitions, etc., and each category can be further divided into sub-categories. Divided by the video length, there are short videos, medium-length videos, and long videos. Divided by the video category, it can be divided into movies and TV shows, popular science, sports, art, life, etc. In terms of the evaluation form, existing datasets mainly use closed-ended questions or open-ended questions, which relatively conform to the capabilities of the large language model.

[0004] However, existing benchmarks focus on video-level understanding or local event-level understanding, and rarely pay attention to global event-level understanding. Global event-level understanding is defined as global temporal understanding, which can be specifically interpreted as the ability of the model to understand all key events in the video. This ability depends on temporal modeling and event recognition, is a fundamental ability, and is also the key to stimulating the general ability of the model. In addition, there is a problem of shortcut learning in closed-ended question answering, that is, the model can execute an exclusion strategy through language clues and common sense clues, which is difficult to truly measure the ability of the model. And open-ended question answering often relies on commercial large models, such as GPT (Generative Pre-trained Transformer), for evaluation, and its reliability is questionable.

[0005] In view of this, the present invention is specifically proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a method, system, device and storage medium for optimizing and evaluating a large video understanding model, which can improve the global temporal understanding ability of the large video understanding model, and at the same time construct an evaluation benchmark data for global temporal understanding, so as to better evaluate the fine-grained understanding ability of the large video understanding model for the entire video.

[0007] The purpose of the present invention is achieved through the following technical solutions:

[0008] A method for optimizing and evaluating a large video understanding model includes:

[0009] Construct a connector including a spatio-temporal downsampler, a local bidirectional Mamba structure and a linear layer connected in sequence, and form a large video understanding model with a visual encoder and a large language model; wherein, a residual structure is also adopted between the spatio-temporal downsampler and the local bidirectional Mamba structure, and Mamba is a selective state space model;

[0010] Perform three-stage progressive training on the large video understanding model; in the first stage, train the local bidirectional Mamba structure with video text pairs and image text pairs in descriptive form; in the second stage, train the entire connector with video text pairs in descriptive form; in the third stage, train the connector and the large language model with video text pairs in multiple-choice form and descriptive form;

[0011] Adopt a semi-automated data generation pipeline to construct a global temporal understanding data set;

[0012] Use the global temporal understanding data set to construct a specific instruction fine-tuning training set, perform fine-tuning training on the large video understanding model after three-stage progressive training, and then use the global temporal understanding data set to evaluate the performance of the fine-tuned large video understanding model.

[0013] A video understanding large model optimization and evaluation system, comprising:

[0014] A model construction unit, configured to construct a connector including a spatio-temporal downsampler, a local bidirectional Mamba structure, and a linear layer connected in sequence, and form a video understanding large model with a visual encoder and a large language model; wherein, a residual structure is also adopted between the spatio-temporal downsampler and the local bidirectional Mamba structure, and Mamba is a selective state space model;

[0015] A three-stage progressive training unit, configured to perform three-stage progressive training on the video understanding large model; in the first stage, the local bidirectional Mamba structure is trained using video text pairs and image text pairs in descriptive form; in the second stage, the entire connector is trained using video text pairs in descriptive form; in the third stage, the connector and the large language model are trained using video text pairs in multiple-choice form and descriptive form;

[0016] A data construction unit, configured to construct a global temporal understanding data set for evaluation and training using a semi-automated data generation pipeline;

[0017] A fine-tuning and evaluation unit, configured to construct a specific instruction fine-tuning training set using the global temporal understanding data set, perform fine-tuning training on the video understanding large model after three-stage progressive training, and then perform performance evaluation on the fine-tuned video understanding large model using the global temporal understanding data set.

[0018] A processing device, comprising: one or more processors; a memory for storing one or more programs;

[0019] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.

[0020] A readable storage medium storing a computer program, which implements the foregoing method when executed by a processor.

[0021] As can be seen from the technical solution provided by the present invention above, a new connector structure is designed to improve the global temporal understanding ability, which consists of a spatio-temporal downsampler, a local bidirectional Mamba (selective state space model) structure, and a linear layer. The spatio-temporal downsampler can reduce the token storage overhead; at the same time, the local bidirectional Mamba structure, on the one hand, makes up for the problem of limited receptive field, and on the other hand, it can model both intra-frame features and inter-frame features; in addition, the training of this connector is low-cost and uses a three-stage progressive training strategy to combat catastrophic forgetting. And, since the existing benchmark datasets for large video understanding models mainly focus on video-level understanding and local event-level understanding, video-level understanding is coarse-grained while local event-level understanding cannot well measure the model's fine-grained understanding ability of the entire video, which is a basic ability for the model. Therefore, the present invention constructs a semi-automated data generation pipeline and proposes a global temporal understanding benchmark based on this pipeline to make up for the evaluation gap in this ability in the existing benchmark field. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 Flowchart of an optimization and evaluation method for a large video understanding model provided by an embodiment of the present invention;

[0024] Figure 2 Schematic diagram of the overall framework of a large video understanding model provided by an embodiment of the present invention;

[0025] Figure 3 Schematic diagram of the semi-automated data generation pipeline provided by an embodiment of the present invention;

[0026] Figure 4 Schematic diagram of an example of a qualitative analysis result provided by an embodiment of the present invention;

[0027] Figure 5 Schematic diagram of another example of a qualitative analysis result provided by an embodiment of the present invention;

[0028] Figure 6 Schematic diagram of a large video understanding model optimization and evaluation system provided by an embodiment of the present invention;

[0029] Figure 7 Schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The following clearly and completely describes the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0031] First, the following explanations are given for the terms that may be used in this article:

[0032] Descriptions with semantic meanings such as "comprising", "including", "containing", "having" or other similar ones should be interpreted as non-exclusive inclusion. For example: including a certain technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) should be interpreted as not only including the clearly listed certain technical feature element, but also including other well-known technical feature elements in the art that are not clearly listed.

[0033] The term "consisting of..." means excluding any technical feature element that is not clearly listed. If this term is used in a claim, this term will make the claim a closed type, so that it does not include technical feature elements other than the clearly listed technical feature elements, except for related conventional impurities. If this term only appears in a certain clause of a claim, then it only limits the elements clearly listed in that clause, and the elements recorded in other clauses are not excluded from the overall claim.

[0034] The following provides a detailed description of a method, system, device, and storage medium for optimizing and evaluating a large video understanding model provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those of ordinary skill in the art. Conditions not specified in the embodiments of the present invention are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. Instruments not specified by the manufacturer in the embodiments of the present invention are all conventional products that can be obtained through commercial purchase.

[0035] Embodiment 1

[0036] The embodiment of the present invention provides a method for optimizing and evaluating a large video understanding model, as Figure 1 shown, which mainly includes the following steps:

[0037] Step 1: Construct a connector including a spatio-temporal downsampler, a local bidirectional Mamba (selective state space model) structure, and a linear layer connected in sequence, and form a large video understanding model with a visual encoder and a large language model.

[0038] Preferably, the spatio-temporal downsampler includes a three-part convolutional structure. The first and third parts are two-dimensional convolutions for spatial modeling within the frame, and the second part is a three-dimensional convolution for temporal modeling between frames.

[0039] In the embodiment of the present invention, two scanning methods are integrated in the local bidirectional Mamba structure, namely, local forward scanning and local reverse scanning within the frame.

[0040] Step 2: Perform three-stage progressive training on the video understanding large model.

[0041] In the embodiment of the present invention, in the first stage, the local bidirectional Mamba structure is trained using video text pairs and image text pairs in descriptive form; in the second stage, the entire connector is trained using video text pairs in descriptive form; in the third stage, the connector and the large language model are trained using video text pairs in multiple-choice form and descriptive form.

[0042] In addition, a residual structure is also adopted between the spatio-temporal downsampler and the local bidirectional Mamba structure to further combat the forgetting problem.

[0043] Step 3: Use a semi-automated data generation pipeline to construct a global temporal understanding dataset.

[0044] Step 4: Use the global temporal understanding dataset to construct a specific instruction fine-tuning training set, and perform fine-tuning training on the video understanding large model after three-stage progressive training, and then use the global temporal understanding dataset to evaluate the performance of the fine-tuned video understanding large model.

[0045] It should be noted that the serial numbers of the above steps are only used to mark different steps and do not represent the execution order between the steps. The specific execution order of the steps can be determined by those skilled in the art according to the content of the steps. For example, Step 1 and Step 2 should be executed sequentially, and Step 4 should be executed based on Step 2 and Step 3. However, there is no strict logical order between Step 3 and Step 1 and Step 2. Therefore, the execution order of Step 3 and Step 1 and Step 2 can be determined according to the actual situation.

[0046] In the above solution provided by the embodiment of the present invention, the data generation pipeline can be used to implement semi-automated benchmark production. In an industrial scenario, it can generate a dataset at low cost, and the global temporal understanding dataset can be applied to the evaluation of the capabilities of the video understanding large model to guide model optimization. The connector proposed by the present invention can be applied to a general video understanding large model, and this video understanding large model can be deployed on a server to complete video understanding tasks. In addition, the progressive training strategy proposed by the present invention can be used in the continued training stage to effectively alleviate the overhead problem and forgetting problem in industrial model training.

[0047] To more clearly present the technical solutions provided by the present invention and the resulting technical effects, the methods provided by the embodiments of the present invention will be described in detail below with specific examples.

[0048] 1. Model Structure Solution.

[0049] Generally speaking, the large video understanding model includes: a visual encoder, a connector, and a large language model. To further improve the global temporal understanding ability of the large video understanding model, a new connector structure is proposed in this paper, as Figure 2 shown in the middle part. This connector consists of three parts, namely a spatio-temporal downsampler, a local bidirectional Mamba, and a linear layer.

[0050] The spatio-temporal downsampler includes three (e.g., three layers) convolutional structures arranged in sequence. The first and third layers are two-dimensional convolutions for spatial modeling within the frame, and the second layer is a three-dimensional convolution for temporal modeling between frames. Each convolution can reduce the number of tokens to alleviate the memory overhead.

[0051] The local bidirectional Mamba is based on the Mamba structure and integrates two scanning methods, one is local forward scanning, and the other is local reverse scanning. These two scanning methods are for within the frame; between frames, they are still unidirectional. These two scanning methods effectively maintain the token order while modeling image features. As Figure 2 shown, the local bidirectional Mamba structure includes: three linear units (responsible for feature mapping), a local forward scanning module, a local reverse scanning module, and an activation layer; the input of the local bidirectional Mamba structure enters the first linear unit D1 and the second linear unit D2. The output of the first linear unit D1 is divided into two paths after passing through the activation layer; the output of the second linear unit D2 is respectively subjected to the local forward scanning module and the local reverse scanning module for global temporal modeling. The outputs of the two scanning methods are separately combined (i.e., combined in a multiplicative manner, using the symbol to represent) with one path of the activation layer, and then after fusion, enter the third linear unit D3. The output of the third linear unit D3 enters the scanning module at the tail to output the scanning mode; for the above two scanning modules, each adjusts the token order according to a specific scanning method, and then sequentially uses one-dimensional convolution, non-linear activation, and a state space model to capture all temporal features. After adjustment, the features obtained by the two scanning methods are fused in an average manner (using the symbol to represent); Figure 2 The dashed box on the right shows the internal structural principle of the two scanning modules, and the symbol As the activation function, the same activation function as that of the activation layer can be used. For example, the SiLU (Sigmoid Linear Unit) function can be selected. The following dashed box shows the meaning of adjusting the token order in two scanning modes. Here, the token is a feature unit, which is a general term in the art and will not be elaborated in the present invention.

[0052] The linear layer is used to further map the features to the language feature space.

[0053] Considering that both the visual encoder and the large language model can be implemented by existing solutions, they will not be elaborated.

[0054] III. Model training scheme.

[0055] Since there are available pre-trained parameters for modules other than Mamba, it is feasible to train Mamba to unify the model. To combat the catastrophic forgetting problem existing in the process of embedding Mamba, the present invention adopts a three-stage progressive training strategy. In the first stage, the present invention uses video text pairs and image text pairs in descriptive form to train only the local bidirectional Mamba to achieve feature alignment. In the second stage, the present invention uses video text pairs in descriptive form to train the entire connector simultaneously to achieve secondary feature alignment. In the third stage, the present invention uses video text pairs in multiple-choice form and descriptive form to train the connector and the language model to achieve overall fusion. In addition, a residual structure is adopted between the spatio-temporal downsampler and the local bidirectional Mamba structure to further combat the forgetting problem. Figure 3 α in is the proportionality coefficient of the output of the local bidirectional Mamba structure during residual connection. This residual structure uses trainable parameters to adjust the fusion degree between features.

[0056] The video text pairs and image text pairs in descriptive form, and the video texts in multiple-choice form and descriptive form involved in this part of training are all existing data forms and can be collected by conventional methods, which will not be elaborated in the present invention.

[0057] III. Global temporal understanding dataset construction and evaluation scheme.

[0058] As mentioned before, since the existing benchmark datasets for video understanding large models mainly focus on video-level understanding and local event-level understanding, video-level understanding is coarse-grained and local event-level understanding cannot well measure the model's fine-grained understanding ability of the entire video, which is a basic ability for the model. Therefore, the present invention constructs a semi-automated data generation pipeline and proposes a global temporal understanding benchmark based on this pipeline to fill the evaluation gap in this ability in the existing benchmark field.

[0059] Specifically, the preferred implementation manner of the present invention for constructing the global temporal understanding dataset using a semi-automated data generation pipeline is as follows:

[0060] 1. Construct a semi - automated data generation pipeline that includes a temporal clue elimination pipeline and a hallucination generation pipeline, as Figure 3 shown, which shows an example of the overall framework of the semi - automated data generation pipeline.

[0061] 2. Collect video data with video event sequences to form an original data set and input it into the semi - automated data generation pipeline. Each sample in the original data set contains a video and a video event sequence; the video event sequence records the event content in the corresponding video. For example, the event content of a certain video is ["cooking", "eating", "washing dishes"]. The above - recorded event content is the label of the video, that is, the video event sequence, and each element in the video event sequence is a real event.

[0062] Exemplarily, the training set and test set (if not available, the validation set is selected) can be integrated from five data sets with video event sequence labels, namely ActivityNet, ViTT, YouCook2, COIN, and VidChapters, to form the original data set.

[0063] 3. The temporal clue elimination pipeline performs semantic clue elimination on each video event sequence in the original data set, and then filters out the video event sequences containing common sense clues, and outputs video event sequences that do not contain common sense clues and have undergone semantic clue elimination.

[0064] Preferably, the temporal clue elimination pipeline includes two parts: semantic clue elimination and common sense clue elimination.

[0065] (3.1) The semantic clue elimination part mainly focuses on language clues with potential temporal logic, such as target reference, temporal words, etc. The implementation process of this part includes: guiding an external large - language model through prompt words for phrase replacement and deletion to eliminate language clues.

[0066] (3.2) The common sense clue elimination part mainly focuses on common event sequences. The implementation process of this part includes: first shuffling the video event sequences after semantic clue elimination, and then guiding an external large - language model through prompt words to sort the shuffled video event sequences; if the sorting is correct, it indicates that the corresponding video event sequence is a common sense sequence, and the corresponding video event sequence is excluded. To ensure thoroughness, this process loops until no more samples are excluded from the data set.

[0067] 4. The video event sequences output by the temporal clue elimination pipeline enter the hallucination generation pipeline to generate hallucination events.

[0068] Preferably, the hallucination generation pipeline includes three parts: hallucination generation, structural clue elimination, and theme consistency correction.

[0069] (4.1) Hallucination generation part, which includes four steps. The first two steps refer to identifying the video event sequence output by the temporal clue elimination pipeline to obtain the corresponding theme and character; then combining the theme and character to generate hallucination events that can be inserted into the corresponding video event sequence. This process covers the latter two steps, that is, the previous step generates relevant hallucination events, and the latter step requires that the generated hallucinations can be inserted into the event sequence to further constrain the correlation between the hallucinations and the original event sequence.

[0070] (4.2) Structural clue elimination part, which determines whether the sentence pattern of the hallucination events generated by the hallucination generation part is consistent with the sentence pattern of the corresponding video event sequence output by the temporal clue elimination pipeline. If not, it is modified.

[0071] (4.3) Theme consistency correction part, which determines whether the theme of the hallucination events generated by the hallucination generation part is consistent with the corresponding video event sequence. If not, it is modified.

[0072] After the sequential processing of the above three parts, the hallucination events output by the hallucination generation part are obtained.

[0073] 5. Integrate and shuffle the video event sequence output by the temporal clue elimination pipeline and the hallucination events generated by the hallucination generation pipeline to obtain a rearranged dataset containing hallucination events, and then construct a global temporal understanding dataset.

[0074] In the embodiment of the present invention, the global temporal understanding dataset includes the above three parts of data, namely: (1) the video event sequence output by the temporal clue elimination pipeline (shuffled to form a temporal dataset); (2) the hallucination events output by the hallucination generation part (inserted into the video event sequence to form a hallucination dataset); (3) the rearranged dataset containing hallucination events (formed by inserting hallucination events into the video sequence and shuffling). The Q&A form of the rearranged dataset containing hallucinations and the temporal dataset is sorting, for example: 1→2→3, while the Q&A form of the hallucination dataset is enumeration, for example: 1&3 (1 and 3).

[0075] In this step, a global temporal understanding dataset is constructed. Each test sample in it includes a video, the processed video event sequence, and the corresponding instruction. The processed video event sequence is any one of the above three parts of data. The instruction requires the video understanding large model to be evaluated to generate a specific form of answer according to the video and the event sequence. For example, the answer form is "1→2→3" or "1&3". This is a semi-open answer form that can achieve a compromise between alleviating shortcut learning and evaluation convenience.

[0076] It should be noted that for the sake of understanding, Figure 3Examples of video event sequences, hallucination events, and integrated and scrambled event sequences are provided. In practical applications, the specific content of various events needs to be determined according to the actual situation, and the present invention does not make any limitations; at the same time, the present invention is applicable to multiple types of languages. Figure 2 The provided text is in English. In practical applications, users can also adjust the language form of the input text according to actual needs.

[0077] The above global temporal understanding dataset can be used to evaluate temporal modeling and event recognition capabilities using a temporal dataset and a hallucination dataset respectively. For the temporal dataset, the instruction requires the model to sort the scrambled event sequences. For the hallucination dataset, the instruction requires selecting hallucinations from the event sequences doped with hallucinations. The event rearrangement dataset containing hallucinations can evaluate both of these capabilities, and the corresponding instruction requires the model to sort the actually existing events.

[0078] In addition, since existing models cannot generate answers in a specific form based on sequence descriptions and instructions, small-scale instruction fine-tuning is necessary. To avoid passing preference knowledge to the model during training, the present invention combines three rules to generate a specific instruction fine-tuning training set. Each data in the specific instruction fine-tuning training set is a question-answer pair, including a training sample and the corresponding answer. The training sample includes a video, a processed and rule-constrained video event sequence, and the corresponding instruction. The processed and rule-constrained video event sequence refers to the video event sequence processed by using set rules. Similarly, the instruction is used to guide the video understanding large model to generate an answer in a specified form; the video event sequence processed by using set rules includes: 1) each video event sequence contains at least two real events (i.e., the events in the corresponding video event sequence in the original dataset); 2) a set proportion (such as 50%) of the question-answer pairs contain hallucination events; 3) when hallucination events are included, the positions of the hallucination events are random. It should be additionally noted that in the training set, the sorting dataset shares the processing method of "shuffling after adding hallucinations".

[0079] During the implementation process, since the insertion of hallucination events and the scrambling of event sequences can be obtained through code, the insertion positions and scrambling methods can be pre-recorded. That is to say, the answers are known, so question-answer pairs can be formed corresponding to the event sequences, which are used to supervise the output of the model during training, and the loss function is calculated accordingly to fine-tune the model. The involved fine-tuning training scheme can be implemented with reference to conventional techniques, and the present invention will not elaborate.

[0080] During the evaluation process, an input test sample and an output answer form a question-answer pair. The question-answer pairs generated from the event rearrangement dataset containing hallucinations and the temporal dataset are called sorting question-answer pairs, and the question-answer pairs generated from the hallucination dataset are called hallucination question-answer pairs.

[0081] For sorted question-answer pairs, the edit distance can be selected as the evaluation metric. For hallucination question-answer pairs, the IOU (Intersection over Union) can be selected as the evaluation metric, and the formula is:

[0082] ;

[0083] where P is the set of predicted sequence numbers, G is the set of ground-truth sequence numbers, is the symbol for calculating the number of elements in the set.

[0084] IV. Performance description of the solution of the present invention.

[0085] The existing benchmark datasets for video understanding large models mainly focus on video-level understanding and local event-level understanding. Video-level understanding is coarse-grained, and local event-level understanding cannot well measure the model's fine-grained understanding ability of the entire video, which is a basic ability for the model. Therefore, the present invention first constructs a semi-automated data generation pipeline and proposes a global temporal understanding benchmark based on this pipeline to fill the evaluation gap in this ability in the existing benchmark field.

[0086] To improve the model's performance in global temporal understanding, a new connector structure is proposed in this paper, which consists of a spatio-temporal downsampler, a local bidirectional Mamba, and a linear layer. The spatio-temporal downsampler can reduce the token storage overhead. At the same time, the local bidirectional Mamba on the one hand makes up for the problem of limited receptive field, and on the other hand, it can model both intra-frame features and inter-frame features. In addition, the training of this connector is low-cost. The present invention proposes a three-stage progressive training strategy based on existing pre-trained parameters to combat catastrophic forgetting.

[0087] The Global Temporal Understanding Benchmark evaluates the performance of existing mainstream large video understanding models, including: VideoLLaMA2, VideoChat2, VideoLLaMA, MovieChat; among them, VideoLLaMA is an audio-visual language model designed to endow large language models with the ability to understand videos and audio; VideoLLaMA2 is a multimodal video large language model designed to enhance video understanding and audio understanding capabilities; VideoChat2 is a multimodal video understanding model designed to achieve in-depth understanding and dialogue interaction of video content by combining a video foundation model and a large language model; MovieChat is an innovative long video understanding framework that combines a visual model and a large language model and is designed to efficiently process long video content of more than 10,000 frames; considering that the above models are all existing models, they will not be elaborated. During the evaluation process, the edit distance is used as the evaluation index. The evaluation result of VideoLLaMA2 is 30.70%, the evaluation result of VideoChat2 is 42.53%, the evaluation result of VideoLLaMA is 26.34%, and the evaluation result of MovieChat is 24.89%. Through the above evaluation, it shows that the existing large video understanding models do not have strong global temporal understanding capabilities, which provides more guidance for future research. The baseline model with the connector structure proposed in the present invention achieved 55.98% on this benchmark, which is a 25.29% improvement compared to the baseline model VideoLLaMA2. This shows that this structure can well improve the global understanding ability. In addition, this connector has generalization ability. The optimized model achieved 56.82% and 45.00% on the general video understanding benchmarks MVBench and VideoMME respectively, which are 3.42% and 1.00% improvements compared to the baseline model. Figure 4 and Figure 5 shows some qualitative analysis results. Each example contains a video, a question, the answers of VideoChat2, VideoLLaMA2, and the optimized model of the present invention. The example also contains the results scored by GPT4o-mini; among them, GPT4o-mini is a mini version of GPT-4o, and GPT-4o is an all-round GPT-4 (fourth-generation generative pre-trained transformer) model; similarly, Figure 4 and Figure 5 also uses text examples in English form. The present invention can answer more video content and more accurately compared to VideoChat2 and VideoLLaMA2. Among them, the solid line marked part is the text that hits the video content, and the dotted line marked part is the hallucinated text. This reflects that the optimization scheme of the present invention is beneficial to global temporal understanding.

[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software or by means of software plus a necessary general hardware platform. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0089] Embodiment 2

[0090] The present invention also provides a video understanding large model optimization system, which is mainly used to implement the methods provided in the foregoing embodiments, such as Figure 6 shown, the system mainly includes:

[0091] A model construction unit, configured to construct a connector including a spatio-temporal downsampler, a local bidirectional Mamba structure, and a linear layer connected in sequence, and form a video understanding large model with a visual encoder and a large language model; wherein, a residual structure is also adopted between the spatio-temporal downsampler and the local bidirectional Mamba structure, and Mamba is a selective state space model;

[0092] A three-stage progressive training unit, configured to perform three-stage progressive training on the video understanding large model; in the first stage, the local bidirectional Mamba structure is trained with video text pairs and image text pairs in descriptive form; in the second stage, the entire connector is trained with video text pairs in descriptive form; in the third stage, the connector and the large language model are trained with video text pairs in multiple-choice form and descriptive form;

[0093] A data construction unit, configured to construct a global temporal understanding data set for evaluation and training by using a semi-automated data generation pipeline;

[0094] A fine-tuning and evaluation unit, configured to construct a specific instruction fine-tuning training set by using the global temporal understanding data set, perform fine-tuning training on the video understanding large model after three-stage progressive training, and then perform performance evaluation on the video understanding large model after fine-tuning training by using the global temporal understanding data set.

[0095] Considering that the main technical details involved in this system have been introduced in detail in the previous embodiments, they will not be elaborated here.

[0096] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional modules as needed, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0097] Embodiment III

[0098] The present invention also provides a processing device, such as Figure 7 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.

[0099] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.

[0100] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:

[0101] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;

[0102] The output device can be a display terminal;

[0103] The memory can be a Random Access Memory (RAM), or a non-volatile memory, such as a disk memory.

[0104] Embodiment IV

[0105] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiments when the computer program is executed by a processor.

[0106] In the embodiments of the present invention, the readable storage medium as a computer-readable storage medium can be disposed in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a Read-Only Memory (ROM), a magnetic disk, or an optical disc.

[0107] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.

Claims

1. A video understanding large model optimization and evaluation method, characterized in that: include: Construct a connector including a spatiotemporal downsampler, a local bidirectional Mamba structure and a linear layer connected in sequence, and form a large video understanding model with a visual encoder and a large language model; wherein a residual structure is also used between the spatiotemporal downsampler and the local bidirectional Mamba structure, and Mamba is a selective state space model; The video understanding large model is trained progressively in three stages; in the first stage, the local bidirectional Mamba structure is trained using video-text pairs and image-text pairs in the form of descriptions; in the second stage, the entire connector is trained using video-text pairs in the form of descriptions; in the third stage, the connector and the large language model are trained using video-text pairs in the form of multiple choices and descriptions; A global temporal understanding dataset is constructed by using a semi-automatic data generation pipeline, including: constructing a semi-automatic data generation pipeline including a temporal clue elimination pipeline and an illusion generation pipeline; collecting video data with video event sequences to form an original dataset, and inputting the original dataset into the semi-automatic data generation pipeline, wherein each sample in the original dataset includes a video and a video event sequence; the temporal clue elimination pipeline performs semantic clue elimination on each video event sequence in the original dataset, and then filters out the video event sequence containing common sense clues, and the output video event sequence is a video event sequence that does not contain common sense clues and has undergone semantic clue elimination; the video event sequence output by the temporal clue elimination pipeline enters the illusion generation pipeline to generate an illusion event; the video event sequence output by the temporal clue elimination pipeline and the illusion event generated by the illusion generation pipeline are integrated and disrupted to construct a global temporal understanding dataset, i.e., a rearranged dataset containing illusion events; A specific instruction fine-tuning training set is constructed using a global temporal understanding dataset, and the large video understanding model after three-stage progressive training is fine-tuned. The global temporal understanding dataset is then used to evaluate the performance of the fine-tuned large video understanding model.

2. A video understanding large model optimization and evaluation method according to claim 1, characterized in that: The spatiotemporal downsampler includes a three-part convolution structure, wherein the first and third parts are two-dimensional convolutions for spatial modeling within a frame; and the second part is a three-dimensional convolution for temporal modeling between frames.

3. A video understanding large model optimization and evaluation method according to claim 1, characterized in that: The local bidirectional Mamba structure integrates two scanning modes, namely, local forward scanning and local reverse scanning within a frame; the local bidirectional Mamba structure includes: three linear units, a local forward scanning module, a local reverse scanning module and an activation layer; the input of the local bidirectional Mamba structure enters the first linear unit and the second linear unit, and the output of the first linear unit is divided into two paths through the activation layer; the output of the second linear unit is respectively subjected to global timing modeling by the local forward scanning module and the local reverse scanning module, and the outputs of the two scanning modes are separately combined with one path of the activation layer, and then fused to enter the third linear unit, and the output of the third linear unit enters the tail scanning module to output the scanning mode.

4. A video understanding large model optimization and evaluation method according to claim 1, characterized in that: The temporal clue elimination pipeline includes two parts: semantic clue elimination and common sense clue elimination; The semantic clue elimination part uses prompt words to guide the external large language model to replace and delete phrases and eliminate language clues; In the common sense clue elimination part, the video event sequence after the semantic clues are eliminated is first shuffled, and the external large language model is guided by prompt words to sort the shuffled video event sequence; if the sorting is correct, it indicates that the corresponding video event sequence is a common sense sequence, and the corresponding video event sequence is eliminated.

5. The method for optimizing and evaluating a large video understanding model according to claim 1, characterized in that: The hallucination generation pipeline includes three parts: hallucination generation, structural clue elimination and subject consistency correction; The hallucination generation part identifies the video event sequence output by the timing clue elimination pipeline, obtains the corresponding theme and role, and then generates a hallucination event that can be inserted into the corresponding video event sequence by combining the theme and role; The structural clue elimination part determines whether the sentence structure of the hallucination event generated by the hallucination generation part is consistent with the sentence structure of the corresponding video event sequence output by the timing clue elimination pipeline, and if not, makes a modification; The theme consistency correction part determines whether the theme of the hallucination event generated by the hallucination generation part is consistent with the corresponding video event sequence. If not, it is modified.

6. A video understanding large model optimization and evaluation method according to claim 1, characterized in that: Each data in the specific instruction fine-tuning training set is a question-answer pair, including a training sample and a corresponding answer, the training sample includes a video, a processed and rule-constrained video event sequence, and a corresponding instruction, the processed and rule-constrained video event sequence refers to a video event sequence processed using set rule constraints, the processed video event sequence is data in the global temporal understanding dataset, and the instruction is used to guide the video understanding large model to generate a specified form of answer; The video event sequence processed by the set rule constraints includes: Each video event sequence contains at least two real events, where the real events refer to events in the corresponding video event sequence in the original data set; A set proportion of question-answer pairs included hallucination events; Also, when hallucination events are included, the locations of the hallucination events are random.

7. A video understanding large model optimization and evaluation system, characterized in that: The method for implementing any one of claims 1 to 6 comprises: A model construction unit is used to construct a connector including a spatiotemporal downsampler, a local bidirectional Mamba structure and a linear layer connected in sequence, and to form a large video understanding model with a visual encoder and a large language model; wherein a residual structure is also used between the spatiotemporal downsampler and the local bidirectional Mamba structure, and Mamba is a selective state space model; A three-stage progressive training unit is used to perform three-stage progressive training on the video understanding large model; in the first stage, the local bidirectional Mamba structure is trained using the video text pairs and image text pairs in the form of description; in the second stage, the entire connector is trained using the video text pairs in the form of description; in the third stage, the connector and the large language model are trained using the video text pairs in the form of multiple choice and description; A data construction unit, which uses a semi-automated data generation pipeline to build a global time series understanding dataset for evaluation and training; The fine-tuning and evaluation unit is used to construct a specific instruction fine-tuning training set using a global temporal understanding dataset, and to perform fine-tuning training on the video understanding large model after three-stage progressive training, and then use the global temporal understanding dataset to perform performance evaluation on the fine-tuned video understanding large model.

8. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.