A Dynamic Iterative Long Video Understanding Method Based on Large Language Models
Through the dynamic iterative long video understanding method, combined with a large language model and a dual Agent system, the problems of high computing complexity and redundant information interference in long video understanding are solved, and efficient and accurate video content understanding and answer generation are achieved.
Patent Information
- Application Number
- CN202510355760.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The prior art has high computational complexity in long video comprehension, making it difficult to dynamically adapt to user needs, and large language model (LLM) is difficult to efficiently and accurately understand due to input length limitations and redundant information interference.
Using a dynamic iterative method based on a large language model, through a self-supervised iterative thinking chain and dual Agent system, combining video frame sampling, text description generation, visual feature extraction and cognitive adaptability evaluation, we simulate human progressive cognitive thinking and optimize the video understanding process.
With limited computing resources, efficient and accurate understanding of long videos is achieved, redundant information processing is reduced, answer accuracy is improved, computing resource consumption is reduced, and multimodal tasks are adapted to multimodal tasks.
Smart Images

Figure CN119863745B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and computer vision, and particularly to a dynamic iterative long video understanding method based on a large language model. Background Art
[0002] Traditional video understanding methods usually rely on dense sampling and global analysis. For example, video frames are evenly extracted and features are extracted using a convolutional neural network (CNN) or a time series model (such as LSTM). Such methods need to process a large amount of frame data, have a high computational complexity, and are difficult to capture the complex semantic logic of long videos. In addition, traditional methods rely on preset rules or fixed templates, lack flexibility, and cannot dynamically adapt to the diverse needs of users (such as open-domain video question answering).
[0003] With the rise of large models, methods based on large language models (LLMs) have significant advantages in semantic understanding and dynamic reasoning. Visual information can be converted into text descriptions through multimodal fusion, and natural language answers can be generated. However, when LLMs are directly applied to long video understanding, there are two major bottlenecks: one is that the input text length is limited and it is difficult to carry all the key information of long videos; the other is that the mixing of redundant information and key segments in long videos easily leads to model reasoning deviation, and continuous retrieval to supplement information will cause a sharp increase in computational resource consumption.
[0004] The deficiencies of the prior art can be summarized as follows: traditional methods are inefficient and lack dynamics, while existing LLM solutions are difficult to balance the efficiency and accuracy of long video understanding due to length limitations and redundant interference. Therefore, there is an urgent need for a method that simulates the progressive cognitive thinking of humans to achieve efficient and accurate understanding of long videos with limited computational resources. Summary of the Invention
[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a dynamic iterative long video understanding method based on a large language model in view of the deficiencies of the prior art. Video understanding technology has a wide range of application scenarios in real life, especially in fields that require quickly extracting key information from a large amount of video data. By combining the powerful language understanding ability and rich prior knowledge of a large language model (LLM), the present invention can efficiently and accurately understand the content of long videos and generate answers to user questions.
[0006] The method includes the following steps:
[0007] Step 1, based on a self-supervised dynamic iterative thought chain, perform mathematical modeling and analysis on the video understanding task;
[0008] Step 2: Preprocess the video input by the user. Perform frame sampling on the video, generate text descriptions for each video frame, and extract the visual features of the video frames. Conduct preliminary reasoning through the Q&A Agent (Agent is an abstract concept. It is usually defined as an entity that can perceive the environment and take actions in the environment to achieve goals), generate preliminary answers by combining the input text and video frames. Use the cognitive adaptability evaluation mechanism to evaluate the answers. If the cognitive adaptability meets the requirements, output the answers; otherwise, proceed to Step 3.
[0009] Step 3: Employ a dual-Agent system for self-supervised information feedback. At each step of the reasoning process, introduce a judgment Agent to recognize the answers. The judgment Agent assists in confirming whether the answers are ambiguous or require further detailed supplementation by detecting the consistency and accuracy of the answers. The Q&A Agent will retrieve key frames and supplement information based on the feedback results, iteratively update the known information, and then return to Step 2 for the next round of reasoning until the cognitive adaptability reaches the preset standard, thereby ensuring the accuracy and consistency of the answers.
[0010] Step 4: Conduct result evaluation. Use the Q&A accuracy rate and the average number of retrieved frames as evaluation metrics to perform quantitative analysis on a recognized open-source dataset to verify the effectiveness of the method. Secondly, perform qualitative analysis using any video and question provided by the user to verify whether the results meet expectations.
[0011] In Step 1, the mathematical modeling and analysis of the video understanding task include: modeling the video understanding process as a Markov process, where the state at each moment corresponds to a stage in the reasoning process, and the state at time t includes the visual features of the video frames and the text descriptions , expressed as:
[0012] ,
[0013] In each round of reasoning, the answer for the next moment will be generated based on the current state , and then continue to search for key information through the existing answers. Let represent the probability of generating the answer for the next moment under the condition of the state , and represent the state at time t + 1. Then the state transition process for the next moment is expressed as:
[0014] ,
[0015] ,
[0016] where where Denote the state transition strategy;
[0017] Use D to denote the evaluation strategy, and use to denote the feedback result after self-evaluation at time t + 1. denote all historical answers from time 0 to time t. The process of self-supervised correction and learning is expressed as:
[0018] ,
[0019] The state transition process is further expressed as:
[0020] ,
[0021] The entire video understanding process is decomposed into question and answer, self-evaluation, feedback and correction, continuously iterating and updating the state. The ultimate goal is to maximize the expected expectation C in the process of dynamically understanding the video. The ultimate optimization goal of the entire thinking chain is:
[0022] ,
[0023] denote finding the variable that maximizes the expected expectation C and .
[0024] Step 2 includes:
[0025] Step 2-1, perform video frame sampling: Uniformly sample the user-input video at a rate of 1 frame per second (fps) and convert it into an image sequence. The frame rate of the video is usually high, such as 30fps or 60fps, which will result in a large amount of image data per second. If all frames are retained during video processing, it will require huge computing and storage resources. Especially in the processing tasks of long videos or large-scale video datasets, the computing cost may be very high. By sampling at a rate of 1 fps, the number of frames to be processed per second is reduced, thus significantly reducing the demand for computing resources while maintaining sufficient scene information. The sampling frequency of 1fps can already capture the basic dynamics and scene changes of the video and avoid unnecessary redundant data;
[0026] Step 2-2, generate text descriptions: Use the existing Bootstrapping Language-Image Pre-training (BLIP) model to generate text descriptions (captions) for each frame of the image. The content of the text description includes the scene, objects, and actions. This text description provides highly structured information for the subsequent reasoning process and can effectively help understand the context and dynamic content of the video;
[0027] Step 2-3, Extract and store image features: Extract visual features from each frame of the image through the visual encoder of the pre-trained model BLIP for guided language images. The visual features include information such as object recognition, scene understanding, and spatial relationships in the image;
[0028] The visual features are saved in the.npy format, providing basic data for fast access in subsequent inference and retrieval. This data storage format aims to optimize processing speed and storage efficiency, supporting large-scale video analysis tasks;
[0029] Step 2-4, Perform Agent inference and cognitive adaptability evaluation.
[0030] Step 2-4 includes: Select 5 initial frames at uniform distribution in the image sequence obtained in Step 2-1. The state at the initial moment is represented as where respectively represent the visual features and text descriptions of the initial frames. From the perspective of the development of events, its dynamic changes can be modeled by the process of "beginning, development, climax, then development, and ending". Based on the time series characteristics of events, considering the changes in events at different stages, the selection of using 5 frames for sampling can be combined with the characteristics of event development to ensure that the initially sampled frames can cover different stages of the entire event development;
[0031] Design a question-and-answer Agent as the main body for inference and generating answers. The question-and-answer Agent includes a language understanding module and an answer generation module;
[0032] The language understanding module uses GPT-4o to provide prior knowledge for the question-and-answer Agent, providing it with strong language understanding capabilities;
[0033] The answer generation module is used to calculate the matching degree between the answer and the known information and generate answers;
[0034] Concatenate the user's question and the corresponding text description into a prompt, and input it into the question-and-answer Agent. The question-and-answer Agent generates options for the preliminary answer according to the prompt and outputs the probability of selecting each option. Use to represent the i-th option, to represent the probability of selecting the i-th option under the condition of knowing . The calculation formula is:
[0035] ,
[0036] where , n represents the total number of options, and , Sim represents the cosine similarity function, exp represents the natural exponential function, and this process corresponds to the modeling process in step 1 , here the generation probability formula of the answer is specified, and according to the current state the probability distribution of each option is generated, and then the answer at the next moment is determined through probability ;
[0037] Use LLM to represent the Q&A Agent. The process of the Q&A Agent generating an answer through the option probability distribution is as follows:
[0038] ,
[0039] Design the cognitive adaptability degree, and use the cognitive adaptability degree as the expected expectation C proposed in step 1. The evaluation of the cognitive adaptability degree includes relative cognitive degree and absolute cognitive degree ;
[0040] Calculate the relative cognitive degree , which is used to reflect the discrimination and difference between options:
[0041] ,
[0042] Calculate the absolute cognitive degree , which is used to measure the certainty and sufficiency of information:
[0043] ,
[0044] Calculate C:
[0045] ,
[0046] Among them, denote the one with the largest probability in as ;
[0047] The relative cognitive degree measures the perception of the differences between answers. The larger the relative cognitive degree, the stronger the difference between options, the clearer the cognition, and the higher the decision-making confidence; the absolute cognitive degree measures the certainty of the overall information based on information entropy. The smaller the information entropy, the larger the absolute cognitive degree, the clearer the cognition, and the higher the certainty of the current task. is a weight coefficient, indicating the importance of relative cognitive degree and absolute cognitive degree in the overall evaluation. Usually 0 < <1; The cognitive adaptability ultimately aggregates the evaluation results of option discrimination and information sufficiency. The cognitive adaptability emphasizes the concept of "cognition" and reflects the comprehensive consideration of information sufficiency and option discrimination in the decision-making process. This is not only a quantification of the model's confidence but also emphasizes the "adaptability" of the model in the face of complex decisions. Maximizing the cognitive adaptability will be used as the goal, corresponding to the expected expectation C in the modeling of step 1;
[0048] If the cognitive adaptability is greater than or equal to the preset threshold (0.6) or exceeds the maximum number of iterations (e.g., 5 times), the answer is directly output; otherwise, go to step 3.
[0049] Step 3 includes:
[0050] Step 3-1, design a judgment Agent for self-supervised feedback;
[0051] Step 3-2, perform dynamic iterative optimization.
[0052] Step 3 includes:
[0053] Step 3-1 includes: The judgment Agent includes an answer checking module and a search interval feedback module;
[0054] The answer checking module uses GPT-4-preview to provide prior knowledge to judge the consistency and correctness of the answer generated in step 2-4. This process corresponds to the modeling in step 1; First, for consistency, check whether the answer at the previous moment is consistent with the current answer ; Second, for correctness, check whether the current known information and the target question can infer the answer (here relying on the prior knowledge of GPT);
[0055] The search interval feedback module is used to predict the key frames in the known frames and record them as known key frames. Accordingly, it gives the feedback video search interval and provides a text description of the expected search and supplementary information: Give feedback according to the judgment result of the answer checking module. If it is consistent and correct, only take a specific number of frames Y1 before and after the current known key frame for key frame search. Here, the calculation formula for the specific number of frames Y1 is:
[0056] Y1 = Y2 / (Y3 - 1),
[0057] where Y2 represents the total number of video frames and Y3 represents the number of known frames;
[0058] Special case description: When the known key frame is the first frame, only search backward; when the known key frame is the last frame, only search forward. This process is similar to humans slowly dragging the video progress bar, the purpose is to refine and supplement the current known information. If it is inconsistent or incorrect, it is necessary to expand the search range to the entire video frame interval, similar to the behavior of humans dragging the video progress bar significantly. The dual-agent verification mechanism realizes self-supervisory feedback, which can greatly reduce the cost of manual labeling.
[0059] Step 3 includes:
[0060] Step 3-2 includes: according to the feedback video search interval and text description given by the judging agent in step 3-1, key frame retrieval and information supplement are performed to dynamically adjust the state , corresponding to the state transition process modeled in step 1 , so far, the formula of dynamic iterative optimization is rewritten as:
[0061] ;
[0062] Strategy The implementation process includes: based on the feedback video search interval and text description, the visual features of all frames in the feedback video search interval are retrieved from the .npy format file obtained in step 2-3, and the frames with the highest similarity between the visual features and the text description features in the feedback video search interval are selected as supplements, and the evaluation agent is set to give m feedback video search intervals. , the formula is:
[0063] ,
[0064] in Indicates the mth feedback video search interval The key frames retrieved from Indicates the mth feedback video search interval The visual features of the jth image in , It is a text description feature;
[0065] The visual features of the retrieved supplementary frame are integrated into the existing information to obtain updated information. At the same time, the description corresponding to the supplementary frame is also updated. Keyframes that represent additional The corresponding text description, where Indicates the mth key frame The corresponding text description, the specific update formula is:
[0066] ,
[0067] ,
[0068] Among them, "concat" represents the concatenation operation;
[0069] After the update is completed, return to steps 2 - 4 to enter the next round of reasoning and iteration, regenerate the answer and evaluate the cognitive adaptability until the output condition is met.
[0070] Step 4 includes:
[0071] Using the EgoSchema_subset dataset, the evaluation metrics include:
[0072] Average number of retrieved frames: This metric measures the average number of frames retrieved to generate an answer for each question when the method is tested on the EgoSchema_subset, and is used to evaluate the computational efficiency. This metric reflects the efficiency of the method in understanding and processing video content. The fewer the number of retrieved frames, the stronger the method's ability to obtain sufficient information in a short time;
[0073] Question - answering accuracy : Used to determine whether the answer output by the model is the same as the annotated answer. The formula is:
[0074] ,
[0075] Among them, TP (True Positive) represents the number of answers generated that exactly match the annotated answers, and FP (False Positive) represents the number of answers generated that do not match the annotated answers. The higher the question - answering accuracy, the stronger the method's ability to generate correct answers;
[0076] In addition, to further verify the effectiveness of the method, qualitative analysis is also carried out using any video and questions provided by users. By comparing the output answers with the users' expected results, it is judged whether the performance of the method in the actual application scenario meets the expectations.
[0077] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the above - mentioned method.
[0078] The present invention also provides a storage medium storing a computer program or instructions. When the computer program or instructions run on a computer, the steps of the above - mentioned method are executed.
[0079] The method of the present invention can be applied to the following fields:
[0080] Video Q&A System: According to the questions raised by users, quickly locate and extract relevant information from long videos to generate accurate answers. For example, in the field of education, students can obtain explanations of specific knowledge points from teaching videos by asking questions; in the medical field, doctors can quickly retrieve key steps in surgical videos through the video Q&A system.
[0081] Video Summary Generation: Automatically generate a concise video summary by understanding the video content to help users quickly understand the core information of the video. For example, in the field of news, the summary of news videos can be automatically generated; in corporate training, the key content of training videos can be generated.
[0082] Video Content Retrieval: Locate specific segments from long videos according to user needs. For example, in film and television production, specific scenes or dialogues can be quickly located.
[0083] Intelligent Monitoring and Security: Automatically identify abnormal behaviors or events by analyzing the content of surveillance videos in real time. For example, in the surveillance systems in public places, the segments where abnormal events occur can be quickly retrieved through video understanding technology, and suspicious behaviors or security hazards can be automatically detected.
[0084] Personalized Recommendation: Generate personalized video recommendations by understanding users' video viewing behaviors and preferences. For example, in video platforms, relevant video segments or complete videos can be recommended based on users' historical viewing records and question content.
[0085] The present invention has the following beneficial effects: 1) By dynamic sampling and cognitive adaptability evaluation, redundant information processing is reduced, and the input length limit of the LLM is broken through; 2) Combining the dual-Agent self-supervised feedback mechanism significantly improves the answer accuracy while reducing manual annotation; 3) Iterative retrieval only supplements information when necessary, optimizing the consumption of computing resources; 4) It can be extended to multi-modal tasks such as video Q&A and summary generation. Description of the Drawings
[0086] Figure 1 is the flowchart of the method of the present invention.
[0087] Figure 2 is the example diagram of the method of the present invention.
[0088] Figure 3 is Figure 2 a clear schematic diagram of all sampled video frames in Detailed Embodiments
[0089] The present invention will be further specifically described below in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0090] As Figure 1As shown in the figure, the embodiment of the present invention provides a dynamic iterative long video understanding method based on a large language model, including:
[0091] Step 1, modeling and analysis at the mathematical level. Drawing on the behavior pattern of humans dragging the progress bar to quickly locate interesting segments when watching videos, analyze the logical thinking chain of humans in dynamically adjusting, finding key segments, and ultimately understanding the video content, and model it from a mathematical perspective. This modeling method can more accurately simulate the thinking and rapid understanding mechanisms of humans, enabling machines to have stronger adaptability and efficiency when understanding videos.
[0092] Model the video understanding process as a Markov process, where the state at each moment corresponds to a stage in the reasoning process, and the state at each moment is composed of multiple pieces of information, including the visual features of video frames , text descriptions , so the state representation is as follows:
[0093] ,
[0094] In each round of reasoning, the answer for the next moment will be generated based on the current state , and then continue to search for key information through the existing answers. Let represent the probability of generating the answer for the next moment under the condition of state , and represent the state at time t + 1. Then the state transition process for the next moment can be expressed as:
[0095] ,
[0096] ,
[0097] where represents the state transition strategy;
[0098] Considering that in the process of humans searching for key video segments, they will make self - judgments and adjustments on whether they are right or wrong based on their own potential knowledge. Use D to represent the judgment strategy, and use to represent the feedback after self - judgment at time t + 1, and to represent all historical answers from time 0 to t. This process of self - supervised correction and learning is expressed as follows:
[0099] ,
[0100] So the state transition process can be further expressed as:
[0101] ,
[0102] The entire video understanding process is broken down into question - answering, self - evaluation, feedback, and correction, continuously iterating and updating the state. The ultimate goal is to maximize the expected value C during the dynamic video understanding process. Combining the above processes, the ultimate optimization goal of the entire thought chain is:
[0103] ,
[0104] The system continuously iterates and updates the state at different stages, gradually optimizing the understanding of video content. This method simulates the cognitive reasoning chain of humans when watching videos, enabling the system to efficiently locate and analyze key segments within limited resources.
[0105] Step 2: Data pre - processing, performing inference, and evaluating cognitive adaptability.
[0106] Step 2 - 1: Video frame sampling: The input video is uniformly sampled at a rate of 1 frame per second (fps) to convert it into an image sequence. The frame rate of videos is usually high, such as 30fps or 60fps, which results in a large amount of image data being generated per second. If all frames are retained during video processing, it will require huge computing and storage resources. Especially in the processing tasks of long - duration videos or large - scale video datasets, the computing cost can be very high. By sampling at a rate of 1 fps, the number of frames to be processed per second is reduced, thus significantly reducing the demand for computing resources while maintaining sufficient scene information. A sampling frequency of 1 fps can capture the basic dynamics and scene changes of the video and avoid unnecessary redundant data. In summary, the sampling method of 1 fps achieves a reasonable balance among computing cost, storage overhead, and information retention, ensuring that the system can efficiently process large - scale video data while retaining sufficient temporal information to support subsequent analysis tasks;
[0107] Step 2-2, Text Description Generation: Use the existing Bootstrapping Language-Image Pre-training (BLIP) model to generate a text description (caption) for each frame of the image. The content of the text description includes the scene, objects, and actions. This text description provides highly structured information for the subsequent reasoning process and can effectively help understand the context and dynamic content of the video. The BLIP model is trained by combining visual and language features and can automatically generate natural language descriptions that conform to the video content, enabling the machine to not only rely on visual features but also use text to enhance the understanding of the scene and events when analyzing the video. Compared with traditional methods that only use visual features, this text enhancement method has two major advantages: 1) Enrich the semantic information of the video: The text description can extract the core events, object relationships, and temporal information in the video, helping the reasoning module understand the scene development more accurately. 2) Improve the utilization rate of multi-modal information. By combining multi-modal information such as visual, text, and context information, the understanding of video content becomes more comprehensive;
[0108] Step 2-3, Image Feature Extraction and Storage: The visual encoder of the BLIP model is used to extract visual features from each frame of the image. These features include, but are not limited to, information such as object recognition, scene understanding, and spatial relationships in the image. The extracted visual features will be saved in the efficient.npy format, providing basic data for fast access during subsequent reasoning and retrieval. This data storage format is designed to optimize processing speed and storage efficiency and support large-scale video analysis tasks. Compared with traditional image storage methods, the.npy format has the following advantages: 1) Fast access and loading: The binary storage method allows data to be directly mapped to memory, improving the efficiency of model reasoning and retrieval. 2) Storage optimization: Supports efficient array storage and compression strategies, reducing data redundancy and occupying less disk space. 3) Strong compatibility: Deeply integrated with the NumPy ecosystem, facilitating seamless docking with various deep learning frameworks (such as PyTorch);
[0109] Step 2-4, Perform Agent Reasoning and Cognitive Adaptability Evaluation: Select 5 initial frames from the image sequence prepared in Step 2-1 according to a uniform distribution. The mathematical representation of the initial state is denoted as , where They are the visual features and text descriptions of the initial frame respectively. From the perspective of event development, its dynamic changes can be modeled by the process of "beginning, development, climax, then development, and ending". Based on the time series characteristics of the event, considering the changes of the event in different stages, the selection of using 5 frames for sampling can be combined with the characteristics of event development to ensure that the initially sampled frames can cover different stages of the entire event development. The sampling points include both important moments and do not overly concentrate on a certain area. If there are too many initialized frames, it will lead to excessive consumption of resources and the input content length exceeding the limit. If there are too few initialized frames, it will increase the number of iterations and the difficulty of frame searching, affecting the efficiency and accuracy of reasoning;
[0110] In the process of reasoning and answer generation, a question-and-answer Agent is designed as the core reasoning entity, which is responsible for understanding the user's question and generating reasonable answers. The question-and-answer Agent includes a language understanding module and an answer generation module. The language understanding module uses GPT-4o to provide prior knowledge for the question-and-answer Agent and endows it with strong language understanding ability. The answer generation module calculates the matching degree between the candidate answers and the known information (visual features and text descriptions), and generates answers according to the calculation results. The user's question and the corresponding text description are concatenated into a prompt and input into the question-and-answer Agent. The question-and-answer Agent generates preliminary answer options according to the prompt and outputs the probability of selecting each option. Let represent the i-th option, represent the probability of selecting the i-th option under the condition of known , where , n represents that there are n options in total and , The calculation formula of
[0111] is as follows:
[0112] where Sim represents the cosine similarity, and this process corresponds to the modeling process in step 1 . Considering the influence of visual features and text descriptions on answer selection comprehensively, the same weight is given to the similarity of the two in the formula. This design shows that the visual information and text description of the video have equal importance in answer selection, ensuring the balanced use of visual and text information. At the same time, by normalizing the similarity calculation of all candidate answers, it is ensured that the sum of the probabilities of all options is 1. Normalization is to make the answer selection follow the probability distribution and avoid a certain answer being overly biased. It effectively helps the question-and-answer Agent generate the most context-compliant answer by calculating the similarity of each candidate answer and converting it into a probability value.
[0113] Let the question-and-answer Agent be represented by LLM. The process of the question-and-answer Agent generating answers through the option probability distribution is:
[0114] ,
[0115] Design the cognitive adaptability degree, and use the cognitive adaptability degree as the predicted expected C proposed in Step 1. The cognitive adaptability degree evaluation includes relative cognitive degree and absolute cognitive degree . Denote the one with the largest probability in as , the one with the second largest probability as
[0116] . The formula and definition of the cognitive adaptability degree are as follows:
[0117] ,
[0118] Relative cognitive degree: Reflect the discrimination and difference between options:
[0119] ,
[0120] Absolute cognitive degree: Measure the certainty and sufficiency of information:
[0121] ,
[0122] Cognitive adaptability degree:
[0123] The relative cognitive degree measures the perception of the differences between answers. The larger the relative cognitive degree, the stronger the difference between options, the clearer the cognition, and the higher the decision-making confidence; the absolute cognitive degree measures the certainty of the overall information based on information entropy. The smaller the information entropy, the larger the absolute cognitive degree, the clearer the cognition, and the higher the certainty of the current task. α is a weight coefficient, indicating the importance of the relative cognitive degree and the absolute cognitive degree in the overall evaluation. Here, it is set to 0.5, that is, the relative cognitive degree and the absolute cognitive degree play an equal role in the final cognitive adaptability degree. The cognitive adaptability degree finally aggregates the evaluation results of both option discrimination and information sufficiency. The cognitive adaptability degree emphasizes the concept of "cognition", reflecting the comprehensive consideration of information sufficiency and option discrimination in the decision-making process. This is not only a quantification of the model's confidence, but also emphasizes the "adaptability" of the model in the face of complex decisions. Maximizing the cognitive adaptability degree will be used as the goal, corresponding to the predicted expected C in the modeling of Step 1;
[0124] If the cognitive adaptability degree is greater than or equal to the preset threshold (0.6) or exceeds the maximum number of iterations (such as 5 times), the answer will be directly output; otherwise, go to Step 3. In this way, the system avoids falling into a long-term iterative process to maximize the target expectation, bringing a lot of unnecessary calculations, and ensures the efficiency of the reasoning process.
[0125] Step 3-1, Self-Supervised Feedback: Design a judging Agent for self-supervised feedback. Use GPT-4-preview to provide prior knowledge to the judging Agent, and for the answer generated in Step 2-4 judge the consistency and correctness of the answer and give a text description of the feedback search range and the expected search and supplementary information, corresponding to the modeling in Step 1 . First, for consistency, check the answer at the previous moment and the current answer to see if they are consistent. If the two are consistent, it indicates that the model's understanding and reasoning direction of the problem are relatively stable, and the coherence of reasoning is relatively high. At this time, the system will take a specific number of frames Y1 before and after the key frame in the current known frame to search for key frames. The calculation formula for the specific number of frames Y1 is:
[0126] Y1 = Y2 / (Y3 - 1),
[0127] Further refine the current reasoning path, similar to when humans view a video and slowly drag the progress bar to obtain more specific content. The more key frames are obtained, the more detailed the search will be. Secondly, for correctness, the system will verify the current answer based on the current known information and the target question, combined with the prior knowledge provided by GPT-4-preview, to see if it matches the question and the existing information. If the current answer is verified as correct, the system will continue to reason on the current path and further supplement the known information; if the answer is inconsistent or incorrect, the system will expand the search range and make a large adjustment to the entire video frame interval, similar to when humans encounter problems in the reasoning process and drag the video progress bar significantly to view other relevant segments in order to obtain more information to correct the reasoning path. The entire process adopts a dual-Agent verification mechanism, that is, one Agent is responsible for generating and reasoning the answer, and the other Agent is responsible for evaluating the correctness and consistency of the answer and giving feedback. This mutual verification between the two Agents enables the model to be self-corrected after each round of reasoning, thereby improving the accuracy and confidence of reasoning. Compared with traditional manual annotation and supervised learning methods, the system using self-supervised feedback and dual-Agent verification mechanism can significantly reduce the dependence on manual annotation and reduce labor costs. In addition, the system can adaptively adjust the reasoning process according to the feedback, thereby improving the accuracy and consistency of reasoning, avoiding the errors caused by manual annotation deviation or deficiency in traditional methods, and enhancing the overall efficiency and decision-making quality;
[0128] Step 3-2, dynamic iterative optimization: According to the feedback video search interval and text description given by the evaluation agent in step 3-1, key frame retrieval and information supplement are performed to dynamically adjust the state , corresponding to the state transition process modeled in step 1 , so far, the formula of dynamic iterative optimization is rewritten as follows:
[0129] ;
[0130] Strategy The implementation process is as follows: Based on the feedback video search interval and text description, the visual features of all frames in the feedback video search interval are retrieved from the .npy format file preprocessed and stored in step 2-3, and the frame with the highest similarity between the visual features and the text description features in the feedback video search interval is selected as a supplement. Assume that the judging agent gives m feedback video search intervals. , the formula is as follows:
[0131] ,
[0132] in Indicates the feedback video search interval The key frames retrieved from Indicates the feedback video search interval The visual features of the jth image in , It is a text description feature;
[0133] The visual features of the retrieved supplementary frame are integrated into the existing information to obtain updated information. At the same time, the corresponding description of the supplementary frame is also updated. Keyframes that represent additional The corresponding text description and specific formula are as follows:
[0134] ,
[0135] ,
[0136] After the update is completed, return to steps 2-4 to enter the next round of reasoning and iteration, regenerate the answer and evaluate the cognitive adaptability until the output conditions are met. Compared with directly inputting the entire video into the large model for reasoning, the iterative optimization method has significant advantages. First of all, directly inputting the entire video may lead to a waste of computing resources and a reduction in reasoning speed because the video data volume is huge and contains a lot of redundant information. The iterative method reduces the amount of input data required for each reasoning by gradually refining information processing and only updating key frames and related descriptions in each round, thus improving efficiency and saving computing resources. In addition, many current large language models (LLMs) have limitations on the input length when processing long texts or large-scale data and cannot process the complete video content at one time. The iterative method effectively avoids this problem by inputting relevant video segments in stages to ensure that the input data does not exceed the input length limit of the model during each round of reasoning. This method can not only focus on the key information related to the current task, improve the accuracy and precision of reasoning, but also avoid information loss caused by the input length limit, thus optimizing the performance and decision-making quality of the model.
[0137] Step 4: Result evaluation. First, the widely recognized open-source EgoSchema_subset dataset is used. This dataset focuses on complex question-and-answer tasks for long videos (with an average duration of 3 minutes), covering various scenarios such as daily life, academic lectures, documentaries, etc., and also includes various types of question classifications, such as causal reasoning, temporal reasoning, object recognition, action recognition, etc., to test the performance of the model in diverse reasoning tasks. There are approximately 5,000 question-answer pairs, all of which are multiple-choice questions. To evaluate the performance of the model, multiple evaluation metrics are used, including the average number of retrieved frames and the question-and-answer accuracy rate. Average number of retrieved frames: This metric measures the average number of frames retrieved to generate the answer to each question when the method is tested on the EgoSchema_subset, and is used to evaluate the computational efficiency. This metric reflects the efficiency of the method in understanding and processing video content. The fewer the number of retrieved frames, the stronger the ability of the method to obtain sufficient information in a short time;
[0138] Question-and-answer accuracy rate : Used to determine whether the answer output by the model is the same as the annotated answer. The formula is:
[0139] ,
[0140] Among them, TP (True Positive) represents the number of generated answers that exactly match the annotated answers, and FP (False Positive) represents the number of generated answers that do not match the annotated answers. The higher the question-answering accuracy rate, the stronger the ability of this method to generate correct answers. The results show that when dealing with complex question answering, this method only needs to retrieve an average of 7.7 frames to achieve an accuracy rate of 65% on the EgoSchema_subset dataset. This result indicates that the model can still effectively generate accurate answers while ensuring a relatively low computational cost and a small number of frames retrieved, demonstrating its efficient reasoning ability and excellent performance. In contrast, many existing methods often need to process dozens or even hundreds of frames to achieve a similar accuracy rate, with lower computational efficiency and performance. This also proves the efficiency and accuracy of this method when facing long videos and complex questions.
[0141] In addition, to further verify the effectiveness of the method, qualitative analysis can also be carried out using any videos and questions provided by users. By comparing the output answers with the expected results of users, it can be judged whether the performance of this method in actual application scenarios meets the expectations. This kind of qualitative analysis not only helps to verify the ability of the model in dealing with real-world video data, but also reveals the performance of the model on different types of questions (such as causal reasoning, action recognition, object recognition, etc.). By comparing with the actual application requirements of users, the model can be further optimized and adjusted to ensure its good adaptability and accuracy in complex environments.
[0142] As Figure 2 shown, the examples of this embodiment include: the user gives a video and asks the question "What was the dog playing with at the beginning?" in the form of a multiple-choice question. As Figure 3As shown, all video frames involved in the examples of this embodiment include: the 1st frame, the 13th frame, the 50th frame, and the 100th frame. Among them, the 1st frame, the 50th frame, and the 100th frame are the initial frames, and the 13th frame is the key frame obtained through retrieval (marked in red). First, 100 frames are sampled from the given video at 1fps, and the visual features of the video frames are extracted using the BLIP model and the text descriptions of the video frames are generated. At initialization, the 1st, 50th, and 100th frames are sampled as the initial information, which includes visual features and text descriptions. Then, the Q&A Agent performs reasoning and calculations based on this information and the question options (A ice, B water, C bag), and gives an answer according to the probability distribution of each option. The cognitive adaptability evaluation module calculates the cognitive adaptability according to the formula. At the first evaluation, the cognitive adaptability is 0.31, which is lower than the threshold of 0.6 and cannot output the result, so it enters the next step. Then, the judgment Agent intervenes, determines that the answer accuracy is insufficient, points out that the current answer does not match "at the beginning" in the question, and relies on the powerful prior knowledge of the large language model to predict that the possible range of the answer is from the 1st to the 50th frame, and the target description is "the dog is playing with something". Based on this feedback, the Q&A Agent supplements and retrieves the 13th frame (whose text description is the dog is playing with a bag) through the similarity-based information retrieval mechanism. Thus, the first round of reasoning ends. Iterate to enter the second round for reasoning again. ( Figure 3 In it, the first-round reasoning path is represented by a blue arrow, and the second-round reasoning is represented by an orange arrow). First, update the state of the known information and add the retrieved key frame. The Q&A Agent gives the probability distribution and the answer again based on the current information. In the cognitive adaptability evaluation of this round, the cognitive adaptability is 0.61, which exceeds the threshold of 0.6 and meets the iteration termination condition, and finally the answer option C is output. In the whole process, only 4 video frames are used, which greatly reduces the consumption of computing resources and improves the reasoning efficiency. This method realizes efficient and accurate long-video Q&A reasoning by means of multi-module collaboration and continuously optimizes the understanding of long-video content and the accuracy of answers in a dynamic iterative manner.
[0143] The present invention provides a dynamic iterative long-video understanding method based on a large language model. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using the prior art.
Claims
1. A dynamic iterative long video understanding method based on large language models, characterized in that, It includes the following steps: Step 1: Based on the self-supervised dynamic iterative thought chain, conduct mathematical modeling and analysis on the video understanding task; Step 2: Preprocess the video input by the user. Perform frame sampling on the video, generate text descriptions for each video frame, and extract the visual features of the video frames. Conduct preliminary reasoning through the Q&A Agent, combine the input text and video frames to generate a preliminary answer. Use the cognitive adaptability evaluation mechanism to evaluate the answer. If the cognitive adaptability meets the requirements, output the answer; otherwise, proceed to Step 3; Step 3: Conduct self-supervised information feedback. At each step of the reasoning process, introduce a judgment Agent to recognize the answer. The judgment Agent assists in confirming whether the answer is ambiguous or requires further detail supplementation by detecting the consistency and accuracy of the answer. The Q&A Agent will perform key frame retrieval and information supplementation based on the feedback results, iteratively update the known information, and then return to Step 2 for the next round of reasoning until the cognitive adaptability reaches the preset standard; Step 4: Conduct result evaluation: Use the Q&A accuracy rate and the average number of retrieved frames as evaluation indicators to conduct quantitative analysis on the open-source dataset to verify the effectiveness of the method; Secondly, conduct qualitative analysis using any video and question provided by the user to verify whether the results meet the expectations; In step 1, the mathematical modeling and analysis of the video understanding task include: modeling the video understanding process as a Markov process, where the state at each moment corresponds to a stage in the reasoning process, and the state S at time t t includes the visual feature V of the video frame t and the text description T t , which is expressed as: S t = {V t , T t}, For each round of reasoning, the next moment's answer A is generated based on the current state S t Then, key information is further searched for through the existing answers. Let P(A t+1 |S t+1 ) denote the probability of generating the next moment's answer A t under the condition of state S t . Let S t+1 represent the state at time t + 1. Then, the state transition process for the next moment is expressed as: t+1 P(A t+1 |S t ) = P(A t+1 |V t , T t ), S t+1 = π(S t , A t+1 ), where π represents the state transition strategy; Let D denote the judgment strategy, and let F t+1 denote the feedback result after self-judgment at time t + 1. A 0:t denotes all historical answers from time 0 to time t. The process of self-supervised correction and learning is expressed as: F t+1 = D(A t+1 , A 0:t , S t ) The state transition process is further expressed as: S t+1 = π(S t , A t+1 , D(A t+1 , A 0:t-1 , S t )) The entire video understanding process is decomposed into question and answer, self-judgment, feedback and correction, continuously iteratively updating the state. The ultimate goal is to maximize the expected expectation C during the dynamic understanding of the video. The final optimization goal of the entire thought chain is: Indicates finding the variable S that maximizes the estimated expected C t and A t+1 ; Step 2 includes: Step 2-1: Conduct video frame sampling: Uniformly sample the video input by the user and convert it into an image sequence; Step 2-2: Generate text descriptions: Use the Bootstrapped Language-Image Pre-training model (BLIP) to generate text descriptions for each frame of the image. The content of the text description includes the scene, objects, and actions; Step 2-3: Extract and store image features: Extract visual features from each frame of the image through the visual encoder of the Bootstrapped Language-Image Pre-training model (BLIP); The visual features are saved in the.npy format; Step 2-4: Execute Agent reasoning and cognitive adaptability evaluation; Step 2-4 includes: Select an initial frame from the image sequence obtained in Step 2-1 according to a uniform distribution. The state S0 at the initial moment is represented as S0 = {V0, T0}, where V0 and T0 represent the visual features and text descriptions of the initial frame respectively; Design a Q&A Agent as the main body for reasoning and generating answers. The Q&A Agent includes a language understanding module and an answer generation module; The language understanding module uses GPT-4o to provide prior knowledge for the Q&A Agent; The answer generation module is used to calculate the matching degree between the answer and the known information and generate an answer; The user's question and the corresponding text description are concatenated into a prompt, which is input into the Q&A Agent. The Q&A Agent generates options for the preliminary answer based on the prompt and outputs the probability of selecting each option, denoted by x i represents the i-th option, and P(x i |V t ,T t ) represents the probability of selecting the i-th option given V t ,T t . The calculation formula is as follows: where \(i\in\{0,1,2,3,\cdots,n - 1\}\), \(n\) represents the total number of options, and \(P(x i |V t ,T t )\in[0,1]\), Sim represents the cosine similarity function, and exp represents the natural exponential function; Represent the Q&A Agent with LLM. The process of the Q&A Agent generating an answer through the option probability distribution is: A t+1 = LLM(P(x i |V t ,T t )), Design the cognitive adaptability degree, and use the cognitive adaptability degree as the expected expectation C proposed in Step 1. The evaluation of the cognitive adaptability degree includes the relative cognitive degree C rel and the absolute cognitive degree C abs ; Calculate the relative recognition C rel , which is used to reflect the discrimination and difference between options: Calculate the absolute awareness C abs , which is used to measure the certainty and sufficiency of information: Calculate C: C = α·C rel +(1 - α)·C abs , Among them, let the one with the highest probability in P(x i |V t ,T t ) be P(x best ), and the one with the second highest probability be P(x secondbest ); α is a weight coefficient; If the cognitive adaptability is greater than or equal to the preset threshold or exceeds the maximum number of iterations, directly output the answer; otherwise, proceed to Step 3; Step 3 includes: Step 3-1, design a judgment Agent for self-supervised feedback; Step 3-2, perform dynamic iterative optimization; Step 3 includes: Step 3-1 includes: the judgment Agent includes an answer checking module and a search range feedback module; The answer checking module uses GPT-4-preview to provide prior knowledge for the answer A generated in steps 2-4 t+1 to judge the consistency and correctness of the answer. First, for consistency, check whether the answer A at the previous moment t is consistent with the current answer A t+1 ; Second, for correctness, check whether the current known information S t and the target question can infer the answer A t+1 ; The search range feedback module is used to predict the key frames in the known frames, denoted as known key frames, give a feedback video search range, and at the same time provide a text description of the expected search and supplementary information: give feedback according to the result judged by the answer checking module. If they are consistent and correct, only take the number of frames Y1 before and after the current known key frame for key frame search. The calculation formula for the number of frames Y1 is: Y1 = Y2 / (Y3 - 1), where Y2 represents the total number of frames of the video, and Y3 represents the number of known frames; When the known key frame is the first frame, only search backward; when the known key frame is the last frame, only search forward; Step 3 includes: Step 3-2 includes: performing key frame retrieval and information supplementation according to the feedback video search interval and text description given by the judging Agent in Step 3-1, and dynamically adjusting the state S t , corresponding to the state transition process S modeled in Step 1 t+1 =π(S t ,A t+1 ,D(A t+1 ,A 0:t-1 ,S t ), thus, the formula for dynamic iterative optimization is rewritten as: {V t+1 ,T t+1} = π({V t ,T t}, LLM(P(x i |V t ,T t )), D(LLM(P(x i |V t ,T t )), A 0:t-1 ,{V t ,T t})), The implementation process of the policy π includes: based on the feedback video search interval and the text description, retrieve the visual features of all frames within the feedback video search interval from the.npy format files obtained in steps 2-3, select the frame with the highest similarity between the visual features within the feedback video search interval and the text description features as a supplement, and assume that the judging Agent gives m feedback video search intervals {I1, I2, …, I m}, and the formula is: wherein represents the key frames retrieved in the m-th feedback video search interval I m and v j represents the visual feature of the j-th image in the m-th feedback video search interval I m and v text is the text description feature; Integrate the visual features of the retrieved supplementary frames into the existing information to obtain updated information. At the same time, update the corresponding description of the supplementary frames, using to represent the supplementary key frames and the corresponding text description, where represents the m-th key frame and the corresponding text description. The specific update formula is: where concat represents the concatenation operation; After the update is completed, return to Step 2-4 to enter the next round of inference and iteration, regenerate the answer and evaluate the cognitive adaptability until the output condition is met; Step 4 includes: Use the EgoSchema_subset dataset, and the evaluation metrics include the average number of retrieved frames and the question-answering accuracy Acc.
2. An electronic device, characterized in that, It includes a processor and a memory. The memory stores program code. When the program code is executed by the processor, the processor executes the steps of the method according to Claim 1.
3. A storage medium, characterized in that, Stores a computer program or instruction. When the computer program or instruction runs on a computer, it executes the steps of the method according to Claim 1.
Citation Information
Patent Citations
Long video question and answer method based on multi-round reasoning of large model agent
CN119202149A