Key frame screening and video question answering method and device based on target existence and storage medium

By combining target existence tables and open vocabulary detection with original pixel stitching, the problems of frame loss and insufficient visual perception in video understanding models are solved, thereby improving the accuracy and robustness of video question answering.

CN120953872APending Publication Date: 2025-11-14ZHEJIANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511042436.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing video understanding models lose fine-grained actions during frame compression, are sensitive to target size and similarity thresholds during video frame selection, and have insufficient visual perception capabilities, leading to key frame loss and description errors.

Method used

By constructing a target existence table and employing a multi-stage frame filtering strategy, an open-vocabulary target detection method is used to replace similarity calculation. Key targets are verified frame by frame, and visual perception is enhanced by combining original pixel stitching, thus constructing a highly reliable video question answering framework.

Benefits of technology

Preserve fine-grained motion changes, improve keyframe recall robustness, eliminate screening bias, enhance visual perception, and improve the accuracy of video question answering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953872A_ABST
    Figure CN120953872A_ABST
Patent Text Reader

Abstract

The invention discloses a key frame screening and video question-answering method and device based on target existence and a storage medium, and the method comprises the steps: (1) carrying out the uniform sampling of an input video stream at a fixed sampling rate, and generating a frame sequence set with a continuous time sequence; (2) generating a target existence table according to the user question and the input video stream, and prompting the large language model to perform frame screening according to the target existence table to obtain candidate frames; (3) splicing the candidate frames into a single composite image according to a time sequence, inputting the spliced image and a question into a large language model, and outputting a refined key frame sequence; and (4) splicing the refined key frame sequence into a single composite image according to the image synthesis method in the step (3), inputting the spliced image and a user question into a large language model, and outputting a JSON format structured response containing answer options, reasoning interpretation and confidence score. By utilizing the method, fine-grained dynamic actions can be preserved, the frame screening sensitivity is eliminated, and the visual perception is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of deep learning and large language models, and in particular to a method, apparatus and storage medium for keyframe screening and video question answering based on target existence. Background Technology

[0002] With breakthroughs in multimodal large-scale models such as GPT-4o and LLaVA, video understanding models are increasingly becoming a core technology empowering applications such as intelligent security and human-computer interaction. Their ability to analyze the spatiotemporal semantics of videos is of great practical significance for achieving high-order scene cognition. Currently, video understanding algorithms based on large language models typically encompass the following core technical aspects:

[0003] First, frame compression strategies. Methods represented by VideoTree (Wang Z, Yu S, Stengel-Eskin E, et al. Videotree: Adaptive tree-based video representation for llm reasoning on long videos[C] Proceedings of the Computer Vision and Pattern Recognition Conference. 2025:3272-3283.) aim to construct compact video representations. Its core mechanism involves extracting keyframes relevant to queries through an iterative process and constructing a hierarchical tree representation accordingly. This method incorporates multi-granular information by having parent nodes carry coarse-grained scene information and child nodes contain fine-grained details. However, its frame aggregation relies heavily on visual feature similarity clustering: aggregating video frames with similar visual features to the same node significantly reduces the number of frames to be processed. However, this strategy has a fundamental limitation: in consecutive frames with similar visual scenes, it cannot effectively capture fine-grained dynamic changes in human actions, leading to the loss of key temporal information and ultimately limiting the upper limit of the model's inference accuracy.

[0004] Second, the frame selection mechanism. Both VideoAgent (Wang X, Zhang Y, Zohar O, et al. Videoagent: Long-form video understanding with large language model as agent[C] / / EuropeanConference on Computer Vision.Cham:Springer Nature Switzerland,2024:58-76.) and LVNet (Awasthi N, Vermeer L, Fixsen LS, et al. LVNet: Lightweight model for leftventricle segmentation for short axis views in echocardiographic imaging[J].IEEE Transactions on Ultrasonics, Ferroelectrics, and Frequency Control,2022,69(6):2115-2128.) adopt a retrieval-based keyframe selection mechanism based on the CLIP model. VideoAgent uses a large language model (LLM) as a central agent to iteratively call the CLIP tool to retrieve frames related to the generated keywords; LVNet designs a hierarchical keyframe selector to optimize retrieval efficiency. However, the frame selection performance of both methods is limited by the inherent semantic alignment defects of the CLIP model: the model's detection results are highly sensitive to the similarity threshold and the size of the target object in the frame. When the key target region occupies a small proportion, or when the similarity threshold is not adjusted properly, the overall similarity between the video frame and the keyword is lower than the threshold, leading to missed detections.

[0005] Third, frame-aware methods. Currently, most video understanding models, including the algorithms mentioned above, use Visual Language Models (VLMs) to generate text descriptions for single frames, replacing the original pixel input. However, this perceptual paradigm has significant problems: the generated descriptive text often contains a large number of static scene details (such as background decorations) unrelated to the core semantics, while omitting crucial dynamic visual cues (such as the continuous changes in a person's movements). More seriously, guiding the model to generate descriptive text through questions can easily induce model illusions, potentially producing false descriptions that do not match the actual scene content, directly contaminating the downstream inference chain and damaging the reliability of the final answer. Summary of the Invention

[0006] This invention provides a keyframe screening and video question answering method, apparatus and storage medium based on target existence, to solve the shortcomings of existing methods such as loss of fine-grained actions during video frame compression, sensitivity to target size and similarity thresholds during video frame screening, and lack of video frame detail perception.

[0007] A keyframe filtering and video question answering method based on target existence includes the following steps:

[0008] (1) The input video stream is uniformly sampled in the time dimension at a fixed sampling rate to generate a set of temporally continuous frame sequences F = {f1, f2, ..., f...} N}, where N is the total number of sampled frames;

[0009] (2) Generate a target existence table based on the user question Q and the input video stream, and prompt the large language model to filter frames based on the target existence table to obtain candidate frames;

[0010] (3) The candidate frames are stitched together into a single synthetic image in chronological order. The stitched image and the user question Q are input into the large language model, and a refined keyframe sequence is output.

[0011] (4) The refined keyframe sequence is stitched together into a single composite image according to the image synthesis method in step (3). The stitched image and the user question Q are then input into the large language model, and a JSON-formatted structured response containing answer options, reasoning explanations and confidence scores is output.

[0012] The specific process of step (2) is as follows:

[0013] (2-1) Input the user question Q and answer options X into the pre-trained large language model, guide the large language model to perform key target extraction through prompt words, and finally output the key target set O = {o1, o2, ..., o3} in JSON format. m};

[0014] (2-2) Using an open vocabulary target detection model, the elements in the key target set O are detected frame by frame in the sampled frame sequence F, and a target existence table T is generated, where the i-th row records the frame f. i The key objectives that exist within;

[0015] (2-3) Input table T and user question Q into the large language model, perform keyframe filtering, and finally output a JSON object containing the reasoning process and candidate frame sequence.

[0016] In step (2-1), when extracting key objectives, the following constraints must be met:

[0017] Text anchoring: Extract only the key objectives explicitly mentioned in question Q or answer option X;

[0018] Visual verifiability: Exclude invisible elements (such as background music);

[0019] Spatiotemporal locatability: The target must possess spatiotemporal distinguishability characteristics;

[0020] Semantic clarity: Exclude vague expressions (such as "something").

[0021] In steps (2-3), the following principles must be met when performing keyframe filtering:

[0022] Spatiotemporal constraint resolution: Map the time descriptors in Q to the predefined intervals of the frame sequence; the time descriptors include start, middle, and end, and the predefined intervals of the frame sequence are the first, middle, and last 30% of the video frames;

[0023] Fault tolerance analysis: If a critical target is detected in consecutive preceding / following frames but is missing in the current frame, the frame is still retained;

[0024] The lenient screening principle is to exclude a frame only if no key target is detected in the frame.

[0025] In step (3), the candidate frames are stitched together into a single composite image in chronological order. The specific process is as follows:

[0026] (3-1) Mark the frame number in the upper left corner of the candidate frame, with black font color and white border;

[0027] (3-2) Fill the target detection boxes in the candidate frames to the aspect ratio of the original video frames and scale them to the size of the original video frames;

[0028] (3-3) The processed candidate frames are stitched together into a single image in chronological order. A grid layout is used to dynamically calculate the number of rows and columns according to the total number of frames, so that the number of rows and columns is close.

[0029] A keyframe filtering and video question answering device based on target existence includes a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they implement the aforementioned keyframe filtering and video question answering method.

[0030] A computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the aforementioned keyframe filtering and video question-answering methods.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] 1. Preserving fine-grained dynamic actions: This invention verifies the existence of key targets frame by frame, avoiding the loss of key frames in visually similar scenes by clustering strategies, and ensuring the integrity of temporal action changes.

[0033] 2. Eliminate frame screening sensitivity: This invention uses open vocabulary target detection to replace similarity calculation, which completely solves the problem of CLIP's sensitivity to similarity threshold and target size, and improves the robustness of keyframe recall.

[0034] 3. Enhanced visual perception: This invention avoids static redundancy and omission of details caused by text description by using original pixel stitching and dynamic focusing mechanism, thus achieving high-fidelity visual information transmission.

[0035] This invention eliminates keyframe loss by using object detection-driven frame existence verification, solves screening bias by replacing similarity calculation with object detection, and enhances visual perception by using raw visual input, thus constructing a highly reliable video question answering framework. Attached Figure Description

[0036] Figure 1 This is a flowchart of a keyframe filtering and video question answering method based on target existence according to the present invention.

[0037] Figure 2 This is a flowchart illustrating the construction of the target existence table in an embodiment of the present invention.

[0038] Figure 3 This is a schematic diagram of the structure of the device of the present invention. Detailed Implementation

[0039] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate the understanding of the present invention and do not constitute any limitation thereof.

[0040] To address the problems of keyframe loss, sensitivity to frame selection parameters and target size, and weak visual perception capabilities in existing video question answering methods when processing long videos, this invention proposes a keyframe filtering and video question answering method based on target existence. This method significantly improves the semantic understanding ability of large language models for long video content by explicitly constructing a target existence table and combining it with a multi-stage frame filtering strategy.

[0041] This invention uses a video and its corresponding question from the NExT-QA public dataset as example input to illustrate the specific implementation steps. However, the implementation of this invention is not limited to a specific dataset and is only used as a reference case.

[0042] In this embodiment, the main content of the video is a man in red hiking in a forest, taking selfies. The video shows both the man in red (the photographer) and his companions wearing other colors. The user question is: What is the man in red doing while filming himself? The options are: A. Turning around B. Hiking C. Cutting ribbon D. Running E. Dancing.

[0043] like Figure 1 As shown, a keyframe filtering and video question answering method based on target existence includes the following steps:

[0044] Step 1, video preprocessing.

[0045] The input video stream is uniformly sampled at a fixed sampling rate of 1fps to extract discrete video frames, generating a temporally continuous set of frame sequences F = {f1, f2, ..., f...}. N}, where N is the total number of sampling frames, and in this embodiment N = 43.

[0046] Step 2, frame filtering based on the target existence table.

[0047] like Figure 2 As shown, a large language model (gpt-4.1-2025-04-14 in this embodiment, and the same model is used in subsequent steps) is used for constraint-based key target extraction. The text containing the question and its options is input into the pre-trained large language model. The model is guided by prompts to extract key targets explicitly mentioned in question Q and option X, which are expected to appear in the associated video frame f. i Visible and locatable people or objects. The structured hints used in this embodiment are as follows:

[0048] Below is a video-related question and its options. Please analyze the people or objects mentioned in the question and options. For each extracted object, please ensure:

[0049] 1. The person or object is indeed mentioned in the text and is not fictional.

[0050] 2. The person or object must be visible in the video; invisible elements such as music are not allowed.

[0051] 3. The person or object has certain characteristics that allow it to be located in the video. For example, the "second woman" cannot be directly located in the video, but the "woman in blue" can.

[0052] 4. The person or object must have a specific meaning and cannot be too broad, such as "something". Only add objects that meet these conditions to the list. Finally, output a JSON string with keys "question", "A", "B", "C", "D", and "E", corresponding to the question and the list of objects mentioned in each option, respectively. If no person or object can be identified from the given information, return an empty list. The question and options are as follows: What did the man in red do while taking a picture of himself? A. Turn around B. Go hiking C. Cut a ribbon D. Run E. Dance

[0053] The large language model filters the identified candidate targets based on the above constraints, retaining the valid targets that meet all conditions, and finally outputs a structured data object in JSON format.

[0054] {"question":["Man in Red"],"A":["Man in Red"],"B":["Man in Red"],"C":["Man in Red","Ribbon"],"D":["Man in Red"],"E":["Man in Red"]}

[0055] For the sampled frame sequence (f1, f2, ..., fk) extracted from the associated video, an open-vocabulary object detection model is used to perform frame-by-frame detection of the key object set. In this embodiment, the object detection model used is Growing-DINO-tiny. Taking frame 12 as an example, the object detection model detects the object category as "man in red", with a confidence score of 0.914, and the coordinates of the lower left corner of the candidate box are (87.07, 304.71) and the upper right corner are (246.14, 477.62). Based on this detection result, an existence table T is generated and constructed. The i-th row of the table records the list of key objects {oi} appearing in video frame fi. In this embodiment, the existence table is constructed as follows: 0: man in red; 2: man in red; 12-24: man in red; 26-30: man in red; 32-42: man in red.

[0056] The existence table T is input into the large language model, and structured instructions guide the model to perform candidate frame filtering. In this embodiment, the model input is:

[0057] Note: Please think through the steps before giving the final answer.

[0058] Below is a video-related question, followed by an object presence table. Each row of the table represents a video clip, and the numbers at the beginning of the table indicate the start and end frame numbers of that clip. If the clip has only one frame, the number represents that frame. The string following the number represents the list of objects present in that clip. Based on the question and the object presence table, analyze which clips (frame numbers) are relevant to the question. Here are some points to note:

[0059] 1. If the question contains descriptions indicating video positions such as "start", "middle" or "end", then they represent the first 30%, middle 30%, and last 30% of the video, respectively.

[0060] 2. The existence table is constructed by the object detection model and may not be completely accurate. For example, if the same object appears in two consecutive frames but is missing in the middle frame, it may be due to missed detection.

[0061] 3. You only need to exclude frames that are clearly irrelevant to the problem, and make sure not to miss any potentially relevant frames. Don't worry about keeping too many frames, as further filtering will be done in later steps.

[0062] Finally, output a JSON string where the "explain" field records your reasoning and analysis process, and the "frame" field contains a JSON list. Each item in the list represents the frame number of the segment related to the question. You can use a single number to represent a single frame, or a string in "start-end" format to represent a range of consecutive frames; in the actual output, replace "start" and "end" with the actual start and end frame numbers. Here is the question: What did the man in red do while filming himself? Here is the object existence table: 0: Man in red\n2: Man in red\n12-24: Man in red\n26-30: Man in red\n32-42: Man in red. Please provide your answer.

[0063] After the large language model completes the filtering inference, it generates standardized output results. The output is a structured JSON object containing two required fields: the "explain" field must provide the complete reasoning process of the large language model's filtering decision; the "frame" field is a list storing the identifiers of all final candidate frames in chronological order. In this embodiment, the model's output is:

[0064] {"explain":"The question asks what the man in red did while filming himself. Therefore, all frames where 'man in red' appears are likely relevant, as the question concerns his behavior during his appearance. According to the object presence table, the man in red appears in frames 0, 2, 12-24, 26-30, and 32-42. Although there are some gaps (e.g., frames 1, 3-11, 25, and 31) where the man in red is not detected, only frames where he is detected are relevant since the question concerns his behavior. No further filtering is needed at this stage."frame":[0,2,"12-24","26-30","32-42"]}

[0065] Step 3: Filter keyframes based on image stitching.

[0066] The candidate frame sequence output in step 2 is stitched together in chronological order to form a single composite image. This image is then input into the large language model, and prompts are provided in conjunction with the original question Q. The LLM is then instructed to output the final keyframe sequence number most relevant to question Q.

[0067] Specifically, before synthesizing each candidate image frame, a corresponding frame number is superimposed on its upper left corner. In this embodiment, taking frame 12 as an example, the annotation text is "Frame 12," using black font with a white border to ensure high recognizability against both dark and light image backgrounds. Then, to improve the density of key information and reduce background interference, the candidate boxes output by the target detection model are sorted according to the original video frame f. i The aspect ratio is filled, and finally scaled to the original frame size. In this example, taking frame 12 as an example, the original frame is 640 wide and 480 high. The candidate box of the target object in this frame is 159 wide and 172 high. The width of the candidate box is filled to 640 / 480*172=299 to make it consistent with the proportion of the original frame. Then it is enlarged proportionally by 480 / 172=2.79 times to make it consistent with the size of the original frame. After that, the processed candidate frames are spatially arranged and stitched in chronological order. The stitching adopts a grid layout from top to bottom and from left to right. For different numbers of candidate frames, the number of rows and columns of the grid is dynamically calculated and set. In this embodiment, there are 21 candidate frames, and the number of rows and columns of the grid is set to 5. Finally, the stitched image and the cue words containing the original question Q are input into the large language model. Based on the question Q and the analysis of the content in the stitched image, the cue large language model identifies and removes redundant frames and highly similar frames, and finally outputs a refined sequence of key frame numbers that is highly relevant to the question Q.

[0068] The input to the model in this embodiment is:

[0069] Note: Please think through the steps before giving the final answer.

[0070] Here is a video-related question, along with relevant frames. The image consists of multiple frames stitched together chronologically, separated by black borders. The top-left corner of each frame indicates its frame number; there are a total of 43 frames.

[0071] **When selecting a frame:**

[0072] 1. If a frame is irrelevant to the problem, ignore it.

[0073] 2. If multiple frames have highly similar content but do not introduce new information relevant to answering the question, only one such frame should be retained.

[0074] Finally, output a JSON string where the "explain" field records your reasoning and analysis process, including the reasons for deleting or keeping frames, and the "frame" field contains a JSON list. Each item in the list represents the frame number of the segment related to the question. You can use a single number to represent a single frame, or a string in "start-end" format to represent a series of consecutive frames; in the actual output, replace "start" and "end" with the actual start and end frame numbers. Here is the question: What did the man in red do while filming himself? Output example: {"frame": [1, "2-4", 5], "explain": Please fill in your explanation here}, please provide your answer.

[0075] The model output is:

[0076] {"frame":[0,13,15,17,19,24,27,30,35,40],"explain":"I examined each frame step by step to explore its relevance to the question: what did the man in red do while filming himself? Frames 0, 13, and 15 show the man in red beginning to film himself walking on the path. Frames 17, 19, 24, 27, 30, 35, and 40 show the man in red continuing to walk, occasionally turning the camera to himself and his companion, and interacting with them. I omitted frames that were almost identical to adjacent frames or did not add new information (e.g., frames where the scene or action did not change). The selected frames collectively show the scene of the man in red walking, filming himself, and interacting with his hiking companions."}

[0077] Step 4: Answer generation.

[0078] The refined keyframe sequence output from step 3 is arranged chronologically and directly stitched together to generate the final synthetic image for question answering. The key difference from previous steps is that, to avoid information loss due to cropping, this step preserves the complete original video frames and no longer performs key target area focusing operations. Subsequently, this synthetic image and the original question Q are input into a large language model, instructing it to analyze the question based on the complete visual content and output a structured response. The response must be in JSON format and contain three elements: answer options (represented by single-letter identifiers such as A / B / C / D / E), reasoning explanation (explaining the basis for eliminating incorrect options and selecting the final answer), and confidence score (using integers 1 / 2 / 3 to represent uncertain / partially certain / very certain, respectively).

[0079] The input to the model in this embodiment is:

[0080] Note: Please think through the steps before giving the final answer.

[0081] This is a video-related question: What did the man in red do while filming himself? A. Turn around B. Go hiking C. Cut a ribbon D. Run E. Dance. The following are the relevant frames. The image is composed of multiple frames stitched together chronologically, with black borders between each frame. The top left corner of each frame indicates the current frame number, with a total of 43 frames. Please answer the question based on the information in the image. Output a JSON string containing three keys: 'answer', 'explain', and 'confidence'. 'answer' corresponds to a single letter (A, B, C, D, E) representing your chosen answer; 'explain' explains how you eliminated incorrect options before choosing the final answer; 'confidence' represents your confidence in the answer, with values ​​of 1, 2, and 3, where 1 indicates uncertainty, 2 indicates partial certainty, and 3 indicates absolute certainty.

[0082] The model output is:

[0083] {“answer":"A","explain":"I analyzed these frames step by step: In frame 0, the camera is shooting forward along a path. By frame 13, the man in red appears in the frame, and by frame 15, he turns his head and body, suggesting a change in the shooting direction. In subsequent frames (frames 17, 24, 27, 30, and 40), the camera's perspective changes, showing people who were previously behind, indicating that the man in red turned back to film them. I ruled out B (hiking) because, in hiking, the specific action asked about is what he did while filming himself. C (ribbon cutting), D (running), and E (dancing) have no visual evidence to support them. The strongest answer is that he turned back while filming.","confidence":3}

[0084] Using the method of this invention, a 0.2% (75.6% to 75.8%) improvement was achieved compared to the state-of-the-art method (Video-Tree) on the NExT-QA dataset, a 5% (66.2% to 71.2%) improvement was achieved compared to the state-of-the-art method (Video-Tree) on the EgoSchema-Subset dataset, and a 1.2% (61.1% to 62.3%) improvement was achieved compared to the state-of-the-art method (Video-Tree) on the EgoSchema-fullset dataset. This demonstrates that the frame filtering and video question answering method based on target existence proposed in this invention can effectively improve the model's ability to perceive videos, thereby improving the accuracy of video question answering.

[0085] Based on the same inventive principle, this invention also provides a keyframe filtering and video question answering device based on target existence, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the keyframe filtering and video question answering method based on target existence mentioned in the above embodiments.

[0086] The embodiments of the keyframe screening and video question-answering method based on target existence of this invention can be applied to any device with data processing capabilities, such as a computer or other similar device. The device embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 3 The diagram shown is a hardware structure diagram of any device with data processing capabilities for the keyframe filtering and video question answering method based on target existence according to the present invention. (Except for...) Figure 3 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.

[0087] Based on the same inventive principle, embodiments of the present invention also provide a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the keyframe screening and video question answering method based on target existence mentioned in the above embodiments.

[0088] The embodiments described above provide a detailed explanation of the technical solutions and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A keyframe filtering and video question answering method based on target existence, characterized in that, Includes the following steps: (1) The input video stream is uniformly sampled in the time dimension at a fixed sampling rate to generate a set of temporally continuous frame sequences F = {f1, f2, ..., f...} N }, where N is the total number of sampled frames; (2) Generate a target existence table based on the user question Q and the input video stream, and prompt the large language model to filter frames based on the target existence table to obtain candidate frames; (3) The candidate frames are stitched together into a single synthetic image in chronological order. The stitched image and the user question Q are input into the large language model, and a refined keyframe sequence is output. (4) The refined keyframe sequence is stitched together into a single composite image according to the image synthesis method in step (3). The stitched image and the user question Q are then input into the large language model, and a JSON-formatted structured response containing answer options, reasoning explanations and confidence scores is output.

2. The keyframe filtering and video question answering method based on target existence according to claim 1, characterized in that, The specific process of step (2) is as follows: (2-1) Input the user question Q and answer options X into the pre-trained large language model, guide the large language model to perform key target extraction through prompt words, and finally output the key target set O = {o1, o2, ..., o3} in JSON format. m }; (2-2) Using an open vocabulary target detection model, the elements in the key target set O are detected frame by frame in the sampled frame sequence F, and a target existence table T is generated, where the i-th row records the frame f. i The key objectives that exist within; (2-3) Input table T and user question Q into the large language model, perform keyframe filtering, and finally output a JSON object containing the reasoning process and candidate frame sequence.

3. The keyframe filtering and video question answering method based on target existence according to claim 2, characterized in that, In step (2-1), when extracting key objectives, the following constraints must be met: Text anchoring: Extract only the key objectives explicitly mentioned in question Q or answer option X; Visual verifiability: Excluding invisible elements; Spatiotemporal locatability: The target must possess spatiotemporal distinguishability characteristics; Semantic clarity: Exclude vague expressions.

4. The keyframe filtering and video question answering method based on target existence according to claim 2, characterized in that, In steps (2-3), the following principles must be met when performing keyframe filtering: Spatiotemporal constraint resolution: Mapping the time descriptors in Q to predefined intervals of the frame sequence; Fault tolerance analysis: If a critical target is detected in consecutive preceding / following frames but is missing in the current frame, the frame is still retained; The lenient screening principle is to exclude a frame only if no key target is detected in the frame.

5. The keyframe filtering and video question answering method based on target existence according to claim 4, characterized in that, The time descriptor includes start, middle, and end, and the predefined interval of the frame sequence is the first, middle, and last 30% of the video frames.

6. The keyframe filtering and video question answering method based on target existence according to claim 2, characterized in that, In step (3), the candidate frames are stitched together into a single composite image in chronological order. The specific process is as follows: (3-1) Mark the frame number in the upper left corner of the candidate frame, with black font color and white border; (3-2) Fill the target detection boxes in the candidate frames to the aspect ratio of the original video frames and scale them to the size of the original video frames; (3-3) The processed candidate frames are stitched together into a single image in chronological order. A grid layout is used to dynamically calculate the number of rows and columns according to the total number of frames, so that the number of rows and columns is close.

7. A keyframe filtering and video question-answering device based on target existence, characterized in that, The device includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the keyframe filtering and video question answering method according to any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, It stores a program that, when executed by a processor, implements the keyframe filtering and video question answering method as described in any one of claims 1-6.

Citation Information

Cited By

  • Spatial intelligent dynamic video frame sampling method and system based on question semantics

    CN121436196A

  • Long video multi-mode understanding method and system for prototype memory enhancement and space-time redundancy suppression

    CN121438192A