Fine-grained reference data set construction method and system for video understanding

By preprocessing and automated annotation of the original video data, combined with iterative optimization of the multimodal large language model, a fine-grained human behavior video understanding benchmark data set is constructed, solving the problem of low reliability of the existing data set and improving the quality and reliability of the data set.

CN120147782AInactive Publication Date: 2025-06-13SUN YAT SEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510291006.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing benchmark datasets on human behavior video understanding are low reliability, lack detailed consideration of human behavior details, and manual labeling is time-consuming and labor-intensive and easy to introduce subjective bias.

Method used

Provide a fine-grained benchmark data set construction method and system for video understanding. By acquiring original video data for preprocessing, multiple character video clips are generated, and frame-level character position tracking and audio content analysis are performed to determine character annotation information. Based on clear and descriptive tasks, clear and descriptive multiple-choice questions are constructed, and the benchmark data set is constructed through manual verification of answers.

Benefits of technology

Through the iterative optimization of automated annotation and multimodal large language model, fine-grained and multimodal character annotation information is generated, which improves the reliability of the benchmark dataset of human behavior video understanding, reduces manual annotation costs, and reduces scoring bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147782A_ABST
    Figure CN120147782A_ABST
Patent Text Reader

Abstract

The invention discloses a video understanding fine-grained reference data set construction method and system, and relates to the technical field of data set construction, and the method comprises the steps: carrying out the preprocessing of original video data, generating a plurality of character video segments, and determining the character annotation information of each character video segment; according to the explicit tasks, explicit question video clips are selected from all the character video clips, and the associated character labeling information is adopted to construct explicit selection questions; according to the description type task, selecting a description type question video clip from each character video clip, and generating a description type selection question in combination with a plurality of multi-modal large language models; and if each manual verification answer is matched with the corresponding explicit answer item or description answer item, constructing a human behavior video reference data set by adopting each explicit choice question and each description choice question. A fine-grained video understanding reference data set about human behaviors is generated through a semi-automatic technology, and the reliability of the reference data set is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dataset construction, and particularly to a method and system for constructing a fine-grained benchmark dataset for video understanding. Background Art

[0002] In the fields of computer vision and multimodal learning, video understanding of human behaviors is an important part of realizing intelligent and automated analysis, and has currently been widely applied to scenarios such as emotion recognition, behavior monitoring, and human-computer interaction.

[0003] Benchmark datasets are an important basis for models to improve their understanding of complex human behaviors. The quality of benchmark datasets is crucial for the reliability of model evaluation. If the deviation is too large or the details are insufficient, it will lead to distorted evaluation results, affecting the correct evaluation of model performance and the selection of improvement directions. Most of the existing benchmark datasets focus on object recognition and simple action classification, lacking careful consideration of the details of human behaviors. In addition, manual annotation is usually relied on for information annotation of human behaviors, which is time-consuming and laborious and prone to introducing subjective biases, resulting in low reliability of benchmark datasets for video understanding of human behaviors. Summary of the Invention

[0004] The present invention provides a method and system for constructing a fine-grained benchmark dataset for video understanding, which solves the technical problem of low reliability of existing benchmark datasets for video understanding of human behaviors.

[0005] A method for constructing a fine-grained benchmark dataset for video understanding provided by the first aspect of the present invention includes:

[0006] Obtain original video data and perform preprocessing to generate multiple human video segments;

[0007] Perform frame-level human position tracking and audio content analysis on each of the human video segments to determine the human annotation information of each of the human video segments;

[0008] Select explicit question video segments from each of the human video segments according to explicit tasks, and use the associated human annotation information to determine explicit distractors and explicit answer items, and construct them into explicit multiple-choice questions;

[0009] After selecting descriptive question video segments from each of the human video segments according to descriptive tasks, use multiple preset large models to iteratively determine the corresponding descriptive distractors and descriptive answer items according to the task design prompt words, and generate descriptive multiple-choice questions;

[0010] Obtain the manually verified answers for each of the multiple-choice questions with definite answers and the multiple-choice questions with descriptive answers. If each of the manually verified answers matches the corresponding definite answer item or descriptive answer item, construct a human behavior video benchmark dataset using each of the multiple-choice questions with definite answers and each of the multiple-choice questions with descriptive answers.

[0011] Optionally, the obtaining of the original video data and the preprocessing of the original video data to determine multiple human video segments include:

[0012] Obtain the original video data, and perform video resolution filtering, video aesthetics filtering, and content suitability filtering on the original video data through a technical filter to determine the target video data;

[0013] Use a scene mapper to perform scene segmentation on the target video data to obtain multiple segmented video segments;

[0014] Perform clip filtering on each of the segmented video segments according to the clip filtering conditions to determine multiple human video segments.

[0015] Optionally, the performing of frame-level human position tracking and audio content analysis on each of the human video segments to determine the human annotation information for each of the human video segments includes:

[0016] Perform face object detection on each frame of each of the human video segments to determine the face bounding boxes of the humans in each frame;

[0017] Construct a sequence of face bounding boxes for each of the humans according to the overlapping degree of the face bounding boxes of any two adjacent frames;

[0018] Perform human object detection on each frame of each of the human video segments to determine the human body bounding boxes of each frame;

[0019] Use each of the sequences of face bounding boxes to match the corresponding human body bounding boxes to determine the sequences of human body bounding boxes for each of the humans;

[0020] Reconstruct specific human video segments for each of the humans from each of the human video segments according to each of the sequences of human body bounding boxes;

[0021] Determine the first human feature information for each of the specific human video segments through a preset multimodal large language model;

[0022] Use a preset audio content classifier to determine the audio types of each of the human video segments;

[0023] If the audio type is speech, match the audio segments of each of the humans according to the sequence of face bounding boxes through a preset speaker detection model;

[0024] Perform speech recognition on each of the audio clips using a preset speech recognition model to determine the second person feature information of each person;

[0025] Integrate the first person feature information and the second person feature information of each person in each of the person video clips into the person annotation information of each of the person video clips.

[0026] Optionally, after selecting the descriptive question videos from each of the person video clips according to the descriptive task, use multiple preset large models to iteratively determine the corresponding descriptive distractors and descriptive answer items according to the task design prompt words, and generate descriptive multiple-choice questions, including:

[0027] Select descriptive question video clips from each of the person video clips according to the descriptive task;

[0028] Generate the first task description of the descriptive question video clip using a preset first text large model according to the task design prompt words;

[0029] Through a preset first multi-modal large model, generate the second task description of the descriptive question video clip according to the task design prompt words, and respectively determine the first task description and the second task description as the descriptive sub-optimal item and the descriptive better item of the descriptive question video clip;

[0030] Based on a preset second multi-modal large model, generate the third task description of the descriptive question video clip according to the task design prompt words, and respectively determine the descriptive better item and the third task description as the descriptive sub-optimal item and the descriptive answer item of the descriptive question video clip;

[0031] Convert each descriptive sub-optimal item into a descriptive distractor through a preset second text large model according to the distractor prompt words;

[0032] Use the descriptive question video clip and the corresponding descriptive distractors and descriptive answer items to construct a descriptive multiple-choice question.

[0033] Optionally, it further includes:

[0034] If any explicit answer item or descriptive answer item does not match the corresponding manually verified answer, determine it as a new explicit distractor or descriptive distractor, and use the corresponding manually verified answer as a new explicit answer item or new descriptive answer item, and then construct a human behavior video benchmark dataset.

[0035] Optionally, the explicit task includes person appearance time detection, speaking person detection, person audio-visual synchronization detection, person speech content matching, and speaking person image-voice matching detection.

[0036] A fine-grained benchmark dataset construction system for video understanding provided by the second aspect of the present invention includes:

[0037] A data preprocessing module for obtaining original video data and performing preprocessing to generate multiple person video segments;

[0038] A person information annotation module for performing frame-level person position tracking and audio content analysis on each of the person video segments to determine the person annotation information of each of the person video segments;

[0039] A multiple-choice question with definite answer construction module for selecting definite-answer question video segments from each of the person video segments according to a definite task, determining definite-answer distractors and definite-answer items using the associated person annotation information, and constructing them into multiple-choice questions with definite answers;

[0040] A multiple-choice question with descriptive answer construction module for selecting descriptive-answer question video segments from each of the person video segments according to a descriptive task, and then using multiple preset large models to iteratively determine the corresponding descriptive-answer distractors and descriptive-answer items according to the task design prompt words, and generating multiple-choice questions with descriptive answers;

[0041] A benchmark dataset construction module for obtaining the manually verified answers of each of the multiple-choice questions with definite answers and the multiple-choice questions with descriptive answers. If each of the manually verified answers matches the corresponding definite-answer item or descriptive-answer item, then constructing a human behavior video benchmark dataset using each of the multiple-choice questions with definite answers and each of the multiple-choice questions with descriptive answers.

[0042] A computer device provided by the third aspect of the present invention includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the method for constructing a fine-grained benchmark dataset for video understanding as described in any one of the above.

[0043] A computer-readable storage medium provided by the fourth aspect of the present invention has a computer program stored thereon. When the computer program is executed, it implements the method for constructing a fine-grained benchmark dataset for video understanding as described in any one of the above.

[0044] A computer program product provided by the fifth aspect of the present invention includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the method for constructing a fine-grained benchmark dataset for video understanding as described in any one of the above.

[0045] It can be seen from the above technical solutions that the present invention has the following advantages:

[0046] The above solution of the present invention provides a method for constructing a fine-grained benchmark dataset for video understanding, which is characterized by including: obtaining original video data and performing preprocessing to generate multiple human video segments; performing frame-level human position tracking and audio content analysis on each human video segment to determine the human annotation information of each human video segment; selecting explicit question video segments from each human video segment according to explicit tasks, determining explicit distractors and explicit answer items using the associated human annotation information, and constructing them into explicit multiple-choice questions; after selecting descriptive question video segments from each human video segment according to descriptive tasks, using multiple preset large models to iteratively determine the corresponding descriptive distractors and descriptive answer items according to the task design prompt words, and generating descriptive multiple-choice questions; obtaining the manual verification answers of each explicit multiple-choice question and descriptive multiple-choice question, and if each manual verification answer matches the corresponding explicit answer item or descriptive answer item, constructing a human behavior video benchmark dataset using each explicit multiple-choice question and each descriptive multiple-choice question. Based on the above solution, the original video data is first preprocessed to ensure that it contains sufficient human-related scenarios, and then fine-grained, multi-modal human annotation information is obtained through automated annotation. These human annotation information provide a basis for generating different evaluation tasks. In addition, for the descriptive question task, the quality and rationality of the multiple-choice questions are ensured through the iterative optimization of multiple advanced multi-modal large language models. The whole process generates a fine-grained video understanding benchmark dataset for human behavior through semi-automated technology, which helps to improve the reliability of the benchmark dataset for video understanding of human behavior. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 It is a flowchart of the steps of a method for constructing a fine-grained benchmark dataset for video understanding provided by an embodiment of the present invention;

[0049] Figure 2 It is a block diagram of the structure of a system for constructing a fine-grained benchmark dataset for video understanding provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The embodiments of the present invention provide a method and a system for constructing a fine-grained benchmark dataset for video understanding, which are used to solve the technical problem of low reliability of the existing benchmark dataset for video understanding of human behavior.

[0051] In order to make the object, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0052] Please refer to Figure 1 , Figure 1 which is a flowchart of the steps of a method for constructing a fine-grained benchmark dataset for video understanding provided by an embodiment of the present invention.

[0053] A method for constructing a fine-grained benchmark dataset for video understanding provided by the present invention includes:

[0054] Step 101: Obtain the original video data and perform preprocessing to generate multiple person video clips.

[0055] A person video clip refers to a video clip that takes a specific person as the core subject and presents a scene related to the specific person.

[0056] It should be noted that the original video data can collect video resources without copyright and with diverse scenes from platforms such as Pexels. These videos become ideal material sources due to their rich content. By performing preprocessing on the original video data, multiple person video clips are obtained to ensure that the selected person video clips are suitable for subsequent question generation.

[0057] Step 101 includes the following sub-steps:

[0058] Obtain the original video data, and perform video resolution filtering, video aesthetics filtering, and content suitability filtering on the original video data through a technical filter to determine the target video data;

[0059] Use a scene mapper to perform scene segmentation on the target video data to obtain multiple segmented video clips;

[0060] Perform clip filtering on each segmented video clip according to the clip filtering conditions to determine multiple person video clips.

[0061] It should be noted that in this embodiment, the preprocessing includes filtering and segmentation. Specifically:

[0062] First, apply multiple technical filters to the original video data to determine the target video data. A technical filter can be understood as a tool that can filter and screen video data based on a specific technical algorithm. The technical filters in this embodiment include, but are not limited to, a video resolution filter, a video aesthetics filter, and a content appropriateness filter. These filters can be implemented using existing Python libraries or open-source models. Among them, through the video resolution filter, videos with a width of at least 1280 pixels and a height of at least 480 pixels can be selected to ensure that the video has sufficient clarity to meet the requirements of high-resolution analysis. Through the video aesthetics filter, the overall visual aesthetics of the video can be evaluated using an algorithm, and videos that do not meet the visual aesthetic standards can be excluded to ensure the video quality. Through the content appropriateness filter, videos containing inappropriate content can be automatically detected and removed to maintain the suitability and compliance of the content.

[0063] Secondly, use an existing scene mapper to split the target video data into multiple segmented video clips according to scene transitions. This step helps to accurately track the people in the video.

[0064] Finally, for each segmented video clip, further apply clip filtering conditions for clip filtering to determine multiple person video clips. The clip filtering conditions can be understood as requirements for screening, excluding, or retaining specific video clips during video editing. The clip filtering conditions in this embodiment include, but are not limited to, a video duration filtering condition, a video motion score filtering condition, and a face frame ratio filtering condition. Among them, the video duration filtering condition can be used to exclude segmented video clips with a video duration shorter than a video duration threshold such as 1 second to avoid the impact of overly short content on the accuracy of subsequent analysis results. The video motion score filtering condition can be used to screen out segmented video clips with a motion score less than a set minimum score such as 1.2. The motion score can be calculated using video optical flow to ensure that the video clip has sufficient dynamic information for analysis. The face frame ratio filtering condition can be used to retain segmented video clips in which the frames with faces appear more than a ratio threshold such as 65% of the total number of frames of the video clip, so as to retain video clips containing obvious facial features of people and provide a good basis for feature annotation based on the people in the video.

[0065] Step 102: Perform frame-level person position tracking and audio content analysis on each person video clip to determine the person annotation information of each person video clip.

[0066] Frame-level person position tracking refers to detecting and positioning the person target in each frame image of the video clip and continuously tracking its position changes between different frames, so as to obtain the movement trajectory of the person in the entire video clip.

[0067] Audio content analysis refers to performing content analysis on the audio of the video clip.

[0068] Person annotation information refers to the information marked for human-related features in video clips.

[0069] Step 102 includes the following sub-steps:

[0070] Perform face object detection frame by frame on each person's video clip to determine the face bounding boxes of the people in each frame;

[0071] Construct a sequence of face bounding boxes for each person according to the overlap degree of the face bounding boxes of any two adjacent frames;

[0072] Perform human object detection frame by frame on each person's video clip to determine the human body bounding boxes of each frame;

[0073] Match each sequence of face bounding boxes with the corresponding human body bounding boxes to determine the sequence of human body bounding boxes for each person;

[0074] Reconstruct the specific person's video clip for each person from each person's video clip according to the sequence of human body bounding boxes;

[0075] Determine the first person feature information of each specific person's video clip through a preset multi-modal large language model;

[0076] Use a preset audio content classifier to determine the audio type of each person's video clip;

[0077] If the audio type is speech, match the audio clips of each person according to the sequence of face bounding boxes through a preset speaker detection model;

[0078] Perform speech recognition on each audio clip using a preset speech recognition model to determine the second person feature information of each person;

[0079] Integrate the first person feature information and the second person feature information of each person in each person's video clip into the person annotation information of each person's video clip.

[0080] It should be noted that in this embodiment, in order to ensure the acquisition of high-quality, high-precision and efficient multi-modal annotation information, person annotation of the person video clip includes two parts: the person feature information determined by video analysis through frame-level person position tracking, that is, the first person feature information, and the person feature information determined by audio content analysis, that is, the second person feature information. All the generated person feature information will be integrated into a unified structured database for subsequent call in the multiple-choice question and its option generation steps;

[0081] For the frame-level human position tracking part: First, for each human video clip, determine the face bounding box of the human in each frame through face object detection; then, construct the face bounding box sequence of each human according to the overlap degree between the current frame and the previous frame (the set overlap ratio threshold is 0.5); then, according to the human body bounding box determined by frame-by-frame human body object detection, use the face bounding box sequence to match the human body bounding box in the corresponding frame. Since the face is usually located in the upper center of the body, after using models such as YOLO11 to detect all human body bounding boxes in a frame, for a certain face bounding box in the same frame, select the body bounding box with the center point directly below it and the closest distance for matching. Based on the face bounding box sequence to which the face of a certain human belongs in each frame for matching, the entire body trajectory of the same human, that is, the human body bounding box sequence, can be obtained. This step not only realizes an approximate count of the number of humans appearing in the video clip but also provides a basis for subsequent human positioning and highlighting; finally, based on the human body bounding box sequence, crop the frame-level human body bounding box from the human video clip and reconstruct a new video clip concentrated on a specific human, that is, a specific human video clip, and use a preset multi-modal large language model. Adopt an instruction to require the model to describe the characteristics of the specified human in the specific human video clip to obtain the first human characteristic information. The first human characteristic information includes but is not limited to appearance description, detailed description of action postures, and detailed description of expression changes. Among them, the appearance description can output gender, age, race, and appearance information in the format of "Gender: --, Age: --, Race: --, Detailed Appearance: --".

[0082] For the audio content analysis part: First, after extracting the audio of each human video clip, use an existing audio content classifier to classify the audio to determine the corresponding audio type. The audio type includes but is not limited to types such as speech, music, and ambient sound, etc.; if an audio is classified as speech, then through an open-source speaker detection model, based on the face bounding box sequence of each human and the audio of the corresponding human video clip, determine whether this person is speaking, determine the audio clip of each human, and use a preset speech recognition model for speech recognition of each audio clip to determine the second human characteristic information of each human. The speech recognition model in this embodiment includes but is not limited to an automatic speech recognition model, a speech emotion recognition model, and a voice gender and age recognition model, which are used to obtain audio paraphrased content, audio emotion, audio gender, and audio age.

[0083] Step 103: Select the explicit question video clips from each human video clip according to the explicit task, determine the explicit distractors and explicit answer items by using the associated human annotation information, and construct them into explicit multiple-choice questions.

[0084] A definite multiple-choice question refers to a multiple-choice question that provides specific and definite answers and options; a definite answer option refers to the option that is the correct answer to a definite multiple-choice question; a definite distractor refers to a confusing option in a definite multiple-choice question, which is opposite to the definite answer option; a definite question video clip refers to the video clip of the person corresponding to the question of a definite multiple-choice question; a definite task refers to the task of constructing a definite multiple-choice question.

[0085] It should be noted that in specific implementation, a series of multiple-choice questions with definite answers (such as restricted categories, numbers or letters) can be constructed according to the definite task using a specially designed template and the person annotation information obtained through step 102. In this embodiment, the definite task includes the detection of the appearance time of a person, the detection of the speaking person, the audiovisual synchronization detection of the person, the matching of the person's speech content, and the matching detection of the speaking person's image - speech.

[0086] For Task 1 - Specified description of the appearance time detection of a person, that is, it is required to identify the appearance time of the person who meets the appearance description in the question in the person video clip: First, select a definite question video clip according to the standard that the video duration exceeds 7 seconds and the time when the target person appears in the person video clip is between one-third and two-thirds of the total length; then, use the frame positioning carried by the person trajectory (sequence of human body bounding boxes or sequence of face bounding boxes) of the target person to determine the appearance time range of the target person (formatted as an integer), which will be used as the correct answer to generate a question about this person, that is, the definite answer option, and the appearance description in the person annotation information corresponding to the target person can be used to construct a definite question; finally, to construct a definite distractor, three random time intervals can be directly generated near the reference fact time interval to ensure that their overlap with the reference fact interval does not exceed a preset overlap duration such as 4 seconds.

[0087] For Task 2 - Detection of the speaking person: According to the corresponding standard of the task, select a suitable definite question video clip based on the person annotation information, such as: the audio label is "speech", the number of people in the video is 2 to 4, the video lasts at least 4 seconds, and there must be and only one speaking person in the video; use the sequence of face bounding boxes to mark all people, and take the only speaking person as the definite answer option, and the others as definite distractors.

[0088] For Task 3 - Audiovisual synchronization detection of a person: Select a suitable definite question video clip according to the corresponding standard of the task, such as: less than 3 people, a duration of more than 8 seconds, the audio type is "speech", and there is at least one speaking person; divide the video clip into three equal parts by duration and mark the time points, select one of the time points as the definite answer option, and the other time points as definite distractors, and then create a "non-audiovisual-synchronized" video clip by reversing the audio before the selected time stamp.

[0089] For Task 4 - Character Speech Content Matching: Select a video clip of a single person with a duration of more than 5 seconds and an audio type of "speech", where one person is the active speaker, the speech content is in English, and the sentence length exceeds 35 characters; the explicit answer option is the paraphrased content of the automatic speech recognition in the character annotation information obtained for the video clip of the person; use a large language model to generate several semantically similar expressions with a sentence length similar to the correct content as explicit distractors;

[0090] For Task 5 - Speaking Character Image - Speech Matching: Select appropriate explicit question video clips. The video conditions include: the audio tag is "speech", the number of people in the video is between 2 and 4, and the video duration is not less than 4 seconds; further, select the target person as the correct answer based on the correlation of age and gender attributes between the audio and the appearance of the people; to enhance the selectivity of the answer, the character types are divided into only three categories: male, female, and child. Specifically, if the audio age belongs to "child", select the only child in the video clip as the target person, and other people as distractors. If the audio age is "adult", select the only adult of the same gender as the audio in the video clip as the target person, and other people as distractors; the explicit question video clips can mark each optional person with a face bounding box or a body bounding box, and use capital letters for distinction.

[0091] Step 104: After selecting descriptive question video clips from each character video clip according to the descriptive task, use multiple preset large models to iteratively determine the corresponding descriptive distractors and descriptive answer options according to the task design prompt words, and generate descriptive multiple-choice questions.

[0092] A descriptive multiple-choice question refers to a multiple-choice question that provides answers in the form of descriptive text or scenarios; a descriptive answer option refers to the option that is the correct answer to a descriptive multiple-choice question; a descriptive distractor refers to a confusing option in a descriptive multiple-choice question, which is opposite to the descriptive answer option; a descriptive question video clip refers to the character video clip corresponding to the question of a descriptive multiple-choice question; a descriptive task refers to the task of constructing a descriptive multiple-choice question, such as expression change description, action change, and causal analysis of character behavior.

[0093] Step 104 includes the following sub-steps:

[0094] Select descriptive question video clips from each character video clip according to the descriptive task;

[0095] Use the preset first text large model to generate the first task description of the descriptive question video clip according to the task design prompt words;

[0096] Generate a second task description for the descriptive question video clip according to the task design prompt words through a preset first multi-modal large model, and determine the first task description and the second task description as the descriptive sub-optimal item and the descriptive better item of the descriptive question video clip respectively;

[0097] Generate a third task description for the descriptive question video clip according to the task design prompt words based on a preset second multi-modal large model, and determine the descriptive better item and the third task description as the descriptive sub-optimal item and the descriptive answer item of the descriptive question video clip respectively;

[0098] Convert each descriptive sub-optimal item into a descriptive interference item according to the interference prompt words through a preset second text large model;

[0099] Construct a descriptive multiple-choice question by using the descriptive question video clip and the corresponding descriptive interference item and descriptive answer item.

[0100] It should be noted that in this embodiment, for the descriptive task, the following framework can be used to iteratively optimize and determine the descriptive answer item and generate the descriptive interference item:

[0101] 1. Screen and determine the descriptive question video clip, and mark the characters that need to be focused on; specifically, screen the question video according to the task characteristics. For multi-person video clips, it is necessary to further determine the characters used for question setting in the video clip, and then use a red bounding box to highlight the target characters so that the subsequent steps can focus on specific objects;

[0102] 2. Generate preliminary questions and answers based on the video description; specifically, use the video clip after highlighting obtained in the previous step, design corresponding task design prompt words for different task types, and let the first text large model generate a task-oriented detailed description of the characters; for example, for the emotion recognition task, the model should describe the emotion state and its changes of the characters in the video. According to the generated detailed description and the task definition, the first text large model further generates a first task description, and the first task description includes specific descriptive questions and preliminary answers (denoted as A_o), ensuring that the question-answer is clear and task-specific;

[0103] 3. Select multiple multimodal large language models, iteratively optimize the best answers, and obtain materials for generating distractors: For the first multimodal large model, first read the question video (with the target person highlighted) and the question, and independently give the answer of this model, denoted as A_c. Then, let the first multimodal large model perform a second inference. In addition to the question video and the question, also give two answers A_c and A_o, and let it select a better answer between A_c and A_o as the descriptive better item A_best, and the eliminated answer as the descriptive sub-optimal item A_e1. For subsequent second or third multimodal large models, repeat the above steps, but A_o is replaced by A_best selected by the previous model. Finally, the final descriptive better item obtained through iterative optimization by multiple models is the descriptive answer item, and the descriptive sub-optimal items A_e1, A_e2, or / and A_e3 are used as raw materials for generating distractors later;

[0104] 4. Use the multiple descriptive sub-optimal items obtained in the previous step and process them with the second text large model to be the descriptive distractors of the multiple-choice questions. Specifically, it can be required that the second text large model modify the original descriptive sub-optimal items by adding interference factors for specific tasks, convert them into distractors of the questions, and ensure that the generated options are significantly different from the correct answers. The modification instructions can be as follows:

[0105] "The following is a ready-made question and its multiple-choice options: (question), correct answer: (answer), distractor 1: (excluded item 1), distractor 2: (excluded item 2), distractor 3: (excluded item 3). This multiple-choice question may have the following problem: The current distractors are not incorrect, they just represent different answers to the question. But this will make it impossible to select the correct answer. Therefore, I hope you can help me add small and different errors to each of the current distractors so that the correct answer is clearly the only correct answer. The following are the types of small errors that can be selected: (designed according to the question type). Remember, the modified distractors must meet the following requirements: 1. Make minor adjustments based on the original distractors, rather than creating new content from scratch. 2. Need to be significantly different from the answer, and small errors can be added. 3. The distractors need to be different from each other. 4. The length of the distractors should be similar to the correct answer. If it is too short, lengthen the description.";

[0106] By using the answers discarded by the multimodal large model to synthesize options, the options can be more confusing to the large model, thus further testing the discrimination ability of the model.

[0107] Step 105. Obtain the manually verified answers of each explicit multiple-choice question and descriptive multiple-choice question. If each manually verified answer matches the corresponding explicit answer item or descriptive answer item, then use each explicit multiple-choice question and each descriptive multiple-choice question to construct a human behavior video benchmark dataset.

[0108] It further includes:

[0109] If any explicit answer item or descriptive answer item does not match the corresponding manually verified answer, it is determined as a new explicit interference item or descriptive interference item. After using the corresponding manually verified answer as a new explicit answer item or new descriptive answer item, a human behavior video benchmark dataset is constructed. It should be noted that...

[0110] It should be noted that due to certain limitations of the annotation model and the model used for question generation, errors may occur in multiple-choice questions. By shuffling the options of the generated multiple-choice questions and asking the verifier to provide manually verified answers; if the manually verified answer matches the answer item automatically generated by the pipeline, it is confirmed that the question conforms to human cognition, and the current explicit multiple-choice questions and descriptive multiple-choice questions are used to construct the human behavior video benchmark dataset; if they do not match, recheck and decide whether to change the answer item. If the best answer is not found, the previously determined answer item is used as a new interference item for the corresponding multiple-choice question, and the corresponding manually verified answer is used as a new answer item to update the explicit multiple-choice questions and descriptive multiple-choice questions, and then construct the human behavior video benchmark dataset; although the dataset generation process still requires manually verified answers, the role of humans in the whole process is changed from question creators to quality inspectors, which not only ensures the quality of the questions but also significantly reduces the labor cost.

[0111] It should be emphasized that the dedicated task models or large language models used in annotation and question construction can be selected from the current state-of-the-art open-source models. For example, the multimodal large language model used for human appearance, action, and expression annotation in this embodiment is ShareGPT4Video, the text large model involved in question creation and interference item adjustment is Qwen2.5, the multimodal large models for iterative optimization of the optimal answer and generation of interference item materials are VideoLLaMA2, CogVLM, or LAVA-OneVision, voice transcription and voice emotion recognition both use SenseVoice, and the speaker detection model is Light-ASD. With the continuous emergence of more powerful and specialized models, the annotation quality and the quality of multiple-choice question generation can be further improved.

[0112] In the embodiments of the present invention, first, the original video data is preprocessed to ensure that it contains sufficient human-related scenarios. Subsequently, fine-grained, multi-modal human annotation information is obtained through an automated annotation pipeline. These human annotation information provide a basis for generating different evaluation tasks. In addition, a dedicated multiple-choice question generation pipeline is designed for the descriptive question task. This process ensures the quality and rationality of the multiple-choice questions through the iterative optimization of multiple advanced multi-modal large language models, and the answers are manually verified to ensure the accuracy of the multiple-choice questions. The entire process generates a fine-grained video understanding benchmark dataset for human behavior through semi-automated technology, which not only reduces the manual annotation cost, but also reduces the scoring bias, provides a more realistic and rich test environment, ensures the objectivity and accuracy of the evaluation results, helps to improve the reliability of the benchmark dataset for video understanding of human behavior, and further helps to promote the performance improvement and development of video processing MLLMs in human behavior understanding.

[0113] Please refer to Figure 2 , Figure 2 which is a structural block diagram of a system for constructing a fine-grained benchmark dataset for video understanding provided by the embodiments of the present invention.

[0114] A system for constructing a fine-grained benchmark dataset for video understanding provided by the present invention includes:

[0115] A data preprocessing module 201, configured to obtain original video data and perform preprocessing to generate multiple human video segments;

[0116] A human information annotation module 202, configured to perform frame-level human position tracking and audio content analysis on each human video segment to determine the human annotation information of each human video segment;

[0117] An explicit multiple-choice question construction module 203, configured to select explicit question video segments from each human video segment according to an explicit task, determine explicit distractors and explicit answer items by using the associated human annotation information, and construct them into explicit multiple-choice questions;

[0118] A descriptive multiple-choice question construction module 204, configured to select descriptive question video segments from each human video segment according to a descriptive task, and then use multiple preset large models to iteratively determine the corresponding descriptive distractors and descriptive answer items according to the task design prompt words, and generate descriptive multiple-choice questions;

[0119] A benchmark dataset construction module 205, configured to obtain the manually verified answers of each explicit multiple-choice question and descriptive multiple-choice question. If each manually verified answer matches the corresponding explicit answer item or descriptive answer item, then use each explicit multiple-choice question and each descriptive multiple-choice question to construct a human behavior video benchmark dataset.

[0120] Optionally, the data preprocessing module 201 is specifically configured to:

[0121] Obtain the original video data, filter the original video data through a technical filter for video resolution filtering, video aesthetics filtering, and content suitability filtering to determine the target video data;

[0122] Use a scene mapper to perform scene segmentation on the target video data to obtain multiple segmented video clips;

[0123] Perform clip filtering on each segmented video clip according to the clip filtering conditions to determine multiple person video clips.

[0124] Optionally, the person information annotation module 202 is specifically configured to:

[0125] Perform face object detection on each frame of each person video clip to determine the face bounding box of the person in each frame;

[0126] Construct a sequence of face bounding boxes for each person according to the overlap degree of the face bounding boxes of any two adjacent frames;

[0127] Perform human body object detection on each frame of each person video clip to determine the human body bounding box of each frame;

[0128] Match each sequence of face bounding boxes with the corresponding human body bounding box to determine the sequence of human body bounding boxes for each person;

[0129] Reconstruct the specific person video clip of each person from each person video clip according to each sequence of human body bounding boxes;

[0130] Determine the first person feature information of each specific person video clip through a preset multi-modal large language model;

[0131] Use a preset audio content classifier to determine the audio type of each person video clip;

[0132] If the audio type is speech, match the audio clips of each person according to the sequence of face bounding boxes through a preset speaker detection model;

[0133] Use a preset speech recognition model to perform speech recognition on each audio clip to determine the second person feature information of each person;

[0134] Integrate the first person feature information and the second person feature information of each person in each person video clip into the person annotation information of each person video clip.

[0135] Optionally, the descriptive multiple-choice question construction module 204 is specifically configured to:

[0136] Select descriptive question video clips from each person video clip according to the descriptive task;

[0137] Generate a first task description for the descriptive question video clip according to the task design prompt words using a preset first text large model;

[0138] Generate a second task description for the descriptive question video clip according to the task design prompt words through a preset first multimodal large model, and respectively determine the first task description and the second task description as the descriptive sub-optimal item and the descriptive better item of the descriptive question video clip;

[0139] Generate a third task description for the descriptive question video clip according to the task design prompt words based on a preset second multimodal large model, and respectively determine the descriptive better item and the third task description as the descriptive sub-optimal item and the descriptive answer item of the descriptive question video clip;

[0140] Convert each descriptive sub-optimal item into a descriptive interference item according to the interference prompt words through a preset second text large model;

[0141] Construct a descriptive multiple-choice question using the descriptive question video clip and the corresponding descriptive interference item and descriptive answer item.

[0142] Optionally, the benchmark dataset construction module 205 is further configured to:

[0143] If any explicit answer item or descriptive answer item does not match the corresponding manually verified answer, determine it as a new explicit interference item or descriptive interference item, and use the corresponding manually verified answer as a new explicit answer item or new descriptive answer item, and then construct a human behavior video benchmark dataset.

[0144] Optionally, the explicit tasks include person appearance time detection, speaking person detection, person audio-visual synchronization detection, person speech content matching, and speaking person image-voice matching detection.

[0145] An embodiment of the present invention further provides a computer device, including a memory and a processor, and a computer program is stored in the memory; when the computer program is executed by the processor, the processor is caused to execute the steps of the method for constructing a fine-grained benchmark dataset for video understanding according to any one of the above embodiments.

[0146] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by the processor, the steps of the method for constructing a fine-grained benchmark dataset for video understanding according to any one of the above embodiments are implemented.

[0147] An embodiment of the present invention further provides a computer program product, including computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the method for constructing a fine-grained benchmark dataset for video understanding according to any one of the above embodiments are implemented.

[0148] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0149] In several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0150] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0151] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0152] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.

[0153] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for constructing a fine-grained benchmark dataset for video understanding, characterized in that: include: Obtaining raw video data and preprocessing it to generate multiple character video clips; Performing frame-level character position tracking and audio content analysis on each of the character video clips to determine character annotation information for each of the character video clips; Selecting explicit question video clips from each of the character video clips according to the explicit task, determining explicit interference items and explicit answer items using associated character annotation information, and constructing them into explicit multiple-choice questions; After selecting descriptive question video clips from each of the character video clips according to the descriptive task, a plurality of preset large models are used to iteratively determine corresponding descriptive interference items and descriptive answer items according to the task design prompt words, and generate descriptive multiple-choice questions; Obtain manually verified answers to the explicit multiple-choice questions and the descriptive multiple-choice questions. If the manually verified answers match the corresponding explicit answer items or descriptive answer items, construct a human behavior video benchmark dataset using the explicit multiple-choice questions and the descriptive multiple-choice questions.

2. The method for constructing a fine-grained benchmark dataset for video understanding according to claim 1, characterized in that: The obtaining of original video data, preprocessing of the original video data, and determining a plurality of character video clips include: Acquire original video data, and perform video resolution filtering, video aesthetic filtering, and content suitability filtering on the original video data through a technical filter to determine target video data; Using a scene mapper to perform scene segmentation on the target video data to obtain a plurality of segmented video clips; The segmented video segments are subjected to editing and filtering according to the editing and filtering conditions to determine a plurality of character video segments.

3. The method for constructing a fine-grained benchmark dataset for video understanding according to claim 1, characterized in that: The performing frame-level character position tracking and audio content analysis on each character video clip to determine character annotation information of each character video clip includes: Performing face target detection on each of the character video clips frame by frame to determine the face boundary box of the character in each frame; Constructing a sequence of face bounding boxes of each of the characters according to the degree of overlap of face bounding boxes of any two adjacent frames; Performing human target detection on each of the character video clips frame by frame to determine a human body bounding box for each frame; Using each of the face bounding box sequences to match the corresponding human body bounding box, to determine a human body bounding box sequence for each of the characters; Reconstructing a specific person video segment of each person from each person video segment according to each person bounding box sequence; Determining first character feature information of each of the specific character video clips by using a preset multimodal large language model; Using a preset audio content classifier to determine the audio type of each of the character video clips; If the audio type is speech, matching the audio clips of each of the characters according to the face bounding box sequence using a preset speaker detection model; Using a preset speech recognition model to perform speech recognition on each of the audio clips to determine second character feature information of each of the characters; The first character feature information and the second character feature information of each character in each character video clip are integrated into the character annotation information of each character video clip.

4. The method for constructing a fine-grained benchmark dataset for video understanding according to claim 1, characterized in that: After selecting descriptive question videos from each of the character video clips according to the descriptive task, a plurality of preset large models are used to iteratively determine corresponding descriptive interference items and descriptive answer items according to the task design prompt words, and generate descriptive multiple-choice questions, including: Selecting descriptive question video clips from each of the character video clips according to the descriptive task; Using a preset first large text model to generate a first task description of the descriptive question video clip according to the task design prompt words; Generate a second task description of the descriptive problem video clip according to the task design prompt words through a preset first multimodal large model, and determine the first task description and the second task description as the descriptive suboptimal item and the descriptive superior item of the descriptive problem video clip respectively; Based on the preset second multimodal large model, a third task description of the descriptive question video clip is generated according to the task design prompt words, and the descriptive preferred item and the third task description are respectively determined as the descriptive suboptimal item and the descriptive answer item of the descriptive question video clip; Convert each descriptive suboptimal item into a descriptive distractor item according to the distractor word through a preset second text large model; Descriptive multiple-choice questions are constructed using the descriptive question video clips and corresponding descriptive distractor items and descriptive answer items.

5. The method for constructing a fine-grained benchmark dataset for video understanding according to claim 1, characterized in that: Also includes: If any explicit answer item or descriptive answer item does not match the corresponding manually verified answer, it is determined to be a new explicit interference item or descriptive interference item, and the corresponding manually verified answer is used as the new explicit answer item or the new descriptive answer item to construct a human behavior video benchmark dataset.

6. The method for constructing a fine-grained benchmark dataset for video understanding according to claim 1, characterized in that: The explicit tasks include character appearance time detection, speaking character detection, character audio-visual synchronization detection, character voice content matching, and speaking character image-voice matching detection.

7. A fine-grained benchmark dataset construction system for video understanding, characterized in that: include: A data preprocessing module is used to obtain and preprocess the original video data to generate multiple character video clips; A character information tagging module is used to perform frame-level character position tracking and audio content analysis on each character video clip to determine character tagging information of each character video clip; A clear-type multiple-choice question construction module is used to select clear-type question video clips from each of the character video clips according to the clear-type task, use the associated character annotation information to determine clear-type interference items and clear-type answer items, and construct them into clear-type multiple-choice questions; A descriptive multiple-choice question construction module is used to select descriptive question video clips from each of the character video clips according to the descriptive task, use multiple preset large models to iteratively determine corresponding descriptive interference items and descriptive answer items according to the task design prompt words, and generate a descriptive multiple-choice question; The benchmark data set construction module is used to obtain the manually verified answers to each of the explicit multiple-choice questions and the descriptive multiple-choice questions. If each of the manually verified answers matches the corresponding explicit answer item or descriptive answer item, each of the explicit multiple-choice questions and the descriptive multiple-choice questions is used to construct a human behavior video benchmark data set.

8. A computer device, characterized in that: It comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the method for constructing a fine-grained benchmark dataset for video understanding as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method for constructing a fine-grained benchmark dataset for video understanding as described in any one of claims 1-6 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method for constructing a fine-grained benchmark dataset for video understanding as described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Data processing method based on artificial intelligence and related device

    CN110488975A

  • Virtual human expression personalized generation method based on multi-modal interaction information

    CN116311456A

  • Online training method and device based on multi-modal large model, storage medium and server

    CN117876170A

  • Data processing systems for data transfer risk identification and related methods

    US20200004968A1