Video processing model training method and video question answering method
By acquiring a sample training dataset for uniqueness verification and target object masking, and generating and labeling prompt video frames, the problem of information loss and localization accuracy in instance-level video understanding of multimodal large language models is solved, thereby improving the model's recognition and reasoning capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHUXING TECH (BEIJING) CO LTD
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-31
AI Technical Summary
Existing multimodal large language models suffer from problems such as information loss, poor target instance localization accuracy, lack of multi-turn interactive reasoning ability, and lack of specialized training paradigms in instance-level video understanding tasks, resulting in poor task performance.
By acquiring a sample training dataset, performing uniqueness verification and target object masking, generating prompt video frames and labeling them, a high-quality training dataset is constructed, and the video question-answering model is trained until the stopping condition is met.
It improves the accuracy and reliability of instance-level video understanding, enhances the model's ability to identify target objects and perform spatiotemporal reasoning, and improves the effectiveness of multi-turn interactive reasoning.
Smart Images

Figure CN122493367A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to training methods for video processing models and video question-answering methods. This specification also relates to a training apparatus for video processing models, a computing device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] In recent years, the rapid development of multimodal large language models (MLLMs) has greatly promoted the development of video understanding technology. In video understanding tasks, related work utilizes enhanced image question answering, localization, and segmentation capabilities. In video processing, some models are applied to video question answering and temporal localization tasks.
[0003] However, despite significant progress made by large language models in video understanding tasks, instance-level video understanding remains largely unsolved. Instance-level video understanding refers to the accurate understanding and localization of specific target instances within complex scenes, enabling the model to answer questions or complete specific tasks based on those instances. Current methods for instance-level video understanding typically involve encoding and compressing the video before processing it within a multimodal large language model. This approach leads to information loss, poor target instance localization accuracy, and a lack of multi-turn interactive reasoning capabilities, resulting in poor performance and low accuracy for multimodal large language models in instance-level video understanding tasks. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a training method for a video processing model and a video question-answering method. This specification also relates to a training apparatus for a video processing model, a computing device, a computer-readable storage medium, and a computer program product, to address the aforementioned problems existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a method for training a video processing model is provided, comprising: Obtain a training data set of samples to be processed, wherein the training data set of samples to be processed includes at least one sample video and sample question-answer pairs corresponding to each sample video; The uniqueness of the target object in the sample video is verified based on the sample question-and-answer pair. Based on the verification result, a reference sample training data set is selected from the sample training data set to be processed. The reference sample training data set includes sample videos. The target description information of the target object in the sample question-and-answer pair is rewritten and replaced with a preset placeholder. Based on the target description information, the prompt video frames are determined from each sample video, and the target objects in the prompt video frames are masked. The target objects in the prompt video frame are marked according to at least two preset tags to generate sample prompt video frames. The preset placeholders in the rewritten sample question and answer pair are corrected according to the marking results to generate target sample question and answer pairs. A target sample training dataset is constructed based on sample videos, target sample question-and-answer pairs, and sample prompt video frames. The video question-and-answer model is then trained based on the target sample training dataset until the model training stops.
[0006] According to a second aspect of the embodiments of this specification, a video question-answering method is provided, comprising: Obtain the video to be processed and the target problem corresponding to the video to be processed; The video to be processed and the target question are input into the video question answering model to obtain the target answer output by the video question answering model. The target answer is generated by the video question answering model through a multi-round video tool call. The video question answering model is trained using the training method of the video processing model described above.
[0007] According to a third aspect of the embodiments of this specification, a training apparatus for a video processing model is provided, comprising: The acquisition module is configured to acquire a set of training data for samples to be processed, wherein the set of training data for samples to be processed includes at least one sample video and sample question-answer pairs corresponding to each sample video; The verification module is configured to perform uniqueness verification on the target object in the sample video based on the sample question-and-answer pair, and select a reference sample training data set from the sample training data set to be processed based on the verification result. The reference sample training data set includes the sample video, rewrites the sample question-and-answer pair and the target description information of the target object, and replaces the description information of the target object in the rewritten sample question-and-answer pair with a preset placeholder. The masking module is configured to determine the prompt video frame from each sample video based on the target description information, and to mask the target object in the prompt video frame; The rewriting module is configured to mark the target objects in the prompt video frames according to at least two preset tags, generate sample prompt video frames, and modify the preset placeholders in the rewritten sample question-and-answer pairs according to the marking results, thereby generating target sample question-and-answer pairs. The training module is configured to construct a target sample training dataset based on sample videos, target sample question-and-answer pairs, and sample prompt video frames, and to train a video question-and-answer model based on the target sample training dataset until the model training stops.
[0008] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.
[0009] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0010] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0011] The training method for the video processing model provided in this specification involves: acquiring a training dataset of samples to be processed, wherein the training dataset includes at least one sample video and corresponding question-and-answer pairs for each sample video; verifying the uniqueness of target objects in the sample videos based on the question-and-answer pairs; selecting a reference training dataset from the training dataset of samples to be processed based on the verification results, wherein the reference training dataset includes sample videos; rewriting the target description information of the target objects in the rewritten question-and-answer pairs, replacing the target object description information in the rewritten question-and-answer pairs with preset placeholders; determining prompt video frames from each sample video based on each target description information, and masking the target objects in the prompt video frames; marking the target objects in the prompt video frames according to at least two preset tags respectively, generating sample prompt video frames; correcting the preset placeholders in the rewritten question-and-answer pairs based on the marking results, generating target sample question-and-answer pairs; constructing a target sample training dataset based on the sample videos, target sample question-and-answer pairs, and sample prompt video frames; and training the video question-and-answer model based on the target sample training dataset until the model training stops.
[0012] This specification provides a training method for a video processing model, which verifies the uniqueness of target objects in sample videos through sample question-and-answer pairs. This ensures the uniqueness of subsequent training data and avoids the problem of the video question-and-answer model being unable to identify target objects during video recognition processing. Simultaneously, the description information of target objects in the sample question-and-answer pairs is replaced with preset placeholders, facilitating faster and more accurate description replacement using preset tags later. Furthermore, this method uses multiple preset tags to mark target objects in prompt video frames, generating sample prompt video frames, thereby expanding the amount of training data. This method constructs a batch of high-quality model training data, and by embedding video processing tools into the framework of the video question-and-answer model, the high-quality training data is used to train the video question-and-answer model until the model training stops. Attached Figure Description
[0013] Figure 1 This is a flowchart of a training method for a video processing model provided in one embodiment of this specification; Figure 2 This is a flowchart illustrating a video question-and-answer method provided in one embodiment of this specification; Figure 3 This is a schematic diagram of the structure of a training device for a video processing model provided in one embodiment of this specification; Figure 4 This is an architecture diagram of a training system for a video processing model provided in one embodiment of this specification; Figure 5 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0018] Multi-modal Large Language Models (MLLMs) are models trained by combining multimodal information such as text, images, audio, and video. They are artificial intelligence models that extend the capabilities of traditional large language models, enabling them to simultaneously process, understand, and generate multiple types of information modalities. Using a large language model as its core inference engine, they break down semantic barriers between different modalities (such as vision, speech, and images) and language by introducing modal encoders and cross-modal alignment mechanisms, achieving cross-modal understanding and generation.
[0019] GRPO (Group Relative Policy Optimization) is a policy optimization algorithm used for reinforcement learning training of large language models. It aims to improve model performance in complex tasks (such as mathematical reasoning, programming, and multi-step reasoning) while reducing computational costs and improving training stability. GRPO does not rely on additionally trained value models to evaluate the quality of individual actions; instead, it estimates advantage through within-group relative comparisons. For the same input cue, it generates multiple candidate outputs, groups these outputs, and calculates the relative advantage of each sample using the statistical properties of within-group rewards.
[0020] RL: Reinforcement Learning is a machine learning method in the fields of artificial intelligence and machine learning that involves an agent interacting with its environment and learning optimal policies based on reward signals. Reinforcement learning is an important branch of machine learning; it is an adaptive learning method where an agent continuously interacts with its environment, performs actions through trial and error, perceives the environmental state, receives immediate reward feedback, optimizes its decision-making strategy based on accumulated rewards, and autonomously learns the optimal behavioral decision rules to maximize long-term gains.
[0021] In recent years, the rapid evolution of multimodal large language models (MLLMs) has greatly promoted the development of video understanding technology. At the same time, research represented by inference large language models has shown that post-training based on reinforcement learning can significantly improve the inference ability of the model, and RL algorithms such as GRPO (Group Relative Policy Optimization) have been extended from the pure text domain to the multimodal domain.
[0022] In image processing, related work utilizes RL to enhance image question answering, localization, and segmentation capabilities. In video processing, RL algorithms are applied to video question answering and temporal localization tasks. Furthermore, multimodal large language models can perform native temporal retrieval and inference by dynamically selecting and re-examining relevant video segments. Multimodal large language models can also utilize a vision toolkit to densely sample video frames as needed during inference.
[0023] However, despite significant progress made by multimodal large language models in general video understanding, they remain insufficiently addressed in the crucial task of instance-level video understanding. Instance-level video understanding refers to the accurate understanding and location of specific target instances within complex scenes. Given a video and a visual cue (such as a target marked with a rectangle, number, arrow, etc.), the model needs to answer questions or complete related tasks for that specific target instance. Unlike general video question answering, instance-level video understanding requires the following two capabilities: 1. Accurate instance recognition capability, that is, the ability to accurately match visual cues with target instances in the video.
[0024] 2. Spatiotemporal reasoning ability, that is, the ability to understand the behavior and state changes of the target instance at different points in time in the video.
[0025] The current mainstream strategy for handling instance-level video understanding tasks is the single-pass paradigm. Specifically, this strategy encodes and compresses candidate videos into a shared context, then generates the answer in a single forward pass using a multimodal large language model. While structurally simple, this design suffers from the following problems in practical applications: I. Information Loss and Evidence Obscuration. Compressing lengthy video content into fixed-size feature representations can easily lead to severe context overload and information loss, thereby obscuring crucial but decisive clues. In particular, target instances marked in visual cue frames are easily overlooked or buried in massive amounts of information, causing the model to be unable to establish a reliable correspondence between visual cues and target instances.
[0026] Second, the accuracy of target instance localization is poor. Compressing the entire video into a shared context severely limits the model's ability to accurately locate evidence. In instance-level video understanding tasks, the model needs to accurately associate visual cues with specific pixel regions in video frames, but the single-pass paradigm cannot support this fine-grained spatial localization, leading to instance recognition errors or localization drift.
[0027] Third, the ability for multi-turn interactive reasoning is indeed lacking. Most existing video reasoning methods based on reinforcement learning still rely on pure text-based thought chain reasoning, which is essentially language-centric. When faced with long videos, text-based intermediate reasoning results cannot effectively access and utilize the original visual information, are prone to visual illusions, and cannot dynamically revisit specific segments of the video based on intermediate reasoning results, resulting in insufficient reliability and verifiability of the reasoning.
[0028] Fourth, there is a lack of specialized training paradigms for instance-level characters. Existing reinforcement learning methods are mainly designed for general video understanding and temporal localization tasks, neglecting the core challenge of instance-level video understanding. Specifically, there is a lack of a training method that can organically combine visual cues, target instances, and spatiotemporal reasoning, resulting in a severe deficiency in the model's generalization ability on instance-level characters.
[0029] As a result, the existing single-pass paradigm is severely limited in its reliability and accuracy in instance-level video understanding tasks because it cannot actively and dynamically acquire and filter multimodal evidence in videos before reasoning, and it lacks specialized training methods for instance-level video understanding.
[0030] Based on this, this specification provides a training method for a video processing model and a video question-answering method. This specification also relates to a training device for a video processing model, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0031] Figure 1 A flowchart illustrating a training method for a video processing model according to an embodiment of this specification is shown, specifically including the following steps: Step 102: Obtain the training data set of the samples to be processed, wherein the training data set of the samples to be processed includes at least one sample video and sample question-answer pairs corresponding to each sample video.
[0032] The training dataset for the samples to be processed refers to the set of training data used to train the video processing model. In the method provided in the embodiments of this specification, the video processing model can perform instance-level video processing tasks, that is, the task of answering questions related to the target object in the video through the video. Based on this, the training dataset for the samples to be processed includes multiple training data, which appear as sample videos and corresponding sample question-and-answer pairs. In practical applications, the training dataset for the samples to be processed includes at least one sample video, and each sample video corresponds to a corresponding sample question-and-answer pair.
[0033] Sample videos are the video foundation for question-and-answer sessions during model training. A sample question-and-answer pair specifically refers to a sample question and its corresponding sample answer for the target object within the sample video. For example, taking sample video A, the sample question in the sample question-and-answer pair could be "Is the red car in the video a truck?" and the sample answer could be "Yes, it is a truck."
[0034] In the method provided in the embodiments of this specification, a training data set of samples to be processed can be obtained. This training data set includes multiple sample videos and corresponding question-and-answer pairs for each sample video. The sample videos and question-and-answer pairs appear in pairs. The training data set of samples to be processed provides the necessary data foundation for subsequent preprocessing of training data.
[0035] In practical applications, there are usually training datasets for multimodal large language models, and these training datasets often include question-answering training data related to videos. However, these training datasets are typically general training sets for videos, not dedicated training sets for instance-level video understanding tasks. Therefore, in order to train for instance-level video understanding tasks in this embodiment, it is necessary to identify a dedicated training set for instance-level video understanding tasks from the general training set of video processing models. Based on this, in a specific embodiment provided in this specification, obtaining the training data set of the samples to be processed includes: Obtain an initial sample training data set, which includes sample videos and corresponding initial sample question-answer pairs; Each initial sample question-and-answer pair is input into the text recognition model to obtain the sample question-and-answer pairs output by the text recognition model, wherein the sample question-and-answer pairs are question-and-answer pairs for the target object in the sample video; The sample question-and-answer pairs and the corresponding sample videos are used as the training data set of the samples to be processed.
[0036] The initial sample training dataset specifically refers to a general training set for the video processing task. This dataset includes multiple sample videos and corresponding initial question-and-answer pairs for each video. Specifically, the initial question-and-answer pairs refer to any question-and-answer pairs related to a sample video.
[0037] In this embodiment, a preliminary screening is performed on the general training set used for video processing tasks to select the training data set of samples to be processed for instance-level video understanding tasks. To improve processing speed, in the method provided in this specification, each initial sample question-and-answer pair is input into a text recognition model. The text recognition model performs semantic understanding and analysis on the initial sample question-and-answer pairs, and selects sample question-and-answer pairs from the initial sample question-and-answer pairs. Specifically, a sample question-and-answer pair refers to a question-and-answer pair about a target object in a sample video.
[0038] In the method provided in the embodiments of this specification, the text recognition model used is a lightweight language model, which only filters the initial sample question-answer pairs of plain text. By understanding the semantic information of the initial sample question-answer pairs, it retains the question-answer pairs that are related to the target object in the sample video and removes the question-answer pairs that are not related to the target object in the sample video.
[0039] For example, if the initial sample question-answer pair is "Question: What is the overall style of the video? Answer: Relaxed and cheerful", and this initial sample question-answer pair is input into the text recognition model, the text recognition model will understand and analyze the initial sample question-answer pair and determine that the question-answer pair does not specifically refer to a particular object in the sample video. Therefore, it can be determined that the initial sample question-answer pair is not suitable for the instance-level video understanding task, and thus the initial sample question-answer pair can be removed.
[0040] For example, the initial sample question-and-answer pair might be: "Question: In the video, a man wearing a green shirt, gray pants, and carrying a white canvas bag enters which store? Answer: Store A." This initial sample question-and-answer pair is input into a text recognition model. The model understands and analyzes this pair, determining that it pertains to a specific object in the sample video. Therefore, it can be determined that this initial sample question-and-answer pair is suitable for instance-level video understanding tasks and can be used as a sample question-and-answer pair.
[0041] After the initial question-and-answer pairs in the initial sample training dataset are filtered by the text recognition model, training data that does not meet the requirements of the instance-level video understanding task can be removed. Only question-and-answer pairs involving specific questions related to a specific target object are retained, and the corresponding sample videos are saved as the training dataset to be processed. Through this filtering step, the data in the general training data can be initially filtered, retaining only the question-and-answer tasks specific to the target object for subsequent processing, thus improving the data quality of the training dataset to be processed.
[0042] Step 104: Perform uniqueness verification on the target object in the sample video based on the sample question-and-answer pair. Based on the verification result, select a reference sample training data set from the sample training data set to be processed. The reference sample training data set includes sample videos. Rewrite the target description information of the sample question-and-answer pair and the target object. The description information of the target object in the rewritten sample question-and-answer pair is replaced with a preset placeholder.
[0043] Specifically, uniqueness verification refers to verifying whether the target object mentioned in the sample question-and-answer pair exists only once in the sample video. In other words, it verifies whether only one target object mentioned in the sample question-and-answer pair exists in the sample video. For example, if the sample question in the sample question-and-answer pair is "What is the person wearing a hat holding in their left hand in the video?", if there are two or more people wearing hats in the sample video, the uniqueness verification result is non-unique. If there is exactly one person wearing a hat in the sample video, the uniqueness verification result is unique.
[0044] Based on the verification results, a reference sample training data set that meets the uniqueness verification requirements can be further selected from the training data set of samples to be processed. That is, in the method provided in the embodiments of this specification, all training data in the reference sample training data set are training data that meet the uniqueness verification requirements.
[0045] The training data in the reference sample training dataset includes sample videos, rewritten sample question-and-answer pairs, and target object description information. In the rewritten sample question-and-answer pairs, the target object description information is replaced with a preset placeholder. Specifically, rewriting a sample question-and-answer pair refers to generating a question-and-answer pair after modifying the original sample question-and-answer pair. In the method provided in the embodiments of this specification, modifying the sample question-and-answer pair specifically means identifying the target object description information in the sample question-and-answer pair and replacing the object description information with a preset placeholder.
[0046] For example, a sample question-and-answer pair might be: "Question: In the video, a man wearing a green shirt, gray pants, and carrying a white canvas bag entered which store? Answer: Store A." The target object's description information is "a man wearing a green shirt, gray pants, and carrying a white canvas bag," which can be achieved using a pre-defined placeholder.<vp\> "Replace the object description information of the target object in the sample question and answer to generate a rewritten sample question and answer pair." Question: In the video<vp\> Which store did you enter? Answer: Store A.
[0047] In this embodiment, the uniqueness of target objects in the sample video is verified based on the sample question-and-answer pair. This aims to ensure the uniqueness of the target object mentioned in the sample question within the sample video, effectively avoiding confusion caused by multiple target objects. Simultaneously, target description information for the target object is generated to provide annotation ranges for pixel-level masks during subsequent processing. Rewriting the object description information in the sample question-and-answer pair according to preset placeholders ensures consistency between the sample question and the target description information, ensuring that they refer to the same target object in subsequent processing.
[0048] With the development of large language model technology, the method provided in the embodiments of this specification can process the training data in the training dataset of the sample to be processed through large language model processing to generate a corresponding reference sample training dataset. Specifically, in a specific embodiment provided in this specification, the uniqueness of the target object in the sample video is verified based on the sample question-and-answer, and the reference sample training dataset is selected from the training dataset of the sample to be processed based on the verification result, including: Input the sample videos and sample question-and-answer pairs into the video filtering model; Based on the video filtering model, the question description information of the target object in the sample question-and-answer pair is obtained. According to the question description information, the number of objects of the target object in the sample video is identified. When the number of objects is one, the target description information of the target object is generated. The positioning time window corresponding to the target object is determined from the sample video. The target description information and the description information of the target object in the sample question-and-answer pair are replaced with preset placeholders.
[0049] In the method provided in the embodiments of this specification, a pre-trained video filtering model can be used to filter the training dataset of samples to be processed. Specifically, each sample video and its corresponding question-and-answer pair in the training dataset of samples to be processed can be input into the video filtering model for uniqueness verification, and the video filtering model outputs the final set of reference sample training data.
[0050] The video selection model preferably uses a pre-trained multimodal large language model, which possesses strong reasoning capabilities. Sample videos and their corresponding question-and-answer pairs are input into the video selection model for processing. The video selection model can perform the following operations on the sample videos and question-and-answer pairs: 1. In this embodiment, the first step of the video filtering model is to verify whether the target object in the sample question-and-answer pair is unique within the sample video. Therefore, the video filtering model first performs semantic analysis on the sample question-and-answer pair to obtain the question description information corresponding to the target object in the sample question-and-answer pair. Specifically, the question description information refers to the description information about the target object in the sample question-and-answer pair. After extracting the question description information, the number of target objects in the sample video can be identified based on this information. The uniqueness verification result is determined based on the number of target objects. If the number of objects is one, it indicates that the target object is unique in the sample video, and further processing can proceed. If the number of objects is at least two, the sample video and its corresponding sample question-and-answer pair are removed. That is, if the number of objects is at least two, it indicates that the target object is not unique in the sample video, and such sample data does not meet the requirements and needs to be removed from the sample data, i.e., the corresponding sample video and its corresponding sample question-and-answer pair are removed.
[0051] 2. Once it is determined that the target object is unique in the sample video, corresponding target description information (Semantic tag) can be generated for the target object in the sample video based on the capabilities of the large language model. This target description information is used to guide the model in subsequent processing so that it can uniquely identify the target object in the sample video.
[0052] 3. In the method provided in the embodiments of this specification, the video screening model also simultaneously determines the positioning time window of the target object from the sample video. The positioning time window can be understood as the time window in which the target object appears in the sample video. For example, if the sample video is 30 seconds long and the target object appears in the sample video from the 3rd to the 7th second, then the time window is "0:03-0:07". In practical applications, the positioning time window specifically refers to the set of times when the target object appears in the sample video. Therefore, the time when the target object appears in the sample video may be discontinuous, and thus, there may be multiple positioning time windows. For example, still taking the sample video as a 30-second example, if the target object appears for the first time from the 3rd to the 7th second, for the second time from the 11th to the 18th second, and for the third time from the 23rd to the 26th second, then there are three positioning time windows, namely "0:03-0:07", "0:11-0:18", and "0:23-0:26".
[0053] 4. Simultaneously, the video filtering model will replace the target description information and the target-related descriptions in the sample question-and-answer pairs with preset placeholders. For example, using the preset placeholder as "<vp\> Using "" as an example for explanation, the target description information is represented by the preset placeholder "<vp\> The replacement indicates that the target description information can be used in subsequent processing.<vp\> "Replace it. At the same time, replace the description information of the target object in the sample question-and-answer pairs with "<vp\> This allows for targeted replacements during subsequent data expansion.
[0054] After the above steps, the video filtering model can identify and filter each sample video and its corresponding question-and-answer pairs. If the sample video and question-and-answer pairs are input into the video filtering model and the model determines that they do not meet the uniqueness requirement, the output result is "False," indicating that the training data does not meet the requirements of the method provided in the embodiments of this specification.
[0055] If the sample video and sample question-and-answer pair are input into the video filtering model for processing, the video filtering model determines that they meet the uniqueness constraint, and the output results are the rewritten sample question-and-answer pair, target description information and positioning time window. Then, based on the corresponding sample video, a reference sample training data set is constructed.
[0056] The steps provided in the embodiments of this specification involve filtering the training data set of samples to be processed, selecting training data that meets the uniqueness verification criteria as the reference sample training data set, and rewriting the sample question-answer pairs and target description information by replacing the target object's description information with preset placeholders. This facilitates improved processing efficiency in subsequent processing. The positioning window is used to locate the segments where the target object appears in subsequent video processing, reducing the amount of data processed and further improving video processing efficiency.
[0057] Step 106: Determine the prompt video frame from each sample video based on the target description information, and mask the target object in the prompt video frame.
[0058] In instance-level video processing tasks, it is necessary to provide multimodal large language models with prompt video frames corresponding to the target objects. Prompt video frames in instance-level video processing tasks are key frames that provide the model with clear visual references, specify the processing target and initialization state. Their core function is to accurately guide the model to locate, identify, segment or track specific object instances in the video, and ensure temporal consistency.
[0059] In practical applications, simply inputting target description information used to represent semantic information into a multimodal large language model for processing is insufficient to support pixel-level visual cue rendering. The method provided in the embodiments of this specification also requires text-driven video diffusion segmentation based on the target description information. Therefore, in the method provided in the embodiments of this specification, the cue video frames corresponding to each sample video are extracted from each sample video according to each target description information, and the target objects in the cue video frames are masked.
[0060] A mask can be understood as a black and white mask image, where the white areas represent the target object to be selected or processed, and the black areas represent the background that does not need to be processed. By using a binary mask, the target object in the prompt video frame is isolated pixel by pixel from the background and other distracting objects, allowing the model to focus on the masked area and ignore irrelevant pixels.
[0061] In one specific embodiment provided in this specification, a prompt video frame is determined from each sample video based on target description information, and the target object in the prompt video frame is masked, including: Input the sample video, the corresponding localization time window of the sample video, and the target description information into the video segmentation model; Based on the video segmentation model, the video segment corresponding to the target object is determined in the sample video according to the positioning time window; the video frame to be processed is extracted from the video segment, and the target object is identified in the video frame to be processed according to the target description information. The target object is then masked at the pixel level to generate a prompt video frame.
[0062] In one specific embodiment provided in this specification, the same processing operation is performed on each sample video. To further clarify the explanation, this embodiment uses one sample video as an example for explanation, and the same processing method can be used for all other sample videos.
[0063] In this embodiment, the sample video, the corresponding localization time window, and the target description information are input into the video segmentation model. The video segmentation model can be understood as a pre-trained basic model used for video segmentation. After the sample video, localization time window, and target description information are input into the video segmentation model, the model can perform corresponding processing.
[0064] Specifically, the video segmentation model can first locate video segments in the sample video where the target object appears based on the localization time window. Then, it extracts multiple initial video frames from the video segments according to a preset sampling frequency, and determines the video frames to be processed from these initial frames. Finally, based on the target description information, it identifies and recognizes the target object from the video frames to be processed, and performs pixel-level masking on the target object.
[0065] In practical applications, multiple initial video frames are extracted from a video segment based on a preset sampling frequency. The video frame to be processed is then determined from these initial frames. The goal is to extract a high-quality video frame for better processing results in subsequent video processing. There are many ways to determine the video frame to be processed from the multiple initial frames; for example, it could be the first initial video frame, a user-defined video keyframe, or an initial video frame whose video quality weight meets a preset threshold, etc.
[0066] In the methods provided in the embodiments of this specification, Segment Anything Model 3 (SAM3) is used as an example for explanation. SAM3 is a general segmentation basic model. Its core breakthrough is the upgrade from single-instance interaction to concept-level open vocabulary segmentation, and it unifies image segmentation, video tracking, and multi-instance detection capabilities. When sample video, localization time window, and target description information are input into the SAM3 model, the SAM3 model can extract the video frame to be processed from the sample video according to the localization time window, determine the target object in the video frame to be processed according to the target description information, mask the target object, and generate a masked prompt video frame.
[0067] By applying the same processing method to all sample videos, corresponding prompt video frames can be generated for each sample video. In this embodiment, during the masking process for the target object, the localization time window of the sample video is simultaneously input, which can help the video segmentation model quickly locate the video segment where the target object appears, thereby reducing the processing time of the video segmentation model and improving its processing speed.
[0068] Step 108: Mark the target objects in the prompt video frame according to at least two preset markers, generate sample prompt video frames, and correct the preset placeholders in the rewritten sample question-and-answer pair according to the marking results, and generate target sample question-and-answer pairs.
[0069] In this step, video cues are rendered and generated within the cue video frames, and these cues are aligned with the target description information. In practical applications, during instance-level video processing tasks, video frames with added cues are used to guide the target object. Furthermore, the cues need to be aligned with the target object's description information so that the video processing model can bind text information to the marker information in the image.
[0070] Based on this, in this embodiment, different preset tags can be used to mark the target objects in the prompt video frames, thereby generating sample prompt video frames corresponding to different preset tags. Simultaneously, the preset placeholders in the sample question-and-answer pairs can be corrected based on the marking results, thereby generating target sample question-and-answer pairs corresponding to different preset tags.
[0071] In one specific embodiment provided in this specification, target objects in the prompt video frame are marked according to at least two preset markers to generate sample prompt video frames, including: A target preset tag is determined from a preset tag library, wherein the preset tag library includes at least two preset tags; The target object in the prompt video frame is marked based on the preset target marker.
[0072] In practical applications, different users may have different video labeling methods. In order to enhance the diversity of training data, the target objects in the prompt video frames can be labeled according to multiple preset labels, thereby generating corresponding sample prompt video frames.
[0073] In this embodiment, a target preset marker can be selected from a pre-set preset marker library, and the target object in the prompt video frame can be marked according to the target preset marker. At this time, the target object in the prompt video frame has already been masked in the above steps. The preset marker library stores at least two preset markers, which can be rectangles, outlines of target objects, ellipses, triangles, doodles, dots, arrows, numbered markers, etc. To accommodate different user habits, different preset markers are needed to mark the target objects in the prompt video frames to generate sample prompt video frames.
[0074] In another specific embodiment provided in this specification, the preset placeholders in the rewritten sample question-and-answer pair are corrected according to the marking results to generate the target sample question-and-answer pair, including: Generate tag description information corresponding to the target object based on the tagging results of the target preset tag; Update the preset placeholders in the rewritten sample question-and-answer pair according to the labeled description information to generate the target sample question-and-answer pair.
[0075] In this embodiment, after marking the target object according to the target preset tag, a tag description information corresponding to the target object can be generated based on the marking result of the target preset tag. This tag description information specifically refers to natural language description information. The preset placeholders in the rewritten sample question-and-answer pair are corrected based on this natural language description information, thereby generating a target sample question-and-answer pair corresponding to the preset tag.
[0076] For example, taking a rectangle as the preset target marker, marking the target object in the prompt video frame with a rectangle can generate a corresponding natural language description of "the object highlighted in the rectangle"; as another example, taking an arrow as the preset target marker, marking the target object in the prompt video frame with an arrow can generate a corresponding natural language description of "the object pointed to by the arrow".
[0077] In the above steps, after rewriting the sample question-and-answer pairs, rewritten question-and-answer pairs are generated. Continuing with the previous example, let the rewritten question-and-answer pair be "Question: In the video..."<vp\> Taking the example of "Which store did the object enter? Answer: Store A," we can further explain that if we mark the target object in the prompt video frame with a rectangle, we can generate its corresponding natural language description as "the object in the highlighted rectangle." This natural language description is then used to correct the preset placeholders in the rewritten question-and-answer pair, generating the target sample question-and-answer pair as "Question: Which store did the object in the highlighted rectangle in the video enter? Answer: Store A."
[0078] At this point, the training data for training the video question answering model is ready and can be used for subsequent model training.
[0079] Step 110: Construct a target sample training dataset based on the sample video, target sample question-and-answer pairs, and sample prompt video frames, and train the video question-and-answer model based on the target sample training dataset until the model training stops.
[0080] In this embodiment, after the above processing, the training data in the training dataset to be processed, after data filtering and preprocessing, generate training data consisting of sample videos, target sample question-and-answer pairs, and sample video prompt frames. This data is used for subsequent training of the video question-and-answer model.
[0081] Specifically, a target sample training dataset can be constructed based on sample videos, target sample question-and-answer pairs, and sample prompt video frames, and the video question-and-answer model can be trained based on this target sample training dataset.
[0082] In practical applications, initial video question-answering models based on a base model have limited capabilities, exhibiting poor instruction compliance and a high error rate in tool invocation. Directly applying reinforcement learning leads to low learning efficiency and introduces significant training noise. Therefore, the method provided in the embodiments of this specification filters the training data used for training the video question-answering model to generate cold-start data for training the model.
[0083] Specifically, in one embodiment provided in this specification, a target sample training dataset is constructed based on sample videos, target sample question-and-answer pairs, and sample prompt video frames, including: The sample video, the sample questions in the target sample question-answer pair, and the sample prompt video frames are input into the multimodal video language model to obtain the predicted answer output by the multimodal video language model. The predicted answer is generated by the multimodal video language model through a multi-round video tool call. If the predicted answer matches the sample answer in the target sample question-answer pair, the sample video, the target sample question-answer pair, and the sample prompt video frame are added to the target sample training dataset.
[0084] In this implementation, a rejection sampling strategy is employed to generate high-quality multi-turn tool call trajectories. Specifically, using the training data generated in the preceding steps, a strong inference model interacts with the environment to execute multiple rounds of video tool calls. The multi-turn tool call trajectory refers to the order in which the training data guides the multimodal video language model to correctly invoke video tools.
[0085] Among them, the rejection sampling strategy is a basic Monte Carlo sampling method. Its core function is to generate candidate samples by means of a simple proposal distribution when it is not possible to sample directly from a complex target distribution, and then filter out samples that conform to the target distribution through an accept or reject mechanism.
[0086] Specifically, in this embodiment, the sample video, sample questions from the target sample question-and-answer pair, and sample prompt video frames are input into the multimodal video language model. The multimodal video language model interacts with the environment (this interaction includes both simulated user interaction with the multimodal video language model and interaction where the multimodal video language model calls video tools). The multimodal video language model calls video frame extraction tools and object cropping tools to view the video frames of the sample video to obtain the spatiotemporal information of the sample video, and generates a predicted answer corresponding to the sample question based on the spatiotemporal information. Then, a discriminator compares this predicted answer with the sample answers in the target sample question-and-answer pair to filter out the correct answer trajectory.
[0087] In practical applications, if the predicted answer matches the sample answer in the target sample question-and-answer pair, it indicates that the training data has passed the correct answer trajectory. Therefore, if the predicted answer matches the sample answer in the target sample question-and-answer pair, the sample video, the target sample question-and-answer pair, and the sample prompt video frame can be added to the target sample training dataset. This method allows for the selection of a batch of high-quality trajectory datasets for cold start training.
[0088] In another specific embodiment provided in this specification, when the predicted answer matches the sample answer in the target sample question-answer pair, the sample video, the target sample question-answer pair, and the sample prompt video frame are added to the target sample training dataset, including: If the predicted answer matches the sample answer in the target sample question-answer pair, verify the question-answer difficulty weight of the target sample question-answer pair; Add sample videos, target sample question-and-answer pairs, and sample prompt video frames that meet the preset weight threshold for question-and-answer difficulty to the target sample training dataset.
[0089] In practical applications, training multimodal large language models requires reinforcement learning. The policy gradient in reinforcement learning is proportional to the variance of the dominance function. When the dominance value is constant, a small variance leads to gradient collapse. In other words, reinforcement learning updates parameters based on the variance of the dominance function. If the dominance values of all samples are similar, the gradient will collapse and fail, causing the model to be unable to learn.
[0090] To address this issue, the method provided in the embodiments of this specification employs Pass@K for data filtering. Pass@K is a strategy for determining the difficulty level of a problem. Simply put, Pass@K means that for the same sample problem or task, the model generates K answers, and if any one of these answers is correct, the sample is considered passed. Psss@K refers to the pass rate for this type of sample. When Psss@K is 0, it indicates that all K answers are wrong, meaning the problem is too difficult and the model's generation will be incorrect, making it impossible to learn. When Psss@K is approximately 1, it indicates that any K generated answers will be correct, meaning the problem is too easy and the model can answer correctly regardless, offering no learning value. Typically, data in the 0.4-0.8 range is considered of medium difficulty; the model's responses are sometimes correct and sometimes incorrect, showing a degree of understanding but not complete mastery, thus providing valuable learning information.
[0091] In the method provided in the embodiments of this specification, the question-answer difficulty weight of the target sample question-answer pair is calculated using the Pass@K strategy. Sample videos, target sample question-answer pairs, and sample prompt video frames that meet a preset weight threshold are added to the target sample training dataset. Within the preset weight threshold, the advantage function of the data exhibits significant variance, ensuring that the dynamic direction of the action space distribution shift points towards the comparison of potentially learnable strategies.
[0092] After the above processing, a high-quality training dataset with moderate difficulty and correct answer trajectories was constructed. The video question-answering model was then trained based on this dataset in subsequent processing.
[0093] To address the problem of single-stage video compression in existing technologies, the method provided in the embodiments of this specification constructs a multi-round interactive visual language model inference framework. This framework is specifically designed for handling instance-level video understanding tasks, and its core lies in enabling the model to actively interact with the environment (such as video and visual cues) through tool calls, gradually resolving complex problems.
[0094] In the inference framework provided in the embodiments of this specification, the video and the user's question are first transformed into an initial context that the model can process. Simultaneously, the video tool is introduced into the inference framework, and tool usage rules are injected. Secondly, the model gradually narrows the scope of the question by "calling tools to obtain more information" until an answer is provided or a preset maximum number of interaction rounds is reached. Finally, the model uniformly returns the answer, interaction trajectory, and tool call history for subsequent debugging and evaluation.
[0095] Specifically, in one embodiment provided in this specification, training a video question-answering model based on the target sample training data set until the model training stops includes: The video question answering model is trained in a supervised manner based on the target sample training data set until the first training stopping condition is met. The video question answering model is trained using reinforcement learning based on the target sample training data set until the second training stopping condition is met.
[0096] In the method provided in the embodiments of this specification, an Agentic RL training scheme for multimodal VLM (Vision-Language Model) is designed based on the training framework of open source large language model reinforcement learning. In this method, a dedicated video tool environment is designed and loaded into the training framework of multimodal large language model.
[0097] The video tool environment includes the `view_visual_prompt` and `crop_video` tools. The `view_visual_prompt` tool returns raw video frame images with visual cues based on frame paths or frame indices. The `crop_video` tool crops a specified frame sequence from the raw video frames based on a location time interval, supporting optional focus cues to indicate the selected cropping region. The video tool environment is deeply integrated with the rollout pipeline in the reinforcement learning training framework, supporting multi-GPU parallel rollouts and asynchronous reward computation.
[0098] During model training, the video question-answering model is first trained using the target sample training set, known as the SFT (Supervised Fine-Tuning) stage. In this stage, the target sample training dataset serves as a model of standard answers, allowing the model to learn by imitation and adaptation. This teaches the video question-answering model the format, wording, basic tasks, and tool usage timing, thus equipping the model with preliminary video processing capabilities.
[0099] Specifically, supervised training of the video question-answering model is performed based on the target sample training data set, including: The sample video, the sample question in the target sample question-answer pair, and the sample prompt video frame are input into the video question answering model to obtain the predicted answer output by the video question answering model. The predicted answer is generated by the video question answering model through a multi-round video tool call. The model loss value is calculated based on the predicted answer and the sample answer in the target sample question-answer pair; The model parameters of the video question answering model are adjusted based on the model loss value.
[0100] In this embodiment, the sample videos, sample questions in the target sample question-answer pairs, and sample prompt video frames from the target sample training dataset are input into the video question-answering model for processing. The video question-answering model is based on a reinforcement learning training framework and uses multiple rounds of interaction to call tools in the video tool environment to generate predicted answers.
[0101] The model loss value is calculated based on the sample answers in the predicted answer and the target sample question-answer pair, and the model parameters of the video question-answering model are adjusted based on the model loss value.
[0102] Repeat the above operations until the first training stopping condition is met. In the method provided in the embodiments of this specification, the first training stopping condition may be that the model loss value is less than a preset loss value threshold, or that the model has reached a preset number of training epochs, or both, etc. In the method provided in the embodiments of this specification, the first training stopping condition is not specifically limited, and the actual application shall prevail.
[0103] After the SFT training phase, the model enters the RL reinforcement learning phase. In this phase, no standard answer is provided; only reward scores are awarded. Good performance earns points, and poor performance deducts points, allowing the model to iterate, evolve optimally, and learn complex reasoning, tool invocation, and multi-round decision-making. In the method provided in the embodiments of this specification, the tool invocation capability is also treated as a reinforcementable ability through a multi-round video tool invocation path during the reinforcement learning phase. This allows the model to learn the ability to proactively use video tools for processing.
[0104] Specifically, in the method provided in the embodiments of this specification, the GRPO (Group Relative Policy Optimization) algorithm is used to train the policy model. GRPO avoids dependence on the value network by calculating the relative advantage within the group, significantly reducing training complexity and improving stability in multimodal scenarios. During the training process, reinforcement learning is divided into three stages: rollout-reward-update. Rollout is the inference stage, relying on SGlang and vLLM to run model inference; reward is the reward scoring stage, relying on a custom reward function for scoring; and update is the parameter update stage, relying on FSDP2 for multi-GPU gradient updates.
[0105] During the rollout phase, the current policy model and the video tools in the video tool environment complete the entire process. Using the SGlang and vLLM high-speed inference engines, batch generation is achieved with high concurrency and minimal GPU memory usage, resulting in high speed. This is used to generate trajectories in large batches in parallel, improving the speed of RL training. In this phase, the high-speed inference engine allows the video question-answering model to automatically call the corresponding tools to provide answers, generating complete interaction trajectories.
[0106] In the reward phase, a custom reward function is used to evaluate accuracy, format, and tool usage. After generating the trajectory answers in the rollout phase, no manual intervention is required. The accuracy of the answers and the conformity of the tool call and output formats are judged by a preset custom reward function, thereby calculating a reward score for each trajectory. The higher the score, the better the performance. In other words, the model's inference trajectories are scored through the reward phase.
[0107] Once we have the trajectory and the score for each trajectory, we can use distributed parallel training technology to compute the policy gradient in parallel on multiple GPUs. We can then use the optimizer in the training framework to update the model weights so that the model can perform better next time. In other words, during the update phase, we can backpropagate and update the video question answering model based on the reward score to complete the learning and evolution.
[0108] Based on this, an Agentic reinforcement learning paradigm for multimodal video processing models was designed, treating tool invocation capability as a learnable reinforcement skill. Through the GRPO algorithm and reward function, the model's ability to proactively use video tools for evidence collection is stimulated. This training method overcomes the limitation of existing reinforcement learning methods that rely solely on textual reasoning chains.
[0109] Unlike reinforcement learning methods that only optimize text output, the method provided in this specification incorporates video tool invocation as part of reinforcement learning, including the timing, location, and parameter selection of tool invocation within the scope of policy optimization. The GRPO algorithm employs a structured design of thinking, tool recall, and response, making tool invocation and answer output discrete and verifiable actions. This effectively prevents multimodal large language models from boosting rewards by generating lengthy or ambiguous text.
[0110] In addition, the method provided in the embodiments of this specification also includes a multi-dimensional reward signal consisting of accuracy rewards, format rewards, and optional tool usage rewards, which guides the model to learn good tool usage habits.
[0111] The method provided in the embodiments of this specification designs a fully automated visual cue training data construction method that requires no manual annotation, efficiently transforming any video and question-answer pair dataset into instance-level video understanding data that relies on visual cues. This method addresses the industry pain point of training data shortage in instance-level video understanding tasks, ensuring the quality and diversity of training data.
[0112] Secondly, the method provided in the embodiments of this specification introduces tool invocation into instance-level video understanding tasks. By coordinating multiple video tools through a video processing model for multi-round, iterative evidence retrieval and understanding, this approach decouples evidence collection from reasoning. Dynamically capturing video frames within specific time intervals by invoking video tools on demand avoids the loss of key clues caused by global video compression. Invoking video tools ensures the model continuously accesses visual cue frames marked with the target instance, thus maintaining awareness of the target object throughout the reasoning process. The multi-round interaction mechanism also allows the model to gradually accumulate visual understanding of the target object at different times, supporting complex temporal behavioral reasoning and multi-hop relationship reasoning. The model also records the tool invocation history, which fully documents the model's reasoning process; any intermediate decisions can be traced back, enhancing the credibility of the reasoning results.
[0113] See Figure 2 , Figure 2 The diagram illustrates a flowchart of a video question-and-answer method according to an embodiment of this specification, as shown below. Figure 2 As shown, the method includes: Step 202: Obtain the video to be processed and the target problem corresponding to the video to be processed.
[0114] Step 204: Input the video to be processed and the target question into the video question answering model to obtain the target answer output by the video question answering model. The target answer is generated by the video question answering model through a multi-round video tool call. The video question answering model is trained using the training method of the video processing model described above.
[0115] In this embodiment, the video question-answering model trained in the above steps is used to perform instance-level video understanding tasks. Specifically, the video to be processed and the corresponding target question are obtained. The target question refers to a question about a specific object in the video to be processed. The video to be processed and the target question are input into the video question-answering model. The video question-answering model calls the video processing tool through multiple rounds of data interaction. First, it identifies the target object from the video to be processed. Then, it calls the video processing tool to extract video segments in the video to be processed that contain the target object. Finally, it analyzes the target object from the video segments to generate the target answer corresponding to the target question.
[0116] The video question-answering model mentioned in the embodiments of this specification specifically refers to a multimodal large language model trained by the training device of the above-mentioned video processing model.
[0117] Corresponding to the above method embodiments, this specification also provides embodiments of a training device for a video processing model. Figure 3A schematic diagram of a training apparatus for a video processing model according to an embodiment of this specification is shown. Figure 3 As shown, the device includes: The acquisition module 302 is configured to acquire a training data set of samples to be processed, wherein the training data set of samples to be processed includes at least one sample video and sample question-answer pairs corresponding to each sample video; The verification module 304 is configured to perform uniqueness verification on the target object in the sample video based on the sample question-and-answer pair, and select a reference sample training data set from the sample training data set to be processed based on the verification result. The reference sample training data set includes the sample video, rewrites the sample question-and-answer pair and the target description information of the target object, and replaces the description information of the target object in the rewritten sample question-and-answer pair with a preset placeholder. The masking module 306 is configured to determine the prompt video frame from each sample video based on each target description information, and to mask the target object in the prompt video frame; The rewriting module 308 is configured to mark the target objects in the prompt video frame according to at least two preset tags, generate sample prompt video frames, and modify the preset placeholders in the rewritten sample question-and-answer pair according to the marking results, thereby generating the target sample question-and-answer pair. The training module 310 is configured to construct a target sample training dataset based on sample videos, target sample question-and-answer pairs, and sample prompt video frames, and to train a video question-and-answer model based on the target sample training dataset until the model training stops.
[0118] In one specific embodiment provided in this specification, the verification module 304 is further configured as follows: Input the sample videos and sample question-and-answer pairs into the video filtering model; Based on the video filtering model, the question description information of the target object in the sample question-and-answer pair is obtained. According to the question description information, the number of objects of the target object in the sample video is identified. When the number of objects is one, the target description information of the target object is generated. The positioning time window corresponding to the target object is determined from the sample video. The target description information and the description information of the target object in the sample question-and-answer pair are replaced with preset placeholders.
[0119] In one specific embodiment provided in this specification, the verification module 304 is further configured to: If there are at least two objects, remove the sample video and the sample question-and-answer pair corresponding to the sample video.
[0120] In one specific embodiment provided in this specification, the mask module 306 is further configured as follows: Input the sample video, the corresponding localization time window of the sample video, and the target description information into the video segmentation model; Based on the video segmentation model, the video segment corresponding to the target object is determined in the sample video according to the positioning time window; the video frame to be processed is extracted from the video segment, and the target object is identified in the video frame to be processed according to the target description information. The target object is then masked at the pixel level to generate a prompt video frame.
[0121] In one specific embodiment provided in this specification, the rewriting module 308 is further configured to: A target preset tag is determined from a preset tag library, wherein the preset tag library includes at least two preset tags; The target object in the prompt video frame is marked based on the preset target marker.
[0122] In one specific embodiment provided in this specification, the rewriting module 308 is further configured to: Generate tag description information corresponding to the target object based on the tagging results of the target preset tag; Update the preset placeholders in the rewritten sample question-and-answer pair according to the labeled description information to generate the target sample question-and-answer pair.
[0123] In one specific embodiment provided in this specification, the rewriting module 308 is further configured to: The sample video, the sample questions in the target sample question-answer pair, and the sample prompt video frames are input into the multimodal video language model to obtain the predicted answer output by the multimodal video language model. The predicted answer is generated by the multimodal video language model through a multi-round video tool call. If the predicted answer matches the sample answer in the target sample question-answer pair, the sample video, the target sample question-answer pair, and the sample prompt video frame are added to the target sample training dataset.
[0124] In one specific embodiment provided in this specification, the rewriting module 308 is further configured to: If the predicted answer matches the sample answer in the target sample question-answer pair, verify the question-answer difficulty weight of the target sample question-answer pair; Add sample videos, target sample question-and-answer pairs, and sample prompt video frames that meet the preset weight threshold for question-and-answer difficulty to the target sample training dataset.
[0125] In one specific embodiment provided in this specification, the training module 310 is further configured as follows: The video question answering model is trained in a supervised manner based on the target sample training data set until the first training stopping condition is met. The video question answering model is trained using reinforcement learning based on the target sample training data set until the second training stopping condition is met.
[0126] In one specific embodiment provided in this specification, the training module 310 is further configured as follows: The sample video, the sample question in the target sample question-answer pair, and the sample prompt video frame are input into the video question answering model to obtain the predicted answer output by the video question answering model. The predicted answer is generated by the video question answering model through a multi-round video tool call. The model loss value is calculated based on the predicted answer and the sample answer in the target sample question-answer pair; The model parameters of the video question answering model are adjusted based on the model loss value.
[0127] In one specific embodiment provided in this specification, the acquisition module 302 is further configured to: Obtain an initial sample training data set, which includes sample videos and corresponding initial sample question-answer pairs; Each initial sample question-and-answer pair is input into the text recognition model to obtain the sample question-and-answer pairs output by the text recognition model, wherein the sample question-and-answer pairs are question-and-answer pairs for the target object in the sample video; The sample question-and-answer pairs and the corresponding sample videos are used as the training data set of the samples to be processed.
[0128] The apparatus provided in the embodiments of this specification features a fully automated method for constructing visual cue training data without manual annotation, efficiently converting arbitrary video and question-answer pair datasets into instance-level video understanding data that relies on visual cues. This method addresses the industry pain point of training data shortage in instance-level video understanding tasks, ensuring the quality and diversity of training data.
[0129] Secondly, the apparatus provided in the embodiments of this specification introduces a tool-calling approach to instance-level video understanding tasks. By coordinating multiple video tools through a video processing model for multi-round, iterative evidence retrieval and understanding, this approach decouples evidence collection from reasoning. Dynamically capturing video frames within specific time intervals by calling video tools on demand avoids the loss of key clues caused by global video compression. Calling video tools ensures the model continuously accesses visual cue frames marked with the target instance, thus maintaining awareness of the target object throughout the reasoning process. The multi-round interaction mechanism also allows the model to gradually accumulate visual understanding of the target object at different times, supporting complex temporal behavioral reasoning and multi-hop relationship reasoning. The model also records the tool-calling history, which fully records the model's reasoning process; any intermediate decisions can be traced back, enhancing the credibility of the reasoning results.
[0130] The above is an illustrative scheme of a training device for a video processing model according to this embodiment. It should be noted that the technical solution of this training device for the video processing model and the technical solution of the video processing model training method described above belong to the same concept. Details not described in detail in the technical solution of the video processing model training device can be found in the description of the technical solution of the video processing model training method described above.
[0131] See Figure 4 , Figure 4 This specification shows an architecture diagram of a training system for a video processing model according to an embodiment of the present specification. The training system for the video processing model may include a client 100 and a server 200. Client 100 is used to send a set of training data for samples to be processed to server 200, wherein the set of training data for samples to be processed includes at least one sample video and sample question-answer pairs corresponding to each sample video; Server 200 is used to perform uniqueness verification on target objects in sample videos based on sample question-and-answer pairs. Based on the verification results, it selects a reference sample training data set from the unprocessed sample training data set. The reference sample training data set includes sample videos. It rewrites the target description information of the sample question-and-answer pairs and the target objects, replacing the target object description information in the rewritten sample question-and-answer pairs with preset placeholders. It determines prompt video frames from each sample video based on each target description and masks the target objects in the prompt video frames. It marks the target objects in the prompt video frames according to at least two preset tags, generating sample prompt video frames. Based on the marking results, it corrects the preset placeholders in the rewritten sample question-and-answer pairs, generating target sample question-and-answer pairs. It constructs a target sample training data set based on the sample videos, target sample question-and-answer pairs, and sample prompt video frames, and trains the video question-and-answer model based on the target sample training data set until the model training stops. Finally, it sends the model parameters of the video question-and-answer model to client 100. Client 100 is also used to receive model parameters sent by server 200 and generate a video question-and-answer model based on the model parameters.
[0132] The training system for the video processing model can include multiple clients 100 and a server 200. Clients 100 can be referred to as edge devices, and the server 200 can be referred to as cloud devices. Multiple clients 100 can establish communication connections through the server 200. In the training scenario of the video processing model, the server 200 is used to provide training services for the video processing model among the multiple clients 100. Each client 100 can act as a sender or receiver, communicating through the server 200.
[0133] Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the training scenario of the video processing model, users can publish data streams to server 200 through client 100, server 200 can generate model parameters of the video question answering model based on the data stream, and push the model parameters of the video question answering model to other clients that have established communication.
[0134] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.
[0135] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on a computing device and depends on the device or certain apps on the device to run. The computing device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.
[0136] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0137] It is worth noting that the training method for the video processing model provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the training method for the video processing model provided in the embodiments of this specification. In other embodiments, the training method for the video processing model provided in the embodiments of this specification may also be executed jointly by the client and the server.
[0138] Figure 5A structural block diagram of a computing device 500 according to an embodiment of this specification is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0139] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0140] In one embodiment of this specification, the above-described components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0141] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.
[0142] The processor 520 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-mentioned video processing model training method and video question answering method.
[0143] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solutions of the video processing model training method and the video question answering method described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the video processing model training method and the video question answering method described above.
[0144] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described video processing model training method and video question-answering method.
[0145] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the video processing model training method and the video question answering method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the video processing model training method and the video question answering method described above.
[0146] An embodiment of this specification also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the above-described video processing model training method and video question answering method.
[0147] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solutions of the video processing model training method and the video question answering method described above. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solutions of the video processing model training method and the video question answering method described above.
[0148] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0149] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0150] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.
[0151] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0152] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. These embodiments have been selected and specifically described in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for training a video processing model, comprising: include: Obtain a training data set of samples to be processed, wherein the training data set of samples to be processed includes at least one sample video and sample question-answer pairs corresponding to each sample video; The uniqueness of the target object in the sample video is verified based on the sample question-and-answer pair. Based on the verification result, a reference sample training data set is selected from the sample training data set to be processed. The reference sample training data set includes sample videos. The target description information of the target object in the sample question-and-answer pair is rewritten and replaced with a preset placeholder. Based on the target description information, the prompt video frames are determined from each sample video, and the target objects in the prompt video frames are masked. The target objects in the prompt video frame are marked according to at least two preset tags to generate sample prompt video frames. The preset placeholders in the rewritten sample question and answer pair are corrected according to the marking results to generate target sample question and answer pairs. A target sample training dataset is constructed based on sample videos, target sample question-and-answer pairs, and sample prompt video frames. The video question-and-answer model is then trained based on the target sample training dataset until the model training stops.
2. The method as described in claim 1, characterized in that, Based on the sample question-and-answer session, the uniqueness of the target object in the sample video is verified. Based on the verification results, a reference sample training dataset is selected from the unprocessed sample training dataset, including: Input the sample videos and sample question-and-answer pairs into the video filtering model; Based on the video filtering model, the question description information of the target object in the sample question-and-answer pair is obtained. According to the question description information, the number of objects of the target object in the sample video is identified. When the number of objects is one, the target description information of the target object is generated. The positioning time window corresponding to the target object is determined from the sample video. The target description information and the description information of the target object in the sample question-and-answer pair are replaced with preset placeholders.
3. The method as described in claim 2, characterized in that, The method further includes: If there are at least two objects, remove the sample video and the sample question-and-answer pair corresponding to the sample video.
4. The method as described in claim 1, characterized in that, Based on the target description information, prompt video frames are determined from each sample video, and the target objects in the prompt video frames are masked, including: Input the sample video, the corresponding localization time window of the sample video, and the target description information into the video segmentation model; Based on the video segmentation model, the video segment corresponding to the target object is determined in the sample video according to the positioning time window; the video frame to be processed is extracted from the video segment, and the target object is identified in the video frame to be processed according to the target description information. The target object is then masked at the pixel level to generate a prompt video frame.
5. The method as described in claim 1, characterized in that, Based on at least two preset markers, the target objects in the prompt video frames are marked respectively, and sample prompt video frames are generated, including: A target preset tag is determined from a preset tag library, wherein the preset tag library includes at least two preset tags; The target object in the prompt video frame is marked based on the preset target marker.
6. The method as described in claim 5, characterized in that, Based on the labeling results, the preset placeholders in the rewritten sample question-and-answer pairs are corrected to generate the target sample question-and-answer pairs, including: Generate tag description information corresponding to the target object based on the tagging results of the target preset tag; Update the preset placeholders in the rewritten sample question-and-answer pair according to the labeled description information to generate the target sample question-and-answer pair.
7. The method as described in claim 1, characterized in that, A target sample training dataset is constructed based on sample videos, target sample question-answer pairs, and sample prompt video frames, including: The sample video, the sample questions in the target sample question-answer pair, and the sample prompt video frames are input into the multimodal video language model to obtain the predicted answer output by the multimodal video language model. The predicted answer is generated by the multimodal video language model through a multi-round video tool call. If the predicted answer matches the sample answer in the target sample question-answer pair, the sample video, the target sample question-answer pair, and the sample prompt video frame are added to the target sample training dataset.
8. The method as described in claim 7, characterized in that, If the predicted answer matches the sample answer in the target sample question-answer pair, the sample video, the target sample question-answer pair, and the sample prompt video frame are added to the target sample training dataset, including: If the predicted answer matches the sample answer in the target sample question-answer pair, verify the question-answer difficulty weight of the target sample question-answer pair; Add sample videos, target sample question-and-answer pairs, and sample prompt video frames that meet the preset weight threshold for question-and-answer difficulty to the target sample training dataset.
9. The method as described in claim 1, characterized in that, The video question-answering model is trained based on the target sample training data set until the model training stops, including: The video question answering model is trained in a supervised manner based on the target sample training data set until the first training stopping condition is met. The video question answering model is trained using reinforcement learning based on the target sample training data set until the second training stopping condition is met.
10. The method as described in claim 9, characterized in that, Supervised training of the video question-answering model is performed based on the target sample training dataset, including: The sample video, the sample question in the target sample question-answer pair, and the sample prompt video frame are input into the video question answering model to obtain the predicted answer output by the video question answering model. The predicted answer is generated by the video question answering model through a multi-round video tool call. The model loss value is calculated based on the predicted answer and the sample answer in the target sample question-answer pair; The model parameters of the video question answering model are adjusted based on the model loss value.
11. The method as described in claim 1, characterized in that, Obtain the training data set of the samples to be processed, including: Obtain an initial sample training data set, which includes sample videos and corresponding initial sample question-answer pairs; Each initial sample question-and-answer pair is input into the text recognition model to obtain the sample question-and-answer pairs output by the text recognition model, wherein the sample question-and-answer pairs are question-and-answer pairs for the target object in the sample video; The sample question-and-answer pairs and the corresponding sample videos are used as the training data set of the samples to be processed.
12. A video question-answering method, characterized in that, include: Obtain the video to be processed and the target problem corresponding to the video to be processed; The video to be processed and the target question are input into the video question answering model to obtain the target answer output by the video question answering model. The target answer is generated by the video question answering model through a multi-round video tool call. The video question answering model is trained by any one of the training methods of claims 1-11.
13. A training device for a video processing model, characterized in that, include: The acquisition module is configured to acquire a set of training data for samples to be processed, wherein the set of training data for samples to be processed includes at least one sample video and sample question-answer pairs corresponding to each sample video; The verification module is configured to perform uniqueness verification on the target object in the sample video based on the sample question-and-answer pair, and select a reference sample training data set from the sample training data set to be processed based on the verification result. The reference sample training data set includes the sample video, rewrites the sample question-and-answer pair and the target description information of the target object, and replaces the description information of the target object in the rewritten sample question-and-answer pair with a preset placeholder. The masking module is configured to determine the prompt video frame from each sample video based on the target description information, and to mask the target object in the prompt video frame; The rewriting module is configured to mark the target objects in the prompt video frames according to at least two preset tags, generate sample prompt video frames, and modify the preset placeholders in the rewritten sample question-and-answer pairs according to the marking results, thereby generating target sample question-and-answer pairs. The training module is configured to construct a target sample training dataset based on sample videos, target sample question-and-answer pairs, and sample prompt video frames, and to train a video question-and-answer model based on the target sample training dataset until the model training stops.
14. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 12.
15. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 12.
16. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 12.