A method, apparatus, storage medium, and electronic device for scene temporal location detection based on multimodal understanding.
By segmenting shot fragments in multimodal videos and constructing a training set, and combining supervised fine-tuning and reinforcement learning to optimize the model, the problems of low accuracy and poor efficiency in label extraction in multimodal video scenarios are solved, achieving high-precision label generation and model stability.
Patent Information
- Application Number
- CN202510998908.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Existing technologies suffer from low accuracy and efficiency in label extraction in multimodal video scenarios, and the lack of stability in reinforcement learning training leads to poor model iteration performance.
By segmenting the video into shot fragments, labeling timecode information, and constructing a training set, and combining text features, visual features, and prompt text to generate input feature sequences, supervised fine-tuning and reinforcement learning are used to optimize the model, and a GRPO reward function is designed to improve model performance.
It improves the accuracy and efficiency of tag extraction in multimodal video scenarios, ensures the semantic integrity and format standardization of model output, and enhances spatiotemporal modeling capabilities.
Smart Images

Figure CN120510555B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of multimodal video understanding and deep learning technology, and more specifically, to a method, apparatus, storage medium, and electronic device for scene temporal location detection based on multimodal understanding. Background Technology
[0002] With the rapid development of artificial intelligence technology, the field of media intelligence is facing unprecedented opportunities and challenges. In today's era of digital information explosion, media data exhibits complex characteristics such as multimodality, massive volume, and strong real-time nature. Traditional methods rely on manual annotation or rule-based classification strategies to assign labels to video clips through pre-set keyword libraries or visual feature matching. While these methods can handle simple scenarios, they often suffer from decreased label accuracy when faced with complex spatiotemporal semantics, multimodal interactions (such as visual-text-action collaborative analysis), or dynamic scene transitions due to insufficient feature representation capabilities or lack of context awareness.
[0003] To address the aforementioned issues, existing research attempts to introduce reinforcement learning to optimize generation strategies or employ progressive learning to reduce reliance on purely manually labeled data. However, these methods still suffer from the following problems:
[0004] 1. Deficiencies in the reward mechanism: A single-dimensional reward signal is insufficient to comprehensively measure the semantic integrity and format standardization of the generated results (such as the compliance of the timestamp format `hh:mm:ss.ff`, missing key fields, etc.), and cannot accurately guide the model to generate high-quality output;
[0005] 2. Insufficient training stability: Directly applying reinforcement learning (such as GRPO) in multimodal scenarios is prone to training failure due to a lack of effective guidance;
[0006] 3. Error accumulation effect: In the multi-stage training process, the deviation of the model in the early stage will be amplified step by step, affecting the stability of the final performance. Summary of the Invention
[0007] The embodiments of this application provide a scene temporal location detection method, apparatus, storage medium and electronic device based on multimodal understanding, to overcome the problems of insufficient multimodal feature fusion, low label generation accuracy and poor model iteration efficiency in the existing technology of video shot label extraction.
[0008] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0009] According to a first aspect of the embodiments of this application, a scene temporal location detection method based on multimodal understanding is provided, including:
[0010] The video is divided into multiple shot segments, and the in and out timecode information and labels of the shot segment transitions are marked to construct training and validation sets.
[0011] Input feature sequence is generated by splicing and combining text features based on timecode information, visual features of shot segments, and designed prompt text.
[0012] Based on the visual language model architecture, a pre-trained model is constructed.
[0013] A supervised fine-tuning strategy is adopted to train the pre-trained model using the training set and optimize the parameters of the pre-trained model.
[0014] The input feature sequence is input into the optimized pre-trained model to obtain the predicted output features, and the gradient optimization of the reinforcement learning algorithm is used to optimize the pre-trained model using the group relative strategy.
[0015] The pre-trained model is comprehensively evaluated using the validation set.
[0016] In some embodiments of this application, based on the foregoing scheme, the step of dividing the video into multiple shot segments, labeling the in / out timecode information and tags of the shot segment transitions, and constructing a training set and a validation set includes:
[0017] The timecode information of the in-and-out points of the video during the camera transition is detected using a lens detection method.
[0018] Establish a tag classification system for intro, short clips, and outro, and classify videos into different categories based on intro and outtro timecode information;
[0019] A dataset is constructed based on a label classification system, and the dataset includes a training set and a validation set.
[0020] In some embodiments of this application, based on the foregoing scheme, the construction of the dataset includes three modes:
[0021] The first method: Randomly shuffle the short clips and then insert the continuous clips corresponding to the video intro between the shuffled clips;
[0022] The second method: randomly select two segments from the short clips and swap them, then insert the continuous clips corresponding to the video intro between these segments;
[0023] Third: Randomly select two adjacent segments from the small clips and swap them, then insert the continuous segments corresponding to the video intro between these segments.
[0024] In some embodiments of this application, based on the foregoing scheme, the step of generating an input feature sequence by concatenating and combining text features based on timecode information, visual features of shot segments, and designed prompt text includes:
[0025] The time code information is converted into text and encoded to obtain text features;
[0026] Extract keyframes from the shot and encode them to obtain visual features;
[0027] Design prompt text;
[0028] Text features, visual features, and prompt text are concatenated and combined to generate an input feature sequence.
[0029] In some embodiments of this application, based on the foregoing scheme, the step of employing a supervised fine-tuning strategy to train the pre-trained model using a training set and optimize the pre-trained model parameters includes:
[0030] Supervised training of the pre-trained model is performed using the training set;
[0031] And the parameters of the pre-trained model are optimized using the loss function L;
[0032] The expression for the loss function L is as follows:
[0033] ;
[0034] In the formula, For input tensors; B The batch size is [value], and the feature dimension is [value]. D ; Indicates the i The input vector of each sample; The output layer weight matrix has a vocabulary size of [size missing]. V The feature dimension is D ; This represents the bias vector, corresponding to the bias term for each category. The output of logits after linear transformation The i The nth sample pair j Unnormalized output for each category; For the first i The true label of each sample.
[0035] In some embodiments of this application, based on the foregoing scheme, the step of optimizing the pre-trained model using the gradient optimization of the reinforcement learning algorithm with a group relative policy includes:
[0036] Design a GRPO reward function, which includes a format reward function and a precision reward function;
[0037] The predicted output features are updated using the gradient of the GRPO reward function.
[0038] In some embodiments of this application, based on the foregoing scheme, the step of comprehensively evaluating the pre-trained model using a validation set includes:
[0039] The optimized pre-trained model was evaluated using a validation set, and its performance and accuracy in video scene recognition were measured using quantitative metrics.
[0040] The quantitative indicators include: format reward function value and accuracy reward function value.
[0041] According to a second aspect of the embodiments of this application, a scene temporal location detection device based on multimodal understanding is provided, comprising:
[0042] The data construction unit is used to divide the video into multiple shot segments, and to annotate the in and out timecode information and labels of the shot segment transitions, and to build training and validation sets.
[0043] The generation unit is used to generate an input feature sequence by splicing and combining text features based on timecode information, visual features of shot segments, and designed prompt text.
[0044] Model building unit, used to build pre-trained models based on visual language model architecture;
[0045] The model training unit is used to train the pre-trained model using a supervised fine-tuning strategy and the training set to optimize the parameters of the pre-trained model.
[0046] The reinforcement learning unit is used to input the input feature sequence into the optimized pre-trained model to obtain the predicted output features, and to optimize the pre-trained model by using the group relative strategy to optimize the gradient of the reinforcement learning algorithm.
[0047] The evaluation unit is used to comprehensively evaluate the pre-trained model using the validation set.
[0048] According to a third aspect of the embodiments of this application, a computer-readable storage medium is provided, the storage medium storing computer instructions that, when executed on a computer, cause the computer to perform the method as described in the first aspect.
[0049] According to a fourth aspect of the embodiments of this application, an electronic device is provided, including: a memory and a processor;
[0050] The memory is used to store computer instructions;
[0051] The processor is configured to invoke computer instructions stored in the memory, causing the electronic device to execute the method described in the first aspect.
[0052] The technical solution of this application, from video segmentation to tag generation and model iteration, revolves around multimodal understanding and efficient learning, and solves the core problems of low accuracy and poor efficiency in tag extraction in complex video scenarios in existing technologies.
[0053] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0054] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0055] Figure 1 A flowchart illustrating a scene temporal location detection method based on multimodal understanding according to an embodiment of this application is shown.
[0056] Figure 2 A block diagram of a scene temporal position detection device based on multimodal understanding according to an embodiment of this application is shown;
[0057] Figure 3 A block diagram of an electronic device according to one embodiment of this application is shown;
[0058] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation
[0059] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0060] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0061] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0062] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0063] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0064] The following detailed description of some embodiments of this application will be provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0065] In existing technologies, single-modal analysis suffers from numerous problems in shot label extraction, such as insufficient semantic integrity and scene coverage, as well as labeling bias when handling dynamic scenes and complex visual elements. Furthermore, traditional methods lack iterative optimization mechanisms, making it difficult for current technologies to meet the demands of diverse video content and adapt to ever-changing application scenarios.
[0066] To address the aforementioned issues, this application proposes a scene temporal location detection method based on multimodal understanding. This method fuses the spatiotemporal features of video with the semantic information of the timecode, and combines supervised fine-tuning and reinforcement learning to optimize model training.
[0067] See Figure 1 The diagram shows a flowchart of a scene temporal location detection method based on multimodal understanding according to an embodiment of this application.
[0068] like Figure 1 As shown, a scene temporal location detection method based on multimodal understanding is demonstrated, which specifically includes steps S100 to S600.
[0069] refer to Figure 1In step S100, the video is divided into multiple shot segments, and the in and out timecode information and labels of the shot segment switching points are marked to construct a training set and a validation set.
[0070] In some feasible embodiments, based on the foregoing scheme, step S100 specifically includes:
[0071] Step S110: Use a lens detection method to detect the timecode information of the in and out points of the video during lens switching;
[0072] Step S120: Establish a tag classification system for the intro, short clips, and outro, and classify the video into different category tags based on the intro / outtro timecode information;
[0073] Step S130: Construct a dataset based on the label classification system. The dataset includes a training set and a validation set.
[0074] In some feasible embodiments, based on the foregoing scheme, the construction of the dataset in step S130 includes three modes:
[0075] The first method: Randomly shuffle the short clips and then insert the continuous clips corresponding to the video intro between the shuffled clips;
[0076] The second method: randomly select two segments from the short clips and swap them, then insert the continuous clips corresponding to the video intro between these segments;
[0077] Third: Randomly select two adjacent segments from the small clips and swap them, then insert the continuous segments corresponding to the video intro between these segments.
[0078] Continue to refer to Figure 1 Step S200: Based on the text features of timecode information, the visual features of the shot clips, and the designed prompt text, an input feature sequence is generated by splicing and combining them.
[0079] In some feasible embodiments, based on the foregoing scheme, step S200 specifically includes:
[0080] Step S210: Convert the time code information into text for encoding and obtain text features;
[0081] Step S220: Extract keyframes from the shot clip and encode them to obtain visual features;
[0082] Step S230: Design the prompt text;
[0083] Step S240: Combine text features, visual features, and prompt text to generate an input feature sequence.
[0084] It should be noted that in this embodiment, since the pre-trained model will undergo supervised training and reinforcement learning, the generated input feature sequence is composed of text features, visual features and prompt text. If the model only undergoes supervised training, then the input feature sequence only needs to be composed of text features and visual features.
[0085] Continue to refer to Figure 1 Step S300: Based on the visual language model architecture, construct a pre-trained model.
[0086] For example, a pre-trained model can be built based on the Qwen2.5-VL model.
[0087] Continue to refer to Figure 1 In step S400, a supervised fine-tuning strategy is adopted to train the pre-trained model using the training set and optimize the parameters of the pre-trained model.
[0088] It should be noted that by performing supervised training and parameter optimization on the pre-trained model through step S400, the pre-trained model can be equipped with video scene recognition capabilities.
[0089] In some feasible embodiments, based on the foregoing scheme, step S400 specifically includes:
[0090] Step S410: Supervised training of the pre-trained model using the training set;
[0091] Step S420, and optimize the parameters of the pre-trained model using the loss function;
[0092] The expression for the loss function L is as follows:
[0093] ;
[0094] In the formula, For input tensors; B The batch size is [value], and the feature dimension is [value]. D ; Indicates the i The input vector of each sample; The output layer weight matrix has a vocabulary size of [size missing]. V The feature dimension is D ; This represents the bias vector, corresponding to the bias term for each category. The output of logits after linear transformation The i The nth sample pair j Unnormalized output for each category; For the first i The true label of each sample.
[0095] Continue to refer to Figure 1 In step S500, the input feature sequence is input into the optimized pre-trained model to obtain the predicted output features, and the pre-trained model is optimized by using the group relative strategy to optimize the gradient of the reinforcement learning algorithm.
[0096] It should be noted that step S500 can improve the model's decision-making ability by optimizing the pre-trained model through the group relative policy optimization reinforcement learning algorithm.
[0097] In some feasible embodiments, based on the foregoing scheme, the step of optimizing the pre-trained model using the gradient optimization of the reinforcement learning algorithm with a group relative strategy includes:
[0098] Design a GRPO reward function, which includes a format reward function and a precision reward function;
[0099] The predicted output features are updated using the gradient of the GRPO reward function.
[0100] Continue to refer to Figure 1 Step S600: Use the validation set to comprehensively evaluate the pre-trained model.
[0101] It should be noted that step S600 can be used to measure the effectiveness and accuracy of the pre-trained model in video scene recognition.
[0102] In some feasible embodiments, based on the foregoing scheme, step S600 specifically includes:
[0103] The optimized pre-trained model was evaluated using a validation set, and its performance and accuracy in video scene recognition were measured using quantitative metrics.
[0104] The quantitative indicators include: format reward function value and accuracy reward function value.
[0105] It should be noted that the format reward function value is used to evaluate the standardization of the model output format, including whether all necessary fields (such as "fragment in-point time code information" and "fragment out-point time code information") are included and whether the format of these fields is correct.
[0106] The accuracy reward function value is used to measure the accuracy of the model's output, including whether the extracted opening and closing timecode information matches the true values, ensuring the model's accuracy in semantic understanding and scene recognition.
[0107] In summary, the method provided in this application employs a two-stage framework integrating supervised fine-tuning training and reinforcement learning optimization. It combines timecode information and multimodal input from keyframes to construct a joint optimization mechanism: supervised fine-tuning is used for pre-training to inject prior knowledge, ensuring basic semantic understanding and format conformity; accuracy and format rewards are designed to achieve synergistic optimization, and joint constraints on output quality improve the accuracy of the group-relative strategy optimization stage; simultaneously, by combining timecode text features and shot visual features into a sequence and inputting it to the decoder to complete dynamic feature fusion, spatiotemporal modeling capabilities are significantly enhanced. This method effectively reduces the semantic fragmentation and format errors problems of traditional methods, providing a scalable technical path for multimodal streaming processing that meets the requirements of real-time video stream processing.
[0108] Table 1 shows the performance evaluation results on 64 validation data points after using different training strategies based on 252 training data points.
[0109] Table 1 shows the performance evaluation results on 64 validation data points after applying different training strategies based on 252 training data points.
[0110]
[0111] The model trained with SFT output 62 correct texts and 2 incorrect ones, achieving an accuracy of 96.88%. Meanwhile, the scene temporal location detection head (referred to as the "detector head") predicted all correct results. The detector head's higher accuracy stems from the fact that its output confidence score is directly constrained during training, and during inference, the time point corresponding to the category with the highest confidence score is directly selected as the prediction result.
[0112] For erroneous samples in the SFT text output, GRPO reinforcement learning training was performed using the model weights initialized after SFT training (labeled "SFT initialization + GRPO training, text output" in Table 1). The results show that the accuracy of the text output reached 100.00%. This indicates that GRPO can further improve model performance based on SFT, a conclusion consistent with the DeepSeek-R1 model and previous research on GRPO training on large language models (LLMs). It should be noted that the GRPO training process did not impose additional constraints on the detection head; therefore, its accuracy data is not included in Table 1.
[0113] In contrast, the performance of directly using GRPO for reinforcement learning training (labeled "Direct GRPO Training, Text Output" in Table 1) was significantly reduced: only 5 predictions were correct at the start time, and none were correct at the end time. This clearly demonstrates that directly applying GRPO reinforcement learning is completely ineffective without effective guidance. This result also confirms the viewpoint of previous research: necessary data guidance through SFT must be combined with GRPO reinforcement learning to achieve better performance.
[0114] The following describes an embodiment of the apparatus described in this application, which can be used to execute a scene temporal position detection method based on multimodal understanding as described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in the above applications.
[0115] Reference Figure 2 As shown, a scene temporal position detection device 200 based on multimodal understanding according to an embodiment of this application includes:
[0116] The data construction unit 201 is used to divide the video into multiple shot segments, and to annotate the in and out timecode information and labels of the shot segment transitions, and to construct training and validation sets.
[0117] The generation unit 202 is used to generate an input feature sequence by splicing and combining text features based on timecode information, visual features of shot segments, and designed prompt text.
[0118] Model building unit 203 is used to build a pre-trained model based on the visual language model architecture;
[0119] The model training unit 204 is used to train the pre-trained model using the training set and optimize the parameters of the pre-trained model by employing a supervised fine-tuning strategy.
[0120] The reinforcement learning unit 205 is used to input the input feature sequence into the optimized pre-trained model to obtain the predicted output features, and to optimize the pre-trained model by using the group relative strategy to optimize the gradient of the reinforcement learning algorithm.
[0121] Evaluation unit 206 is used to comprehensively evaluate the pre-trained model using the validation set.
[0122] like Figure 3 As shown, this application embodiment also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of the above-mentioned scene temporal position detection method based on multimodal understanding.
[0123] Since the electronic device described in this embodiment is the device used to implement the scene timing position detection device based on multimodal understanding in the embodiments of this application, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the method described in the embodiments of this application. Therefore, how the electronic device implements the method in the embodiments of this application will not be described in detail here. Any device used by those skilled in the art to implement the method in the embodiments of this application is within the scope of protection of this application.
[0124] In practice, when the computer program 311 is executed by the processor, it can implement any of the embodiments corresponding to the first aspect.
[0125] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.
[0126] It should be noted that, Figure 4 The computer system 400 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0127] like Figure 4 As shown, the computer system 400 includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 402 or programs loaded from storage portion 408 into Random Access Memory (RAM) 403, such as performing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 403. The CPU 401, ROM 402, and RAM 403 are interconnected via bus 404. An Input / Output (I / O) interface 405 is also connected to bus 404.
[0128] The following components are connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 410 as needed so that computer programs read from it can be installed into storage section 408 as needed.
[0129] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit (CPU) 401, it performs various functions defined in the system of this application.
[0130] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0132] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0133] In another aspect, this application also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the scene temporal location detection method based on multimodal understanding described in the above embodiments.
[0134] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement the scene temporal position detection method based on multimodal understanding described in the above embodiments.
[0135] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0136] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this application.
[0137] Other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. It should be understood that this application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A scene temporal location detection method based on multimodal understanding, characterized in that, include: The video is divided into multiple shot segments, and the in and out timecode information and labels of the shot segment transitions are marked to construct training and validation sets. Input feature sequence is generated by splicing and combining text features based on timecode information, visual features of shot segments, and designed prompt text. Based on the visual language model architecture, a pre-trained model is constructed. A supervised fine-tuning strategy is adopted to train the pre-trained model using the training set and optimize the parameters of the pre-trained model. The input feature sequence is input into the optimized pre-trained model to obtain the predicted output features, and the gradient optimization of the reinforcement learning algorithm is used to optimize the pre-trained model using the group relative strategy. The pre-trained model is comprehensively evaluated using the validation set; The supervised fine-tuning strategy, which uses the training set to train the pre-trained model and optimize its parameters, includes: Supervised training of the pre-trained model is performed using the training set; And the parameters of the pre-trained model are optimized using the loss function L; in, ; In the formula, For input tensors; B The batch size is [value], and the feature dimension is [value]. D ; Indicates the first i The input vector of each sample; The output layer weight matrix has a vocabulary size of [size missing]. V The feature dimension is D ; This represents the bias vector, corresponding to the bias term for each category. The output of logits after linear transformation The i The nth sample pair j Unnormalized output for each category; For the first i The true label of each sample; The method of optimizing the pre-trained model using the gradient optimization of the reinforcement learning algorithm with a group relative strategy includes: Design a GRPO reward function, which includes a format reward function and a precision reward function; The predicted output features are updated using the gradient of the GRPO reward function.
2. The method according to claim 1, characterized in that, The process of dividing the video into multiple shot segments, labeling the in and out timecode information and tags of the shot segment transitions, and constructing training and validation sets includes: The timecode information of the in-and-out points of the video during the camera transition is detected using a lens detection method. Establish a tag classification system for intro, short clips, and outro, and classify videos into different categories based on intro and outtro timecode information; A dataset is constructed based on a label classification system, and the dataset includes a training set and a validation set.
3. The method according to claim 2, characterized in that, There are three modes for building datasets: The first method: Randomly shuffle the short clips and then insert the continuous clips corresponding to the video intro between the shuffled clips; The second method: randomly select two segments from the short clips and swap them, then insert the continuous clips corresponding to the video intro between these segments; Third: Randomly select two adjacent segments from the small clips and swap them, then insert the continuous segments corresponding to the video intro between these segments.
4. The method according to claim 1, characterized in that, The input feature sequence is generated by concatenating and combining text features based on timecode information, visual features of shot segments, and designed prompt text, including: The time code information is converted into text and encoded to obtain text features; Extract keyframes from the shot and encode them to obtain visual features; Design prompt text; Text features, visual features, and prompt text are concatenated and combined to generate an input feature sequence.
5. The method according to claim 1, characterized in that, The comprehensive evaluation of the pre-trained model using the validation set includes: The optimized pre-trained model was evaluated using a validation set, and its performance and accuracy in video scene recognition were measured using quantitative metrics. The quantitative indicators include: format reward function value and accuracy reward function value.
6. A scene temporal position detection device based on multimodal understanding, applied to the method as described in any one of claims 1-5, characterized in that, include: The data construction unit is used to divide the video into multiple shot segments, and to annotate the in and out timecode information and labels of the shot segment transitions, and to build training and validation sets. The generation unit is used to generate an input feature sequence by splicing and combining text features based on timecode information, visual features of shot segments, and designed prompt text. Model building unit, used to build pre-trained models based on visual language model architecture; The model training unit is used to train the pre-trained model using a supervised fine-tuning strategy and the training set to optimize the parameters of the pre-trained model. The reinforcement learning unit is used to input the input feature sequence into the optimized pre-trained model to obtain the predicted output features, and to optimize the pre-trained model by using the group relative strategy to optimize the gradient of the reinforcement learning algorithm. The evaluation unit is used to comprehensively evaluate the pre-trained model using the validation set.
7. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-5.
8. An electronic device, characterized in that, include: memory and processor; The memory is used to store computer instructions; The processor is configured to invoke computer instructions stored in the memory, causing the electronic device to perform the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Video classification method and device and storage medium
CN114238690A
News video dynamic abstract extraction method, equipment, medium and system
CN116416549A