A Dense Video Description Training Method, Device, Equipment and Medium
By using large language models and pre-trained vision-language models in intensive video description technology, combined with pseudo-boundary enhancement and refining algorithms, the pseudo-boundary noise and high computational complexity in unlabeled video data is solved, and the model performance is significantly improved.
Patent Information
- Application Number
- CN202410316982.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-03-20
AI Technical Summary
When existing intensive video description technology uses unlabeled video data, there are problems of pseudo-boundary noise and high computational complexity, resulting in limited improvement in model performance.
Using large language models and pre-trained vision-language models, high-quality event descriptions and corresponding pseudo-event boundary boundaries are generated through pseudo-boundary enhancement and refining algorithms, and combined with online refining algorithms, the performance of the model on intensive video description tasks is improved.
It realizes the performance of dense video description models without manual annotation, achieves the best performance in weakly supervised and supervised methods, and is compatible with the existing labeled training framework.
Smart Images

Figure CN118334677B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dense video description, and particularly to a dense video description training method, device, equipment and medium. Background Art
[0002] Dense video description plays an important role in many fields such as long video analysis, video understanding and video recommendation. Different from traditional video description tasks that only require generating a single text description for a short video clip, dense video description needs to locate multiple events in a long video sequence and provide more detailed descriptions for each event. Therefore, this task exhibits higher complexity and poses more challenges. Among them, the temporal boundaries of events are the key to dense video description, which provides accurate event localization for long videos, thus ensuring that the generated dense descriptions are consistent and contextually coherent. However, video data containing accurate temporal boundaries of events is very scarce, and at the same time, annotating event boundaries is a costly, time-consuming and laborious process.
[0003] Some existing technologies attempt to solve this problem of dense video description from the annotation level. One school of thought focuses on dense video description technology under weak supervision, that is, only using video-event description data, without using event boundary annotations to train a dense video description model, and expecting to approximate the performance of dense video description under full supervision. However, these methods need to redesign a completely new training process or inference framework, so they cannot be directly combined with existing dense video description methods. At the same time, these methods will inevitably introduce more complex models or structural designs, bringing higher computational complexity or slower inference speed. More importantly, these technologies do not directly conduct in-depth analysis on video-description data pairs without time boundary annotations from the data level.
[0004] Recently, a technique has emerged that emphasizes leveraging large-scale unannotated video data at the data level to train dense video description models, thereby significantly enhancing model performance. Specifically, this method designs a single-stage dense video description model that can simultaneously output text descriptions of all events and their corresponding temporal boundaries. Based on this model, they further collect a large number of unannotated videos, transcribe the audio dialogues in these videos into corresponding captions and timelines, and use different sentences as corresponding pseudo-text descriptions and event boundaries to form the pseudo-labels required for training. Finally, this method uses a large number of video-pseudo-label data pairs for training, effectively improving the performance of dense video description without manual annotation. However, the original video caption content used in this technique usually contains some meaningless dialogues, personalized conversations, or background sounds. In addition, the video transcribed caption content and the video visual content are often not strictly aligned in the time domain, which will further introduce a large amount of noise in event learning. Therefore, although this technical solution has successfully utilized unannotated videos to enhance dense video description, there is still significant room for improvement in the utilization rate and effectiveness of large-scale videos. Summary of the Invention
[0005] The purpose of the present invention is to provide a dense video description training method, device, equipment, and medium. With the help of a large language model, a large number of unannotated videos are utilized, combined with a pseudo-boundary enhancement and refinement algorithm, to generate high-quality event descriptions and corresponding pseudo-event boundaries, and combined with an online refinement algorithm for pseudo-boundaries, large-scale unannotated videos are used for pre-training of dense video description tasks, significantly improving model performance.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] According to the first aspect of the present invention, there is provided a dense video description training method based on pseudo-boundary enhancement and refinement using unannotated videos, including the following steps:
[0008] S1, collect unannotated video data, and extract the original caption text from the videos;
[0009] S2, based on the large language model, design a specific task text, input the original caption text into the large language model for summarization, and output the ordered event description text in the refined video;
[0010] S3, based on the pre-trained vision-language model, calculate the similarity matrix between the event description text and the original video frames, use the joint optimization algorithm for event description-boundary to generate the pseudo-event boundary corresponding to each event description, and optimize the event description text;
[0011] S4. Use the generated event description text and the corresponding pseudo-event boundaries as learning objectives in the optimization process of the dense video description model. Calculate the loss function based on the event description and boundary results predicted by the model and the event description text and pseudo-event boundaries respectively, and pre-train the dense video description model based on the loss function value;
[0012] S5. During the pre-training process of the dense video description model, generate multiple candidate boundaries based on the pseudo-event boundaries, evaluate the multiple candidate boundaries using the dense video description model, refine the pseudo-event boundaries corresponding to the video iteratively according to the evaluation results, and use the updated pseudo-event boundaries as the learning objective for the next round of the dense video description model;
[0013] S6. Fine-tune the pre-trained dense video description model on the downstream dense video description dataset with event boundary annotations.
[0014] In the step S2, the extraction process of the large language model for the ordered event description text includes the following steps:
[0015] Test the large language model with a small number of real annotated event descriptions, and design the optimal step extraction task prompt text through the model feedback;
[0016] Append the original caption text to the task prompt text and input it into the large language model;
[0017] Use multiple large language models to generate multiple candidate event description texts.
[0018] In the step S3, the generation process of the pseudo-event boundaries includes the following steps:
[0019] S31. Based on the pre-trained vision-language model, calculate the similarity matrix between the original video frames and the event description text;
[0020] S32. Based on the assumption of event continuity and uniform distribution, divide the entire video evenly into N segments as the initialized pseudo-boundaries;
[0021] S33. Adjust each pseudo-boundary based on the similarity matrix:
[0022] S331. Calculate the new boundary center: In the range near the current pseudo-boundary, collect the k frames of images that are most similar to the event description text based on the similarity matrix, calculate the total distance between each frame and the other k - 1 frames, and select the frame with the shortest total distance as the position where the similar frames are most dense, that is, the new boundary center;
[0023] S332. Calculate the new boundary width: Consider the k frames of images that are most similar to the event description text in the entire video, calculate the standard deviation of the k frames of images based on the new boundary center, and determine the new boundary width based on the standard deviation;
[0024] S333. Update the center and width of the current boundary to generate a new pseudo-boundary, and evaluate the quality of the generated pseudo-boundary based on the cost function;
[0025] S334. Repeat the above steps S331 - S333 to iteratively adjust the pseudo-boundary, and select the pseudo-boundary with the minimum cost function during the iteration as the final pseudo-event boundary.
[0026] The cost function is:
[0027] L = s kn *distance(c k , box)
[0028] where s kn is the similarity between the k-th frame and the current event description text, c k is the position of the k-th frame, distance represents the distance from c k to the current pseudo-boundary box. If c k is within the pseudo-boundary range, distance is negative; otherwise, it is positive:
[0029] distance = -min(c k - left, right - c k ), if c k is in box,
[0030] distance = max(left - c k , index - right), if c k is not in box
[0031] left represents the left boundary of the pseudo-boundary, and right represents the right boundary of the pseudo-boundary.
[0032] In step S3, the optimization of the event description text is specifically as follows: Based on the pseudo-boundaries generated from each candidate event description text, calculate the cost function respectively, and select the candidate event description text with the minimum cost function as the final event description text to complete the optimization.
[0033] In step S4, the loss function includes the event description text loss and the pseudo-boundary loss. Among them, the event description text loss uses the cross-entropy loss, and the pseudo-boundary loss uses the L2 loss.
[0034] The training process in step S4 can be combined with most existing dense description video technologies.
[0035] In step S5, the refinement process of the pseudo-event boundary includes the following steps:
[0036] Based on the original pseudo-event boundary positions, randomly perturb to generate multiple candidate boundaries;
[0037] Based on the intermediate results of the dense video description model at each candidate boundary, evaluate each candidate boundary to obtain the final score of each candidate boundary, where the intermediate results include the confidence score for predicting the event boundary according to the video features and the confidence score for predicting the event text description according to the video features;
[0038] Normalize the scores of all candidate boundaries;
[0039] Select and fuse candidate boundaries based on the normalized scores;
[0040] Replace the original pseudo-event boundaries with the fused candidate boundaries.
[0041] According to the second aspect of the present invention, there is provided a dense video description training device based on pseudo-boundary enhancement and refinement using unannotated videos, including:
[0042] A data collection and preprocessing module for collecting unannotated video data and extracting the original caption text from the video;
[0043] An event description text generation module for designing specific task texts based on a large language model, inputting the original caption text into the large language model for summarization, and outputting the ordered event description text in the refined video;
[0044] A pseudo-boundary generation module for calculating the similarity matrix between the event description text and the original video frames based on a pre-trained vision-language model, using the joint optimization algorithm of event description-boundary to generate the pseudo-event boundaries corresponding to each event description, and optimizing the event description text;
[0045] A dense video description model optimization module for performing the following steps: using the generated event description text and the corresponding pseudo-event boundaries as the learning objectives in the optimization process of the dense video description model, calculating the loss functions based on the event description and boundary results predicted by the model respectively and the event description text and the pseudo-event boundaries, and pre-training the dense video description model based on the loss function values; during the pre-training process of the dense video description model, generating multiple candidate boundaries based on the pseudo-event boundaries, evaluating the multiple candidate boundaries using the dense video description model, and iteratively refining the pseudo-event boundaries corresponding to the video online according to the evaluation results, and using the updated pseudo-event boundaries as the learning objectives for the next round of the dense video description model;
[0046] A model fine-tuning module for fine-tuning the pre-trained dense video description model on a downstream dense video description dataset with event boundary annotations.
[0047] According to a third aspect of the present invention, there is provided an electronic device, including a memory and a processor, where a computer program is stored on the memory, and when the processor executes the program, the method described above is implemented.
[0048] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] 1. The present invention can well combine labeled data and unlabeled data, be compatible with current existing labeled training frameworks, and implement pre-training on a large amount of unlabeled data.
[0051] 2. From the test results of the publicly available dataset, it can be seen that the present invention achieves the best performance among weakly supervised and supervised methods, that is, it has the best performance in event localization and description generation.
[0052] 3. By using large language models and pre-trained multi-modal models, the present invention efficiently utilizes a large amount of unlabeled data, thereby further improving the performance of the model in dense event prediction tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a flowchart of the method of the present invention;
[0054] Figure 2 is a comparison between the processing method in step S2 of the present invention and the prior art;
[0055] Figure 3 is a schematic diagram of generating pseudo-boundaries corresponding to events in step S3 of the present invention;
[0056] Figure 4 is a schematic diagram of the video description result of the optimized dense video description model of the present invention in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.
[0058] This embodiment provides a method for training dense video description based on pseudo-boundary enhancement and refinement using unlabeled videos, as Figure 1 shown, including the following steps:
[0059] S1. Collect unlabeled video data and extract original caption text from the videos.
[0060] The video data is mainly selected and collected from the publicly available dataset HowTo100M, which is collected from the video website YouTube. The video types are mainly step-by-step teaching videos, such as cooking tutorials in the kitchen scene, etc. In this embodiment, 50,000 videos are selected from HowTo100M, and the Google audio transcription tool Speech-to-Text is used to generate corresponding subtitle files for each video data. The subtitle files contain subtitle text and the timestamp corresponding to each sentence of text.
[0061] S2. Based on the large language model, design specific task text, input the original subtitle text into the large language model for summarization, and output the numbered event description text in the refined video.
[0062] Specifically, it includes the following steps:
[0063] S21. Use a small number of real annotated event descriptions to test the large language model, and design the optimal step extraction task prompt text through the model feedback.
[0064] To obtain the prompts for completing this task, this embodiment adopts a loop prompt strategy: first provide a small number of real event descriptions, and ask the language model what kind of prompts can be used to extract similar results from the video subtitles; then test and manually correct the initial prompts given by the language model, and iterate continuously until the generated prompts meet the requirements. The final designed prompt format is as follows:
[0065] Task: Only extract concise step-by-step operational instructions from the video subtitles, with each step represented by a sentence, which needs to clearly reflect the operation behavior and exclude content unrelated to the video.
[0066] S22, as Figure 2 shown, append the original subtitle text to the task prompt text and input it into the large language model to obtain the summarized numbered step-by-step event text description.
[0067] S23. Repeat the above steps using multiple different (number of parameters, style types) large language models (such as LLAMA, ChatGPT, etc.) to generate multiple candidate event description texts for joint optimization in step S3.
[0068] S3. Based on the pre-trained vision-language model, calculate the similarity matrix between the event description text and the original video frames, use the event description-boundary joint optimization algorithm to generate the pseudo-event boundaries corresponding to each event description, and optimize the event description text.
[0069] This step aims to effectively identify and determine the boundaries of each event description to ensure its accuracy and coherence on the video timeline. Specifically, it includes the following steps:
[0070] S31. Calculate the similarity matrix between video frames and event descriptions based on pre-trained vision-language models such as CLIP and UniVL.
[0071] S32. Initialize pseudo-boundaries: Considering that the N events in a video are in a sequential order (N is the number of event descriptions), the entire video is evenly divided into N segments as the initialized pseudo-boundaries. As shown in Figure 3 the figure, then for each event boundary, adjust it using the similarity matrix. The steps are as follows:
[0072] S33. Adjust each pseudo-boundary based on the similarity matrix:
[0073] S331. Calculate the new boundary center: In the range near the current pseudo-boundary, collect the k frames of images that are most similar to the event description text based on the similarity matrix. Calculate the total distance between each frame and the other k - 1 frames, and select the frame with the shortest total distance as the position where the similar frames are most concentrated, that is, the new boundary center.
[0074] In this embodiment, the near range is defined as: Let the current pseudo-boundary be represented as [start, end], then the near range of the current pseudo-boundary is represented as [start – scale * width, end + scale * width], where scale is a hyperparameter, width is the average length of each step, that is, how many frames of images correspond to each event description on average, width = num_frame / num_steps, num_frame represents the video frame data, and num_steps represents the number of event description texts.
[0075] S332. Calculate the new boundary width: Consider the k frames of images that are most similar to the event description text in the entire video, calculate the standard deviation std of these k frames of images based on the new boundary center, and determine the new boundary width based on the standard deviation: New boundary width = ratio * std, where ratio is a hyperparameter, and the value range is generally 1 - 2.
[0076] S333. Update the center and width of the current boundary, generate a new pseudo-boundary, and evaluate the quality of the generated pseudo-boundary based on the cost function.
[0077] According to the new boundary center, taking the new boundary center as the reference, determine all the frame images within the width range according to the new boundary width. For example, if k1, k2, k3, k4 are included within the width range, then take the position of the earliest frame image k1 among these frame images as the left boundary, and the position of the last frame image k4 as the right boundary, thereby generating a new pseudo-boundary.
[0078] In this embodiment, the cost function is as follows:
[0079] L = s kn *distance(c k , box)
[0080] where s kn is the similarity between the k-th frame and the current event description text, c k is the position of the k-th frame, distance represents the distance from c k to the current pseudo-boundary box. If c k is within the pseudo-boundary, distance is negative, indicating the distance of this frame from leaving the current pseudo-boundary; otherwise, it is positive, indicating the distance of this frame approaching the pseudo-boundary. Specifically, it is expressed as:
[0081] distance = -min(c k - left, right - c k ), if c k is in the box,
[0082] distance = max(left - c k , index - right), if c k is not in the box
[0083] left represents the left boundary of the pseudo-boundary, and right represents the right boundary of the pseudo-boundary.
[0084] S334. Repeat the above steps S331 - S333, iteratively adjust the pseudo-boundary, and select the pseudo-boundary with the minimum cost function during the iteration as the final pseudo-event boundary.
[0085] In addition, for each pseudo-boundary generated based on each candidate event description text, calculate the cost function respectively, and select the candidate event description text with the minimum cost function as the final event description text to complete the optimization.
[0086] S4. Use the generated event description text and the corresponding pseudo-event boundary as the learning objectives during the optimization of the dense video description model. Calculate the loss function based on the event description and boundary results predicted by the model and the event description text and pseudo-event boundary respectively, and pre-train the dense video description model based on the loss function value.
[0087] Since the present invention does not involve the design of the structure of the dense video description model, the existing fully supervised dense video description technology can be used for the dense video description model here. Specifically, the present invention combines the existing technology 1 to train and infer the dense video description task. During the training process, the generated event description text and the corresponding pseudo-event boundaries are used as the learning objectives in the model training process. The cross-entropy loss is calculated between the event description result predicted by the model and the event description text, and the L2 loss is calculated between the boundary result predicted by the model and the pseudo-boundaries. The two are added together to calculate the final loss function. After calculating the value of the loss function, the stochastic gradient descent algorithm is used to update the model parameters. The learning rate of the model parameters decays by 0.1 times every 10 iteration times, and the entire training process includes 30 complete iteration times.
[0088] S5. During the pre-training process of the dense video description model, multiple candidate boundaries are generated based on the pseudo-event boundaries, and the dense video description model is used to evaluate the multiple candidate boundaries. According to the evaluation results, the pseudo-event boundaries corresponding to the video are refined online in an iterative manner, and the updated pseudo-event boundaries are used as the learning objectives for the next round of the dense video description model.
[0089] In the step S5, the refinement process of the pseudo-event boundaries includes the following steps:
[0090] S51. Based on the original pseudo-event boundary positions, random perturbations are made to generate multiple candidate boundaries;
[0091] S52. Based on the intermediate results of the dense video description model at each candidate boundary, each candidate boundary is evaluated to obtain the final score of each candidate boundary. Among them, the intermediate results include the event confidence score for predicting the event boundary based on the video features and the description confidence score for predicting the event text description based on the video features. The event confidence score and the description confidence score are multiplied to obtain the final score;
[0092] S53. Normalize the scores of all candidate boundaries;
[0093] S54. Select and fuse the candidate boundaries based on the normalized scores: First, select all candidate pseudo-boundaries, pick out the K candidate boundaries with the highest scores, and then fuse these K boundaries according to their scores;
[0094] S55. Replace the original pseudo-event boundaries with the fused candidate boundaries.
[0095] S6. Fine-tune the pre-trained dense video description model on the downstream dense video description dataset with event boundary annotations.
[0096] Finally, the model trained using the above process can be fine-tuned on the fully annotated video data. Specifically, the model is trained using the real-annotated event boundaries and event descriptions as the learning objectives. During the fine-tuning process, the learning rate of the model parameters is set to 0.0001, and the fine-tuning process includes 10 complete training processes.
[0097] As shown in Tables 1 and 2, the present invention achieves the optimal performance in weakly supervised and supervised settings on the widely used YouCook2 and ActivityNet datasets.
[0098] Table 1 Comparison of supervised performance on YouCook2
[0099]
[0100] Table 2 Comparison of weakly supervised and supervised performance on ActivityNet
[0101]
[0102] The evaluation metrics used in Tables 1 and 2 are METEOR (Metric for Evaluation of Translation with Explicit ORdering), CIDEr (Consensus-based Image Description Evaluation), SODA_c (Story-Oriented Dense Video Captioning Evaluation Metric), Recall, and Precision. The larger these metric values are, the better the performance.
[0103] The prior art 1 adopts the solution in the following literature: Teng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng, Ran Cheng, and Ping Luo. End-to-end dense videocaptioning with parallel decoding. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 6847–6857, 2021.
[0104] The solution adopted by the prior art 2 is the one in the following literature: Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 10714–10726, 2023.
[0105] The solution adopted by the prior art 3 is the one in the following literature: Shaoxiang Chen and Yu-Gang Jiang. Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 8425–8435, 2021
[0106] The visualization results are as Figure 4 shown. Compared with the true event annotations, it can be seen that the present invention enables the dense video captioning model to predict accurate event boundaries. At the same time, the text descriptions of the events can also more accurately describe the video visual content in the corresponding time intervals. In addition, the text content between events is also more coherent and consistent.
[0107] The above is the introduction of the method embodiment. The following further illustrates the solution of the present invention through the device embodiment.
[0108] This embodiment provides a training device for dense video captioning based on pseudo-boundary enhancement and refinement using unannotated videos, including:
[0109] A data collection and preprocessing module, configured to collect unannotated video data and extract the original caption text from the videos;
[0110] An event description text generation module, which is used to design specific task texts based on a large language model, input the original caption text into the large language model for summarization, and output the ordered event description text in the refined video;
[0111] A pseudo-boundary generation module, which is used to calculate the similarity matrix between the event description text and the original video frames based on a pre-trained vision-language model, and use the joint optimization algorithm of event description-boundary to generate the pseudo-event boundaries corresponding to each event description and optimize the event description text;
[0112] A dense video description model optimization module, which is used to perform the following steps: use the generated event description text and the corresponding pseudo-event boundaries as the learning objectives in the optimization process of the dense video description model, calculate the loss functions based on the event descriptions and boundary results predicted by the model and the event description text and the pseudo-event boundaries respectively, and pre-train the dense video description model based on the loss function values; during the pre-training process of the dense video description model, generate multiple candidate boundaries based on the pseudo-event boundaries, use the dense video description model to evaluate the multiple candidate boundaries, and iteratively refine the pseudo-event boundaries corresponding to the video online according to the evaluation results, and use the updated pseudo-event boundaries as the learning objectives for the next round of the dense video description model;
[0113] A model fine-tuning module, which is used to fine-tune the pre-trained dense video description model on a dense video description dataset with event boundary annotations given downstream.
[0114] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.
[0115] The electronic device of the present invention includes a processor, which can be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processing unit (CPU), an application processor (AP), a digital signal processor (DSP), a graphics processing unit (GPU), a neural-network processing unit (NPU), etc. The processor can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) or the computer program instructions loaded from the storage unit into the random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The processor, ROM, and RAM are connected to each other via a bus. The input / output (I / O) interface is also connected to the bus.
[0116] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0117] The processor executes the various methods and processes described above, such as method S1~S6. For example, in some embodiments, method S1~S6 can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the processor, one or more steps of method S1~S6 described above can be executed. Alternatively, in other embodiments, the processor can be configured to execute method S1~S6 in any other suitable manner (for example, by means of firmware).
[0118] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), and the like.
[0119] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0120] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0121] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A dense video description training method based on pseudo-boundary enhancement and refinement using unlabeled videos, characterized in that: The following steps are involved: S1, collect unlabeled video data and extract original subtitle text from the video; S2, based on the large language model, designs specific task text, inputs the original subtitle text into the large language model for summary and induction, and outputs the ordered event description text in the refined video; S3, based on the pre-trained visual-language model, calculate the similarity matrix between the event description text and the original video frame, use the event description-boundary joint optimization algorithm to generate a pseudo event boundary corresponding to each event description, and optimize the event description text, the event description-boundary joint optimization algorithm selects the pseudo boundary with the smallest cost function in the iterative process as the final pseudo event boundary, and calculates the cost function based on the pseudo boundary generated by each candidate event description text, and selects the candidate event description text with the smallest cost function as the final event description text; wherein the cost function is: , in, It is k The similarity between the frame and the current event description text, For the k The position of the frame, distance express To the current pseudo boundary box If the distance Within the pseudo-boundary range, distance is a negative value, ,if Not within the pseudo boundary, distance is a positive value, , left represents the left boundary of the pseudo boundary, right represents the right boundary of the pseudo boundary, index express The position index of S4, taking the generated event description text and the corresponding pseudo event boundary as the learning target in the optimization process of the dense video description model, calculating the loss function with the event description text and the pseudo event boundary respectively based on the event description and boundary results predicted by the dense video description model, and pre-training the dense video description model based on the loss function value; S5, during the pre-training process of the dense video description model, multiple candidate boundaries are generated based on the pseudo event boundaries, the multiple candidate boundaries are evaluated using the dense video description model, the pseudo event boundaries corresponding to the video are refined online in an iterative manner based on the evaluation results, and the updated pseudo event boundaries are used as the learning target of the next round of the dense video description model; S6, fine-tune the pre-trained dense video description model on the downstream given dense video description dataset with event boundary annotations.
2. The method for dense video description training based on pseudo boundary enhancement and refinement using unlabeled videos according to claim 1, characterized in that: The step S2 specifically includes the following steps: Use real annotated event descriptions to test the large language model, and use model feedback to design the optimal steps to extract task prompt text; Append the original subtitle text to the task prompt text and input it into the large language model; Multiple large language models are used to generate multiple candidate event description texts.
3. The method for dense video description training based on pseudo boundary enhancement and refinement using unlabeled videos according to claim 1, characterized in that: In step S3, the process of generating the pseudo event boundary includes the following steps: S31, based on the pre-trained visual-language model, calculates the similarity matrix between the original video frame and the event description text; S32, based on the assumption that events are continuous and evenly distributed, the entire video is evenly divided into N segments as initial pseudo boundaries; S33, based on the similarity matrix, adjust each pseudo boundary: S331, calculating the new boundary center: within the vicinity of the current pseudo boundary, based on the similarity matrix, collect the k frames of images that are most similar to the event description text, calculate the total distance between the current frame and the other k-1 frames for each frame, and select the frame with the shortest total distance as the location with the most similar frames, that is, the new boundary center; S332, calculating the new border width: considering the k frames of images that are most similar to the event description text in the entire video, calculating the standard deviation of the k frames of images based on the center of the new border, and determining the new border width based on the standard deviation; S333, updating the center and width of the current boundary, generating a new pseudo boundary, and evaluating the quality of the generated pseudo boundary based on the cost function; S334, repeat the above steps S331-S333, iteratively adjust the pseudo boundary, and select the pseudo boundary with the smallest cost function in the iterative process as the final pseudo event boundary.
4. The method for dense video description training based on pseudo boundary enhancement and refinement using unlabeled videos according to claim 1, characterized in that: In step S4, the loss function includes event description text loss and pseudo boundary loss, wherein the event description text loss adopts cross entropy loss and the pseudo boundary loss adopts L2 loss.
5. The method for dense video description training based on pseudo boundary enhancement and refinement using unlabeled videos according to claim 1, characterized in that: In step S5, the refining process of the pseudo event boundary includes the following steps: Based on the original pseudo event boundary position, randomly perturb to generate multiple candidate boundaries; Based on the intermediate results of the dense video description model at each candidate boundary, each candidate boundary is evaluated to obtain a final score for each candidate boundary, wherein the intermediate results include a confidence score for predicting the event boundary based on the video features and a confidence score for predicting the event text description based on the video features; Normalize the scores of all candidate boundaries; Select and merge candidate boundaries based on the normalized scores; Replace the original pseudo-event boundaries with the fused candidate boundaries.
6. A dense video description training device based on pseudo-boundary enhancement and refinement using unlabeled videos, characterized in that: include: Data collection and preprocessing module, used to collect unlabeled video data and extract original subtitle text from the video; The event description text generation module is used to design specific task text based on the large language model, input the original subtitle text into the large language model for summarization, and output the ordered event description text in the refined video; The pseudo boundary generation module is used to calculate the similarity matrix between the event description text and the original video frame based on the pre-trained visual-language model, generate the pseudo event boundary corresponding to each event description using the event description-boundary joint optimization algorithm, and optimize the event description text. The event description-boundary joint optimization algorithm selects the pseudo boundary with the smallest cost function in the iterative process as the final pseudo event boundary, and calculates the cost function based on the pseudo boundary generated by each candidate event description text, and selects the candidate event description text with the smallest cost function as the final event description text; wherein the cost function is: , in, It is k The similarity between the frame and the current event description text, For the k The position of the frame, distance express To the current pseudo boundary box If the distance Within the pseudo-boundary range, distance is a negative value, ,if Not within the pseudo boundary, distance is a positive value, , left represents the left boundary of the pseudo boundary, right represents the right boundary of the pseudo boundary, index express The position index of The dense video description model optimization module is used to perform the following steps: using the generated event description text and the corresponding pseudo-event boundary as the learning target in the dense video description model optimization process, calculating the loss function based on the event description and boundary results predicted by the dense video description model and the event description text and the pseudo-event boundary, respectively, and pre-training the dense video description model based on the loss function value; in the dense video description model pre-training process, generating multiple candidate boundaries based on the pseudo-event boundary, using the dense video description model to evaluate the multiple candidate boundaries, refining the pseudo-event boundary corresponding to the video online in an iterative manner according to the evaluation results, and using the updated pseudo-event boundary as the next round of learning target of the dense video description model; The model fine-tuning module is used to fine-tune the pre-trained dense video description model on the dense video description dataset with event boundary annotations given downstream.
7. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Techniques for dense video descriptions
CN110709855A
Lable feature near-duplicated video detection method based on convolutional neural network semantic classification
CN111723692A