Weak supervision video time sequence content positioning method and system based on action nominating and denoising
By using action nomination denoising method in video timing content positioning, the network is reconstructed by using the timing diffusion network and semantic text, combining noise mixing and combinatorial sorting learning, the problem of event boundary blurring and semantic inconsistency between modes in weakly supervised video timing content positioning is solved, and higher positioning accuracy and lower computational complexity are achieved.
Patent Information
- Application Number
- CN202510150661.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-10
AI Technical Summary
Video timing content positioning based on weakly supervised learning faces problems such as blurred event time boundaries, poor continuity and semantic inconsistency between modes, resulting in low positioning accuracy and high computational complexity.
Using a method based on action nomination denoising, video-level text description supervision, time-sequence diffusion network and semantic text reconstruction network, combining noise mixing and combinatorial sorting learning, and optimizing video clip positioning.
It improves the accuracy and boundary clarity of video positioning, reduces the cost of computing complexity and timing annotation, and enhances the generalization ability of the model.
Smart Images

Figure CN120126045A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video understanding, relates to multi-modal fusion and video content localization technologies, and particularly relates to a weak-supervised video temporal content localization method and system based on action nomination denoising. Background Art
[0002] With the development of deep learning technologies, video understanding has become one of the important research directions in the field of computer vision. As a carrier form of multi-modal information, video integrates dynamic information in the time dimension and visual content in the space dimension, and has extensive applications in fields such as action recognition, event localization, and multi-modal retrieval. Compared with images, videos consist of a continuous sequence of frames in the time dimension and have higher complexity, which makes the research on video modalities more challenging. In recent years, with the rapid development of large models and high computing power, significant progress has been made in related research in the field of video understanding. However, these models with excellent performance rely on large-scale finely labeled datasets, resulting in expensive data annotation costs and being difficult to expand in practical applications. In real life, videos are mostly of different lengths and have complex content, and frame-by-frame annotation of events occurring in long videos incurs expensive time and labor costs. Therefore, weak-supervised learning methods have become a solution for analyzing and understanding long videos. Video temporal content localization based on weak-supervised learning is an important task in video understanding, and its goal is to locate the time segments related to the coarse-grained text description in the video. Different from the full-supervised method, the method based on weak-supervised learning only provides video-level text annotations without providing precise time boundaries, which greatly reduces the annotation cost and shows great application potential in practical scenarios such as intelligent video surveillance, sports analysis, and video retrieval.
[0003] Due to the limitations of weak-supervised annotation and the complexity of video content itself, video temporal content localization based on weak-supervised learning faces many challenges: on the one hand, the time boundaries of events in the video are blurred and the continuity is poor, making it difficult for the model to accurately distinguish adjacent events, and the model lacks dynamic perception modeling of the time dimension; on the other hand, videos and text descriptions belong to different modalities, with semantic ambiguity and inconsistency, which requires the model to fully understand the text semantics for supplementation and inference. The weak-supervised video temporal content localization method based on action nomination denoising proposed by the present invention relies on video-level text description supervision and proposes solutions to the above two types of problems, enabling the model to fully mine natural language text and integrate it into the understanding of the time and space content of the video modality, thereby improving the accuracy of video localization. Summary of the Invention
[0004] The object of the present invention is to provide a method for weakly supervised video temporal content localization based on action nomination denoising for general problems under weakly supervised learning conditions, aiming to solve the problems of under-localization and over-localization of video segments caused by the subjectivity and ambiguity of weakly supervised text annotations. The steps include:
[0005] Obtain the video to be processed and the text description corresponding to the video; generate positive examples and negative examples of the text description as text examples;
[0006] After frame-cutting the video, obtain video segments, and extract video features; extract text features and text description features based on the text examples and the text description;
[0007] Obtain a noise representation through the text features, and input the noise representation into a noise mixer to combine with the original noise to obtain mixed noise;
[0008] Input the mixed noise, video features, and text features into a temporal diffusion network to obtain a hidden feature representation;
[0009] Input the hidden feature representation into a temporal nomination generation network to obtain Gaussian modeling parameters, and then construct a temporal segment mask;
[0010] Input the temporal segment mask, the video features, and the text description features into a semantic text reconstruction network for semantic text reconstruction and combinatorial ranking learning;
[0011] Based on the temporal segment mask and Gaussian modeling parameters, obtain the positions of video segments related to the text description, thereby completing content localization.
[0012] Further, generate the text examples through a large language model based on predetermined prompt words.
[0013] Further, extract the video features and the text features through a CLIP neural network.
[0014] Further, obtain predicted noise based on the hidden feature representation, and then optimize the video features.
[0015] Further, the temporal segment mask includes a target segment mask related to positive examples and a background target segment mask related to negative examples.
[0016] Further, the temporal diffusion network is a Transformer model with a two-stage structure; the temporal nomination generation network is a single-layer linear mapper; the semantic text reconstruction network is a Transformer model with a two-stage structure.
[0017] Further, the strategy of the combinatorial ranking learning adopts a Margin Ranking loss.
[0018] The present invention also provides a weakly supervised video temporal content localization system based on action nomination denoising, including:
[0019] A data acquisition module for acquiring a video to be processed and a text description corresponding to the video; generating positive and negative examples of the text description as text examples;
[0020] A feature extraction module for splitting frames of the video to obtain video segments and extracting video features; extracting text features and text description features based on the text examples and the text description;
[0021] A noise mixing module for obtaining a noise representation through the text features, and inputting the noise representation into a noise mixer to combine with the original noise to obtain mixed noise;
[0022] A temporal diffusion module for inputting the mixed noise, video features and text features into a temporal diffusion network to obtain a hidden feature representation;
[0023] A temporal nomination generation module for inputting the hidden feature representation into a temporal nomination generation network to obtain Gaussian modeling parameters, and further constructing a temporal segment mask;
[0024] A semantic text reconstruction module for inputting the temporal segment mask, the video features and the text description features into a semantic text reconstruction network for semantic text reconstruction and combined sorting learning;
[0025] A content localization module for obtaining the positions of video segments related to the text description based on the temporal segment mask and the Gaussian modeling parameters, thereby completing content localization.
[0026] The present invention also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the above method.
[0027] The present invention also provides a storage medium storing a computer program, and when the computer program is executed by a computer, the above method is implemented.
[0028] The beneficial effects of the present invention are as follows:
[0029] Using the method of the present invention, the target video segment that best matches the statement description can be found in a long video according to the text description. Compared with the prior art, it has the following advantages:
[0030] 1. The present invention proposes a method for action nomination denoising, which improves the quality of action nomination by optimizing the initial noise and the denoising process guided by text descriptions, making the time boundaries clear and having a high continuity, so as to assist in video segment localization.
[0031] 2. The present invention uses a combined ranking learning strategy to decouple the negative sample generation process, reduce the problem of misjudgment of temporal boundaries caused by semantic ambiguity of text descriptions, and improve the localization accuracy and generalization of the model.
[0032] 3. The present invention uses a weakly supervised learning mechanism to learn the model, and only uses video-level language description labels for training without using temporal labels, which greatly reduces the computational complexity and the time of temporal annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart using an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0034] The following further describes the present invention in detail through specific implementation examples and drawings.
[0035] The weakly supervised video temporal content localization method based on action nomination denoising provided by the present invention is applicable to temporal content localization of long videos, and the specific process of this method is as Figure 1 shown.
[0036] S100 Obtain the video to be processed and text descriptions.
[0037] The present invention uses the Charades-STA dataset as the data to be processed, and this dataset contains text-video pairs.
[0038] S200 Generate text examples.
[0039] In this step, a predefined prompt is used to guide a large language model (such as GPT-3.5, Gwen2) to expand text examples, including positive examples and negative examples. Positive examples are texts related to the given text annotation, obtained by inputting the given text description and a specific prompt template, and negative examples are texts unrelated to the given text annotation, obtained by replacing a certain proportion of verbs or nouns in the given text annotation.
[0040] The predefined prompt is: prompt=f"The following is a sentence describing the action that occurs in the video.Please provide 5 positive sample sentences (with the same meaning but significantly different structures) and 5 negative sample sentences (with completely different actions but similar scenes).The sentence is:{annotation}". Here, annotation is the given text description;
[0041] S300 extracts features.
[0042] For the video, perform RGB video frame slicing to obtain video segments.
[0043] For the text samples, perform word segmentation.
[0044] For the video segments, the text samples after word segmentation, and the text description, extract the features of the video and text modalities through the pre-trained CLIP network to obtain video features, text features, and text description features.
[0045] S400 processes noise.
[0046] Use a single-layer linear layer to transform the feature dimension and map the text features into the noise space to obtain the noise representation ∈ s . Then, the output noise representation ∈ s is combined with the original Gaussian noise · g through a noise mixer to form the mixed noise ∈ hybrid . The noise mixer is in the form of a weighted combination and is expressed as follows:
[0047]
[0048] where η is a hyperparameter.
[0049] S500 inputs the temporal diffusion network.
[0050] The above mixed noise and the features of the two modalities are input into the temporal diffusion network for cross-modal semantic feature interaction, and the fused latent encoded features are output. Then, the predicted noise is output through the noise predictor. The predicted noise indirectly optimizes the video features through the MSE loss function with the mixed noise. The temporal diffusion network is a two-level Transformer model, and the noise predictor is a single-layer linear mapper.
[0051] The noise addition process of the temporal diffusion network is shown as follows:
[0052]
[0053] where z t is the video feature at a certain time step t, β t is a predefined hyperparameter.
[0054] S600 is input into the temporal nomination generation network.
[0055] The obtained latent encoded features are input into the temporal nomination generation network of the single-layer linear mapper, and the central value and width value required for Gaussian mask modeling are output. Based on the central value and width value, a temporal segment mask related to the text example is constructed.
[0056] The target segment mask constructed based on the positive example is as follows:
[0057]
[0058] where, are the generated k-th temporal segment masks respectively, and are the central value and width of the k-th temporal segment mask, and σ is a hyperparameter. N is the number of video frames.
[0059] The background target segment masks constructed based on the negative examples are respectively expressed as:
[0060]
[0061] where, and respectively represent the temporal segment masks represented by the negative examples with low and high proportion replacements.
[0062] S700 is input into the semantic text reconstruction network.
[0063] For each temporal segment mask, it is input into the semantic text reconstruction network together with the video features and text description features for text semantic reconstruction and combined sorting learning to optimize the temporal nomination generation network. The text reconstructor is a two-level Transformer model.
[0064] The combined sorting learning strategy adopts a Margin Ranking loss, expressed as:
[0065]
[0066] where Y is the text description feature, and S p , and are the semantic text reconstruction expressions related to the positive example and two negative examples respectively, m 0 and m 1 are distance parameters, and d is a loss function for measuring the feature distribution, expressed as:
[0067]
[0068] S800 content positioning.
[0069] This step combines the mixed noise ∈ hybrid and the positive example of the text description into the time step in the denoising process, and applies S600 to obtain the accurate temporal Gaussian modeling parameters ( and ) to predict the center position and variance of the target video segment, and then obtain the start position and end position where the video segment related to the text description occurs, improving the average accuracy of video temporal content positioning.
[0070] The specific denoising process is expressed as follows:
[0071]
[0072] where m is a predefined hyperparameter, T is the denoising time step, and the DN function is expressed as:
[0073]
[0074] The process of predicting the start position and end position (τ s , τ e ) of the video segment is expressed as follows:
[0075]
[0076] where T v is the time length of the entire video.
[0077] The present invention utilizes a large language model to expand text information and generate corresponding positive and negative text examples to prompt and guide a temporal diffusion network to perceive corresponding inter-frame transformations, thereby perceiving dynamic events in a video. Among them, an optimized hybrid noise mechanism is proposed in the training stage to map text prompts into the noise space and combine them with traditional noise to encourage the model to perceive text-related temporal frame changes; in the denoising stage, text prompts are added as conditional features to the denoising process to generate action nominations related to text semantics. In this process, the present invention establishes a combined sorting learning strategy to decouple positive and negative samples, thereby refining the boundary of target event localization.
[0078] To evaluate the effectiveness of the method of the present invention, the temporal language localization evaluations of the present invention and the prior art are calculated respectively. R@n, IoU=m represents the proportion of results with an intersection over union (IoU) metric greater than m (∈(0,1]) among the top n returned results in the total n returned results. The larger the value of the evaluation index, the better the performance of the method. Taking the Charades-STA dataset as an example, the results of temporal language localization are shown in Table 1:
[0079] Method R@1, IoU = 0.3 R@1, IoU = 0.5 R@1, IoU = 0.7 SCN 42.46 23.58 9.97 CPL 66.40 49.24 22.39 The method of the present invention 69.41 50.51 23.21
[0080] Table 1
[0081] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A weakly supervised video temporal content localization method based on action nomination denoising, the steps of which include: Obtain a video to be processed and a text description corresponding to the video; generate positive samples and negative samples of the text description as text samples; Cut the video into frames to obtain video clips, and extract video features; Extracting text features and text description features based on the text sample and the text description; Obtaining noise expression through the text features, inputting the noise expression into a noise mixer and combining it with the original noise to obtain mixed noise; Inputting the mixed noise, video features and text features into a temporal diffusion network to obtain latent feature expression; Inputting the latent feature expression into a temporal nomination generation network to obtain Gaussian modeling parameters, and then constructing a temporal segment mask; inputting the temporal segment mask, the video feature and the text description feature into a semantic text reconstruction network to perform semantic text reconstruction and combination sorting learning; The position of the video segment related to the text description is obtained based on the temporal segment mask and Gaussian modeling parameters, thereby completing content positioning.
2. The method according to claim 1, characterized in that The text sample is generated based on a predetermined prompt word by a large language model.
3. The method according to claim 1, characterized in that: The video features and the text features are extracted by using a CLIP neural network.
4. The method according to claim 1, characterized in that: The predicted noise is obtained based on the latent feature expression, and then the video feature is optimized.
5. The method according to claim 1, characterized in that The temporal segment masks include a target segment mask associated with a positive example and a background target segment mask associated with a negative example.
6. The method according to claim 1, characterized in that The temporal diffusion network is a Transformer model with a two-level structure; the temporal nomination generation network is a single-layer linear mapper; and the semantic text reconstruction network is a Transformer model with a two-level structure.
7. The method according to claim 1, characterized in that The strategy of combined ranking learning adopts MarginRanking loss.
8. A weakly supervised video temporal content localization system based on action nomination denoising, comprising: A data acquisition module, used to acquire a video to be processed and a text description corresponding to the video; Generating positive examples and negative examples of the text description as text examples; A feature extraction module, used to obtain video segments after cutting the video into frames, and extract video features; Extracting text features and text description features based on the text sample and the text description; A noise mixing module, used for obtaining a noise expression through the text feature, and inputting the noise expression into a noise mixer to combine it with the original noise to obtain a mixed noise; A time diffusion module, used for inputting the mixed noise, video features and text features into a time diffusion network to obtain latent feature expression; A temporal nomination generation module, used for inputting the latent feature expression into the temporal nomination generation network to obtain Gaussian modeling parameters, and then constructing a temporal segment mask; A semantic text reconstruction module, used for inputting the temporal segment mask, the video feature and the text description feature into a semantic text reconstruction network for semantic text reconstruction and combination sorting learning; The content positioning module is used to obtain the position of the video segment related to the text description based on the temporal segment mask and Gaussian modeling parameters, thereby completing content positioning.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.
10. A storage medium storing a computer program, wherein when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.