A Weakly Supervised Temporal Action Localization Method and Device Based on Pseudo-Label Generation
Through the dual-branch framework and pseudo-label proposal fusion strategy, combined with the uncertainty masking mechanism, the problems of low quality of pseudo-label proposals and noise pseudo-label interference in timing action positioning are solved, and high-precision and robust timing action positioning are achieved.
Patent Information
- Application Number
- CN202510217146.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In the prior art, in the timing action positioning, there are problems such as low quality of pseudo-label proposals, insufficient utilization of prior information, and difficulty in combining with the full supervision framework, resulting in inaccurate positioning and serious interference from noise pseudo-labels.
Using a dual-branch framework, through the combination of weakly supervised branches and full-supervised branches, the pseudo-label proposal fusion strategy and uncertainty masking mechanism are used to dynamically optimize the full-supervised branches, generate high-quality pseudo-label proposals, and reduce interference from noise tags during training.
The quality and positioning accuracy of the pseudo-label proposal are improved, the robustness and stability of the model are enhanced, and the higher precision timing action positioning is achieved under the condition of less labeled data.
Smart Images

Figure CN119723678B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of temporal action localization, and particularly to a weakly supervised temporal action localization method and device based on pseudo-label generation. Background Art
[0002] Temporal Action Localization (TAL) is an important task in the field of video analysis, aiming to identify all actions in a video and accurately locate their start and end times. With the rapid development of deep learning technologies, especially the application of technologies such as convolutional neural networks and long short-term memory networks, significant progress has been made in the field of temporal action localization. In traditional TAL methods based on handcrafted features, action recognition and localization rely on manually designed features, and these methods usually require a large amount of manual intervention and cannot handle complex video scenarios.
[0003] In recent years, deep learning-based TAL methods have gradually become the mainstream. Current TAL methods can generally be divided into two-stage methods and single-stage methods. Two-stage methods locate action boundaries by generating candidate action proposals and performing refined adjustments; for example, works based on proposal generation, such as Boundary-Sensitive Network (BSN) and Boundary-Matching Network (BMN). Single-stage methods directly perform action classification and localization from video sequences, simplifying the design and training of the model. For example, the Transformer-based temporal action localization model ActionFormer and the single-stage framework TriDet for temporal action detection.
[0004] Temporal action localization is a complex and challenging visual task, with the goal of accurately classifying and localizing actions in untrimmed videos. However, temporal action localization has always faced the problem of inaccurate localization, which makes it difficult to accurately predict the start and end times of actions. Although existing fully supervised and weakly supervised methods have been improved to a certain extent, they still face problems such as relying on annotation information and noisy pseudo-labels. These problems not only affect the generalization ability and accuracy of the model, but also limit the application of the method on large-scale video datasets.
[0005] Fully supervised temporal action localization methods rely on complete annotation data, including the category, start time, and end time of each action. This method can usually provide relatively accurate localization results because it makes full use of accurate annotation information. However, the high dependence of fully supervised methods on annotation data makes large-scale applications difficult. The collection of annotation data is not only time-consuming and costly, especially when a large amount of video data is required, it is almost impossible to achieve. The generalization ability of fully supervised methods is weak. When the annotation data is limited, the training results of the model are difficult to generalize to unseen data. Therefore, the insufficient generalization ability of fully supervised methods becomes their main limiting factor.
[0006] Different from fully supervised methods, weakly supervised temporal action localization methods are trained by using relatively simple global labels or a small number of segment labels. Although this method significantly reduces the annotation cost, it still faces great challenges in terms of accuracy and stability. Due to the lack of accurate segment-level labels, the model can only use global information for learning during the training process. This makes its localization accuracy relatively low in complex scenarios, especially when the boundaries between actions are blurred and actions overlap, and it cannot be compared with fully supervised methods.
[0007] The temporal action localization method based on pseudo-label generation is a bridge connecting fully supervised temporal action localization and weakly supervised temporal action localization. When the samples are unlabeled, generating pseudo-labels through a confidence score threshold is a common pseudo-label generation method, and this method is also used in traditional weakly supervised methods. However, this direct strategy may not be able to accurately capture the complex temporal information in the video, resulting in low-quality generated proposals. In addition, complex proposal fusion methods are highly dependent on hyperparameters. If the hyperparameters are set improperly, it may lead to over-segmentation or omission of important proposals, affecting the final localization accuracy. This dependence on hyperparameter adjustment increases the complexity of the model and limits its adaptability in different datasets and scenarios.
[0008] The quality of pseudo-label proposals is crucial for the final performance of the model. Existing methods fail to effectively address the problem of pseudo-label noise. They directly use pseudo-labels when training the model with pseudo-labels without effectively processing the low-quality boundary regions, which often contain a lot of noise, resulting in the negative impact of pseudo-label noise on the training process. Especially in the early stage of training, the early prediction results of the model may contain a lot of noise, directly affecting the stability of the model. Due to the lack of a dynamic adjustment and screening mechanism for noisy pseudo-labels, these methods are prone to learning incorrect boundary information during the training process, resulting in unsatisfactory training effects of the model.
[0009] In the existing technology, there is a lack of a weakly supervised temporal action localization method that can generate high-quality pseudo-labels, make full use of prior information, and effectively combine fully supervised methods. Summary of the Invention
[0010] To solve the technical problems of insufficient utilization of prior information, low quality of pseudo-label proposals, and difficulty in combining with the full supervision framework in the existing technology, an embodiment of the present invention provides a weakly supervised temporal action localization method and device based on pseudo-label generation. The technical solution is as follows:
[0011] On the one hand, a weakly supervised temporal action localization method based on pseudo-label generation is provided. This method is implemented by a weakly supervised temporal action localization device, and the method includes:
[0012] Obtain an unclipped video containing actions; perform segment division on the unclipped video to obtain a set of video segments;
[0013] According to the set of video segments, perform feature extraction through a feature extractor to obtain a set of per-segment features;
[0014] According to the set of per-segment features, perform preliminary action classification through a weakly supervised branch to obtain a classification attention sequence and a multi-scale set of action proposals;
[0015] According to the classification attention sequence and the set of action proposals, use a proposal fusion strategy to optimize the proposal fusion and obtain a set of pseudo-label proposals;
[0016] Based on a preset dilation ratio and a preset contraction ratio, generate a set of uncertainty masks at the time boundaries of each proposal in the set of action proposals;
[0017] Based on the attention mechanism, according to the set of pseudo-label proposals and the set of uncertainty masks, optimize and train the full supervision branch through a dynamic optimization mechanism to obtain a second optimized full supervision branch;
[0018] Obtain a video of the action to be located; based on the feature extractor, the weakly supervised branch, and the second optimized full supervision branch, perform action localization according to the video of the action to be located.
[0019] On the other hand, a weakly supervised temporal action localization device based on pseudo-label generation is provided. This device is applied to the weakly supervised temporal action localization method based on pseudo-label generation, and the device includes:
[0020] A video segment acquisition module, configured to obtain an unclipped video containing actions; perform segment division on the unclipped video to obtain a set of video segments;
[0021] A video feature extraction module, configured to perform feature extraction through a feature extractor according to the set of video segments to obtain a set of per-segment features;
[0022] A weakly supervised branch classification module, which is used to perform preliminary action classification through a weakly supervised branch according to the set of segment-by-segment features, and obtain a classification attention sequence and a multi-scale set of action proposals;
[0023] A pseudo-label proposal generation module, which is used to perform proposal fusion optimization using a proposal fusion strategy according to the classification attention sequence and the set of action proposals, and obtain a set of pseudo-label proposals;
[0024] An uncertainty mask generation module, which is used to generate a set of uncertainty masks at the time boundaries of each proposal in the set of action proposals based on a preset dilation ratio and a preset contraction ratio;
[0025] A fully supervised branch optimization module, which is used to optimize and train the fully supervised branch through a dynamic optimization mechanism based on an attention mechanism according to the set of pseudo-label proposals and the set of uncertainty masks, and obtain a second optimized fully supervised branch;
[0026] A video action localization module, which is used to obtain an action video to be localized; and perform action localization on the action video to be localized based on the feature extractor, the weakly supervised branch, and the second optimized fully supervised branch.
[0027] On the other hand, a weakly supervised temporal action localization device is provided. The weakly supervised temporal action localization device includes: a processor; a memory, and a computer-readable instruction is stored on the memory. When the computer-readable instruction is executed by the processor, any method in the weakly supervised temporal action localization method based on pseudo-label generation as described above is implemented.
[0028] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any method in the weakly supervised temporal action localization method based on pseudo-label generation as described above.
[0029] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0030] The present invention proposes a weakly supervised temporal action localization method based on pseudo-label generation. Through a dual-branch framework, the gap between weakly supervised and fully supervised temporal action localization in terms of framework and performance is bridged; by using the information in the weakly supervised branch and the fully supervised branch at the same time, the dual-branch framework can achieve more accurate temporal action localization under the condition of less labeled data; a pseudo-label proposal fusion strategy is adopted to effectively use the prior information in the pseudo-label to generate pseudo-label proposals with high confidence and accurate boundaries; this strategy significantly improves the quality of the proposal and the accuracy of the model through the fusion of multi-scale information; an uncertainty mask is introduced, and a new optimization mechanism is proposed for effective learning in the presence of noisy labels. The mechanism effectively reduces the interference of noisy labels during the training process by dynamically adjusting the quality of pseudo-labels, enhances the robustness of the regression model, and ensures that the model can converge stably during the training process. The present invention is a weakly supervised temporal action localization method for generating high-quality pseudo-labels that fully utilizes prior information and effectively combines the full supervision method. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0032] Figure 1 It is a flow chart of a weakly supervised temporal action localization method based on pseudo-label generation provided by an embodiment of the present invention;
[0033] Figure 2 It is a block diagram of a weakly supervised sequential action positioning device based on pseudo-label generation provided by an embodiment of the present invention;
[0034] Figure 3 It is a structural diagram of a weakly supervised sequential action positioning device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0035] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0036] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0037] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.
[0038] In the embodiments of the present invention, sometimes subscripts such as W 1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0039] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0040] The embodiments of the present invention provide a weakly supervised temporal action localization method based on pseudo-label generation. This method can be implemented by a weakly supervised temporal action localization device, which can be a terminal or a server. As Figure 1 shown in the flowchart of the weakly supervised temporal action localization method based on pseudo-label generation, the processing flow of this method can include the following steps:
[0041] S1. Obtain an unclipped video containing actions; perform segment division on the unclipped video to obtain a set of video segments.
[0042] In a feasible implementation, the present invention performs segmentation operations on the input unclipped video. The main purpose is to reduce the amount of data that the model needs to process while retaining local temporal information. The video segments obtained in this step are regarded as the smallest processing units that cannot be further divided in subsequent steps.
[0043] S2. According to the set of video segments, perform feature extraction through a feature extractor to obtain a set of per-segment features.
[0044] In a feasible implementation, in the feature extraction module, the present invention uses a deep learning model for video action recognition (Inflated 3D ConvNet, I3D) pre-trained on the Kinetics dataset to extract RGB and optical flow features. The extracted feature dimension is 1024, and a segment feature is extracted every 16 frames.
[0045] S3. According to the set of per-segment features, perform preliminary action classification through a weakly supervised branch to obtain a classification attention sequence and a multi-scale set of action proposals.
[0046] Optionally, according to the set of per-segment features, preliminary action classification is performed through a weakly supervised branch to obtain a classification attention sequence and a set of multi-scale action proposals, including:
[0047] Based on the attention mechanism, according to the set of per-segment features, adaptive action classification is performed through a base model to obtain an attention score sequence and a classification score matrix;
[0048] Weighted fusion is performed according to the attention score sequence and the classification score matrix to obtain a classification attention sequence;
[0049] Based on a preset score threshold, thresholding is performed on the classification attention sequence to obtain a set of multi-scale action proposals; the set of action proposals includes a set of time boundaries and a set of action categories.
[0050] In a feasible implementation manner, the present invention uses a baseline model trained based on multi-instance learning as a basic module of the weakly supervised learning framework, and uses the basic module to obtain the features and action recognition results of each video segment.
[0051] Among them, the base model is a baseline model based on the multi-instance learning training method.
[0052] In a feasible implementation manner, in the present invention, the weakly supervised branch adopts a framework for weakly supervised temporal action localization (Dual-Evidential Learning for Uncertainty modeling, DELU). This framework gradually focuses on the entire action instance by introducing video-level and segment-level uncertainties, thereby solving the problems of background noise and action-background blur.
[0053] S4. According to the classification attention sequence and the set of action proposals, use a proposal fusion strategy to optimize proposal fusion to obtain a set of pseudo-label proposals.
[0054] Optionally, according to the classification attention sequence and the set of action proposals, use a proposal fusion strategy to optimize proposal fusion to obtain a set of pseudo-label proposals, including:
[0055] Based on the set of action proposals, calculations are performed according to the classification attention sequence to obtain the mean of the internal attention scores of the proposals and the mean of the external attention scores of the proposals;
[0056] Based on the mean of the internal attention scores of the proposals and the mean of the external attention scores of the proposals, confidence calculations are performed according to the classification attention sequence to obtain a set of proposal confidence scores;
[0057] According to the set of action proposals and the set of proposal confidence scores, fusion optimization is performed through a proposal fusion strategy to obtain a set of pseudo-label proposals; the proposal fusion strategy refers to an optimization strategy that fuses the action proposals mapped to the global shared space with the confidence scores corresponding to the action proposals; the proposal fusion strategy is used to filter the misclassified results of the action proposals.
[0058] In a feasible implementation manner, the design focus of the present invention is the generation of temporal pseudo-labels, and the key to this step lies in how to use the constraint information existing in the video and introduce artificial prior knowledge to improve the quality and reliability of the pseudo-labels.
[0059] Taking the output of the base model as the input of the temporal pseudo-label construction module, using the constraint information existing in the video temporal sequence and the artificially introduced prior information, optimizing the pseudo-supervised information through the proposal fusion strategy, fusing information of different scales into a global shared space, reducing noise and improving the quality of the pseudo-labels.
[0060] This method generates accurate and reliable pseudo-labels by performing threshold processing on the proposals. The obtained pseudo-supervised information has a form similar to that of real temporal supervision labels, that is, it can provide a confidence score similar to the intersection over union for any action proposal. Through maximizing regression training on this score, the localization task can be introduced into the model training, reducing the gap between the training task and the test task of the weakly supervised video action understanding model.
[0061] The proposals output by the base model of the weakly supervised branch differ in scale, and a higher threshold will result in a smaller scale perception. The proposals are generated separately, ignoring the overall relationship between different proposals and scales. In addition, there is overlap between the action proposals, which makes it impossible to accurately describe the boundaries of the action proposals, thereby affecting the training of the fully supervised branch.
[0062] Global perception is obtained by mapping all proposals to a unified space, enhancing spatial understanding and simplifying the representation. The present invention uses the Ricker wavelet, which emphasizes the central region of each proposal while suppressing the peripheral region, effectively fusing proposals of different perception scales. This method enhances the effect of temporal action localization by achieving balanced and robust multi-scale proposal alignment.
[0063] For each action proposal, map it to the Ricker distribution space as shown in the following formula (1):
[0064] (1);
[0065] Where represents the Ricker distribution value at a certain time point . For the th action proposal , Indicates the start time of the action proposal of the action proposal, Indicates the start time of the action proposal of the action proposal, Indicates the end time of the action proposal of the action proposal, Indicates the length of the action proposal of the action proposal,
[0066] The Ricker wavelet can naturally focus on the central part of the proposal while suppressing the boundary region. Its suppression effect gradually increases in the region far from the center and then gradually weakens, which helps to reduce noise and reduce the uncertainty of the boundary region. Such characteristics are highly consistent with the positioning results because it ensures that each proposal can capture the core action with a high confidence level, while the influence of the boundary part is effectively suppressed.
[0067] Fuse all proposals into a shared global space, and the process is as follows in Equation (2):
[0068] (2);
[0069] where represents the final fused wavelet distribution of the proposal with category for integrating the global representation of all proposals. represents the number of all proposals, represents the category of the th proposal, represents a specified action category; all action proposals are fused together with the corresponding confidence score and the misclassified predictions will be filtered out. The final pseudo-label is obtained from this, and the generated pseudo-proposals will have better quality and more reliable boundaries.
[0070] Compared with other pseudo-label proposal fusion methods, the existing pseudo-label proposal fusion methods are highly dependent on hyperparameters. If the hyperparameters are set improperly, it may lead to over-segmentation or omission of important proposals, affecting the final positioning accuracy. This dependence on hyperparameter adjustment increases the complexity of the method and limits its adaptability in different datasets and scenarios.
[0071] To improve the quality of pseudo-label proposals, an efficient proposal fusion method proposed by the present invention reduces noise and enhances the correlation between proposals by fusing proposals of different scales into a unified global space. The confidence of action proposals is evaluated using the external-internal contrast score of the classification attention sequence and mapped to the Ricker wavelet distribution to fuse multi-scale information into a unified representation. High-quality proposals are generated through threshold processing. Compared with traditional methods, the method proposed by the present invention can effectively integrate predictions from different scales, reduce the influence of noise, and significantly improve the accuracy and localization precision of pseudo-label proposals.
[0072] S5. Generate a set of uncertainty masks at the temporal boundaries of each proposal in the action proposal set based on a preset dilation ratio and a preset contraction ratio.
[0073] In a feasible implementation, to improve the training process based on noisy pseudo-labels, the present invention utilizes a prior assumption based on proposal accuracy: the central region of a proposal is usually more reliable, while there is more uncertainty in the boundary part. Based on this assumption, an uncertainty mask is introduced, and the mathematical expression of the uncertainty mask is as follows in Equation (3):
[0074] (3);
[0075] Where, denotes the uncertainty mask for the action proposal , denotes the interval of the uncertain time region of the proposal , which contains two sub-intervals and , denotes the uncertain region near the start time of the proposal , denotes the uncertain region near the end time of the proposal , and are the preset dilation ratio and the preset contraction ratio respectively. For the th proposal , denotes the start time of the proposal , denotes the end time of the proposal . The region around the boundary of each proposal is identified as the uncertainty region, and the uncertainty region in the action proposal may contain more noise than other regions.
[0076] All the masks are combined, and the process is as follows in Equation (4):
[0077] (4);
[0078] Among them, represents the global uncertainty mask, represents the number of anchor boxes connecting all layers, represents the number of proposals for all layers. The uncertainty mask is generated at multiple scales after assigning proposals to different feature pyramid levels.
[0079] During subsequent training, the boundaries of each proposal will be partially masked to avoid interference with model training. The uncertainty mask is designed to be adjusted over time. As the model's confidence in the proposal boundaries increases, the uncertain regions gradually decrease through a simple and effective linear function. As training progresses, the accuracy of the proposal boundaries improves and the mask gradually fades, enabling the entire proposal to participate in training, ultimately reducing the impact of noisy labels and improving the overall performance of the model.
[0080] S6. Based on the attention mechanism, according to the pseudo-label proposal set and the uncertainty mask set, optimize and train the fully supervised branch through a dynamic optimization mechanism to obtain a second optimized fully supervised branch.
[0081] Optionally, based on the attention mechanism, according to the pseudo-label proposal set and the uncertainty mask set, optimize and train the fully supervised branch through a dynamic optimization mechanism to obtain a second optimized fully supervised branch, including:
[0082] Based on the morphological preprocessing method, filter and optimize the pseudo-label proposal set to obtain an optimized action proposal set;
[0083] Input the optimized action proposal set into the encoder for data processing to obtain an encoded pseudo-label proposal set;
[0084] According to the encoded pseudo-label proposal set, perform multi-scale feature extraction through a feature pyramid network to obtain a pseudo-label proposal feature map;
[0085] Based on the uncertainty mask set, occlude the pseudo-label proposal feature map to obtain a first unoccluded proposal feature map;
[0086] Based on a preset first training round, according to the first unoccluded proposal feature map, perform the first-stage training through the fully supervised branch to obtain a first predicted action proposal set;
[0087] Based on the attention mechanism, according to the pseudo-label proposal set, dynamically optimize the first predicted action proposal set and the uncertainty mask set to obtain a second predicted action proposal set and an updated uncertainty mask set;
[0088] Calculate the loss function according to the second predicted action proposal set and the pseudo-label proposal set to obtain a first action loss;
[0089] Optimize the parameters of the fully supervised branch according to the first action loss to obtain the first optimized fully supervised branch;
[0090] Based on the updated set of uncertainty masks, occlude the pseudo-label proposal feature map to obtain the second unoccluded proposal feature map;
[0091] Based on the preset second training round, according to the second unoccluded proposal feature map, perform the second-stage training through the first optimized fully supervised branch to obtain the third set of predicted action proposals;
[0092] Calculate the loss function according to the third set of predicted action proposals and the set of pseudo-label proposals to obtain the second action loss;
[0093] Optimize the parameters of the first optimized fully supervised branch according to the second action loss to obtain the second optimized fully supervised branch.
[0094] In a feasible implementation, the pseudo-label proposals are simply preprocessed through morphological preprocessing. The basic assumption it relies on is that video actions have a certain temporal continuity and will not start and end repeatedly in a short period of time.
[0095] Improve the performance of the model by fusing multiple prior information such as the classification attention sequence and pseudo-label proposals. After sharing the extracted segment features with the weakly supervised branch, first use the transformer encoder to encode the features.
[0096] Use a feature pyramid network to perform multi-scale processing on the encoded features to generate anchor boxes. According to the duration of each proposal, all pseudo-label proposals are assigned to different scales. And predict the category of each anchor box through the classification head. For each anchor box, the start time and end time of the proposal are decoded through the regression head.
[0097] If these pseudo-label proposals containing noise are directly used for training, it may mislead the model. Therefore, a denoising mechanism needs to be introduced during the training process to reduce the impact of noisy pseudo-labels. The present invention realizes dynamic denoising processing of pseudo-labels by designing and introducing an uncertainty mask module. During the training of the fully supervised branch, the mask module will be dynamically adjusted according to the confidence of the pseudo-label proposal boundary, excluding the influence of those low-confidence regions. When the model calculates the loss, the loss generated in the low-confidence region will be ignored, and the focus will be on the loss in the high-confidence region, thus effectively avoiding the interference of the noise information in the pseudo-labels on the model training.
[0098] Compared with a single fully supervised or weakly supervised framework, in order to effectively learn various prior knowledge in pseudo-labels, bridge the framework gap between weakly supervised and fully supervised temporal action localization, and learn from both proposal-level and segment-level information.
[0099] In the fully-supervised branch, a regression model similar to the mainstream fully-supervised temporal action localization architecture is introduced to utilize the proposal-level pseudo-labels generated in the weakly-supervised branch. It can learn the boundaries and durations of actions like fully-supervised temporal action localization methods. Additionally, by learning from the classification attention sequence in the weakly-supervised branch, it can enrich its encoded features with segment-level classification information. Each pseudo-label contains natural priors introduced through the multi-instance learning process and artificial priors added by design, enhancing the overall performance of the proposed method.
[0100] The focus of this step is on the design of a denoising mechanism based on the uncertainty mask. The parameters in the mask can dynamically adjust the model's attention to pseudo-labels. In the first-stage training, it mainly relies on the high-confidence parts of the pseudo-labels for learning and excludes less reliable label information. In the second-stage training, the model's confidence in the pseudo-labels gradually increases, and the applicable range of the mask gradually shrinks, ultimately achieving the effective utilization of the entire pseudo-labels. Through this mechanism, the model can adaptively identify and exclude the noise in the pseudo-labels, improving the robustness and accuracy of the training process.
[0101] Compared with traditional methods for processing noisy pseudo-labels, traditional methods lack an adaptive noisy label training mechanism and often cannot dynamically identify and exclude unreliable pseudo-label boundary regions, resulting in noise having a greater impact on the model's performance.
[0102] In addition to the uncertainty mask and dynamic optimization mechanism, the present invention introduces an attention mechanism to achieve the dynamic evaluation of pseudo-labels during the training process. When calculating the training loss of the model, it ignores the loss generated by pseudo-labels with attention below the threshold and focuses on the loss generated by pseudo-labels with attention above the threshold. Thus, during the training process, it dynamically avoids the misleading of the model training by the error information in pseudo-supervision.
[0103] Optionally, based on the attention mechanism, according to the pseudo-label proposal set, the first predicted action proposal set and the uncertainty mask set are dynamically optimized to obtain a second predicted action proposal set and an updated uncertainty mask set, including:
[0104] Based on the attention mechanism, calculate according to the pseudo-label proposal set and the first predicted action proposal set to obtain a first action proposal attention sequence;
[0105] Based on a preset attention threshold, screen the first action proposal attention sequence to obtain a second action proposal attention sequence;
[0106] Based on the second action proposal attention sequence, select the corresponding action set in the first predicted action proposal set to obtain a second predicted action proposal set;
[0107] Based on a preset linear function, according to the second action proposal attention sequence, dynamically linearly decay the mask weights of the uncertainty mask set to obtain an updated uncertainty mask set; the range of the mask weights is from 1 to 0.
[0108] In a feasible implementation, during the training phase of the fully supervised branch, the pseudo-label proposals and uncertainty masks are optimized after the first-phase training. The pseudo-label proposals in the first-phase training are generated by the base model. After the first-phase training, the proposals generated by the regression model are considered to have more reliable boundaries and confidence scores. In each training iteration, the proposals output by the regression model in the fully supervised branch are combined with the initial pseudo-label proposals and updated. The updated proposal set is used to optimize the uncertainty mask, thereby helping the model to more accurately locate the temporal boundaries.
[0109] After the training enters the second-phase training, the parameters in the mask are linearly decayed. This change indicates that as the training progresses, the reliability and accuracy of the proposal boundaries are gradually improved. Therefore, the model will gradually reduce its dependence on the uncertainty mask and instead rely more on the optimized proposals to make the final prediction.
[0110] S7. Obtain the video of the action to be located; based on the feature extractor, the weakly supervised branch, and the second optimized fully supervised branch, perform action localization according to the video of the action to be located.
[0111] Optionally, performing action localization according to the video of the action to be located based on the feature extractor, the weakly supervised branch, and the second optimized fully supervised branch includes:
[0112] Performing temporal action localization prediction based on the feature extractor, the weakly supervised branch, and the second optimized fully supervised branch to obtain a set of predicted action proposals;
[0113] Based on the attention mechanism, according to the set of predicted action proposals, filter out redundant action proposals through the soft non-maximum suppression algorithm to obtain the filtered action proposals;
[0114] Perform action localization according to the action category and temporal boundaries of the filtered action proposals.
[0115] In a feasible implementation, after two-phase training, use the regression-based student model in the fully supervised branch for inference. To obtain the video-level classification result, use the attention scores predicted by the attention head, select the top k scores for each category, and average them. Use the predicted attention scores to filter the video categories and only retain the proposals within the predicted categories. For categories below the threshold, filter out their proposals. And apply soft non-maximum suppression to filter out redundant action proposals, and finally obtain the predicted action proposal result.
[0116] In a feasible implementation, the present invention uses the DELU model as the base model for the weakly supervised branch and constructs the regression model for the fully supervised branch with reference to the ActionFormer and TriDet models. In the feature extraction module, the I3D model pre-trained on the Kinetics dataset is used to extract RGB and optical flow features, and the extracted feature dimension is 1024. For the THUMOS14 dataset, according to the method of the ActionFormer model, a segment feature is extracted every 16 frames with a stride of 4 (the base model stride is 16); for the ActivityNet1.3 dataset, a segment feature is extracted every 16 frames. For the THUMOS14 dataset, the second training round is set to 38; for the ActivityNet1.3 dataset, the second training round is set to 16, and the first rounds are 20 and 10 respectively. The learning rates for THUMOS14 and ActivityNet1.3 are 1e-4 and 1e-3 respectively. For the uncertainty mask, the dilation ratio for the THUMOS14 dataset is 0.1, and the contraction ratio is 0; for the ActivityNet1.3 dataset, the dilation ratio is 0.05, and the contraction ratio is 0.05.
[0117] In a feasible implementation, the performance of the method described in the present invention is compared with that of the current advanced methods on the THUMOS14 dataset, and the results are shown in Table 1 (Comparison Table of Multi-Method Action Localization Results on the THUMOS14 Dataset). It can be seen from Table 1 that the mean average precision of the method of the present invention in the interval (0.1:0.7) is increased by 5.5% compared with the existing benchmark model, the double evidence learning framework, and the mean average precision in other intervals also shows a similar improvement. Compared with the method guided by prior information, the present invention significantly outperforms in all intersection over union thresholds (from 0.1 to 0.7), and the gap is large. It can be seen from Table 1 that the performance of the method adopted by the present invention is close to that of the early point supervision method, indicating that the gap between weakly supervised temporal action localization and fully supervised temporal action localization has been narrowed.
[0118] Table 1
[0119]
[0120] The performance of the present invention and other existing methods on the ActivityNet1.3 dataset is compared. The results are shown in Table 2 (Comparison of action localization results of multiple methods on the ActivityNet1.3 dataset). The results show a similar trend: the performance of the present invention is improved by 0.9% compared with the current best method based on prior information guidance. The method proposed by the present invention exceeds some recent point supervision methods. As can be seen from Table 2, the excellent performance of the present invention method on the challenging ActivityNet1.3 dataset further confirms the strong robustness and high generalization ability.
[0121] In tests conducted on two benchmark test sets, THUMOS14 and ActivityNet1.3, the proposed method significantly outperforms existing methods in performance, has wide applicability, and bridges the gap between weakly supervised and fully supervised temporal action localization.
[0122] Table 2
[0123]
[0124] The present invention proposes a weakly supervised temporal action localization method based on pseudo-label generation. Through a dual-branch framework, the gap between weakly supervised and fully supervised temporal action localization in terms of framework and performance is bridged; by using the information in the weakly supervised branch and the fully supervised branch at the same time, the dual-branch framework can achieve more accurate temporal action localization under the condition of less labeled data; a pseudo-label proposal fusion strategy is adopted to effectively use the prior information in the pseudo-label to generate pseudo-label proposals with high confidence and accurate boundaries; this strategy significantly improves the quality of the proposal and the accuracy of the model through the fusion of multi-scale information; an uncertainty mask is introduced, and a new optimization mechanism is proposed for effective learning in the presence of noisy labels. The mechanism effectively reduces the interference of noisy labels during the training process by dynamically adjusting the quality of pseudo-labels, enhances the robustness of the regression model, and ensures that the model can converge stably during the training process. The present invention is a weakly supervised temporal action localization method for generating high-quality pseudo-labels that fully utilizes prior information and effectively combines the full supervision method.
[0125] Figure 2 1 is a block diagram of a weakly supervised temporal action positioning device based on pseudo-label generation according to an exemplary embodiment, wherein the device is used in a weakly supervised temporal action positioning method based on pseudo-label generation. Figure 2 The device includes a video clip acquisition module 210, a video feature extraction module 220, a weakly supervised branch classification module 230, a pseudo-label proposal generation module 240, an uncertainty mask generation module 250, a fully supervised branch optimization module 260 and a video action positioning module 270. Among them:
[0126] The video clip acquisition module 210 is used to acquire an unclipped video containing actions; segment the unclipped video to obtain a set of video clips;
[0127] The video feature extraction module 220 is used to extract features through a feature extractor according to the set of video clips to obtain a per-clip feature set;
[0128] The weakly-supervised branch classification module 230 is used to perform preliminary action classification through the weakly-supervised branch according to the per-clip feature set to obtain a classification attention sequence and a multi-scale action proposal set;
[0129] The pseudo-label proposal generation module 240 is used to fuse and optimize the proposals using a proposal fusion strategy according to the classification attention sequence and the action proposal set to obtain a pseudo-label proposal set;
[0130] The uncertainty mask generation module 250 is used to generate an uncertainty mask set at the temporal boundaries of each proposal in the action proposal set based on a preset dilation ratio and a preset contraction ratio;
[0131] The fully-supervised branch optimization module 260 is used to optimize and train the fully-supervised branch through a dynamic optimization mechanism based on the attention mechanism according to the pseudo-label proposal set and the uncertainty mask set to obtain a second optimized fully-supervised branch;
[0132] The video action localization module 270 is used to acquire the video of the action to be localized; perform action localization on the video of the action to be localized based on the feature extractor, the weakly-supervised branch, and the second optimized fully-supervised branch.
[0133] Optionally, the weakly-supervised branch classification module 230 is further used for:
[0134] Based on the attention mechanism, perform adaptive action classification through a base model according to the per-clip feature set to obtain an attention score sequence and a classification score matrix;
[0135] Perform weighted fusion according to the attention score sequence and the classification score matrix to obtain a classification attention sequence;
[0136] Based on a preset action threshold, perform thresholding on the classification attention sequence to obtain a multi-scale action proposal set; the action proposal set includes a temporal boundary set and an action category set.
[0137] Wherein, the base model is a baseline model based on a multi-instance learning training method.
[0138] Optionally, the pseudo-label proposal generation module 240 is further used for:
[0139] Based on the set of action proposals, calculate according to the classification attention sequence to obtain the mean of the attention scores inside the proposals and the mean of the attention scores outside the proposals;
[0140] Based on the mean of the attention scores inside the proposals and the mean of the attention scores outside the proposals, calculate the confidence according to the classification attention sequence to obtain the set of proposal confidence scores;
[0141] According to the set of action proposals and the set of proposal confidence scores, perform fusion optimization through the proposal fusion strategy to obtain the set of pseudo-label proposals; the proposal fusion strategy refers to the optimization strategy of fusing the action proposals mapped to the global shared space with the confidence scores corresponding to the action proposals; the proposal fusion strategy is used to filter the misclassified results of the action proposals.
[0142] Optionally, the fully supervised branch optimization module 260 is further used for:
[0143] Based on the morphological preprocessing method, filter and optimize the set of pseudo-label proposals to obtain the optimized set of action proposals;
[0144] Input the optimized set of action proposals into the encoder for data processing to obtain the encoded set of pseudo-label proposals;
[0145] According to the encoded set of pseudo-label proposals, perform multi-scale feature extraction through the feature pyramid network to obtain the pseudo-label proposal feature map;
[0146] Based on the set of uncertainty masks, occlude the pseudo-label proposal feature map to obtain the first unoccluded proposal feature map;
[0147] Based on the preset first training round, according to the first unoccluded proposal feature map, perform the first-stage training through the fully supervised branch to obtain the first set of predicted action proposals;
[0148] Based on the attention mechanism, according to the set of pseudo-label proposals, dynamically optimize the first set of predicted action proposals and the set of uncertainty masks to obtain the second set of predicted action proposals and the updated set of uncertainty masks;
[0149] Calculate the loss function according to the second set of predicted action proposals and the set of pseudo-label proposals to obtain the first action loss;
[0150] According to the first action loss, optimize the parameters of the fully supervised branch to obtain the first optimized fully supervised branch;
[0151] Based on the updated set of uncertainty masks, occlude the pseudo-label proposal feature map to obtain the second unoccluded proposal feature map;
[0152] Based on a preset second training round, according to the second unoccluded proposal feature map, perform second-stage training through the first optimized fully supervised branch to obtain a third set of predicted action proposals;
[0153] Calculate the loss function based on the third set of predicted action proposals and the set of pseudo-label proposals to obtain the second action loss;
[0154] Optimize the parameters of the first optimized fully supervised branch according to the second action loss to obtain a second optimized fully supervised branch.
[0155] Optionally, the fully supervised branch optimization module 260 is further used for:
[0156] Based on the attention mechanism, calculate according to the set of pseudo-label proposals and the first set of predicted action proposals to obtain a first action proposal attention sequence;
[0157] Based on a preset attention threshold, screen the first action proposal attention sequence to obtain a second action proposal attention sequence;
[0158] Based on the second action proposal attention sequence, select the corresponding action set in the first set of predicted action proposals to obtain a second set of predicted action proposals;
[0159] Based on a preset linear function, dynamically linearly decay the mask weights of the uncertainty mask set according to the second action proposal attention sequence to obtain an updated uncertainty mask set; the range of the mask weights is from 1 to 0.
[0160] Optionally, the video action localization module 270 is further used for:
[0161] Perform temporal action localization prediction based on the feature extractor, the weakly supervised branch, and the second optimized fully supervised branch to obtain a set of predicted action proposals;
[0162] Based on the attention mechanism, filter redundant action proposals through the soft non-maximum suppression algorithm according to the set of predicted action proposals to obtain filtered action proposals;
[0163] Perform action localization according to the action category and time boundary of the filtered action proposals.
[0164] The present invention proposes a weakly supervised temporal action localization method based on pseudo-label generation, which bridges the gap between weakly supervised and fully supervised temporal action localization in terms of framework and performance through a dual-branch framework; meanwhile, by utilizing the information in the weakly supervised branch and the fully supervised branch, the dual-branch framework can achieve higher-precision temporal action localization under the condition of less labeled data; adopting a pseudo-label proposal fusion strategy, it effectively utilizes the prior information in the pseudo-labels to generate pseudo-label proposals with high confidence and accurate boundaries; this strategy significantly improves the quality of the proposals and enhances the accuracy of the model through the fusion of multi-scale information; an uncertainty mask is introduced, and a new optimization mechanism is proposed for effective learning in the presence of noisy labels. This mechanism effectively reduces the interference of noisy labels during the training process by dynamically adjusting the quality of the pseudo-labels, enhances the robustness of the regression model, and ensures the stable convergence of the model during the training process. The present invention is a weakly supervised temporal action localization method that makes full use of prior information for high-quality pseudo-label generation and effectively combines fully supervised methods.
[0165] Figure 3 FIG. is a schematic structural diagram of a weakly supervised temporal action localization device provided by an embodiment of the present invention, as Figure 3 shown, the weakly supervised temporal action localization device may include the above Figure 2 shown weakly supervised temporal action localization device based on pseudo-label generation. Optionally, the weakly supervised temporal action localization device 310 may include a first processor 2001.
[0166] Optionally, the weakly supervised temporal action localization device 310 may further include a memory 2002 and a transceiver 2003.
[0167] Wherein, the first processor 2001, the memory 2002, and the transceiver 2003 may be connected through a communication bus.
[0168] Next, in conjunction with Figure 3 each component of the weakly supervised temporal action localization device 310 will be specifically introduced:
[0169] Among them, the first processor 2001 is the control center of the weakly-supervised temporal action localization device 310, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0170] Optionally, the first processor 2001 can execute various functions of the weakly-supervised temporal action localization device 310 by running or executing software programs stored in the memory 2002 and invoking the data stored in the memory 2002.
[0171] In a specific implementation, as an embodiment, the first processor 2001 can include one or more CPUs, such as Figure 3 CPU0 and CPU1 shown in
[0172] In a specific implementation, as an embodiment, the weakly-supervised temporal action localization device 310 can also include multiple processors, such as Figure 3 the first processor 2001 and the second processor 2004 shown in. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0173] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0174] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 3 not shown) of the weakly supervised temporal action localization device 310. The embodiments of the present invention do not make specific limitations thereto.
[0175] The transceiver 2003 is used to communicate with a network device or with a terminal device.
[0176] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0177] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 3 not shown) of the weakly supervised temporal action localization device 310. The embodiments of the present invention do not make specific limitations thereto.
[0178] It should be noted that Figure 3 the structure of the weakly supervised temporal action localization device 310 shown does not constitute a limitation to the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0179] In addition, the technical effects of the weakly supervised temporal action localization device 310 may refer to the technical effects of the weakly supervised temporal action localization method based on pseudo-label generation described in the above method embodiments, and will not be elaborated here.
[0180] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0181] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0182] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0183] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0184] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0185] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0186] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0187] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0188] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0189] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0190] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0191] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0192] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A weakly supervised temporal action localization method based on pseudo-label generation, characterized in that: The method comprises: Obtaining an uncut video containing an action; dividing the uncut video into segments to obtain a set of video segments; According to the video segment set, feature extraction is performed by a feature extractor to obtain a segment-by-segment feature set; According to the segment-by-segment feature set, preliminary action classification is performed through a weak supervision branch to obtain a classified attention sequence and a multi-scale action proposal set; According to the classified attention sequence and the action proposal set, a proposal fusion strategy is used to perform proposal fusion optimization to obtain a pseudo-label proposal set; Among them, in order to solve the technical problem that the overlap between the action proposals in the action proposal set makes it impossible for the action proposals to accurately describe the boundaries, thus affecting the training of the fully supervised branch; Using Ricker wavelet, for each action proposal, it is mapped to the Ricker distribution space as follows (1): (1) in, Indicates a point in time Ricker distribution value; for the Action Proposal , Indicates action proposal The start time, Indicates action proposal The end time of Indicates action proposal Length, Indicates action proposal The midpoint of Then all proposals are merged into a shared global space. The process is as follows: (2) in, Indicates the category The final fused wavelet distribution of the proposals is used to integrate the global representation of all proposals. represents the number of all proposals, Indicates The categories of proposals, Represents a specified action category; all action proposals are associated with corresponding confidence scores By fusing them together, misclassified predictions are filtered out, and a set of high-quality pseudo-label proposals with reliable boundaries is obtained. Based on a preset expansion ratio and a preset contraction ratio, generating an uncertainty mask set for a time boundary of each proposal in the action proposal set; Based on the attention mechanism, according to the pseudo-label proposal set and the uncertainty mask set, the full supervision branch is optimized and trained through a dynamic optimization mechanism to obtain a second optimized full supervision branch; Acquire a video of the action to be located; perform action location according to the video of the action to be located based on the feature extractor, the weak supervision branch and the second optimized full supervision branch; The method of performing action positioning according to the action video to be positioned based on the feature extractor, the weak supervision branch and the second optimized full supervision branch includes: Perform temporal action location prediction based on the feature extractor, the weak supervision branch, and the second optimized full supervision branch to obtain a set of predicted action proposals; Based on the attention mechanism, according to the predicted action proposal set, redundant action proposals are filtered through a soft non-maximum suppression algorithm to obtain filtered action proposals; Action localization is performed according to the action category and time boundary of the filtered action proposals.
2. The weakly supervised temporal action localization method based on pseudo-label generation according to claim 1, characterized in that: According to the segment-by-segment feature set, preliminary action classification is performed through a weak supervision branch to obtain a classification attention sequence and a multi-scale action proposal set, including: Based on the attention mechanism, adaptive action classification is performed through the basic model according to the segment-by-segment feature set to obtain an attention score sequence and a classification score matrix; Performing weighted fusion according to the attention score sequence and the classification score matrix to obtain a classification attention sequence; Based on a preset action threshold, the classified attention sequence is thresholded to obtain a multi-scale action proposal set; the action proposal set includes a time boundary set and an action category set.
3. The weakly supervised temporal action localization method based on pseudo-label generation according to claim 2 is characterized in that: The basic model is a baseline model based on a multiple instance learning training method.
4. The weakly supervised temporal action localization method based on pseudo-label generation according to claim 1, characterized in that: According to the classified attention sequence and the action proposal set, a proposal fusion strategy is used to perform proposal fusion optimization to obtain a pseudo-label proposal set, including: Based on the action proposal set, calculate according to the classified attention sequence to obtain the mean internal attention score of the proposal and the mean external attention score of the proposal; Based on the mean of the proposal internal attention score and the mean of the proposal external attention score, confidence calculation is performed according to the classified attention sequence to obtain a proposal confidence score set; According to the action proposal set and the proposal confidence score set, a proposal fusion strategy is used to perform fusion optimization to obtain a pseudo-label proposal set; the proposal fusion strategy refers to an optimization strategy that fuses the action proposals mapped to the global shared space with the confidence scores corresponding to the action proposals; the proposal fusion strategy is used to filter out the classification misdetection results of action proposals.
5. The weakly supervised temporal action localization method based on pseudo-label generation according to claim 1, characterized in that: The method of optimizing and training the full supervision branch based on the attention mechanism and the pseudo-label proposal set and the uncertainty mask set through a dynamic optimization mechanism to obtain a second optimized full supervision branch includes: Based on the morphological preprocessing method, the pseudo-label proposal set is filtered and optimized to obtain an optimized action proposal set; Inputting the optimized action proposal set into the encoder for data processing to obtain an encoded pseudo-label proposal set; According to the encoded pseudo-label proposal set, multi-scale feature extraction is performed through a feature pyramid network to obtain a pseudo-label proposal feature map; Based on the uncertainty mask set, masking the pseudo-label proposal feature map to obtain a first unmasked proposal feature map; Based on the preset first training round, according to the first unoccluded proposal feature map, a first stage of training is performed through a fully supervised branch to obtain a first predicted action proposal set; Based on the attention mechanism, dynamically optimize the first predicted action proposal set and the uncertainty mask set according to the pseudo-label proposal set to obtain a second predicted action proposal set and update the uncertainty mask set; Calculate the loss function according to the second predicted action proposal set and the pseudo label proposal set to obtain a first action loss; According to the first action loss, optimizing the parameters of the fully supervised branch to obtain a first optimized fully supervised branch; Based on the updated uncertain mask set, the pseudo-label proposal feature map is masked to obtain a second unmasked proposal feature map; Based on the preset second training round, according to the second unoccluded proposal feature map, the second stage training is performed through the first optimized full supervision branch to obtain a third predicted action proposal set; Calculate the loss function according to the third predicted action proposal set and the pseudo label proposal set to obtain a second action loss; According to the second action loss, parameters of the first optimized fully supervised branch are optimized to obtain a second optimized fully supervised branch.
6. The weakly supervised temporal action localization method based on pseudo-label generation according to claim 5, characterized in that: The method of dynamically optimizing the first predicted action proposal set and the uncertainty mask set based on the attention mechanism and according to the pseudo-label proposal set to obtain a second predicted action proposal set and update the uncertainty mask set includes: Based on the attention mechanism, a first action proposal attention sequence is obtained by performing calculation according to the pseudo-label proposal set and the first predicted action proposal set; Based on a preset attention threshold, the first action proposal attention sequence is screened to obtain a second action proposal attention sequence; Based on the second action proposal attention sequence, select a corresponding action set in the first predicted action proposal set to obtain a second predicted action proposal set; Based on a preset linear function, according to the second action proposal attention sequence, the mask weights of the uncertainty mask set are dynamically linearly attenuated to obtain an updated uncertainty mask set; the range of the mask weights is 1 to 0.
7. A weakly supervised temporal action positioning device based on pseudo-label generation, wherein the weakly supervised temporal action positioning device based on pseudo-label generation is used to implement the weakly supervised temporal action positioning method based on pseudo-label generation as claimed in any one of claims 1 to 6, characterized in that: The device comprises: The video segment acquisition module is used to acquire an uncut video containing actions; divide the uncut video into segments to obtain a video segment set; A video feature extraction module, used to extract features through a feature extractor according to the video segment set to obtain a segment-by-segment feature set; A weakly supervised branch classification module, configured to perform preliminary action classification through a weakly supervised branch according to the segment-by-segment feature set, and obtain a classified attention sequence and a multi-scale action proposal set; A pseudo-label proposal generation module is used to perform proposal fusion optimization using a proposal fusion strategy according to the classification attention sequence and the action proposal set to obtain a pseudo-label proposal set; An uncertainty mask generation module, configured to generate an uncertainty mask set at a time boundary of each proposal in the action proposal set based on a preset expansion ratio and a preset contraction ratio; A fully supervised branch optimization module, used to optimize and train the fully supervised branch based on the attention mechanism, according to the pseudo-label proposal set and the uncertainty mask set, through a dynamic optimization mechanism, to obtain a second optimized fully supervised branch; The video action localization module is used to obtain a video of the action to be located; based on the feature extractor, the weak supervision branch and the second optimized full supervision branch, the action is located according to the video of the action to be located.
8. A weakly supervised sequential action positioning device, characterized in that: The weakly supervised sequential action positioning device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method for eliminating uncertainty of self-supervised three-dimensional reconstruction
CN113592913A
Weak supervision video anomaly detection method for text interaction context features
CN119478756A