Video motion detection pseudo label denoising method and system based on self-training

The video features are extracted by pre-training visual encoder and combined with feature fusion and mapping modules to generate high-quality pseudo-labels. The self-finishing module is used to denoise the low-quality pseudo-labels, which solves the problem of unstable pseudo-label quality in the intelligent video monitoring of power system, and improves the accuracy and robustness of video action detection.

CN120356138AActive Publication Date: 2025-07-22GUIZHOU POWER GRID CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510847370.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The prior art uses noise-containing pseudo-label training in power system video intelligent safety monitoring, the results are not ideal, the quality of pseudo-labels is unstable, the feature processing is insufficient, and the dynamic optimization capability of the model is limited.

Method used

The timing and spatial features of the video are extracted through a pre-trained visual encoder, combined with feature fusion and mapping modules to optimize the space-time dimensions, use the detection model to generate high-quality pseudo-labels, and denoising the low-quality pseudo-labels through the self-finement module. The self-finement module is used to denoise and optimize the low-quality pseudo-labels.

Benefits of technology

It improves the accuracy and robustness of video action detection, solves the problems of unstable pseudo-label quality and insufficient feature processing, and improves the dynamic optimization capability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356138A_ABST
    Figure CN120356138A_ABST
Patent Text Reader

Abstract

The invention discloses a video motion detection pseudo label denoising method and system based on self-training, relates to the field of power system video intelligent safety supervision, and is used for improving the ability of a video motion detection model to utilize unlabeled video data. Comprising the following steps: using a semi-supervised framework based on a teacher-student model to reduce the noise of low-quality pseudo labels without adding extra parameter limitation, and improving the utilization efficiency of label-free video data. The pseudo labels are ranked from high to low according to the evaluation quality, and differential learning of the model on the pseudo labels with different qualities is guaranteed; the precision of the low-quality false label is improved through the self-fine adjustment module; on the basis of the improved pseudo tag, a training method based on time sequence prediction invariance is further provided to help the model learn more robust spatial-temporal characteristics. Compared with a general semi-supervised method, the active learning and pseudo label self-fine adjustment technology is introduced in the field of semi-supervised video motion detection for the first time, and a pseudo label with higher precision than that under default self-training setting can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video intelligent safety supervision in power systems, and specifically to a self-training-based video action detection pseudo-label denoising method and system applicable to power grid operation scenarios. Background Art

[0002] The training of temporal action detection models relies on a large amount of manually annotated data. The annotation of videos is more complicated and time-consuming than that of images. Moreover, in the context of power intelligent safety supervision, the annotation of video content requires some professional knowledge, which makes the annotation of video data in power operation scenarios difficult. Therefore, it is necessary to combine pseudo-label learning technology. However, learning effective representations of tasks from unlabeled videos is a difficult task. Existing methods are mostly based on self-learning and pseudo-labels. Due to problems such as class errors and temporal boundary deviations in pseudo-labels, the effect of learning unlabeled videos from pseudo-labels has always been unsatisfactory. Summary of the Invention

[0003] In view of the above problems, the present invention is proposed.

[0004] Therefore, the technical problem solved by the present invention is: how to solve the problem that the existing technology has an unsatisfactory effect when training with noisy pseudo-labels.

[0005] To solve the above technical problem, the present invention provides the following technical solution: A self-training-based video action detection pseudo-label denoising method, which includes the following steps. Extract the spatio-temporal semantic features of the video. Through a pre-trained visual encoder, extract visual features containing temporal information and spatial information to obtain a temporal feature vector and a spatial feature vector. Map the video feature vector to a subspace suitable for the temporal action detection task. Introduce the temporal feature vector and spatial feature vector containing explicit semantic information, and design a feature fusion and mapping module to obtain a spatio-temporal feature vector corresponding to the video length. According to the spatio-temporal feature vector corresponding to the video length, initialize two identical detection models, one as a supervised model for generating pseudo-labels, and the other as a training model for updating weights. The detection model includes a multi-layer feature encoding transformer, a multi-layer group normalization layer, an action category prediction branch, and a temporal prediction branch. Obtain a preliminary action prediction through the inference of the supervised model. Sort the preliminary action predictions from high to low according to the pseudo-label evaluation sorting method. Denoise the pseudo-labels of the top 50% after quality sorting using a self-refinement module. The self-refinement module reuses the inference module of the supervised model, performs several inferences and aggregates the inference results to obtain denoised pseudo-labels.

[0006] As a preferred solution of a pseudo-label denoising method for video action detection based on self-training according to the present invention, wherein: extracting the spatio-temporal semantic features of the video, through a pre-trained visual encoder, extracting visual features including temporal information and spatial information, obtaining a temporal feature vector and a spatial feature vector includes, Input action video sequence , where M represents the total number of frames, represents the th video frame of the action video sequence; Processing the action video sequence into a sequence of T video clips : ; Wherein, is the total number of video clips, each clip contains 16 frames, and the interval between clips is 4 frames; Extracting features based on the obtained clip sequence, extracting RGB frames and optical flow frames for the clips, and respectively obtaining d-dimensional vectors , , the expression of the video feature vector sequence is: , , , Wherein, is the RGB feature of the t-th clip, is the optical flow feature of the t-th clip, concat represents concatenation in the channel dimension, is the spatio-temporal feature of the t-th clip, is the video semantic vector space, is the temporal length of the feature, is the number of channels of the feature.

[0007] As a preferred solution of a pseudo-label denoising method for video action detection based on self-training according to the present invention, wherein: mapping the video feature vector to a subspace suitable for the temporal action detection task, introducing a temporal feature vector and a spatial feature vector containing explicit semantic information, designing a feature fusion and mapping module, and obtaining a spatio-temporal feature vector corresponding to the video length includes, Based on the transpose of the obtained video feature vector sequence F, through a feature fusion and mapping module composed of stacked one-dimensional convolutional neural networks, aggregating local context information in the temporal dimension, and obtaining a visual feature vector , wherein, is the D-dimensional feature vector of the t-th clip.

[0008] As a preferred solution of a method for denoising pseudo-labels in video action detection based on self-training according to the present invention, wherein: mapping the video feature vector to a subspace suitable for the temporal action detection task, introducing temporal feature vectors and spatial feature vectors containing explicit semantic information, and designing a feature fusion and mapping module to obtain spatio-temporal feature vectors corresponding to the video length includes, Based on the obtained transposed video feature vector sequence F, through a feature fusion and mapping module composed of stacked one-dimensional convolutional neural networks, aggregating local context information in the temporal dimension to obtain visual feature vectors , wherein, is the D-dimensional feature vector of the t-th segment.

[0009] As a preferred solution of a method for denoising pseudo-labels in video action detection based on self-training according to the present invention, wherein: initializing two identical detection models according to the spatio-temporal feature vectors corresponding to the video length, one as a supervised model for generating pseudo-labels, and the other as a training model for updating weights includes, The detection model includes a multi-layer feature encoding transformer, and each layer adds a transformer module with one-dimensional convolutional modules of different sizes. The input of the first layer is the visual feature vector , and the output of the first layer is the visually feature vector of further processed temporal downsampling by a factor of two , the current visual feature vector is used as the input of the second layer, and propagates forward layer by layer, and the outputs of each layer are combined as , where L is the number of layers of the multi-layer feature encoding transformer, is the output of the l-th layer , is the output of the L-th layer. If T cannot be divided by , extend the time dimension of to by padding with 0, is the smallest integer greater than or equal to T and divisible by ; The detection model includes a multi-layer group normalization layer, which includes L layers of group normalization layers, and the outputs of each layer are respectively passed through the group normalization layers to obtain a deep representation of the video features , wherein, is the output of the l-th layer of group normalization layer, , is the output of the L-th layer of group normalization layer; The detection model includes an action category prediction branch, which includes two layers of one-dimensional convolutional neural networks and one layer of fully connected network. Each layer is connected by a normalization layer and a ReLU activation layer. For the output of each layer of group normalization layer , the output result through the action category prediction branch is , where C is the number of action categories to be predicted; The detection model includes a temporal prediction branch, which contains two layers of one-dimensional convolutional neural networks and one layer of fully connected network. Between each layer, there are a normalization layer and a ReLU activation layer. For the output of each layer of group normalization layer , the output result of the temporal boundary prediction branch is . In the two dimensions, the first dimension represents the start time point of the action predicted for the current segment, and the second dimension represents the end time point of the action predicted for the current segment; A single video obtains several prediction results, and the representation of the j-th action prediction is , where is the action category, is the start time of the action in the video, is the end time of the action in the video; The output results of the action category prediction branches of each layer are concatenated to obtain , where is the action category output by the first layer, is the action category output by the second layer, is the action category output by the L-th layer, to clarify the action category ; To obtain and , the output results of the temporal prediction branches of each layer are concatenated to obtain , where is the temporal boundary output by the first layer, is the temporal boundary output by the second layer, is the temporal boundary output by the L-th layer. On the temporal dimension, N action predictions with the highest confidence are selected according to A, and N corresponding temporal intervals are selected. The initial action predictions are obtained by filtering using a confidence threshold, to clarify the start time of the action in the video and the end time of the action in the video. All action predictions included in a video are regarded as a pseudo-label.

[0010] As a preferred solution of the method for denoising pseudo-labels for video action detection based on self-training according to the present invention, where: The initial action predictions are sorted from high to low according to the pseudo-label evaluation and sorting method, including Several pseudo-labels obtained from several videos are used to calculate the category uncertainty of each pseudo-label and the sample difficulty of each pseudo-label; A number of pseudo-labels obtained from a number of videos, as well as corresponding A and B. For an action prediction j, obtain all time series intervals with the same category and overlapping time series intervals from A and B , calculate the temporal boundary uncertainty of each pseudo-label, where represents the start time of the action in the pseudo-label, represents the end time of the action in the pseudo-label, and K represents the number of times of adding random offsets to an action interval; Complete the comprehensive evaluation according to the category uncertainty of each pseudo-label, the sample difficulty of each pseudo-label, and the temporal boundary uncertainty of each pseudo-label.

[0011] As a preferred solution of a pseudo-label denoising method for video action detection based on self-training according to the present invention, wherein: the 50% of the pseudo-labels after quality sorting are denoised using a self-refinement module, and the self-refinement module reuses the inference module of the supervised model. Performing several inferences and aggregating the inference results to obtain the denoised pseudo-labels includes, According to the pseudo-labels, add random offsets to each predicted time series interval in the 50% of the pseudo-labels after confidence sorting, and add offsets on the basis of the prediction to obtain a new offset interval , where is the start time of the new offset interval, is the end time of the new offset interval; Repeat the operation K times to obtain a set of adjacent intervals with random offsets , where is the start time of the i-th adjacent interval of the j-th pseudo-label, is the end time of the i-th adjacent interval of the j-th pseudo-label. Expand the interval to twice the interval length, generate a set of expanded intervals, extract the features at the corresponding time series positions, add noise, and re-infer through the time series prediction branch of the supervised model to obtain K new time series predictions. Average the K new time series predictions to obtain the final denoised time series interval, and update each prediction of the pseudo-label.

[0012] As a preferred solution of a pseudo-label denoising method for video action detection based on self-training according to the present invention, wherein: the 50% of the pseudo-labels after quality sorting are denoised using a self-refinement module, and the self-refinement module reuses the inference module of the supervised model. Performing several inferences and aggregating the inference results to obtain the denoised pseudo-labels further includes, Use a temporal scaling enhancement strategy based on Gaussian sampling for all extracted temporal features, use a classification loss function to train the action category prediction branch of the training model, use a temporal interval loss function to train the temporal prediction branch of the training model, and according to the pseudo-label evaluation ranking, first use 50% of the pseudo-labels with a comprehensive evaluation score higher than the threshold to train the training model for E rounds; The obtained updated low-scoring pseudo-labels are added to the training process, and the training model is continuously trained for E epochs. The weights of the training model are updated to the supervised model by exponential moving average.

[0013] Another object of the present invention is to provide a self-training based video action detection pseudo-label denoising system, which can extract visual features containing temporal information and spatial information in the video by introducing a pre-trained visual encoder module, optimize the features in the spatio-temporal dimension by combining a feature fusion and mapping module, generate action prediction results by using an inference module, and then generate high-quality pseudo-labels through a detection model module. At the same time, a self-fine-tuning module is used to denoise and optimize low-quality pseudo-labels, solving the problems of unstable pseudo-label quality, insufficient feature processing, and limited model dynamic optimization ability in the prior art, thereby effectively improving the accuracy and robustness of video action detection.

[0014] To solve the above technical problems, the present invention provides the following technical solutions: A self-training based video action detection pseudo-label denoising system, comprising: a pre-trained visual encoder module, a feature fusion and mapping module, an inference module, a detection model module, and a self-fine-tuning module; The pre-trained visual encoder module inputs an action video sequence, divides the video into multiple segment sequences, each segment containing temporal information and spatial information, extracts RGB frames and optical flow frames for each segment respectively, and encodes them through a pre-trained 3D convolutional neural network, embeds the temporal and spatial information of the video into a feature vector, generates a temporal feature vector and a spatial feature vector, provides a feature expression containing rich semantics, and outputs a visual feature vector sequence;

[0015] The feature fusion and mapping module maps the visual feature sequence of the video to a subspace suitable for temporal action detection, performs feature fusion for different characteristics of temporal and spatial information, designs a one-dimensional convolutional neural network as a feature fusion component, processes the local context information of the temporal dimension layer by layer, aggregates the results of multiple layers of convolution, and outputs a spatio-temporal feature vector consistent with the video length; The inference module generates preliminary action predictions by inputting spatio-temporal features, including action categories, time intervals, and confidence levels, and uses a method of layer-by-layer calculation for prediction inference, combining the outputs of the action category prediction branch and the temporal prediction branch; The detection model module works collaboratively through two identical detection models, which are respectively responsible for pseudo-label generation and model weight update, generates pseudo-labels by combining all features, including action categories and their corresponding time intervals, and sorts and evaluates them; The self - fine - tuning module is aimed at pseudo - labels with quality rankings lower than the threshold. It uses the inference ability of the supervised model to perform denoising processing, adds random offsets to the low - quality pseudo - labels to generate multiple perturbed temporal intervals, obtains different prediction results through inference, aggregates the prediction results to generate a denoised temporal interval, and uses the denoised pseudo - labels to further optimize the training model. Through the temporal scaling enhancement strategy based on Gaussian sampling and the classification loss and temporal interval loss functions, the dynamic improvement of the model performance is completed.

[0016] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above - mentioned method for denoising pseudo - labels in video action detection based on self - training are implemented.

[0017] A computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above - mentioned method for denoising pseudo - labels in video action detection based on self - training are implemented.

[0018] The beneficial effects of the present invention: The present invention extracts the temporal features and spatial features of the video through a pre - trained visual encoder module, optimizes the features in the spatio - temporal dimension by combining the feature fusion and mapping module, generates action prediction results using the inference module, and generates high - quality pseudo - labels through the detection model module. At the same time, the self - fine - tuning module is used to denoise and optimize the low - quality pseudo - labels, solving the problems of unstable pseudo - label quality, insufficient feature processing, and low denoising efficiency in the prior art, thereby effectively improving the accuracy and robustness of video action detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is the overall flowchart of a method for denoising pseudo - labels in video action detection based on self - training provided by the first embodiment of the present invention.

[0021] Figure 2 It is the schematic flowchart of the pseudo - label evaluation and ranking method in the method for denoising pseudo - labels in video action detection based on self - training provided by the second embodiment of the present invention.

[0022] Figure 3 It is the schematic flowchart of the pseudo - label self - fine - tuning method in the method for denoising pseudo - labels in video action detection based on self - training provided by the second embodiment of the present invention. Detailed implementation manners

[0023] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0024] Example 1, refer to Figure 1 For an embodiment of the present invention, a method for denoising pseudo-labels of video action detection based on self-training is provided, including: Step 1: Extract the spatio-temporal semantic features of the video. Through a pre-trained visual encoder, extract visual features containing temporal information and spatial information to obtain a temporal feature vector and a spatial feature vector.

[0025] Input the action video sequence , where M represents the total number of frames, represents the th video frame of the action video sequence. Further process the action sequence into a sequence of T video segments , , where represents the tth video segment, each segment contains 16 frames, and the interval between segments is 4 frames.

[0026] Extract features from the obtained segment sequence. Extract RGB (Red Green Blue original video frames) frames and optical flow frames for the segments, and respectively pass them through a 3D convolutional neural network to obtain d-dimensional vectors , . The video feature vector sequence is , where , , is the video semantic vector space.

[0027] Step 2: Map the video feature vectors to a subspace suitable for the temporal action detection task. Using the obtained temporal feature vectors and spatial feature vectors containing explicit semantic information, design a feature fusion and mapping module to obtain spatio-temporal feature vectors corresponding to the video length.

[0028] Transpose the obtained video feature vector sequence F to obtain . Through a feature fusion and mapping module composed of stacked one-dimensional convolutional neural networks, aggregate the local context information in the temporal dimension to obtain a visual feature vector , where is the D-dimensional feature vector of the tth segment.

[0029] Step 3. Initialize two identical detection models according to the obtained spatio-temporal feature vector corresponding to the video length. One is used as a supervised model to generate pseudo-labels, and the other is used as a training model to update the weights. The model includes a multi-layer feature encoding transformer, a multi-layer group normalization layer, an action category prediction branch, and a temporal prediction branch; obtain the preliminary action prediction through the inference of the supervised model.

[0030] First, the detection model includes a multi-layer feature encoding transformer, and each layer adds a transformer module with one-dimensional convolutional modules of different sizes. The input of the first layer is the visual feature vector , and the output of the first layer is the visually feature vector of the further processed temporal downsampling by a factor of two . This visually feature vector serves as the input of the second layer and propagates forward layer by layer. The outputs of all layers are merged into , where L is the number of layers of the multi-layer feature encoding transformer, is the abbreviation of the output of the l-th layer . If T cannot be divided by , extend the time dimension of to by padding with 0s. is the smallest integer greater than or equal to T and divisible by ; Subsequently, the detection model includes a multi-layer group normalization layer, which includes L layers of group normalization layers. The outputs of each layer in step 3.1) are respectively passed through the group normalization layers to obtain the deep representation of the video features , where ; Secondly, the detection model includes an action category prediction branch, which includes two layers of one-dimensional convolutional neural networks and one layer of fully connected network, with a normalization layer and a ReLU activation layer between each layer. For the output of each layer of the group normalization layer, the output result through the action category prediction branch is , where C is the number of action categories to be predicted; Thirdly, the detection model includes a temporal prediction branch, which includes two layers of one-dimensional convolutional neural networks and one layer of fully connected network, with a normalization layer and a ReLU activation layer between each layer. For the output of each layer of the group normalization layer, the output result through the action category prediction branch is . In the two dimensions, the first dimension represents the start time point of the action predicted by the current segment, and the second dimension represents the end time point of the action predicted by the current segment; Finally, a single video may obtain multiple prediction results, and the representation of the j-th action prediction is , the three elements are the action category, the start time in the video, and the end time in the video. To obtain , it is necessary to splice the output results of the action category prediction branches of each layer to obtain , splice the output results of the temporal prediction branches of each layer to obtain , select the top K action predictions with the highest confidence according to A in the temporal dimension, and select the corresponding top K temporal intervals, and use the confidence threshold to filter to obtain preliminary action predictions. All the action predictions included in a video are regarded as a pseudo-label.

[0031] Among them, Topk is the number of action predictions selected according to the confidence.

[0032] Step 4. Refer to Figure 2 According to the obtained preliminary action predictions, sort the preliminary action predictions from high to low according to the pseudo-label evaluation sorting method. The pseudo-label evaluation sorting method includes three evaluation criteria, namely category uncertainty, sample difficulty, and temporal boundary uncertainty.

[0033] First, according to the N pseudo-labels obtained from N videos, calculate the category uncertainty of each pseudo-label: , where J is the number of predictions in the pseudo-label, C is the number of actions, is the confidence score of the i-th action in the j-th prediction; Subsequently, according to the N pseudo-labels obtained from N videos, calculate the sample difficulty of each pseudo-label: , where is the confidence threshold of the prediction, is the confidence score vector of the j-th prediction, ; Then, according to the N pseudo-labels obtained from N videos and the corresponding A and B, for an action prediction j, obtain all the temporal intervals with the same category and overlapping temporal intervals from A and B , where is the start time of the temporal interval, is the end time of the temporal interval, and calculate the temporal boundary uncertainty of each pseudo-label: , where is the start time point of the j-th action prediction, is the end time point of the j-th action prediction, represents calculating the standard deviation of a vector; Finally, based on the calculated , calculate the comprehensive evaluation score: , where is the category uncertainty, is the difficulty of the pseudo-label sample, represents the boundary uncertainty, , , are respectively weights, are respectively the maximum values of among all pseudo-labels; let the weights satisfy , M is the comprehensive score of each pseudo-label, sort all pseudo-labels in descending order according to the calculated M, and consider the first 50% as high-quality pseudo-labels and the last 50% as low-quality pseudo-labels.

[0034] Step 5. Refer to Figure 3 According to the obtained pseudo-label sorting, use the self-finetuning module to denoise the pseudo-labels in the last 50% of the quality sorting. The self-finetuning module reuses the inference module of the supervised model, performs several inferences and aggregates the inference results to obtain the denoised pseudo-labels.

[0035] First, according to the obtained pseudo-labels, add random offsets to each predicted time series interval in the pseudo-labels in the last 50% of the confidence sorting: , where is the random offset intensity of the start time of the interval, is the random offset intensity of the end time of the interval. Add the offset on the basis of the prediction to obtain a new offset interval , and repeat the operation of adding the offset K times to obtain a set of randomly offset adjacent intervals ; Subsequently, according to the obtained set of randomly offset adjacent intervals , represents the i-th adjacent interval generated by the j-th pseudo-label, and expand the interval to twice the interval length, so as to generate a set of expanded intervals: , where is the start time of the expanded interval, is the end time of the expanded interval. Extract the features corresponding to the time series positions of this set of expanded intervals: , where represents the feature corresponding to the first expanded interval, represents the feature corresponding to the second expanded interval, Representing the features corresponding to the Kth expanded interval, respectively obtained by in the corresponding time series interval , , by intercepting.

[0036] Finally, based on the features of the K expanded intervals obtained for an action prediction, these features are added with noise and then re-inferred through the time series prediction branch of the supervised model to obtain K new time series predictions. The average of these K new time series predictions is taken to obtain the final denoised time series interval: , where is the start time of the denoised time series interval, is the end time of the denoised time series interval, and model is the time series prediction branch of the supervised model.

[0037] After obtaining the final time series interval, update each prediction of the pseudo-label: , where are the category, action start time, and action end time of the pseudo-label, respectively.

[0038] Step Six: In the final model training stage, based on the obtained updated pseudo-labels, use them to supervise the training of the model and update the weights.

[0039] First, use the time series scaling enhancement strategy based on Gaussian sampling for all the extracted time series features, and use the classification loss function to train the action category prediction branch of the training model, where is related to whether it is a positive or negative sample, and γ is the modulation factor for focusing on difficult samples. The larger γ is, the smaller the influence of simple samples; use the time series interval loss function to train the time series prediction branch of the training model, where ρ is the Euclidean distance between the prediction interval of the training model and the time series interval of the corresponding pseudo-label, and c is the union length of the prediction interval of the training model and the time series interval of the corresponding pseudo-label; Subsequently, according to the evaluation and ranking of the obtained pseudo-labels, first use 50% of the pseudo-labels with a comprehensive evaluation score higher than the threshold to train the training model for E rounds; Then, add the obtained updated low-score pseudo-labels to the training process and continue to train the training model for E rounds; Finally, the weights of the training model are updated to the supervised model by means of exponential smoothing average: , where is the weight of the training model, is the weight of the supervised model, is the weight updated for the training model, is the weight update coefficient.

[0040] The present invention is mainly used to make full use of the labeled videos and unlabeled videos to improve the accuracy of the action detection model in the case of insufficient professional annotators.

[0041] Embodiment 2, which is an embodiment of the present invention, provides a system for a self-training-based video action detection pseudo-label denoising method, including: a pre-trained visual encoder module, a feature fusion and mapping module, an inference module, a detection model module, and a self-fine-tuning module; The pre-trained visual encoder module inputs an action video sequence, divides the video into multiple segment sequences, and each segment contains temporal information and spatial information. RGB frames and optical flow frames are respectively extracted from the segments and encoded by a pre-trained 3D convolutional neural network. The temporal and spatial information of the video is embedded into the feature vectors to generate temporal feature vectors and spatial feature vectors, providing a feature expression rich in semantics, and outputting a visual feature vector sequence; The feature fusion and mapping module maps the visual feature sequence of the video to a subspace suitable for temporal action detection, performs feature fusion for the different characteristics of temporal and spatial information, designs a one-dimensional convolutional neural network as the feature fusion component, processes the local context information of the temporal dimension layer by layer, aggregates the results of multiple layers of convolution, and outputs a spatio-temporal feature vector consistent with the video length; The inference module generates preliminary action predictions by inputting spatio-temporal features, including action categories, time intervals, and confidence levels, and uses a method of layer-by-layer calculation for prediction inference, combining the outputs of the action category prediction branch and the temporal prediction branch; The detection model module works collaboratively through two identical detection models, which are respectively responsible for pseudo-label generation and model weight update, generates pseudo-labels by combining all features, including action categories and their corresponding time intervals, and sorts and evaluates them; The self-fine-tuning module performs denoising processing on the pseudo-labels with quality rankings lower than the threshold by using the inference ability of the supervised model, adds random offsets to the low-quality pseudo-labels to generate multiple perturbed temporal intervals, obtains different prediction results through inference, aggregates the prediction results to generate denoised temporal intervals, and further optimizes the training model using the denoised pseudo-labels. Through a temporal scaling enhancement strategy based on Gaussian sampling and classification loss and temporal interval loss functions, the dynamic improvement of the model performance is completed.

[0042] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0043] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.

[0044] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), fiber optic devices, and portable compact disc read-only memories (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0045] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0046] Example 3. In this example, in order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments. In this example, experiments are respectively carried out on the existing traditional method and the method of this example.

[0047] This example proposes a self-training model. Compared with the existing methods, PLD can generate higher-quality pseudo-labels.

[0048] The present invention uses two commonly used video action detection datasets: THUMOS14 and ActivityNet v1.3. In addition, the present invention also uses 4019 video data collected in a real power operation scenario, including 3319 videos in the training set, 700 videos in the validation set, and 697 videos in the test set, with a total of five action categories.

[0049] The video preprocessing steps are as follows: First, RGB and optical flow data in the form of frames are extracted from the video. Subsequently, the I3D model pre-trained on the ImageNet image dataset is used to extract features from the RGB frames and optical flow frames. A segment consisting of every 16 frames is used as the unit for feature extraction, and the interval between segments is 4 frames. For example, the first three segments respectively take the information of frames [0~15], [4~19], and [8~23]. The output feature channel number of I3D is 1024. Finally, the RGB features and optical flow features are concatenated to obtain a 2048-dimensional feature tensor.

[0050] The configuration of the proposed model is as follows: 1. The visual encoder uses a two-stream I3D model; 2. The feature fusion and mapping module of the detection model uses a two-layer one-dimensional convolutional block, and the feature encoder uses a stack of seven-layer feed-forward transformers for multi-scale feature extraction. The action category prediction branch and the temporal prediction branch each contain two layers of one-dimensional convolutional neural networks and one layer of fully connected networks; 4. The Adam optimizer with an initial learning rate of 1 × 10−4 and a weight decay rate of 5 × 10−2 is used to optimize the PLD model during the training and inference processes.

[0051] The present invention uses the mean average precision (mAP) to evaluate the performance of the PLD model. The calculation method of mAP is as follows: First, for each prediction result, calculate its intersection over union (IoU) with all ground truths (GTs). Subsequently, set multiple IoU thresholds according to the task requirements. The IoU thresholds are [0.3:0.7:0.1] on THUMOS14 and [0.5:0.95:0.05] on ActivityNet v1.3. Subsequently, for each prediction result, if its IoU with a certain GT is greater than or equal to the set threshold and the GT has not been matched by a previous prediction result, it is determined as a true positive (TP). If the IoU of the prediction result with the GT is less than the threshold, or although the IoU is greater than the threshold but the GT has been matched by a previous prediction result, it is determined as a false positive (FP). According to the numbers of TP, FP, and false negatives (FN), calculate the precision at the current IoU threshold. Average the class precisions at all IoU thresholds to obtain the average precision (AP) for each class. Average the APs of all classes to obtain the mean average precision mAP. mAP ranges from 0 to 1, and the larger it is, the higher the detection precision.

[0052] In the three datasets, compared with the state-of-the-art methods, the PLD model proposed by the present invention significantly improves the mAP, indicating that the PLD model of the present invention can generate higher-quality speech.

[0053] Table 1: Mean average precision of the model , Referring to Table 1, the results show that: by using the three datasets of THUMOS14, ActivityNet v1.3, and real power scenarios for verification, the mAP results of the PLD model proposed by the present invention are better than those of other models, proving that the PLD model can effectively generate high-quality pseudo-labels.

[0054] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for denoising pseudo-labels of video action detection based on self-training, characterized in that Including: Extract the spatio-temporal semantic features of the video. Through a pre-trained visual encoder, extract visual features containing temporal information and spatial information to obtain a temporal feature vector and a spatial feature vector; Map the video feature vector to a subspace suitable for the temporal action detection task. Introduce the temporal feature vector and spatial feature vector containing explicit semantic information, design a feature fusion and mapping module, and obtain a spatio-temporal feature vector corresponding to the video length; According to the spatio-temporal feature vector corresponding to the video length, initialize two identical detection models, one as a supervision model for generating pseudo-labels and the other as a training model for updating weights; The detection model includes a multi-layer feature encoding transformer, a multi-layer group normalization layer, an action category prediction branch, and a temporal prediction branch. Obtain a preliminary action prediction through the inference of the supervision model; Sort the preliminary action predictions from high to low according to the pseudo-label evaluation and sorting method; Denoise the 50% of the pseudo-labels after quality sorting using a self-refinement module. The self-refinement module reuses the inference module of the supervision model, performs several inferences and aggregates the inference results to obtain denoised pseudo-labels.

2. The pseudo-label denoising method for video action detection based on self-training according to claim 1, characterized in that: The extracting the spatio-temporal semantic features of the video. Through a pre-trained visual encoder, extract visual features containing temporal information and spatial information to obtain a temporal feature vector and a spatial feature vector includes, Input action video sequence , where M represents the total number of frames, represents the th video frame of the action video sequence; Process an action video sequence into a sequence of T video segments : ; Among them, is the total number of video segments, each segment contains 16 frames, and the interval between segments is 4 frames; Extract features based on the obtained segment sequences, extract RGB frames and optical flow frames for the segments, and respectively obtain d-dimensional vectors through 3D convolutional neural networks , , and the expression of the video feature vector sequence is: , , , Among them, is the RGB feature of the t-th segment, is the optical flow feature of the t-th segment, and concat represents concatenation in the channel dimension. is the spatio-temporal feature of the t-th segment, is the video semantic vector space, is the temporal length of the feature, is the number of channels of the feature.

3. A pseudo-label denoising method for video action detection based on self-training according to claim 2, characterized in that: The mapping the video feature vector to a subspace suitable for the temporal action detection task. Introduce the temporal feature vector and spatial feature vector containing explicit semantic information, design a feature fusion and mapping module, and obtain a spatio-temporal feature vector corresponding to the video length includes, Based on the obtained video feature vector sequence \(F^T\), through the feature fusion and mapping module composed of stacked one-dimensional convolutional neural networks, aggregate the local context information in the temporal dimension to obtain the visual feature vector , where is the \(D\)-dimensional feature vector of the \(t\)-th segment.

4. A self-training based video action detection pseudo-label denoising method according to claim 3, characterized in that: The according to the spatio-temporal feature vector corresponding to the video length, initialize two identical detection models, one as a supervision model for generating pseudo-labels and the other as a training model for updating weights includes, The detection model includes a multi-layer feature encoding transformer. Each layer adds a transformer module with one-dimensional convolutional modules of different sizes. The input of the first layer is a visual feature vector , and the output of the first layer is a temporally downsampled visual feature vector by a factor of two for further processing . The current visual feature vector serves as the input to the second layer and propagates forward layer by layer. The outputs of all layers are merged as , where L is the number of layers of the multi-layer feature encoding transformer, is the output of the l-th layer , is the output of the L-th layer. If T is not divisible by , the time dimension of is extended to by padding with zeros. is the smallest integer greater than or equal to T and divisible by ; The detection model includes a multi-layer group normalization layer, which contains L group normalization layers. The outputs of each layer are respectively passed through the group normalization layers to obtain a deep representation of the video features. , where is the output of the l-th group normalization layer, , is the output of the L-th group normalization layer; The detection model contains an action category prediction branch, which includes two layers of one-dimensional convolutional neural networks and one layer of fully connected network. Between each layer, there are a normalization layer and a ReLU activation layer. For the output of each layer's group normalization layer , the output result of the action category prediction branch is , where C is the number of action categories to be predicted; The detection model contains a temporal prediction branch, which includes two layers of one-dimensional convolutional neural networks and one layer of fully connected network. Between each layer, there are a normalization layer and a ReLU activation layer. For the output of each group normalization layer , the output result through the temporal boundary prediction branch is , in two dimensions, the first dimension represents the start time point of the action predicted for the current segment, and the second dimension represents the end time point of the action predicted for the current segment; A single video yields several prediction results, where the representation of the j-th action prediction is , where is the action category, is the start time of the action in the video, is the end time of the action in the video; The output results of the action category prediction branches of each layer are concatenated to obtain , where is the action category output by the first layer, is the action category output by the second layer, is the action category output by the L-th layer, and the action category is determined ; To obtain and , the output results of the timing prediction branches of each layer are concatenated to obtain , where is the timing boundary output by the first layer, is the timing boundary output by the second layer, is the timing boundary output by the L-th layer. N action predictions with the highest confidence are selected according to A in the timing dimension, and N corresponding timing intervals are selected. The initial action predictions are obtained by filtering using the confidence threshold, and the start time of the action in the video and the end time of the action in the video are determined. All action predictions included in a video are regarded as a pseudo-label.

5. A self-training based video action detection pseudo-label denoising method as claimed in claim 4, wherein: The sorting the preliminary action predictions from high to low according to the pseudo-label evaluation and sorting method includes, For the several pseudo-labels obtained from several videos, calculate the class uncertainty of each pseudo-label and the sample difficulty of each pseudo-label; A number of pseudo-labels obtained from a number of videos, and the corresponding A and B. For an action prediction j, all time intervals with the same category and overlapping time series intervals are obtained from A and B , calculate the temporal boundary uncertainty of each pseudo-label, where represents the start time of the action in the pseudo-label, represents the end time of the action in the pseudo-label, and K represents the number of times of adding random offsets to an action interval; Complete a comprehensive evaluation according to the class uncertainty of each pseudo-label, the sample difficulty of each pseudo-label, and the temporal boundary uncertainty of each pseudo-label.

6. A pseudo-label denoising method for video action detection based on self-training according to claim 5, characterized in that: The denoising the 50% of the pseudo-labels after quality sorting using a self-refinement module. The self-refinement module reuses the inference module of the supervision model, performs several inferences and aggregates the inference results to obtain denoised pseudo-labels includes, According to the pseudo-labels, a random offset is added to each predicted time series interval in the 50% of the pseudo-labels after confidence ranking, and a new offset interval is obtained by adding the offset on the basis of the prediction. , where is the start time of the new offset interval, is the end time of the new offset interval; Repeat the operation K times to obtain a set of adjacent intervals with random offsets , where is the start time of the i-th adjacent interval of the j-th pseudo-label, is the end time of the i-th adjacent interval of the j-th pseudo-label. Expand the interval to twice its length, generate a set of expanded intervals, extract the features at the corresponding time series positions, add noise, and then re-infer through the time series prediction branch of the supervised model to obtain K new time series predictions. Average the K new time series predictions to obtain the final denoised time series interval, and update each prediction of the pseudo-label.

7. A method for denoising pseudo-labels of video action detection based on self-training, characterized in that: The denoising the 50% of the pseudo-labels after quality sorting using a self-refinement module. The self-refinement module reuses the inference module of the supervision model, performs several inferences and aggregates the inference results to obtain denoised pseudo-labels also includes, Apply a temporal scaling enhancement strategy based on Gaussian sampling to all the extracted temporal features. Use a classification loss function to train the action category prediction branch of the training model, use a temporal interval loss function to train the temporal prediction branch of the training model. According to the pseudo-label evaluation and sorting, first use 50% of the pseudo-labels with a comprehensive evaluation score higher than the threshold to train the training model for E rounds; The obtained updated low-score pseudo-labels are added to the training process, and the training model is continuously trained for E rounds. The weights of the training model are updated to the supervised model by means of exponential smoothing average.

8. A system for a pseudo-label denoising method of video action detection based on self-training, applying a pseudo-label denoising method of video action detection based on self-training as described in any one of claims 1 to 7, characterized in that It includes a pre-trained visual encoder module, a feature fusion and mapping module, an inference module, a detection model module, and a self-fine-tuning module; The pre-trained visual encoder module inputs an action video sequence, divides the video into multiple segment sequences, and each segment contains temporal information and spatial information. RGB frames and optical flow frames are respectively extracted from the segments and encoded by a pre-trained 3D convolutional neural network, and the temporal and spatial information of the video is embedded into the feature vectors to generate temporal feature vectors and spatial feature vectors, providing a feature expression rich in semantics, and outputting a sequence of visual feature vectors. The feature fusion and mapping module maps the visual feature sequence of the video to a subspace suitable for temporal action detection, performs feature fusion for different characteristics of temporal and spatial information, designs a one-dimensional convolutional neural network as a feature fusion component, processes the local context information of the temporal dimension layer by layer, aggregates the results of multiple layers of convolution, and outputs a spatio-temporal feature vector consistent with the video length. The inference module generates preliminary action predictions by inputting spatio-temporal features, including action categories, time intervals, and confidence levels, and uses a method of layer-by-layer calculation for prediction inference, combining the outputs of the action category prediction branch and the temporal prediction branch. The detection model module works collaboratively through two identical detection models, which are respectively responsible for pseudo-label generation and model weight update, generates pseudo-labels by combining all features, including action categories and their corresponding time intervals, and sorts and evaluates them. The self-fine-tuning module performs denoising processing on pseudo-labels with quality rankings lower than the threshold by using the inference ability of the supervised model, adds random offsets to the low-quality pseudo-labels to generate multiple perturbed temporal intervals, obtains different prediction results through inference, aggregates the prediction results to generate denoised temporal intervals, and further optimizes the training model using the denoised pseudo-labels. Through a temporal scaling enhancement strategy based on Gaussian sampling and classification loss and temporal interval loss functions, the dynamic improvement of the model performance is completed.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for denoising pseudo-labels of video action detection based on self-training according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for denoising pseudo-labels of video action detection based on self-training according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Weak supervision video behavior detection method and system based on iterative learning

    CN111797771A

  • Semi-supervised time sequence action positioning method, system, equipment and medium

    CN116363755A

  • Weak supervision video anomaly detection method based on label noise perception strategy

    CN120071213A

  • End-to-end video action detection and positioning system

    WO2022134655A1