Weak supervision time sequence action positioning method for fuzzy fragment enhancement and false positive suppression

Through the methods of fuzzy fragment enhancement and false positive inhibition, the problems of inaccurate prediction of action boundary and many false positive fragments in weakly supervised timing action positioning are solved, and higher accuracy and robustness of action positioning are achieved.

CN119942055APending Publication Date: 2025-05-06TIANJIN UNIVERSITY OF TECHNOLOGY +6
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510057291.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing weakly supervised timing action positioning methods still have shortcomings in improving the accuracy of action boundary prediction, especially in distinguishing semantic similar fragments and suppressing false positive fragments.

Method used

Through the methods of fuzzy fragment enhancement and false positive inhibition, the fuzzy fragments are mined using the foreground attention scores of the two modalities RGB and optical flow, and their discriminability is enhanced by comparative learning loss constraints. At the same time, action background separation is performed through clustering operations and cross-entropy loss constraints, and false positive suppression is performed using false positive fragment masks.

Benefits of technology

The discrimination of fuzzy fragments and the separation accuracy of the action background are improved, false positive fragments are effectively suppressed, and the accuracy and robustness of action positioning are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942055A_ABST
    Figure CN119942055A_ABST
Patent Text Reader

Abstract

The invention relates to a weak supervision time sequence action positioning method for fuzzy fragment enhancement and false positive suppression, and belongs to the field of computer vision. The method comprises the following steps: acquiring data; classifying foreground attention scores and fragment-level actions; enhancing fuzzy fragments; separating action backgrounds; false positive inhibition; video-level action classification and localization. According to the method, positive and negative sample pairs are constructed for fuzzy fragments, and the semantic correlation between the fuzzy fragments and discriminable actions and background fragments is increased by adopting contrast learning loss constraints, so that the discriminability of the fuzzy fragments is enhanced, and foreground and background separation is better carried out; in addition, false positive suppression is carried out on the original activation sequence according to a false positive fragment mask and a calculated false positive score, the obtained false positive suppressed activation sequence is used as a false label for supervising loss constraint, the original activation sequence is corrected, the purpose of suppressing the false positive fragment is achieved, and a more accurate action positioning effect can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of computer vision, and in particular relates to a weakly supervised temporal action localization method for blur segment enhancement and false positive suppression. Background Art

[0002] With the growth of online video data, how to understand and utilize video content efficiently and accurately has become an urgent problem to be solved. The study of weakly supervised temporal action localization provides an effective method for mining valuable information in large-scale video data. It uses only video-level labels for temporal action localization, which simplifies the data collection process. Due to the lack of precise temporal boundaries, existing work on WTAL mainly follows the localization-classification pipeline. The temporal class activation sequence of each input video is learned by feeding features into the classifier, which represents the probability that each frame in the video belongs to each action class. In the test phase, the boundaries of action instances in the video can be located by the threshold score on the temporal class activation sequence. In order to effectively distinguish the action background, many studies have improved the attention mechanism, suppressed the background activation score, highlighted the action activation score, or mined more action-related motion and scene information through feature enhancement to improve the accuracy of localization. Some work also uses video information to generate pseudo labels to improve the quality of class activation sequences. Although these methods have achieved significant improvements on WTAL, improving the accuracy of action boundary prediction is still a problem that needs to be solved.

[0003] On the one hand, there are still some semantically similar clips in the video that are difficult to distinguish, which makes it impossible to accurately separate the action background, resulting in incomplete and inaccurate action localization. On the other hand, since action localization is performed according to the classification localization pipeline, the generated activation sequence is often disturbed by category-related scenes, resulting in a large number of false positive clips in the prediction results, resulting in positioning errors. Summary of the invention

[0004] In order to solve the above problems, the present invention provides a weakly supervised temporal action localization method for blur segment enhancement and false positive suppression.

[0005] In order to achieve the above object, the present invention is implemented through the following technical solutions: The present invention provides a weakly supervised temporal action localization method for blur segment enhancement and false positive suppression, comprising the following steps: S1. Data acquisition: obtain video data and perform image analysis on the video data to obtain RGB enhancement features and optical flow enhancement features; S2. Foreground attention score and segment-level action classification: The RGB enhanced features and the optical flow enhanced features are sent to the attention module to obtain the foreground attention scores of the two modalities of RGB and optical flow, and then the two features are spliced ​​in the channel dimension, and then the features are fused to obtain the fused features, and then the fused features are sent to the classification module to obtain the temporal class activation sequence; S3. Blurred segment enhancement: We mine blurred segments and discriminative action background segments based on the foreground attention scores of the two modalities of RGB and optical flow, then establish positive and negative sample pairs of blurred action and background, and use contrastive learning loss to constrain blurred segments; S4. Action-background separation: Cluster the weighted foreground attention scores of the two modalities RGB and optical flow to obtain the foreground attention score segments belonging to the action and the background. Then, calculate the video-level action score and background score respectively according to the foreground attention score. Then, calculate the cross entropy loss with the video-level action-background label constraint. Perform the same clustering operation on the temporal class activation sequence to obtain the video-level action score and background score. S5. False positive suppression: The intersection of the segments classified as background by clustering operation based on foreground attention score and the segments classified as action by clustering operation based on temporal class activation sequence is taken as false positive segments to obtain false positive segment masks. The temporal class activation sequence obtained in step S2 is subjected to false positive suppression to obtain a false positive suppressed temporal class activation sequence. The temporal class activation sequence is then used as a pseudo-label with supervised loss constraint to correct the original activation sequence during training. S6. Video-level action classification and localization: Multiply the temporal class activation sequence with the foreground attention score of the optical flow modality to obtain the temporal class activation sequence with background suppression. Use the top-k aggregation strategy to obtain the video-level action classification probability and the action localization result. The training process is supervised by the video-level category label.

[0006] Further, step S1 is specifically: Obtain video data, divide each video into non-overlapping segments of 16 frames, sample a fixed number of segments to represent the video, use the TVL1 algorithm to extract RGB images and optical flow images in the video, use the I3D network pre-trained on the Kinetics dataset to perform image analysis on the RGB images and optical flow images to obtain RGB features and optical flow features; use DDG-Net as the basic network to perform data enhancement on RGB features and optical flow features to obtain RGB enhanced features and optical flow enhancement features .

[0007] Furthermore, step S2 is specifically as follows: Enhance the RGB feature and optical flow enhancement features The foreground attention scores of the RGB modality and the optical flow modality are respectively sent to the attention module to obtain the foreground attention scores of the RGB modality and the optical flow modality; the RGB enhanced features are and optical flow enhancement features Splicing is performed in the channel dimension to obtain a splicing feature with a channel dimension of 2048, and then a convolution operation is used to fuse the features to obtain a fused feature. , and then the fusion features Send it to the classification module to get the time-series class activation sequence , the formula is as follows: , , , in, represents the foreground attention score in RGB modality, represents the foreground attention score of the optical flow modality, represents the attention module, Represents the fusion feature, where N represents the number of samples, T represents the time dimension, and D represents the feature dimension. Indicates concatenating two features in the channel dimension. represents the RELU activation function, represents the convolution operation, represents the temporal class activation sequence, Represents a classifier.

[0008] Furthermore, step S3 is specifically as follows: According to the threshold hyperparameters set , the part where the foreground attention scores of the two modalities are both greater than the threshold is taken as the action segment, and the part where the foreground attention scores of the two modalities are both less than the threshold is taken as the background segment, and finally the discriminable action segment is obtained and discriminable background fragments , the formula is as follows: , , Among them, t represents the fragment, represents the foreground attention score of the RGB modality of segment t, represents the foreground attention score of the optical flow modality of segment t, is the threshold hyperparameter, Represents the features enhanced by DDG-Net; For features For the remaining segments in step S2, according to the temporal class activation sequence obtained in step S2, the k segments with the largest activation scores and the k segments with the smallest scores are selected in the time dimension as fuzzy action segments. and blur background clips , and finally according to the blurred action clip and blur background clips Construct positive and negative sample pairs, and set the positive samples of the blurred action clips as discriminable action clips , negative samples are set to discriminate background fragments , set the positive samples of blurred background segments as discriminable background segments , negative samples are set to discriminate action clips , the formula is as follows: , Constrained Blurred Action Clips Using Contrastive Learning Loss and blur background clips , enhance the discriminability of blurred fragments, and the loss function formula is as follows: , in, represents the contrastive learning loss, represents the dimension hyperparameter, Indicates the number of blurred segments, Represents an exponential function.

[0009] Furthermore, step S4 is specifically as follows: Foreground attention score for RGB modality and foreground attention score in optical flow modality Take a weighted average to get a weighted foreground attention score , the time-series class activation sequence obtained in step S2 Remove the last column of the last dimension, then sum it up in the channel dimension and perform a Sigmoid operation to get the segment-level action score , the weighted attention score and segment-level action scores Perform clustering operations separately to obtain the foreground attention score fragments belonging to the action , the foreground attention score fragment belonging to the background , the temporal activation sequence fragment belonging to the action And the temporal activation sequence fragments belonging to the background , the process formula is as follows: , , , , in, represents the first hyperparameter, represents the second hyperparameter, It means adding the time-series class activation sequence S in the last dimension. Represents a sequential class activation sequence Remove the last column of the last dimension, Sigmoid represents the activation function, represents the segment-level action score, cluster represents the clustering operation, represents the foreground attention score fragment belonging to the action, represents the foreground attention score fragment belonging to the background, represents a temporal class activation sequence fragment belonging to an action, represents a temporal class activation sequence fragment belonging to the background; The specific process of clustering operation is as follows: Sort the foreground attention score A in descending order in the time dimension, take the first value after sorting as the action class, and the last value as the background class, and initialize the action class center and background center , and then starting from the second value, process the sorted values ​​one by one, and calculate the distance from the current value to the center of the action class and the center of the background class. The formula is as follows: , , , , in, represents the foreground attention score after sorting in the time dimension, represents the foreground attention score of the sorted fragment t in the time dimension, represents the distance between segment t and the center of the action class, represents the distance between segment t and the background center, T represents the time dimension, Represents the first value after sorting, Indicates the last value after sorting, sort indicates sorting in descending order according to the last dimension, and Respectively represent the currently initialized action center and background center; determine whether the current segment belongs to the action class or the background class according to the distance from the current segment to the action center and the background center, then update the action center and the background center and update the number of segments of the two clustering centers until all segments are processed and the clustering operation ends. The formula is as follows: , , in, represents the foreground attention score of segment t in the time dimension, represents the foreground attention score of the segment k in the current action class in the temporal dimension, represents the foreground attention score of the segment m in the current background class in the temporal dimension, and Represent the number of clips in the action class and the number of clips in the background class respectively; the final action class center score and background class center score As the video-level action score and video-level background score respectively, finally according to the video-level action background label ={1,1}, using cross entropy loss Constraints are made, and the formula is expressed as follows: , , , in, and They represent the video-level action score and the video-level background score respectively, C represents the number of categories, and cat represents the concatenation operation; For clip-level action scores , the same clustering operation is performed to obtain the temporal activation sequence fragments belonging to the action and the temporal class activation sequence fragments belonging to the background , and then used as the video-level action score and video-level background score respectively, according to the video-level action background label ={1,1}, and use cross entropy loss for constraints, the formula is as follows: .

[0010] Furthermore, step S5 is specifically as follows: Set the position of the foreground attention score fragment belonging to the action to 1, and set the foreground attention score fragment belonging to the background to 0, and get the action fragment mask based on the foreground attention score and background fragment mask Similarly, the position of the temporal class activation sequence fragment belonging to the action is set to 1, and the temporal class activation sequence fragment belonging to the background is set to 0, and the action fragment mask based on the temporal class activation sequence is obtained. , then and The intersection of is taken as the false positive fragment to obtain the false positive fragment mask , and then get the temporal class activation sequence after false positive suppression , the formula is as follows: , , , , Among them, mask represents the false positive fragment mask, represents the temporal class activation sequence after false positive suppression of the action class, represents the temporal class activation sequence after false positive suppression of the background class, Represents a splicing operation; The temporal class activation sequence after false positive suppression As pseudo-labels, using supervised loss Constraints, the time series class activation sequence S is corrected, and the loss function is as follows: Where T represents the time dimension, represents the temporal class activation sequence after pseudo-positive inhibition in the i-th time dimension, It indicates that the i-th is the time-series class activation sequence of the time dimension.

[0011] Furthermore, step S6 is specifically as follows: Activate the sequence of the time series class Foreground Attention Score with Optical Flow Modality Multiply to get the temporal class activation sequence of background suppression , and then use the top-k aggregation strategy to get the video-level action score for each action category , represents the score of each action category on the video, and then the softmax operation is performed to obtain the video-level classification probability P, and finally the classification and positioning results are obtained. The training process is supervised by the video-level action classification loss, and the formula is as follows: , , in, The background is suppressed in the time series class activation sequence, N represents the number of samples, C represents the number of categories, represents the video-level truth label, represents the video-level action classification loss, represents the action score of each category in each video, n represents the video in a sample, and c represents the action category.

[0012] The present invention also provides a weakly supervised temporal action localization system for fuzzy segment enhancement and false positive suppression, comprising: Data acquisition module: used to acquire video data and perform image analysis on the video data to obtain RGB enhancement features and optical flow enhancement features; Foreground attention score and segment-level action classification module: used to send the RGB enhanced features and optical flow enhanced features into the attention module to obtain the foreground attention scores of the two modalities of RGB and optical flow, then concatenate the two features in the channel dimension, and then perform feature fusion to obtain fused features, and then send the fused features into the classification module to obtain the temporal class activation sequence; Blurry segment enhancement module: used to mine blurry segments and discriminative action background segments based on the foreground attention scores of the two modalities of RGB and optical flow, then establish positive and negative sample pairs of blurred action and background, and use contrastive learning loss to constrain blurry segments; Action-background separation module: It is used to cluster the weighted foreground attention scores of the two modalities RGB and optical flow to obtain the foreground attention score segments belonging to the action and the background, and then calculate the video-level action score and background score respectively according to the foreground attention score. Then, the cross entropy loss is calculated using the video-level action-background label constraint. The same clustering operation is performed on the temporal class activation sequence to obtain the video-level action score and background score. False positive suppression module: It is used to take the intersection of the segments that are clustered as background based on the foreground attention score and the segments that are clustered as action based on the temporal class activation sequence as false positive segments, obtain false positive segment masks, and suppress false positives on the temporal class activation sequence to obtain the temporal class activation sequence with false positive suppression. Then, the temporal class activation sequence is used as a pseudo-label with the supervision loss constraint to correct the original activation sequence during the training process. Video-level action classification and localization module: It is used to multiply the temporal class activation sequence with the foreground attention score of the optical flow modality to obtain the temporal class activation sequence with background suppression, and use the top-k aggregation strategy to obtain the video-level action classification probability to obtain the action localization result.

[0013] The advantages of the present invention are: The present invention increases the semantic relevance of action and background segments in blurred segments by performing contrast loss constraints in a blurred segment enhancement module and establishing positive and negative sample pairs of blurred actions and backgrounds, thereby enhancing the discriminability of blurred segments and better performing forward background separation; the present invention uses a false positive suppression module to perform clustering operations based on foreground scores and activation sequence scores to obtain false positive segment masks, then performs false positive suppression on the original activation sequence through the mask and the false positive score to obtain a false positive suppressed activation sequence, then uses the new activation sequence as a pseudo-label and uses a supervised loss constraint to correct the original activation sequence, thereby achieving the purpose of suppressing false positive segments and obtaining robust positioning and classification features. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0015] Figure 1 is a flow chart of the steps of the method of the present invention; Figure 2 This is a comparison chart of the visualization results of the method of the present invention on the THUMOS14 dataset. DETAILED DESCRIPTION

[0016] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0017] Example 1 In this embodiment, Figure 1 As shown, the present invention provides a weakly supervised temporal action localization method for blur segment enhancement and false positive suppression, and the specific steps include: S1. Data acquisition: obtain video data and perform image analysis on the video data to obtain RGB enhancement features and optical flow enhancement features; Specifically, each video is divided into non-overlapping segments of 16 frames each, and a fixed number of segments are sampled to represent the video. The TVL1 algorithm is used to extract the RGB image and optical flow image in the video. The I3D network pre-trained on the Kinetics dataset is used to perform image analysis on the RGB image and optical flow image to obtain RGB features and optical flow features. DDG-Net is used as the basic network to perform data enhancement on the RGB features and optical flow features to obtain RGB enhanced features. and optical flow enhancement features .

[0018] S2. Foreground attention score and segment-level action classification: The RGB enhanced features and the optical flow enhanced features are sent to the attention module to obtain the foreground attention scores of the two modalities of RGB and optical flow, and then the two features are spliced ​​in the channel dimension, and then the features are fused to obtain the fused features, and then the fused features are sent to the classification module to obtain the temporal class activation sequence; Specifically, the RGB enhancement feature and optical flow enhancement features The foreground attention scores of the RGB modality and the optical flow modality are respectively sent to the attention module to obtain the foreground attention scores of the RGB modality and the optical flow modality; the RGB enhanced features are and optical flow enhancement features Splicing is performed in the channel dimension to obtain a splicing feature with a channel dimension of 2048, and then a convolution operation is used to fuse the features to obtain a fused feature. , and then the fusion features Send it to the classification module to get the time-series class activation sequence , the formula is as follows: , , , in, represents the foreground attention score in RGB modality, represents the foreground attention score of the optical flow modality, represents the attention module, Represents the fusion feature, where N represents the number of samples, T represents the time dimension, and D represents the feature dimension. Indicates concatenating two features in the channel dimension. represents the RELU activation function, represents the convolution operation, represents the temporal class activation sequence, Represents a classifier.

[0019] S3. Blurred segment enhancement: According to the foreground attention scores of the two modalities of RGB and optical flow, blurred segments and discriminative action background segments are mined, and then positive and negative sample pairs of blurred action and background are established. The contrastive learning loss is used to constrain the blurred segments, learn the feature information of positive samples, and enhance the discriminability of blurred segments; Specifically, according to the threshold hyperparameters set , the part where the foreground attention scores of the two modalities are both greater than the threshold is taken as the action segment, and the part where the foreground attention scores of the two modalities are both less than the threshold is taken as the background segment, and finally the discriminable action segment is obtained and discriminable background fragments , the formula is as follows: , , Among them, t represents the fragment, represents the foreground attention score of the RGB modality of segment t, represents the foreground attention score of the optical flow modality of segment t, is the threshold hyperparameter, Represents the features enhanced by DDG-Net; For features For the remaining segments in step S2, according to the temporal class activation sequence obtained in step S2, the k segments with the largest activation scores and the k segments with the smallest scores are selected in the time dimension as fuzzy action segments. and blur background clips , and finally according to the blurred action clip and blur background clips Construct positive and negative sample pairs, and set the positive samples of the blurred action clips as discriminable action clips , negative samples are set to discriminate background fragments , set the positive samples of blurred background segments as discriminable background segments , negative samples are set to discriminate action clips , the formula is as follows: , Constrained Blurred Action Clips Using Contrastive Learning Loss and blur background clips , enhance the discriminability of blurred fragments, and the loss function formula is as follows: , in, represents the contrastive learning loss, represents the dimension hyperparameter, Indicates the number of blurred segments, Represents an exponential function.

[0020] S4. Action-background separation: Cluster the weighted foreground attention scores of the two modalities RGB and optical flow to obtain the foreground attention score segments belonging to the action and the background. Then, calculate the video-level action score and background score respectively according to the foreground attention score. Then, calculate the cross entropy loss with the video-level action-background label constraint. Perform the same clustering operation on the temporal class activation sequence to obtain the video-level action score and background score. Specifically, the foreground attention score for the RGB modality and foreground attention score in optical flow modality Take a weighted average to get a weighted foreground attention score , the time-series class activation sequence obtained in step S2 Remove the last column of the last dimension, then sum it up in the channel dimension and perform a Sigmoid operation to get the segment-level action score , the weighted attention score and segment-level action scores Perform clustering operations separately to obtain the foreground attention score fragments belonging to the action , the foreground attention score fragment belonging to the background , the temporal activation sequence fragment belonging to the action And the temporal activation sequence fragments belonging to the background , the process formula is as follows: , , , , in, represents the first hyperparameter, represents the second hyperparameter, It means adding the time-series class activation sequence S in the last dimension. Represents a sequential class activation sequence Remove the last column of the last dimension, Sigmoid represents the activation function, represents the segment-level action score, cluster represents the clustering operation, represents the foreground attention score fragment belonging to the action, represents the foreground attention score fragment belonging to the background, represents a temporal class activation sequence fragment belonging to an action, represents a temporal class activation sequence fragment belonging to the background; The specific process of clustering operation is as follows: Sort the foreground attention score A in descending order in the time dimension, take the first value after sorting as the action class, and the last value as the background class, and initialize the action class center and background center , and then starting from the second value, process the sorted values ​​one by one, and calculate the distance from the current value to the center of the action class and the center of the background class. The formula is as follows: , , , , in, represents the foreground attention score after sorting in the time dimension, represents the foreground attention score of the sorted fragment t in the time dimension, represents the distance between segment t and the center of the action class, represents the distance between segment t and the background center, T represents the time dimension, Represents the first value after sorting, Indicates the last value after sorting, sort indicates sorting in descending order according to the last dimension, and Respectively represent the currently initialized action center and background center; determine whether the current segment belongs to the action class or the background class according to the distance from the current segment to the action center and the background center, then update the action center and the background center and update the number of segments of the two clustering centers until all segments are processed and the clustering operation ends. The formula is as follows: , , in, represents the foreground attention score of segment t in the time dimension, represents the foreground attention score of the segment k in the current action class in the temporal dimension, represents the foreground attention score of the segment m in the current background class in the temporal dimension, and Represent the number of clips in the action class and the number of clips in the background class respectively; the final action class center score and background class center score As the video-level action score and video-level background score respectively, finally according to the video-level action background label ={1,1}, using cross entropy loss Constraints are made, and the formula is expressed as follows: , , , in, and They represent the video-level action score and the video-level background score respectively, C represents the number of categories, and cat represents the concatenation operation; For clip-level action scores , the same clustering operation is performed to obtain the temporal activation sequence fragments belonging to the action and the temporal class activation sequence fragments belonging to the background , and then used as the video-level action score and video-level background score respectively, according to the video-level action background label ={1,1}, and use cross entropy loss for constraints, the formula is as follows: .

[0021] S5. False positive suppression: The intersection of the segments classified as background by clustering operation based on foreground attention score and the segments classified as action by clustering operation based on temporal class activation sequence is taken as false positive segments to obtain false positive segment masks. The temporal class activation sequence obtained in step S2 is subjected to false positive suppression to obtain a false positive suppressed temporal class activation sequence. The temporal class activation sequence is then used as a pseudo-label with supervised loss constraint to correct the original activation sequence during training. Specifically, the position of the foreground attention score fragment belonging to the action is set to 1, and the foreground attention score fragment belonging to the background is set to 0, and the action fragment mask based on the foreground attention score is obtained. and background fragment mask Similarly, the position of the temporal class activation sequence fragment belonging to the action is set to 1, and the temporal class activation sequence fragment belonging to the background is set to 0, and the action fragment mask based on the temporal class activation sequence is obtained. , then and The intersection of is taken as the false positive fragment to obtain the false positive fragment mask , and then get the temporal class activation sequence after false positive suppression , the formula is as follows: , , , , Among them, mask represents the false positive fragment mask, represents the temporal class activation sequence after false positive suppression of the action class, represents the temporal class activation sequence after false positive suppression of the background class, Represents a splicing operation; The temporal class activation sequence after false positive suppression As pseudo-labels, using supervised loss Constraints, the time series class activation sequence S is corrected, and the loss function is as follows: Where T represents the time dimension, represents the temporal class activation sequence after pseudo-positive inhibition in the i-th time dimension, It indicates that the i-th is the time-series class activation sequence of the time dimension.

[0022] S6. Video-level action classification and localization: Multiply the temporal class activation sequence with the foreground attention score of the optical flow modality to obtain the temporal class activation sequence with background suppression. Use the top-k aggregation strategy to obtain the video-level action classification probability and the action localization result. The training process is supervised by the video-level category label.

[0023] Specifically, the timing class activation sequence Foreground Attention Score with Optical Flow Modality Multiply to get the temporal class activation sequence of background suppression , and then use the top-k aggregation strategy to get the video-level action score for each action category , represents the score of each action category on the video, and then the softmax operation is performed to obtain the video-level classification probability P. The training process is supervised by the video-level action classification loss, and the formula is as follows: , , in, The background is suppressed in the time series class activation sequence, N represents the number of samples, C represents the number of categories, represents the video-level truth label, represents the video-level action classification loss, represents the action score of each category in each video, n represents the video in a sample, and c represents the action category. Then, according to the video-level classification probability P, the categories with video-level classification probability lower than a certain threshold are first discarded. For the remaining categories, the background segments are first discarded by setting a threshold on the foreground attention score, and then the continuous part of the remaining segments is used as a candidate action segment. After obtaining the proposal, the action confidence score q of each proposal is calculated, and finally the non-maximum suppression strategy is used to remove overlapping proposals to obtain the final classification and positioning results of each proposal, which is expressed as , where c represents the action category, Indicates the start time of the action. Indicates the end time of the action.

[0024] Example 2 During the testing phase, classes whose video-level class scores are lower than a certain threshold are first discarded. For the remaining classes, we first discard background segments by setting a threshold on the attention weights, and then take the continuous part of the remaining segments as a candidate action segment. After obtaining the proposal, the action confidence score q of each proposal is calculated. Finally, the soft non-maximum suppression strategy is used to remove overlapping proposals to obtain the final classification and positioning results and confidence scores of each proposal, expressed as (c, ts, te, q). The comparison of the experimental results of the present invention with other methods on the THUMOS14 dataset is shown in Table 1, and the comparison of the experimental results of the present invention with other methods on the ActivityNet1.2 dataset is shown in Table 2. Among them, Avg(0.1:0.5), Avg(0.1:0.7), and Avg(0.5:0.95) represent the mean average precision (mAP) of temporal intersection over Union (tIoU) from 0.1 to 0.5, 0.1 to 0.7, and 0.5 to 0.95, respectively. Fully represents the fully supervised method, Weakly+ represents the weakly supervised method with additional annotations or additional training data, and Weakly represents the weakly supervised method with only video-level label supervision.

[0025] Table 1 Comparison results of the experimental results of the present invention and other methods on the THUMOS14 dataset Table 2 Comparison of the experimental results of the present invention with other methods on the ActivityNet1.2 dataset It can be seen from Tables 1 and 2 that on the THUMOS14 dataset, the effect achieved by the present invention as a weakly supervised method with only video-level label supervision exceeds that of similar methods, reaching 49.0% on the average mAP when tIoU ranges from 0.1 to 0.7. Compared with the fully supervised method and the weakly supervised method with additional annotations or additional training data, the positioning effect of the present invention is significantly better. On the ActivityNet1.2 dataset, the effect achieved by the present invention also exceeds that of most methods, reaching 29.9% on the average mAP when tIoU ranges from 0.5 to 0.95.

[0026] In order to better demonstrate the effectiveness of this method, some comparisons of the detection results visualized on the THUMOS14 dataset are given. The visualization results are shown in Figure 2As shown in the figure. Two representative videos are selected from the THUMOS14 dataset for analysis. The gray area in the first row "Action" indicates the position of the real action clip in the video. The second row "Basic Branch" is the predicted probability of the baseline algorithm on the action category. The baseline algorithm is the DDG-Net method that removes the three modules of blurry clip enhancement, action background separation and false positive suppression. The third row is the predicted probability of the algorithm in this paper. In the first video clip, it can be seen from the position marked by the dotted box that our method can better identify the foreground clip and increase the probability of the foreground clip compared with the DDG-Net method that removes the three modules of blurry clip enhancement, action background separation and false positive suppression. This is due to the fact that the blurry clip enhancement module increases the feature discriminability and can more completely locate the action clip. In the second video clip, it can be seen from the part marked by the dotted box that compared with the DDG-Net method, the method in this paper can suppress some noisy clip information, remove clips that are not related to the action, effectively suppress false positive clips, and improve the accuracy of action positioning. The above analysis shows that our model can more accurately identify some blurry clips, and can effectively suppress the activation of false positive clips, which can effectively improve the performance of the action positioning task.

[0027] Example 3 This embodiment provides a weakly supervised temporal action localization system for blur segment enhancement and false positive suppression, including: Data acquisition module: used to acquire video data and perform image analysis on the video data to obtain RGB enhancement features and optical flow enhancement features; Foreground attention score and segment-level action classification module: used to send the RGB enhanced features and optical flow enhanced features into the attention module to obtain the foreground attention scores of the two modalities of RGB and optical flow, then concatenate the two features in the channel dimension, and then perform feature fusion to obtain fused features, and then send the fused features into the classification module to obtain the temporal class activation sequence; Blurry segment enhancement module: used to mine blurry segments and discriminative action background segments based on the foreground attention scores of the two modalities of RGB and optical flow, then establish positive and negative sample pairs of blurred action and background, and use contrastive learning loss to constrain blurry segments; Action-background separation module: It is used to cluster the weighted foreground attention scores of the two modalities RGB and optical flow to obtain the foreground attention score segments belonging to the action and the background, and then calculate the video-level action score and background score respectively according to the foreground attention score. Then, the cross entropy loss is calculated using the video-level action-background label constraint. The same clustering operation is performed on the temporal class activation sequence to obtain the video-level action score and background score. False positive suppression module: It is used to take the intersection of the segments that are clustered as background based on the foreground attention score and the segments that are clustered as action based on the temporal class activation sequence as false positive segments, obtain false positive segment masks, and suppress false positives on the temporal class activation sequence to obtain the temporal class activation sequence with false positive suppression. Then, the temporal class activation sequence is used as a pseudo-label with the supervision loss constraint to correct the original activation sequence during the training process. Video-level action classification and localization module: It is used to multiply the temporal class activation sequence with the foreground attention score of the optical flow modality to obtain the temporal class activation sequence with background suppression, and use the top-k aggregation strategy to obtain the video-level action classification probability to obtain the action localization result.

[0028] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A weakly supervised temporal action localization method with fuzzy segment enhancement and false positive suppression, characterized in that: The following steps are involved: S1. Data acquisition: obtain video data and perform image analysis on the video data to obtain RGB enhancement features and optical flow enhancement features; S2. Foreground attention score and segment-level action classification: The RGB enhanced features and the optical flow enhanced features are sent to the attention module to obtain the foreground attention scores of the two modalities of RGB and optical flow, and then the two features are spliced ​​in the channel dimension, and then the features are fused to obtain the fused features, and then the fused features are sent to the classification module to obtain the temporal class activation sequence; S3. Blurred segment enhancement: We mine blurred segments and discriminative action background segments based on the foreground attention scores of the two modalities of RGB and optical flow, then establish positive and negative sample pairs of blurred action and background, and use contrastive learning loss to constrain blurred segments; S4. Action-background separation: Cluster the weighted foreground attention scores of the two modalities RGB and optical flow to obtain the foreground attention score segments belonging to the action and the background. Then, calculate the video-level action score and background score respectively according to the foreground attention score. Then, calculate the cross entropy loss with the video-level action-background label constraint. Perform the same clustering operation on the temporal class activation sequence to obtain the video-level action score and background score. S5. False positive suppression: The intersection of the segments classified as background by clustering operation based on foreground attention score and the segments classified as action by clustering operation based on temporal class activation sequence is taken as false positive segments to obtain false positive segment masks. The temporal class activation sequence obtained in step S2 is subjected to false positive suppression to obtain a false positive suppressed temporal class activation sequence. The temporal class activation sequence is then used as a pseudo-label with supervised loss constraint to correct the original activation sequence during training. S6. Video-level action classification and localization: Multiply the temporal class activation sequence with the foreground attention score of the optical flow modality to obtain the temporal class activation sequence with background suppression. Use the top-k aggregation strategy to obtain the video-level action classification probability and the action localization result. The training process is supervised by the video-level category label.

2. The weakly supervised temporal action localization method with blur segment enhancement and false positive suppression according to claim 1 is characterized in that: Step S1 is specifically as follows: Get video data, divide each video into non-overlapping segments of 16 frames, sample a fixed number of segments to represent the video, use the TVL1 algorithm to extract RGB images and optical flow images from the video, use the I3D network pre-trained on the Kinetics dataset to perform image analysis on the RGB images and optical flow images, and obtain RGB features and optical flow features; DDG-Net is used as the basic network to enhance the RGB features and optical flow features to obtain RGB enhanced features. and optical flow enhancement features .

3. The weakly supervised temporal action localization method with blur segment enhancement and false positive suppression according to claim 2 is characterized in that: Step S2 is specifically as follows: Enhance the RGB feature and optical flow enhancement features The foreground attention scores of the RGB modality and the optical flow modality are respectively sent to the attention module to obtain the foreground attention scores of the RGB modality and the optical flow modality; the RGB enhanced features are and optical flow enhancement features Splicing is performed in the channel dimension to obtain a splicing feature with a channel dimension of 2048, and then a convolution operation is used to fuse the features to obtain a fused feature. , and then the fusion features Send it to the classification module to get the time-series class activation sequence , the formula is as follows: , , , in, represents the foreground attention score in RGB modality, represents the foreground attention score of the optical flow modality, represents the attention module, Represents the fusion feature, where N represents the number of samples, T represents the time dimension, and D represents the feature dimension. Indicates concatenating two features in the channel dimension. represents the RELU activation function, represents the convolution operation, represents the temporal class activation sequence, Represents a classifier.

4. The weakly supervised temporal action localization method with blur segment enhancement and false positive suppression according to claim 3 is characterized in that: Step S3 is specifically as follows: According to the threshold hyperparameters set , the part where the foreground attention scores of the two modalities are both greater than the threshold is taken as the action segment, and the part where the foreground attention scores of the two modalities are both less than the threshold is taken as the background segment, and finally the discriminable action segment is obtained and discriminable background fragments , the formula is as follows: , , Among them, t represents the fragment, represents the foreground attention score of the RGB modality of segment t, represents the foreground attention score of the optical flow modality of segment t, is the threshold hyperparameter, Represents the features enhanced by DDG-Net; For features For the remaining segments in step S2, according to the temporal class activation sequence obtained in step S2, the k segments with the largest activation scores and the k segments with the smallest scores are selected in the time dimension as fuzzy action segments. and blur background clips , and finally according to the blurred action clip and blur background clips Construct positive and negative sample pairs, and set the positive samples of the blurred action clips as discriminable action clips , negative samples are set to discriminate background fragments , set the positive samples of blurred background segments as discriminable background segments , negative samples are set to discriminate action clips , the formula is as follows: , Constrained Blurred Action Clips Using Contrastive Learning Loss and blur background clips , enhance the discriminability of blurred fragments, and the loss function formula is as follows: , in, represents the contrastive learning loss, represents the dimension hyperparameter, Indicates the number of blurred segments, Represents an exponential function.

5. The weakly supervised temporal action localization method with blur segment enhancement and false positive suppression according to claim 4 is characterized in that: Step S4 is specifically as follows: Foreground attention score for RGB modality and foreground attention score in optical flow modality Take a weighted average to get a weighted foreground attention score , the time-series class activation sequence obtained in step S2 Remove the last column of the last dimension, then sum it in the channel dimension and perform a Sigmoid operation to get the segment-level action score , the weighted attention score and segment-level action scores Perform clustering operations separately to obtain the foreground attention score fragments belonging to the action , the foreground attention score fragment belonging to the background , the temporal activation sequence fragment belonging to the action And the temporal activation sequence fragments belonging to the background , the process formula is as follows: , , , , in, represents the first hyperparameter, represents the second hyperparameter, It means adding the time-series class activation sequence S in the last dimension. Represents a sequential class activation sequence The time-series class activation sequence after removing the last column of the last dimension, Sigmoid represents the activation function, represents the segment-level action score, cluster represents the clustering operation, represents the foreground attention score fragment belonging to the action, represents the foreground attention score fragment belonging to the background, represents a temporal class activation sequence fragment belonging to an action, represents a fragment of temporal class activation sequence belonging to the background; The specific process of clustering operation is as follows: Sort the foreground attention score A in descending order in the time dimension, take the first value after sorting as the action class, and the last value as the background class, and initialize the action class center and background center , and then starting from the second value, process the sorted values ​​one by one, and calculate the distance from the current value to the center of the action class and the center of the background class. The formula is as follows: , , , , in, represents the foreground attention score after sorting in the time dimension, represents the foreground attention score of the sorted fragment t in the time dimension, represents the distance between segment t and the center of the action class, represents the distance between segment t and the background center, T represents the time dimension, Represents the first value after sorting, Indicates the last value after sorting, sort indicates sorting in descending order according to the last dimension, and Respectively represent the currently initialized action center and background center; determine whether the current segment belongs to the action class or the background class according to the distance from the current segment to the action center and the background center, then update the action center and the background center and update the number of segments of the two clustering centers until all segments are processed and the clustering operation ends. The formula is as follows: , , in, represents the foreground attention score of segment t in the time dimension, represents the foreground attention score of the segment k in the current action class in the temporal dimension, represents the foreground attention score of the segment m in the current background class in the temporal dimension, and Represent the number of clips in the action class and the number of clips in the background class respectively; the final action class center score and background class center score As the video-level action score and video-level background score respectively, finally according to the video-level action background label ={1,1}, using cross entropy loss Constraints are made, and the formula is expressed as follows: , , , in, and They represent the video-level action score and the video-level background score respectively, C represents the number of categories, and cat represents the concatenation operation; For clip-level action scores , the same clustering operation is performed to obtain the temporal activation sequence fragments belonging to the action and the temporal class activation sequence fragments belonging to the background , and then used as the video-level action score and video-level background score respectively, according to the video-level action background label ={1,1}, and use cross entropy loss for constraints, the formula is as follows: 。 6. The weakly supervised temporal action localization method with blur segment enhancement and false positive suppression according to claim 5 is characterized in that: Step S5 is specifically as follows: The position of the foreground attention score fragment belonging to the action is set to 1, and the foreground attention score fragment belonging to the background is set to 0, and the action fragment mask based on the foreground attention score is obtained. and background fragment mask Similarly, the position of the temporal class activation sequence fragment belonging to the action is set to 1, and the temporal class activation sequence fragment belonging to the background is set to 0, and the action fragment mask based on the temporal class activation sequence is obtained. , then and The intersection of is taken as the false positive fragment to obtain the false positive fragment mask , and then get the temporal class activation sequence after false positive suppression , the formula is as follows: , , , , Among them, mask represents the false positive fragment mask, represents the temporal class activation sequence after false positive suppression of the action class, represents the temporal class activation sequence after false positive suppression of the background class, Represents a splicing operation; The temporal class activation sequence after false positive suppression As pseudo-labels, using supervised loss Constraints, the time series class activation sequence S is corrected, and the loss function is as follows: Where T represents the time dimension, represents the temporal class activation sequence after pseudo-positive inhibition in the i-th time dimension, It indicates that the i-th is the time-series class activation sequence of the time dimension.

7. The weakly supervised temporal action localization method with blur segment enhancement and false positive suppression according to claim 6 is characterized in that: Step S6 is specifically as follows: Activate the sequence of the time series class Foreground Attention Score with Optical Flow Modality Multiply to get the temporal class activation sequence of background suppression , and then use the top-k aggregation strategy to get the video-level action score for each action category , represents the score of each action category on the video, and then the softmax operation is performed to obtain the video-level classification probability P, and finally the classification and positioning results are obtained. The training process is supervised by the video-level action classification loss, and the formula is as follows: , , in, The background is suppressed in the time series class activation sequence, N represents the number of samples, C represents the number of categories, represents the video-level truth label, represents the video-level action classification loss, represents the action score of each category in each video, n represents the video in a sample, and c represents the action category.

8. A system using the weakly supervised temporal action localization method for blur segment enhancement and false positive suppression according to claim 1, characterized in that: include: Data acquisition module: used to acquire video data and perform image analysis on the video data to obtain RGB enhancement features and optical flow enhancement features; Foreground attention score and segment-level action classification module: used to send the RGB enhanced features and optical flow enhanced features into the attention module to obtain the foreground attention scores of the two modalities of RGB and optical flow, then concatenate the two features in the channel dimension, and then perform feature fusion to obtain fused features, and then send the fused features into the classification module to obtain the temporal class activation sequence; Blurry segment enhancement module: used to mine blurry segments and discriminative action background segments based on the foreground attention scores of the two modalities of RGB and optical flow, then establish positive and negative sample pairs of blurred action and background, and use contrastive learning loss to constrain blurry segments; Action-background separation module: It is used to cluster the weighted foreground attention scores of the two modalities RGB and optical flow to obtain the foreground attention score segments belonging to the action and the background, and then calculate the video-level action score and background score respectively according to the foreground attention score. Then, the cross entropy loss is calculated using the video-level action-background label constraint. The same clustering operation is performed on the temporal class activation sequence to obtain the video-level action score and background score. False positive suppression module: It is used to take the intersection of the segments that are clustered as background based on the foreground attention score and the segments that are clustered as action based on the temporal class activation sequence as false positive segments, obtain false positive segment masks, and suppress false positives on the temporal class activation sequence to obtain the temporal class activation sequence with false positive suppression. Then, the temporal class activation sequence is used as a pseudo-label with the supervision loss constraint to correct the original activation sequence during the training process. Video-level action classification and localization module: It is used to multiply the temporal class activation sequence with the foreground attention score of the optical flow modality to obtain the temporal class activation sequence with background suppression, and use the top-k aggregation strategy to obtain the video-level action classification probability to obtain the action localization result.