Video action detection method based on random mask

Through random mask pre-training and symmetric decoder design, the problems of spatiotemporal redundancy and locality traps in temporal action detection are solved, the positioning accuracy and robustness of the model are improved, and it is adapted to multi-scale video scenes.

CN120708123APending Publication Date: 2025-09-26ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510805346.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing video action detection methods have problems of spatiotemporal redundancy interference and local traps in temporal action detection, and self-supervised pre-training methods are not adaptable enough in the video field and cannot effectively restore high-level semantic features.

Method used

A random mask-based pre-training method is adopted to pre-train the encoder by constructing a mask training set. Combined with the hierarchical masking strategy and symmetric decoder design, the positioning accuracy and cross-architecture versatility of the model are improved.

Benefits of technology

It significantly improves the positioning accuracy and model robustness of temporal action detection, takes into account efficiency, accuracy and generalization, and adapts to the spatiotemporal characteristics of different video scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708123A_ABST
    Figure CN120708123A_ABST
Patent Text Reader

Abstract

The invention discloses a video action detection method based on a random mask. The method comprises the following steps: firstly, constructing an original training set; then, performing random mask processing on the original training set to obtain a mask training set; then, a first neural network model is constructed, and the first neural network model comprises an encoder and a decoder which are connected; then training the first neural network model by using the mask training set until the training is completed, and obtaining a trained encoder; constructing a second neural network model by using the trained encoder, training the second neural network model by using the original training set until the training is completed, and obtaining a time sequence action detection neural network model; and finally, preprocessing a to-be-detected action video, inputting the preprocessed to-be-detected action video into the time sequence action detection neural network model, and outputting a corresponding time sequence action detection result by the model. According to the method, the accuracy of time sequence action detection can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video action detection method, in particular to a video action detection method based on random mask. Background Art

[0002] As a core technology in video understanding, temporal action detection aims to accurately locate the start and end times and categories of action instances in long, untrimmed videos. It is widely used in scenarios such as intelligent surveillance and video retrieval. Existing mainstream methods typically adopt a two-stage paradigm: a pre-trained feature extractor generates a low-dimensional feature sequence, which is then fed into a detection head for prediction.

[0003] However, such methods face two core challenges: first, spatiotemporal redundancy interference, that is, a large amount of irrelevant background noise is mixed in the features, resulting in the dilution of key action semantics; second, locality trap, mainstream architectures have difficulty modeling long-range action dependencies due to local receptive fields or fixed window mechanisms.

[0004] Although self-supervised pre-training alleviates the dependence on data annotation through mask reconstruction tasks, its direct migration to the video field has significant limitations: the general mask strategy ignores the multi-scale temporal characteristics of the video, and the random mask is difficult to adapt to the temporal and spatial measurement differences of action instances; the lightweight decoder is good at pixel-level reconstruction, but cannot restore the high-level semantic features required for temporal action detection, exacerbating the domain differences between pre-training and downstream tasks.

[0005] Therefore, it is necessary to propose a video action detection method based on a new pre-training method. Summary of the Invention

[0006] To address the incompatibility of existing pre-training methods for sequential action detection, this paper proposes a video action detection method based on random masking. This method significantly improves the positioning accuracy of the sequential action detection model without changing the original model.

[0007] The technical solution adopted in the present invention is as follows:

[0008] 1. A video action detection method based on random mask

[0009] 1) Construct an original training set based on action videos;

[0010] 2) After performing random masking on the original training set, a masked training set is obtained;

[0011] 3) constructing a first neural network model, the first neural network model including a connected encoder and decoder; then training the first neural network model using the masked training set until the training is completed to obtain a trained encoder;

[0012] 4) Using the encoder trained in 3) to build a second neural network model, and then using the original training set to train the second neural network model until the training is completed, thereby obtaining a time series action detection neural network model;

[0013] 5) The action video to be detected is preprocessed and then input into the temporal action detection neural network model, and the model outputs the corresponding temporal action detection results.

[0014] The specific embodiment of 1) is:

[0015] First, each acquired action video is divided into several image sequences of equal length. Then, a feature extractor is used to extract features from each image sequence to obtain the corresponding original feature sequence. The features in each original feature sequence are then preprocessed to obtain the original training set.

[0016] The preprocessing includes downsampling and random cropping operations.

[0017] Said 2) is specifically:

[0018] First, a low-resolution mask template is established. Then, a random mask of a preset ratio is applied to the low-resolution mask template to obtain an initialized low-resolution mask template. The initialized low-resolution mask template is then upsampled to a high-resolution mask template through bilinear interpolation, and the high-resolution mask template is then dimensionally expanded to obtain the final mask template. Finally, the final mask template is applied to the original training set to obtain the mask training set.

[0019] In the above 3), the encoder is an encoder that outputs a time series pyramid feature, and the time series pyramid feature is composed of a first time series feature, a second time series feature, a third time series feature and a fourth time series feature; the decoder includes an upsampling convolution layer, a main module and a convolution mapping layer, the first time series feature is input to the first upsampling convolution layer, and the first upsampling convolution layer outputs the first upsampling time series feature; the first upsampling time series feature is used as the input of the first main module, and the first main module outputs the first upsampling main time series feature; the first upsampling main time series feature and the second time series feature are added together to obtain a second fused time series feature; the second fused time series feature is used as the input of the second upsampling convolution layer, and the second upsampling convolution layer outputs the second upsampling time series feature; the second upsampling time series feature is used as the input of the second main module, and the second main module outputs the second upsampling main time series feature; the second upsampling main The third fused temporal feature is obtained by adding the temporal feature to the third temporal feature; the third fused temporal feature is used as the input of the third upsampling convolution layer, and the third upsampling convolution layer outputs the third upsampling temporal feature; the third upsampling temporal feature is used as the input of the third main module, and the third main module outputs the third upsampling main temporal feature; the third upsampling main temporal feature and the fourth temporal feature are added to obtain the fourth fused temporal feature; the fourth fused temporal feature is used as the input of the fourth upsampling convolution layer, and the fourth upsampling convolution layer outputs the fourth upsampling temporal feature; the fourth upsampling temporal feature is used as the input of the fourth main module, and the fourth main module outputs the fourth upsampling main temporal feature; the fourth fused main temporal feature is used as the input of the first convolution mapping layer, the first convolution mapping layer is connected to the second convolution mapping layer, and the second convolution mapping layer outputs the reconstructed feature as the output of the decoder.

[0020] The second neural network model includes a trained encoder and a classification head connected to the trained encoder.

[0021] In the above 3), the loss function in the training process of the first neural network model is the MSE loss calculated based on the output of the first neural network model and the unmasked part in the mask training set.

[0022] 2. A computer device

[0023] The device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method for video action detection based on random mask when executing the computer program.

[0024] 3. A computer-readable storage medium

[0025] The medium stores a computer program, which, when executed by a processor, implements the steps of the method for detecting video motion based on random masks.

[0026] 4. A computer program product

[0027] The product includes a computer program / instruction, which, when executed by a processor, implements the steps of the method for video action detection based on random mask.

[0028] The beneficial effects of the present invention are:

[0029] 1) The present invention constructs a mask training set by applying random masks, and uses the mask training set to pre-train the encoder in the first neural network model. This can adapt to the mainstream temporal action detection method architecture and has strong cross-architecture versatility and robustness.

[0030] 2) This paper proposes a hierarchical masking strategy (i.e., the process of generating mask templates from low-resolution to high-resolution mask templates), actively masking redundant areas and forcing the model to learn the global temporal structure from local observations, which can significantly improve the model's ability to focus on key semantics.

[0031] 3) The present invention constructs a symmetric decoder during the reconstruction process. The symmetric decoder design ensures the consistency of high-level semantic reconstruction and promotes the alignment of pre-trained features with downstream tasks.

[0032] In summary, the present invention provides a pre-training method for temporal action detection that balances efficiency, accuracy, and generalization. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a network structure diagram of the first neural network model adopted by the present invention.

[0034] Figure 2 This is the network structure diagram of the decoder in the first neural network model.

[0035] Figure 3 This is a network structure diagram of the sequential action detection neural network model used in the present invention.

[0036] Figure 4 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0037] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0038] The specific embodiments of the present invention are as follows:

[0039] like Figure 4 As shown, the present invention proposes a video action detection method based on random mask, which includes the following steps:

[0040] 1) Construct an original training set based on the acquired action videos;

[0041] 1) Specifically:

[0042] First, each action video is segmented into non-overlapping 16-frame segments, forming an image sequence. The I3D network is used as a feature extractor to extract features from each image sequence, obtaining a raw feature sequence. The features in each raw feature sequence are then preprocessed to form the original training set. Preprocessing includes random cropping and downsampling. Specifically, each image sequence is temporally downsampled to obtain unjoined pre-extracted features. All segments of the same video are spliced ​​temporally using the I3D network to obtain the raw feature sequence. The raw feature sequence is divided into training and validation sets using official annotations. Then, after applying a random cropping range of 0.9 to 1 during the pre-training and training phases, pre-extracted feature segments of 256 temporal length are randomly obtained as input. During the validation phase, no cropping is applied, and pre-extracted feature segments of 256 temporal length are randomly obtained as input. The shape of the pre-extracted feature segments is 256×768.

[0043] The action video dataset of the present invention adopts the THUMOS14 temporal action detection dataset. There are unsegmented videos of 20 categories of actions in the dataset, including 200 verification set videos (containing 3007 behavior fragments) and 213 test set videos (containing 3358 behavior fragments). These annotated unsegmented videos can be used to train and test temporal behavior detection models. There are an average of 150 action timing annotations for each category on the verification set used for training, and the average duration of each action is 4.04 seconds. There are a total of 3007 action timing annotations, and the annotated actions last a total of 12159.8 seconds. There are an average of 167.9 action timing annotations for each category on the test set, and the average duration of each action is 4.47 seconds. There are a total of 3358 action timing annotations, and the annotated actions last a total of 15040.3 seconds.

[0044] 2) After performing random masking on the original training set, a masked training set is obtained;

[0045] 2) Specifically:

[0046] First, a low-resolution mask template with all zeros and a length of 16 is created. Next, a random mask with a preset ratio (e.g., 75%) is applied to the low-resolution mask template. This initialization of the low-resolution mask template is achieved by setting 75% of the content in the low-resolution mask template to 1, with 1s representing masked regions and 0s representing retained regions. This low-resolution mask template is then upsampled to a high-resolution mask template with a length of 256 using bilinear interpolation, ensuring that the masked regions are continuously aligned along the temporal axis. The high-resolution mask template is then dimensionally expanded to obtain a final mask template with a shape of 256×768. Finally, the final mask template is applied to each image sample in the original training set to obtain the mask training set. The application of the mask template ensures that the features at any resolution and at any time step have the same mask ratio.

[0047] 3) constructing a first neural network model, the first neural network model including a connected encoder and decoder; then training the first neural network model using the masked training set until the training is completed to obtain a trained encoder;

[0048] 3) In the Figure 1 As shown, the encoder is an encoder that outputs a temporal pyramid feature, and the temporal pyramid feature is composed of a first temporal feature, a second temporal feature, a third temporal feature, and a fourth temporal feature; unlike the traditional mask pre-training method which requires the construction of a lightweight decoder, the decoder constructed by the present invention is as follows Figure 2As shown, it has a similar parameter amount to that of the encoder, which can better realize the reconstruction of the timing features. The parameter amount of the decoder is similar to that of the encoder. The decoder includes an upsampling convolution layer, a main module and a convolution mapping layer. The first timing feature is input to the first upsampling convolution layer, and the first upsampling convolution layer outputs the first upsampling timing feature; the first upsampling timing feature is used as the input of the first main module, and the first main module outputs the first upsampling main timing feature; the first upsampling main timing feature and the second timing feature are added to obtain the second fused timing feature; the second fused timing feature is used as the input of the second upsampling convolution layer, and the second upsampling convolution layer outputs the second upsampling timing feature; the second upsampling timing feature is used as the input of the second main module, and the second main module outputs the second upsampling main timing feature; the second upsampling main timing feature is added to the third timing feature to obtain the third fused timing feature. features; using the third fused time series features as the input of the third upsampling convolutional layer, and the third upsampling convolutional layer outputs the third upsampling time series features; using the third upsampling time series features as the input of the third main module, and the third main module outputs the third upsampling main time series features; adding the third upsampling main time series features and the fourth time series features to obtain the fourth fused time series features; using the fourth fused time series features as the input of the fourth upsampling convolutional layer, and the fourth upsampling convolutional layer outputs the fourth upsampling time series features; using the fourth upsampling time series features as the input of the fourth main module, and the fourth main module outputs the fourth upsampling main time series features; using the fourth fused main time series features as the input of the first convolutional mapping layer, the first convolutional mapping layer is connected to the second convolutional mapping layer, and the second convolutional mapping layer outputs the reconstructed features as the output of the decoder.

[0049] The first upsampling convolution layer, the second upsampling convolution layer, the third upsampling convolution layer and the fourth upsampling convolution layer are all transposed convolutions with a kernel size of 3, a stride of 2 and a padding of 1.

[0050] The first convolutional mapping layer and the second convolutional mapping layer are convolutional layers with a convolution kernel size of 3, a stride of 1, and a padding of 1.

[0051] The main module in the present invention can be the main module of any temporal action detection model. The present invention is described using the main module adopted by the Actionformer model.

[0052] The encoder in the present invention can be an encoder of any temporal action detection model. The present invention is described using an encoder of the Actionformer model.

[0053] The temporal pyramid features used in the present invention are intermediate features generated by all temporal action detection models.

[0054] like Figure 3As shown, the second neural network model includes a trained encoder and a classification head connected to the trained encoder.

[0055] 3), the loss function during the training of the first neural network model is the MSE loss calculated based on the output of the first neural network model (i.e., the reconstructed features output by the decoder) and the unmasked portion of the mask training set. While the reconstruction target of traditional mask pre-training is the masked portion of the mask training set, the reconstruction target of the present invention is exactly the opposite. This approach can effectively avoid the feature leakage problem that may be caused by the main module of the convolutional base.

[0056] 4) Using the encoder trained in 3) to construct a second neural network model, i.e., the encoder in the second neural network model has the same structure as the encoder in the first neural network model, and the trained weights of the encoder are also imported into the second neural network model. The second neural network model is then trained using the original training set until training is complete, thereby obtaining a time-series action detection neural network model;

[0057] The time sequence action detection neural network of the present invention can be any time sequence action detection network. The present invention is described using the Actionformer model.

[0058] 5) Input the test set into the temporal action detection neural network model, and the model outputs the corresponding temporal action detection results.

[0059] The pre-extracted feature fragments to be detected in the test set in 1) are input into the trained neural network model for testing and output evaluation indicators. The evaluation indicators used are mainly the performance of the mean average precision (mAP) under a specific tIoU threshold and the average mAP. On the THUMOS14 dataset, the tIoU threshold is {0.3, 0.4, 0.5, 0.6, 0.7}. On the THUMOS14 dataset, the performance comparison of the method of the present invention and the existing method (that is, the Actionformer model only performs ordinary basic training and does not perform the mask training set training process in the present invention) is shown in Table 1. It can be seen from Table 1 that the baseline network using the method of the present invention can be improved.

[0060] Table 1 is a performance comparison table of the method of the present invention and the existing method

[0061] mAP@0.3 mAP@0.4 mAP@0.5 mAP@0.6 mAP@0.7 ave.mAP Existing methods 85.14 80.06 73.85 63.09 48.76 70.18 Method of the present invention 85.97 81.56 74.85 64.83 49.8 71.4

[0062] Finally, it should be noted that the above embodiments and explanations are intended only to illustrate the technical solutions of the present invention and are not intended to limit the present invention. It should be understood by those skilled in the art that modifications or equivalent substitutions to the technical solutions of the present invention may be made without departing from the spirit and scope of the technical solutions disclosed herein, and all such modifications or equivalent substitutions shall be encompassed within the scope of protection of the claims of the present invention.

Claims

1. A video action detection method based on random mask, characterized in that: The steps include: 1) Construct an original training set based on action videos; 2) After performing random masking on the original training set, a masked training set is obtained; 3) constructing a first neural network model, the first neural network model including a connected encoder and decoder; then training the first neural network model using the masked training set until the training is completed to obtain a trained encoder; 4) Using the encoder trained in 3) to build a second neural network model, and then using the original training set to train the second neural network model until the training is completed, thereby obtaining a time series action detection neural network model; 5) The action video to be detected is preprocessed and then input into the temporal action detection neural network model, and the model outputs the corresponding temporal action detection results.

2. The method for video action detection based on random mask according to claim 1, characterized in that: The specific embodiment of 1) is: First, each acquired action video is divided into several image sequences of equal length. Then, a feature extractor is used to extract features from each image sequence to obtain the corresponding original feature sequence. The features in each original feature sequence are then preprocessed to obtain the original training set.

3. The method for video action detection based on random mask according to claim 2, characterized in that: The preprocessing includes downsampling and random cropping operations.

4. The method for video action detection based on random mask according to claim 1, characterized in that: Said 2) is specifically: First, a low-resolution mask template is established. Then, a random mask of a preset ratio is applied to the low-resolution mask template to obtain an initialized low-resolution mask template. The initialized low-resolution mask template is then upsampled to a high-resolution mask template through bilinear interpolation, and the high-resolution mask template is then dimensionally expanded to obtain the final mask template. Finally, the final mask template is applied to the original training set to obtain the mask training set.

5. The method for video action detection based on random mask according to claim 1, characterized in that: In the above 3), the encoder is an encoder that outputs a time series pyramid feature, and the time series pyramid feature is composed of a first time series feature, a second time series feature, a third time series feature and a fourth time series feature; the decoder includes an upsampling convolution layer, a main module and a convolution mapping layer, the first time series feature is input to the first upsampling convolution layer, and the first upsampling convolution layer outputs the first upsampling time series feature; the first upsampling time series feature is used as the input of the first main module, and the first main module outputs the first upsampling main time series feature; the first upsampling main time series feature and the second time series feature are added together to obtain a second fused time series feature; the second fused time series feature is used as the input of the second upsampling convolution layer, and the second upsampling convolution layer outputs the second upsampling time series feature; the second upsampling time series feature is used as the input of the second main module, and the second main module outputs the second upsampling main time series feature; the second upsampling main The third fused temporal feature is obtained by adding the temporal feature to the third temporal feature; the third fused temporal feature is used as the input of the third upsampling convolution layer, and the third upsampling convolution layer outputs the third upsampling temporal feature; the third upsampling temporal feature is used as the input of the third main module, and the third main module outputs the third upsampling main temporal feature; the third upsampling main temporal feature and the fourth temporal feature are added to obtain the fourth fused temporal feature; the fourth fused temporal feature is used as the input of the fourth upsampling convolution layer, and the fourth upsampling convolution layer outputs the fourth upsampling temporal feature; the fourth upsampling temporal feature is used as the input of the fourth main module, and the fourth main module outputs the fourth upsampling main temporal feature; the fourth fused main temporal feature is used as the input of the first convolution mapping layer, the first convolution mapping layer is connected to the second convolution mapping layer, and the second convolution mapping layer outputs the reconstructed feature as the output of the decoder.

6. The method for video action detection based on random mask according to claim 1, characterized in that: The second neural network model includes a trained encoder and a classification head connected to the trained encoder.

7. The method for video action detection based on random mask according to claim 1, characterized in that: In the above 3), the loss function in the training process of the first neural network model is the MSE loss calculated based on the output of the first neural network model and the unmasked part in the mask training set.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the video action detection method based on random mask described in any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the video action detection method based on random mask described in any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the video action detection method based on random mask described in any one of claims 1 to 7 are implemented.