A Sparse Temporal Action Detection Method Based on Dynamic Instance Interaction Heads
By introducing dynamic instance interaction headers and sparse interaction processes in timing behavior detection, combined with set prediction loss and feature embedding modules, the problems of high computing burden, susceptible to human parameters and difficult to detect short instances in the prior art are solved, and more efficient and accurate timing action detection is achieved.
Patent Information
- Application Number
- CN202210579421.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-25
AI Technical Summary
The existing timing behavior detection methods have problems such as high computational burden, susceptible to human parameters, and difficulty in detecting short instances when processing the duration of action instances in videos.
Using a sparse timing action detection method based on dynamic instance interaction heads, through the sparse interaction process of dynamic instance interaction heads, the proposal features only interact with the corresponding fragment features, reducing the calculation of global features, using the Hungarian algorithm's set prediction loss and the feature embedding module randomly embeds feature proposals, reducing the impact of human parameters.
The quality of timing action detection is improved, the calculation burden is reduced, the influence of human parameters is avoided, short action instances can be better detected, and detection performance is improved.
Smart Images

Figure CN114998989B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sparse temporal action detection model for temporal action localization (TAL) based on a dynamic instance interactive head. The dynamic instance interactive head interacts the candidate features and the corresponding features of interest one by one, continuously enhancing the foreground features to obtain the final prediction features. The present invention completely adopts an end-to-end form and has a good effect on temporal action behavior detection. Background Art
[0002] With the rapid development of computer technology and network technology, multimedia information has grown explosively. Among them, video, as an important information carrier, is more and more favored by people, and more information is transmitted through video. However, the processing of a large amount of video information has become a difficult problem, and the traditional manual detection method is very inefficient and boring. With the rise of deep learning technology, processing videos by automatically extracting effective information in videos by a computer can greatly improve work efficiency and save human resources. Therefore, the task of temporal action detection (TAD) has received extensive attention. Temporal action detection is a key technology in video understanding. For a given untrimmed long video, its goal is to locate the time period when the action occurs and predict the category of the action.
[0003] In recent years, after deep learning has achieved good results in various fields, it has been widely used in fields such as object detection, image generation, and video analysis. Compared with traditional machine learning algorithms, deep neural networks extract and fuse features by building network models adapted to different tasks, and then adopt different strategies to solve corresponding problems for different tasks. As the main method of computer vision at present, deep learning has many advantages: 1. It can better represent features according to the adaptive learning of network layers; 2. It has good generalization ability through the learning of big data; 3. It can express features layer by layer, from low-level raw data to high-level semantic information. By combining deep learning with temporal behavior detection, the current temporal behavior detection methods are mainly divided into two types: anchor-based methods and boundary-based methods.
[0004] Anchor-based methods: First, multi-scale anchors are designed on each grid of the feature sequence. Then, the network performance classification and boundary regression are performed on these candidate objects. Since the durations of the ground truth instances vary significantly in different videos, these methods require a large amount of computation when placing dense candidate proposals and may have inaccurate temporal boundaries.
[0005] Boundary-based method: It processes inaccurate boundary problems in a bottom-up manner, where each matching pair of the video sequence is evaluated. The regression process is abandoned, and confidence scores are directly generated for densely distributed proposals. However, this method can only be used for generating temporal action candidate boxes, so an external classifier is required for action classification.
[0006] These two effective methods have been continuously improved and proven effective with excellent performance. However, these two methods still have some limitations. First, they rely to a large extent on dense candidate proposals, which will bring a heavy computational burden. Second, they are vulnerable to artificial parameters, such as anchor design and confidence thresholds. Finally, in the task of temporal action detection (TAD), an important problem is that the duration of action instances in the video ranges from a few seconds to several minutes, and it is difficult for the network to detect short instances. Feature Pyramid Network (FPN) has been widely used in image object detection to solve the problem of large object scale variations. The method based on recent queries, RTD-Net, adopts global attention between query features and globally encoded features, and the quadratic computational complexity prevents it from constructing multi-scale features. Some other works have constructed Temporal Feature Pyramid Network (TFPN) to alleviate the difficulty of temporal boundary localization. Nevertheless, these are all based on constructing TFPN from the features extracted from the last layer of the backbone, which contains high-level representations of video segments. The downsampling operation in the TFPN architecture will further lose the information of short action instances, making it difficult to perform precise temporal boundary regression. These problems and limitations are exactly the existing difficult problems in temporal behavior detection. Summary of the Invention
[0007] The present invention proposes a sparse temporal action detection method based on a dynamic instance interaction head to solve the above three difficult problems.
[0008] 1. Set prediction. Adopt the set prediction loss based on the Hungarian algorithm to optimize the entire network end-to-end from the original two-stream input terminals given RGB frames and optical flow, and output predictions without delay fusion. The target candidate boxes output by the network are the final prediction boxes, without the need for non-maximum suppression post-processing. It bypasses the many-to-one label assignment problem and achieves one-to-one label matching. It solves the limitation that the previous methods rely to a large extent on dense candidate proposals and reduces the computational burden.
[0009] 2. Sparse proposals. Use the feature embedding module to randomly embed N (e.g., 50) feature proposals, getting rid of complex manual design. Through a large number of experiments, it is proved that the final experimental results will not be affected by the initialization of feature proposals. It solves the problem that experimental results are affected by artificial parameters, and researchers no longer have to worry about anchor design and confidence thresholds.
[0010] 3. Sparse Interaction. Based on the sparse interaction process of the dynamic instance interaction head, the proposed features only interact with the corresponding segment features without computing on the global features. Therefore, we can directly use the output of the intermediate layer of the backbone network to construct the hierarchical feature map. Since the intermediate features have a higher temporal resolution, it is possible to retain the details of action instances with a longer duration of change, improving the quality of temporal action detection.
[0011] A sparse temporal action detection method based on a dynamic instance interaction head, comprising the following steps:
[0012] Step (1), data preprocessing, extracting the initial spatio-temporal features of the video data;
[0013] First, extract the image frames and optical flow of the video data; secondly, extract the corresponding features based on the extracted image frames and optical flow respectively; then, stack the extracted features in the temporal dimension and use a sliding window method to extract video segments of equal length.
[0014] Step (2), constructing a dynamic instance interactive head (Dynamic instance interactive head) network model based on a temporal feature pyramid network structure (Temporal Feature Pyramid Networks, TFPN);
[0015] The dynamic instance interactive head network model based on the temporal feature pyramid network structure includes a temporal feature pyramid and a dynamic instance interactive head.
[0016] The temporal feature pyramid consists of two parts: bottom-up and top-down. The bottom-up part extracts features through a traditional convolutional network, and the top-down path is used for feature fusion to construct a higher resolution on the low-resolution feature layer with rich semantics, and a lateral connection method is adopted to solve the problem of target offset caused by continuous upsampling and downsampling. The feature pyramid obtains a total of five layers of outputs: P1, P2, P3, P4, and P5. In order to obtain more video information, four feature layers of the pyramid P1 - P4 are extracted to predict the key points of behavior at multiple scales.
[0017] The dynamic instance interactive head receives the multi-level features generated by the temporal feature pyramid network and then predicts the time period and action category of the action instance. The input of the dynamic instance interactive head includes three contents: one is the multi-scale features output by the temporal pyramid network; the second is the learnable proposal box; the third is the learnable proposal feature. The proposal box is a two-dimensional parameter representing the normalized center position and duration of the time period. The proposal box can be set to any size and is randomly placed on the feature sequence during initialization to avoid complex candidate proposal design. The proposal feature encodes rich instance information for each proposal candidate.
[0018] Step (3), model training;
[0019] The candidate boxes of unified size pass through the fully connected layer to obtain feature vectors of a fixed size, and N unordered sets are output. Each set element includes classification and localization information. Using the cascade idea, the output candidate boxes are adjusted, and the output information of each cascade stage is trained using the optimal bipartite matching and classification regression loss until the entire network model converges.
[0020] Step (4), generating localization detection results;
[0021] According to the optimal bipartite matching method, one-to-one label matching is performed on the feature vectors output by the model. The candidate boxes output by the model training are the final prediction boxes.
[0022] The data preprocessing described in step (1), extracting the initial spatio-temporal features of the video data, is specifically as follows:
[0023] For each input video v in the video dataset V n , first extract image frames at 30 FPS, and at the same time use the TVL-1 algorithm to extract the optical flow of the video. Feature extraction is performed on the extracted images and optical flow, and the I3D model pre-trained based on the Kinetics dataset is used to extract the features corresponding to the images and optical flow respectively and where N represents different videos with different temporal lengths, and 1024 represents the feature dimension output after each video segment is extracted by the pre-trained I3D model. In order to integrate the appearance features and motion features of the input video, the image feature F rgb and the optical flow feature F flow are stacked in the temporal dimension, and the initial spatio-temporal features are obtained Then, a sliding window slides on the temporal length N with an overlap rate of 50%, and finally the spatio-temporal features of the window are obtained where T = 256.
[0024] The dynamic instance interaction head network model based on the temporal feature pyramid structure described in step (2) is specifically as follows:
[0025] 2-1. Temporal feature pyramid;
[0026] The traditional bottom-up path in the pyramid structure is essentially a feed-forward calculation of a downsampling convolutional neural network. The graph attention convolution (GraphAttentionNetwork, GAT) plus the max pooling operation with a stride of 2 is used to replace the original simple one-dimensional convolution. The specific formula is:
[0027] F high= Maxpooling(GAT(F cur )) (1)
[0028] where F high represents the output of the high-level feature map after the current graph convolution, and F cur represents the input features of the current layer. Then comes the top-down path, which is essentially to increase the resolution of the feature map with high-level semantic information. Upsample the feature map with a large receptive field at the top, and the stride is the same as the max-pooling operation, both being 2. Linear interpolation is used during upsampling. After upsampling, it is horizontally concatenated with the feature map of the same size during the bottom-up convolution. When fusing, the form of adding corresponding elements is adopted. The specific formula can be expressed as:
[0029] F low = Interpolate(conv(F cur )) (2)
[0030] where conv is a 1×3 convolution used to mitigate the aliasing effect of upsampling. The top-down path transmits better semantic information, and the bottom-up path transmits better localization information. By fusing them through horizontal concatenation, features with better localization information and better semantic information can be obtained. Outputs of different layers can obtain features sensitive to different scales and identify behaviors at different time scales.
[0031] 2-2. Dynamic Instance Interaction Head;
[0032] The input of the dynamic instance interaction head includes three parts: one is F 1 , F 2 , F 3 , F 4 in the multi-scale feature layers output by the temporal pyramid network, where Fea = 2048 is the feature dimension; the second is the learnable proposal box; the third is the learnable proposal feature. The output of the final dynamic instance interaction head includes two parts: one is the class prediction, and the other is the boundary prediction.
[0033] The learnable proposal boxes mentioned above are finally used as candidate proposals. These proposal boxes are initialized with two-dimensional parameters in the range of 0-1, representing the normalized center coordinates and the length of the action duration. During training, the parameters of the proposal boxes will be updated using the Back-Propagation (BP) algorithm. The number of candidate proposals is greater than the maximum number of ground-truth action instances in all video clips of the video dataset. Due to its learnability, the influence of initialization is small, making the proposal boxes more flexible. Conceptually, the learnable proposal boxes are statistical information about potential action locations in the training set, an initial guess of the regions in the video that are most likely to contain actions, regardless of the input.
[0034] Although the two-dimensional proposal box is a simple and clear representation of the action range, it only provides a rough localization of the action duration and loses a lot of detailed information, such as the category of the action and the relevant information about the actor. Therefore, proposal features are introduced, which are high-dimensional latent vectors that will encode rich action instances. The number of proposal features is the same as the number of proposal boxes.
[0035] The initialized proposal boxes are mapped to the unit time range of 0-1. Before being input into the dynamic instance interaction head, their weights are initialized and scaled to sizes of 0-256 frames, 0-128 frames, 0-64 frames, and 0-32 frames according to the scale sizes output by the temporal pyramid network. The SOI features R soi (T×l, l = 16) are extracted from the temporal feature pyramid using the proposal boxes through the SOI-Align module. Each SOI feature will be used in its own dedicated head for action classification and localization, and each head is conditioned on a specific proposal feature. PK performs self-attention to generate the convolutional kernel parameter PK conv , and then the generated convolutional kernel parameter PK conv interacts sparsely with R soi to filter out invalid units and output the final prediction feature F fin . The specific interaction process is shown in the following formula:
[0036] F fin = norm 3 (drop 3 (forw(norm 2 (drop 2 (inter(R soi , norm 1 (drop 1 (PK)+PK conv )))+PK)))+PK) (3)
[0037] where norm 1 and norm 2, norm 3 is the fully connected layer in the neural network, drop 1 , drop 2 , drop 3 is gradient clipping, forw is the feed-forward neural network, and the specific content is as shown in the formula:
[0038] forw(x) = Linear 2 (relu(drop(Linear 1 (x)))) (4)
[0039] Linear 1 , Linear 2 are fully connected networks, and relu is the activation function. The sparse interaction part in formula (3) can be expressed as the following formula:
[0040] inter(x,y) = relu(norm(bmm(x,Linear(y)))) (5)
[0041] bmm performs matrix multiplication on the two input parameters. Therefore, the interaction process can be regarded as the SOI feature passing through two one-dimensional convolutional layers, which is beneficial for the model to make full use of the high-resolution features of the intermediate layer. Finally, two parallel branches, namely the classification branch and the regression branch, are established on the dynamic instance interaction head to obtain the final classification score and boundary regression prediction of the action instance. The classification branch is a linear layer with Sigmoid activation, which is used to predict the probability of each action category. The regression branch consists of a three-layer feed-forward network, which has a RELU activation function for time boundary regression. Stacking all the dynamic instance interaction heads together, the predictions of each head are obtained. The predicted proposal boxes and proposal features at each stage will be used as the initial proposal boxes and proposal features of the next stage for continuous improvement.
[0042] Step (3) Model training is as follows:
[0043] The prediction finally generated by the dynamic instance interaction head is called the action instance set ψ p , which contains N instances, and the value of N is greater than the number M of real action instances in the video dataset V. All real action instances constitute the real target set ψ g , and by padding categories, the real target set ψ g is expanded to N. The set prediction loss is adopted on these two fixed-size sets, and the set-based prediction loss generates the best bipartite matching between the predicted value and the real value. The definition of the matching cost is as follows:
[0044] L = λ cls ·L cls +λ L1·L L1 +λ iou ·L iou (6)
[0045] L cls is the focal loss between the true class label and the predicted class, L L1 and L iou are the l1 loss and IoU loss between the center coordinates and action duration of the predicted bounding box and the true bounding box. λ cls 、λ L1 and λ iou are the weight coefficients of each type of loss respectively. Except for being executed only on matching pairs, the training loss is the same as the matching loss, and the final loss is the sum of all pairs normalized by the number of objects in the training batch. The specific formula of L cls is shown as follows:
[0046] L cls (p t )=-α t (1 - p t ) γ log(p t ) (7)
[0047] where p t represents the probability of being predicted as a key point, α t represents the corresponding weight of positive and negative samples, and the role of γ is to reduce the loss of simple samples and force the model to pay more attention to difficult-to-select samples. Using one-to-one matching based on set loss bypasses the many-to-one matching problem, which is an end-to-end attempt in video detection.
[0048] Generating the localization detection results described in step (4) is specifically as follows:
[0049] According to the predicted features and temporal prediction bounding boxes obtained in step (2), directly perform the best bipartite matching using the matching loss in step (3) to obtain the predicted labels. The predicted labels and temporal prediction bounding boxes are the final predicted classes and action boundaries, and the final performance is calculated using the mean average precision (mAP).
[0050] Furthermore, the number N of the candidate proposals is 50.
[0051] The beneficial effects of the present invention are as follows:
[0052] The present invention proposes a sparse temporal action detection method based on a dynamic instance interaction head. Although the existing temporal behavior detection methods have achieved good results, a large number of manually designed anchor boxes not only affect the computational time complexity, but also the prediction results are affected by humans. The present invention uses a query-based method to initialize N proposal features and proposal boxes, solving the complexity problem of anchor boxes. The present invention also introduces a dynamic instance interaction head module based on a temporal feature pyramid. Using the temporal feature pyramid can better predict behaviors of different scales, solving the influence on the experimental results caused by different time spans of each behavior; at the same time, different from the query method, the dynamic instance interaction head does not require the query features of each group to interact with the global features to better learn the features. The dynamic instance interaction head module only sparsely interacts the proposal features with the local features, and can well learn valuable information, greatly reducing the computational amount. Finally, the present invention uses the best bipartite matching based on the set prediction loss, which can perform label matching one by one, and finally only outputs N candidate boxes equal to the number of initial proposal boxes. Without using non-maximum suppression post-processing before computing performance, it can be directly output as a prediction box. The method of the present invention has achieved a large performance improvement compared with the traditional temporal behavior detection method. Description of the Drawings
[0053] Figure 1 It is a complete flowchart of an embodiment of the present invention. Detailed Embodiment
[0054] The technical solution of the present invention will be further specifically described below in conjunction with the drawings and embodiments.
[0055] As Figure 1 shown, the present invention provides a sparse temporal action detection method based on a dynamic instance interaction head, including the following steps:
[0056] Step (1), data preprocessing, extracting the initial spatio-temporal features of video data;
[0057] Preprocessing of the video dataset V: For each input video v in the video dataset V n, first extract image frames at 30 FPS, and at the same time use the TVL-1 algorithm to extract the optical flow of the video. For the extracted images and optical flow, use the I3D model pre-trained on the Kinetics dataset to extract the corresponding features of the images and optical flow respectively, and then stack these two features in the temporal dimension to integrate the appearance features and motion features of the input video, taking into account both temporal information while ensuring spatial information, and obtaining the final initial spatio-temporal features. Since the lengths of each video are different, in order to facilitate the unified input of features into the network model, a sliding window form is adopted, and video segments of the same length are slid out with a certain overlap rate on the basis that the window size can contain almost all instances.
[0058] Here, the THUMOS’14 dataset is used as the training and test data.
[0059] For each input video v in the THUMOS’14 dataset n , first extract image frames at 30 FPS, and then use the TVL-1 algorithm in the OpenCV library to extract the optical flow of the video. For the extracted images and optical flow, to unify the image size, while maintaining the aspect ratio, scale the minimum side of each image to 256 pixels, and at the same time center-crop to 224×224 pixels. Uniformly sample each video into 750 video segments, and then use the I3D model pre-trained on the Kinetics dataset to extract the corresponding features of the images and optical flow respectively, and then stack these two features in the temporal dimension to integrate the appearance features and motion features of the input video, and obtain the final initial features Since the temporal lengths of each video are different, for the convenience of feature extraction, use a window size of T = 256 and a stride of stride = 128 to slide and extract feature segments of the same size. Obtain the initial feature size
[0060] Step (2): Construct a dynamic instance interaction head model based on the temporal feature pyramid structure;
[0061] Due to the limitation of GPU memory, first reduce the dimension of the original 2048-dimensional feature through a common convolution to obtain a convolutional feature dimension of 256 dimensions.
[0062] 2-1. The pyramid structure uses a total of 5 layers. For the initial feature with a temporal length of 256, first obtain different temporal lengths T n , T n = 128, 64, 32, 16 in the bottom-up direction. Then, in the bottom-up direction and with horizontal connections, return from T n = 16 back to T n= 32, 64, 128, 256, a total of 5 scales of feature blocks are obtained. Finally, the feature layers F1, F2, F3, and F4 among them are used for subsequent feature extraction.
[0063] 2-2. Use SOI-Align to pool the proposal boxes on the multi-scale feature layers. The boundary length of the proposal boxes is uniformly pooled to the size of 16 frames, and the region of interest feature R is obtained. soi (256×16), R soi And the proposal feature PK conv (N×256) performs sparse interaction, that is, matrix multiplication operation, to obtain candidate features of size (N×16). We use the idea of iteration. A total of 6 instance interaction heads are used, and the number of dynamic instance interaction heads in each iteration is N, to ensure that there are different interaction heads for each different proposal box and proposal feature, and dynamically learn the proposal feature and boundary feature. In each layer of iteration, the result output by the dynamic instance interaction head is used to calculate the regression prediction using a perception method with a relu activation function and three hidden layers, and a linear projection layer is used to calculate the classification prediction. The predicted features and predicted boxes output in the previous iteration process will be used as the proposal features and proposal boxes for initialization input in the next iteration process. The results output by the 6 iteration processes will all be saved, but we finally only take the results of the last iteration for the classification of labels and the prediction of bounding boxes.
[0064] Step (3), model training;
[0065] 3-1. The action instance set ψ generated by the dynamic instance interaction head p Contains N instances, and the value of N is greater than the number of real action instances in the dataset. The larger the value of N, the higher the accuracy of the experiment. However, considering the performance issue of the experiment, in the present invention, the value of N is uniformly taken as 50. By padding The category expands the real target set ψ g To N, and the set prediction loss is adopted on these two fixed-size sets. The set-based prediction loss generates the best bipartite matching between the predicted value and the true value.
[0066] 3-2. According to the focal loss formula:
[0067] FL(p t ) = -α t (1 - p t ) γ log(p t )
[0068] α tSet it to 0.75 at the positive sample and 0.25 at the negative sample, and set γ to 2. Select 5 positive samples and 10 negative samples from each layer of each video segment for training. If there are not enough positive samples, fill them with negative samples.
[0069] 3-3. The formula of the loss function is the same as that of the matching function. The specific formula is as follows:
[0070] L = λ cls ·L cls +λ L1 ·L L1 +λ iou ·L iou
[0071] Finally, the total training loss is normalized according to the number of objects in the training batch. Input it into the network using backpropagation until the loss converges.
[0072] Step (4) generates the localization detection results, which are specifically as follows:
[0073] According to the predicted features and temporal prediction boxes obtained in step (2), directly perform bipartite matching using the matching loss in step (3) to obtain the predicted labels. The predicted labels and temporal prediction boxes are the final predicted categories and action boundaries. On THUMOS14, use the tIoU thresholds [0.3:0.1:0.7] and the mean average precision (mAP) to calculate the final performance.
[0074] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, they can also make several substitutions or modifications to these described embodiments, and these substitution or modification methods should all be regarded as belonging to the protection scope of the present invention.
[0075] The parts not detailed in the present invention belong to the well-known technologies in the art.
Claims
1. A sparse temporal action detection method based on a dynamic instance interaction head, characterized in that, it includes the following steps: Step (1), data preprocessing, extracting the initial spatio-temporal features of video data; First, extract the image frames and optical flow of the video data; secondly, extract the corresponding features based on the extracted image frames and optical flow respectively; then, stack the extracted features in the temporal dimension and use a sliding window method to take out video segments of equal length; Step (2), constructing a dynamic instance interaction head network model based on a temporal feature pyramid structure; The dynamic instance interaction head network model based on the temporal feature pyramid structure includes a temporal feature pyramid and a dynamic instance interaction head; The temporal feature pyramid consists of two parts: bottom-up and top-down. The bottom-up part is to extract features through a traditional convolutional network, and the top-down path is for feature fusion. Higher resolutions are constructed on the low-resolution feature layers with rich semantics, and lateral connections are used to solve the problem of target offset caused by continuous upsampling and downsampling; the feature pyramid obtains a total of five layers of outputs, namely P1, P2, P3, P4, and P5. In order to obtain more video information, four feature layers of the pyramid P1 - P4 are extracted to predict the key points of actions at multiple scales; The dynamic instance interaction head receives the multi-level features generated by the temporal feature pyramid network and then predicts the time period and action category of the action instance; the input of the dynamic instance interaction head includes three contents: one is the multi-scale features output by the temporal pyramid network; the second is the learnable proposal box; the third is the learnable proposal feature; The proposal box is a two-dimensional parameter representing the normalized central position and duration of the time period; the proposal box can be set to any size and is randomly placed on the feature sequence during initialization to avoid complex candidate proposal design; the proposal feature encodes rich instance information for each proposal candidate; Step (3), model training; Candidate boxes of the same size pass through a fully connected layer to obtain fixed-size feature vectors, and N unordered sets are output. Each set element includes classification and localization information; using the cascade idea, the output candidate boxes are adjusted, and the output information of each cascade stage is trained using the optimal bipartite matching and classification regression loss until the entire network model converges; Step (4), generating localization detection results; According to the optimal bipartite matching method, one-to-one label matching is performed on the feature vectors output by the model; the candidate boxes output by model training are the final prediction boxes.
2. A sparse temporal action detection method based on a dynamic instance interaction head according to claim 1, characterized in that, the data preprocessing in step (1) for extracting the initial spatio-temporal features of video data is as follows: For each input video v in the video dataset V n , first extract image frames at 30 FPS, and at the same time use the TVL-1 algorithm to extract the optical flow of the video; perform feature extraction on the extracted images and optical flow, and use the I3D model pre-trained on the Kinetics dataset to extract the features corresponding to the images and optical flow respectively and where N represents different videos with different temporal lengths, and 1024 represents the feature dimension output after each video segment is extracted by the pre-trained I3D model; in order to integrate the appearance features and motion features of the input video, the image feature F rgb and the optical flow feature F flow are stacked in the temporal dimension to obtain the initial spatio-temporal features Then, use a sliding window to slide on the temporal length N with an overlap rate of 50%, and finally obtain the spatio-temporal features of the window where T = 256.
3. A sparse temporal action detection method based on a dynamic instance interaction head according to claim 2, characterized in that, the dynamic instance interaction head network model based on the temporal feature pyramid structure in step (2) is as follows: 2-1. Temporal feature pyramid; The traditional bottom-up path in the pyramid structure is essentially the feed-forward calculation of a downsampling convolutional neural network. The graph attention convolution plus the max pooling operation with a stride of 2 is used to replace the original simple one-dimensional convolution. The specific formula is as follows: F high = Maxpooling(GAT(F cur ))(1) Among them, F high represents the output of the high-level feature map after the current graph convolution, and F cur represents the input feature of the current layer; then there is a top-down path, which is essentially to increase the resolution of the feature map with high-level semantic information; the feature map with a large receptive field at the top is upsampled, and the stride is the same as that of the max pooling operation, both being 2, and linear interpolation is used during upsampling; after upsampling, it is horizontally connected to the feature map with the same size during bottom-up convolution, and element-wise addition is used during fusion. The specific formula can be expressed as: F low = Interpolate(conv(F cur ))(2) where conv is a 1×3 convolution used to mitigate the aliasing effect in the upsampling; 2-2. Dynamic instance interaction head; The input of the dynamic instance interaction head includes three contents: one is F in the multi-scale feature layers output by the temporal pyramid network 1 , F 2 , F 3 , F 4 , where Fea = 2048 is the feature dimension; Second, learnable proposal boxes; third, learnable proposal features; Finally, the output of the dynamic instance interaction head includes two parts: one is class prediction, and the other is boundary prediction; The learnable proposal boxes mentioned above are finally used as candidate proposals; these proposal boxes are initialized as two-dimensional parameters between 0 and 1, representing the normalized center coordinates and action duration lengths; during training, the parameters of the proposal boxes will be updated using the backpropagation algorithm; the number of candidate proposals is greater than the maximum number of ground-truth action instances in all video clips in the video dataset; Although the two-dimensional proposal box is a simple and clear representation of the action range, it only provides a rough localization of the action duration and loses a lot of detailed information; therefore, proposal features are introduced, which is a high-dimensional latent vector that will encode rich action instances; The number of proposal features is the same as the number of proposal boxes; The initialized proposal box is mapped to the unit time of 0-1. Before being input into the dynamic instance interaction head, its initial weights are given and scaled to the sizes of 0-256 frames, 0-128 frames, 0-64 frames, and 0-32 frames respectively according to the scale sizes output by the temporal pyramid network; the SOI-Align module uses the proposal box to extract the SOI feature R from the temporal feature pyramid soi , and each SOI feature will be used in its own dedicated head for action classification and localization, and each head is conditioned on a specific proposal feature; PK performs self-attention to generate the convolutional kernel parameter PK conv , and then the generated convolutional kernel parameter PK conv is sparsely interacted with R soi to filter out invalid units and output the final predicted feature F fin ; the specific interaction process is shown in the following formula: F fin = norm 3 (drop 3 (forw(norm 2 (drop 2 (inter(R soi ,norm 1 (drop 1 (PK)+PK conv )))+PK)))+PK) (3) Among them, norm 1 , norm 2 , norm 3 are fully connected layers in the neural network, drop 1 , drop 2 , drop 3 are gradient clipping, and forw is a feedforward neural network. The specific content is shown in the formula: forw(x) = Linear 2 (relu(drop(Linear 1 (x))))(4) Linear 1 and Linear 2 are fully connected networks, and relu is the activation function; the sparse interaction part in formula (3) can be expressed as the following formula: inter(x,y)=relu(norm(bmm(x,Linear(y))))(5) bmm performs matrix multiplication on the two input parameters; therefore, the interaction process can be regarded as the SOI feature passing through two one-dimensional convolutional layers, which is beneficial for the model to fully utilize the high-resolution features of the intermediate layer; finally, two parallel branches, namely the classification branch and the regression branch, are established on the dynamic instance interaction head to obtain the final classification score and boundary regression prediction of the action instance; The classification branch is a linear layer with Sigmoid activation, used to predict the probability of each action category; the regression branch consists of a three-layer feed-forward network with a RELU activation function for temporal boundary regression; stack all the dynamic instance interaction heads to obtain the predictions of each head; the predicted proposal boxes and proposal features at each stage will be used as the initial proposal boxes and proposal features for the next stage to continuously improve.
4. A sparse temporal action detection method based on a dynamic instance interaction head according to claim 3, characterized in that step (3) model training is as follows: The prediction finally generated by the dynamic instance interaction head is called the action instance set ψ p , which contains N instances, and the value of N is greater than the number M of real action instances in the video dataset V; all real action instances constitute the real target set ψ g , by padding categories, the real target set ψ g is expanded to N, and the set prediction loss is adopted on these two fixed-size sets. The set-based prediction loss produces the best bipartite matching between the predicted value and the real value, and the matching cost is defined as follows: L = λ cls ·L cls + λ L1 ·L L1 + λ iou ·L iou (6) L cls is the focal loss between the true class label and the predicted class, L L1 and L iou are the l1 loss and the IoU loss between the center coordinates and the action duration of the predicted bounding box and the true bounding box; λ cls 、λ L1 and λ iou are the weight coefficients of each loss respectively; the matching loss is the same as the training loss except that it is only performed on the matching pairs, and the final loss is the sum of all pairs normalized by the number of objects in the training batch; the specific formula of L cls is shown as follows: L cls (p t )=-α t (1 - p t ) γ log(p t )(7) where p t represents the probability predicted as a key point, and α t represents the corresponding weights of positive and negative samples. The role of γ is to reduce the loss of simple samples and force the model to pay more attention to difficult-to-select samples.
5. A sparse temporal action detection method based on a dynamic instance interaction head according to claim 4, characterized in that step (4) generating the localization detection result is as follows: According to the predicted features and temporal prediction boxes obtained in step (2), directly perform the best bipartite matching using the matching loss in step (3) to obtain the predicted labels. The predicted labels and temporal prediction boxes are the final predicted categories and action boundaries, and the average precision is used to calculate the final performance.
6. A sparse temporal action detection method based on a dynamic instance interaction head according to claim 3, characterized in that the number N of the candidate proposals is 50.
Citation Information
Patent Citations
Behavior detection method and device based on graph network
CN112347964A
Graph attention network time sequence action positioning method based on pyramid structure
CN113255443A