A Human Action Recognition Method, System and Device Based on a Split Attention Network

Through the human body movement recognition method based on the shunt attention network, the human body movement appearance characteristics of the channel domain and the time-space domain are extracted, and the time difference timing characteristics are extracted in combination with BiLSTM and the time-difference self-attention module, which solves the problems of large calculation overhead and insufficient feature information in the prior art, and achieves the effect of significantly improving the accuracy of the action recognition without increasing the calculation amount.

CN114627555BActive Publication Date: 2025-06-03HUAIYIN INSTITUTE OF TECHNOLOGY +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210252373.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-06-03
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

The existing human body movement recognition technology has major problems in computing overhead and training time, and the signal-to-noise ratio of the skeleton sequence is low and the characteristic information is insufficient, making it difficult to effectively apply on machinery and equipment.

Method used

The human body movement recognition method based on the shunt attention network is adopted to extract the appearance characteristics of the human body movements in the channel domain and the time and space domain through the shunt attention network, and combine BiLSTM and the time difference self-attention module to extract the time difference timing characteristics, and ultimately improve the accuracy of action recognition.

Benefits of technology

Without increasing the calculation amount, the accuracy of human body movement recognition is significantly improved, and the problems of large computing overhead and insufficient feature information in the prior art are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627555B_ABST
    Figure CN114627555B_ABST
Patent Text Reader

Abstract

The present invention discloses a human action recognition method, system and device based on a split attention network. The method includes: S1, parsing the videos in the human action recognition dataset into a sequence of human action frames; S2, preprocessing the parsed sequence of human action frames and sampling to obtain a training dataset and a test dataset; S3, inputting the T-frame sequence after data preprocessing into the split attention network to extract the human action appearance features in the channel domain and the spatio-temporal domain; S4, inputting the human action appearance features into a temporal sequence network model to extract the time difference temporal features of the human action; S5, training a human action feature model based on the time difference temporal features, and inputting the test dataset into the trained human action feature model to obtain the final classification result of the human action. The split attention network and the time difference dot product self-attention module proposed by the present invention can further improve the accuracy of human action recognition without increasing the computational amount.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to human action recognition technology in the field of computer vision, and specifically to a human action recognition method, system and device based on a split attention network. Background Art

[0002] In recent years, computer vision technology based on deep neural networks has been widely used in human action recognition and analysis tasks, but there are still huge problems and challenges for practical applications.

[0003] First, 3D action recognition methods based on RGB features have achieved good recognition effects by stacking 3D convolutions, such as the 3D network proposed by Tran et al.: Learning Spatiotemporal Features with 3D Convolutional Networks [C] Proceedings of the IEEE International Conference on Computer Vision. 2015: 4489-4497. The P3D network proposed by Qiu et al.: Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks [C]. Proceedings of the IEEE International Conference on Computer Vision. 2017: 5533-5541. However, these methods have problems such as many network parameters, large computational overhead, and long training time, and cannot be well applied to machine devices. Due to the characteristics of 3D convolution, this problem has not been well solved.

[0004] Second, action recognition methods based on skeleton features construct three-dimensional human skeletons by using methods such as graph convolution and encoder-decoder to enhance the action recognition effect. Such as the Spatial Temporal Graph Convolutional Networks for Skeleton-based Action Recognition [C] proposed by Yan et al. Thirty-second AAAI Conference on Artificial Intelligence. 2018. However, it still has great challenges for the problems of low signal-to-noise ratio and insufficient feature information in the skeleton sequence. Summary of the Invention

[0005] Objective of the Invention: An objective of the present invention is to provide a human action recognition method based on a split attention network, which can further improve the accuracy of human action recognition without increasing the computational complexity.

[0006] Another objective of the present invention is to provide a human action recognition system and device based on a split attention network.

[0007] Technical Solution: A human action recognition method based on a split attention network of the present invention includes the following steps:

[0008] S1. Parse the videos in the human action recognition dataset into a sequence of human action frames, wherein the videos in the human action recognition dataset are labeled human action videos;

[0009] S2. Perform random flipping and transformation operations on the parsed sequence of human action frames for data augmentation to obtain a preprocessed sequence of human action frames, and sample to obtain a training dataset and a test dataset;

[0010] S3. Input the training dataset in step S2 into the split attention network to extract the human action appearance features in the channel domain and the spatio-temporal domain;

[0011] S4. Input the human action appearance features obtained in step S3 into a temporal network model combined with a BiLSTM recurrent neural network and a time-difference dot product self-attention module to extract the time-difference temporal features of human actions;

[0012] S5. Train a human action feature model based on the time-difference temporal features obtained in step S4, and input the test dataset into the trained human action feature model to obtain the final classification result of human actions.

[0013] Further, the sampling method of the training dataset in step S2 is: randomly select a sampling interval and a starting frame from the preprocessed sequence of human action frames as the training dataset;

[0014] The sampling method of the test dataset is: uniformly sample from the first frame of the preprocessed sequence of human action frames as the test dataset.

[0015] Further, the split attention network in step S3 includes a backbone network module, and the backbone network module includes 5 sequentially connected residual blocks, and each residual block includes a 7×7 convolutional layer, a channel domain attention module, a spatio-temporal domain attention module, a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer;

[0016] The 7×7 convolutional layer of the first residual block in the backbone network module extracts the underlying human action features R in the training dataset and outputs them to the channel domain attention module and the spatio-temporal domain attention module respectively;

[0017] The channel domain attention module uses spatial adaptive average pooling to sequentially infer a one-dimensional channel domain attention mask M ca ; and multiplies and adds the one-dimensional channel domain attention mask M ca and the underlying human action features R to obtain the channel domain attention features R ca ;

[0018] The spatio-temporal domain attention module uses channel average and max pooling to sequentially infer a one-dimensional spatio-temporal domain attention mask M' sta ; and multiplies and adds the one-dimensional spatio-temporal domain attention mask M' sta and the underlying human action features R to obtain the spatio-temporal domain attention features R sta ;

[0019] For the obtained channel domain attention features R ca and spatio-temporal domain attention features R sta are added together to obtain the channel domain and spatio-temporal domain human action mixed features;

[0020] The mixed features of human actions are successively passed through a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer and then input into the second residual block; then successively passed through the third residual block, the fourth residual block, and the fifth residual block to output the appearance features of human actions.

[0021] Furthermore, the channel domain attention module enhances the influence of channel domain features by compressing spatial features and further enhances the expressive ability of channel features through local cross-channel interaction, specifically including the following steps:

[0022] S301. For the underlying human action features R extracted by the 7×7 convolutional layer of the first residual block in the backbone network module, the channel domain attention module uses spatial adaptive average pooling to perform spatial feature encoding on all channels, compresses the spatial features into a global feature, and compresses the global spatial information into the channel descriptor to obtain the channel domain feature information; the spatial adaptive average pooling formula used is:

[0023]

[0024] where the input feature of the channel domain attention module is the underlying human action features R∈R NT×C×H×W , NT is the number of underlying human action feature maps, C is the number of channels of each underlying human action feature map, H is the height of each underlying human action feature map, W is the width of each underlying human action feature map, and the output feature F of the channel domain attention module ∈RNT×C×1×1 ;

[0025] S302. Use the 2D convolutional layer k 1 Compress the number of channels of the output feature F of the channel attention module at a ratio r to further reduce the number of parameters. The formula used is:

[0026] F r = k 1 * F

[0027] where k 1 is a 1×1 2D convolutional layer, and F r is the compressed channel feature. Reshape F r into

[0028] S303. Input the compressed channel feature F' reshaped in step S302 r into the 1D convolutional layer k 2 for cross-channel interaction. The formula used is:

[0029] F temp = k2 * F' r

[0030] where k 2 is a 1×1 1D convolutional layer, and F temp is the interaction channel feature. Reshape F temp into

[0031] S304. Use the 2D convolutional layer k 3 to decompress the interaction channel feature F' reshaped in step S303 temp and feed it into the Sigmoid activation function. The formulas used are respectively:

[0032] F c = k 3 * F' temp

[0033] M ca = δ(F c )

[0034] where F c is the channel mask, F c ∈ R NT×C×1×1 , M ca is the one-dimensional channel domain attention mask, M ca ∈ R NT ×C×1×1 , and δ(·) is the Sigmoid activation function. Finally, the obtained channel domain attention feature Rca is:

[0035] R ca = R + R⊙M ca .

[0036] Furthermore, the spatio-temporal domain attention module enhances the influence of spatial features by compressing channel features and further enhances the temporal expression ability of spatial features through a 3D convolutional layer. The specific steps are as follows:

[0037] S311. Reshape the underlying feature R∈R extracted by the 7×7 convolutional layer of the first residual block in the backbone network module into R′∈R NT×C×H×W ; N×T×C×H×W ;

[0038] S312. The spatio-temporal domain attention module uses channel average pooling and channel max pooling to perform channel feature encoding over all spaces, compressing the channel features into global features F avg and F max , respectively, and compresses the global channel information into the spatial descriptor to obtain spatio-temporal feature information. The average pooling and max pooling formulas used are as follows:

[0039]

[0040] F max = max(R′[:,:,i,:,:])

[0041] where F avg is the spatial attention average pooling feature, F avg ∈R N×T×1×H×W , F max is the spatial attention max pooling feature, F max ∈R N×T×1×H×W ; C is the number of channels of the underlying features of human actions;

[0042] Reshape the spatial attention average pooling feature F avg into F′ avg ∈R N×1×T×H×W ;

[0043] Reshape the spatial attention max pooling feature F max into F′ max ∈R N×1×T×H×W ;

[0044] S313. Use the 3D convolutional layer k 4 to perform spatio-temporal feature extraction on the input reshaped spatial attention average pooling feature F′ avg and the reshaped spatial attention max pooling feature F′ max respectively. The specific formulas are as follows:

[0045] F″ avg = k 4* F′ avg

[0046] F″ max = k 4* F′ max

[0047] where k 4 is a 1×1 3D convolutional layer, F″ avg is the spatio-temporal attention average pooling feature, F″ avg ∈R N ×1×T×H×W and F″ max is the spatio-temporal attention max pooling feature, F″ max ∈R N×1×T×H×W ,

[0048] Reshape the spatio-temporal attention average pooling feature F″ avg into

[0049] Reshape the spatio-temporal attention max pooling feature F″ max into

[0050] S314. Fuse the reshaped spatio-temporal attention average pooling feature and the reshaped spatio-temporal attention max pooling feature and feed them into the Sigmoid activation function; the formula used is:

[0051]

[0052] where M sta is a one-dimensional spatio-temporal domain attention mask;

[0053] S315. Reshape the one-dimensional spatio-temporal domain attention mask M sta ∈R N×T×1×H×W into a 2D convolutional feature M′ sta ∈R NT×1×H×W , and perform the final feature weighting; the formula used is:

[0054] R sta = R + R ⊙ M′ sta

[0055] where R sta is the spatio-temporal domain attention feature;

[0056] S316. The obtained channel domain attention feature R ca and the spatio-temporal domain attention feature R staPerform addition to obtain the human action hybrid feature X in the channel domain and spatio-temporal domain;

[0057] S317. Input the hybrid feature X of the human action into the second residual block after passing it through a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer in sequence; then output the appearance feature of the human action after passing through the third residual block, the fourth residual block, and the fifth residual block in sequence.

[0058] Further, step S4 includes the following steps:

[0059] S41. The BiLSTM recurrent neural network learns the temporal feature representation between the frames of the human action appearance feature sequence from both forward and backward directions, concatenates the features learned by the forward LSTM vector and the features learned by the backward LSTM vector to obtain the human action hybrid temporal feature, and then inputs it into the temporal dot product self-attention module;

[0060] S42. The temporal dot product self-attention module obtains three feature matrices Q, K, and V by passing the input human action hybrid temporal feature through three identical linear mappings respectively;

[0061] S43. Calculate the dot product similarity between the feature matrix Q and the transpose of the feature matrix K to obtain the weight matrix;

[0062] S44. Normalize the obtained weight matrix using the Softmax function;

[0063] S45. Perform temporal difference feature weighting on the normalized weight matrix in step S44, so that the T-frame sequences in the input training dataset are assigned different weights due to the different input time sequences;

[0064] S46. Calculate the dot product of the weight matrix obtained in step S45 and the feature matrix V and perform weighted summation to obtain the final temporal difference temporal feature.

[0065] A human action recognition system based on a shunt attention network of the present invention includes:

[0066] A video parsing module for parsing the videos in the human action recognition dataset into frame sequences;

[0067] A data preprocessing module for performing data augmentation by randomly flipping and transforming the parsed human action frame sequences;

[0068] A dataset partitioning module for sampling a training dataset and a test dataset from the preprocessed human action frame sequences;

[0069] A feature extraction module for extracting the appearance features of human actions in the channel domain and spatio-temporal domain, as well as the temporal difference temporal features;

[0070] A model training module that trains a human action feature model using the extracted time difference time series features;

[0071] A model testing module that tests the trained human action feature model using a test data set to obtain the final classification result of the human action.

[0072] A device of the present invention includes a memory and a processor, wherein:

[0073] The memory is used to store a computer program that can run on the processor;

[0074] The processor is used to execute the steps of the above-mentioned human action recognition method based on the split attention network when running the computer program.

[0075] A storage medium of the present invention stores a computer program, and when the computer program is executed by at least one processor, the steps of the above-mentioned human action recognition method based on the split attention network are implemented.

[0076] Beneficial effects: Compared with the prior art, the channel domain attention module of the present invention further reduces the number of network parameters while performing cross-channel interaction by introducing a 1D convolutional layer; the spatio-temporal domain attention module uses average and max pooling to maximize the aggregation of spatial features, and makes up for the problem of insufficient extraction of spatio-temporal features by the network by introducing a 3D convolutional layer; the time difference dot product self-attention module weights the time difference features on the normalized weight matrix, so that the input T-frame sequence gets different weight distributions in time sequence, so that the attention focuses on features with more information, further improving the accuracy of action recognition. In addition, the split attention network and the time difference dot product self-attention module proposed by the present invention can further improve the accuracy of human action recognition without increasing the amount of calculation. Description of the Drawings

[0077] Figure 1 It is a block diagram of the human action recognition method of the present invention;

[0078] Figure 2 It is a structural diagram of the channel domain attention module;

[0079] Figure 3 It is a structural diagram of the spatio-temporal domain attention module;

[0080] Figure 4 It is a structural diagram of the split attention network;

[0081] Figure 5 It is a structural diagram of the time series network combining BiLSTM and the time difference dot product self-attention module;

[0082] Figure 6The structural diagram of the time-difference dot product self-attention module. Detailed implementation manners

[0083] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only for explaining the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. All equivalent or modified forms made according to the spirit of the present invention should be covered within the protection scope of the present invention.

[0084] As Figure 1 shown, a human action recognition method based on a split attention network of the present invention includes the following steps:

[0085] S1. Parse the videos in the human action recognition dataset into a sequence of human action frames, where the videos in the human action recognition dataset are labeled human action videos;

[0086] S2. Perform random flipping and transformation operations on the parsed sequence of human action frames for data augmentation to obtain a preprocessed sequence of human action frames, and sample to obtain a training dataset and a test dataset;

[0087] To avoid overfitting during the training process, perform random flipping and transformation operations on the parsed sequence of human action frames for data augmentation to obtain a preprocessed sequence of human action frames.

[0088] Randomly select a sampling interval and a starting frame from the preprocessed sequence of human action frames as the training dataset; uniformly sample from the first frame of the preprocessed sequence of human action frames as the test dataset.

[0089] S3. Input the training dataset in step S2 into the split attention network to extract the human action appearance features in the channel domain and the spatio-temporal domain;

[0090] In the embodiment of the present invention, the split attention network includes a backbone network module. The backbone network module uses a ResNet-50 residual network pre-trained on the ImageNet dataset, which contains 5 residual blocks with the same structure, and the depths are 64, 128, 256, 512, 512 respectively, and the stride is 2. The 7×7 convolutional layer of the first residual block in the pre-trained ResNet-50 residual network extracts the underlying human action features R in the training dataset and outputs them to the channel domain attention module and the spatio-temporal domain attention module respectively; the channel domain attention module uses spatial adaptive average pooling to sequentially infer a one-dimensional channel domain attention mask M ca ; and the one-dimensional channel domain attention mask M caMultiply it with the underlying human action feature R and then sum it to obtain the channel-domain attention feature R ca ; The spatio-temporal domain attention module uses channel average and max pooling to sequentially infer a one-dimensional spatio-temporal domain attention mask M′ sta ; And the one-dimensional spatio-temporal domain attention mask Ms sta Multiply it with the underlying human action feature R and then sum it to obtain the spatio-temporal domain attention feature R sta ; Add the obtained channel-domain attention feature R ca and the spatio-temporal domain attention feature R sta to obtain the mixed feature of human action; Finally, pass the mixed feature of human action through a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer in sequence and then input it into the second residual block; Then pass through the third residual block, the fourth residual block, and the fifth residual block in sequence to output the appearance feature of human action. Specifically:

[0091] Figure 2 is the structure diagram of the channel-domain attention module, which enhances the influence of channel-domain features by compressing spatial features and further enhances the expression ability of channel features through local cross-channel interaction. The specific steps are as follows:

[0092] S301. For the underlying human action feature R extracted by the 7×7 convolutional layer of the first residual block in the ResNet-50 residual network pre-trained on the ImageNet dataset, the channel-domain attention uses spatial adaptive average pooling to perform spatial feature encoding on all channels, compresses the spatial features into a global feature, and compresses the global spatial information into the channel descriptor, so as to obtain sufficient channel-domain feature information. The average pooling formula used is:

[0093]

[0094] where the input feature of the channel-domain attention module is the underlying human action feature R∈R NT×C×H×W , NT is the number of underlying human action feature maps, C is the number of channels of each underlying human action feature map, H is the height of each underlying human action feature map, W is the width of each underlying human action feature map, and the output feature F of the channel-domain attention module ∈R NT×C×1×1 .

[0095] S302. Use a 2D convolutional layer k 1 to compress the number of channels of the output feature F of the channel-domain attention module at a ratio r (here r = 16), further reducing the number of parameters. The formula it uses is:

[0096] F r = k 1 *F(2)

[0097] Among them, k 1 is a 1×1 2D convolutional layer, and F r is the compressed channel feature. To input the 2D convolutional feature into a 1D convolutional layer, so F r is reshaped into

[0098] S303. Input the compressed channel feature F′ reshaped in step S302 r into the 1D convolutional layer k 2 to perform cross-channel interaction. The formula used is:

[0099] F temp = k 2 * F′ r (3)

[0100] Among them, k 2 is a 1×1 1D convolutional layer, and F temp is the interactive channel feature. To input the 1D convolutional feature into a 2D convolutional layer, so F temp is reshaped into

[0101] S304. Use the 2D convolutional layer k 3 to decompress the interactive channel feature F′ reshaped in step S303 temp and feed it into the Sigmoid activation function. The formulas used respectively are:

[0102] F c = k 3* F′ temp (4)

[0103] M ca = δ(F c ) (5)

[0104] Among them, F c is the channel mask, F c ∈ R NT×C×1×1 and M ca is the one-dimensional channel-domain attention mask, M ca ∈ R NT ×C×1×1 , and δ(·) is the Sigmoid activation function. The Sigmoid activation function can make F c have stronger non-linear expression ability; the finally output channel-domain attention feature R ca is:

[0105] R ca = R + R ⊙ M ca (6).

[0106] Figure 3 It is the structural diagram of the spatio-temporal domain attention module, which enhances the influence of spatial features by compressing channel features and further enhances the temporal expression ability of spatial features through a 3D convolutional layer. The specific steps are as follows:

[0107] S311. Reshape the underlying human action features R∈R extracted from the 7×7 convolutional layer of the first residual block in the ResNet-50 residual network pre-trained on the ImageNet dataset NT×C×H×W to R′∈R N×T×C×H×W ;

[0108] S312. The spatio-temporal domain attention module uses channel average pooling and channel max pooling to perform channel feature encoding over all spaces, compressing the channel features into global features F avg and F max , respectively, and compressing the global channel information into the spatial descriptor to obtain sufficient spatio-temporal feature information. The average pooling and max pooling formulas used are as follows:

[0109]

[0110] F max =max(R′[:,:,i,:,:]) (8)

[0111] where F avg is the spatial attention average pooling feature, F avg ∈R N×T×1×H×W , F max is the spatial attention max pooling feature, F max ∈R N×T×1×H×W . To adapt to the characteristics of 3D convolution, reshape the spatial attention average pooling feature F avg to F′ avg ∈R N×1×T×H×W , and reshape the spatial attention max pooling feature F max to Fr max ∈R N×1×T×H×W .

[0112] S313. Use the 3D convolutional layer k 4 to perform spatio-temporal feature extraction on the input reshaped spatial attention average pooling feature F′ avg and the reshaped spatial attention max pooling feature F′ max , respectively. The specific formulas are as follows:

[0113] F″ avg =k 4 *F′ avg (9)

[0114] F″max = k 4 * F′ max (10)

[0115] where k 4 is a 1×1 3D convolutional layer, and F″ avg is the spatio-temporal attention average pooling feature, and F″ avg ∈ R N ×1×T×H×W , and F″ max is the spatio-temporal attention max pooling feature, and F″ max ∈ R N×1×T×H×W . To restore the features of the 2D convolution, the spatio-temporal attention average pooling feature F″ avg is reshaped into The spatio-temporal attention max pooling feature F″ max is reshaped into

[0116] S314. For the reshaped spatio-temporal attention average pooling feature output in step S313 and the reshaped spatio-temporal attention max pooling feature are fused and fed into the Sigmoid activation function. The formula used is:

[0117]

[0118] where M sta is a one-dimensional spatio-temporal domain attention mask;

[0119] S315. The one-dimensional spatio-temporal domain attention mask M in step S314 sta ∈ R N×T×1×H×W is reshaped into a 2D convolutional feature M′ sta ∈ R NT×1×H×W , and final feature weighting is performed. The formula used is:

[0120] R sta = R + R ⊙ M′ sta (12)

[0121] where R sta is the spatio-temporal domain attention feature;

[0122] S316. The obtained channel domain attention feature R ca and the spatio-temporal domain attention feature R sta are added together to obtain the mixed feature X of human actions;

[0123] S317. The mixed feature X of the human body movement is successively passed through a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer, and then input into the second residual block; then it is successively passed through the third residual block, the fourth residual block, and the fifth residual block to output the appearance feature of the human body movement.

[0124] Figure 4 As shown in , the split attention network includes a backbone network module. The backbone network module includes 5 sequentially connected residual blocks. Each residual block includes a 7×7 convolutional layer, a channel domain attention module, a spatio-temporal domain attention module, a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer. The training dataset first passes through a 7×7 convolutional layer to extract the underlying feature R of the human body movement, then passes through a human body movement feature extraction structure with a channel domain attention module and a spatio-temporal domain attention module in parallel, and finally passes through a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer in sequence.

[0125] S4. Input the appearance feature of the human body movement obtained in step S3 into a temporal network model combined with a BiLSTM recurrent neural network and a time-dot product self-attention module to extract the time difference temporal feature Y of the human body movement;

[0126] Figure 5 As shown in , BiLSTM consists of two layers of LSTM, which learn from the forward and backward directions respectively, so that the output at the current moment is related not only to the previous state but also possibly to the future state, enabling it to better capture bidirectional time series features; the time-dot product self-attention module reduces the dependence on external parameters by dynamically adjusting the weights of the time series features, which helps BiLSTM capture global features.

[0127] As Figure 5 shown, the BiLSTM layer learns the temporal feature representation between frame sequences from both the forward and backward directions, and splices the features learned by the forward LSTM vector and the features learned by the backward LSTM vector and then inputs them into the time-dot product self-attention layer. The calculation formula is as follows:

[0128]

[0129]

[0130]

[0131] Among them, is the forward LSTM, is the backward LSTM, X i is the mixed feature of the i-th frame of the human body movement extracted by the split attention network, and T represents the sequence of T frame images in the input training dataset. is the output result of the forward LSTM, is the output result of the backward LSTM, and h is the temporal feature.

[0132] Figure 6 is the structural diagram of the time-difference dot product self-attention module. As a variant of the attention mechanism, self-attention can reduce the dependence on external information by dynamically adjusting the feature weights, pay more attention to the internal feature correlation, and effectively capture the long-distance related features; the time-difference dot product self-attention weights the time-difference features on the normalized weight matrix, so that the T-frame image sequences in the input training dataset are assigned different weights according to the time sequence, so that the attention focuses on the part with more information, further improving the accuracy of action recognition.

[0133] The specific steps are as follows:

[0134] S41. Input the mixed temporal features extracted by the BiLSTM recurrent neural network into the time-difference dot product self-attention module, and obtain three feature matrices Q, K, and V through three identical linear mappings respectively;

[0135] S42. Calculate the dot product similarity between the feature matrix Q and the transpose of the feature matrix K to obtain the feature similarity weight matrix S; S = Q · K T ;

[0136] S43. Normalize the obtained weight matrix using the Softmax function; the normalized weight matrix Att_S is: Att_S = Softmax(S);

[0137] S44. Perform time-difference feature weighting on the normalized weight matrix in step S43, so that the T-frame image sequences in the input training dataset are assigned different weights due to the different input times;

[0138] S45. Perform dot product weighted summation on the time-difference weight matrix obtained in step S44 and the feature matrix V to obtain the final time-difference temporal feature Att_TD, Att_TD = sum(TD · V), where TD is the weight matrix after time-difference feature weighting.

[0139] S5. Train the human action feature model based on the time-difference temporal feature obtained in step 4, input the test dataset into the trained human action feature model, and obtain the final classification result of the human action.

[0140] A human action recognition system based on a shunt attention network, including:

[0141] A video parsing module for parsing the videos in the human action recognition dataset into frame sequences;

[0142] A data preprocessing module for performing random flipping and transformation operations on the parsed human action frame sequence for data augmentation;

[0143] A dataset partitioning module for sampling a training dataset and a test dataset from the preprocessed human action frame sequence;

[0144] A feature extraction module for extracting appearance features in the channel domain and spatio-temporal domain of human actions, as well as time-difference temporal features;

[0145] A model training module for training a human action feature model using the extracted time-difference temporal features;

[0146] A model testing module for testing the trained human action feature model using the test dataset to obtain the final classification result of human actions.

[0147] A device, including a memory and a processor, wherein:

[0148] The memory is used to store a computer program that can run on the processor;

[0149] The processor is used to execute the steps of the above-mentioned human action recognition method based on the split attention network when running the computer program, and achieve the same technical effects as the above method.

[0150] A storage medium stores a computer program, and when the computer program is executed by at least one processor, it implements the steps of the above-mentioned human action recognition method based on the split attention network, and achieves the same technical effects as the above method.

Claims

1. A human action recognition method based on a split attention network, characterized in that, it includes the following steps: S1. Parse the videos in the human action recognition dataset into a sequence of human action frames, where the videos in the human action recognition dataset are labeled human action videos; S2. Perform random flipping and transformation operations on the parsed sequence of human action frames for data augmentation to obtain a preprocessed sequence of human action frames, and sample to obtain a training dataset and a test dataset; S3. Input the training dataset in step S2 into the split attention network to extract the human action appearance features in the channel domain and the spatio-temporal domain; the split attention network includes a backbone network module, and the backbone network module includes 5 sequentially connected residual blocks, and each residual block includes a 7×7 convolutional layer, a channel domain attention module, a spatio-temporal domain attention module, a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer; The 7×7 convolutional layer of the first residual block in the backbone network module extracts the underlying human action features R in the training dataset and outputs them to the channel domain attention module and the spatio-temporal domain attention module respectively; The channel domain attention module uses spatial adaptive average pooling to sequentially infer a one-dimensional channel domain attention mask M ca ; and multiply and add the one-dimensional channel domain attention mask M ca and the underlying human action feature R to obtain the channel domain attention feature R ca ; The spatio-temporal domain attention module infers a one-dimensional spatio-temporal domain attention mask M′ through channel average and max pooling in sequence sta ; and multiplies and adds the one-dimensional spatio-temporal domain attention mask M′ sta and the underlying features R of human actions to obtain spatio-temporal domain attention features R sta ; For the obtained channel domain attention feature R ca and the spatio-temporal domain attention feature R sta perform addition to obtain the human action hybrid feature in the channel domain and the spatio-temporal domain; The mixed features of the human action pass through a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer in sequence and then are input into the second residual block; then they pass through the third residual block, the fourth residual block, and the fifth residual block in sequence to output the appearance features of the human action; S4. Input the human action appearance features obtained in step S3 into a temporal sequence network model combined with a BiLSTM recurrent neural network and a time difference dot product self-attention module to extract the time difference temporal features of the human action; S5. Train a human action feature model based on the time difference temporal features obtained in step S4, input the test dataset into the trained human action feature model, and obtain the final classification result of the human action.

2. A human action recognition method based on a split attention network according to claim 1, characterized in that, the sampling method of the training dataset in step S2 is: randomly select a sampling interval and a starting frame from the preprocessed sequence of human action frames as the training dataset; the sampling method of the test dataset is: uniformly sample starting from the first frame from the preprocessed sequence of human action frames as the test dataset.

3. A human action recognition method based on a split attention network according to claim 1, characterized in that, The channel domain attention module enhances the influence of the channel domain features by compressing the spatial features and further enhances the expression ability of the channel features through local cross-channel interaction, specifically including the following steps: S301. For the underlying human action features R extracted by the 7×7 convolutional layer of the first residual block in the backbone network module, the channel domain attention module uses spatial adaptive average pooling to perform spatial feature encoding on all channels, compresses the spatial features into a global feature, and compresses the global spatial information into a channel descriptor to obtain channel domain feature information; the spatial adaptive average pooling formula used is: Among them, the input feature of the channel domain attention module is the underlying feature R of human body actions, R ∈ R NT×C×H×W , NT is the number of underlying feature maps of human body actions, C is the number of channels of each underlying feature map of human body actions, H is the height of each underlying feature map of human body actions, W is the width of each underlying feature map of human body actions, and the output feature F of the channel domain attention module is F ∈ R NT×C×1×1 ; S302. Use the 2D convolutional layer k 1 Compress the number of channels of the output feature F of the channel attention module at a ratio r to further reduce the number of parameters. The formula used is: F r = k 1 * F where k 1 is a 1×1 2D convolutional layer, and F r is the compressed channel feature, reshapes F r into S303. Input the compressed channel feature F' reshaped in step S302 r into the 1D convolutional layer k 2 for cross-channel interaction. The formula used is as follows: F temp = k 2 * F' r where k 2 is a 1×1 1D convolutional layer, F temp is the interactive channel feature, reshapes F temp into S304. Use the 2D convolutional layer k 3 Uncompress the interactive channel feature F' reshaped in step S303 temp and feed it into the Sigmoid activation function; the formulas used are respectively: F c = k 3 * F' temp M ca = δ(F c ) Among them, F c is the channel mask, F c ∈R NT×C×1×1 , M ca is a one-dimensional channel domain attention mask, M ca ∈R NT×C×1×1 , δ(·) is the Sigmoid activation function, and the finally obtained channel domain attention feature R ca is: R ca = R + R ⊙ M ca .

4. A human action recognition method based on a split attention network according to claim 1, characterized in that, The spatio-temporal domain attention module enhances the influence of spatial features by compressing channel features and further enhances the temporal expression ability of spatial features through a 3D convolutional layer. The specific steps are as follows: S311. Reshape the underlying feature R∈R extracted by the 7×7 convolutional layer of the first residual block in the backbone network module into R′∈R NT×C×H×W ; N×T×C×H×W ; S312. The spatio-temporal domain attention module uses channel average pooling and channel max pooling to perform channel feature encoding over all spaces, compressing the channel features into global features F avg and F max , and compresses the global channel information into the spatial descriptor to obtain spatio-temporal feature information. The average pooling and max pooling formulas used are as follows: F max = max(R′[:,:,i,:,:]) Among them, F avg is the spatial attention average pooling feature, and F avg ∈R N×T×1×H×W ; F max is the spatial attention maximum pooling feature, and F max ∈R N×T×1×H×W ; C is the number of channels of the underlying features of human actions; Average-pool the spatial attention feature F avg Reshape it into F' avg ∈R N×1×T×H×W ; Reshape the spatial attention maximum pooling feature F max into F' max ∈ R N×1×T×H×W ; S313. Use the 3D convolutional layer k 4 Perform spatio-temporal feature extraction on the reshaped spatial attention average pooling feature F' avg and the reshaped spatial attention max pooling feature F' max respectively; the specific formulas are as follows: F″ avg = k 4 * F′ avg F″ max = k 4 * F′ max where k 4 is a 1×1 3D convolutional layer, and F″ avg is the spatio-temporal attention average pooling feature, and F″ avg ∈R N×1×T×H×W , and F″ max is the spatio-temporal attention maximum pooling feature, and F″ max ∈R N×1×T×H×W , Spatio-temporal attention average pooling feature F″ avg Reshape it into Reshape the spatio-temporal attention maximum pooling feature F″ max into S314. Reshape the spatio-temporal attention average pooling features output in step S313 and the reshaped spatio-temporal attention max pooling features for fusion and feed them into the Sigmoid activation function; the formula used is: Among them, M sta is a one-dimensional spatio-temporal domain attention mask; S315. Reshape the one-dimensional spatio-temporal domain attention mask M in step S314 sta ∈R N×T×1×H×W into a 2D convolutional feature M' sta ∈R NT×1×H×W , and perform final feature weighting; the formula used is: R sta = R + R ⊙ M' sta Among them, R sta is the spatio-temporal domain attention feature; S316. Add the obtained channel domain attention feature R ca and the spatio-temporal domain attention feature R sta to obtain the human action hybrid feature X of the channel domain and the spatio-temporal domain; S317: The mixed features X of the human action are successively input into the second residual block after passing through a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer; then, they successively pass through the third residual block, the fourth residual block, and the fifth residual block to output the appearance features of the human action.

5. A human action recognition method based on a split attention network according to claim 1, characterized in that, step S4 includes the following steps: S41: The BiLSTM recurrent neural network learns the temporal feature representation between the frames of the human action appearance features from both forward and backward directions, and splices the features learned by the forward LSTM vector and the features learned by the backward LSTM vector to obtain the mixed temporal features of the human action, and then inputs them into the time-difference dot product self-attention module; S42: The time-difference dot product self-attention module respectively obtains three feature matrices Q, K, and V by passing the input mixed temporal features of the human action through three identical linear mappings; S43: Calculate the dot product similarity between the feature matrix Q and the transpose of the feature matrix K to obtain the weight matrix; S44: Normalize the obtained weight matrix using the Softmax function; S45: Perform time-difference feature weighting on the normalized weight matrix in step S44, so that the T-frame sequences in the input training data set are assigned different weights due to the different input time sequences; S46: Calculate the dot product of the weight matrix obtained in step S45 and the feature matrix V and perform weighted summation to obtain the final time-difference temporal features.

6. A human action recognition system based on a split attention network, characterized in that, it includes: A video parsing module for parsing the videos in the human action recognition data set into frame sequences; A data preprocessing module for performing random flipping and transformation operations on the parsed human action frame sequences for data augmentation; A data set partitioning module for sampling from the preprocessed human action frame sequences to obtain a training data set and a test data set; A feature extraction module for inputting the training data set obtained by the data set partitioning module into the split attention network to extract the appearance features of human actions in the channel domain and the spatio-temporal domain, and inputting the appearance features of human actions into a temporal network model combined with a BiLSTM recurrent neural network and a time-difference dot product self-attention module to extract the time-difference temporal features of human actions; the split attention network includes a backbone network module, and the backbone network module includes 5 sequentially connected residual blocks, and each residual block includes a 7×7 convolutional layer, a channel domain attention module, a spatio-temporal domain attention module, a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer; The 7×7 convolutional layer of the first residual block in the backbone network module extracts the underlying features R of the human action in the training data set and outputs them to the channel domain attention module and the spatio-temporal domain attention module respectively; The channel attention module uses spatial adaptive average pooling to sequentially infer a one-dimensional channel attention mask M ca ; And multiply the one-dimensional channel domain attention mask M ca and the underlying human action feature R, and then sum them to obtain the channel domain attention feature R ca ; The spatio-temporal domain attention module infers a one-dimensional spatio-temporal domain attention mask M′ through channel average and max pooling in sequence sta ; and multiply and add the one-dimensional spatio-temporal domain attention mask M′ sta and the underlying features R of human actions to obtain spatio-temporal domain attention features R sta ; For the obtained channel-domain attention feature R ca and the spatio-temporal domain attention feature R sta perform addition to obtain the human action hybrid feature in the channel domain and the spatio-temporal domain; The mixed features of human body movements are sequentially input into the second residual block after passing through a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer; then they pass through the third residual block, the fourth residual block, and the fifth residual block in sequence to output the appearance features of human body movements; The model training module uses the extracted time difference time series features to train the human body movement feature model; The model testing module uses the test data set to test the trained human body movement feature model to obtain the final classification result of human body movements.

7. A device, characterized in that, it includes a memory and a processor, wherein: The memory is used to store a computer program that can run on the processor; The processor is used to execute the steps of a human body movement recognition method based on a split attention network according to any one of claims 1-5 when running the computer program.

8. A storage medium, characterized in that, a computer program is stored on the storage medium, and when the computer program is executed by at least one processor, the steps of a human body movement recognition method based on a split attention network according to any one of claims 1-5 are implemented.

Citation Information

Patent Citations

  • Human body behavior identification method based on Bi-LSTM-Attention model

    CN109784280A

  • Attention weight calculation method and device based on convolutional neural network

    CN110909862A

  • Video feature extraction method and device, equipment and storage medium

    CN113515994A