A Behavior Recognition Method Based on Channel-Time Features
By inserting channel attention module, long-term timing module and short-term timing module into the CNN network, combining channel information and timing information, the problem of information extraction and combination in behavior recognition is solved, and the recognition performance is significantly improved.
Patent Information
- Application Number
- CN202210614670.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-05-30
AI Technical Summary
In the prior art, in behavior recognition tasks, it is difficult to effectively extract and combine channel information and timing information, resulting in low recognition performance.
Insert the channel attention module, long-term timing module and short-term timing module with complementary functions in the CNN network to establish a connection between channel information and timing information, and improve network performance through multi-module fusion.
Through the synergy between multiple modules, the performance of behavior recognition is significantly improved, and the problem of channel information and timing information characterization is solved.
Smart Images

Figure CN114937225B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and relates to a method for recognizing the actions of targets in video files, and particularly relates to a behavior recognition method based on channel-time features. Background Art
[0002] Behavior recognition is a very challenging topic in the field of computer vision. The main research object is the actions of targets in videos, such as determining whether a person is walking, jumping or waving, etc. Behavior recognition methods have important applications in scenarios such as video surveillance, video recommendation and human-computer interaction. In recent decades, with the rise of neural networks, many network structures have been applied to perform behavior recognition tasks. Different from object recognition, behavior recognition requires not only analyzing the spatial dependence relationship of the target, but also analyzing the historical information of the target change. Therefore, how to better extract spatio-temporal features is the key problem to improve the performance of behavior recognition.
[0003] Convolutional Neural Network (CNN) is a feedforward neural network. Due to its network structure characteristics, CNN has superior performance in image processing, especially in large-scale image processing. In addition, CNN also has obvious advantages in terms of computational complexity compared with other network structures. The Temporal Shift Module (TSM) performs efficient temporal modeling by moving the feature map along the time dimension, achieving powerful temporal modeling capabilities, solving the problem that traditional 2D convolution cannot capture the relationship in the time dimension, and the high deployment cost of the method based on 3D convolution. The SE attention mechanism models the dependencies of each channel to improve the channel information representation ability of the network, and can adjust the features channel by channel, so that the network can learn to selectively enhance the features containing useful information and suppress useless features through global information. Inserting the SE block into the main network will only increase a small amount of computational consumption, but can greatly improve the network performance. Inserting the temporal shift module and the attention mechanism into the convolutional neural network to represent the channel information and temporal information of the file to be recognized can achieve the behavior recognition task. However, when constructing a recognition model that combines multiple modules in the prior art, multiple modules are usually inserted in an independent manner, having disadvantages such as weak inter-module correlation and redundant network structure. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention proposes a behavior recognition method based on channel-time features. By inserting a channel attention module, a long-term temporal module and a short-term temporal module with complementary functions into the CNN network, while representing the channel information and temporal information, the connection between the two is established, so that the detection model can learn the information in the video more fully and improve the recognition performance.
[0005] A behavior recognition method based on channel-time features, specifically including the following steps:
[0006] Step 1: Divide the video file containing the target to be recognized into T video segments with equal durations, and then randomly extract one frame image from each video segment as the recognition data set.
[0007] Step 2: Construct a channel attention module for selectively enhancing some features. The channel attention module is obtained by inserting a 1D convolutional layer in the middle of the two fully connected layers of the SE attention mechanism and replacing the two fully connected layers with two 2D convolutional layers respectively.
[0008] Step 3: Construct a short-term temporal module for characterizing local motion information. The short-term temporal module compresses the number of channels of the input data, and then calculates the frame difference information of the features at the t-1 moment, t+1 moment and t moment respectively. After adding them, they pass through a pooling layer, a 2D convolutional layer with a size of 1x1 and a Sigmoid activation layer respectively, and then the processing results are convolved and added to the input data before channel compression in turn to output the characterization result of local motion information.
[0009] Step 4: Construct a long-term temporal module for time shift processing. The long-term temporal module divides the input data into 8 equal parts according to the channel order, and then shifts the first part and the third part of the data backward in the temporal dimension, shifts the second part of the data forward in the temporal dimension, and does not process other data.
[0010] Step 5: Construct a multi-module block based on the channel attention module, short-term temporal module and long-term temporal module obtained in Steps 2-4. The multi-module block inputs the data into the channel attention module and the long-term temporal module respectively. The output of the channel attention module passes through the short-term temporal module and then is added to the output of the channel attention module, and then passes through two convolutional modules and is added to the output of the long-term temporal module, and the calculation result is used as the output of the multi-module block.
[0011] Step 6: Construct a ResNet-50 neural network, use the multi-module block obtained in Step 5 to replace the blocks in the 2nd to 5th layers of ResNet-50 to obtain a multi-module behavior recognition model. Use the data with known labels to train this model, and then input the recognition data set obtained in Step 1 into the trained model to obtain the action classification of the target in the video image, and complete the behavior recognition.
[0012] The present invention has the following beneficial effects:
[0013] Model the channel information using the channel attention module, and model the temporal information using the long- and short-term temporal modules, which solves the two major problems of representing channel information and temporal information in the action recognition task. The above modules are fused in a block in serial, summation, multi-branch and other ways, and then inserted into the main network to achieve functional complementarity of multiple modules, so that multiple modules cooperate to improve the network performance. Description of the Drawings
[0014] Figure 1 It is a flowchart of the action recognition method;
[0015] Figure 2 It is a schematic structural diagram of the channel attention module in the embodiment;
[0016] Figure 3 It is a schematic structural diagram of the short-term temporal module in the embodiment;
[0017] Figure 4 It is a schematic structural diagram of the long-term temporal module in the embodiment. Detailed Embodiment
[0018] The present invention will be further explained below with reference to the accompanying drawings.
[0019] As Figure 1 shown, an action recognition method based on channel-time features specifically includes the following steps:
[0020] Step 1: Divide the video file V containing the target to be recognized into T video segments with equal duration, V=(V1, V2, … V T ), and then randomly extract one frame of image from each video segment. After data augmentation of these T frames of images, uniformly resize them to 224x224 as the recognition image data for subsequent input to the model.
[0021] Step 2: Construct a channel attention module as Figure 2 shown for selectively strengthening some features. The channel attention module is modified based on the SE attention mechanism. Specifically, a 1D convolutional layer is inserted in the middle of two fully connected layers, and two 2D convolutional layers are used to replace the two fully connected layers respectively. The processing steps of the channel attention module for the input data X∈R N ×T×C×H×W are as follows:
[0022] s2.1: Obtain the global spatial information X1 of the input data X through the pooling layer:
[0023]
[0024] where, X1∈R N×T×C×1×1, H, W, and C represent the height, width, and number of channels of the input data X respectively, N represents the batch size, and X[i, j] represents the pixel at the i-th row and j-th column.
[0025] s2.2. Compress X1 into X2 using a 2D convolutional layer with a kernel size of 1x1:
[0026] X2 = K1 * X1 (2)
[0027] where, K1 represents the convolutional kernel of size 1x1, and r represents the compression ratio.
[0028] s2.3. Process X2 sequentially using a 1D convolutional layer with a kernel size of 3, a 2D convolutional layer with a kernel size of 1x1, and a Sigmoid activation layer to obtain X4:
[0029] X4 = K3 * X3 (3)
[0030] X3 = K2 * X2 (4)
[0031] where, X4 ∈ R N×T×C×1×1 , K2 represents the convolutional kernel of size 3, and K3 represents the convolutional kernel of size 1x1 and the Sigmoid activation operation.
[0032] s2.4. According to the dependencies of each channel obtained in the above steps, selectively enhance the features containing useful information and suppress the useless features, and finally output X O1 ∈ R N×T×C×H×W :
[0033] X O1 = X + X ⊙ X4 (5)
[0034] where, ⊙ represents the matrix multiplication operation between feature maps.
[0035] Step 3. Construct a short-term temporal module as shown in Figure 3 to represent local motion information. The processing steps of the short-term temporal module for the input data X ∈ R N×T×C×H×W are as follows:
[0036] s3.1. Compress the number of channels of the input data X using a 2D convolutional layer with a kernel size of 1x1 to obtain
[0037] s3.2. Calculate the frame difference information of the features at time t - 1, t + 1, and t respectively, and then add the frame difference information:
[0038] X6 = (K * X5[t + 1] - X5[t]) + (K * X5[t] - X5[t - 1]) (6)
[0039] Among them, is the frame difference at the feature level, representing local motion information. K represents a 2D convolution kernel of size 3×3, and X5[t-1], X5[t], and X5[t+1] represent the features at time t-1, time t, and time t+1, respectively.
[0040] S3.3. For the frame difference information X6 obtained in step S3.2, use the pooling layer to obtain the global spatial information X7∈R N ×T×C×1×1 , and then pass through a 2D convolution layer with a convolution kernel size of 1x1 and a Sigmoid activation layer in sequence to obtain X8∈R N×T×C×1×1 .
[0041] S3.4. Finally, splice the features at different levels to obtain the output X of the short-term temporal module O2 ∈R N×T×C×H×W :
[0042] X O2 = X + X⊙X8 (7)
[0043] Step 4. Construct a long-term temporal module as shown in Figure 4 for time shift processing. The long-term temporal module divides the input data into 8 equal parts according to the channel order, i = 1, 2, …, 8, and then shifts the first and third data X1 and X3 backward in the temporal dimension, shifts the second data X2 forward in the temporal dimension, and does not process other data.
[0044] Step 5. Construct a multi-module block based on the channel attention module, short-term temporal module, and long-term temporal module obtained in steps 2 to 4. The multi-module block inputs the data into the channel attention module and the long-term temporal module respectively to obtain X O1 and X O2 . After the output X O1 of the channel attention module is processed by the time shift of the short-term temporal module, it is added to the unshifted X O1 , and then passes through two convolution modules, and then added to X O2 , and the final result is used as the output of the multi-module block.
[0045] Step 6. Construct a ResNet-50 neural network, use the multi-module block obtained in step 5 to replace the blocks in the 2nd to 5th layers of ResNet-50 to obtain a multi-module action recognition model. Use the data with known labels to train this model, and then input the recognition data set obtained in step 1 into the trained model to obtain the action classification of the target in the video image, and complete the action recognition.
Claims
1. A behavior recognition method based on channel-time features, characterized in that: Specifically, it includes the following steps: Step 1: Divide the video file containing the target to be recognized into T video segments with equal durations, and then randomly extract one frame image from each video segment as the recognition data set; Step 2: Construct a channel attention module for selectively enhancing some features; the channel attention module is obtained by inserting a 1D convolutional layer in the middle of the two fully connected layers of the SE attention mechanism and replacing the two fully connected layers with two 2D convolutional layers respectively; Step 3, construct a short-term temporal module to represent local motion information; the processing steps of the short-term temporal module for the input data X∈R N ×T×C×H×W are as follows: S3.
1. Compress the number of channels of the input data X using a 2D convolutional layer with a kernel size of 1x1 to obtain s3.2: Calculate the frame difference information of the features at the t-1 moment, t+1 moment and t moment respectively, and then add the frame difference information: X6 = (K * X5[t+1] - X5[t]) + (K * X5[t] - X5[t-1]) (1) Among them, is the frame difference at the feature level, representing local motion information; K represents a 2D convolution kernel of size 3×3, and X5[t-1], X5[t], and X5[t+1] represent the features at time t-1, time t, and time t+1, respectively; S3.
3. For the frame difference information X6 obtained in step S3.2, use the pooling layer to obtain the global spatial information X7 ∈ R N×T×C×1×1 , and then successively pass through a 2D convolutional layer with a convolutional kernel size of 1x1 and a Sigmoid activation layer to obtain X8 ∈ R N×T×C×1×1 ; S3.
4. Finally, splice the features at different levels to obtain the output X of the short-term time series module O2 ∈R N×T×C×H×W , that is, the result of representing local motion information: X O2 = X + X ⊙ X8 (2) Step 5: Construct a long-term temporal module for time shift processing; the long-term temporal module divides the input data into 8 equal parts according to the channel order, then shifts the first and third parts of the data backward in the temporal dimension, shifts the second part of the data forward in the temporal dimension, and does not process other data; Step 6: Construct a multi-module block based on the channel attention module, short-term temporal module, and long-term temporal module obtained in Steps 2-4; the multi-module block inputs the data into the channel attention module and the long-term temporal module respectively; the output of the channel attention module is added to the output of the channel attention module after passing through the short-term temporal module, and then after passing through two convolutional modules, it is added to the output of the long-term temporal module, and the calculation result is used as the output of the multi-module block; Step 7: Construct a ResNet-50 neural network, use the multi-module block obtained in Step 5 to replace the blocks in the 2nd to 5th layers of ResNet-50 to obtain a multi-module action recognition model; use the data with known labels to train this model, and then input the recognition data set obtained in Step 1 into the trained model to obtain the action classification of the target in the video image, and complete the action recognition.
2. The behavioral recognition method based on channel-time features according to claim 1, wherein: After enhancing the recognition data set obtained in Step 1, resize it to 224x224 and then input it into the trained model.
3. The method for behavior recognition based on channel-time features according to claim 1, wherein: The processing steps of the channel attention module for the input data X∈R N×T×C×H×W are as follows: s2.1: Obtain the global spatial information X1 of the input data X through a pooling layer: where X1 ∈ R N×T×C×1×1 , H, W, and C represent the height, width, and number of channels of the input data X, N represents the batch size, and X[i, j] represents the pixel at the i-th row and j-th column; s2.2: Compress X1 into X2 using a 2D convolutional layer with a convolutional kernel size of 1x1: X2 = K1 * X1 (4) Among them, K1 represents a 1x1 convolutional kernel, and r represents the compression ratio; s2.3: Process X2 sequentially using a 1D convolutional layer with a convolutional kernel size of 3, a 2D convolutional layer with a convolutional kernel size of 1x1, and a Sigmoid activation layer to obtain X4: X4 = K3 * X3 (5) X3 = K2 * X2 (6) Among them, X4 ∈ R N×T×C×1×1 , K2 represents a convolutional kernel of size 3, and K3 is a convolutional kernel of size 1x1 and a Sigmoid activation operation; S2.
4. Based on the dependencies of each channel obtained in the above steps, selectively enhance the features containing useful information and suppress the useless features, and finally output X O1 ∈R N×T×C×H×W : X O1 = X + X ⊙ X4 (7) Where, ⊙ represents the matrix multiplication operation between feature maps.
Citation Information
Patent Citations
Action recognition method based on double-flow convolution attention
CN112926396A
KR20220050758A