A Human Action Recognition Method Based on Dual-Stream Non-Local Spatiotemporal Convolutional Neural Network
Through the dual-stream non-local spatiotemporal convolutional neural network, combined with the recognition results of spatial flow and temporal flow sub-network, high-accuracy human behavior recognition is achieved, solving the problem of low accuracy in the prior art.
Patent Information
- Application Number
- CN202210460726.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-04-28
AI Technical Summary
The prior art has low accuracy in human behavior recognition, and it is difficult to effectively combine the spatial and temporal structures in videos.
The dual-flow non-local spatiotemporal convolution neural network is used to process the RGB image sequence and the optical flow image sequence respectively through the spatial flow and the temporal flow subnet, and the final human behavior type prediction is obtained through mean fusion.
It realizes end-to-end human behavior recognition, has high accuracy, and can effectively combine the spatial and temporal structure in the video.
Smart Images

Figure CN114898461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network, belonging to the technical field of computer vision. Background Art
[0002] In recent years, human behavior recognition in videos has become a research hotspot in the field of computer vision. Currently, the research methods in this field can be divided into two categories, including machine learning methods based on manually designed features and methods based on deep neural networks. Representative methods in the methods based on manually designed features include interest point detection method, sparse and dense sampling, etc. The earliest action recognition work used 3D models to describe actions and understand and interpret human behaviors. The holistic representation method similar to the structural model of the human body is more likely to retain the spatial and temporal structure of actions. However, currently, deep learning methods are favored, and using deep learning to process image and video data is a research hotspot. For example, convolutional neural networks do not require manual feature extraction, can obtain underlying feature information from training samples, and then obtain high-level feature information through multi-layer convolution, which is applied to the processing of data such as images and videos. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network, which can achieve end-to-end human behavior recognition and has a high accuracy rate.
[0004] To achieve the above purpose, the present invention is implemented by the following technical solutions:
[0005] The present invention provides a human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network, including:
[0006] Obtain a video to be detected containing human behaviors;
[0007] Extract the RGB image sequence frame by frame from the video to be detected, and input the RGB image sequence into the trained spatial stream convolutional neural network NST-CNN to obtain the spatial stream human behavior type prediction;
[0008] Input the video to be detected into the trained PWC-Net network to generate an optical flow image sequence, and input the optical flow image sequence into the trained temporal stream convolutional neural network NST-CNN to obtain the temporal stream human behavior type prediction;
[0009] Perform mean fusion according to the spatial stream and temporal stream human behavior type predictions to obtain the final human behavior type prediction.
[0010] Optionally, the step of inputting the video to be detected into the trained PWC-Net network to generate an optical flow image sequence includes: separately inputting the images of adjacent frames of the video to be detected into a six-level feature pyramid network to obtain feature maps at various scales; performing cost volume calculation, warping operation, optical flow extraction layer, upsampling, and context network layer on the feature maps to generate an optical flow image sequence; wherein, the PWC-Net network is trained using the optical flow datasets Flying Chairs and Flying Things3D as training datasets.
[0011] Optionally, both the spatial flow and temporal flow convolutional neural network NST-CNN include a non-local spatio-temporal convolutional layer, a first non-local spatio-temporal convolutional block, a second non-local spatio-temporal convolutional block, a third non-local spatio-temporal convolutional block, a fourth non-local spatio-temporal convolutional block, a fifth non-local spatio-temporal convolutional block, a 3D pooling layer, a fully connected layer, a Dropout layer, and a Softmax layer, which are connected in sequence.
[0012] Optionally, the first non-local spatio-temporal convolutional block includes a spatio-temporal convolutional layer; the second non-local spatio-temporal convolutional block includes a non-local module and three residual blocks; the third non-local spatio-temporal convolutional block includes a non-local module and four residual blocks; the fourth non-local spatio-temporal convolutional block includes a non-local module and six residual blocks; the fifth non-local spatio-temporal convolutional block includes a non-local module and three residual blocks.
[0013] Optionally, the residual block includes a spatio-temporal convolutional layer, a batch normalization layer, a Leaky ReLU activation function layer, and a spatio-temporal convolutional layer, which are connected in sequence. The input and output of the residual block are directly connected, and a Leaky ReLU activation function layer is connected after each residual block.
[0014] Optionally, the spatio-temporal convolutional layer includes a spatial convolutional layer, a batch normalization layer, a Leaky ReLU activation function layer, and a temporal convolutional layer, which are connected in sequence.
[0015] Optionally, the Leaky ReLU activation function is:
[0016]
[0017] where x is the input and λ is the parameter.
[0018] Optionally, the digital representation of the non-local module is:
[0019]
[0020] where x i and x j are the feature values at positions i and j of the input signal respectively, and Zi The output eigenvalue at position i of the input signal; f(x i , x j ) is the correlation function of the eigenvalues at positions i and j of the input signal; g(x j ) = W g x j , W g and W Z are weight matrices, C(x) is a normalization parameter, and:
[0021]
[0022]
[0023] where θ(x i ) = W θ x i , Φ(x j ) = W Φ x j , and W θ 、W Φ are weight matrices.
[0024] Optionally, the training of the spatial stream and temporal stream convolutional neural network NST-CNN includes: using the Kinetics-400 dataset as the training dataset to perform a first training on the spatial stream and temporal stream convolutional neural network NST-CNN; using the UCF101 dataset and the HMDB51 dataset as the training dataset to perform a second training on the spatial stream and temporal stream convolutional neural network NST-CNN; during the first training and the second training, the training is optimized by adopting the improved stochastic gradient descent algorithm with momentum based on the gradient centering algorithm.
[0025] Optionally, the obtaining of the final human behavior type prediction includes:
[0026]
[0027] where y average is the final human behavior type prediction, and x t , x s are the human behavior type predictions of the spatial stream and temporal stream respectively.
[0028] Compared with the prior art, the beneficial effects achieved by the present invention:
[0029] A human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network provided by the present invention combines the respective characteristics of a 3D network and a two-stream network, and at the same time introduces a non-local module into the network, and proposes a two-stream non-local spatio-temporal convolutional neural network for human behavior recognition. The two-stream network consists of a temporal stream subnet and a spatial stream subnet, and the recognition networks of both parts adopt a non-local spatio-temporal convolutional neural network (NST-CNN); the input of the spatial stream subnet is an RGB image sequence extracted from the video frames to be detected; the input of the temporal stream subnet is an optical flow map sequence extracted from the video to be detected by using an optical flow image estimation network PWC-Net; the recognition results of the spatial stream subnet and the temporal stream subnet are fused by using a mean fusion method to achieve human behavior recognition. In addition, the network is trained by using a stochastic gradient descent algorithm with momentum based on gradient centering improvement. The present invention can achieve end-to-end human behavior recognition and has a high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 FIG. is a flowchart of a human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network provided by an embodiment of the present invention;
[0031] Figure 2 FIG. is a schematic structural diagram of a spatio-temporal convolutional neural network (NST-CNN) provided by an embodiment of the present invention;
[0032] Figure 3 FIG. is a schematic structural diagram of a non-local spatio-temporal convolutional block and a residual block (ResBlock) provided by an embodiment of the present invention;
[0033] Figure 4 FIG. is a schematic structural diagram of a non-local module (Nonlocal Block) provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0035] Embodiment 1:
[0036] As Figure 1 shown, an embodiment of the present invention provides a human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network, including the following steps:
[0037] 1. Obtain a video to be detected containing human behaviors.
[0038] 2. Extract the input video frame by frame to generate an RGB image sequence, and input the RGB image sequence into the trained spatial flow convolutional neural network NST-CNN to obtain the prediction of the human behavior type in the spatial flow.
[0039] 3. Input the video to be detected into the trained PWC-Net network to generate an optical flow image sequence, and input the optical flow image sequence into the trained temporal flow convolutional neural network NST-CNN to obtain the prediction of the human behavior type in the temporal flow.
[0040] Among them, inputting the video to be detected into the trained PWC-Net network to generate an optical flow image sequence includes: respectively sending the images of adjacent frames of the video to be detected into a six-level feature pyramid network to obtain feature maps at various scales; performing cost volume calculation, warping operation, optical flow extraction layer, upsampling, and context network layer on the feature maps to generate an optical flow image sequence; among them, the PWC-Net network uses the optical flow datasets Flying Chairs and Flying Things3D as training datasets for training.
[0041] 4. Perform mean fusion based on the predictions of the human behavior types in the spatial flow and temporal flow to obtain the final prediction of the human behavior type. Obtaining the final prediction of the human behavior type includes:
[0042]
[0043] Among them, y average is the final prediction of the human behavior type, and x t and x s are the predictions of the human behavior types in the spatial flow and temporal flow respectively.
[0044] As Figures 2 - 4 shown, both the spatial flow and temporal flow convolutional neural networks NST-CNN include a non-local spatio-temporal convolutional layer, a first non-local spatio-temporal convolutional block, a second non-local spatio-temporal convolutional block, a third non-local spatio-temporal convolutional block, a fourth non-local spatio-temporal convolutional block, a fifth non-local spatio-temporal convolutional block, a 3D pooling layer, a fully connected layer, a Dropout layer, and a Softmax layer connected in sequence.
[0045] The first non-local spatio-temporal convolutional block includes a spatio-temporal convolutional layer; the spatio-temporal convolutional layer includes a spatial convolutional layer, a batch normalization layer, a Leaky ReLU activation function layer, and a temporal convolutional layer connected in sequence; there are 45 spatial convolutional layers with a size of 1×7×7, and 64 temporal convolutional layers with a size of 3×1×1. The stride of the spatial convolutional layer and the temporal convolutional layer are 1×2×2 and 1×1×1 respectively.
[0046] The second non-local spatio-temporal convolution block includes a non-local module and three residual blocks; the residual block includes a spatio-temporal convolution layer, a batch normalization layer, a Leaky ReLU activation function layer, and a spatio-temporal convolution layer connected in sequence. The input and output of the residual block are directly connected, and a Leaky ReLU activation function layer is connected after each residual block. There are 144 spatio-temporal convolution layers with a size of 1×3×3, and 64 temporal convolution layers with a size of 3×1×1.
[0047] The third non-local spatio-temporal convolution block includes a non-local module and four residual blocks; the residual block includes a spatio-temporal convolution layer, a batch normalization layer, a Leaky ReLU activation function layer, and a spatio-temporal convolution layer connected in sequence. The input and output of the residual block are directly connected, and a Leaky ReLU activation function layer is connected after each residual block. The first spatial convolution layer of the first residual block has 230, and the first temporal convolution layer has 128; the second spatial convolution layer has 288, and the second temporal convolution layer has 128; the spatial convolution layers of the other three residual blocks have 288, and the temporal convolution layers have 128; the size of the spatial convolution layer is 1×3×3, and the size of the temporal convolution layer is 3×1×1.
[0048] The fourth non-local spatio-temporal convolution block includes a non-local module and six residual blocks; the residual block includes a spatio-temporal convolution layer, a batch normalization layer, a Leaky ReLU activation function layer, and a spatio-temporal convolution layer connected in sequence. The input and output of the residual block are directly connected, and a Leaky ReLU activation function layer is connected after each residual block. The first spatial convolution layer of the first residual block has 460, and the first temporal convolution layer has 256; the second spatial convolution layer has 576, and the second temporal convolution layer has 258; the spatial convolution layers of the other five residual blocks have 576, and the temporal convolution layers have 258; the size of the spatial convolution layer is 1×3×3, and the size of the temporal convolution layer is 3×1×1.
[0049] The fifth non-local spatio-temporal convolution block includes a non-local module and three residual blocks. The residual block includes a spatio-temporal convolution layer, a batch normalization layer, a Leaky ReLU activation function layer, and a spatio-temporal convolution layer connected in sequence. The input and output of the residual block are directly connected, and a Leaky ReLU activation function layer is connected after each residual block. The first spatial convolution layer of the first residual block has 921, and the first temporal convolution layer has 512; the second spatial convolution layer has 1152, and the second temporal convolution layer has 512; the spatial convolution layers of the other two residual blocks have 1152, and the temporal convolution layers have 518; the size of the spatial convolution layer is 1×3×3, and the size of the temporal convolution layer is 3×1×1.
[0050] Among them, the Leaky ReLU activation function is:
[0051]
[0052] Among them, x is the input, and λ is a parameter, generally set to 0.02.
[0053] The digital representation of the non-local module is:
[0054]
[0055] Among them, x i and x j are the eigenvalues at positions i and j of the input signal respectively, Z i is the output eigenvalue at position i of the input signal; f(x i , x j ) is the correlation function of the eigenvalues at positions i and j of the input signal; g(x j ) = W g x j , W g and W Z are weight matrices, C(x) is the normalization parameter, and:
[0056]
[0057]
[0058] Among them, θ(x i ) = W θ x i , Φ(x j ) = W Φ x j , and W θ 、W Φ are weight matrices.
[0059] NST-CNN places the Dropout layer after the fully connected layer and before the Softmax layer, that is, after all BN (batch normalization) layers, in order to avoid the variance shift phenomenon that may occur during the test phase when the Dropout layer is placed before the BN layer; and taking the HMDB51 dataset as an example, when the Dropout dropout rates of the spatial stream sub-network and the temporal stream sub-network are 0.3 and 0.2 respectively, the recognition accuracy of the network is the highest;
[0060] The training of the spatio-temporal convolutional neural network NST-CNN for spatial and temporal streams includes: using the Kinetics-400 dataset as the training dataset to conduct a first training on the spatio-temporal convolutional neural network NST-CNN for spatial and temporal streams; using the UCF101 dataset and the HMDB51 dataset as the training datasets to conduct a second training on the spatio-temporal convolutional neural network NST-CNN for spatial and temporal streams; during the first training and the second training, the improved stochastic gradient descent algorithm with momentum based on the gradient centering algorithm is adopted to optimize the training. The settings of the stochastic gradient descent algorithm are: setting the weight decay to 0.0005, the momentum to 0.9, the initial learning rate to 0.0001, and updating the learning rate with whether the loss decreases as the index, and the learning patience value is 10. And, according to the experimental conditions, when the input frame length of the network is 8, the batch size is 10, and when the input frame length is 16, the batch size is 5.
[0061] The present invention innovatively proposes a human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network. The present invention combines the respective characteristics of the 3D network and the two-stream network, and at the same time introduces a non-local module into the network, and proposes a two-stream non-local spatio-temporal convolutional neural network for human behavior recognition. The two-stream network is composed of a temporal stream subnet and a spatial stream subnet, and the recognition networks of both parts adopt a non-local spatio-temporal convolutional neural network (NST-CNN); the input of the spatial stream subnet is an RGB image sequence extracted from the video frames to be detected; the input of the temporal stream subnet is an optical flow map sequence extracted from the video to be detected by using the optical flow image estimation network PWC-Net; finally, the recognition results of the spatial stream subnet and the temporal stream subnet are fused by using the mean fusion method to realize human behavior recognition. In addition, the network is trained by using the improved stochastic gradient descent algorithm with momentum based on gradient centering. The present invention can realize end-to-end human behavior recognition and has a high accuracy.
[0062] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0063] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network, characterized in that, Including: Obtain a video to be detected containing human behaviors; Extract the video to be detected frame by frame to generate an RGB image sequence, and input the RGB image sequence into the trained spatial stream convolutional neural network NST-CNN to obtain the prediction of the human behavior type in the spatial stream; Input the video to be detected into the trained PWC-Net network to generate an optical flow image sequence, and input the optical flow image sequence into the trained temporal stream convolutional neural network NST-CNN to obtain the prediction of the human behavior type in the temporal stream; Perform mean fusion based on the predictions of the human behavior types in the spatial stream and the temporal stream to obtain the final prediction of the human behavior type; Wherein, both the spatial stream and the temporal stream convolutional neural network NST-CNN include a non-local spatio-temporal convolutional layer, a first non-local spatio-temporal convolutional block, a second non-local spatio-temporal convolutional block, a third non-local spatio-temporal convolutional block, a fourth non-local spatio-temporal convolutional block, a fifth non-local spatio-temporal convolutional block, a 3D pooling layer, a fully connected layer, a Dropout layer, and a Softmax layer connected in sequence; The first non-local spatio-temporal convolutional block includes a spatio-temporal convolutional layer; the second non-local spatio-temporal convolutional block includes a non-local module and three residual blocks; the third non-local spatio-temporal convolutional block includes a non-local module and four residual blocks; the fourth non-local spatio-temporal convolutional block includes a non-local module and six residual blocks; the fifth non-local spatio-temporal convolutional block includes a non-local module and three residual blocks; The digital representation of the non-local module is: where x i and x j are the eigenvalues at positions i and j of the input signal, respectively, and Z i is the output eigenvalue at position i of the input signal; f(x i , x j ) is the correlation function of the eigenvalues at positions i and j of the input signal; g(x j ) = W g x j , where W g and W Z are weight matrices, C(x) is the normalization parameter, and: where, θ(x i ) = W θ x i , Φ(x j ) = W Φ x j , and W θ and W Φ are weight matrices.
2. The human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network according to claim 1, wherein, The step of inputting the video to be detected into the trained PWC-Net network to generate an optical flow image sequence includes: respectively sending the images of adjacent frames of the video to be detected into a six-level feature pyramid network to obtain feature maps at various scales; performing cost volume calculation, warping operation, optical flow extraction layer, upsampling, and context network layer on the feature maps to generate an optical flow image sequence; wherein, the PWC-Net network uses the optical flow datasets Flying Chairs and Flying Things 3D as training datasets for training.
3. A human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network according to claim 1, characterized in that, The residual block includes a spatio-temporal convolutional layer, a batch normalization layer, a Leaky ReLU activation function layer, and a spatio-temporal convolutional layer connected in sequence. There is a direct connection between the input and output of the residual block, and a LeakyReLU activation function layer is connected after each residual block.
4. The human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network according to claim 3, characterized in that The spatio-temporal convolutional layer includes a spatial convolutional layer, a batch normalization layer, a Leaky ReLU activation function layer, and a temporal convolutional layer connected in sequence.
5. A human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network according to claim 4, wherein the Leaky ReLU activation function is: Among them, x is the input and λ is the parameter.
6. The human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network according to claim 3, characterized in that The training of the spatial-temporal convolutional neural network NST-CNN includes: using the Kinetics-400 dataset as the training dataset to conduct a first training on the spatial-temporal convolutional neural network NST-CNN; using the UCF101 dataset and the HMDB51 dataset as the training datasets to conduct a second training on the spatial-temporal convolutional neural network NST-CNN; during the first training and the second training, the training is optimized by adopting the improved stochastic gradient descent algorithm with momentum based on the gradient centering algorithm.
7. A human behavior recognition method based on a two-stream non-local spatio-temporal convolutional neural network according to claim 1, characterized in that, The obtaining of the final prediction of the human behavior type includes: Among them, y average is the final prediction of human behavior type, and x t , x s are the predictions of human behavior type for the spatial stream and the temporal stream respectively.
Citation Information
Patent Citations
Methods for training a CRNN and for semantic segmentation of an inputted video using said crnn
EP3608844A1
Motion recognition method based on feature interactive learning, and terminal device
WO2022073282A1