Action Recognition Method Based on Dynamic Local-Global Graph Convolutional Neural Network
By adopting a dynamic local-global graph convolutional neural network in action recognition, using attention mechanism and self-attention fusion information, and introducing channel attention modules, the problem of unoptimization of graph topology structure and neglecting local spatial relationships in the existing technology is solved, and a higher accuracy of action recognition is achieved.
Patent Information
- Application Number
- CN202210703550.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-06-21
AI Technical Summary
The existing action recognition algorithm based on graph convolution network ignores the connection between non-physical connection nodes, resulting in the graph topology cannot be optimized during training, and the edge weight cannot be updated, which affects the discrimination of action differences. At the same time, these methods ignore subtle connections in local spatial relationships and the importance of different frameworks and channels for action recognition.
A method based on dynamic local-global graph convolution neural network is adopted to dynamically allocate weights to the adjacency matrix under three partitioning strategies through the attention mechanism to generate a learnable transformation matrix. Combined with the improved Transformer self-attention fusion of local and global information, and introduces a channel attention module to enhance the extraction of important feature information.
The accuracy of action recognition is improved, and by dynamically adjusting the adjacency matrix and fusing local and global information, the expression ability of feature modeling and the extraction ability of action-related features are enhanced.
Smart Images

Figure CN114998525B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a graph convolutional network, and in particular to an action recognition method based on a dynamic local-global graph convolutional neural network. Background Art
[0002] Human action recognition is a hot topic in the field of computer vision. In particular, skeleton-based action recognition has received increasing attention. Compared with RGB data, skeleton data is considered to be a more robust representation of human action dynamics. At the same time, skeleton data is very compact in terms of data size. This makes it possible to design more lightweight models. Skeleton data can be easily captured by a depth camera (such as dynamics) or estimated using a human pose estimation algorithm. Among them, action recognition algorithms based on spatiotemporal graph convolution have achieved good results.
[0003] Existing action recognition algorithms based on graph convolutional networks only construct natural connection graphs of the human body, ignoring the connections between non-physically connected joints, which makes it impossible to add new connections to the graph. The topological structure of the graph is determined by the adjacency matrix and the mask. The adjacency matrix is fixed for different action samples, it cannot be optimized during the training process, and the weights of the edges cannot be updated. This affects the discrimination of the differences in actions of different categories. In addition, the previous graph convolutional network used non-local operations to measure the degree of dependence between any two skeleton joints. The graph constructed by this method has certain limitations: it may not be able to capture the subtle connections between two joints in local spatial relationships. In addition, most methods based on graph convolutional networks ignore the different importance of different frames and channels for action recognition. In skeleton sequences, more attention should be paid to representative action feature frames. Summary of the invention
[0004] Purpose of the invention: The purpose of the present invention is to provide an action recognition method based on a dynamic local-global graph convolutional neural network, thereby improving the accuracy of action recognition.
[0005] Technical solution: The action recognition method based on dynamic local-global graph convolutional neural network described in the present invention has the following principle: The innovative content of the present invention is to use the attention mechanism to dynamically assign weights to the adjacency matrices under the three partitioning strategies, and weight the three adjacency matrices to obtain a learnable transformation matrix. Different weight parameters encode different features in the spatial dimension, increasing the expressive power of feature modeling in the skeleton graph; in addition, local and global information are fused by using improved Transformer self-attention. The local operation divides the number of skeleton nodes into blocks and focuses on the most relevant parts of the adjacent nodes in each area. By combining the spatial positions and feature attributes of different adjacent points, appropriate attention weights are assigned to them. The global operation applies self-attention to each frame, independently calculates the correlation between each pair of joints in each frame, and extracts low-level features embedded in the relationship between the various parts of the body; in addition, the introduction of channel attention makes the model pay more attention to important channel features, further improves the performance of the model, and makes the classification prediction results more accurate.
[0006] The present invention comprises the following steps:
[0007] (1) Use the posture estimation algorithm to process the video data into human skeletal structure data. The original skeleton sequence is represented by the three-dimensional coordinates of all human joints in each frame;
[0008] (1.1) For a skeleton sequence with N nodes and T frames, an undirected graph G = (V, E) is constructed on the skeleton sequence; where V = {v ti |t=1,2,…,T,i=1,2,…,N} represents a node set, t represents the frame number, i represents the node, and the feature information of each node is represented by a feature vector composed of spatial coordinates (x, y, z). E is E s and E t The edge set composed of E s Indicates that the joints in the same frame are naturally connected, which is an intra-frame connection; E t It represents the connection between the same joint point in adjacent frames, which is an inter-frame connection;
[0009] (1.2) The NTU+RCB+D dataset is used to define the human body as the three-dimensional coordinates of 25 key joints. While obtaining the spatiotemporal graph, the coordinates of each joint and its confidence are also obtained. These data are stored in a text file for subsequent use.
[0010] (2) Obtain the bone information, node information and adjacency matrix A from step (1); the joint information is a feature vector composed of the spatial coordinates (x, y, z) of each joint point; since each bone is bound to two joints, the joint close to the bone's center of gravity (the center of gravity is on the chest of the human skeleton diagram) is defined as the source joint, and the joint far from the center of gravity is defined as the target joint; each bone represents a vector pointing from its source joint to its target joint, which contains length information and direction information; for example, given a bone v with a source joint 1 =(x 1 ,y 1 ,z 1 ) and its target joint v 2 =(x 2 ,y 2 ,z 2 ), then the bone vector is Because the central joint is not assigned to any bone, the number of joints is one more than the number of bones, so an empty bone with a value of 0 is added to the central joint so that the bone can use the same network as the joint; the adjacency matrix A is a matrix that describes whether points and edges are connected or not, and its value is fixed; the information in this step is used in step (3);
[0011] (3) Build a basic framework for a dynamic local-global graph convolutional neural network with channel attention;
[0012] (3.1) Building a dynamic local-global graph convolution layer: In an end-to-end learning manner, the network topology is optimized together with other parameters of the network. The skeleton graph is unique for different layers and samples, thereby increasing the flexibility of the model; as shown in formula (1):
[0013]
[0014] where f Dynamic GCN (·) represents the dynamic local-global graph convolution output feature map, f in (·) represents the input feature map, represents the dynamic adjacency matrix, B represents the global self-attention matrix, and C represents the local self-attention matrix; || represents the concat operation, and S(·) converts the dynamic adjacency matrix Rearrange and reshape; W V1 and W V2 is the weight of the 1×1 convolution kernel; the three partitioning strategies mentioned above are: 1. The vertex itself; 2. The centripetal subset, which contains the adjacent vertices close to the centroid; 3. The centrifugal subset, which contains the adjacent vertices far from the centroid;
[0015] is a dynamic adjacency matrix with a dimension of B×N×N; it dynamically learns the connection strength between two vertices in the three partitioning strategies from the input feature graph, increasing the flexibility and personalization of the graph structure; specifically, assuming that the input feature graph (B is batch, C is in is the number of channels, T is the number of frames, and N is the number of nodes). First, adaptive average pooling and adaptive maximum pooling are used in parallel to transform the dimension of the input feature map into B×C in ; Then it passes through a fully connected layer to compress the number of channels to C in / 4, and then through an activation function (ReLU) and a fully connected layer to get an f d ∈R B×3 The feature map is normalized to 0-1 through a normalization function softmax, and dynamically matched with the adjacency matrix as the weight; then it is matrix multiplied with the physical adjacency matrix (A) 3×N×N to obtain a dynamic adjacency matrix A of B×N×N d ; Through the above operations, three weights are dynamically assigned to different skeleton graphs to adaptively fuse the adjacency matrices of the three partitions; In addition, in order to connect the multi-level semantic features, A d and the dynamic adjacency matrix of the previous layer Add and average to get the final dynamic adjacency matrix According to formula 3, we can calculate
[0016] f d =softmax(φ(θ(f in ))) (2)
[0017]
[0018] Among them, φ(·) represents linear change, θ(·) performs adaptive pooling and compression operations; A represents the three physical adjacency matrices under the three partitioning strategies, which are related to the feature map f d Fusion is performed in a weighted summation manner;
[0019] B is the global self-attention matrix, which helps the model better model the dynamics of each sample; specifically, given an input feature map First, two 2D convolutional layers are used to transform f in Map and rearrange to reshape and The matrix is then multiplied by a normalization function to obtain a B×N×N similarity matrix B:
[0020] B=softmax((f in W Q1 )(fin W K1 ) T ) (4)
[0021] Where W Q1 , W K1 is the convolution kernel weight of the two convolutional layers;
[0022] C is the local self-attention matrix; the present invention proposes two combined schemes to divide the human skeleton into multiple body parts to extract their different local features: (1) When the human body performs some actions, the amplitude from the trunk to the limbs is different, so the skeleton diagram is divided into three parts; (2) The human body is divided into five parts, including two arms, two legs and the trunk; some actions are completed by several parts of the body; for example, only one or two arms are involved in the waving action, and the rest of the body parts are still. Divide the N skeleton nodes into α blocks according to the above two schemes, pay attention to the spatial relationship between the N / α nodes in each block, and capture more subtle connections; given an input feature map Use 1×1 convolution to reshape it into and The T dimension is moved into the channel dimension, effectively sharing parameters along the time dimension, and is calculated separately on each frame:
[0023] C = softmax((f in W Q2 )(f in W K2 ) T ) (5)
[0024] Where W Q2 , W K2 is the convolution kernel weight of the two convolutional layers;
[0025] (3.2) Build a dynamic local-global graph convolution module: After the dynamic local-global graph convolution layer, there is a batch normalization (BN) layer, an activation function (ReLU) layer and an additional random dropout layer. The dropout rate is set to 0.5, and the output feature map is used in step (3.3);
[0026] (3.3) Build the temporal convolution module: The output of step (3.2) is processed through a standard 2D convolution (TCN) to process the feature information in the time dimension. The temporal convolution (TCN) uses 1×K t The convolution kernel of the input dimension C out ×T×N performs convolution operation in two dimensions, where K t is the number of frames considered within the kernel receptive field; the output f out The dimension size is C out ×Tout ×N feature map; after the temporal convolution, there is a batch normalization (BN) layer, an activation function (ReLU) layer, and an additional random dropout (Dropout) layer, and the Dropout rate is set to 0.5;
[0027] (3.4) After steps (3.2) and (3.3), the extracted spatial dimension features and temporal dimension features are obtained. In order to obtain better action feature representation, a super channel attention module (ECA) for deep CNN is built. It is added after the dynamic local-global graph convolution layer and temporal convolution to recalibrate the channel features to improve the recognition accuracy; given an input feature map X∈R C×T×V , first perform global average pooling on each channel to extract information to obtain a dimension of 1×1×C, then use a fast one-dimensional convolution with a convolution kernel size of 5 to connect each channel and its adjacent channels to obtain cross-channel interaction information and optimize channel information; then use the activation function to get σ(y); σ(y) is the importance of each feature channel, and finally multiply σ(y) by the input X feature map to get the output Out of the channel attention module. The operation is as follows:
[0028] Z=GAP(X) (6) y=Conv(Z) (7)
[0029] Out=X×σ(y) (8)
[0030] Where Conv(·) is a fast one-dimensional convolution function, GAP(·) is a global average pooling, and σ is an activation function, whose function form is
[0031] (3.5) Build a dynamic local-global graph convolution module with channel attention: A basic block is a dynamic local-global graph convolution module, a temporal convolution module, and a channel attention module. In order to stabilize the training, a residual connection is added to each block;
[0032] (3.6) Build a dynamic local-global graph convolutional neural network with channel attention: The dynamic local-global graph convolutional neural network with channel attention is a stack of step (3.5), with a total of 9 modules, and the number of output channels of each module is 64, 64, 64, 128, 128, 128, 256, 256 and 256; add a data BN layer at the beginning to standardize the input data, perform a global average pooling layer (Global MaxPooling) to pool the feature maps of different samples to the same size, and finally pass the SoftMax classifier to obtain the prediction.
[0033] (4) Build a two-stream dynamic local-global graph convolutional neural network model with channel attention and train it to see its effect: input the skeleton information and node information in step (2) as temporal features and spatial features into the dynamic local-global graph convolutional neural network with channel attention built in step (3), obtain the prediction score through the softmax classifier, and then add the two scores to get the final classification result; the final classification score is S, and its expression is shown in formula (9):
[0034] S=W 1 S 1 +W 2 S 2 (9)
[0035] Where S 1 , S 2 Respectively represent the prediction scores of the two sub-networks, ranging from 0 to 1; W 1 and W 2 represents their weights, W 1 +W 2 =1, adjust its value according to the result; the final classification score S is also between 0 and 1;
[0036] (5) Training the model of the present invention: first preprocess the data, reorganize the data structure in the public data set NTU-RGB+D, and input the data of step (2) into step (3); adopt the stochastic gradient descent method (SGD) with Nesterov momentum of 0.9 as the optimization strategy; its batch size (Batch_Size) is 64, the weight decay is 0.0001, and the cross entropy is selected as the loss function to back propagate the gradient, and the number of training times is 64; obtain the final accurate classification result score S.
[0037] A computer storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned action recognition method based on a dynamic local-global graph convolutional neural network.
[0038] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned action recognition method based on a dynamic local-global graph convolutional neural network is implemented.
[0039] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0040] 1. The present invention uses a dynamic adjacency matrix to assign different weights to different sample skeleton graphs under three partitioning strategies, thereby increasing the expressive power of feature modeling. It also uses self-attention to fuse local and global information and focus on the most relevant parts of adjacent nodes. In addition, the channel attention module is used to effectively enhance the ability to extract more important feature information.
[0041] 2. The dynamic local-global graph convolutional neural network with channel attention of the present invention helps to extract features that are more relevant to the action, thereby improving the accuracy of action recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flow chart of the steps of the present invention;
[0043] Figure 2 It is a schematic diagram of space-time diagram;
[0044] Figure 3 A schematic diagram of key joint points defined in the public dataset NTU+RCB+D according to an embodiment of the present invention;
[0045] Figure 4 Schematic diagram of the dynamic local-global graph convolution layer;
[0046] Figure 5 It is a schematic diagram of the dynamic adjacency matrix;
[0047] Figure 6 The block division scheme for the skeleton graph is 1;
[0048] Figure 7 Scheme 2 for dividing the skeleton graph;
[0049] Figure 8 Schematic diagram of the super channel attention module;
[0050] Fig. 9 Schematic diagram of a dynamic local-global graph convolution module with channel attention;
[0051] Fig.10 Schematic diagram of a dynamic local-global graph convolutional network with channel attention;
[0052] Fig.11 Schematic diagram of the model structure of an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0054] An action recognition method based on dynamic local-global graph convolutional neural network, the innovation content is mainly composed of three parts: the first part is to use the attention mechanism to dynamically assign weights to the adjacency matrix under the three partitioning strategies, and weight the three adjacency matrices to obtain a learnable transformation matrix. Different weight parameters encode different features in the spatial dimension, increasing the expressive power of feature modeling in the skeleton graph; the second part uses the improved Transformer self-attention to fuse local and global information; the third part introduces channel attention, so that the model pays more attention to important channel features, further improving the performance of the model and making the classification prediction results more accurate.
[0055] like Figure 1 As shown, the present invention comprises the following steps:
[0056] Step 1: Input a video as a sample to test the algorithm of the present invention. Use the posture estimation algorithm to process the video data into human skeletal structure data. The original skeleton sequence is represented by the three-dimensional coordinates of all human joints in each frame. Since the human skeleton is a topological structure, it is used Figure 2 The spatiotemporal graph models the joints in the skeleton sequence in the time and space dimensions respectively, where the dots represent skeleton joints, the black dot connections represent the first frame skeleton diagram, the gray dot connections represent the second frame skeleton diagram, and the white dot connections represent the third frame skeleton diagram; the gray lines represent the natural physical connections of the human body, and the black lines represent the temporal connections of the same joint point in adjacent frames.
[0057] Step 1.1: For a skeleton sequence with N nodes and T frames, construct an undirected graph G = (V, E) on the skeleton sequence. Where V = {v ti |t=1,2,…,T,i=1,2,…,N} represents a node set, t represents the frame number, i represents the node, and the feature information of each node is represented by a feature vector composed of spatial coordinates (x, y, z). E is E s and E t The edge set composed of E s Indicates that the joints in the same frame are naturally connected, which is an intra-frame connection; E t It represents the connection between the same joint point in adjacent frames, which is an inter-frame connection.
[0058] Step 1.2: Figure 3 The figure shows the key joints of the human body defined by the NTU+RCB+D dataset. It defines the human body as the three-dimensional coordinates of 25 key joints. While obtaining the spatiotemporal graph, the coordinates of each joint point and its confidence and other features are also obtained, and these data are stored in a text file for subsequent use.
[0059] Step 2: Get the bone information, node information and adjacency matrix A from step 1. The joint information is a feature vector consisting of the spatial coordinates (x, y, z) of each joint point. Since each bone is bound to two joints, the joint close to the center of gravity of the bone (the center of gravity is in the chest of the human skeleton diagram) is defined as the source joint, and the joint far from the center of gravity is defined as the target joint. Each bone represents a vector pointing from its source joint to its target joint, which contains length information and direction information. For example, given a bone v with a source joint 1 =(x 1 ,y 1 ,z 1 ) and its target joint v 2 =(x 2 ,y 2 ,z 2 ), then the bone vector is Because the central joint is not assigned to any bone, the number of joints is one more than the number of bones, so an empty bone with a value of 0 is added to the central joint so that the bone can use the same network as the joint. The adjacency matrix A is a matrix that describes whether points and edges are connected or not, and its value is fixed. Use the information in this step for step 3.
[0060] Step 3: Build the basic framework of a dynamic local-global graph convolutional neural network with channel attention. The specific steps are as follows:
[0061] Step 3.1: Build a dynamic local-global graph convolution layer, such as Figure 4 As shown in the figure, it optimizes the network topology together with other parameters of the network in an end-to-end learning manner. The skeleton graph is unique for different layers and samples, which greatly increases the flexibility of the model. Taking node information as an example (the same applies to skeleton information), it is input into the dynamic local-global graph convolution layer, and its input feature map (B is batch, C is in is the number of channels, T is the number of frames, and N is the number of nodes). The first branch will calculate the B is added to get the matrix (B, N, N), and then the dimension is transformed to (B, C in T, N) and then undergo a 2D convolution to change the number of channels to obtain (B, C out ,T,N); the other branch is to By rearranging and reshaping and adding to C, we get Matrix, and then transform it into Multiply, and then change the number of channels through a 2D convolution to get (B, C out ,T,N). Concatenate the two branches and then use a skip connection to get the final output f out Specifically, according to formula (1):
[0062]
[0063] where f Dynamic GCN (·) represents the dynamic local-global graph convolution output feature map, f in (·) represents the input feature map, represents the dynamic adjacency matrix, B represents the global self-attention matrix, and C represents the local self-attention matrix. ||| represents the concat operation, S(·) converts the dynamic adjacency matrix Rearrange and reshape. V1 and W V2 is the weight of the 1×1 convolution kernel. The three partitioning strategies mentioned above are: 1. The vertex itself; 2. The centripetal subset, which contains the adjacent vertices close to the centroid; 3. The centrifugal subset, which contains the adjacent vertices far from the centroid.
[0064] is a dynamic adjacency matrix with a dimension of B×N×N. It dynamically learns the connection strength between two vertices in the three partitioning strategies from the input feature graph, increasing the flexibility and personalization of the graph structure. Figure 5 Shown Specifically, assuming that the input feature map (B is batch, C is in is the number of channels, T is the number of frames, and N is the number of nodes). First, adaptive average pooling and adaptive maximum pooling are used in parallel to transform the dimension of the input feature map into B×C in Then it passes through a fully connected layer to compress the number of channels to C in / 4, and then through an activation function (ReLU) and a fully connected layer to get an f d ∈R B×3 The feature map is normalized to 0-1 through a normalization function softmax, and used as a weight to dynamically match the adjacency matrix. Then it is matrix multiplied with the physical adjacency matrix (A) 3×N×N to obtain a dynamic adjacency matrix A of B×N×N d Through the above operations, we dynamically assign three weights to different skeleton graphs to adaptively fuse the adjacency matrices of the three partitions. In addition, in order to connect the multi-level semantic features, we will A d and the dynamic adjacency matrix of the previous layer Add and average to get the final dynamic adjacency matrix According to formula 3, we can calculate
[0065] f d =softmax(φ(θ(f in ))) (2)
[0066]
[0067] Among them, φ(·) represents linear change, and θ(·) performs adaptive pooling and compression operations. A represents the three physical adjacency matrices under the three partitioning strategies, which are related to the feature map f d The fusion is performed in a weighted summation manner.
[0068] B is the global self-attention matrix, which helps the model better model the dynamics of each sample. Specifically, given an input feature map First, two 2D convolutional layers are used to transform f in Map and rearrange to reshape and The matrix is then multiplied by a normalization function to obtain a B×N×N similarity matrix B:
[0069] B=softmax((f in W Q1 )(f in W K1 ) T ) (4)
[0070] Where W Q1 , W K1 are the convolution kernel weights of the two convolutional layers.
[0071] C is the local self-attention matrix. The present invention proposes two combined schemes for dividing the human skeleton into multiple body parts to extract different local features: (1) When the human body is doing some actions, the amplitude from the trunk to the limbs is different. Therefore, Figure 6 As shown in the figure, we divide the skeleton diagram into three parts; (2) Figure 7 As shown in Figure 1, we divide the human body into five parts, including two arms, two legs, and the torso. Some actions are performed by several parts of the body. For example, only one or two arms are involved in the waving action, and the rest of the body is stationary. We divide the N skeleton nodes into α blocks according to the above two schemes, focusing on the spatial relationship between the N / α nodes in each block to capture more subtle connections. Given an input feature map Use 1×1 convolution to reshape it into and The T dimension is moved into the channel dimension, effectively sharing parameters along the time dimension and calculating them separately on each frame.
[0072] C = softmax((f in W Q2 )(f in W K2 ) T ) (5)
[0073] Where W Q2 , W K2 are the convolution kernel weights of the two convolutional layers.
[0074] Step 3.2: Build a dynamic local-global graph convolution module. After the dynamic local-global graph convolution layer, there is a batch normalization (BN) layer, an activation function (ReLU) layer, and an additional random dropout layer with a dropout rate of 0.5. The output feature map is used in step 3.3.
[0075] Step 3.3: Build a temporal convolution module. The output of step 3.2 is processed through a standard 2D convolution (TCN) to process the feature information in the time dimension. The temporal convolution (TCN) uses 1×K t The convolution kernel of the input dimension C out ×T×N performs convolution operation in two dimensions, where K t is the number of frames considered within the kernel receptive field. The output f is obtained out The dimension size is C out ×T out × N feature maps. After the temporal convolution, there is a batch normalization (BN) layer, an activation function (ReLU) layer, and an additional random dropout layer (Dropout rate) set to 0.5.
[0076] Step 3.4: After step 3.2 and step 3.3, the extracted spatial dimension features and temporal dimension features are obtained. In order to obtain better action feature representation, the present invention builds a super channel attention module (ECA) for deep CNN, which is added after the dynamic local-global graph convolution layer and temporal convolution to recalibrate the channel features to improve recognition accuracy. The structure of this module is as follows Figure 8 As shown. Given an input feature map X∈R C×T×V First, perform global average pooling on each channel to extract information to obtain a dimension of 1×1×C. Then use a fast one-dimensional convolution with a convolution kernel size of 5 to connect each channel and its adjacent channels to obtain cross-channel interaction information and optimize channel information. Then use the activation function to get σ(y). σ(y) is the importance of each feature channel. Finally, multiply σ(y) by the input X feature map to get the output Out of the channel attention module. The operation is as follows:
[0077] Z=GAP(X) (6) y=Conv(Z) (7)
[0078] Out=X×σ(y) (8)
[0079] Where Conv(·) is a fast one-dimensional convolution function, GAP(·) is a global average pooling, and σ is an activation function, whose function form is
[0080] Step 3.5: Build a dynamic local-global graph convolution module with channel attention. Fig. 9 As shown in Figure 1, a basic block is a dynamic local-global graph convolution module, a temporal convolution module, and a channel attention module. In order to stabilize the training, a residual connection is added to each block.
[0081] Step 3.6: Build a dynamic local-global graph convolutional neural network with channel attention. The dynamic local-global graph convolutional neural network with channel attention is a stack of step 3.5, such as Fig.10 As shown, there are 9 modules in total, and the number of output channels of each module is 64, 64, 64, 128, 128, 128, 256, 256 and 256. A data BN layer is added at the beginning to standardize the input data. The input data performs each operation of the adaptive graph convolution module with attention mechanism in step 3.5, and then performs a global average pooling layer (Global MaxPooling) to pool the feature maps of different samples to the same size. The final output passes through the SoftMax classifier to obtain the prediction. The index corresponding to the number with the highest score is taken from the 60 action categories, and its corresponding category is the action category predicted by the network.
[0082] Step 4: Build a two-stream dynamic local-global graph convolutional neural network model with channel attention and train it to see its effect. Fig.11 The model proposed by the present invention is shown in FIG. The present invention inputs the skeleton information and node information in step 2 as temporal features and spatial features into the dynamic local-global graph convolutional neural network with channel attention built in step 3, obtains the prediction score through the softmax classifier, and then adds the two scores to obtain the final classification result. The final classification score is S, and its expression is shown in formula (9):
[0083] S=W 1 S 1 +W 2 S 2 (9)
[0084] Where S 1 , S 2 Respectively represent the prediction scores of the two sub-networks, ranging from 0 to 1. 1 and W 2 represents their weights, W 1 +W 2=1, and its value can be adjusted according to the result. The final classification score S is also between 0 and 1.
[0085] Step 5: Train the model of the present invention. First, preprocess the data, reorganize the data structure in the public data set NTU-RGB+D, and input the data of step 2 into step 3. The model of the present invention uses stochastic gradient descent (SGD) with Nesterov momentum of 0.9 as the optimization strategy. The batch size (Batch_Size) is 64, the weight decay is 0.0001, and the cross entropy is selected as the loss function to back propagate the gradient. The number of training times is 64. The final accurate classification result score S is obtained.
[0086] By implementing the model of this patent, the final classification accuracy can be improved. The dynamic adjacency matrix is used to assign different weights to different sample skeleton graphs under the three partitioning strategies, which increases the expressiveness of feature modeling; and by fusing local and global information, the most relevant parts of adjacent nodes are focused on; in addition, the channel attention module effectively enhances the ability to extract more important feature information. The dynamic local-global graph convolutional neural network with channel attention helps extract features that are more relevant to the action, thereby improving the accuracy of action recognition.
Claims
1. An action recognition method based on a dynamic local-global graph convolutional neural network, characterized in that: The following steps are involved: (1) Use the posture estimation algorithm to process the video data into human skeletal structure data. The original skeleton sequence is represented by the three-dimensional coordinates of all human joints in each frame; (2) Obtain skeleton information, node information and adjacency matrix A from step (1); The joint information is a feature vector composed of the spatial coordinates (x, y, z) of each joint point; since each bone is bound to two joints, the joint close to the center of gravity of the bone is defined as the source joint, and the joint far from the center of gravity is defined as the target joint; each bone represents a vector pointing from its source joint to its target joint, and the vector contains length information and direction information; because the central joint is not assigned to any bone, the number of joints is one more than the number of bones, so an empty bone with a value of 0 is added to the central joint so that the bones can use the same network as the joints; the adjacency matrix A is a matrix that describes whether points and edges are connected, and its value is fixed; the information of this step is used in step (3); (3) Build a basic framework for a dynamic local-global graph convolutional neural network with channel attention; (4) Build a two-stream dynamic local-global graph convolutional neural network model with channel attention and train it to see its effect: input the skeleton information and node information in step (2) as temporal features and spatial features into the dynamic local-global graph convolutional neural network with channel attention built in step (3), obtain the prediction score through the softmax classifier, and then add the two scores to get the final classification result; the final classification score is S, and its expression is shown in formula (9): S=W1S1+W2S2 (9) Where S1 and S2 represent the prediction scores of the two sub-networks, ranging from 0 to 1; W1 and W2 represent their weights, W1+W2=1, and their values are adjusted according to the results; the final classification score S is also between 0 and 1; (5) Training the model of the present invention: first preprocess the data, reorganize the data structure in the public data set NTU-RGB+D, and input the data of step (2) into step (3); adopt the stochastic gradient descent method with Nesterov momentum of 0.9 as the optimization strategy; the batch size is 64, the weight decay is 0.0001, and the cross entropy is selected as the loss function to back propagate the gradient, and the number of training times is 64; obtain the final accurate classification result score S.
2. The method for action recognition based on a dynamic local-global graph convolutional neural network according to claim 1, characterized in that: The step (1) is specifically: (1.1) For a skeleton sequence with N nodes and T frames, an undirected graph G = (V, E) is constructed on the skeleton sequence; where V = {v ti |t=1,2,…,T,i=1,2,…,N} represents a node set, t represents the frame number, i represents the node, and the feature information of each node is represented by a feature vector composed of spatial coordinates (x,y,z). E is E s and E t The edge set composed of E s Indicates that joints on the same frame are naturally connected, which is an intra-frame connection; E t It represents the connection between the same joint point in adjacent frames, which is an inter-frame connection; (1.2) The NTU+RCB+D dataset is used to define the human body as the three-dimensional coordinates of 25 key joints. While obtaining the spatiotemporal graph, the coordinates of each joint and its confidence are also obtained. These data are stored in a text file for subsequent use.
3. The method for action recognition based on dynamic local-global graph convolutional neural network according to claim 1, characterized in that: The step (3) is specifically: (3.1) Building a dynamic local-global graph convolution layer: In an end-to-end learning manner, the network topology is optimized together with other parameters of the network. The skeleton graph is unique for different layers and samples, thereby increasing the flexibility of the model; as shown in formula (1): where f Dynamic GCN (·) represents the dynamic local-global graph convolution output feature map, f in (·) represents the input feature map, represents the dynamic adjacency matrix, B represents the global self-attention matrix, and C represents the local self-attention matrix; || represents the concat operation, and S(·) converts the dynamic adjacency matrix Rearrange and reshape; W V1 and W V2 is the weight of the 1×1 convolution kernel; the three partitioning strategies are:
1. The vertex itself; 2. The centripetal subset, which contains the adjacent vertices close to the centroid; 3. The centrifugal subset, which contains the adjacent vertices far from the centroid; is a dynamic adjacency matrix with a dimension of B×N×N; it dynamically learns the connection strength between two vertices in the three partitioning strategies from the input feature graph, increasing the flexibility and personalization of the graph structure; specifically, assuming that the input feature graph First, adaptive average pooling and adaptive maximum pooling are used in parallel to transform the dimension of the input feature map into B×C in ; Then it passes through a fully connected layer to compress the number of channels to C in / 4, and then through an activation function and a fully connected layer to get f d ∈R B×3 The feature map is normalized to 0-1 through a normalization function softmax, and dynamically matched with the adjacency matrix as the weight; then it is matrix multiplied with the physical adjacency matrix (A) 3×N×N to obtain a dynamic adjacency matrix A of B×N×N d ; Through the above operations, three weights are dynamically assigned to different skeleton graphs to adaptively fuse the adjacency matrices of the three partitions; In addition, in order to connect the multi-level semantic features, A d and the dynamic adjacency matrix of the previous layer Add and average to get the final dynamic adjacency matrix According to formula 3, we can calculate f d =softmax(φ(θ(f in ))) (2) Among them, φ(·) represents linear change, θ(·) performs adaptive pooling and compression operations; A represents the three physical adjacency matrices under the three partitioning strategies, which are related to the feature map f d Fusion is performed in a weighted summation manner; B is the global self-attention matrix, which helps the model better model the dynamics of each sample; specifically, given an input feature map First, two 2D convolutional layers are used to transform f in Map and rearrange to reshape and The matrix is then multiplied by a normalization function to obtain a B×N×N similarity matrix B: B=softmax((f in W Q1 )(f in W K1 ) T ) (4) Where W Q1 , W K1 is the convolution kernel weight of the two convolutional layers; C is the local self-attention matrix; the present invention proposes two combined schemes for dividing the human skeleton into multiple body parts to extract different local features: (1) when the human body performs some actions, the amplitude from the trunk to the limbs is different, therefore, the skeleton graph is divided into three parts; (2) the human body is divided into five parts, including two arms, two legs and the trunk; some actions are performed by several parts of the body; N skeleton nodes are divided into α blocks according to the above two schemes, focusing on the spatial relationship between N / α nodes in each block to capture more subtle connections; given an input feature map Use 1×1 convolution to reshape it into and The T dimension is moved into the channel dimension, effectively sharing parameters along the time dimension, and is calculated separately on each frame: C=softmax((f in W Q2 )(f in W K2 ) T ) (5) Where W Q2 , W K2 is the convolution kernel weight of the two convolutional layers; (3.2) Build a dynamic local-global graph convolution module: After the dynamic local-global graph convolution layer, there is a batch normalization layer, an activation function layer and an additional random dropout layer. The dropout rate is set to 0.5, and the output feature map is used in step (3.3); (3.3) Build the temporal convolution module: The output of step (3.2) is processed through a standard 2D convolution to process the feature information in the temporal dimension. The temporal convolution uses 1×K t The convolution kernel of the input dimension C out ×T×N performs convolution operation in two dimensions, where K t is the number of frames considered within the kernel receptive field; the output f out The dimension size is C out ×T out ×N feature map; after the temporal convolution, there is a batch normalization layer, an activation function layer and an additional random dropout layer, and the dropout rate is set to 0.5; (3.4) After steps (3.2) and (3.3), the extracted spatial dimension features and temporal dimension features are obtained. In order to obtain better action feature representation, a super channel attention module for deep CNN is built, which is added after the dynamic local-global graph convolution layer and temporal convolution to recalibrate the channel features to improve recognition accuracy; given an input feature map X∈R C×T×V , first perform global average pooling on each channel to extract information to obtain a dimension of 1×1×C, then use a fast one-dimensional convolution with a convolution kernel size of 5 to connect each channel and its adjacent channels to obtain cross-channel interaction information and optimize channel information; then use the activation function to get σ(y); σ(y) is the importance of each feature channel, and finally multiply σ(y) by the input X feature map to get the output Out of the channel attention module. The operation is as follows: Z=GAP(X) (6) y=Conv(Z) (7) Out=X×σ(y) (8) Where Conv(·) is a fast one-dimensional convolution function, GAP(·) is a global average pooling, and σ is an activation function, whose function form is (3.5) Build a dynamic local-global graph convolution module with channel attention: A basic block is a dynamic local-global graph convolution module, a temporal convolution module, and a channel attention module. In order to stabilize the training, a residual connection is added to each block; (3.6) Build a dynamic local-global graph convolutional neural network with channel attention: The dynamic local-global graph convolutional neural network with channel attention is a stack of step (3.5), with a total of 9 modules, and the number of output channels of each module is 64, 64, 64, 128, 128, 128, 256, 256 and 256; a data BN layer is added at the beginning to standardize the input data, and a global average pooling layer is performed to pool the feature maps of different samples to the same size, and the final output is passed through a SoftMax classifier to obtain a prediction.
4. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, an action recognition method based on a dynamic local-global graph convolutional neural network is implemented as described in any one of claims 1 to 3.
5. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements an action recognition method based on a dynamic local-global graph convolutional neural network as described in any one of claims 1-3.
Citation Information
Patent Citations
Human skeleton action recognition method and system and medium
CN110490035A
Convolutional neural network non-local information construction method
CN112329801A