Action recognition method based on multi-branch three-dimensional graph convolution and LSTM
By designing a three-dimensional graph convolution MBA_3DGCN with multi-branch attention in action recognition and combining an LSTM network, the problem of difficulty in extracting the sequence space and timing characteristics of the human body skeleton in action recognition is solved, and a higher accuracy of action recognition is achieved.
Patent Information
- Application Number
- CN202510569924.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The prior art is difficult to extract spatial and temporal features in human skeleton sequences at the same time, resulting in insufficient accuracy in action recognition.
A three-dimensional graph convolution MBA_3DGCN with multi-branch attention is designed, combined with the LSTM network, which is used to simultaneously extract the spatial and local timing characteristics of the skeleton sequence, and extract the global timing variation characteristics through LSTM.
It improves the graph convolution feature extraction ability, enhances the anti-interference, reduces the interference of single-frame error data, and significantly improves the accuracy of action recognition.
Smart Images

Figure CN120088863A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning and human action recognition, and particularly relates to an action recognition method based on multi-branch three-dimensional graph convolution and LSTM. Background Art
[0002] Human action recognition is an advanced and abstract visual perception task, belonging to the category of pattern recognition, aiming to understand and judge human behavior activities. Action recognition has a wide range of applications in all walks of life and various fields. According to different input information for action recognition, it is mainly divided into two types of methods: based on images / videos and based on skeleton sequences.
[0003] With the rapid development of human skeleton pose estimation algorithms, such as the emergence of OpenPose, DeepPose, AlphaPose, etc., it has become easier and more feasible to directly obtain human skeleton sequence data in real time from videos. Therefore, action recognition based on skeleton data has received more attention. Currently, directly using a deep network model to extract the temporal features contained in the skeleton sequence to achieve end-to-end action classification and recognition has become the mainstream method.
[0004] Therefore, how to extract the spatial pose features and temporal dynamic change features contained in the skeleton sequence is the key to action recognition based on skeleton sequences. For the extraction of dynamic temporal features between frames, the currently mainly used methods are models based on the long short-term memory (LSTM) with variable length of the recursive neural network (RNN) and the temporal convolutional network TCN. For the extraction of spatial features of nodes within a frame, currently, the GCN network model is mostly used to extract the spatial dependence features of human joints. However, graph convolution often only focuses on local connection features and cannot pay attention to the spatial dependence relationship between non-adjacent nodes.
[0005] Currently, the existing patent applications related to action recognition based on skeleton sequences, such as the related patent contents disclosed in Patent Document 1, Patent Document 2, etc., are usually single-frame graph convolutions, that is, they can only extract the skeleton spatial information of a single frame, and the extraction of temporal information is then implemented by using a TCN module.
[0006] For the extraction of temporal features and spatio-temporal features, they can also be extracted simultaneously instead of separately. Currently, the relevant literatures disclosed in the prior art have not involved designing a three-dimensional graph convolution operation to simultaneously extract the spatial features and temporal features of adjacent multiple frames, and the skeleton motion information of consecutive multiple frames is an important information source for action recognition.
[0007] Related Technical Literature Patent Document 1 Chinese Invention Patent Application, Publication No.: CN 113688765 A, Publication Date: November 23, 2021; Patent Document 2 Chinese invention patent application, publication number: CN 114998525 A, publication date: September 2, 2022. Summary of the Invention
[0008] The purpose of the present invention is to propose an action recognition method based on multi-branch 3D graph convolution and LSTM. It designs a multi-branch attention 3D graph convolution MBA_3DGCN for the skeleton sequence of adjacent multiple frames to simultaneously extract the spatial and local temporal features of the skeleton sequence, and uses the LSTM network to extract the entire temporal change features of the skeleton sequence. Through the combination of MBA_3DGCN and LSTM, it is beneficial to improve the accuracy of action recognition.
[0009] In order to achieve the above purpose, the present invention adopts the following technical solutions: An action recognition method based on multi-branch 3D graph convolution and LSTM, comprising the following steps: Step 1. Preprocess the input human body three-dimensional pose sequence to obtain node flow information data, including node position information, node movement speed information, and node movement acceleration information, and fuse the three kinds of information. Step 2. For the skeleton sequence of adjacent multiple frames, design a multi-branch attention 3D graph convolution MBA_3DGCN to extract features for the nodes of the current frame and adjacent frames, the nodes connected inward, the nodes connected outward, and all other nodes respectively, and perform feature aggregation through corresponding learnable attention matrices to achieve graph convolution operations. Step 3. Use MBA_3DGCN and LSTM to build an action recognition model based on the skeleton sequence. The action recognition model first extracts the spatial and local temporal features of the skeleton sequence in the MBA_3DGCN module, and then uses the LSTM network to extract the entire temporal change features of the skeleton sequence for action recognition. Step 4. Use the node flow information data as the input of the action recognition model, train the built action recognition model, and use the trained action recognition model for action recognition to obtain the final recognition result.
[0010] The present invention has the following advantages: As described above, the present invention relates to an action recognition method based on multi-branch three-dimensional graph convolution and LSTM. For the skeleton sequences of adjacent multiple frames, the method designs a three-dimensional graph convolution MBA_3DGCN based on local temporal multi-branch attention, and combines it with a global temporal LSTM network to build a model structure for human action recognition. The method of the present invention realizes spatio-temporally unified three-dimensional graph convolution. The present invention first uses MBA_3DGCN to simultaneously extract the spatial and local multi-frame temporal features of the skeleton sequence, improves the graph convolution feature extraction ability while increasing the anti-interference ability, and reduces the interference of single-frame error data. On this basis, the LSTM network is used to extract the entire temporal change features of the skeleton sequence for action recognition, giving full play to the ability of LSTM to extract long temporal features. The combination of MBA_3DGCN and LSTM significantly improves the accuracy of action recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a network structure diagram of the action recognition method based on multi-branch three-dimensional graph convolution and LSTM in an embodiment of the present invention; Figure 2 It is a human body structure diagram in an embodiment of the present invention; wherein, Figure 2 (a) is a schematic diagram of joint points of the NTU_D skeleton data set, Figure 2 and (b) in it is a schematic diagram of body parts formed by adjacent nodes; Figure 3 It is a network structure diagram of the multi-branch attention three-dimensional graph convolution MBA_3DGCN in an embodiment of the present invention; Figure 4 It is a network structure diagram of LSTM in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] The present invention will be further described in detail below with reference to the drawings and specific embodiments: Embodiment 1 This embodiment proposes an action recognition method based on multi-branch 3D graph convolution and LSTM. Considering the fact that action recognition is to recognize the continuous change information of multiple joints of the human body's skeleton, and at the same time, due to the natural graph structure of the human body, graph convolution is considered to extract features. Between adjacent consecutive frames of the skeleton sequence, it can not only represent the human body's posture, but also contain the change information of the posture, which is important information for action recognition. Therefore, the present invention designs a multi-branch attention 3D graph convolution (Multi Branch Attention 3 Dimension GraphConvolution) operation to simultaneously extract the spatial features and local temporal features in the human skeleton sequence. The multi-branch attention structure uses multiple transformation matrices to extract information contributed to different nodes. The global attention therein breaks through the limitation that graph convolution only focuses on local features. At the same time, the features of each node not only focus on the information of other nodes in this frame, but also focus on the information of nodes in the previous frame and the next frame, that is, the information of consecutive frame nodes. Therefore, the spatial and temporal transformation features of the skeleton sequence can be simultaneously extracted, improving the feature extraction ability of graph convolution. Due to the problem of computational complexity, the length of the time frame of the 3D graph convolution MBA_3DGCN is generally limited, and it mainly extracts local temporal transformation features. Therefore, the present invention also uses MBA_3DGCN and LSTM to build an action recognition model based on the skeleton sequence. First, the spatial and local temporal features of the skeleton extracted by the MBA_3DGCN graph convolution are used, then the global skeleton change features are extracted by LSTM, and finally, action classification recognition is performed through a fully connected layer to output the action recognition result.
[0013] As Figure 1 shown, the action recognition method based on multi-branch 3D graph convolution and LSTM in this embodiment includes the following steps: Step 1. Preprocess the input three-dimensional human body posture sequence to obtain node stream information data, including node position information, node motion speed information, and node motion acceleration information, and fuse the three kinds of information.
[0014] The input three-dimensional human body posture sequence is represented as .
[0015] where N represents the number of joint points of the human body posture information, T is the length of the skeleton sequence, that is, the number of key frames, D is the dimension of the input joint point information, and for a three-dimensional joint point sequence, D = 3.
[0016] Since the dataset used in this embodiment is the NTU_D dataset, therefore, D = 3, T = the number of key frames (the longest is 300 frames, the shortest is 32 frames, and the average length is 82.9 frames), N = 25.
[0017] Joint point set The data is obtained after centering and normalization preprocessing .
[0018] Among them, the centering process uses the selected waist joint coordinates Figure 2 In (a) among them, the node numbered 21 is used as the center to calculate the position differences of each node coordinate relative to , and then perform normalization processing. The formula is as follows: .
[0019] The node motion speed information, which is the first-order information of the node position information, is expressed as: .
[0020] The node motion acceleration information, which is the second-order information of the node position information, is expressed as: .
[0021] Fuse the information of all joint points as the input information of the node branch, and the formula is expressed as: .
[0022] Herein is the fusion of the above three kinds of information (position information, speed information and acceleration information), and is used as the input of the action recognition model based on multi-branch three-dimensional graph convolution and LSTM built in step 3 below.
[0023] Step 2. For the skeleton sequences of adjacent multiple frames, design a multi-branch attention three-dimensional graph convolution (MultiBranch Attention 3 Dimension Graph Convolution, MBA_3DGCN).
[0024] This graph convolution adopts a multi-branch structure (MB, Multi-Branch, specifically 4 branches), and extracts features respectively for the nodes of the current frame and adjacent frames provided to its own nodes, inward-connected nodes, outward-connected nodes and all other nodes, and aggregates the features through corresponding learnable attention matrices (A, attention) to implement the graph convolution operation.
[0025] As Figure 3 shown, the design process of this multi-branch attention three-dimensional graph convolution MBA_3DGCN is as follows: Step 2.1. Construct a multi-branch attention graph convolution for a single frame.
[0026] For the single-frame node information, a feature transformation matrix with four branches is set, namely the self-node feature transformation matrix , the inward-connected adjacent node feature transformation matrix , the outward-connected adjacent node feature transformation matrix and the global other node feature transformation matrix ; where , is the dimension of the input feature , is the dimension of the output feature . Among them is used to extract the self-node features, , , are used to extract the global other node features. The multi-branch local feature extracted by adding the attention mechanism is represented by the following graph convolution formula: (1) where represents the attention matrix for focusing on self-features, i.e., the self-attention matrix, which is initialized with the self-link adjacency matrix, i.e., the identity matrix ; represents the attention matrix for focusing on centripetal adjacent nodes, i.e., the inward attention matrix, which is initialized with the centripetal adjacency matrix , is the matrix representation of the directed graph of the human body skeleton's connection from outside to inside; represents the attention matrix for focusing on centrifugal adjacent nodes, i.e., the outward attention matrix, which is initialized with the centrifugal adjacency matrix , is the matrix representation of the directed graph of the human body skeleton's connection from inside to outside.
[0027] The weights therein are adjusted through adaptive learning. These matrices are only initialized with the adjacency matrix, but the learned parameters are not limited to the edges defined by these adjacency matrices, increasing the flexibility of the local attention matrix.
[0028] For the single-frame node information, node features are extracted from a global perspective, and a global feature transformation matrix of the single frame is set, and then max pooling MaxPool is performed on the input node information, with the set stride being 2.
[0029] For a dataset where the number of nodes is not a multiple of 2, such as the NTU_D dataset with 25 skeleton nodes, the positions of its 25 nodes are as shown in Figure 2 (a). In this embodiment, the following processing is performed: Duplicate the centroid node 21 nodes to obtain 26 nodes. MaxPool extracts the features of a certain part of the body, and the part is as shown in Figure 2 (b) in, obtaining the features of 13 body parts, and then through the attention matrix to focus on the features of these 13 body parts, and the global feature with the attention mechanism The graph convolution formula extracted is expressed as: (2) Then combine the multi-branch local features and the global features to obtain the output of the multi-branch attention graph convolution of a single frame as: (3).
[0030] Step 2.2. Expand the multi-branch attention graph convolution of a single frame obtained in Step 2.1 to the time domain to obtain the expression form of the output of the t-th frame of the three-dimensional graph convolution of this multi-branch attention.
[0031] Specifically, expand the multi-branch attention graph convolution of a single frame shown in formula (3) to multiple consecutive frames adjacent to the t-th moment, that is, expand 1 frame to multiple consecutive frames at the t-th moment. The process is as follows: Let the convolution time kernel length be , indicating that the feature extraction window span for capturing time-dependent relationships is , for the consecutive frames adjacent to the t-th moment, are required for each frame , , , attention matrices and transformation matrices for each frame.
[0032] Then the output of the t-th frame of the three-dimensional graph convolution of this multi-branch attention is expressed as: (4) where represents time, [ ] is the rounding operation, represents the input information of the node branch , the superscript in in represents input, and the subscript in represents short-term time series;
[0033] Step 2.3. Simplify the expression form of the output of the t-th frame of the 3D graph convolution in Step 2.2.
[0034] Since it is considered that regardless of which frame, the aggregated features from adjacent nodes include temporal features and spatial location features, the same feature transformation matrix can be used for transformation.
[0035] Simplify formula (4). For each branch, regardless of which frame in the time series, that is, for the four different branches, the continuous multi-frame action features at the t-th moment of each branch at any time are respectively set with four shared feature transformation matrices. Then the simplified 3D graph convolution formula is expressed as: (5).
[0036] In this way, at the t-th moment, after each branch aggregates the features of the previous frames and then performs unified feature transformation, that is, all frames of each branch share a W feature transformation matrix, significantly reducing the number of parameters of the 3D graph convolution.
[0037] Step 2.4. Optimize the operation of the 3D graph convolution through the normalization of the attention matrix.
[0038] To improve the trainability of the attention matrix, for the feature aggregation of all frames of each branch in formula (5) , denoted as , , , , which have the same operation, the following optimization is performed:
[0039] (1) Expand the feature aggregation, and equivalently express it as: (6)
[0040] (2) Perform row normalization on the concatenated attention matrix operation, and express it as: (7) The operation uses the softmax function and defines formula (7) as the Δ operation, then rewrite it as: (8) In formula (8), i is the local time series parameter, from the first frame when i = 0 to the last frame when Substitute formula (8) into formula (5) to get: (9) Write the formula (9) in the form of a separated formula: (10).
[0041] The proposed multi-branch attention-based adaptive 3D graph convolution MBA_3DGCN of the present invention (as shown in formula (10)) takes into account the global features of the space of the body structure and the local characteristics of the time series. The attention matrix is obtained through adaptive learning, and the number of parameters is saved by sharing the feature transformation matrix. It can not only achieve spatial aggregation but also achieve the aggregation of nodes between different frames (temporally consecutive frames), that is, temporal local feature aggregation. Therefore, it has higher anti-interference ability.
[0042] Step 3. Use MBA_3DGCN and LSTM to build an action recognition model based on the skeleton sequence. The action recognition model built in this embodiment includes MBA_3DGCN, LSTM, as well as a fully connected layer FC and Softmax.
[0043] This action recognition model first extracts the spatial and local temporal features of the skeleton sequence in the MBA_3DGCN module, and then uses the LSTM network to extract the overall temporal change features of the skeleton sequence.
[0044] After the global temporal feature extraction by LSTM, the recognition of human actions is realized through the fully connected layer FC and Softmax.
[0045] The network structure of LSTM realizes the modeling of sequence data by controlling the transmission and forgetting of information, as Figure 4 shown. F t out , C t , h t are the input, cell memory state, and output vector of LSTM at time t, respectively. f t , i t , o t are the forget gate, input gate, and output gate of the LSTM cell, respectively. These gates have fully connected layers with the activation function of σ (the activation function uses softmax).
[0046] The formula of the LSTM module is expressed as: ; ; ; ; ; .
[0047] where W is the input weight and U is the weight of the previous cell state. The local feature vectors generated by MBA_3DGCN are used as the input of the LSTM network, and the global temporal features are extracted through the LSTM.
[0048] Step 4. Use the node flow information data as the input of the action recognition model, train the built action recognition model, and use the trained action recognition model for action recognition to obtain the final recognition result.
[0049] The process of training the built action recognition model is as follows: Step 4.1. Obtain the standard human action recognition NTU_D dataset.
[0050] Step 4.2. Obtain the human pose joint point sequence data from the NTU_D dataset and perform preprocessing, that is, perform joint sequence data under centering and normalization parameters, and calculate the motion speed information and motion acceleration information of the joint points.
[0051] Then fuse to obtain the data input to the action recognition model , that is, the three-dimensional node data sequence.
[0052] Step 4.3. Use the three-dimensional node data sequence obtained in Step 4.2 as the input of the action recognition model.
[0053] According to the adjacency matrix representing the connection of the human skeleton Initialize , and , and randomly initialize K learnable global attention matrices . Initialize all the weights W of the LSTM with the uniform function and randomly initialize the value of h. Use the NTU_D dataset to pre-train the entire action recognition model, with a batch size of 64, a weight decay set to 0.0001, and a learning rate set to 0.1; finally obtain the trained model parameters.
[0054] The loss function uses the classical classification loss function, the negative log-likelihood loss with label smoothing .
[0055] The formula of is expressed as:
[0056] where, is the number of all samples, is the number of labels, is the smoothing exponent, taking , is the true label for the i-th sample , is the probability value of the sample after passing through the model and being predicted to belong to the label.
[0057] The training output of MBA_3DGCN is obtained as the input of the LSTM network for joint training, and finally the trained model parameters are obtained for 3D pose estimation. After obtaining the trained action recognition model, the model is actually deployed.
[0058] After the model is deployed, the general process of action recognition is as follows: First, the human body bone sequence is obtained through a human pose estimation algorithm or a depth camera; then the bone sequence data is preprocessed according to the preprocessing process in step 1; finally, the skeleton sequence is input into the action recognition model built by MBA_3DGCN and LSTM for human action recognition.
[0059] To verify the performance of the proposed MBA_3DGCN+LSTM action recognition model of the present invention, it is verified on the NTU_D dataset. The MBA_3DGCN+LSTM model designed by the present invention is compared with ST-GCN and agcn of the baseline method. To better compare its performance, the present invention all adopts the single-stream 1s method, that is, only using the node information of the skeleton data and not using its bone information. The results of the comparison method and the method of the present invention are shown in Table 1.
[0060] Table 1 Experimental results of cross_subject (CS) evaluation and cross_view (CV) evaluation on the NTU_D dataset
[0061] It can be easily seen from Table 1 above that the performance of the action recognition model based on MBA_3DGCN and LSTM designed by the present invention has been significantly improved. The effect of all using 3D multi-branch GCN is worse than that of 2D multi-branch GCN, probably because in the later stage of the model, after the features are extracted and fused in the previous stages, the graph structure between adjacent frames becomes less important. Therefore, the performance of the action recognition model based on MBA_3DGCN and LSTM adopted by the present invention has been significantly improved.
[0062] Through the experimental results on the NTU_D dataset, the effectiveness of the method proposed by the present invention is effectively proved.
[0063] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.
Claims
1. An action recognition method based on multi-branch three-dimensional graph convolution and LSTM, characterized in that: The steps include: Step 1. Preprocess the input human body 3D posture sequence to obtain node flow information data, including node position information, node motion speed information and node motion acceleration information, and fuse the three types of information; Step 2. For the skeleton sequence of adjacent frames, a multi-branch attention three-dimensional graph convolution MBA_3DGCN is designed to extract features from the current frame and adjacent frames for its own nodes, inwardly connected nodes, outwardly connected nodes, and all other nodes, and perform feature aggregation through the corresponding learnable attention matrix to realize the graph convolution operation; Step 3. Use MBA_3DGCN and LSTM to build an action recognition model based on skeleton sequence; The action recognition model first extracts the spatial and local temporal features of the skeleton sequence based on the MBA_3DGCN module, and then uses the LSTM network to extract the entire temporal change features of the skeleton sequence for action recognition. Step 4. Use the node flow information data as the input of the action recognition model, train the built action recognition model, and use the trained action recognition model to perform action recognition to obtain the final recognition result.
2. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, characterized in that: The step 1 is specifically as follows: The input human body 3D posture sequence is expressed as ; Where N represents the number of joint points of human body posture information, T is the length of the skeleton sequence, i.e. the number of key frames, and D is the dimension of the input joint point information. For a three-dimensional joint point sequence, D=3; The node position information is obtained by centering and normalizing the node coordinate data vector. The node motion speed information is the first-order information of the node position information, and the node motion acceleration information is the second-order information of the node position information. The position information, velocity information and acceleration information of each joint point are fused as the input information of the node branch.
3. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.
1. Construct a multi-branch attention graph convolution for a single frame; Step 2.
2. Expand the single-frame multi-branch attention graph convolution obtained in step 2.1 to the temporal domain to obtain the expression of the t-th frame output of the three-dimensional graph convolution of the multi-branch attention; Step 2.
3. Simplify the expression of the t-th frame output of the three-dimensional graph convolution in step 2.2; Step 2.
4. Optimize the 3D graph convolution operation by normalizing the attention matrix.
4. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 3, characterized in that: The step 2.1 is specifically as follows: For the single-frame node information, the feature conversion matrices of four branches are set, namely, the self-node feature conversion matrix , connect the neighboring node feature conversion matrix inward , connect neighboring node feature conversion matrix outward And the global other node feature conversion matrix , , is the input feature The dimension of is the output feature Dimensions; Used to extract features from nodes, , , Used to extract other global node features; Multi-branch local features with attention mechanism The extracted graph convolution formula is expressed as: (1) in, The attention matrix representing attention to its own features; The attention matrix representing the attention to the centripetal adjacent nodes; represents the attention matrix that focuses on the centrifugal adjacent nodes; For single-frame node information, extract node features from a global perspective and set the global feature conversion matrix of a single frame. , and then perform the maximum pooling MaxPool on the input node information, setting the stride to 2; MaxPool extracts the features of a certain part of the body, obtains the features of multiple body parts, and then uses the attention matrix To focus on the characteristics of each body part; Then add the global features of the attention mechanism The extracted graph convolution formula is expressed as: (2) Then, the multi-branch local features and global features are combined to obtain the output of the multi-branch attention graph convolution of a single frame: (3)。 5. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 4, characterized in that: In step 2.2, the multi-branch attention graph convolution of a single frame shown in formula (3) is extended to multiple consecutive frames adjacent to time t, that is, one frame is extended to multiple consecutive frames at time t, and the process is as follows: Assume the convolution time kernel length is , for the consecutive adjacent Frame, required For each frame , , , Attention matrix and for each frame Transformation matrix; Then the t-th frame output of the 3D graph convolution of the multi-branch attention is It is expressed as: (4) in Indicates time, [ ] is a rounding operation, Represents the input information of the node branch , The superscript in in the string indicates input. Subscript in Indicates short-term timing; Realizes short-term timing The features of each frame are aggregated and transformed.
6. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 5, characterized in that: The step 2.3 is specifically as follows: Formula (4) is simplified, and different frames use the same feature transformation matrix. Four shared feature transformation matrices are set for the four branches respectively. , then the simplified 3D graph convolution formula is expressed as: (5)。 7. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 6, characterized in that: The step 2.4 is specifically as follows: Aggregate the features of all frames of each branch in formula (5) , here Refers to , , , , they have the same operations and are optimized as follows: (1) Expand the feature aggregation and express it equivalently as follows: (6) (2) For the concatenated attention matrix Perform row normalization Operation processing, expressed as: (7) in, The operation adopts the softmax function and defines formula (7) as a Δ operation, which can be rewritten as: (8) In formula (8), i is the local time series parameter, from the first frame when i=0 to The last frame at time; Substituting formula (8) into formula (5), we get: (9) Write formula (9) in split formula form: (10)。 8. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, characterized in that: The action recognition model also includes a fully connected layer FC and a Softmax; wherein, after the global temporal feature is extracted by the LSTM network, the human action is recognized by the fully connected layer FC and the Softmax.
9. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, characterized in that: In step 4, the training process of the built action recognition model is as follows: Step 4.
1. Obtain the standard human action recognition NTU_D dataset; Step 4.
2. Obtain human body posture joint sequence data from the dataset NTU_D and perform preprocessing, that is, perform centering and normalization of the joint sequence data under parameters, and calculate the motion velocity information and motion acceleration information of the joint points; Then the data input by the action recognition model is fused , i.e., a three-dimensional node data sequence; Step 4.
3. Use the 3D node data sequence obtained in step 4.2 as the input of the action recognition model; According to the adjacency matrix representing the connection of human bones initialization , and , randomly initialize K learnable global attention matrices ; Use the uniform function to initialize all weights W of LSTM and randomly initialize the value of h; Use the NTU_D dataset to pre-train the entire action recognition model with a batch size of 64, weight decay set to 0.0001, and learning rate set to 0.1; Finally, obtain the trained model parameters.
10. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, characterized in that: In step 4, after obtaining the trained action recognition model, firstly obtain the human skeleton sequence through the human posture estimation algorithm or the depth camera; then preprocess the skeleton sequence data according to the preprocessing process of step 1; finally, input the skeleton sequence into the action recognition model built by MBA_3DGCN and LSTM to perform human action recognition.
Citation Information
Patent Citations
Attention mechanism-based action recognition method of adaptive graph convolutional network
CN113688765A
Action recognition method based on dynamic local-global graph convolutional neural network
CN114998525A
Three-dimensional human body posture estimation method based on multi-branch attention graph convolution
CN116030537A
Human skeleton behavior recognition method based on self-attention graph convolution
CN117238025A
Skeleton behavior recognition method based on self-adaptive multi-dimensional dynamic graph convolutional network
CN118587479A