An Action Recognition Method Based on Multi-Branch 3D Graph Convolution and LSTM

The integration of multi-branch attention-based three-dimensional graph convolution with LSTM networks addresses the limitations of existing methods by simultaneously extracting spatial and temporal features from skeletal data, significantly improving action recognition accuracy.

CN120088863BActive Publication Date: 2025-07-15SHANDONG UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510569924.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-15
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract the spatial posture characteristics and timing dynamic changes of the skeleton sequence at the same time. Graph convolution only focuses on local connection characteristics and cannot capture the spatial dependence relationship between non-adjacent nodes. Single-frame graph convolution cannot extract skeleton motion information of multiple consecutive frames.

Method used

Multi-branch 3D graph convolution MBA_3DGCN combined with LSTM network is used to design a multi-branch attention three-dimensional graph convolution MBA_3DGCN, which extracts spatial and local timing features for the skeleton sequence of adjacent multi-frames, and extracts the entire timing change features through the LSTM network to build an action recognition model.

Benefits of technology

It improves the accuracy of action recognition, enhances the feature extraction ability of graph convolution, reduces interference with single-frame error data, fully utilizes the long-term feature extraction ability of LSTM, and improves the overall accuracy of action recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088863B_ABST
    Figure CN120088863B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of human action recognition, and discloses an action recognition method based on multi-branch three-dimensional graph convolution and LSTM. The present invention designs a three-dimensional graph convolution operation with multi-branch attention to simultaneously extract spatial features and local temporal features in the human skeleton sequence. The multi-branch attention structure uses multiple transformation matrices to extract information contributed to different nodes. The global attention therein breaks through the limitation that graph convolution only focuses on local features. At the same time, the features of each node not only focus on the information of other nodes in the current frame, but also focus on the node information of the previous frame and the next frame, that is, the continuous frame node information. Therefore, it can simultaneously extract the spatial and temporal transformation features of the skeleton sequence, and improves the feature extraction ability of graph convolution. The present invention first uses three-dimensional graph convolution to extract skeleton spatial and local temporal features, then uses LSTM to extract global skeleton change features, and finally performs action classification and recognition through a fully connected layer, improving the accuracy of action recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and human action recognition, and particularly relates to an action recognition method based on multi-branch three-dimensional graph convolution and LSTM. Background Art

[0002] Human action recognition is an advanced and abstract visual perception task, belonging to the category of pattern recognition, aiming to understand and judge human behavior activities. Action recognition has a wide range of applications in all walks of life and various fields. According to different input information of action recognition, it is mainly divided into two types of methods: based on image / video and based on skeleton sequence.

[0003] With the rapid development of human skeleton pose estimation algorithms, such as the emergence of OpenPose, DeepPose, AlphaPose, etc., it has become easier and more feasible to directly obtain human skeleton sequence data in real time from videos. Therefore, action recognition based on skeleton data has received more attention. Currently, directly using a deep network model to extract the temporal features contained in the skeleton sequence to achieve end-to-end action classification and recognition has become the mainstream method.

[0004] Therefore, how to extract the spatial pose features and temporal dynamic change features contained in the skeleton sequence is the key to action recognition based on skeleton sequence. For the extraction of dynamic temporal features between frames, the currently mainly used methods are models based on the variable-length long short-term memory (LSTM) of the recurrent neural network (RNN) and the temporal convolutional network TCN network. For the extraction of spatial features of in-frame nodes, currently, the GCN network model is mostly used to extract the spatial dependence features of human joints. However, graph convolution often only focuses on local connection features and cannot pay attention to the spatial dependence relationship between non-adjacent nodes.

[0005] Currently, existing patent applications related to action recognition based on skeleton sequence, such as the related patent contents disclosed in Patent Document 1, Patent Document 2, etc., are usually single-frame graph convolution, that is, only the skeleton spatial information of a single frame can be extracted, and the extraction of temporal information is then implemented by using a TCN module.

[0006] For the extraction of temporal features and spatio-temporal features, they can also be extracted simultaneously instead of separately. Currently, the relevant literatures disclosed in the prior art have not involved designing a three-dimensional graph convolution operation to simultaneously extract the spatial features and temporal features of adjacent multiple frames, and the skeleton motion information of continuous multiple frames is an important information source for action recognition.

[0007] Related Technical Literature

[0008] Patent Document 1 Chinese Invention Patent Application, Publication No.: CN 113688765 A, Publication Date: November 23, 2021;

[0009] Patent Document 2 Chinese Invention Patent Application, Publication No.: CN 114998525 A, Publication Date: September 2, 2022. Summary of the Invention

[0010] The object of the present invention is to propose an action recognition method based on multi-branch three-dimensional graph convolution and LSTM. A multi-branch attention three-dimensional graph convolution MBA_3DGCN is designed for the skeleton sequence of adjacent multiple frames to simultaneously extract the spatial and local temporal features of the skeleton sequence, and the LSTM network is used to extract the entire temporal change features of the skeleton sequence. Through the combination of MBA_3DGCN and LSTM, it is beneficial to improve the accuracy of action recognition.

[0011] In order to achieve the above object, the present invention adopts the following technical solutions:

[0012] An action recognition method based on multi-branch three-dimensional graph convolution and LSTM, comprising the following steps:

[0013] Step 1. Preprocess the input human three-dimensional pose sequence to obtain node flow information data, including node position information, node motion speed information, and node motion acceleration information, and fuse the three kinds of information;

[0014] Step 2. For the skeleton sequence of adjacent multiple frames, design a multi-branch attention three-dimensional graph convolution MBA_3DGCN, and extract features for its own nodes, inwardly connected nodes, outwardly connected nodes, and all other nodes of the current frame and adjacent frames respectively, and perform feature aggregation through corresponding learnable attention matrices to implement graph convolution operations;

[0015] Step 3. Use MBA_3DGCN and LSTM to build an action recognition model based on the skeleton sequence;

[0016] The action recognition model first extracts the spatial and local temporal features of the skeleton sequence in the MBA_3DGCN module, and then uses the LSTM network to extract the entire temporal change features of the skeleton sequence for action recognition;

[0017] Step 4. Use the node flow information data as the input of the action recognition model, train the built action recognition model, and use the trained action recognition model for action recognition to obtain the final recognition result.

[0018] The present invention has the following advantages:

[0019] As described above, the present invention relates to an action recognition method based on multi-branch three-dimensional graph convolution and LSTM. This method designs a three-dimensional graph convolution MBA_3DGCN based on local temporal multi-branch attention for the skeleton sequences of adjacent multiple frames, and combines it with a global temporal LSTM network to build a model structure for human action recognition. Among them, the method of the present invention realizes spatio-temporal unified three-dimensional graph convolution. The present invention first uses MBA_3DGCN to extract the spatial and local multi-frame temporal features of the skeleton sequence simultaneously, improves the graph convolution feature extraction ability while increasing the anti-interference ability, and reduces the interference of single-frame error data. On this basis, the LSTM network is used to extract the entire temporal change features of the skeleton sequence for action recognition, giving full play to the ability of LSTM to extract long temporal features. The combination of MBA_3DGCN and LSTM significantly improves the accuracy of action recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is the network structure diagram of the action recognition method based on multi-branch three-dimensional graph convolution and LSTM in the embodiment of the present invention;

[0021] Figure 2 It is the human body structure diagram in the embodiment of the present invention; among them, Figure 2 (a) is the schematic diagram of the joint points of the NTU_D skeleton data set in it, Figure 2 and (b) in it is the schematic diagram of the body parts formed by adjacent nodes;

[0022] Figure 3 It is the network structure diagram of the multi-branch attention three-dimensional graph convolution MBA_3DGCN in the embodiment of the present invention;

[0023] Figure 4 It is the network structure diagram of LSTM in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0025] Embodiment 1

[0026] This embodiment proposes an action recognition method based on multi-branch 3D graph convolution and LSTM. Considering the fact that action recognition is to recognize the continuous change information of multiple joints of the human body's bones, and at the same time, due to the natural graph structure of the human body, graph convolution is considered to extract features. Between adjacent consecutive frames of the skeleton sequence, it can not only represent the human body's posture, but also contain the change information of the posture, which is important information for action recognition. Therefore, the present invention designs a multi-branch attention 3D graph convolution (Multi Branch Attention 3 Dimension GraphConvolution) operation to simultaneously extract the spatial features and local temporal features in the human skeleton sequence. The multi-branch attention structure uses multiple transformation matrices to extract information contributed to different nodes. The global attention therein breaks through the limitation that graph convolution only focuses on local features. At the same time, the features of each node not only focus on the information of other nodes in this frame, but also focus on the information of nodes in the previous frame and the next frame, that is, the continuous frame node information. Therefore, the spatial and temporal transformation features of the skeleton sequence can be simultaneously extracted, improving the feature extraction ability of graph convolution. Due to the problem of computational complexity, the length of the time frame of the 3D graph convolution MBA_3DGCN is generally limited, and it mainly extracts local temporal transformation features. Therefore, the present invention also uses MBA_3DGCN and LSTM to build an action recognition model based on the skeleton sequence. First, the spatial and local temporal features of the skeleton extracted by the MBA_3DGCN graph convolution are used, then the global skeleton change features are extracted by LSTM, and finally, action classification recognition is performed through a fully connected layer to output the action recognition result.

[0027] As Figure 1 shown, the action recognition method based on multi-branch 3D graph convolution and LSTM in this embodiment includes the following steps:

[0028] Step 1. Preprocess the input three-dimensional human body posture sequence to obtain node stream information data, including node position information, node movement speed information, and node movement acceleration information, and fuse the three types of information.

[0029] The input three-dimensional human body posture sequence is expressed as .

[0030] Where N represents the number of joint points of the human body posture information, T is the length of the skeleton sequence, that is, the number of key frames, D is the dimension of the input joint point information, and for a three-dimensional joint point sequence, D = 3.

[0031] Since the dataset used in this embodiment is the NTU_D dataset, therefore, D = 3, T = the number of key frames (the longest is 300 frames, the shortest is 32 frames, and the average length is 82.9 frames), N = 25.

[0032] Set of joint points After centering and normalization preprocessing, the data is obtained .

[0033] Among them, the centering process is based on the selected waist joint coordinates Figure 2 In (a), the node numbered 21 is used As the center, calculate the position difference of each node coordinate Relative to And then perform normalization processing. The formula is as follows:

[0034] .

[0035] The node motion speed information, which is the first-order information of the node position information, is expressed as:

[0036] .

[0037] The node motion acceleration information, which is the second-order information of the node position information, is expressed as:

[0038] .

[0039] Fuse the information of all joint points as the input information of the node branch. The formula is expressed as:

[0040] .

[0041] Here, Is the fusion of the above three kinds of information (position information, speed information and acceleration information), and is used as the input of the action recognition model based on multi-branch three-dimensional graph convolution and LSTM built in step 3 below.

[0042] Step 2. For the skeleton sequences of adjacent multiple frames, design a multi-branch attention three-dimensional graph convolution (MultiBranch Attention 3 Dimension Graph Convolution, MBA_3DGCN).

[0043] This graph convolution adopts a multi-branch structure (MB, Multi-Branch, specifically 4 branches), and extracts features for its own nodes, inwardly connected nodes, outwardly connected nodes and all other nodes provided by the current frame and adjacent frames respectively, and aggregates the features through the corresponding learnable attention matrix (A, attention) to implement the graph convolution operation.

[0044] As Figure 3 Shown, the design process of this multi-branch attention three-dimensional graph convolution MBA_3DGCN is as follows:

[0045] Step 2.1. Construct a multi-branch attention graph convolution for a single frame.

[0046] For the single-frame node information, the present invention sets four branch feature transformation matrices, namely the self-node feature transformation matrix , the inwardly connected adjacent node feature transformation matrix , the outwardly connected adjacent node feature transformation matrix and the global other node feature transformation matrix ; where , is the dimension of the input feature , is the dimension of the output feature . Among them is used to extract the self-node features, , , are used to extract the global other node features. The graph convolution formula for extracting the multi-branch local features with the attention mechanism added is expressed as:

[0047] (1)

[0048] where represents the attention matrix for focusing on its own features, i.e., the self-attention matrix, which is initialized with the self-link adjacency matrix, i.e., the identity matrix ; represents the attention matrix for focusing on the centripetal adjacent nodes, i.e., the inward attention matrix, which is initialized with the centripetal adjacency matrix , is the matrix representation of the directed graph of the human body's bones connected from the outside to the inside; represents the attention matrix for focusing on the centrifugal adjacent nodes, i.e., the outward attention matrix, which is initialized with the centrifugal adjacency matrix , is the matrix representation of the directed graph of the human body's bones connected from the inside to the outside.

[0049] Adjust the weights therein through adaptive learning. These matrices are only initialized with the adjacency matrix, but the learned parameters are not limited to the edges defined by these adjacency matrices, increasing the flexibility of the local attention matrix.

[0050] For the single-frame node information, extract the node features from a global perspective, set the global feature transformation matrix of the single frame, and then perform max pooling MaxPool on the input node information, with the set stride being 2.

[0051] For a dataset where the number of nodes is not a multiple of 2, such as the NTU_D dataset with 25 skeletal nodes, the positions of its 25 nodes are as shown in Figure 2 (a) in it. The following processing is performed in this embodiment:

[0052] Duplicate the centroid node 21 to get 26 nodes. MaxPool extracts the features of a certain part of the body, and the part is as shown in Figure 2 (b) in it, obtaining the features of 13 body parts. Then, through the attention matrix to focus on the features of these 13 body parts, and add the global features with the attention mechanism The graph convolution formula extracted is expressed as:

[0053] (2)

[0054] Then, synthesize the multi-branch local features and global features to obtain the output of the multi-branch attention graph convolution for a single frame as:

[0055] (3).

[0056] Step 2.2. Expand the multi-branch attention graph convolution for a single frame obtained in Step 2.1 to the temporal domain to obtain the expression form of the output of the t-th frame of the three-dimensional graph convolution of this multi-branch attention.

[0057] Specifically, expand the multi-branch attention graph convolution for a single frame shown in formula (3) to the continuous multi-frames adjacent to the t-th moment, that is, expand 1 frame to the continuous multi-frames at the t-th moment. The process is as follows:

[0058] Let the convolution time kernel length be , indicating that the feature extraction window span for capturing temporal dependence is . For the continuous frames adjacent to the t-th moment, attention matrices for each frame and , , , and transformation matrices for each frame are required.

[0059] Then, the output of the t-th frame of the three-dimensional graph convolution of this multi-branch attention is expressed as:

[0060] (4)

[0061] where represents time, [ ] is the floor operation, represents the input information of the node branch ,​ The superscript in in represents the input, and the subscript in represents short-term timing; The feature aggregation of each frame of multiple frames of short-term timing is realized. If a

[0062] is set for each frame,

[0063] it will significantly increase the number of model parameters and is not conducive to learning. Therefore, the formula is simplified.

[0064] Step 2.3. Simplify the expression form of the output of the t-th frame of the 3D graph convolution in Step 2.2. Since it is considered that regardless of which frame, the aggregated features from the adjacent nodes include timing features and spatial position features, the same feature transformation matrix can be used for transformation.

[0065] (5).

[0066] At time t, each branch aggregates the features of the previous frames and then performs unified feature transformation, that is, all frames of each branch share a W feature transformation matrix, significantly reducing the number of parameters of the 3D graph convolution.

[0067] Step 2.4. Optimize the operation of the 3D graph convolution through the normalization process of the attention matrix.

[0068] To improve the trainability of the attention matrix, for the feature aggregation of all frames of each branch in formula (5) denoted as

[0069] (1) Expand the feature aggregation and equivalently represent it as:

[0070] (6)

[0071] (2) Perform row normalization on the concatenated attention matrix operation, which is expressed as:

[0072] ​​​​​ (7)

[0073] If the operation uses the softmax function and defines formula (7) as a Δ operation, it is rewritten as:

[0074] (8)

[0075] In formula (8), i is the local time series parameter, from the first frame at i = 0 to the last frame at

[0076] Substituting formula (8) into formula (5) gives:

[0077] (9)

[0078] Write formula (9) in the form of a separate formula:

[0079] (10).

[0080] The adaptive three-dimensional graph convolution MBA_3DGCN with multi-branch attention proposed in the present invention (as shown in formula (10)) considers the global features of the space of the body structure and the local characteristics of the time series. The attention matrix is obtained through adaptive learning, and the parameter quantity is saved by sharing the feature transformation matrix. It can not only achieve spatial aggregation, but also achieve the aggregation of nodes between different frames (temporally consecutive frames), that is, temporal local feature aggregation. Therefore, it has higher anti-interference ability.

[0081] Step 3. Use MBA_3DGCN and LSTM to build an action recognition model based on the skeleton sequence. The action recognition model built in this embodiment includes MBA_3DGCN, LSTM, as well as a fully connected layer FC and Softmax.

[0082] This action recognition model first extracts the spatial and local temporal features of the skeleton sequence in the MBA_3DGCN module, and then uses the LSTM network to extract the entire temporal change feature of the skeleton sequence.

[0083] After the global temporal feature extraction by LSTM, the recognition of human actions is realized through the fully connected layer FC and Softmax.

[0084] The network structure of LSTM realizes the modeling of sequence data by controlling the transmission and forgetting of information, as Figure 4 shown. F t out 、C t 、h t are the input, unit memory state, and output vector of LSTM at time t, respectively, f t, i t , o t are the forget gate, input gate, and output gate of the LSTM unit respectively. These gates have fully connected layers with the activation function of σ (the activation function uses softmax).

[0085] The LSTM module is expressed by the formula as follows:

[0086] ;

[0087] ;

[0088] ;

[0089] ;

[0090] ;

[0091] .

[0092] Among them, W is the input weight, and U is the weight of the previous cell state. The local feature vector generated by MBA_3DGCN is used as the input of the LSTM network, and the global temporal features are extracted through the LSTM.

[0093] Step 4. Use the node flow information data as the input of the action recognition model, train the built action recognition model, and use the trained action recognition model to perform action recognition to obtain the final recognition result.

[0094] The process of training the built action recognition model is as follows:

[0095] Step 4.1. Obtain the standard human action recognition NTU_D dataset.

[0096] Step 4.2. Obtain the human pose joint point sequence data from the dataset NTU_D and perform preprocessing, that is, perform joint sequence data under centering and normalization parameters, and calculate the motion speed information and motion acceleration information of the joint points.

[0097] Then fuse to obtain the data input to the action recognition model , that is, the three-dimensional node data sequence.

[0098] Step 4.3. Use the three-dimensional node data sequence obtained in Step 4.2 as the input of the action recognition model.

[0099] According to the adjacency matrix representing the connection situation of the human skeleton Initialize , and , randomly initialize K learnable global attention matrices . Initialize all the weights W of the LSTM using the uniform function and randomly initialize the value of h. Use the NTU_D dataset to pre-train the entire action recognition model with a batch size of 64, a weight decay set to 0.0001, and a learning rate set to 0.1; finally obtain the trained model parameters.

[0100] The loss function uses the classic classification loss function, the negative log-likelihood loss with label smoothing .

[0101] The formula is expressed as:

[0102] .

[0103] Among them, is the number of all samples, is the number of labels, is the smoothing exponent, taking , is the true label of the i-th sample , is the probability value that the sample after passing through the model and is predicted to belong to the label.

[0104] Obtain the training output of MBA_3DGCN as the input of the LSTM network for joint training, and finally obtain the trained model parameters for 3D pose estimation. After obtaining the trained action recognition model, deploy the model in practice.

[0105] After the model is deployed, the general process of action recognition is as follows: First, obtain the human body bone sequence through the human body pose estimation algorithm or depth camera; then preprocess the bone sequence data according to the preprocessing process in step 1; finally, input the skeleton sequence into the action recognition model built by MBA_3DGCN and LSTM for human action recognition.

[0106] To verify the performance of the proposed MBA_3DGCN+LSTM action recognition model of the present invention, it is verified on the NTU_D dataset. The MBA_3DGCN+LSTM model designed by the present invention is compared with ST-GCN and agcn of the baseline method. To better compare its performance, the present invention all adopts the single-stream 1s method, that is, only using the node information of the skeleton data and not using its bone information. The results of the comparison method and the method of the present invention are shown in Table 1.

[0107] Table 1 Experimental Results of Cross-Subject (CS) and Cross-View (CV) Evaluations on the NTU_D Dataset

[0108]

[0109] As can be easily seen from Table 1 above, the performance of the action recognition model designed by the present invention based on MBA_3DGCN and LSTM has been significantly improved. The effect of using all 3D multi-branch GCNs is worse than that of 2D multi-branch GCNs. Perhaps in the later stage of the model, after the feature extraction and fusion in the previous stages, the graph structure between adjacent frames becomes less important. Therefore, the performance of the action recognition model based on MBA_3DGCN and LSTM adopted by the present invention has been significantly improved.

[0110] The experimental results on the NTU_D dataset effectively prove the effectiveness of the method proposed by the present invention.

[0111] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.

Claims

1. An action recognition method based on multi-branch 3D graph convolution and LSTM, characterized in that The method includes the following steps: Step 1. Preprocess the input three-dimensional human body pose sequence to obtain node stream information data, including node position information, node motion speed information, and node motion acceleration information, and fuse the three types of information; Step 2. For the adjacent multi-frame skeleton sequences, design a multi-branch attention three-dimensional graph convolution MBA_3DGCN, and extract features respectively for the nodes of the current frame and adjacent frames provided to its own nodes, inward-connected nodes, outward-connected nodes, and all other nodes, and perform feature aggregation through corresponding learnable attention matrices to implement graph convolution operations; Step 3. Use MBA_3DGCN and LSTM to build an action recognition model based on the skeleton sequence; The action recognition model first extracts the spatial and local temporal features of the skeleton sequence in the MBA_3DGCN module, and then uses the LSTM network to extract the entire temporal change features of the skeleton sequence for action recognition; Step 4. Use the node stream information data as the input of the action recognition model, train the built action recognition model, and use the trained action recognition model for action recognition to obtain the final recognition result.

2. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, wherein the specific content of step 1 is: The input three-dimensional human body pose sequence, expressed as ; where N represents the number of human body pose information joint points, T is the length of the skeleton sequence, i.e., the number of key frames, and D is the dimension of the input joint point information. For a three-dimensional joint point sequence, D = 3; The node position information is obtained by centralizing and normalizing the node coordinate data vector. The node motion speed information is the first-order information of the node position information, and the node motion acceleration information is the second-order information of the node position information; Fuse the position information, speed information, and acceleration information of each joint point as the input information of the node branch.

3. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, wherein the specific content of step 2 is: Step 2.

1. Construct a single-frame multi-branch attention graph convolution; Step 2.

2. Expand the single-frame multi-branch attention graph convolution obtained in step 2.1 to the temporal domain to obtain the expression form of the output of the t-th frame of the multi-branch attention three-dimensional graph convolution; Step 2.

3. Simplify the expression form of the output of the t-th frame of the three-dimensional graph convolution in step 2.2; Step 2.

4. Optimize the operation of the three-dimensional graph convolution through the normalization process of the attention matrix.

4. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 3, wherein the specific content of step 2.1 is: Set the feature transformation matrices of four branches for single-frame node information, namely the self-node feature transformation matrix , the inwardly connected adjacent-node feature transformation matrix , the outwardly connected adjacent-node feature transformation matrix and the global other-node feature transformation matrix , , is the dimension of the input feature ; is the dimension of the output feature ; is used to extract self-node features, , , are used to extract global other-node features; Multi-branch Local Features with Attention Mechanism The extracted graph convolution formula is expressed as: (1) Among them, represents the attention matrix focusing on its own features; represents the attention matrix focusing on centripetal adjacent nodes; represents the attention matrix focusing on centrifugal adjacent nodes; For single-frame node information, extract node features from a global perspective and set the global feature transformation matrix for a single frame , then perform max pooling (MaxPool) on the input node information with a set stride of 2; MaxPool extracts features of a certain body part, obtaining features of multiple body parts, and then uses the attention matrix to focus on the features of each body part; The global features with the attention mechanism added The graph convolution formula extracted is expressed as: (2) Then synthesize the multi-branch local features and global features, and the output of the single-frame multi-branch attention graph convolution is: (3)。 5. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 4, wherein In step 2.2, the multi-branch attention graph convolution of a single frame shown in formula (3) is extended to consecutive multi-frames adjacent at time t, that is, one frame is extended to consecutive multi-frames at time t. The process is as follows: Let the length of the convolutional time kernel be , for the consecutive frames adjacent to the t-th moment, attention matrices for each frame and , , , transformation matrices for each frame are required; ​ The output of the t-th frame of the 3D graph convolution with multi-branch attention is expressed as: (4) Among them represents time, and [ ] is the rounding operation, represents the input information of the node branch , The superscript "in" in represents input, and the subscript in represents short-term time series; realizes the conversion of the aggregated features of each frame of the short-term time series frames.

6. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 5, wherein: Step 2.3 is specifically as follows: Simplify formula (4). The same feature transformation matrix is used for different frames, and four shared feature transformation matrices are set for the four branches respectively. , then the simplified 3D graph convolution formula is expressed as: (5)。 7. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 6, wherein: Step 2.4 is specifically as follows: Feature aggregation for all frames of each branch in formula (5) , where refers to , , , , which have the same operations and are processed with the following optimizations: (1) Expand the feature aggregation, and equivalently represent it as: (6) (2) For the spliced attention matrix perform row normalization operation processing, expressed as: (7) Among them, The operation uses the softmax function and defines formula (7) as the Δ operation, then it is rewritten as: (8) In formula (8), i is the local time series parameter, from the first frame at i = 0 to the last frame at Substitute formula (8) into formula (5) to get: (9) Write formula (9) in the form of a separated formula: (10)。 8. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, wherein: The action recognition model further includes a fully connected layer FC and Softmax; among them, after global temporal feature extraction by the LSTM network, the recognition of human actions is realized through the fully connected layer FC and Softmax.

9. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, wherein: In step 4, the process of training the built action recognition model is as follows: Step 4.

1. Obtain the standard human action recognition NTU_D dataset; Step 4.

2. Obtain the human body pose joint point sequence data of the dataset NTU_D and perform preprocessing, that is, perform joint sequence data under centering and normalization parameters, and calculate the motion speed information and motion acceleration information of the joint points; Then, the data input to the action recognition model is obtained through fusion , that is, the three-dimensional node data sequence; Step 4.

3. Use the three-dimensional node data sequence obtained in step 4.2 as the input of the action recognition model; According to the adjacency matrix representing the connection of the human body bones Initialize 、 and , randomly initialize K learnable global attention matrices ; Initialize all the weights W of the LSTM with the uniform function and randomly initialize the value of h; Use the NTU_D dataset to pre-train the entire action recognition model, with a batch size of 64, a weight decay set to 0.0001, and a learning rate set to 0.1; Finally, obtain the trained model parameters.

10. The action recognition method based on multi-branch three-dimensional graph convolution and LSTM according to claim 1, wherein: In step 4, after obtaining the trained action recognition model, first obtain the human body bone sequence through a human body pose estimation algorithm or a depth camera; then preprocess the bone sequence data according to the preprocessing process in step 1; finally, input the skeleton sequence into the action recognition model built by MBA_3DGCN and LSTM to perform human action recognition.

Citation Information

Patent Citations

  • Attention mechanism-based action recognition method of adaptive graph convolutional network

    CN113688765A

  • Action recognition method based on dynamic local-global graph convolutional neural network

    CN114998525A

  • Human skeleton behavior recognition method based on self-attention graph convolution

    CN117238025A

  • Skeleton behavior recognition method based on self-adaptive multi-dimensional dynamic graph convolutional network

    CN118587479A