Skeleton Action Recognition Method Based on Spatio-Temporal Adaptive Feature Fusion Graph Convolutional Network
The spatio-temporal adaptive feature fusion graph convolutional network addresses limitations in GCN models by enhancing feature extraction and prediction accuracy through joint and skeletal motion flows, improving action recognition precision.
Patent Information
- Application Number
- CN202310609183.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-05-29
AI Technical Summary
The existing skeleton action recognition method based on graph convolution networks has problems such as limited feature extraction capability, simple multi-stream fusion, large parameter calculation, and difficult to extract edges of adjacent matrices, resulting in low accuracy of action recognition.
The spatiotemporal adaptive feature fusion graph convolution network is used to perform multi-scale adaptive feature extraction and attention mechanism fusion of joint flow, bone flow and limb flow data, combined with more joint data with obvious characteristics, and use the spatiotemporal adaptive feature fusion module and global average pooling layer for prediction.
Without increasing the calculation amount, the accuracy of human behavior prediction is improved, richer and more detailed motion representation is achieved, and the accuracy of motion recognition is improved.
Smart Images

Figure CN116665300B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and deep learning, and particularly to a skeleton action recognition method based on a spatio-temporal adaptive feature fusion graph convolutional network. Background Art
[0002] Skeleton data is a time series of two-dimensional or three-dimensional coordinate positions of multiple human body skeletal joints, which can be extracted from video images using pose estimation methods or directly collected using sensor devices. Compared with traditional RGB video recognition methods, action recognition based on skeleton data can effectively reduce the influence of interference factors such as illumination changes, environmental backgrounds, and occlusions during the recognition process, and has strong adaptability to dynamic environments and complex backgrounds.
[0003] Currently, a typical method for using skeletons for action recognition is to construct graph convolutional networks (GCNs). However, the current mainstream models based on GCNs still have the following deficiencies: (1) Limited feature extraction ability. Modules closer to the input (low-level modules) have relatively small receptive fields. Therefore, compared with low-level modules, the focus of high-level modules has a more global view of the input skeleton sequence. Therefore, for time-scale learning, it is difficult to solve problems such as skeleton action semantics by simply using convolutional kernels or dilation rates of fixed sizes in each layer of the network to obtain more effective modeling; (2) The method of multi-stream fusion for specific behavior patterns is simple. Currently, classic multi-stream framework models usually directly add the softmax scores of each stream to obtain the final prediction result. However, in fact, the prediction effects of each stream are significantly different, and simply adding scores is difficult to obtain accurate prediction results, and the parameter calculation amount is large. (3) Generating an adjacency matrix with semantic significance for edges is particularly important in this task. Traditional spatial topology graphs are affected by physical connectivity, and the extraction of edges remains a challenging problem. Summary of the Invention
[0004] The purpose of the present invention is to propose a skeleton action recognition method based on a spatio-temporal adaptive feature fusion graph convolutional network for the above problems, which can more fully extract context information of different scales, and combine more joint data with more obvious features to achieve human behavior prediction without increasing the amount of calculation, which helps to improve the prediction accuracy of human behavior.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] A skeleton action recognition method based on a spatio-temporal adaptive feature fusion graph convolutional network, comprising the following steps:
[0007] S1. Perform data preprocessing and data augmentation on a large-scale original skeleton action sequence human action recognition data set;
[0008] S2. Process the enhanced skeleton data to obtain the second-order information of the skeleton data, where X represents the joint flow state X joint , which is the first-order skeleton information, and C, T, and N are the characteristic dimensions, sequence frame numbers, and joint numbers of the joints respectively.
[0009] The second-order skeleton information includes the bone flow state X bone , the joint motion flow state X joint-motion and the bone motion flow state X bone-motion data, and the formula is as follows:
[0010] X bone = x[:, :, i] - x[:, :, i nei |i = 1, 2,..., N
[0011] X joint-motion = x[:, t + 1, :] - x[:, t, :]|t = 1, 2,..., T, x ∈ X joint
[0012] X joint-motion = x[:, t + 1, :] - x[:, t, :]|t = 1, 2,..., T, x ∈ X bone
[0013] where i represents the i-th joint, and i nei represents the adjacent joint of the i-th joint on the same frame, and t represents the t-th frame of the sequence.
[0014] S3. The original dataset contains 25 human joints. The joint motion flow state and the bone motion flow state are integrated in the channel dimension through an aggregation method to form a limb flow, and the limb flow only contains a total of 22 joints on the four limbs.
[0015] S4. Input the joint flow state X joint , the bone flow state, and the limb flow data into the spatio-temporal adaptive feature fusion graph convolutional network for training to obtain the corresponding initial prediction results and softmax scores, and finally fuse and output the final prediction results by adding weights.
[0016] Preferably, the spatio-temporal adaptive feature fusion graph convolutional network model includes a spatio-temporal adaptive feature fusion module, a global average pooling layer, a fully connected layer, and a softmax classifier connected in sequence. The spatio-temporal adaptive feature fusion module includes ten feature extraction modules with output channels of 64, 64, 64, 64, 128, 128, 128, 256, 256, and 256 in sequence.
[0017] Preferably, each layer of the feature extraction module includes a spatial attention graph convolution module, a BN+ReLU layer, and a temporal adaptive feature fusion module connected in sequence. At the same time, the skeleton data is input into a 1×1 convolutional layer and multiplied by the output of the spatial attention graph convolution module and then input into the BN+ReLU layer. The skeleton data is added to the output of the BN+ReLU layer with a residual connection and then input into the temporal adaptive feature fusion module.
[0018] Preferably, the spatial attention graph convolution module contains two parallel branches. Each branch contains a 1×1 convolutional layer and a temporal pooling module. The pooling outputs of the two branches are subtracted, and then passed through a Tanh module and a 1×1 convolutional layer in sequence to construct a feature map. This feature map is added to a predefined adjacency matrix A to obtain A cwt which satisfies the following formula:
[0019] A cwt = αQ(X in ) + A
[0020] where α is a learnable parameter, and A cwt is the topological graph of a specific channel. The definition of Q is as follows:
[0021] Q(X in ) = σ(Tanh(TP(φ(X in )) - TP(ψ(X in ))))
[0022] where σ, φ, and ψ are the 1×1 convolutional layers, and TP is the temporal pooling module.
[0023] Preferably, the temporal adaptive feature fusion module contains four branches, an attention feature fusion module M, and an attention feature fusion module 1 - M. Each branch contains a 1×1 convolutional layer to reduce the channel dimension. The first three branches contain two dynamic temporal convolutions with a convolutional kernel size of ks×1, dynamic dilated convolutions with dilation rates of 1 and dr respectively, and a max - pooling layer. The calculation formula of ks is as follows:
[0024]
[0025] where abs represents taking the absolute value, and C l is the output channel dimension of the l - th layer feature extraction module. gamma and b are set to 2 and 1 respectively. The dynamic convolutional kernel and dynamic dilation rate can be obtained through t as follows:
[0026] ks = t if t % 2 else t + 1
[0027] The outputs of the four branches are aggregated through the Concat function to obtain the multi-scale temporal feature X1, and the initial skeleton data is input into the attention feature fusion module M in the form of residuals. The output of the attention feature fusion module M is multiplied by the initial skeleton feature X to obtain the initial attention fusion feature. The multi-scale temporal feature X1 is input into the attention feature fusion module 1-M, and the output of the attention feature fusion module 1-M is multiplied by X1 to obtain the temporal attention fusion feature. The initial attention fusion feature and the temporal attention fusion feature are added to output the feature X′, and the formula is expressed as:
[0028]
[0029] where M(·) is expressed as:
[0030]
[0031] where g(·) and l(·) represent the global context and the local context respectively.
[0032] Preferably, the data preprocessing is that during the training process, the entire skeleton sequence is evenly divided into 20 segments, and one frame is randomly selected from each segment as a new sequence of 20 frames.
[0033] Preferably, the data augmentation is that during the training process, the three-dimensional skeleton sequence is randomly rotated to enhance the robustness to view changes.
[0034] The present invention has the following characteristics and beneficial effects:
[0035] This method adopts a spatio-temporal adaptive feature fusion graph convolutional network model. Firstly, multi-scale adaptive feature extraction is performed on the aggregated spatio-temporal topology to obtain a larger receptive field, and then the attention mechanism is used for feature fusion. The time modeling module can adaptively achieve topological feature fusion to help complete the modeling of actions. On the basis of the existing multi-stream processing method, a body-part-based stream processing method called limb stream is proposed, which can achieve a richer and more refined representation. Without increasing the amount of computation, more joint data with more obvious features are combined to achieve human behavior prediction, effectively improving the final prediction accuracy of human behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 This is the flowchart of the skeleton action recognition method based on the spatio-temporal adaptive feature fusion graph convolutional network of the present invention;
[0038] Figure 2 This is the framework diagram of the spatio-temporal adaptive feature fusion graph convolutional network of the present invention;
[0039] Figure 3 This is the structural schematic diagram of the spatio-temporal adaptive feature fusion graph convolutional network model of the present invention;
[0040] Figure 3 (a) This is the structural schematic diagram of the single-fluid input of the spatio-temporal adaptive feature fusion graph convolutional network of the present invention;
[0041] Figure 3 (b) This is the structural schematic diagram of the feature extraction module of the present invention;
[0042] Figure 3 (c) This is the structural schematic diagram of the spatial attention graph convolutional module of the present invention;
[0043] Figure 3 (d) This is the structural schematic diagram of the time adaptive feature fusion module of the present invention. Detailed implementation manners
[0044] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0045] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0046] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0047] The present invention provides a skeleton action recognition method based on a spatio-temporal adaptive feature fusion graph convolutional network, as Figures 1 - 3 shown, which includes the following steps:
[0048] S1. Perform data preprocessing and data augmentation on the human action recognition data set of the large-scale original skeleton action sequence.
[0049] In one embodiment, the data set is an available public skeleton data set made by a depth sensing camera, including 56,800 skeleton action sequences. The data set has 60 action categories, and the number of each human skeleton joint is 25; the data preprocessing is that during the training process, the entire skeleton action sequence is evenly divided into 20 segments, and one frame is randomly selected from each segment as a new sequence of 20 frames. The data augmentation is that during the training process, the three-dimensional skeleton action sequence is randomly rotated to enhance the robustness to view changes.
[0050] S2. Process the enhanced skeleton data to obtain the second-order information of the skeleton data, where X represents the joint flow state X joint , which is the first-order skeleton information, and C, T, and N are the feature dimension, the number of sequence frames, and the number of joints of the joints respectively.
[0051] It should be noted that the method for processing the enhanced skeleton data is: generating based on joint data, which is the vector difference between two joint points. This processing method is a conventional technical means, so no specific description is given here.
[0052] Specifically, the second-order skeleton information includes the bone flow state X bone , the joint motion flow state X joint-motion and the bone motion flow state X bone-motion data, and the formula is as follows:
[0053] X bone = x[:, :, i] - x[:, :, i nei | i = 1, 2,..., N
[0054] X joint-motion= x[:, t+1, :] - x[:, t, :] | t = 1, 2, ..., T, x ∈ X joint
[0055] X joint-motion = x[:, t+1, :] - x[:, t, :] | t = 1, 2, ..., T, x ∈ X bone
[0056] where i represents the i-th joint, and i nei represents the adjacent joint of the i-th joint in the same frame, and t represents the t-th frame of the sequence.
[0057] For the action recognition task of the skeleton, the first-order skeleton information (coordinates of joints), second-order skeleton information (directions and lengths of bones), and their motion information are all helpful for action recognition. Combining more data with more obvious features helps improve the accuracy of action recognition.
[0058] S3. The original dataset contains 25 human joints. The joint motion flow and the bone motion flow are integrated to form a limb flow, and the limb flow only contains a total of 22 joints on the four limbs.
[0059] S4. Respectively input the joint flow state X joint , the bone flow state and the limb flow data into the spatio-temporal adaptive feature fusion graph convolutional network for training, obtain the corresponding initial prediction results and softmax scores, and finally fuse and output the final prediction results, as Figure 2 shown.
[0060] In this embodiment, the spatio-temporal adaptive feature fusion graph convolutional network model includes a spatio-temporal adaptive feature fusion module, a global average pooling layer, a fully connected layer, and a softmax classifier connected in sequence. The spatio-temporal adaptive feature fusion module includes ten feature extraction modules with output channels of 64, 64, 64, 64, 128, 128, 128, 256, 256, and 256 in sequence. For the joint and bone flow states, ten feature extraction modules are used, and for the limb flow state, eight of the feature extraction modules are used.
[0061] As Figure 3 (a) shown, taking a single flow state input as an example, different flow state data pass through feature extraction modules with different numbers of layers, and finally global average pooling and full connection are performed to obtain the output scores.
[0062] As Figure 3 (b) shown, each layer of the feature extraction module includes a spatial attention graph convolutional module, a BN + ReLU layer, and a temporal adaptive feature fusion module, and at the same time, the skeleton data Input into a 1×1 convolutional layer and multiplied by the output of the spatial attention map convolutional module, then input into the BN+ReLU layer to process the skeleton data Added to the output of the BN+ReLU layer with residual connection and input into the temporal adaptive feature fusion module.
[0063] As Figure 3 shown in (c), the spatial attention map convolutional module takes as input the skeleton data This module contains two parallel branches. Each branch contains a 1×1 convolutional layer and a temporal pooling module. The pooling outputs of the two branches are subtracted, and then passed through a Tanh module and a 1×1 convolutional layer in sequence to construct a feature map. This feature map is added to the predefined adjacency matrix A to obtain A cwt , satisfying the following formula:
[0064] A cwt =αQ(X in ) + A
[0065] where α is a learnable parameter, and A cwt is the topological map of a specific channel. The definition of Q is as follows:
[0066] Q(X in ) = σ(Tanh(TP(φ(X in )) - TP(ψ(X in ))))
[0067] where σ, φ, and ψ are the 1×1 convolutional layers, and TP is the temporal pooling module.
[0068] As Figure 3 shown in (d), the temporal adaptive feature fusion module contains four branches. Each branch contains a 1×1 convolutional layer to reduce the channel dimension. The first three branches contain two dynamic temporal convolutions with kernel size ks×1, dynamic dilated convolutions with dilation rates of 1 and dr respectively, and a max pooling layer. The calculation formula of ks is as follows:
[0069]
[0070] where abs represents taking the absolute value, C l is the output channel dimension of the l-th layer feature extraction module, and gamma and b are set to 2 and 1 respectively. The dynamic convolution kernel and dynamic dilation rate can be obtained through t as follows:
[0071] ks = tift % 2 elset + 1
[0072] The outputs of the four branches are aggregated through the Concat function to obtain the multi-scale temporal feature X1. The initial skeleton feature X is input into the attention feature fusion module M in a residual form. The output of the attention feature fusion module M is multiplied by the initial skeleton feature X to obtain the initial attention fusion feature. The multi-scale temporal feature X1 is input into the attention feature fusion module 1-M. The output of the attention feature fusion module 1-M is multiplied by X1 to obtain the temporal attention fusion feature. The initial attention fusion feature and the temporal attention fusion feature are added together to output the feature X′, which is expressed by the formula:
[0073]
[0074] where M(·) is expressed as:
[0075]
[0076] where g(·) and l(·) represent the global context and the local context respectively. Initial feature fusion is performed on the input features X and X1. After passing through the sigmoid activation function, the output value is between 0 and 1. Through training, the network can determine their respective weights.
[0077] S5. In this embodiment, all experiments are carried out under the PyTorch deep learning framework, and two NVIDIA A800 GPUs are used for training. The training parameters are as follows: the initial learning rate is set to 0.1, the weight decay is set to 0.0004, stochastic gradient descent (SGD) with Nesterov momentum of 0.9 is used to adjust the parameters, the maximum number of training epochs is set to 80, and the learning rate is divided by 10 at the 35th and 55th training stages. Training the model is a well-known technique for those skilled in the art and will not be elaborated here.
[0078] The following are the specific experiments added and the descriptions:
[0079] This embodiment is compared with the advanced models on the NTU-RGB+D 60 and NTU-RGB+D 120 datasets. As shown in Tables 1 and 2, our model has obtained state-of-the-art results in almost all benchmark tests.
[0080] Table 1: Comparison of top-1 accuracy (%) on the NTU-RGB+D dataset with state-of-the-art methods
[0081]
[0082]
[0083] Table 2: Comparison of top-1 accuracy (%) on the NTU-RGB+D120 dataset with state-of-the-art methods
[0084]
[0085]
[0086] S6. This embodiment demonstrates the effectiveness of the multi-modal adaptive feature fusion network. All ablation experiments are conducted on the NTURGB+D 60 and NTU RGB+D 120 Cross Subject benchmarks.
[0087] Table 3 demonstrates Figure 3 (d) The effectiveness of the temporal adaptive feature fusion module
[0088]
[0089] In Table 3, the backbone network is the CTR-GCN model. The present invention improves on it and the accuracy is increased by 0.7% and 0.8% respectively on the Cross Subject benchmarks of NTU60 and NTU120. It shows that the temporal adaptive fusion module can guide the model to better learn action classification.
[0090] Table 4 Performance comparison between the present invention and traditional independent streams on different data streams.
[0091]
[0092] In Table 4, when compared with advanced methods, the CTR-GCN uses joint stream, bone stream, joint motion stream, and bone motion stream. The experimental reproduction results are 89.8%, 90.2%, 87.4%, and 86.9% respectively. The present invention uses data from joint stream, bone stream, and limb stream. Compared with the CTR-GCN, the joint stream and bone stream are improved by 0.4% and 0.5% respectively. The limb stream is the fusion of joint motion stream and bone motion stream, and it has better effects than single joint motion stream and bone motion stream.
[0093] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principles and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments including components still fall within the protection scope of the present invention.
Claims
1. A skeleton action recognition method based on a spatio-temporal adaptive feature fusion graph convolutional network, characterized in that It includes the following steps: S1. Obtain the original dataset of the human skeleton action sequence and perform data preprocessing and data augmentation; S2. Process the skeleton data obtained after preprocessing and data augmentation to obtain the second-order skeleton information of the skeleton data, where X represents the joint flow state X joint , which is the first-order skeleton information. C, T, and N are the feature dimension, sequence frame number, and number of joints of the joints respectively. The second-order skeleton information includes skeleton flow state, joint motion flow state, and skeleton motion flow state; S3. Integrate the joint motion flow state and the bone motion flow state in the channel dimension by aggregation to form a limb flow; S4. Construct a spatio-temporal adaptive feature fusion graph convolutional network, and the spatio-temporal adaptive feature fusion graph convolutional network model includes a spatio-temporal adaptive feature fusion module, a global average pooling layer, a fully connected layer, and a softmax classifier connected in sequence; The spatio-temporal adaptive feature fusion module includes ten feature extraction modules with output channels of 64, 64, 64, 64, 128, 128, 128, 256, 256, and 256 in sequence; Each layer of the feature extraction module includes a 1×1 convolutional layer, a spatial attention graph convolutional module, a BN+ReLU layer, and a temporal adaptive feature fusion module. At the same time, the skeleton data is input into a 1×1 convolutional layer and a spatial attention graph convolutional module, and the outputs of the two are multiplied and input into the BN+ReLU layer. The skeleton data is added to the output of the BN+ReLU layer through residual connection and input into the temporal adaptive feature fusion module; The time adaptive feature fusion module contains four branches, an attention feature fusion module M, and an attention feature fusion module 1 - M. Each branch contains a 1×1 convolutional layer to reduce the channel dimension. The first three branches contain two dynamic temporal convolutions with a convolutional kernel size of ks×1, a dynamic dilated convolution with dilation rates of 1 and dr respectively, and a max pooling layer. The outputs of the four branches are aggregated through the Concat function to obtain the multi-scale temporal feature X1, and the initial skeleton data are input into the attention feature fusion module M in the form of residuals. The output of the attention feature fusion module M is multiplied by the initial skeleton feature X to obtain the initial attention fusion feature. The multi-scale temporal feature X1 is input into the attention feature fusion module 1-M. The output of the attention feature fusion module 1-M is multiplied by X1 to obtain the temporal attention fusion feature. The initial attention fusion feature and the temporal attention fusion feature are added to output the feature X′, and the formula is expressed as: Where M(·) is expressed as: Where g(·) and l(·) represent the global context and the local context respectively; S5. Respectively input the joint flow state X joint , bone flow state, and limb flow data into the spatio-temporal adaptive feature fusion graph convolutional network for training, obtain the corresponding initial prediction results and softmax scores, and finally fuse and output the final prediction results by weighted summation.
2. The skeleton action recognition method based on the spatio-temporal adaptive feature fusion graph convolutional network according to claim 1, wherein, In the step S1, the preprocessing method of the original dataset is: evenly divide the entire skeleton action sequence into 20 segments, and randomly select one frame from each segment as a new sequence of 20 frames.
3. The skeleton action recognition method based on the spatio-temporal adaptive feature fusion graph convolutional network according to claim 2, wherein In the step S1, data augmentation is performed on the preprocessed skeleton action sequence by randomly rotating the skeleton action sequence.
4. The skeleton action recognition method based on the spatio-temporal adaptive feature fusion graph convolutional network according to claim 1, characterized in that The spatial attention graph convolution module contains two parallel branches. Each branch contains a 1×1 convolutional layer, a temporal pooling module, and a Tanh module. The outputs of the temporal pooling modules of the two branches are subtracted, and then passed through a Tanh module and a 1×1 convolutional layer in sequence to construct a feature map. The feature map is added to the predefined adjacency matrix A to obtain A cwt , which satisfies the following formula: A cwt = αQ(X in ) + A where α is a learnable parameter, and A cwt is the topological graph of a specific channel, and Q is defined by the following formula: Q(X in ) = σ(Tanh(TP(φ(X in )) - TP(ψ(X in )))) Where σ, φ, and ψ are the 1×1 convolutional layers, and TP is the time pooling module.
5. The skeleton action recognition method based on the spatio-temporal adaptive feature fusion graph convolutional network according to claim 4, wherein The calculation formula of the ks is as follows: where abs represents taking the absolute value, and C l is the output channel dimension of the l-th layer feature extraction module. gamma and b are set to 2 and 1 respectively. The dynamic convolution kernel and the dynamic dilation rate can be obtained through t as follows: ks = t if t%2 else t + 1.
6. The method for skeleton action recognition based on a spatio-temporal adaptive feature fusion graph convolutional network according to claim 4, wherein The original dataset contains 25 human joints, and the limb flow contains a total of 22 joints on the four limbs.
Citation Information
Patent Citations
Skeleton action recognition method based on multi-gravity-center space-time attention graph convolutional network
CN116012950A
Skeleton action recognition method based on multidimensional dynamic topology learning graph convolution
CN116092182A