A Skeleton Action Recognition Method Based on Multidimensional Dynamic Topology Learning Graph Convolution

By using multi-dimensional dynamic topology learning graph convolution and dynamic skeleton topology learning methods in skeleton action recognition, the problem of limited learning ability of graph convolution networks in the existing technology is solved, and more efficient feature extraction and improvement of skeleton action recognition performance is achieved.

CN116092182BActive Publication Date: 2025-06-17JIANGXI UNIV OF SCI & TECH

Patent Information

Application Number
CN202211582847.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-06-17
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

Existing skeleton action recognition methods based on graph convolution use the same topology when feature aggregation, limiting the learning ability of the network, especially in timing and channel dimensions.

Method used

A method based on multidimensional dynamic topology learning graph convolution is proposed, which enhances the learning ability of the network by dynamically modeling the human skeleton topology with timing specificity and channel specificity, and uses multidimensional dynamic topology learning graph convolution (MD2TL-GC) and dynamic skeleton topology learning (DSTL) to enhance the learning ability of the network.

Benefits of technology

It realizes more flexible and efficient feature extraction, and can learn advanced spatial and temporal features under shallow network depth, significantly improving the performance of skeleton action recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116092182B_ABST
    Figure CN116092182B_ABST
Patent Text Reader

Abstract

The present invention provides a skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution, which relates to the field of computer vision. The skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution comprises three components: pure node topology learning graph convolution, dynamic temporal specificity topology learning graph convolution, and channel specificity topology learning graph convolution. Among them, in the dynamic temporal specificity topology learning graph convolution, we propose a dynamic skeleton topology modeling method to efficiently model the dynamic skeleton topology rich in global spatio-temporal topology features. And multi-scale temporal convolution is used to obtain multi-scale temporal features. In addition, in order to supplement the spatial information of the skeleton data, we additionally introduce relative node data and relative bone data for the fusion of the multi-stream network. This model can extract more comprehensive and effective action features and perceive more subtle action differences, which is worthy of vigorous promotion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and relates to the improvement of a skeleton action recognition model, and the action recognition and simulation implementation of skeleton video data. Background Art

[0002] Action recognition is a research hotspot in the field of computer vision, and has a wide range of applications in the fields of human-computer interaction, public security monitoring, film and television production, and medical rehabilitation. Due to the development of depth sensors and real-time human pose estimation technology, skeleton data has become more widespread and inexpensive. Compared with the original RGB and RGB+D data, the skeleton data has a high information density and provides high-semantic-level structural information. Therefore, using skeleton data for action recognition can improve the computational efficiency and recognition performance. Especially in complex scenarios, it has stronger robustness. However, when the existing graph convolution-based skeleton action recognition method performs graph convolution, all temporal frames and channels use the same topological structure for feature aggregation, which limits the learning ability of the graph convolution network. Because different temporal frames represent different moments of action execution, and different channels store different motion pattern features. The relationship between joint points is not always the same at different moments of motion execution and different motion patterns. Therefore, using the same topological structure for graph convolution in different temporal and channel dimensions limits the learning ability of the network. Thus, how to improve the feature extraction ability of graph convolution is still a difficult problem to be solved urgently in the skeleton action recognition task.

[0003] To solve the above problems, the present invention proposes a skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution, which can dynamically model the human skeleton topology with both temporal specificity and channel specificity. In the dynamic temporal dimension-specific topology modeling, by learning the dynamic skeleton topology graph, the global spatio-temporal dynamic information is enhanced. In addition, we also propose relative node data flow and relative bone data flow to supplement the skeleton information. And after modeling through the spatial topological structure, spatio-temporal features are obtained through multi-scale time modeling, and finally the final classification result is obtained through the fusion of the classification results of multi-flow data input. Summary of the Invention

[0004] 1. Technical problems to be solved:

[0005] Aiming at the deficiencies of the prior art, the present invention provides a skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution, which solves the problem that the existing graph convolution method uses a shared graph topology for feature aggregation on all frames or channels, thus greatly limiting the representation ability of the graph convolution network.

[0006] 2. Technical solutions:

[0007] The present invention proposes a skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution. In order to simultaneously model channel and temporal specific topologies, it includes a novel multi-dimensional dynamic topology learning graph convolution, abbreviated as MD 2 TL-GC; it simultaneously models the topological structures in the node dimension, temporal dimension, and channel dimension, enabling each frame and each channel to have independent and dynamic feature aggregation parameters, thereby making the model have greater flexibility and stronger learning ability. In addition, in order to reduce the network depth required to obtain global semantic information, the present invention proposes dynamic skeleton topology learning, abbreviated as DSTL; it expresses the dynamic skeleton sequence as a single dynamic skeleton graph, and performs graph convolution on this basis to aggregate features from a global temporal perspective, enabling the network to learn high-level spatio-temporal features with only a relatively shallow depth. MD 2 TL-GC adaptively models the topological structure in each data dimension using global spatio-temporal information. The network constructed by it only needs 5 layers to outperform the current mainstream methods. In addition, the present invention proposes a method for expressing skeleton information of relative nodes and relative bones, providing complementary information for the original skeleton sequence data, and achieving higher performance improvement compared with the classical node motion flow and bone motion flow.

[0008] The skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution described in the present invention includes the following steps:

[0009] S1. In the input skeleton sequence data, first input it into the multi-dimensional dynamic topology learning graph convolution for spatial modeling. This graph convolution consists of three main parts, namely the pure node topology learning graph convolution, abbreviated as J-GC; the dynamic temporal specificity topology learning graph convolution, abbreviated as DTW-GC; and the channel specificity topology learning graph convolution, abbreviated as CW-GC. Specifically, after the input features are divided in channels, they are respectively input into J-GC and DTW-GC to learn the pure node topological structure and the dynamic temporal specificity topological structure in parallel, and then the learned feature maps are concatenated in channels, and then a 1×1 convolution is used to fully interact and fuse the time dimension and space dimension features; then, input into CW-GC for channel specificity topology learning to obtain a feature map with rich spatial topology structure information;

[0010] S2. Input the spatial modeling features obtained in S1 into a multi-scale temporal convolution, abbreviated as MS-TC. This temporal convolution consists of four branches. In the first three branches, a 1×1 convolution is first used to reduce the number of channels to 1 / 4 of the input channel number, and then dilated convolutions with a dilation rate of 1 and a kernel size of 5, a dilation rate of 2 and a kernel size of 5, and a temporal dimension max pooling operation are respectively used to obtain features with different temporal receptive fields. The fourth branch, as a residual connection, only uses a 1×1 convolution to adjust the number of channels to 1 / 4 of the input channel number. Finally, the results of all branches are concatenated along the channel dimension to obtain the final output, and a multi-scale spatio-temporal feature map can be obtained.

[0011] S3. Input the skeleton sequence into six input data types, including node flow data input, bone flow data input, relative node flow data input, relative bone flow data input, node motion flow data input, and bone motion flow input. Repeat the spatio-temporal modeling process of S1 and S2 for these six data inputs respectively, add and fuse the six different feature outputs, and then train them to obtain the final classification result.

[0012] 3. Beneficial effects:

[0013] The present invention discloses a skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution, in which the proposed MD 2 TL-GC dynamically models the topological structures of each dimension in a hierarchical manner, thus realizing flexible and effective topological relationship modeling and efficient feature extraction. Moreover, the proposed dynamic skeleton topology modeling method, abbreviated as DSTL, learns the global spatio-temporal expression of the skeleton sequence data. Based on this global dynamic topological structure, graph convolution can be performed to enhance the global features in multi-dimensional topology learning and achieve better performance with a shallower network. In addition, the present invention introduces a brand-new relative node and relative bone modality data, further improving the performance of action recognition. A large number of experiments have been carried out on two large skeleton action recognition datasets, NTU-RGB+D and NTU-RGB+D 120. The results show that the network MD 2 TL-GC constructed by using MD 2 TL-GCN has achieved superior performance. Compared with ST-GCN, the network proposed in the present invention has obtained significant improvements of +11.14% and +8.06% respectively on the CS and CV benchmarks of the NTU-RGB+D dataset. Brief Description of the Drawings

[0014] Figure 1 It is the overall process framework of the present invention.

[0015] Figure 2 It is the structural diagram of the J-GC module of the present invention.

[0016] Figure 3 This is the structural diagram of the DSTL and DTW-GC modules of the present invention.

[0017] Figure 4 This is the structural diagram of the CW-GC module of the present invention.

[0018] Figure 5 This is the illustration diagram of multi-stream data input of the present invention.

[0019] Figure 6 This is the structural diagram of the multi-stream network fusion model of the present invention.

[0020] Figure 7 This is the comparison diagram of the accuracy of the present invention and other advanced methods for difficult classes. Specific implementation manners

[0021] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners. A skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution, and its specific implementation steps are as follows:

[0022] S1: Design a multi-dimensional dynamic topology learning graph convolution module.

[0023] In order to achieve flexible and effective topology relationship modeling and efficient feature extraction, the present invention proposes a novel multi-dimensional dynamic topology learning graph convolution, abbreviated as MD 2 TL-GC; it dynamically models the topological structures of each dimension in a hierarchical manner. As Figure 1 shown, MD 2 TL-GC mainly includes three parts: pure node topology learning graph convolution, abbreviated as J-GC; dynamic time series specific topology learning graph convolution, abbreviated as DTW-GC; channel specific topology learning graph convolution, abbreviated as CW-GC. They respectively learn spatio-temporal dimension specific topologies and channel specific topologies through hierarchical learning, aiming to further refine the features using channel specific topologies after obtaining rich spatio-temporal features. Specifically, the input features are respectively input into J-GC and DTW-GC after channel division to learn the pure node topological structure and dynamic time series specific topological structure in parallel, and then the learned feature maps are concatenated on the channels, and then a 1×1 convolution is used to fully interact and fuse the time dimension and space dimension features; afterwards, the input is fed into CW-GC for channel specific topology learning.

[0024] MD 2 The overall calculation process of TL-GC is shown in the following formula.

[0025]

[0026] For different actions and different data samples, in order to obtain a flexible topological structure between nodes, we use J-GC to learn the pure topological structure between nodes. As Figure 2 shown, J-GC learns two types of topological structures: action-specific pure node topological structure and data-specific pure node topological structure Among them has nothing to do with the input features and is directly trained through the action classification loss, where The calculation process is shown in the following formula.

[0027]

[0028] J-GC utilizes the action-specific pure node topological structure and the data-specific pure node topological structure The process of convolution is described in the following formula.

[0029]

[0030] In the skeleton sequence, on different temporal frames, the topological relationship between nodes is different. To learn the specific topological structure in the temporal dimension, the present invention proposes dynamic temporal-specific topological learning graph convolution, simply referred to as DTW-GC. As Figure 3 (c) shown, DTW-GC contains two branches: the temporal-specific branch, simply referred to as TIB; the temporal-global branch, simply referred to as TGB. They are respectively used to learn two types of topological structures: the temporal-specific topology G t and the temporal-global dynamic topology G d . For the stability of training, as Figure 3 (b) shown, in the first layer of MD 2 TL-GCN, the fusion of G t and G d adopts the fusion of the direct topological structure; while in the second layer to the fifth layer, as Figure 3 (c) shown, the concatenation fusion of the channel dimension is performed on the graph convolution results based on G t and G d . The temporal-specific topology G t enables the network to use different feature aggregation methods on different temporal frames, while the temporal-global dynamic topology G d provides the network with global dynamic topological structure information, enabling the graph convolution to have a global receptive field in the time dimension. The temporal-specific topology branch

[0031] As Figure 3 (c) shown, TIB contains the temporal-specific topology and the action-specific topological structure for adjusting the temporal topology The calculation process is shown as follows.

[0032]

[0033] Temporal-specific topology Each temporal frame corresponds to a different topological structure. Therefore, the graph convolution calculation process of TIB is shown as follows.

[0034]

[0035] In TGB, in order to learn the global dynamic skeleton topology, the present invention designs a DSTL module to learn the dynamic skeleton topological structure. The DSTL calculation process is as Figure 3 (a) shown. For the input features {f1, f2, …, f T} of the skeleton sequence, first calculate the approximate sorting pooling coefficient: a t = 2t – T – 1, and then use this coefficient for approximate sorting pooling to obtain the dynamic skeleton, that is, as shown in the following formula:

[0036]

[0037] Subsequently, perform convolution mapping on this dynamic skeleton to obtain the topological structure information containing global spatio-temporal features, and its calculation process is as shown in the following formula:

[0038]

[0039] The dynamic skeleton topology G learned by DSTL d has a global temporal dimension receptive field. Based on this, graph convolution can be performed to efficiently learn high-level semantic features.

[0040] Considering that the shallow layer of the network lacks high-level semantic features, directly performing temporal-specific graph convolution on it will introduce a large variance. Therefore, for the first layer of MD 2 TL-GCN, simply referred to as L1; as Figure 3 (b) shown, fuse the dynamic skeleton topology G containing global temporal information d with the temporal dimension-specific topology G t in an additive manner, and then perform graph convolution on this basis. The fusion topological structure of the L1 layer is calculated as shown in the following formula:

[0041]

[0042] where represents the learnable action-specific topological structure, and both α t and β are learnable parameters used to adjust the proportion between each topological structure.

[0043] For MD 2 For the other layers of TL-GCN, such as Figure 3 (c) shows that the graph convolution results of TIB and TGB are concatenated and fused in the channel dimension. First, the TGB branch performs graph convolution based on the dynamic skeleton topology G learned by DSTL d The process is shown in the following formula

[0044]

[0045] where represents the action-specific learnable topology structure, f2 is the residual connection after adjusting the channel dimension, and α d is a learnable parameter used to adjust the relative importance between the two topology structures

[0046] Then, after concatenating the graph convolution results of the two branches in the channel dimension, they are input into a 1×1 convolution for further learning of feature fusion and channel dimension adjustment

[0047] After concatenating the convolution results of J-GC and DTW-GC, the channel dimension in the obtained feature map contains rich spatio-temporal dynamic features and motion pattern features. Therefore, in order to further refine the features, we propose channel-specific topology learning graph convolution, abbreviated as CW-GC, to further enhance the learning ability of the network by learning the specific topology structure in the channel dimension

[0048] As Figure 4 shown, the learning process of the channel-specific topology structure is shown in the following formula

[0049]

[0050] where represents the learnable parameters of the 1×1 convolution for feature learning and channel dimension adjustment. In order to construct a channel-specific topology structure with rich features, after matrix multiplication and dimension conversion, the obtained Then, an average pooling operation is performed along the second-to-last dimension to obtain a topology structure with channel specificity. The self-correlation of the channel dimension is fused into the adjacency matrix. After that, use to adjust the dimension to obtain the channel-specific topology structure

[0051] CW-GC also uses an action-specific learnable adjacency matrix for adaptive adjustment and its graph convolution process is shown in the following formula

[0052]

[0053] wherein are graph convolution parameters; α c is a learnable parameter used to adjust the relative importance between the two topological structures. The CW-GC network can learn the channel-specific topological structures for the feature maps with rich spatio-temporal information, enabling each channel to have independent feature aggregation parameters and achieving feature refinement.

[0054] S2: Multi-scale temporal feature modeling.

[0055] Considering the different speeds and durations of actions, in temporal modeling, we use multi-scale temporal convolution, abbreviated as MS-TC. The structure of MS-TC is as Figure 1 shown. It contains 4 branches. In the first 3 branches, 1×1 convolution is first used to reduce the number of channels to 1 / 4 of the input channel number, and then dilated convolutions with dilation rate D = 1 and convolution kernel size K t = 5, dilated convolutions with dilation rate D = 2 and K t = 5, and temporal dimension max pooling operations are used to obtain features with different temporal receptive fields; while the fourth branch, as a residual connection, only uses 1×1 convolution to adjust the number of channels to 1 / 4 of the input channel number. Finally, the results of all branches are concatenated along the channel dimension to obtain the final output. In the second and fourth layers of MD 2 TL-GCN, MS-TC uses convolution with a stride of 2 in the temporal dimension, achieving halving of the temporal dimension, while the other layers maintain the size of the temporal dimension unchanged.

[0056] S3: Multi-stream data input result fusion.

[0057] In the skeleton action recognition task, in addition to using the initial node data, skeleton data and motion information are also very important. By using a multi-stream architecture composed of node stream, skeleton stream and their motion information, the network recognition ability can be greatly enhanced. Figure 5 Common node data, skeleton data, and the relative node data and relative skeleton data proposed by the present invention are exemplified in

[0058] Each edge obtained by the natural connection of human body joints is used as a skeleton, which is represented by the difference in 3D coordinates of the two adjacent nodes it associates. For the calculation of node motion information, it can be obtained by the 3D coordinate difference of the joint points in adjacent time frames.

[0059] In addition, the position of the node relative to the human body center of gravity and the position of the skeleton relative to the human body center of gravity are also very important for action recognition, so we introduce relative node data and relative skeleton data. As Figure 5As shown in (d) and (e), they are obtained by calculating the relative coordinate vectors between each node and the central node and the relative coordinate vectors between each bone and the central bone respectively.

[0060] Six-stream network architecture 6s-MD 2 TL-GCN is as Figure 6 shown, and respectively includes node data stream, relative node data stream, node motion data stream, bone data stream, relative bone data stream and bone motion data stream. They can be trained and tested in parallel, and finally the softmax scores of the six streams are added and fused as the final classification result. The effects of the present invention will be described in detail below in conjunction with the embodiments of the target detection effect diagram.

[0061] Verify the MD proposed by the present invention in Table 1 2 the effectiveness of the TL-GC module. We reproduce the 5-layer ST-GCN network as a comparative baseline network, denoted as ST-GCN_L5. The experimental results are shown in Table 1. In Table 1, "ST-GCN_L5+MS-TC" means adding a residual connection to ST-GCN_L5 and replacing its temporal convolution with MS-TC; "ST-GCN_L5+MS-TC+iMD 2 " means that on the basis of "ST-GCN_L5+MS-TC", the first i graph convolutional layers are replaced with MD2TL-GC. At the same time, we also reproduce the channel-specific graph convolutional network CTR-GCN, and compare the CTR-GCN with only 5-layer graph convolution, denoted as "CTR-GCN_L5" and the CTR-GCN with 10-layer graph convolution to verify the superiority of the multi-dimensional dynamic topology graph convolution MD2TL-GC proposed by the present invention.

[0062] As can be seen from Table 1, when the graph convolutional layers of ST-GCN_L5 are gradually replaced with MD2TL-GC, the recognition accuracy of the network is gradually improved. After all replacements, the recognition accuracy of MD2TL-GCN is 3.23% higher than that of ST-GCN_L5+MS-TC. Compared with CTR-GCN_L5 with 5-layer graph convolution, the recognition accuracy of MD2TL-GCN is 0.56% higher, and 0.21% higher than that of CTR-GCN with 10-layer graph convolution. These experimental results fully verify the MD 2 the effectiveness of TL-GC.

[0063] Table 1 MD 2 Effectiveness analysis of TL-GC

[0064]

[0065] Tables 2 and 3 compare the method proposed by the present invention, namely MD 2Comparison of the recognition accuracy of TL-GCN and other methods on two large-scale skeleton action recognition datasets, NTU-RGB+D and NTU-RGB+D 120. The accuracies of other methods are all from the results reported in their papers. As can be seen from the results in Table 2 and Table 3, MD 2 TL-GCN outperforms all current mainstream methods in almost all metrics. Especially in the X-Sub evaluation metric of the NTU-RGB+D 120 dataset and the CS evaluation metric of the NTU-RGB+D dataset, compared with the current state-of-the-art method CTR-GCN, the recognition accuracy has increased by 0.39% and 0.24% respectively.

[0066] In Figure 7 we also further compared the CS recognition accuracy of MD 2 TL-GCN and ST-GCN, AAGCN, and CTR-GCN algorithms on the "writing", "reading", "eating", "playing with mobile phone or tablet", "touching head or headache", "sneezing or coughing", "drinking water", and "nausea or vomiting" classes in the NTU-RGB+D dataset using only node flow as input. These classes are action classes with relatively low performance reported in the literature and are recognized as relatively difficult actions to identify. The experimental comparison results are as Figure 7 shown. MD 2 TL-GCN outperforms the other three models in terms of performance on all difficult classes. Especially in the "playing with mobile phone or tablet", MD 2 TL-GCN has increased the recognition accuracy by 20%, 8.73%, and 3.28% compared with the ST-GCN, AAGCN, and CTR-GCN algorithms respectively. This shows that compared with other graph convolution algorithms, the MD 2 TL-GCN we proposed can effectively extract the features of various subtle motion patterns in long-term actions and classify them more accurately.

[0067] Table 2 Comparison of NTU-RGB+D dataset

[0068]

[0069] Table 3 Comparison of NTU-RGB+D 120 dataset

[0070]

[0071] A skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution proposed by the present invention can dynamically model the human skeleton topology with both temporal specificity and channel specificity. In the dynamic temporal dimension specific topology modeling, by learning the dynamic skeleton topology graph, the global spatio-temporal dynamic information is enhanced. In addition, we also propose relative node data flow and relative bone data flow to supplement the skeleton information. The experimental results on two large-scale public skeleton action recognition datasets NTU-RGB+D and NTU-RGB+D 120 prove that the dynamic skeleton topology and multi-dimensional specific topology can learn more subtle action differences, and the proposed network model MD 2 TL-GCN can outperform all current mainstream methods when only stacking 5 layers.

Claims

1. A skeleton action recognition method based on multi-dimensional dynamic topology learning graph convolution, comprising the following steps: S1. In the input skeleton sequence data, it is first input into the multi-dimensional dynamic topology learning graph convolution for spatial modeling. This graph convolution consists of three main parts, namely the pure node topology learning graph convolution, simply referred to as J-GC; the dynamic temporal specificity topology learning graph convolution, abbreviated as DTW-GC; and the channel specificity topology learning graph convolution, simply called CW-GC. Specifically, after the input features are divided by channels, they are respectively input into J-GC and DTW-GC to learn the pure node topology structure and the dynamic temporal specificity topology structure in parallel. Then, the learned feature maps are concatenated on the channels, and then a 1×1 convolution is used to fully interact and fuse the time dimension and space dimension features. After that, it is input into CW-GC for channel specificity topology learning to obtain a feature map with rich spatial topology structure information. S2. The spatial modeling features obtained in S1 are input into the multi-scale time convolution, simply called MS-TC. This time convolution consists of four branches. Among them, the first 3 branches first use 1×1 convolution to reduce the number of channels to 1 / 4 of the input channel number, and then use dilated convolutions with a dilation rate of 1 and a kernel size of 5, a dilation rate of 2 and a kernel size of 5, and a temporal dimension max pooling operation respectively to obtain features with different temporal receptive fields. The fourth branch, as a residual connection, only uses 1×1 convolution to adjust the number of channels to 1 / 4 of the input channel number. Finally, the results of all branches are concatenated on the channels to obtain the final output, thus obtaining the multi-scale spatio-temporal feature map. S3. The skeleton sequence input respectively includes 6 types of input data, namely node flow data input, bone flow data input, relative node flow data input, relative bone flow data input, node motion flow data input, and bone motion flow input. The 6 types of data inputs respectively repeat the spatio-temporal modeling process of S1 and S2, add and fuse the 6 different feature outputs, and then train them to obtain the final classification result.

Citation Information

Patent Citations

  • Human body behavior recognition method and system based on graph convolution network

    CN110796110A

  • Method for constructing human body behavior recognition model based on graph convolution network

    CN111652124A

Cited By

  • Human skeleton extraction method based on spatio-temporal feature enhancement with timing enhancement constraint

    CN122510955A