Skeleton action recognition method based on space-time dependency enhanced network
By constructing the AMGC-IBM-CNN network, the difficulty of cross-time information exchange caused by the separation of time and space convolution layers in the traditional model is solved, and the effective acquisition of cross-time and space dependencies of the joint nodes and the accuracy of action recognition is improved.
Patent Information
- Application Number
- CN202510154523.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
AI Technical Summary
In traditional models, the separation of time and space convolution layers hinders the exchange of cross-time and space information and makes it difficult to obtain cross-time and space dependencies of link nodes.
The AMGC-IBM-CNN network is constructed, and the spatial and temporal graphs are extracted through the AMGC network, and combined with the IBM module and the CNN network to realize direct communication of cross-temporal information.
Eliminate redundant dependencies during long-distance joint feature aggregation, improve the model receptive field, pay more flexibly to different joint nodes, and improve the accuracy of action recognition.
Smart Images

Figure CN120088855A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of action recognition, and particularly relates to a skeleton action recognition method based on a spatio-temporal dependence enhancement network. Background Art
[0002] Human action behavior recognition is a task with many practical significances; in particular, skeleton-based human action recognition.
[0003] Some scholars have proposed graph convolutional neural networks, which use skeleton spatio-temporal graphs to model human joints; the skeleton spatio-temporal graph is an aggregation of body bone graphs in different time frames, and is a more specific representation of time and space information in joints; in recent years, behavior recognition methods based on graph convolutional neural networks have been continuously improved and optimized, and the recognition effect has also been greatly improved; however, most existing graph convolutional neural network methods use temporal convolution and spatial convolution separately, similar to factorized 3D convolution; the typical model structure is to first use a graph convolutional layer to obtain the spatial relationship within each time frame; then use a recurrent or 1D convolutional layer to simulate temporal dynamics.
[0004] Traditional models such as the ST-GCN model proposed by Yan et al. and the Dynamic Graph Convolutional Network (DynamicGCN) proposed by Ye et al.; Figure 4 As shown in the upper half of the figure, although this method allows obtaining a larger receptive field, the separate temporal and spatial convolutional layers prevent direct cross-spatio-temporal information exchange, and it is difficult to obtain the cross-spatio-temporal dependence relationship of joint points; obtaining this cross-spatio-temporal dependence relationship of joint points is a crucial step in aggregating joint features on the spatio-temporal graph; for example, for the action of "getting up", it is generally completed by the upper body and the lower body together, and there is a close relationship between the forward movement of the upper body and the future standing action of the lower body. Summary of the Invention
[0005] Aiming at the deficiencies of the existing methods, the present invention solves the problem that the separate temporal and spatial convolutional layers in traditional models prevent cross-spatio-temporal information exchange and it is difficult to obtain the cross-spatio-temporal dependence relationship of joint points.
[0006] The technical solution adopted by the present invention is: a skeleton action recognition method based on a spatio-temporal dependence enhancement network includes the following steps:
[0007] Step 1: Obtain a graph containing skeleton action joint points;
[0008] Step 2: Construct the AMGC-IBM-CNN network: Using joints as nodes and bones as edges to build a spatio-temporal graph, and using the AMGC network to extract global feature vectors and local feature vectors from the spatio-temporal graph. After inputting the global feature vectors into the IBM module and then into the CNN network, the output feature vectors of the CNN network are added to the local feature vectors, followed by global average pooling, FC, and finally the action category is output through a classifier.
[0009] As a preferred embodiment of the present invention, the AMGC network includes two branches. The first branch is the operation of Conv1 and the first adjacency matrix, and the second branch is the operation of Conv1 and the second adjacency matrix.
[0010] As a preferred embodiment of the present invention, the formula for the adjacency matrix is:
[0011]
[0012] where d(v i ,v j ) gives the shortest distance between v i and v j . v i and v j are the i-th and j-th joint points respectively, and v i , v j ∈V; A K is the K-order generalization of the adjacency matrix A to other neighborhoods; K is the shortest distance between joint points.
[0013] As a preferred embodiment of the present invention, the first adjacency matrix K = 3, and the second adjacency matrix K = 5.
[0014] As a preferred embodiment of the present invention, the output of the AMGC network is:
[0015]
[0016] where W K represents the weight function, X represents the input of the AMGC network, θ K represents the learnable adjacency matrix corresponding to each K layer, and σ(.) represents the activation function.
[0017] As a preferred embodiment of the present invention, the output of the IBM module is:
[0018]
[0019] where A’ is the adjacency matrix predicted from the global features of the AMGC network, and X' represents the input bone sequence.
[0020] As a preferred embodiment of the present invention, the output of the CNN network is:
[0021]
[0022] Wherein, X represents the output features of the IBM module, and Y m (v ti ) represents the output features of the CNN network; W(m, c, q, j) represents the weights of the 2D convolution, m represents the filter, c represents the channel, b is the convolution bias, t and q represent the time frame numbers, i and j represent the joint numbers, and γ is a predefined parameter representing the size of the time window.
[0023] As a preferred embodiment of the present invention, the classifier is a Softmax classifier.
[0024] As a preferred embodiment of the present invention, a skeleton action recognition system based on a spatio-temporal dependence enhancement network includes: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement a skeleton action recognition method based on a spatio-temporal dependence enhancement network.
[0025] As a preferred embodiment of the present invention, a computer-readable medium storing computer program code, the computer program code implementing a skeleton action recognition method based on a spatio-temporal dependence enhancement network when executed by a processor.
[0026] Advantages of the present invention:
[0027] 1. The present invention constructs an AMGC network to eliminate the redundant dependence relationship on nearby joints during the aggregation of distant joint features; solves the problem that the local features dominate when aggregating global features, restricting the receptive field of the model;
[0028] 2. Enables effective long-range modeling with a larger K value; at the same time, through the weight value W K biased weighting of the features obtained at each scale, enabling the model to more flexibly focus on different joints during the recognition of different actions;
[0029] 3. Constructs a dense adjacency matrix, which can not only reflect the joint structure information within a single frame, but also reflect the joint structure information in different time frames;
[0030] 4. Through spatio-temporal modeling of the skeleton data by the AMGC network, combining the AMGC network with the CNN network using the data bridge module, an AMGC-IBM-CNN network is constructed, and through multiple inter-frame edges connection, direct communication of cross-spatio-temporal information is achieved. Description of the Drawings
[0031] Figure 1 is the overall framework diagram of the present invention;
[0032] Figure 2 This is a comparison diagram of feature aggregation between the present invention and existing methods;
[0033] Figure 3 This is a schematic diagram of the dense matrix of the present invention;
[0034] Figure 4 This is a comparison schematic diagram of spatio-temporal information exchange between the spatio-temporal information exchange of the present invention and the traditional model. Detailed implementation manners
[0035] The present invention will be further described below with reference to the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.
[0036] As Figure 1 shown, a skeleton action recognition method based on a spatio-temporal dependence enhanced network includes the following steps:
[0037] Step 1: Obtain a graph containing skeleton action joint points;
[0038] The graph of action joint points is extracted frame by frame from the video.
[0039] Step 2: Construct a spatio-temporal dependence enhanced network (AMGC-IBM-CNN). Among them, the construction of the adaptive multi-scale graph convolutional network (AMGC) includes:
[0040] Step 21: In AMGC, first, a spatio-temporal graph is established with joints as nodes and bones as edges for the skeleton data, and then the features of the joints are aggregated through the edges of the spatio-temporal graph to capture the structural relationship between the joints;
[0041] The adaptive multi-scale graph convolutional network (AMGC) adopts a two-layer parallel structure. In the upper layer, K = 3 is taken to extract local information; in the lower layer, K = 5 is taken to extract global information, where K is the shortest distance between joint points.
[0042] The AMGC network includes two branches. The first branch is the first graph convolution, and the first graph convolution is the operation of Conv1 and the adjacency matrix A 3 to obtain a local feature vector; the second branch is the second graph convolution, and the second graph convolution is the operation of Conv1 and the adjacency matrix A 5 to obtain a global feature vector;
[0043] The size and stride of the Conv1 convolution kernel are default values;
[0044] The input of the AMGC network is the feature vector C in xTxV, C inis the number of input feature channels, T is the time dimension, V is the joint feature; the feature vector C in xTxV is obtained from the graph of action joints;
[0045] The output local feature vector and global feature vector of the AMGC network, and the global feature vector is the skeleton sequence X;
[0046] The K - adjacency matrix A K is defined as:
[0047]
[0048] where d(v i , v j ) gives the shortest distance between v i and v j , v i , v j are the i - th and j - th joints respectively, v i , v j ∈V; A k is the K - order generalization of the adjacency matrix A to other neighborhoods; K is the shortest distance between joints, K = 3 and 5; at this time, aggregating multi - scale structural information can be formulated as:
[0049]
[0050] where Y is the output feature vector of the AMGC network. When K = 3, Y is the local feature vector, and when K = 5, Y is the global feature vector; W K represents the weight function, X represents the input of the AMGC network, θ K represents the learnable adjacency matrix corresponding to each K - layer, and σ(.) represents the activation function.
[0051] Compared with the traditional model, the method of the present invention eliminates the redundant dependence relationship of nearby joints during the aggregation of long - distance joint features; therefore, it solves the problem that the local features dominate when aggregating global features, restricting the receptive field of the model, and enables effective long - range modeling with a larger K value; at the same time, through the weight value W K biased weighting of the features obtained at each scale, the model can more flexibly focus on different joints during the recognition of different actions.
[0052] Such as Figure 2The first line indicates the method of using a high - order adjacency matrix for feature aggregation; the second line indicates the method of using a multi - scale approach for feature aggregation; the third line indicates the method of using the AMGC method of the present invention for feature aggregation. Using the multi - scale approach for feature aggregation solves the problem that the weights of neighboring nodes are large and the weights of distant nodes are small compared with using the high - order adjacency matrix method. Compared with the multi - scale method, AMGC no longer uses the same weights when aggregating joint features and pays more attention to important joint point information.
[0053] Step 22: Use the data bridge module to construct a dense adjacency matrix, as Figure 3 shown;
[0054] Construct a data bridge module (IBM) to allow the structural information between joints to be reflected in the pseudo - image. First, predict the adjacency matrix A of the joints from the bone sequence of AMGC, and use this to construct a dense adjacency matrix The formula is:
[0055]
[0056] where A’ is the adjacency matrix predicted from the global features of the AMGC network;
[0057] Specifically, each element in the dense adjacency matrix The adjacency matrix constructed in this way can not only reflect the joint structure information within a single frame, but also reflect the joint structure information in different time frames; In Figure 3 the dense adjacency matrix constructed is multiplied by the joint coordinate tensor. In this way, the joint structure information can be reflected in the convolutional neural network. The obtained adjacency matrix is not fixed and will adaptively change according to different input data. Even for joint points that are not naturally connected in the human body, as long as there is a correlation in actions, this correlation will be captured and reflected in the input feature map of the convolutional neural network through the predicted adjacency matrix.
[0058]
[0059] where X' represents the input bone sequence; Y is the feature vector processed by the data bridge module, with dimensions (N, C, T, V), with the number of joints V as the width, the number of time frames T as the height, and N as the input batch. Each joint point will be closer to other relevant joint points within the row and between adjacent rows to facilitate the convolutional neural network to aggregate the information between joints.
[0060] Step 23: Cross - space - time information exchange modeling, as Figure 1 shown;
[0061] Simple spatio-temporal modeling of skeleton data is performed through the AMGC network. At this time, the inter-frame edges in the spatio-temporal graph only connect the same joints in adjacent frames, and the connections between different joints are cut off. Combining with the data bridge module, the AMGC network is combined with the convolutional neural network CNN, integrating a spatio-temporal dependence enhancement network. Through multiple inter-frame edge connections, direct communication of cross-spatio-temporal information is achieved.
[0062] Figure 1 The convolutional neural network CNN has a total of four stages, and each stage has two bottleneck layers. The number of input channels in the four stages is 128, 256, 512, and 1024 respectively. Average pooling operations are added in the second and third stages to reduce the input features by half.
[0063] CNN maps the skeletal data into a pseudo-image, aggregating joint features through 2D or 3D convolution. The node distribution in the pseudo-image is arranged in sequence according to the joint index order, and this index cannot express the structural information of human joints. For example, "left hand" and "left thumb" are naturally connected in the human body. However, in the NTU RGB+D and NTU-120RGB+D datasets, their joint indices are 12 and 25 respectively. Therefore, it is difficult for the CNN network to obtain joint-like structural information from the pseudo-image arranged by joint indices. Because of the different processing methods for skeleton data, directly transmitting data between the AMGC network and CNN will cause the loss of important joint structure information.
[0064] Before connecting to the convolutional neural network, the sampling area B(v t,i ) in the time dimension of the individual AMGC network can be formulated as:
[0065]
[0066] Among them, i represents the joint label; q and t represent time frames, γ is a predefined parameter representing the size of the time window; B(v t,i ) represents the sampling area centered on the i-th joint point in the t-th frame; v q,i represents the i-th joint point in the q-th frame.
[0067] Combining AMGC and CNN through the IBM module, the sampling area B(v t,i ) of the integrated spatio-temporal dependence enhancement network can be formulated as:
[0068]
[0069] Among them, both i and j represent different joint labels.
[0070] The output formula of the CNN network is:
[0071]
[0072] Among them, X represents the output features of the IBM module, and Y m (v ti ) represents the output features of the CNN network; W(m, c, q, j) represents the weights of the 2D convolution, where m represents the filter, c represents the channel, b is the convolution bias, t and q represent the time frame numbers, i and j represent the joint numbers, and γ is a predefined parameter representing the size of the time window; it can be seen from this formula that the output at node v ti is the weighted sum of all nodes within the spatio-temporal region, reflecting the direct information exchange of joints across space and time
[0073] After adding the output features of the CNN network and the local features of the AMGC network, the action category is output through the FC and Softmax classifiers.
[0074] Such as Figure 4 is a comparison schematic diagram of the spatio-temporal information exchange between the method of the present invention and traditional models.
[0075] Test results:
[0076] After training the model using two mainstream skeleton datasets (NTU RGB+D and NTU-120RGB+D), and testing through the test set, any skeleton data can be input into the network to recognize the actions in the skeleton data; to ensure fairness in comparison, the present invention adopts the same stream fusion strategy as other methods; specifically, the experimental results of the joint stream, skeleton stream, joint movement stream, and skeleton movement stream are obtained respectively, and the four experimental results are fused to obtain the final prediction; this method allows for making full use of the advantages of each data stream to improve the overall recognition effect. The comparison results are shown in Tables 1 and 2:
[0077] Table 1 Comparison of the present invention with advanced methods on the NTU-RGB+D dataset
[0078]
[0079] Table 2 Comparison with advanced methods on the NTU120-RGB+D dataset
[0080]
[0081] Table 1 shows the experimental results under two evaluation criteria, cross-subject (CS) and cross-view (CV), on the NTU RGB+D dataset; Table 2 shows the experimental results under two evaluation criteria, cross-subject (CS) and cross-setting (CE), on the NTU120-RGB+D dataset.
[0082] Inspired by the above-described ideal embodiments of the present invention, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A skeleton action recognition method based on spatiotemporal dependency enhancement network, characterized in that: The following steps are involved: Step 1: Get a graph containing skeleton action joint points; Step 2: Construct an AMGC-IBM-CNN network: Use joints as nodes and bones as edges to build a space-time graph. Use the AMGC network to extract global and local feature vectors from the space-time graph. Input the global feature vector into the IBM module and then into the CNN network. Add the output feature vector of the CNN network to the local feature vector, perform global average pooling and FC, and then output the action category through the classifier.
2. The skeleton action recognition method based on spatiotemporal dependency enhanced network according to claim 1 is characterized in that: The AMGC network includes two branches, the first branch is Conv1 and the first adjacency matrix operation, and the second branch is Conv1 and the second adjacency matrix operation.
3. The skeleton action recognition method based on spatiotemporal dependency enhanced network according to claim 2 is characterized in that: The formula for the adjacency matrix is: Among them, d(v i , v j ) gives v i and v j The shortest distance between i 、v j are the i-th and j-th joint points respectively, v i 、v j ∈V;A K It is the K-order generalization of the adjacency matrix A to other neighborhoods; K is the shortest distance between joint points.
4. The skeleton action recognition method based on spatiotemporal dependency enhanced network according to claim 3 is characterized in that: The first adjacency matrix K=3, and the second adjacency matrix K=5.
5. The skeleton action recognition method based on spatiotemporal dependency enhanced network according to claim 1, characterized in that: The output of the AMGC network is: Among them, W K represents the weight function, X represents the input of the AMGC network, θ K represents the learnable adjacency matrix corresponding to each K layer, and σ(.) represents the activation function.
6. The skeleton action recognition method based on spatiotemporal dependency enhanced network according to claim 1, characterized in that: The output of the IBM module is: Among them, A' is the adjacency matrix predicted from the global features of the AMGC network, and X' represents the input bone sequence.
7. The skeleton action recognition method based on spatiotemporal dependency enhanced network according to claim 1, characterized in that: The output of the CNN network is: Among them, X represents the output feature of the IBM module, Y m (v ti ) represents the output features of the CNN network; W(m, c, q, j) represents the weight of the 2D convolution, m represents the filter, c represents the channel, b is the convolution bias, t and q represent the time frame number, i and j represent the joint point number, and γ is a predefined parameter representing the size of the time window.
8. The skeleton action recognition method based on spatiotemporal dependency enhanced network according to claim 1, characterized in that: The classifier is Softmax classifier.
9. A skeleton action recognition system based on a spatiotemporal dependency enhanced network, characterized in that: include: a memory for storing instructions executable by a processor; A processor, configured to execute instructions to implement the skeleton action recognition method based on a spatiotemporal dependency enhanced network as described in any one of claims 1 to 8.
10. A computer readable medium storing computer program code, characterized in that: When the computer program code is executed by a processor, the method for skeleton action recognition based on a spatiotemporal dependency enhanced network as described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Non-contact heart rate measurement method and system based on space-time enhancement network
CN121370107A
A non-contact heart rate measurement method and system based on spatiotemporal augmentation networks
CN121370107B