Human Skeleton Action Recognition Method Based on Hierarchical Spatiotemporal Attention Network

Through the human body skeleton action recognition method based on hierarchical spatiotemporal attention network, problems such as redundant data interference, feature confusion, and large calculation amount in the prior art are solved, and the effect of efficient extraction of multi-level features and global time features is achieved.

CN116343338BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310351817.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-04
Publication Date
2025-07-01
Estimated Expiration
2043-04-04

AI Technical Summary

Technical Problem

The existing human skeleton action recognition methods have problems such as redundant data interference, confusion between temporal features and spatial features, large calculations, and difficulty in extracting multi-level features and global temporal features.

Method used

A method of human skeleton action recognition based on hierarchical space-time attention network is proposed. By dividing the human skeleton sequence into time segments, a hierarchical space attention module and a hierarchical time attention module are constructed, spatial features and temporal features are separated and extracted, and a full connection layer is used as a classification module.

Benefits of technology

It reduces the interference of redundant data on spatial and temporal features, avoids the confusion between temporal and spatial features, reduces the amount of calculation, extracts multi-level human body movement features, and improves the ability to obtain global temporal features of the complete human skeleton sequence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116343338B_ABST
    Figure CN116343338B_ABST
Patent Text Reader

Abstract

The present invention discloses a human skeleton action recognition method based on a hierarchical spatio-temporal attention network, which mainly solves the problems of confusing temporal features with spatial features and lacking multi-level action representations in the prior art. The implementation solution is as follows: Obtain the human skeleton sequence to form human skeleton data, and construct a training sample set and a test sample set; Divide the joint points in the training sample set and the test sample set into body parts, and divide the human skeleton sequence into time segments; Construct a human skeleton action recognition model composed of a position encoding module, a hierarchical spatio-temporal attention module, and a classification module; Use the divided training sample set to train the human skeleton action recognition model; Input the divided test sample set into the trained action recognition model to obtain the human skeleton action recognition result. The present invention can extract multi-level action features of human actions, and has high recognition accuracy, and can be widely applied to fields such as video understanding, human-computer interaction, intelligent monitoring, and medical rehabilitation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a method for human skeleton action recognition, which can be used in video understanding, human-computer interaction, intelligent monitoring, and medical rehabilitation. Background Art

[0002] The purpose of human skeleton action recognition is to determine the type of human actions occurring in a clipped human skeleton sequence. Human skeleton action recognition has broad application prospects in fields such as video understanding, human-computer interaction, intelligent monitoring, and medical rehabilitation. With the progress of pose estimation algorithms and body sensor camera technologies, it has become possible to accurately and quickly obtain human skeleton data at a low cost, which has promoted the application of artificial intelligence methods in human skeleton action recognition. A human skeleton sequence is represented by the two-dimensional or three-dimensional coordinates of human joint points, has high-level feature representation while having low complexity, and is robust to external conditions such as background and lighting. These advantages have continuously attracted researchers to develop action recognition methods based on human skeleton sequences.

[0003] Currently, common human skeleton action recognition methods can be divided into the following categories: action recognition based on handcrafted features, action recognition based on convolutional neural networks, action recognition based on long short-term memory networks, action recognition based on graph convolutional neural networks, and action recognition based on attention networks, where:

[0004] Handcrafted features include joint point coordinates, joint point coordinate variants, and statistical features, etc. These features retain all the information of the human skeleton sequence and have high interpretability. These low-level features can also be used as different input streams in deep learning methods, but these features are difficult to reflect the deep connections between joint points during actions.

[0005] Methods based on convolutional neural networks represent the skeleton sequence as a matrix, then quantize the matrix into an image and input it into a convolutional neural network model for feature extraction and action recognition, but convolutional neural networks are difficult to handle the complex graph topology relationship of human skeletons.

[0006] Methods based on long short-term memory networks are suitable for dealing with time series problems. However, due to the large computational amount of long short-term memory networks and the inability to process in parallel, the practical application value of long short-term memory networks in skeleton action recognition tasks is affected.

[0007] Methods based on graph convolutional neural networks overcome the defect that convolutional neural networks are difficult to handle graph topology relationships and are the mainstream methods for current skeleton action recognition. However, the aggregation of features by graph neural networks expands point by point along the graph topology structure and has the defect of locality, while in human actions, even joint points that are far apart may have strong correlations.

[0008] The method based on the attention network, whose core idea is to simulate human attention, can learn the global dependencies among input data with relatively low computational complexity and good parallelism, so as to discover the hidden features in human actions. In addition, since this method does not need to know the internal relationships between data, it provides greater flexibility for mining discriminative action representations in human body movements.

[0009] Some scholars designed three self-supervised learning tasks in STST: Spatial-Temporal Specialized Transformer for Skeleton-based Action Recognition published in ACM MULTIMEDIA 2021 to capture the features of skeleton sequences to assist model training. However, due to the relatively complex model design, the training takes a long time.

[0010] In addition, the above existing human skeleton action recognition methods also have the following four common problems:

[0011] First, the minimum unit and sampling frequency for modeling human actions are not selected specifically according to the process of human movement, resulting in being easily interfered by noise and redundant data when aggregating spatio-temporal features and prone to information loss.

[0012] Second, the spatio-temporal differences in the human skeleton action recognition problem are ignored, confusing the time features and space features of human skeleton data.

[0013] Third, only processing human skeleton data at the joint point level is difficult to reflect the deep features of human actions and has poor interpretability.

[0014] Fourth, it is difficult to model the global correlation of skeleton data during the action process, and the mined time features are local, that is, only the local time relationships within a single frame or several frames can be analyzed by one layer of the model, making it difficult to obtain the global time features of the complete sequence in the human skeleton action recognition task. Summary of the Invention

[0015] Aiming at the deficiencies of the above existing technologies, the purpose of the present invention is to propose a human skeleton action recognition method based on a hierarchical spatio-temporal attention network to reduce the interference of redundant data on the aggregated spatio-temporal features, avoid the confusion between time features and space features, be able to extract multi-level features of human actions while reducing the amount of calculation, and improve the ability to obtain spatio-temporal features of the complete human skeleton sequence.

[0016] The technical solution of the present invention to achieve the above purpose includes the following:

[0017] (1) Obtain the human skeleton sequence, form the human skeleton data with the human skeleton sequence, divide the human skeleton data into a training sample set and a test sample set according to a ratio of 2:1, and perform view standardization and scale normalization on all human skeleton sequences in the training sample set and the test sample set to obtain the standardized human skeleton sequence;

[0018] (2) Divide the human joint points into body parts, convert all human skeleton sequences with more than 120 frames into 120 frames using the downsampling method, convert all human skeleton sequences with less than 120 frames into 120 frames using the upsampling method, and then evenly divide all 120-frame human skeleton sequences into time segments of 6 frames;

[0019] (3) Construct a human skeleton action recognition model based on a hierarchical spatio-temporal attention network:

[0020] (3a) Construct a position encoding module composed of the cascade of time position encoding and space position encoding;

[0021] (3b) Cascade the hierarchical spatial attention module and the hierarchical temporal attention module, and add layer normalization and residual connection to form a spatio-temporal attention layer. Then stack multiple spatio-temporal attention layers to construct a hierarchical spatio-temporal attention module;

[0022] (3c) Select a fully connected layer as the classification module;

[0023] (3d) Cascade the position encoding module, the hierarchical spatio-temporal attention module and the classification module to form a human skeleton action recognition model based on a hierarchical spatio-temporal attention network;

[0024] (4) Input the training sample set into the human skeleton action recognition model, and use the Adam optimization algorithm to perform iterative training for 90 rounds to obtain a trained human skeleton action recognition model;

[0025] (5) Input the test sample into the trained action recognition model to obtain the action recognition results of the test sample set.

[0026] The present invention has the following advantages compared with the prior art:

[0027] 1. Reduce the interference of redundant data on the aggregated spatio-temporal features

[0028] The present invention reduces the interference of data redundancy caused by the similarity of frames within the time segment on the aggregated spatio-temporal features by dividing the human skeleton sequence into time segments and calculating the attention in the deep layer of the hierarchical spatio-temporal attention module in units of time segments.

[0029] 2. Avoid the confusion between time features and space features

[0030] In view of the spatio-temporal differences in human skeleton action recognition problems, this invention distinguishes the spatio-temporal differences in human skeleton action recognition problems, constructs a hierarchical spatial attention module and a hierarchical temporal attention module, separates the extraction processes of spatial features and temporal features, avoids the confusion between temporal features and spatial features during the feature extraction process, and improves the action recognition accuracy.

[0031] 3. It can extract multi-level human action features while reducing the computational load

[0032] Compared with only using joint points as the minimum unit for action recognition, since this invention calculates attention in units of joint points in the shallow layer of the hierarchical spatio-temporal attention module and calculates attention in units of body parts in the deep layer of the hierarchical spatio-temporal attention module, the process of calculating attention on multiple joint points can be transformed into the process of calculating attention on a small number of body parts, reducing the computational load; at the same time, since action features at the joint point level and action features at the body part level are extracted, and body parts that are more in line with human intuition are used as action recognition units, this invention not only has stronger interpretability, but also reduces the impact of single joint point noise in the data on the action recognition performance.

[0033] 4. It improves the ability to obtain global temporal features of the complete human skeleton sequence

[0034] Based on the attention model, this invention uses the characteristic that the attention model can model global information to globally consider the inter-frame relationship in the complete human skeleton sequence, explore the deep spatio-temporal information during the human action process, obtain the temporal relationship between all frames of the complete human skeleton sequence, solves the problem that existing methods can only analyze the local temporal relationship within a single frame or a few frames, and improves the human skeleton action recognition accuracy. Brief Description of the Drawings

[0035] Figure 1 is the implementation flowchart of this invention;

[0036] Figure 2 is the structural schematic diagram of the human skeleton action recognition model based on the hierarchical spatio-temporal attention network.

[0037] Figure 3 is the structural schematic diagram of the hierarchical spatial attention module in this invention. Detailed Embodiment

[0038] The following describes the implementation steps of this invention in detail with reference to the drawings.

[0039] With the development of science and technology, humans have gained more understanding of human skeleton action recognition and designed many artificial intelligence algorithms to capture the spatio-temporal features of skeleton sequences. However, when using these algorithms for human skeleton action recognition, the classification accuracy and efficiency are often poor due to the failure to select appropriate data inputs according to the data characteristics and the lack of targeted design of algorithm models that can explore global spatio-temporal features. In response to this situation, through exploration and experimentation, the present invention proposes a method for human skeleton action recognition based on a hierarchical spatio-temporal attention network for human skeleton action recognition.

[0040] See Figure 1 , the implementation steps of this example are as follows:

[0041] Step 1: Obtain the human skeleton sequence and standardize it.

[0042] 1.1) Obtain the human skeleton sequence:

[0043] Use a depth camera to obtain the human skeleton sequence, or use a pose estimation method to extract the human skeleton sequence from a human action video;

[0044] Use the obtained human skeleton sequences to form a human skeleton database or obtain a publicly available human skeleton database from the network, and divide the human skeleton data into a training sample set and a test sample set in a ratio of 2:1;

[0045] 1.2) Standardize the human skeleton sequence:

[0046] Since the method of obtaining the human skeleton sequence affects the perspective and position of the human skeleton sequence and cannot be directly input into the action recognition model, it is necessary to perform perspective normalization and scale normalization on the joint point coordinates of the human skeleton sequence. In this example, all human skeleton sequences in the training sample set and the test sample set are subjected to perspective normalization and position normalization of the joint point coordinates to obtain the standardized human skeleton sequence. By performing the same perspective normalization and position normalization operations on all joint point coordinates in the human skeleton sequence, the relative positions of the joint point coordinates can be ensured to remain unchanged. The specific implementation is as follows:

[0047] 1.2.1) Let the coordinates of the nth joint point in the human skeleton sequence at the pos-th frame be The coordinates of the joint point "hip center" in the pos-th frame are T0 is the total number of frames of the human skeleton sequence, is the coordinate of the nth joint point in the pos-th frame of the human skeleton sequence, ψ is the set of human torso joint points, then the torso matrix M of the human skeleton sequence can be expressed as:

[0048]

[0049] 1.2.2) Calculate the average value o of the coordinates of the joint point "hip center" in all frames of the human skeleton sequence according to the set parameters:

[0050]

[0051] 1.2.3) Calculate the first principal component of the torso matrix M by the principal component analysis method and normalize it to obtain the Z-axis direction vector γ in the new coordinate system. Calculate the second principal component of M and normalize it to obtain the Y-axis direction vector β in the new coordinate system. Calculate the cross product α of β and γ;

[0052] 1.2.6) Use the dimensionality elevation [x, y, z, 1] of the original joint point coordinates [x, y, z] of the human skeleton sequence and α, β, γ, o to calculate the dimensionality elevation [x', y', z', 1] of the joint point coordinates [x', y', z'] of the standardized human skeleton sequence, which is expressed as the product of the following four-dimensional matrix and [x, y, z, 1]:

[0053]

[0054] In the formula, (α - o) T , (β - o) T and (γ - o) T are the X-axis perspective normalization vector, the Y-axis perspective normalization vector, and the Z-axis perspective normalization vector respectively. o is the coordinate of the standardized coordinate origin in the original coordinate system, and T represents matrix transpose.

[0055] In this embodiment, three mainstream skeleton action recognition benchmark datasets are selected for experiments, including the NW-UCLA dataset, the NTU-RGB+D 60 dataset, and the NTU-RGB+D 120 dataset;

[0056] The NW-UCLA dataset contains 1494 video clips, covering 10 action categories, and is simultaneously captured by three Kinect cameras from multiple viewpoints; each action is performed by 10 different subjects, and each subject has 20 joint positions; the videos captured by the first two cameras are used as the training sample set, and the videos captured by the third camera are used as the test sample set;

[0057] The NTU-RGB+D 60 dataset is a large-scale human action recognition dataset containing 56880 skeleton action sequences, with 60 action categories, and is simultaneously captured by three Kinect v2 cameras from multiple viewpoints; each action is performed by 40 subjects, and each subject has 25 joint positions;

[0058] The NTU-RGB+D 120 dataset is an extension of the NTU-RGB+D 60 dataset, with 60 additional actions added on the basis of NTU-RGB+D 60. It is also captured simultaneously from multiple viewpoints by three Kinect v2 cameras and has 25 joint positions.

[0059] Step 2: Divide body parts and time segments.

[0060] 2.1) Divide body parts:

[0061] The division of body parts is carried out according to three conditions set according to the human body topology and the fineness of action recognition tasks:

[0062] The first condition is that the joint points included in the divided body parts have only one connected subgraph in the human body topology structure;

[0063] The second condition is that the divided body parts can represent the finest-grained actions in the human action recognition task;

[0064] The third condition is that the number of joint points included in all body parts is the same. If they are not the same, they can be made the same by filling with 0 or copying some joint points;

[0065] Compared with the existing method that only takes joint points as the smallest unit for action recognition, in this example, since the joint points are aggregated into body parts, the process of calculating attention on multiple joint points can be transformed into the process of calculating attention on a small number of body parts, reducing the computational amount; at the same time, since the action features at the joint point level and the action features at the body part level are extracted, and introducing body parts into the action recognition algorithm process is more in line with human cognition of human movement, not only makes the present invention have stronger interpretability, but also reduces the influence of the noise of individual joint points in the data on action recognition;

[0066] 2.2) Divide time segments:

[0067] 2.2.1) For a human skeleton sequence with a total number of frames T0 greater than 120, a downsampling method is used to generate a human skeleton sequence with a total number of frames of 120, where the k-th frame of the generated human skeleton sequence is at the frame number t of the original human skeleton sequence k is:

[0068]

[0069] 2.2.2) For a human skeleton sequence with a total number of frames T0 less than 120, an upsampling method is used to generate a human skeleton sequence with a total number of frames of 120. Let the coordinates of the k-th frame of the generated human skeleton sequence be [x k , y k , z k from the t-th frame of the original human skeleton sequencek Frame coordinates and the t k +1 frame coordinates are obtained by trilinear interpolation, i.e.:

[0070]

[0071] 2.2.3) Uniformly divide the human skeleton sequences of all 120 frames into time segments of 6 frames;

[0072] In this example, for the NW-UCLA dataset, the joint points are divided into 5 body parts according to the refinement degree required by the action recognition task, namely the torso, left arm, right arm, left leg, and right leg. Each body part contains 5 joint points; the training batch size is set to 16;

[0073] For the NTU-RGB+D 60 and NTU-RGB+D 120 datasets, the joint points are divided into 7 body parts according to the refinement degree required by the action recognition task, namely the torso, left arm, left hand, right arm, right hand, left leg, and right leg. Each body part contains 4 joint points. Among them, the left wrist joint is in both the left hand and the left arm body parts, and the right wrist joint is in both the right hand and the right arm body parts, so that the number of joint points in all body parts is the same. The training batch size is set to 32 for both.

[0074] Step 3, construct a human skeleton action recognition model based on a hierarchical spatio-temporal attention network.

[0075] As Figure 2 shown, the specific implementation of this step is as follows:

[0076] 3.1) Establish a position encoding module:

[0077] According to the property of permutation invariance of the existing attention model, it is necessary to add the temporal information of each frame to the human skeleton sequence using position encoding to capture the temporal information of the input sequence; and the temporal features and spatial features of the human skeleton data are different. Therefore, the present invention designs temporal position encoding and spatial position encoding respectively for this difference:

[0078] 3.1.1) Temporal position encoding:

[0079] The time position encoding selects the cosine position encoding method. This encoding method can calculate a unique, deterministic, and bounded encoding value for each frame, and retains the generalization ability on the basis of adding the time order information of the frames to the human skeleton sequence. Since the time information contained in all joint points within the same frame should be the same, within the same frame, all joint points share the same time position encoding. That is, let d be the feature dimension of the human skeleton sequence, then the time position encoding of all joint points in the pos-th frame of the human skeleton sequence The value of the i-th dimension are all:

[0080]

[0081] 3.1.2) Spatial position encoding:

[0082] The spatial position encoding selects the learnable position encoding method. This encoding method can represent the inherent connection of joint points during human actions, and this inherent connection can be trained during the training process of the human skeleton action recognition model; the spatial position encoding is initialized as a zero vector, and within a body part, all frames in the human skeleton sequence share the same spatial position encoding;

[0083] 3.1.3) Concatenate the time position encoding and the spatial position encoding to form a position encoding module;

[0084] 3.2) Establish a hierarchical spatio-temporal attention module:

[0085] The hierarchical spatio-temporal attention module is based on the attention network. The attention mechanism in the attention network simulates human attention, and can learn the global dependence relationship between input data with relatively low computational complexity and good parallelism, so as to discover the hidden features in human actions. The specific implementation is as follows:

[0086] 3.2.1) Establish a hierarchical spatial attention module that calculates spatial attention with joint points as units:

[0087] Select an existing attention model, and set the three linear layers in this attention model to be respectively and Let the action features of the n-th joint point in all t frames or time segments be Connect the action features of all frames or time segments to obtain the joint point action feature

[0088] Let the extraction method of the attention model for the joint point action feature be as follows:

[0089] Take the joint point action feature as the input of the attention model, and according to the attention model, calculate the joint point query value respectively Joint point key value and joint point information value

[0090] Let the human body joint point topology matrix be E n , calculate the similarity between the joint point query value and the joint point key value and sum it with the human body joint point topology matrix, and then use this as the weight to weighted sum the joint point information value to calculate the deep joint point action feature

[0091]

[0092] where T represents matrix transpose, C n represents the number of channels of K n ;

[0093] Take the attention model that extracts the joint point action feature as a hierarchical spatial attention module that calculates spatial attention in units of joint points;

[0094] 3.2.2) Establish a hierarchical spatial attention module that calculates spatial attention in units of body parts:

[0095] Select an existing attention model, and set the three linear layers in this attention model to be and Let the joint point action features of p joint points included in the m-th body part at the pos-th frame or time segment be Connect these joint point action features to obtain the body part action feature

[0096] Let the extraction method of the attention model for the body part action feature be as follows:

[0097] Take the body part action feature as the input of the attention model, and according to the attention model, calculate the body part query value body part key value and body part information value

[0098] Let the body part topology matrix be E m , calculate the similarity between the body part query value and the body part key value and sum it with the body part topology matrix, and then use this as the weight to weighted sum the body part information value to obtain the deep body part action feature

[0099]

[0100] where T represents matrix transpose, Cm Denote K m as the number of channels;

[0101] Use the attention model for updating the action features of body parts as a hierarchical spatial attention module that calculates spatial attention in units of joint points;

[0102] 3.2.3) Establish a hierarchical temporal attention module that calculates temporal attention in units of frames:

[0103] Select an existing attention model, and set the three linear layers in this attention model to be respectively and Let the action features of all k joint points or body parts in the pos-th frame be respectively Connect the action features of these joint points or body parts to obtain the frame action feature

[0104] Let the extraction method of the attention model for the frame action feature be as follows:

[0105] Update the frame action feature as follows:

[0106] Take the frame action feature as the input of the attention model, and calculate the frame query value and respectively according to the three linear layers frame key value and frame information value

[0107] Find the similarity between the frame query value and the frame key value Then use this as the weight to weighted sum the frame information values to obtain the deep-level frame action feature

[0108]

[0109] where T represents matrix transpose, and C pos denotes K pos as the number of channels;

[0110] Use the attention model for updating the frame action feature as a hierarchical temporal attention module that calculates temporal attention in units of frames;

[0111] 3.2.4) Establish a hierarchical temporal attention module that calculates temporal attention in units of time segments:

[0112] Select an existing attention model, and set the three linear layers in this attention model to be respectively and Let the action features of the joints or body parts in the r-th time segment, which consist of 6 frames, be respectively Connecting the action features of these joints or body parts gives the action feature of the time segment

[0113] Let the extraction method of the attention model for the action feature of the time segment be as follows:

[0114] Taking the action feature of the time segment as the input of the attention model, according to the three linear layers in the attention model and calculate the query value of the time segment the key value of the time segment and the information value of the time segment

[0115] Find the similarity between the query value of the time segment and the key value of the time segment Then, using this as the weight, perform a weighted sum of the information values of the time segment to obtain the deep action feature of the time segment

[0116]

[0117] where T represents matrix transpose, and C r represents the number of channels of K r ;

[0118] Regarding the attention model used to update the action feature of the time segment as a hierarchical time attention module that calculates time attention in units of time segments;

[0119] 3.2.5) Cascade the hierarchical spatial attention module that calculates spatial attention in units of joints and the hierarchical time attention module that calculates time attention in units of frames, add layer normalization and residual connections, and then stack them as the shallow layer of the hierarchical spatio-temporal attention module;

[0120] 3.2.6) Cascade the hierarchical spatial attention module that calculates spatial attention in units of body parts and the hierarchical time attention module that calculates time attention in units of time segments, add layer normalization and residual connections, and then stack them as the deep layer of the hierarchical spatio-temporal attention module;

[0121] 3.2.7) Cascade the shallow layer of the hierarchical spatio-temporal attention module and the deep layer of the hierarchical spatio-temporal attention module to form the hierarchical spatio-temporal attention module, as Figure 3 shown;

[0122] 3.3) Select a fully connected layer as the classification module:

[0123] In this example, the fully connected layer is implemented using a 1×1 convolutional layer;

[0124] (3.4) Cascade the position encoding module, hierarchical spatio-temporal attention module, and classification module to form a human skeleton action recognition model based on the hierarchical spatio-temporal attention network;

[0125] In this example, for the NW-UCLA dataset, the position encoding dimension of the human skeleton action recognition model is set to 32, the attention dimension is set to 128, the number of shallow layers in both the shallow layer of the hierarchical spatio-temporal attention module and the hierarchical spatio-temporal attention module is set to 1, the classification module dimension is set to 128, and dropout is set to 0.1;

[0126] For the NTU-RGB+D 60 and NTU-RGB+D 120 datasets, the position encoding dimension of the action recognition model is set to 64, the attention dimension is set to 256, the number of shallow layers in both the shallow layer of the hierarchical spatio-temporal attention module and the hierarchical spatio-temporal attention module is set to 4, the classification module dimension is set to 256, and dropout is set to 0.1;

[0127] Initialize all parameters in the model using a truncated normal distribution.

[0128] Step 4, train the human skeleton action recognition model.

[0129] Input the training sample set into the human skeleton action recognition model and use the Adam optimization algorithm to train it as follows:

[0130] (4a) Let the learning rate be \(l\) r , the cumulative gradient \(m\) τ at the \(\tau\)-th round of training, its decay rate \(\beta_1\), the update direction \(v\) τ at the \(\tau\)-th round of training, its decay rate \(\beta_2\), the learnable parameters \(\theta\) τ of the model, a small constant \(\epsilon\), and the maximum number of training rounds is 90;

[0131] In this example, the learning rate \(l\) r is set to 0.001, the decay rate \(\beta_1\) of the cumulative gradient \(m\) τ is set to 0.2, the decay rate \(\beta_2\) of the update direction \(v\) τ is set to 0.3, and the small constant \(\epsilon\) is set to \(1\times10\) -5 ;

[0132] (4b) Use the position encoding module to add the sequential information of the frames to the human skeleton sequence, and the resulting human skeleton sequence with position encoding is \(X\) in :

[0133] \(X\) in = \(X + PE\) time + \(PE\) space ;

[0134] Among them, X is the human skeleton sequence in the training sample set, and PE time is the temporal position encoding of all frames and joints, and PE space is the spatial position encoding of all frames and joints;

[0135] (4c) Use the hierarchical spatio-temporal attention module to extract the action features of X in at the τ-1 round;

[0136] (4d) Process the action features with the classification module to obtain the action recognition result of the τ-1 round model's learnable parameter θ τ-1 ;

[0137] (4e) Calculate the gradient of the current round where denotes gradient derivation, and J(θ τ-1 ) represents the difference between the action recognition result of the τ-1 round model's learnable parameter θ τ-1 and the true category;

[0138] (4f) Update the cumulative gradient m τ =β1m τ-1 +(1 - β1)g τ , which is the weighted sum of the τ-1 round cumulative gradient m τ-1 and the current round gradient g τ ;

[0139] (4g) Update v τ =β2v τ-1 +(1 - β2)(g τ ) 2 , which is the weighted sum of the τ-1 round update direction v τ-1 and the square of the current round cumulative gradient (g τ ) 2 ;

[0140] (4h) Calculate the bias correction of m τ and the bias correction of v respectively τ to reduce the impact of the bias towards 0 on model training and make the model easier to train;

[0141] (4i) Update the learnable parameters(4i) Update the learnable parameters

[0142] (4j) Repeat (4b) to (4i) until τ reaches the maximum number of training rounds, stop iteration, complete the update of the learnable parameter θ τ and obtain the trained action recognition model.

[0143] Step 5, test the human skeleton action recognition model.

[0144] The test sample set is input into the trained action recognition model. Through the position encoding module, the temporal information of each frame is added to the human skeleton sequence. The hierarchical spatio-temporal attention module extracts action features, and the classification module obtains the action recognition results of the test sample set, completing the recognition of human skeleton actions.

[0145] The effects of the present invention can be further illustrated by the following simulation results:

[0146] I. Simulation conditions

[0147] The simulation test platform of the present invention is a PC with a 3.6 GHz Intel Core i7-9700K CPU and a Nvidia RTX2080Ti graphics card. The simulation platform is the Ubuntu18.04 operating system, using the Pytorch deep learning framework and implemented in the Python language.

[0148] The simulation test data sets of the present invention are the NW-UCLA data set, the NTU-RGB+D 60 data set, and the NTU-RGB+D 120 data set. Both the NTU-RGB+D 60 data set and the NTU-RGB+D 120 data set adopt two training sample set and test sample set division methods, namely the Xsub division method and the Xview division method. The training sample set of the Xsub division method comes from 40 subjects, and the test sample set comes from the other 40 subjects; the training sample set of the Xview division method is the video captured by the first two cameras, and the test sample set is the video captured by the third camera.

[0149] II. Simulation content

[0150] Under the above simulation conditions, the action recognition method of the present invention is used to perform action recognition on the three data sets respectively, and the action recognition accuracy p of the test sample set is calculated:

[0151]

[0152] where; num is the number of human skeleton sequences included in the test sample set, y a is the action recognition result of the a-th test sample, y' a is the true category of the a-th test sample, g(y a , y' a ) is a function with a value of 1 when y a = y' a and a value of 0 when y a ≠y' a .

[0153] The calculation results are shown in Table 1.

[0154] Table 1

[0155]

[0156] As can be seen from Table 1, the human skeleton action recognition model based on the hierarchical spatio-temporal attention network proposed by the present invention utilizes the advantages of the attention model for global information modeling to explore the spatio-temporal features of the human skeleton sequence during the human action process, thereby enabling high-precision human skeleton action recognition.

[0157] The above are only examples of the present invention and do not impose any formal limitations on the present invention. Any simple modifications and equivalent changes made to the above embodiments based on the technical essence of the present invention all fall within the protection scope of the present invention.

Claims

1. A human skeleton action recognition method based on a hierarchical spatio-temporal attention network, characterized in that, It includes the following steps: (1) Obtain the human skeleton sequence, form human skeleton data with the human skeleton sequence, divide the human skeleton data into a training sample set and a test sample set according to a ratio of 2:1, and perform view standardization and scale normalization on all human skeleton sequences in the training sample set and the test sample set to obtain the standardized human skeleton sequence; (2) Divide the human joint points into body parts, convert all human skeleton sequences with more than 120 frames into 120 frames using the downsampling method, convert all human skeleton sequences with less than 120 frames into 120 frames using the upsampling method, and then evenly divide all 120-frame human skeleton sequences into time segments of 6 frames; (3) Construct a human skeleton action recognition model based on a hierarchical spatio-temporal attention network: (3a) Construct a position encoding module composed of the cascading of time position encoding and spatial position encoding; (3b) Construct a hierarchical spatial attention module and a hierarchical time attention module, cascade the hierarchical spatial attention module and the hierarchical time attention module, add layer normalization and residual connections and stack them to construct a hierarchical spatio-temporal attention module; (3c) Select a fully connected layer as the classification module; (3d) Cascade the position encoding module, the hierarchical spatio-temporal attention module and the classification module to form a human skeleton action recognition model based on a hierarchical spatio-temporal attention network; (4) Input the training sample set into the human skeleton action recognition model, and use the Adam optimization algorithm to perform iterative training for 90 rounds to obtain a trained human skeleton action recognition model; (5) Input the test sample set into the trained human skeleton action recognition model to obtain the action recognition results of the test sample set.

2. The method according to claim 1, characterized in that In step (1), the view standardization and scale normalization of all human skeleton sequences in the training sample set and the test sample set are realized as follows: (1a) Calculate the average value o of the coordinates of the joint point "hip center" in all frames of the human skeleton sequence. Let T0 be the total number of frames of the human skeleton sequence. is the coordinate of the nth joint point in the pos-th frame of the human skeleton sequence. ψ is the set of human torso joint points. Then the torso matrix M of the human skeleton sequence can be expressed as: (1b) Calculate and normalize the first principal component of the torso matrix M using the principal component analysis method to obtain the Z-axis direction vector γ in the new coordinate system, calculate and normalize the second principal component of M to obtain the Y-axis direction vector β in the new coordinate system, and calculate the cross product α of γ and β; (1c) Use the dimensionality increase [x, y, z, 1] of the original joint point coordinates [x, y, z] of the human skeleton sequence and α, β, γ, o to calculate the dimensionality increase [x', y', z', 1] of the joint point coordinates [x', y', z'] of the standardized human skeleton sequence, which is expressed as the product of the following four-dimensional matrix and [x, y, z, 1]: where (α-o) T , (β-o) T and (γ-o) T are the normalized vectors of the X-axis view, the Y-axis view, and the Z-axis view respectively, o is the coordinate of the normalized coordinate origin in the original coordinate system, and T represents matrix transpose.

3. The method according to claim 1, characterized in that, In step (2), the human joint points are divided into body parts and the human skeleton sequence is divided into time segments, which are realized as follows: The division of the body parts is to aggregate the human joint points into body parts according to the human topology and the recognition fineness of the action recognition task; The division of the time segments is to use the downsampling method for human skeleton sequences with the number of frames T greater than 120, and use the downsampling method for human skeleton sequences with the number of frames T greater than 120 to generate human skeleton sequences with 120 frames, and then evenly divide all 120-frame human skeleton sequences into time segments of 6 frames.

4. The method according to claim 1, characterized in that, The time position encoding and spatial position encoding in step (3a) are realized as follows: The time position encoding uses the cosine position encoding method, and within one frame, all joint points share the same time position encoding. Let d be the dimension of the human skeleton sequence feature, then the time position encoding of all joint points in the pos-th frame of the human skeleton sequence The value of the i-th dimension is all: The spatial position encoding selects a learnable position encoding method, is initialized as a zero vector, is trained during the training process of the human skeleton action recognition model, and within a body part, all frames in the human skeleton sequence share the same spatial position encoding.

5. The method according to claim 1, wherein The hierarchical spatial attention module and the hierarchical temporal attention module in step (3b) are implemented as follows: The hierarchical spatial attention module selects an existing attention model, calculates spatial attention in units of joint points in the shallow layer of the hierarchical spatio-temporal attention module and calculates deep joint point action features based on this, and calculates spatial attention in units of body parts in the deep layer of the hierarchical spatio-temporal attention module and calculates deep body part action features based on this; The hierarchical temporal attention module selects an existing attention model, calculates temporal attention in units of frames in the shallow layer of the hierarchical spatio-temporal attention module and calculates deep frame action features based on this, and calculates temporal attention in units of time segments in the deep layer of the hierarchical spatio-temporal attention module and calculates deep time segment action features based on this.

6. The method according to claim 1, characterized in that, Step (4) selects the Adam optimization algorithm to perform iterative training on the action recognition model, and is implemented as follows: (4a) Set the learning rate l r , the cumulative gradient m in the τ-th round of training t and its decay rate β1, the update direction v in the τ-th round of training τ and its decay rate β2, the model's learnable parameter θ τ , a small constant ε, and the maximum number of training rounds is 90. Then the model training process for each round is as follows: (4b) Calculate the gradient g of the current round τ ; (4c) Update the cumulative gradient m τ = β1m τ-1 + (1 - β1)g τ ; (4d) Update v τ = β2v τ-1 + (1 - β2)(g τ ) 2 ; (4e) Calculate m τ for deviation correction and v τ for deviation correction (4f) Update learnable parameters (4g) Repeat (4b) to (4f) until τ reaches the maximum number of training epochs, stop the iteration, and obtain the trained action recognition model.

7. The method according to claim 5, wherein Calculate spatial attention in units of joint points and calculate deep joint point action features based on this, and is implemented as follows: Let the action features of the nth joint point in all t frames or time segments be respectively Connect the action features of all frames or time segments to obtain the joint point action feature According to three linear layers in the attention model and calculate the joint query value joint key value and joint information value Let the human body joint point topological matrix be E n , and find the similarity between the joint point query value and the joint point key value And sum it with the human body joint point topological matrix, and then use this as the weight to weighted sum the joint point information values to obtain the updated joint point motion feature where T represents matrix transpose, and C n represents the number of channels of K n .

8. The method according to claim 5, characterized in that Calculate spatial attention in units of body parts and calculate deep body part action features based on this, and is implemented as follows: Let the joint action features of the p joint points included in the m-th body part at the pos-th frame or time segment be respectively Connect these joint action features to obtain the body part action feature According to the three linear layers in the attention model and calculate the body part query value body part key value and body part information value Let the body part topology matrix be E m , and calculate the similarity between the body part query value and the body part key value Then sum it with the body part topology matrix, and use this as the weight to sum the body part information values weighted to obtain the updated body part action feature where T represents matrix transpose, and C m represents the number of channels of K m .

9. The method according to claim 5, characterized in that Calculate temporal attention in units of frames and calculate deep frame action features based on this, and is implemented as follows: Let the action features of all k joint points or body parts in the pos-th frame be respectively Connect the action features of these joint points or body parts to obtain the frame action feature According to the three linear layers in the attention model and calculate the frame query value the frame key value and the frame information value Find the similarity between the frame query value and the frame key value Then, use this as a weight to perform a weighted sum of the frame information values to obtain the updated frame action feature where T represents matrix transpose, and C pos represents the number of channels of K pos .

10. The method according to claim 5, characterized in that, Calculate spatial attention in units of time segments and calculate deep time segment action features based on this, and is implemented as follows: Let the joint or body part action features of the 6 frames included in the r-th time segment be respectively Connect the action features of these joints or body parts to obtain the time segment action feature According to three linear layers in the attention model and calculate the time segment query value the time segment key value and the time segment information value Find the similarity between the time segment query value and the time segment key value Then, use this as a weight to perform a weighted sum on the time segment information value to obtain the updated time segment action feature where T represents matrix transpose, and C r represents the number of channels of K r .

Citation Information

Patent Citations

  • Motion recognition method and system based on fusion graph convolutional network and Transform network

    CN115100574A

  • Human body behavior recognition method based on human body skeleton data

    CN115841647A