Human skeleton modeling method based on dynamic graph generation and adaptive graph convolution
Through dynamic graph generation and adaptive graph convolution, the human skeleton modeling method is solved in the existing technology with limited bone structure fixation and time modeling capabilities, achieving higher recognition accuracy and stability, and is suitable for complex human movement recognition.
Patent Information
- Application Number
- CN202510304744.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, human skeleton modeling relies on predefined adjacency matrix, resulting in the fixed bone structure and the inability to adapt to the connection between different individuals or dynamically adjust key points. The time modeling ability is limited and the calculation complexity is high.
The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution is adopted. The dynamic graph generator dynamically generates an optimized adjacency matrix adapted to different individuals and action modes, and combines the spatial-temporal graph convolution to extract spatial-temporal joint features, and finally outputs the action category probability through classification heads.
It effectively improves the accuracy and stability of the model, is suitable for complex human body movement recognition tasks, shows stronger generalization ability and application value, can adaptively learn bone topology, dynamically adjust the connection relationship between key points, and adapt to different individuals and action patterns.
Smart Images

Figure CN119963744A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and relates to a human skeleton modeling method based on dynamic graph generation and adaptive graph convolution. Background Art
[0002] In recent years, human action recognition has become an active research field, and it plays an important role in video understanding. Usually, human action recognition has multiple modalities, such as appearance, depth, optical flow, and body skeleton. Among these modalities, dynamic human skeletons can usually complement other modalities and convey important information. In the form of 2D or 3D coordinates, the dynamic skeleton modality can be naturally represented by a time series of human joint positions. Then, human action recognition can be achieved by analyzing its action pattern. In recent years, the graph convolutional network (GCN) that generalizes the convolutional neural network (CNN) to arbitrary structural graphs has received increasing attention and has been successfully applied to image classification, document classification, semi-supervised learning and other fields. There are also dynamic graph models of GCN on large-scale datasets that are applied to the modeling of human skeleton sequences.
[0003] At present, in the skeleton modeling of human action recognition, commonly used techniques include graph representation methods, classic graph convolution (GCN) modeling methods, spatiotemporal graph convolution (ST-GCN) modeling methods, and graph convolution network models with adaptive graph size. At present, the most effective method for human skeleton modeling is the human skeleton modeling method based on spatiotemporal graph convolution mentioned in Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition published by Yan, S. et al. in 2018. This method constructs a spatiotemporal graph based on a sequence of body joints in a given 2D or 3D coordinate system, in which human joints correspond to nodes of the graph, and the connectivity of the human body structure and the connectivity in time correspond to two types of edges in the graph. Therefore, the input of ST-GCN is the joint coordinate vector of the graph nodes. It is equivalent to an image-based CNN simulation, in which the input is formed by the pixel intensity vector on the 2D image grid, which makes up for the defect that the previous method cannot capture features on the time frame.
[0004] However, this method relies on a predefined adjacency matrix, whose skeleton structure is fixed and cannot adapt to different individuals or dynamically adjust the connection between key points. Existing adaptive adjacency matrix designs mostly rely on fixed graph structures or simple attention mechanisms, which are difficult to capture complex spatiotemporal dependencies. Moreover, the temporal modeling capability is limited and the computational complexity is high. Summary of the invention
[0005] The purpose of the present invention is to provide a human skeleton modeling method based on dynamic graph generation and adaptive graph convolution, which solves the problems existing in the prior art that the skeleton structure is fixed due to reliance on a predefined adjacency matrix, the inability to adapt to different individuals or dynamically adjust the connection between key points, and the weak ability to capture long-term dependencies.
[0006] The technical solution adopted by the present invention is a human skeleton modeling method based on dynamic graph generation and adaptive graph convolution, and the steps are as follows: Step 1, obtaining human behavior skeleton point map data as the original data set; Step 2, preprocessing and dividing the original data set; Step 3: Design a dynamic graph generator to dynamically generate an optimized adjacency matrix that adapts to different individuals and action modes by using the preprocessed data set through node feature encoding, graph structure generation, and adjacency matrix optimization. Step 4: Input the preprocessed data set and the optimized adjacency matrix into the spatiotemporal graph convolution to extract spatiotemporal joint features; Step 5: Input the spatiotemporal joint features into a classification head, output the model’s predicted action category probability, compare it with the true category of the input data, calculate the cross entropy loss and back propagate, obtain the trained human action recognition model, and output the final recognition result.
[0007] The present invention is characterized in that: Step 1 is specifically as follows: obtaining NTU RGB+D 60 / 120 human behavior skeleton point map data in the NTU dataset as the original dataset.
[0008] Step 2 is as follows: Step 2.1, normalize the human behavior skeleton point map data, specifically: calculate the mean and standard deviation of all samples in each coordinate channel, and use them to standardize the data so that the coordinate values of all skeleton points are distributed in the same range; the normalized data maintains the original time series structure, each frame corresponds to a set of skeleton key points, and each key point contains information from multiple channels; The data format after normalization is a three-dimensional tensor dataset M´: M´=reshape(M,(N,t,V,C))(1) Among them, M is the original data set, M , N is the batch size, t is the time step, V is the number of key points, and C is the number of coordinate channels, including normalized coordinates, bone vectors, and motion speed; Step 2.2: Add random time jitter, that is, randomly discard some frames or insert extra frames; Step 2.3: Randomly hide some key points; Step 2.4: Generate pseudo targets, including frame sequence prediction and node prediction; Step 2.5: Divide the preprocessed dataset into training set X train and the test set X test , the format of the divided data set is as follows: (2).
[0009] Step 3 is specifically as follows: the dynamic graph generator includes a node feature encoder, a graph structure generator, and a weight matrix optimization; The node feature encoder encodes the skeleton node features in the input data set, and performs nonlinear transformation on these features through a multi-layer perceptron to generate high-dimensional features H. The formula is as follows: H = σ(W2⋅σ(W1X+b1)+b 2 ) (3) Where: X is the input dataset, D represents the dimension of the high-dimensional space, W1 and W2 are the learnable weight matrices of the first and second layers respectively; b1 and b2 are bias terms; σ represents the nonlinear activation function.
[0010] The features of bone nodes include spatial coordinates and motion speed information; The graph structure generator calculates the correlation weights between nodes through a multi-head attention mechanism and generates a dynamic adjacency matrix. The formula is as follows: Calculate the weighted attention matrix between nodes based on the high-dimensional feature H : (4) Where W is a learnable weight matrix used to generate queries and keys, D1 represents the feature dimension, and T represents the matrix transpose. Dynamic adjacency matrix The formation is as follows: (5) in represents an initialized learnable parameter matrix; Weight Matrix Optimization for Dynamic Adjacency Matrix Perform weighted screening to remove redundant connections and capture higher-order graph structure relationships, specifically: (6) Where L represents the number of layers, γ represents the weight parameter, which is used to balance the adjacency matrix of the current layer and the previous layer. The adjacency matrix after screening and optimization for the Lth layer can adapt to different individuals and action modes.
[0011] Step 4 is as follows: Step 4.1: Extract spatial features of skeleton data through CTR-GCN: (7) Where X is the input bone dataset, is the adjacency matrix after optimization in step 3, is the extracted spatial feature; Specifically: (8) in That is, add the self-connected adjacency matrix, I is the identity matrix, yes The degree matrix of , W is the learnable weight matrix, (·) is the activation function, and F is the extracted spatial feature; Step 4.2: Capture dependencies at different time scales through a multi-scale spatiotemporal attention mechanism; Specifically, the time scale is divided into three scales: short, medium, and long. For each time scale, defined as s, the attention weight is calculated as follows: (9) in , , , respectively represent query, key and value, where W represents the learnable weight matrix, D1 represents the feature dimension, and then weighted fusion is performed, the formula is as follows: (10) in, Refers to the weight of each time scale, which is a learnable parameter. Refers to the extracted temporal features; Step 4.3, the extracted spatial features and temporal features are fused and normalized to generate the final spatiotemporal joint features: (11) in, Refers to the final spatiotemporal joint features, and LayerNorm refers to layer normalization.
[0012] Step 5 is as follows: The spatiotemporal joint features obtained in step 4 are The input is input into a classification head containing a global average pooling layer and a fully connected layer to output the model's predicted action category probability. The model's predicted action category probability is compared with the true category of the input data, the cross entropy loss is calculated, and the network parameters are iteratively trained and optimized through the Adam optimizer and back propagation algorithm. After the training is completed, the trained human behavior recognition model is obtained and the model is saved and output.
[0013] The cross entropy loss calculation formula is as follows: (12) Where L is the cross entropy loss, c represents the category index, is the one-hot encoding of the true category. If category c is the correct category, then y c =1, otherwise y c =0; It is the action category probability predicted by the model, that is, the Softmax output.
[0014] Network parameters include: batch size, learning rate, weight decay, optimizer, and training cycle.
[0015] The beneficial effects of the present invention are: The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution of the present invention combines a dynamic graph generation network with adaptive graph convolution (GCN), long-term feature extraction and spatiotemporal joint learning, breaking through the limitation of the fixed adjacency matrix of the traditional method, effectively improving the accuracy and stability of the model, and being suitable for complex human motion recognition tasks, showing stronger generalization ability and application value; it can adaptively learn the skeleton topology structure, dynamically generate an adjacency matrix adapted to different individuals and motion modes, dynamically adjust the connection relationship between key points, adapt to different individuals and motion modes, and improve the generalization ability of the model; Compared with methods such as ST-GCN, the human skeleton modeling method based on dynamic graph generation and adaptive graph convolution of the present invention further enhances the adaptive adjacency strategy, makes skeleton modeling more flexible, and effectively improves the recognition accuracy. In addition, traditional GCN is difficult to capture long-term dependencies. This method extracts global timing information through a multi-scale spatiotemporal attention mechanism, overcomes the defect that TCN only relies on local timing modeling, can capture long-term dependencies and local details, and makes the model more advantageous in processing action sequences with long time spans. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flow chart of a human skeleton modeling method based on dynamic graph generation and adaptive graph convolution of the present invention; Figure 2 It is a skeleton point structure diagram of the human skeleton modeling method of the present invention; Figure 3 It is a curve showing the change of the loss function value of the human skeleton modeling method of the present invention with the number of training iterations. DETAILED DESCRIPTION
[0017] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] Example 1 The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution of the present invention has the following process: Figure 1 As shown, the specific steps include: Step 1, obtaining human behavior skeleton point map data as the original data set; Step 2: After preprocessing the original data set, divide it into training set and test set; Step 3: Design a dynamic graph generator to dynamically generate an optimized adjacency matrix that adapts to different individuals and action modes by using the preprocessed data set through node feature encoding, graph structure generation, and weight matrix optimization. Step 4: Input the preprocessed data set and the optimized adjacency matrix into the spatiotemporal graph convolution to extract spatiotemporal joint features; Step 5: Input the spatiotemporal joint features into the classification head, output the model’s predicted action category probability, compare it with the true category of the input data, calculate the cross entropy loss and back propagate, obtain the trained human action recognition model, and output the final recognition result.
[0019] Example 2 Based on Example 1, step 1 is specifically as follows: The NTU RGB+D 60 / 120 dataset is used as the experimental data source. This dataset contains 60 categories of human actions, a total of 56,880 samples, covering multimodal information such as RGB video, depth map, 3D skeleton data and infrared video (120 categories is double). The process of obtaining the dataset is as follows: Step 1.1: Visit the official website of the NTU dataset and fill out the dataset use application form, stating that the research purpose is patent development and behavior recognition algorithm research.
[0020] Step 1.2: After submitting the application, you will be granted the download permission for the dataset after it is reviewed and approved.
[0021] Step 1.3, download the NTU RGB+D 60 / 120 dataset, decompress it and store it on the local server.
[0022] Step 1.4 verifies the data integrity, ensures that all samples and annotation files are correct, and obtains the original data set M of human behavior skeleton point graph data. Figure 2 As shown, during use, the use agreement of the data provider is strictly followed to ensure that the data is only used for non-commercial research purposes.
[0023] The specific process of step 2 is as follows Figure 1 As shown: Step 2.1: Standardize the human behavior skeleton point map data, that is, normalize the key point coordinates and unify the skeleton size; First, the original human skeleton data is normalized to ensure that the feature scales of different samples are consistent. The specific method is to calculate the mean and standard deviation of all samples in each coordinate channel, and use them to standardize the data so that the coordinate values of all skeleton points are distributed in the same range. The normalized data still maintains the original time series structure, and each frame corresponds to a set of skeleton key points, and each key point contains information from multiple channels.
[0024] The data format after normalization is a three-dimensional tensor dataset M´: M´=reshape(M,(N,t,V,C))(1) Among them, M is the original data set, M , N is the batch size, t is the time step, V is the number of key points, and C is the number of coordinate channels, including normalized coordinates, bone vectors, and motion speed.
[0025] Step 2.2: Add random time jitter (randomly discard some frames or insert extra frames) to enhance temporal robustness; Step 2.3: Randomly hide some key points to improve the model's adaptability to missing data; Step 2.4: Generate pseudo targets, such as frame order prediction and node prediction, for self-supervised learning to obtain the preprocessed dataset X.
[0026] Step 2.5: Randomly select 80% of the samples from the result dataset of step 2.4 as the training set X train , the remaining 20% is used as the test set X test ; Ensure that the samples in the training set and the test set are evenly distributed to avoid class imbalance problems; the training set is used for model training, and the test set is used for final performance evaluation.
[0027] The format of the divided data set is as follows: (2) The training set accounts for 80% of the total data, and the test set accounts for 20% of the total data. The tensor of the training set will be used as the input of the dynamic graph generator in step 3.
[0028] Example 3 Based on Example 2, step 3 is specifically as follows: design a dynamic graph generator (DGG) to adapt to the changes of different individuals and action modes by dynamically generating an adjacency matrix. Traditional graph convolutional networks (GCNs) usually rely on predefined adjacency matrices. This fixed structure cannot adapt to the differences in skeletal topology of different individuals, nor can it dynamically adjust the connection relationship between key points. Therefore, an adaptive adjacency matrix generation method based on a neural network is designed.
[0029] Step 3.1, the dynamic graph generator is a lightweight neural network module that is used to dynamically generate an adjacency matrix based on the input human skeleton sequence. It consists of three parts: node feature encoder, graph structure generator, and weight matrix optimization. The data is passed through the designed dynamic graph generator to generate a new adjacency matrix.
[0030] Step 3.1.1, the dynamic graph generator first encodes the features of the skeleton nodes in the input data set. The features of each skeleton node include its spatial coordinates, movement speed and other information. The training set in the preprocessed data set X is input, and these features are nonlinearly transformed through the multi-layer perceptron (MLP) to generate high-dimensional features H. The formula is as follows: H = σ(W2⋅σ(W1X+b1)+b 2 ) (3) Where: X is the input dataset, D represents the dimension of the high-dimensional space, W1 and W2 are the learnable weight matrices of the first and second layers respectively; b1 and b2 are bias terms; σ represents a nonlinear activation function, such as ReLU.
[0031] The purpose of this step is to map the original skeleton node features into a higher-dimensional feature space so that the subsequent graph structure generator can better capture the complex relationships between nodes.
[0032] Step 3.1.2, the graph structure generator is the core part of the dynamic graph generator. It dynamically generates the adjacency matrix through the learnable parameter matrix and attention mechanism, and calculates the attention weights between nodes based on H obtained in step 3.1.1 , the specific formula is as follows: (4) Among them, W is a learnable weight matrix used to generate queries and keys, D1 represents the feature dimension, T represents the matrix transpose, and finally the weight attention matrix between nodes is obtained. In the dynamic graph generator, the graph structure generator calculates the correlation weights between nodes through the multi-head attention mechanism, generates a dynamic adjacency matrix, and initializes a learnable parameter matrix again. , used to adjust the connection strength between nodes. Dynamic adjacency matrix The formation is as follows: (5) in This design enables the adjacency matrix to be adaptively adjusted according to the input data and to adapt to the changes of different individuals and action patterns.
[0033] Step 3.1.3, the dynamic graph generator combines the learnable parameter matrix and the weight matrix generated by the attention mechanism. The adjacency matrix generated in the previous step needs to be further optimized. Weight Matrix Optimization for Dynamically Generated Adjacency Matrix Screening is performed to remove redundant connections and capture higher-order graph structure relationships. The specific method of weighted screening at the Lth layer is as follows: (6) Where L represents the number of layers, γ represents the weight parameter to balance the adjacency matrix of the current layer and the previous layer. The adjacency matrix optimized for the Lth layer can adapt to different individuals and action modes. This matrix is applied to the network model in step 4.
[0034] Example 4 Based on Example 3, the specific process of step 4 is as follows: Step 4.1: After obtaining the adjacency matrix in the previous stage, input the preprocessed data set and the adaptive adjacency matrix generated in step 3 here. , spatial features of skeleton data are extracted through the channel topology refinement graph convolutional network CTR-GCN.
[0035] Step 4.1.1, first is spatial feature extraction, the formula is as follows: (7) Where X is the input bone dataset, is the adjacency matrix after optimization in step 3, is the extracted spatial feature; Step 4.1.2, the usage formula of CTR-GCN is as follows: (8) in That is, add the self-connected adjacency matrix, I is the identity matrix, yes The degree matrix of , W is the learnable weight matrix, (·) is the activation function, and F is the extracted spatial feature; Step 4.2, temporal modeling captures dependencies at different time scales through the multi-scale spatiotemporal attention mechanism (MST-Attention). The time scale is divided into three scales: short, medium, and long. For each time scale, defined as s, the attention weight is calculated as follows: (9) in , , , respectively represent query, key and value, where W represents the learnable weight matrix and D1 represents the feature dimension. Next, we need to perform weighted fusion, the formula is as follows: (10) in Refers to the weight of each time scale, which is a learnable parameter. Refers to the extracted temporal features.
[0036] Step 4.3, spatiotemporal joint features The spatial features and temporal features are fused and normalized to generate the final spatiotemporal joint features. The specific formula is as follows: (11) in It refers to the final spatiotemporal joint features, and LayerNorm refers to layer normalization, which is used to stabilize the training process and obtain the final result.
[0037] Example 5 Based on Example 4, step 5 is specifically: the final result obtained in step 4 is Input into a classification head containing a global average pooling layer and a fully connected layer to output the model's predicted action category probability; The classification head refers to the part of the neural network used for final classification, including: Global Average Pooling (GAP): reduces the dimension of spatiotemporal features and extracts global information; Fully Connected Layer (FC): used to map to specific action categories; Softmax layer: used to calculate the final category probability.
[0038] The model predicts the action category probability and the actual category of the input data to calculate the cross entropy loss. The formula is as follows: (12) Where L is the cross entropy loss, c represents the category index, is the one-hot encoding of the true category. If category c is the correct category, then y c =1, otherwise y c =0; It is the action category probability predicted by the model, that is, the Softmax output.
[0039] The network parameters are optimized through iterative training using the Adam optimizer and back-propagation algorithm. The network parameters are set as follows: batch_size is 16, the initial learning rate is 0.001, the weight decay is 5e-4, the optimizer selected is Adam, the epochs, or training cycle, is 80 rounds, and the loss change curve is as follows: Figure 3 As shown, after the training is completed, a trained human behavior recognition model is obtained and the final recognition result is output.
[0040] Input the test set into the trained human behavior recognition model to test the model performance.
[0041] Example 6 In order to verify the effectiveness of the human skeleton modeling method based on dynamic graph generation and adaptive graph convolution in Example 5 of the present invention in human skeleton modeling and action recognition, the following comparative experiments were conducted: The NTU-RGB+D 60-class dataset is used, which contains 60 types of human action skeleton graphs. The batch_size is set to 16, the initial learning rate is 0.001, the weight decay is 5e-4, and the Top-1 accuracy and Top-5 accuracy are used as the main evaluation indicators. In order to fully reflect the effectiveness of the method of the present invention, the method is compared with the ST-GCN and CTR-GCN methods trained and tested on the same dataset.
[0042] Experimental environment: hardware NVIDIA 3060GPU, software PyTorch 2.0, CUDA 12.0; A cross-validation strategy was used to perform a five-fold cross-validation. During the test, the Top-1 accuracy and Top-5 accuracy were recorded, and the computational complexity (GFLOPs) and inference speed (FPS) were calculated. The comparison results are shown in Table 1 below: Table 1 Accuracy on the NTU RGB-D60 dataset
[0043] As can be seen from Table 1, the method of Example 5 of the present invention has a Top-1 accuracy of 87.6% and a Top-5 accuracy of 96.5% on the NTU RGB+D 60 dataset, which are both higher than ST-GCN (83.4%, 94.1%) and CTR-GCN (85.8%, 95.3%). This shows that the present method has higher recognition accuracy in human skeleton modeling and action recognition tasks. Compared with ST-GCN and CTR-GCN, the present method enables the adjacency matrix to be adaptively adjusted through a dynamic graph generation network, thereby more flexibly capturing changes in human skeleton structure. In addition, the introduction of adaptive graph convolution improves the modeling ability of long-term dependencies, making the model perform better in complex action recognition scenarios. The experimental results verify the effectiveness and advantages of the present method.
[0044] The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution of the present invention solves the complexity and accuracy problems of human skeleton modeling by redesigning the adjacency matrix that reflects the positional relationship of human skeleton points and combining time and space modeling, thereby improving recognition accuracy; dynamically generates an adjacency matrix that adapts to different individuals and action modes, which can not only adaptively adjust the connection relationship between key points, but also capture long-term dependencies and local details.
Claims
1. A human skeleton modeling method based on dynamic graph generation and adaptive graph convolution, characterized in that: Here are the steps: Step 1, obtaining human behavior skeleton point map data as the original data set; Step 2, preprocessing and dividing the original data set; Step 3: Design a dynamic graph generator to dynamically generate an optimized adjacency matrix that adapts to different individuals and action modes by using the preprocessed data set through node feature encoding, graph structure generation, and adjacency matrix optimization. Step 4: Input the preprocessed data set and the optimized adjacency matrix into the spatiotemporal graph convolution to extract spatiotemporal joint features; Step 5: Input the spatiotemporal joint features into a classification head, output the model’s predicted action category probability, compare it with the true category of the input data, calculate the cross entropy loss and back propagate, obtain the trained human action recognition model, and output the final recognition result.
2. The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution according to claim 1, characterized in that: The step 1 specifically includes: obtaining NTU RGB+D 60 / 120 human behavior skeleton point map data in the NTU data set as the original data set.
3. The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1, normalize the human behavior skeleton point map data, specifically: calculate the mean and standard deviation of all samples in each coordinate channel, and use them to standardize the data so that the coordinate values of all skeleton points are distributed in the same range; the normalized data maintains the original time series structure, each frame corresponds to a set of skeleton key points, and each key point contains information from multiple channels; The data format after normalization is a three-dimensional tensor dataset M´: M´=reshape(M,(N,t,V,C))(1) Among them, M is the original data set, M , N is the batch size, t is the time step, V is the number of key points, and C is the number of coordinate channels, including normalized coordinates, bone vectors, and motion speed; Step 2.2: Add random time jitter, that is, randomly discard some frames or insert extra frames; Step 2.3: Randomly hide some key points; Step 2.4: Generate pseudo targets, including frame sequence prediction and node prediction; Step 2.5: Divide the preprocessed dataset into training set X train and the test set X test , the format of the divided data set is as follows: (2)。 4. The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution according to claim 1, characterized in that: The step 3 is specifically as follows: the dynamic graph generator includes a node feature encoder, a graph structure generator, and a weight matrix optimization; The node feature encoder encodes the skeleton node features in the input data set, and performs nonlinear transformation on these features through a multi-layer perceptron to generate high-dimensional features H. The formula is as follows: H=σ(W2⋅σ(W1X+b1)+b 2 ) (3) Where: X is the input data set, D represents the dimension of the high-dimensional space, W1 and W2 are the learnable weight matrices of the first and second layers respectively; b1 and b2 are bias terms; σ represents a nonlinear activation function; The skeleton node features include spatial coordinates and motion speed information; The graph structure generator calculates the correlation weights between nodes through a multi-head attention mechanism and generates a dynamic adjacency matrix. The formula is as follows: Calculate the weighted attention matrix between nodes based on the high-dimensional feature H : (4) Where W is a learnable weight matrix used to generate queries and keys, D1 represents the feature dimension, and T represents the matrix transpose. Dynamic adjacency matrix The formation is as follows: (5) in represents an initialized learnable parameter matrix; The weight matrix optimization is applied to the dynamic adjacency matrix Perform weighted screening to remove redundant connections and capture higher-order graph structure relationships, specifically: (6) Where L represents the number of layers, and γ represents the weight parameter, which is used to balance the adjacency matrix of the current layer and the previous layer. The adjacency matrix after screening and optimization for the Lth layer can adapt to different individuals and action modes.
5. The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution according to claim 1, characterized in that: The step 4 is specifically as follows: Step 4.1: Extract spatial features of skeleton data through CTR-GCN: (7) Where X is the input bone dataset, is the adjacency matrix after optimization in step 3, is the extracted spatial feature; Specifically: (8) in That is, add the self-connected adjacency matrix, I is the identity matrix, yes The degree matrix of , W is the learnable weight matrix, (·) is the activation function, and F is the extracted spatial feature; Step 4.2: Capture dependencies at different time scales through a multi-scale spatiotemporal attention mechanism; Specifically, the time scale is divided into three scales: short, medium, and long. For each time scale, defined as s, the attention weight is calculated as follows: (9) in , , , respectively represent query, key and value, where W represents the learnable weight matrix, D1 represents the feature dimension, and then weighted fusion is performed, the formula is as follows: (10) in, Refers to the weight of each time scale, which is a learnable parameter. Refers to the extracted temporal features; Step 4.3, the extracted spatial features and temporal features are fused and normalized to generate the final spatiotemporal joint features: (11) in, Refers to the final spatiotemporal joint features, and LayerNorm refers to layer normalization.
6. The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution according to claim 1, characterized in that: The step 5 is specifically as follows: The spatiotemporal joint features obtained in step 4 are The input is input into a classification head containing a global average pooling layer and a fully connected layer to output the model's predicted action category probability. The model's predicted action category probability is compared with the true category of the input data, the cross entropy loss is calculated, and the network parameters are iteratively trained and optimized through the Adam optimizer and back propagation algorithm. After the training is completed, the trained human behavior recognition model is obtained and the model is saved and output.
7. The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution according to claim 6, characterized in that: The cross entropy loss calculation formula is as follows: (12) Where L is the cross entropy loss, c represents the category index, is the one-hot encoding of the true category. If category c is the correct category, then y c =1, otherwise y c =0; It is the action category probability predicted by the model, that is, the Softmax output.
8. The human skeleton modeling method based on dynamic graph generation and adaptive graph convolution according to claim 6, characterized in that: The network parameters include: batch size, learning rate, weight decay, optimizer, and training cycle.
Citation Information
Cited By
Rapid nondestructive detection method and system for lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning
CN121207914A
Action recognition method, model training method, electronic device and program product
CN122244955A
Method of recognizing action, method of training model, electronic device, program product
CN122244955B