Human Motion Posture Recognition Method Based on Deep Learning and Intelligent Wearable Device

By adopting deep learning methods in human posture recognition, combining gamma correction, nonlinear operation and three-dimensional convolutional network, the problems of large amount of calculation and low recognition accuracy are solved, and high-precision and real-time human motion posture recognition are achieved.

CN119229520BActive Publication Date: 2025-06-24HEXI (GUANGZHOU) INFORMATION TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411078357.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2025-06-24
Estimated Expiration
2044-08-07

AI Technical Summary

Technical Problem

The calculation amount is large and the recognition accuracy is low in human posture recognition.

Method used

The human body motion posture recognition method is adopted based on deep learning, and the spatial and temporal features and graph structural features are extracted through technologies such as gamma correction, nonlinear operation, three-dimensional maximum pooling and three-dimensional graph convolution, combined with a hybrid model of three-dimensional convolution neural network and three-dimensional graph neural network.

Benefits of technology

It improves the accuracy and robustness of human posture recognition, reduces the amount of calculation, and realizes high-precision recognition and real-time feedback of human movement postures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229520B_ABST
    Figure CN119229520B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a human motion posture recognition method and an intelligent wearable device based on deep learning, which relates to the technical field of image recognition. The method includes the following steps: Step S1: Obtain an image of the original motion data and perform gamma correction processing; Step S2: Obtain the feature vector after feature extraction; Step S3: Perform a non-linear operation; Step S4: Perform a three-dimensional maximum pooling operation to obtain the downsampled feature vector; Step S5: Repeat the above three-dimensional maximum pooling operation several times; Step S6: Extract graphic features from the downsampled feature vector; Step S7: Perform a three-dimensional TopK graph pooling operation to obtain the downsampled feature matrix; Step S8: Repeat the above three-dimensional TopK graph pooling operation several times; Step S9: Map the downsampled feature matrix to the classification space to obtain the classification probability of the motion posture. The present application solves the technical problems of large computational complexity and low recognition accuracy in human posture recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a human motion posture recognition method and a smart wearable device based on deep learning. Background Art

[0002] Human motion posture recognition is an important research direction in the field of computer vision and artificial intelligence, and has broad application prospects, such as health monitoring, motion analysis, virtual reality, and human-computer interaction. Traditional posture recognition methods mainly rely on image processing and machine learning techniques. However, these methods often show limitations such as large computational complexity and low recognition accuracy when faced with complex motion patterns and diverse environmental changes.

[0003] In recent years, with the development of deep learning technology, convolutional neural networks and graph neural networks have demonstrated powerful capabilities in feature extraction and pattern recognition, but a single network structure still has certain limitations when processing high-dimensional, spatiotemporally correlated motion data.

[0004] Therefore, the large amount of computation and low recognition accuracy in human posture recognition have become technical problems that need to be solved urgently. Summary of the invention

[0005] The present application provides a human motion posture recognition method and an intelligent wearable device based on deep learning to solve the technical problems of large computational complexity and low recognition accuracy in human posture recognition.

[0006] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0007] In a first aspect, the present invention provides a method for human motion posture recognition based on deep learning, comprising the following steps:

[0008] Step S1: obtaining an image of original motion data, and performing gamma correction on the image of the original motion data to obtain motion data after gamma correction processing;

[0009] Step S2: extracting features from the motion data after gamma correction to obtain a feature vector after feature extraction;

[0010] Step S3: performing a nonlinear operation on the feature vector after feature extraction to obtain a feature vector after nonlinear operation;

[0011] Step S4: performing a three-dimensional maximum pooling operation on the feature vector after the nonlinear operation to obtain a downsampled feature vector;

[0012] Step S5: Repeat the above three-dimensional maximum pooling operation several times to obtain the downsampled feature vector output by the three-dimensional maximum pooling operation;

[0013] Step S6: Extract graphic features from the downsampled feature vectors based on the topological relationship of the graph neural network nodes;

[0014] Step S7: Perform a three-dimensional TopK graph pooling operation on the graphic features to obtain a downsampled feature matrix;

[0015] Step S8: Repeat the above three-dimensional TopK graph pooling operation several times, and correspondingly obtain the downsampled feature matrices output after the three-dimensional TopK graph pooling operation;

[0016] Step S9: Map the downsampled feature matrix in Step S8 to the classification space to obtain the classification probability of the motion posture.

[0017] A further technical solution lies in: In the said Step S1, perform gamma correction on the image;

[0018]

[0019] wherein, X is the motion data after gamma correction processing; I in is the input image; γ is the gamma value, usually taken from 0.5 to 2.5.

[0020] A further technical solution lies in: In the said Step S2, extract spatio-temporal features from the input three-dimensional motion data, and extract local spatio-temporal features in the motion data through convolution operations;

[0021] Y = X * W + b

[0022] wherein: Y is the feature vector after feature extraction; X is the motion data after gamma correction; W is the weight of the convolution operation; b is the bias.

[0023] A further technical solution lies in: In the said Step S3, introducing an activation function can effectively alleviate the vanishing gradient and help the model converge faster;

[0024] Y′ = max(0, x)

[0025] wherein, Y′ is the feature vector after non-linearly operating on the feature vector Y after feature extraction, indicating taking the maximum value between 0 and x; x is the input data, that is, the feature vector Y after feature extraction.

[0026] A further technical solution lies in: In the said Step S4, introducing a three-dimensional max pooling operation, and the process of the three-dimensional max pooling operation includes: dividing the input non-linearly operated feature vector Y′ into several rectangular regions, and the rectangular regions are pooling windows or filters, and then selecting the maximum value in the sub-regions within each rectangular region as the output of the rectangular region;

[0027] Yt,h,w,d = max t′,h′,w′ (X t+t′,h+h′,w+w′,d )

[0028] Wherein, Y t,h,w,d is the node feature matrix, that is, the feature vector after downsampling; X t+t′,h+h′,w+w′,d is the eigenvalue of each local area; max() means to select the largest element value as the representative value of the area within each rectangular area; max t′,h′,w′ (X t+t′,h+h′,w+w′,d ) means to select the node with the largest eigenvalue from each local area X t+t′,h+h′,w+w′,d to form a new node feature matrix Y t,h,w,d .

[0029] A further technical solution lies in that: in the step S6, the topological relationship of the graph neural network nodes is constructed according to the human body bone graph;

[0030] H′ = σ(AHW)

[0031] Wherein, H′ is the graph feature; σ is the activation function LeakyReLU; A is the topological relationship of the graph neural network nodes, that is, the adjacency matrix; H is the feature vector after downsampling in step S5, that is, the node feature matrix; W is the graph convolution weight matrix.

[0032] A further technical solution lies in that: the step S7 specifically includes the following steps:

[0033] Step S701: Calculate the importance score of the node;

[0034] s i = MLP(h i )

[0035] Wherein, MLP is a multi-layer perceptron Multi-LayerPerceptron scoring function, h i is the feature vector of the i-th node in the graph feature H′, and s i is the importance score of each node;

[0036] Step S702: Sort the importance scores from largest to smallest and select the top k nodes;

[0037] Indices = TopK(s i , k)

[0038] Wherein, Indices is the index of the top k nodes with the highest scores returned by the TopK function, that is, select the top k nodes with the highest scores to form a new node set; s i is the importance score of each node; k is the node;

[0039] Step S703: Reconstruct the selected nodes to form a new adjacency matrix A' and a node feature matrix H'';

[0040] H'' = H'[Indices]

[0041] A' = A[Indices,Indices]

[0042] Wherein, A' is the new adjacency matrix; H'' is the downsampled feature matrix, i.e., the node feature matrix; Indices is the index of the top k nodes with the highest scores.

[0043] A further technical solution lies in: In the said step S9,

[0044] P = softmax(W f H'' + b f )

[0045] Wherein, P is the classification probability; W f is the weight matrix of the fully connected layer; b f is the bias of the fully connected operation.

[0046] A further technical solution lies in: The neural network model structure adopted includes a gamma correction module, a three-dimensional convolutional layer feature extraction module, an activation function ReLU, a three-dimensional max pooling layer, a three-dimensional graph convolutional layer, a three-dimensional TopK graph pooling layer, and a fully connected layer connected in sequence; wherein,

[0047] Input the image of the obtained original motion data into the gamma correction module, the gamma correction module outputs the motion data after gamma correction processing to the three-dimensional convolutional layer feature extraction module, the three-dimensional convolutional layer feature extraction module outputs the feature vector after feature extraction to the activation function ReLU, the activation function ReLU outputs the feature vector after non-linear operation to the three-dimensional max pooling layer, the three-dimensional max pooling layer outputs the downsampled feature vector to the three-dimensional graph convolutional layer, the topological relationship of the graph neural network nodes is input to the three-dimensional graph convolutional layer, the three-dimensional graph convolutional layer outputs the graph features to the three-dimensional TopK graph pooling layer, the three-dimensional TopK graph pooling layer outputs the downsampled feature matrix to the fully connected layer, and the fully connected layer outputs the classification probability of the motion posture.

[0048] In a second aspect, the present invention provides an intelligent wearable device, including a front-end sensor component, a wireless communication module, and a display module;

[0049] The sensor component includes: a nine-axis gyroscope, an acceleration sensor, an angular velocity sensor, a magnetometer, and a depth camera, and is used to capture the motion data of the wearer's limbs;

[0050] The wireless communication module is used to transmit the captured motion data of the wearer's limbs to the back-end processing system. Among them, the processor in the back-end processing system is used to execute the deep learning-based human motion posture recognition method described in the first aspect, for highly accurate recognition of the wearer's human motion posture, and transmit the motion posture recognition result back to the wireless communication module through the wireless network;

[0051] The display module is used to display the final motion posture recognition result of the wearer.

[0052] The beneficial effects produced by adopting the above technical solutions are as follows:

[0053] The deep learning-based human motion posture recognition method of the present invention has high recognition accuracy in human posture recognition through gamma correction, etc., and has a small amount of calculation in human posture recognition through non-linear operations, etc.; it has high recognition accuracy in human posture recognition through a gamma correction module, etc., and has a small amount of calculation in human posture recognition through an activation function ReLU, etc. By combining a hybrid model of a three-dimensional convolutional neural network and a three-dimensional graph neural network, it effectively extracts the spatio-temporal features in the motion video and the graph structure features of the human motion skeleton, realizes highly accurate recognition and real-time feedback of the human motion posture, and effectively improves the accuracy and robustness of the recognition by using gamma correction technology. During the recognition process, it can analyze the wearer's motion posture in real time and give accurate feedback, such as posture correction suggestions, motion effect evaluation, emotion detection, fall warning, and atrial fibrillation monitoring, etc.

[0054] See the description in the specific implementation part for details. Description of the Drawings

[0055] Figure 1 is a flowchart of the deep learning-based human motion posture recognition method according to an embodiment of the present invention;

[0056] Figure 2 is a schematic diagram of the neural network model structure according to an embodiment of the present invention;

[0057] Figure 3 is Figure 1 a flowchart of step S7 in

[0058] Figure 4 is a schematic diagram of the structure of the smart wearable device according to an embodiment of the present invention. Specific Embodiments

[0059] To make the objectives, technical solutions and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts fall within the scope of protection of this application.

[0060] In addition, the term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after.

[0061] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Many specific details are set forth in the following description to facilitate a thorough understanding of this application, but this application may also be implemented in other ways different from those described herein. Those skilled in the art may make similar extensions without departing from the connotation of this application, so this application is not limited by the specific embodiments disclosed below.

[0062] First, for the front-end data of this application, intelligent wearable devices integrated with a variety of non-invasive sensors are used to capture the motion data of the limbs, such as acceleration sensors, gyroscopes, depth sensors, and inertial measurement units (IMUs), so as to collect the motion data of the wearer in real time; these data are transmitted to the back-end processing system through a wireless communication module; the back-end processing system uses deep learning technology to perform high-precision recognition of human motion postures. Specifically, the system first performs processing such as cropping, padding, and normalization on the received raw motion data through a preprocessing module to obtain standardized data. Then, a pose recognition network model based on deep learning is used to extract features and classify the standardized data, and finally, the motion pose recognition result of the wearer is output.

[0063] The backend processing system adopts an improved three-dimensional convolutional graph neural network 3DCNN-GNN to recognize action postures based on human skeletal movements. The three-dimensional convolutional neural network 3DCNN is used to extract spatio-temporal features from the motion data, and the graph neural network GNN is used to model these features to better understand the motion relationships of various parts of the human body. In addition, the present invention also adopts gamma correction technology to effectively improve the accuracy of target recognition. The pose recognition network model is trained and optimized with a large amount of training data with pose labels, and has strong generalization ability and robustness. During the recognition process, it can analyze the wearer's motion posture in real time and give accurate feedback, such as pose correction suggestions, motion effect evaluation, emotion detection, fall warning, and atrial fibrillation monitoring, etc.

[0064] This application effectively extracts spatio-temporal features and graph structure features of human motion skeletons in motion videos through a hybrid model combining a three-dimensional convolutional neural network and a three-dimensional graph neural network, realizes high-precision recognition and real-time feedback of human motion postures, and effectively improves the accuracy and robustness of recognition by using gamma correction technology.

[0065] Embodiment 1:

[0066] As Figures 1 to 3 shown, the present invention discloses a method for recognizing human motion postures based on deep learning. The neural network model structure adopted in the recognition process is described in detail as follows:

[0067] As Figure 2 shown, the neural network model structure for recognizing human motion postures is an improved three-dimensional convolutional graph neural network 3DCNN-GNN structure. This neural network model structure includes a gamma correction module, a three-dimensional convolutional layer feature extraction module, an activation function ReLU, a three-dimensional max pooling layer, a three-dimensional graph convolutional layer, a three-dimensional TopK graph pooling layer, and a fully connected layer connected in sequence. The image of the original motion data obtained is input into the gamma correction module, and the gamma correction module outputs the motion data after gamma correction processing to the three-dimensional convolutional layer feature extraction module. The three-dimensional convolutional layer feature extraction module outputs the feature vector after feature extraction to the activation function ReLU. The activation function ReLU outputs the feature vector after non-linear operation to the three-dimensional max pooling layer. The three-dimensional max pooling layer outputs the downsampled feature vector to the three-dimensional graph convolutional layer. The topological relationship of the graph neural network nodes is input into the three-dimensional graph convolutional layer. The three-dimensional graph convolutional layer outputs the graph features to the three-dimensional TopK graph pooling layer. The three-dimensional TopK graph pooling layer outputs the downsampled feature matrix to the fully connected layer. The fully connected layer outputs the classification probability of the motion posture.

[0068] Specifically, as Figure 1 shown, the present invention discloses a method for recognizing human motion postures, including the following steps:

[0069] Step S1: Obtain an image of the original motion data, and perform gamma correction on the image of the original motion data to obtain the motion data after gamma correction processing.

[0070] In this step S1, when the gamma correction module obtains the image of the original motion data, it performs gamma correction on the image of the original motion data to obtain the motion data after gamma correction processing.

[0071] Performing gamma correction on the image can adjust the brightness of the image, enhance the contrast of the image, and improve the visual quality.

[0072] Input: Input motion data, original sensor data I in , with a size of T×H×W×C.

[0073] T: Number of time steps.

[0074] H: Feature height.

[0075] W: Feature brightness.

[0076] C: Number of channels, i.e., the number of sensors.

[0077] Formula:

[0078] I in : Input image.

[0079] γ: Gamma value, usually taken from 0.5 to 2.5.

[0080] Function: Adjust the brightness and contrast of the image, and improve the visual quality.

[0081] Output: Motion data X after gamma correction processing.

[0082] Step S2: Extract features from the motion data after gamma correction processing to obtain the feature vector after feature extraction.

[0083] In this step S2, when the three-dimensional convolutional layer feature extraction module obtains the motion data after gamma correction processing, it extracts features from the motion data after gamma correction processing to obtain the feature vector after feature extraction.

[0084] Three-dimensional convolutional layer feature extraction is used to extract spatio-temporal features from the input three-dimensional motion data. Local spatio-temporal features in the motion data are extracted through convolution operations, and the size and number of convolutional kernels directly affect the effect of feature extraction.

[0085] Input: Gamma-corrected motion data cube X, with a size of T×H×W×C.

[0086] Output: The feature vector Y after feature extraction, with a size of T′×H′×W′×D1. Value range: Normalized to [0, 1]. The formula is:

[0087] Y = X * W + b

[0088] In the formula, W: The convolution kernel, the weight of the convolution operation, used to extract features from the input data. The convolution kernel determines which features to extract, and appropriate weights can extract key motion features. The weights are continuously adjusted through training to capture important patterns in the data. Value range: Usually randomly sampled from a Gaussian distribution or a uniform distribution during initialization. The value range varies according to the initialization strategy. For the Gaussian distribution, such as N(0, 0.01).

[0089] b: The bias, the bias of the convolution operation, used to adjust the output value. The bias value is used to adjust the convolution output and helps the model better fit the data. Value range: Usually set to 0 during initialization, or randomly sampled from a uniform distribution, with a relatively small value range.

[0090] T′, H′, W′: The sizes calculated through the stride and padding.

[0091] D1: The number of output channels.

[0092] Step S3: Perform a non-linear operation on the feature vector after feature extraction to obtain the feature vector after the non-linear operation.

[0093] In this step S3, when the activation function ReLU obtains the feature vector after feature extraction, a non-linear operation is performed to obtain the feature vector after the non-linear operation.

[0094] An activation function is introduced. The activation function ReLU can effectively alleviate the problem of gradient disappearance and help the model converge faster. Non-linearity is introduced to enable the model to learn complex features.

[0095] Input: The feature vector Y after feature extraction, with a size of T′×H′×W′×D1.

[0096] Output: The feature vector Y′ after the non-linear operation, with the same size as the input, T′×H′×W′×D1. max(0, x), that is, non-negative.

[0097] Y′ = max(0, x)

[0098] In the formula, Y′ is the feature vector after non-linearly operating on the feature vector Y after feature extraction, representing taking the maximum value between 0 and x; x is the input data, that is, the feature vector Y after feature extraction.

[0099] Step S4: Perform a three-dimensional max pooling operation on the feature vector after the non-linear operation to obtain the downsampled feature vector.

[0100] In step S4, when the three-dimensional max pooling layer obtains the feature vector after the non-linear operation, a three-dimensional max pooling operation is performed to obtain the downsampled feature vector.

[0101] The three-dimensional max pooling operation is introduced. The pooling operation is used for downsampling. Its operation process can be summarized as follows: the input matrix or feature map is divided into several rectangular regions, which are also called pooling windows or filters, and then the maximum value is selected in the sub-regions of each rectangular region as the output of the rectangular region, reducing the size of the feature map, retaining the main features, and reducing the computational amount.

[0102] Input: The feature vector Y' after non-linear operation, with a size of T'×H'×W'×D1.

[0103] Output: The downsampled feature vector, with a size of T″×H″×W″×D1.

[0104] Y t,h,w,d = max t′,h′,w′ (X t+t′,h+h′,w+w′,d )

[0105] In the formula, Y t,h,w,d is the node feature matrix, that is, the downsampled feature vector; X t+t′,h+h′,w+w′,d is the eigenvalue of each local region; max() means that in each rectangular region, the maximum element value is selected as the representative value of the region; max t′,h′,w′ (X t+t′,h+h′,w+w′,d ) means selecting the node with the largest eigenvalue from each local region X t+t′,h+h′,w+w′,d to form a new node feature matrix Y t,h,w,d .

[0106] T″, H″, W″ are the dimensions calculated through the pooling window and stride. The pooling window and stride control the range and stride of the pooling operation, determining the degree of downsampling. Value range: The common pooling window size is 2×2×2, and the stride is 2.

[0107] Step S5: Repeat the above three-dimensional max pooling operation several times until a smaller feature tensor is obtained, corresponding to obtaining the downsampled feature vector output by the three-dimensional max pooling operation.

[0108] The dimensions of different experimental samples are different, and the parameter thresholds are also different. The specific parameter adjustment depends on the actual experimental situation, corresponding to obtaining the downsampled feature vector output by the three-dimensional max pooling operation; taking the input feature tensor size of 8×8×8 as an example, the pooling window size is 2×2×2, and the stride is 2. The change in the size of the feature tensor after one pooling operation is as follows:

[0109] Number of operations Feature tensor size Initial input 8×8×8 First pooling 4×4×4

[0110] Step S6: Extract graphic features from the downsampled feature vectors based on the topological relationship of the graph neural network nodes.

[0111] In this step S6, when the 3D graph convolutional layer obtains the downsampled feature vectors, graphic features are extracted based on the topological relationship of the graph neural network nodes.

[0112] Construct the topological relationship of the graph neural network nodes according to the human body bone graph, and extract graphic features through the 3D graph convolutional layer.

[0113] Input: Downsampled feature vectors and the topological relationship of the graph neural network nodes

[0114] T″: The number of nodes, i.e., the number of joints.

[0115] D″: The number of features.

[0116] H″: Feature height.

[0117] W″: Feature brightness.

[0118] H: The downsampled feature vectors, i.e., the node feature matrix.

[0119] A: The topological relationship of the graph neural network nodes, i.e., the adjacency matrix, representing the connection relationship of the human body bone graph, i.e., the adjacency matrix of each node and edge. The nodes are joints and the edges are bones. The adjacency matrix defines the graph structure and affects the way features propagate in the graph. Value range: Determined according to the bone structure of the human body bone graph, with values of 0 or 1. 0 indicates no connection and 1 indicates connection.

[0120] Output: Graphic features

[0121] H′ = σ(AHW)

[0122] In the formula, H′ is the graphic feature; σ is the activation function LeakyReLU; A is the adjacency matrix; H is the node feature matrix; W is the graph convolution weight matrix.

[0123] W: The graph convolution weight matrix is used to learn the transfer relationship between node features. The learning effect of the weights directly affects the accuracy of pose recognition. Value range: Randomly sampled from a Gaussian distribution during initialization.

[0124] σ: Introduce the activation function LeakyReLU, which is an improved version of the ReLU function. When the input is negative, the output is no longer zero but has a small slope, such as 0.01, which helps avoid the problem of neuron death, increases the robustness of the neural network, and has a similar computational complexity to ReLU. The activation function can effectively alleviate the vanishing gradient problem, help the model converge faster, introduce non-linearity, and enable the model to learn complex features. The value range is: the output is max(0,x), that is, non-negative.

[0125] H”′: The number of output features.

[0126] Step S7: Perform 3D TopK graph pooling operation on the graph features to obtain the downsampled feature matrix.

[0127] Step S7: 3D TopK graph pooling layer.

[0128] When the 3D TopK graph pooling layer obtains the graph features, perform a downsampling operation to obtain the downsampled feature matrix.

[0129] The 3D TopK graph pooling layer reduces the dimension and aggregates the node features by performing a downsampling operation on the graph structure, aggregates the high-dimensional node features into a lower-dimensional representation, thereby retaining the main features, reducing the computational amount, and improving the generalization ability of the model.

[0130] Input: Graph features

[0131] Output: Downsampled feature matrix

[0132] T”′: The number of pooled nodes.

[0133] D”′: The dimension of the pooled features.

[0134] H”′: The height of the pooled features.

[0135] W”′: The brightness of the pooled features.

[0136] As Figure 3 shown, Step S7 specifically includes the following steps:

[0137] Step S701: Calculate the importance score of the nodes.

[0138] s i = MLP(h i )

[0139] In the formula, MLP is a multi-layer perceptron Multi-LayerPerceptron scoring function, h i is the feature vector of the i-th node in the graph feature H′, si is the importance score of each node.

[0140] Step S702: Sort the importance scores from largest to smallest and select the top k nodes.

[0141] Indices = TopK(s i , k)

[0142] where Indices is the index of the top k nodes with the highest scores returned by the TopK function, that is, select the top k nodes with the highest scores to form a new set of nodes; s i is the importance score of each node; k is the node.

[0143] Step S703: Reconstruct the selected nodes to form a new adjacency matrix A′ and node feature matrix H″.

[0144] H″ = H′[Indices]

[0145] A′ = A[Indices, Indices]

[0146] where A′ is the new adjacency matrix; H″ is the node feature matrix; Indices is the index of the top k nodes with the highest scores.

[0147] Step S8: Repeat the above three-dimensional TopK graph pooling operation several times, similar to Step S5, until a smaller feature representation is obtained, corresponding to the downsampled feature matrix output after the three-dimensional TopK graph pooling operation.

[0148] Step S9: Map the downsampled feature matrix in Step S8 to the classification space to obtain the classification probability of the motion posture.

[0149] In this Step S9, when the fully connected layer obtains the downsampled feature matrix H″ in Step S8, it is mapped to the classification space to obtain the classification probability P of the motion posture.

[0150] Mapping the high-dimensional features to the classification space determines the final posture classification result.

[0151] Input: Output of the graph convolutional layer.

[0152] Output: Classification probability P of the motion posture.

[0153] Formula: P = softmax(W f H″ + b f )

[0154] where P is the classification probability; W f is the weight matrix of the fully connected layer; b f is the bias of the fully connected operation.

[0155] W f : The weight matrix of the fully connected layer, which is used to map high-dimensional features to the output space. The weights determine how features are mapped to the classification results, and appropriate weights can improve the classification accuracy. Value range: Randomly sampled from a Gaussian distribution during initialization.

[0156] b f : The bias of the fully connected operation. The bias is used to adjust the output of the fully connected layer and helps the model better fit the data. Value range: Set to 0 during initialization.

[0157] The above embodiments of the present application, through a hybrid model combining a three-dimensional convolutional neural network and a three-dimensional graph neural network, effectively extract spatio-temporal features in motion videos and graph structure features of human motion skeletons, achieve high-precision recognition and real-time feedback of human motion postures, and effectively improve the accuracy and robustness of recognition by using gamma correction technology.

[0158] Embodiment 2

[0159] As Figure 4 shown, the present invention discloses an intelligent wearable device 200, including a front-end sensor component 210, a wireless communication module 220, and a display module 230;

[0160] The sensor component 210 includes: a nine-axis gyroscope, an acceleration sensor, an angular velocity sensor, a magnetometer, and a depth camera, which are used to capture the motion data of the wearer's limbs;

[0161] The wireless communication module 220 is used to transmit the captured motion data of the wearer's limbs to the back-end processing system; wherein, the processor in the back-end processing system is used to execute the deep learning-based human motion posture recognition method as in Embodiment 1, for high-precision recognition of the wearer's human motion posture, and transmit the motion posture recognition result back to the wireless communication module 220 through a wireless network;

[0162] The display module 230 is used to display the final motion posture recognition result of the wearer, and the display information may include the current motion posture and accurate feedback for the motion posture information, such as posture correction suggestions, motion effect evaluation, emotion detection, fall warning, and atrial fibrillation monitoring, etc.

[0163] According to this embodiment, the smart wearable device 200 can analyze the wearer's motion posture in real time and give accurate feedback, such as posture correction suggestions, exercise effect evaluation, emotion detection, fall warning, and atrial fibrillation monitoring. By combining a hybrid model of a three-dimensional convolutional neural network and a three-dimensional graph neural network, the spatio-temporal features in the motion video and the graph structure features of the human motion skeleton are effectively extracted, realizing high-precision recognition and real-time feedback of the human motion posture, and the gamma correction technology is used to effectively improve the accuracy and robustness of the recognition.

[0164] The specific embodiments are as follows:

[0165] By means of a standardized experimental environment and unified experimental parameter settings, the fairness and repeatability of the experimental results can be guaranteed. These detailed configuration and parameter descriptions help other researchers to repeat the experiment under the same or similar conditions and verify the results.

[0166] As shown in Table 1, it is the software and hardware experimental environment table used in this experimental algorithm. As shown in Table 2, it is the parameter setting of the algorithm. As shown in Table 3, it is the result of the comparative experiment.

[0167] Table 1: Experimental Environment Table

[0168]

[0169]

[0170] Table 2: Experimental Parameter Table

[0171]

[0172]

[0173] Table 3: Comparative Experiment Result Table

[0174]

[0175] It can be seen from the above results that the improved three-dimensional convolutional graph neural network 3DCNN-GNN shows better performance than the graph convolutional neural network in various evaluation criteria, and can more effectively extract spatio-temporal features, improving the accuracy, recall rate and overall performance of motion posture recognition.

[0176] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitations are imposed herein.

[0177] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A method for human motion posture recognition based on deep learning, characterized in that: The steps include: Step S1: obtaining an image of original motion data, and performing gamma correction on the image of the original motion data to obtain motion data after gamma correction processing; Step S2: extracting features from the motion data after gamma correction to obtain feature vectors after feature extraction; in the step S2, extracting spatiotemporal features from the input three-dimensional motion data, and extracting local spatiotemporal features in the motion data through convolution operation; ; Where: is the feature vector after feature extraction; is the gamma-corrected motion data; is the weight of the convolution operation; is bias; Step S3: performing a nonlinear operation on the feature vector after feature extraction to obtain a feature vector after nonlinear operation; Step S4: performing a three-dimensional maximum pooling operation on the feature vector after the nonlinear operation to obtain a downsampled feature vector; Step S5: Repeat the above three-dimensional maximum pooling operation several times to obtain the downsampled feature vector output by the three-dimensional maximum pooling operation; Step S6: Based on the topological relationship of the graph neural network nodes, the graph features are extracted from the downsampled feature vectors; Step S7: Perform a three-dimensional TopK graph pooling operation on the graph features to obtain a downsampled feature matrix; Step S8: Repeat the above three-dimensional TopK graph pooling operation several times, and obtain the downsampled feature matrix output after the three-dimensional TopK graph pooling operation; Step S9: Mapping the feature matrix downsampled in step S8 to the classification space to obtain the classification probability of the motion posture; The neural network model structure adopted by the method includes a gamma correction module, a three-dimensional convolution layer feature extraction module, an activation function ReLU, a three-dimensional maximum pooling layer, a three-dimensional graph convolution layer, a three-dimensional TopK graph pooling layer and a fully connected layer connected in sequence; wherein, The image of the obtained original motion data is input into the gamma correction module, the gamma correction module outputs the motion data after gamma correction to the three-dimensional convolutional layer feature extraction module, the three-dimensional convolutional layer feature extraction module outputs the feature vector after feature extraction to the activation function ReLU, the activation function ReLU outputs the feature vector after nonlinear operation to the three-dimensional maximum pooling layer, the three-dimensional maximum pooling layer outputs the downsampled feature vector to the three-dimensional graph convolution layer, the topological relationship of the graph neural network nodes is input into the three-dimensional graph convolution layer, the three-dimensional graph convolution layer outputs the graphic features to the three-dimensional TopK graph pooling layer, the three-dimensional TopK graph pooling layer outputs the downsampled feature matrix to the fully connected layer, and the fully connected layer outputs the classification probability of the motion posture.

2. The method for human motion posture recognition based on deep learning according to claim 1, characterized in that: In the step S1, gamma correction is performed on the image to enhance the contrast of the image; ; In the formula, The motion data is processed for gamma correction; is the input image; is the gamma value, ranging from 0.5 to 2.

5.

3. The method for human motion posture recognition based on deep learning according to claim 2, characterized in that: In step S3, introducing the activation function can effectively alleviate the gradient disappearance and help the model converge faster; ; In the formula, The feature vector after feature extraction The eigenvector after nonlinear operation, indicating the value of 0 and The maximum value between is the input data, that is, the feature vector after feature extraction .

4. The method for human motion posture recognition based on deep learning according to claim 3, characterized in that: In step S4, a three-dimensional maximum pooling operation is introduced. The three-dimensional maximum pooling operation process includes: inputting the feature vector after nonlinear operation Divide into several rectangular areas, the rectangular areas are pooling windows or filters, and then select the maximum value in the sub-area within each rectangular area as the output of the rectangular area; ; In the formula, is the node feature matrix, that is, the feature vector after downsampling; is the characteristic value of each local area; max() means that in each rectangular area, the largest element value is selected as the representative value of the area; Indicates that from each local area Select the node with the largest eigenvalue to form a new node feature matrix .

5. The method for human motion posture recognition based on deep learning according to claim 4, characterized in that: In the step S6, the topological relationship of the graph neural network nodes is obtained according to the human skeleton graph; ; In the formula, is a graphic feature; is the activation function LeakyReLU; It is the topological relationship of the nodes of the graph neural network, i.e. the adjacency matrix; is the feature vector after downsampling in step S5, i.e., the node feature matrix; is the graph convolution weight matrix.

6. The method for human motion posture recognition based on deep learning according to claim 5, characterized in that: The step S7 specifically includes the following steps: Step S701: Calculate the importance score of the node; ; In the formula, MLP is a multi-layer perceptron scoring function. It is a graphic feature The feature vector of the i-th node in , is the importance score of each node; Step S702: sort the importance scores from large to small and select the first k nodes; =TopK(s i ,k); In the formula, The TopK function returns the indexes of the top k nodes with the highest scores, that is, the top k nodes with the highest scores are selected to form a new node set; is the importance score of each node; k is the node; Step S703: Reconstruct the selected nodes to form a new adjacency matrix and the node feature matrix ; ; = ; In the formula, is the new adjacency matrix; is the feature matrix after downsampling, i.e., the node feature matrix; are the indices of the first k nodes with the highest scores.

7. The method for human motion posture recognition based on deep learning according to claim 6, characterized in that: In the step S9, ; In the formula, is the classification probability; is the weight matrix of the fully connected layer; is the bias for fully connected operation.

8. A smart wearable device, characterized in that: Including front-end sensor components, wireless communication modules, and display modules; The sensor assembly includes: a nine-axis gyroscope, an acceleration sensor, an angular velocity sensor, a magnetometer, and a depth camera, which are used to capture the motion data of the wearer's limbs; The wireless communication module is used to transmit the captured motion data of the wearer's limbs to a back-end processing system; wherein the processor in the back-end processing system is used to execute the human motion posture recognition method based on deep learning according to any one of claims 1 to 7, for high-precision recognition of the wearer's human motion posture, and transmit the motion posture recognition result back to the wireless communication module through a wireless network; The display module is used to display the final wearer's motion posture recognition result.

Citation Information

Patent Citations

  • Illumination and head posture robust expression recognition method and device and storage medium

    CN112541422A

  • Gesture collaborative graph convolution gait recognition method

    CN116486437A

  • Dynamic gesture recognition method combining multi-modal inter-frame movement and shared attention weight

    CN118072395A