Millimeter wave radar motion evaluation method, system and equipment based on human skeleton model

Through the improved DBSCAN clustering algorithm and PointNet neural network model, the human skeleton structure is reconstructed, combined with the ST-GCN network to analyze the motion trajectory, the problem of low accuracy in motion evaluation of existing millimeter wave radars is solved, and high-precision motion monitoring and evaluation in complex environments is achieved.

CN120392055APending Publication Date: 2025-08-01UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510496855.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing millimeter-wave radar technology is difficult to accurately extract joints and bone nodes of the human body during motion evaluation, resulting in low accuracy of motion evaluation and poor results in traditional methods in complex environments.

Method used

The improved DBSCAN clustering algorithm combined with weighted distance calculation is used to reconstruct the human skeleton structure through the PointNet neural network model, and the ST-GCN network is used to analyze the motion trajectory to achieve high-precision skeleton node estimation and motion evaluation.

Benefits of technology

It improves the accuracy and robustness of bone node extraction, can work stably in complex environments, provides high-precision motion monitoring and evaluation, and improves user comfort and operation convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120392055A_ABST
    Figure CN120392055A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of millimeter-wave radars, in particular to a millimeter-wave radar action evaluation method, system and equipment based on a human skeleton model, and the method comprises the following steps: S1, millimeter-wave radar point cloud sampling: employing millimeter-wave radar equipment to collect three-dimensional point cloud data of a human body, and extracting target points related to human skeleton nodes; s2, skeleton node estimation: predicting three-dimensional coordinates of each joint of the human body by adjusting an output structure of a PointNet neural network model, so as to reconstruct and fit a skeleton structure of the human body; and S3, human body action evaluation: constructing the obtained three-dimensional coordinate data of the multiple frames of human body skeleton nodes into time sequence data, and analyzing a motion track by using an evaluation system based on an ST-GCN network. According to the invention, complex dynamic changes and behavior modes in human motion can be accurately captured, and high-precision real-time evaluation and feedback are provided for the fields of motion monitoring, rehabilitation training, health evaluation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of millimeter-wave radar, and in particular to a millimeter-wave radar action evaluation method, system and device based on a human body bone model. Background Art

[0002] With the continuous growth of the global demand for health management and medical rehabilitation, action evaluation has gradually become an important research direction. Most current action evaluation systems rely on wearable sensors or visual capture technologies (such as infrared cameras and Kinect, etc.) to monitor human actions. However, these technologies have certain limitations. Wearable devices often affect the natural activities of users, while visual capture systems are highly dependent on factors such as lighting conditions and installation environments, and cannot achieve effective action capture and analysis in complex environments. To solve these problems, researchers have begun to focus on using millimeter-wave radar for human action evaluation. Due to its characteristics of high penetration, non-contact, and insensitivity to lighting conditions, millimeter-wave radar can perform accurate target detection in a variety of environments, and thus has become an emerging direction in the field of action evaluation.

[0003] Currently, in the field of action evaluation, millimeter-wave radar is mainly used to detect the position information of the human body and achieve simple action recognition. Common solutions include obtaining the reflected point cloud data of the human body through millimeter-wave radar and analyzing the data to identify some simple actions or static postures. However, the point cloud data obtained in the prior art is often sparse, and since the point cloud cannot clearly outline the entire human body, it is difficult for the prior art to effectively extract detailed information such as joints and bone nodes of the whole human body, thereby limiting the accuracy and application scenarios of action evaluation. Therefore, how to improve the parsing accuracy of point cloud data, obtain accurate bone node coordinates, make full use of emerging technologies such as deep learning and apply them to complex action evaluation scenarios is still a technical difficulty faced by this field. Summary of the Invention

[0004] The present invention provides a millimeter-wave radar action evaluation method, system and device based on a human body bone model.

[0005] A millimeter-wave radar action evaluation method based on a human body bone model includes the following steps:

[0006] S1, millimeter-wave radar point cloud sampling: Use a millimeter-wave radar device to collect three-dimensional point cloud data of the human body, and adopt an improved DBSCAN clustering algorithm to optimize the sampling accuracy of different dimensions (horizontal, vertical, and hierarchical) by introducing a weighting factor, and extract target points related to human bone nodes;

[0007] S2, skeleton node estimation: Based on the extracted target points related to the human skeleton nodes, the 3D coordinates of each human joint are predicted by adjusting the output module of the PointNet neural network model, thereby reconstructing and fitting the human skeleton structure;

[0008] S3, human motion evaluation: The acquired three-dimensional coordinate data of multiple frames of human skeleton nodes are constructed into time series data, and the motion trajectory is analyzed using an evaluation system based on the ST-GCN network to achieve evaluation and feedback of human motion behavior.

[0009] Optionally, the millimeter-wave radar point cloud sampling in S1 includes:

[0010] S11, millimeter-wave radar point cloud two-dimensional projection: The millimeter-wave radar point cloud is projected two-dimensionally along the radar's front view. After projection, the coordinates of each target point are converted from three-dimensional to two-dimensional plane coordinates, while retaining the point cloud layer information, spherical sequence number and horizontal angle. The target point information is defined as a five-tuple (x, z, y, Kn, θ h ,θ q ), where x, z are the positions of the target point in the two-dimensional plane, y is the distance from the target point to the radar, Kn is the spherical number of the target point, θ h is the horizontal angle of the target point, θ q is the pitch angle of the target point;

[0011] S12, target point distance calculation improvement: introduce weighting factors to optimize the calculation method of the distance between target points, define the target point A (x1, z1, y1, K n1 ,θ h1 ,θ q1 ) and B(x2,z2,y2,K n2 ,θ h2 ,θ q2 ), expressed as:

[0012]

[0013] Δz=α z (z2-z1);

[0014]

[0015] Among them, α z is the weighting factor along the z and x directions, adjusting the weights in different directions, α y is the weighting factor along the y direction, balancing the effect of the sphere spacing, and d is the fixed distance between adjacent spheres;

[0016] S13, Clustering process: After improving the calculation of the distances to the target points, the improved DBSCAN algorithm is used to perform clustering analysis on the depth point cloud data. By calculating the neighborhood radius ε of each point, clustering clusters of points are formed.

[0017] Optionally, the clustering process in S13 includes:

[0018] S131, Parameter setting: According to the density characteristics of the point cloud, the minimum number of samples MinPts is set.

[0019] S132, Neighborhood radius: For each target point P, by introducing an angle-radius coupling model, the neighborhood radius is dynamically corrected. The neighborhood radius ε can be expressed by the following formula:

[0020] ε = ε0·(1 + k·(|θ h | + |θ q |))

[0021] Where k is the angle sensitivity coefficient, which controls the influence weight of the angle on the neighborhood radius. When the incident angle of the radar beam increases, the projection density of the point cloud in the local area decreases, and the neighborhood radius needs to be expanded to ensure the robustness of clustering.

[0022] S133, Point sequence sorting and clustering formation: Starting from the core points, all density-reachable points are recursively merged to form the clustering cluster Ck.

[0023] S134, Clustering result analysis: Points that do not belong to any cluster are marked as noise points and removed from the output. The distances from the sample points in all clusters to the hierarchical center are calculated and sorted, and the priority decreases gradually from far to near.

[0024] Optionally, the skeleton node estimation in S2 includes:

[0025] S21, Data preprocessing: After clustering, the three-dimensional coordinates of the target points are represented as (x, y, z). The target points are normalized, and data augmentation is performed through rotation, translation, and noise perturbation.

[0026] S22, Network structure design: By adjusting the output structure of the PointNet neural network model, the three-dimensional coordinates of each joint of the human body are predicted. Specifically, it includes:

[0027] Input layer: Accept the target points after clustering.

[0028] Spatial transformation layer: The learnable parameterized transformation matrix (256×9) implicitly learns the spatial mapping relationship between millimeter-wave radar point clouds and human bone nodes. One-dimensional convolution with shared weights (Conv1D, kernel_size = 1) is used to extract high-dimensional features point by point, generating an intermediate feature tensor of n×1024. After max pooling, it is compressed into a global feature vector of 1×1024, and then reduced to 256 through a fully connected layer and multiplied by the transformation matrix to be reconstructed into a 3×3 rotation matrix;

[0029] Feature extraction layer: High-dimensional spatial features are gradually extracted from the normalized point cloud. The input point cloud first learns the local spatial association of a single point through a 1×3 convolution kernel (64 channels), generating a primary feature tensor of n×64; the spatial transformation network further learns a 64×64 rotation matrix to orthogonally constrain the feature space and enhance geometric consistency. Subsequently, it is gradually expanded to 128 and 1024 channels through two levels of 1×1 convolution, deeply mining the implicit bone topological structure information in the point cloud, and finally outputting a high-dimensional feature tensor of n×1024 to provide support for the three-dimensional coordinate regression of human bone nodes;

[0030] Max pooling layer: Maps local features to the global representation space. The input is an n×1024-dimensional point-level feature matrix, and the maximum value sampling is performed channel by channel along the point set dimension, generating a 1×1024-dimensional global feature vector independent of the point order, only retaining the most significant response value of each feature channel;

[0031] Output layer: Outputs the predicted three-dimensional coordinates of 11 key human bone nodes;

[0032] Regression output layer: The input is a 1024-dimensional global feature vector, which is reduced through a fully connected layer with 512 neurons, and the ReLU activation function is used to introduce non-linear expression ability; during the training phase, 30% of the Dropout layer is used to randomly mask some neuron connections to suppress the risk of overfitting and enhance the generalization ability to radar point cloud dynamic noise; the normalization layer normalizes the 512-dimensional features to zero mean to alleviate the abnormal fluctuation of the gradient to accelerate the training convergence; finally, through a 33-dimensional fully connected layer linear mapping, the predicted values of the three-dimensional spatial coordinates of 11 human bone nodes are output, which are strictly aligned with the Kinect data. This design balances the model capacity and computational efficiency through feature compression (1024→512→33), and the fully connected weight matrix can be interpreted as a linear decoder from global features to spatial coordinates, directly establishing the mathematical relationship between the point cloud distribution pattern and the human bone geometric structure.

[0033] S23, Training and optimization: The mean square error (MSE) is used as the loss function to minimize the error between the predicted coordinates and the true coordinates. The Adam optimizer is used to adaptively adjust the learning rate, and the loss change is monitored on the validation set;

[0034] S24, Prediction and fitting results: After the training of the PointNet neural network model is completed, the millimeter-wave radar point cloud data collected in real time is input into the model, and the three-dimensional coordinates of 11 key bone nodes are output to realize the fitting and reconstruction of the human bone structure, and the superposition effect of the bone nodes and the point cloud data is visually displayed to verify the prediction accuracy.

[0035] Optionally, the human motion evaluation in S3 includes:

[0036] S31, Data construction and preprocessing: Continuously collect the three-dimensional coordinates of multiple frames of bone nodes, set a time window, arrange each frame of data in chronological order to form a complete motion time series, and calculate the relative positions, velocities, and accelerations between bone nodes;

[0037] S32, ST-GCN network model design: The ST-GCN network model includes a graph convolutional layer and a temporal convolutional layer. After the ST-GCN backbone network, n independent fully connected classification heads are added, corresponding to the binary classification tasks of n indicators of the evaluated actions respectively. The weighted multi-task loss function is used to solve the problem of importance differences between indicators;

[0038] S33, Model training: Use the cross-entropy loss function to handle the classification task. At the same time, the classification losses are weighted and combined into a comprehensive loss function to achieve balance. In terms of the optimization strategy, the AdamW optimizer is adopted, combined with the learning rate warm-up and cosine annealing strategy to optimize the training effect;

[0039] S34, Evaluation results and feedback: Output the binary classification results of multiple indicators of human actions. <s

[0040] Optionally, the ST-GCN network model design in S32 includes:

[0041] S321, Graph structure definition: Model the human skeleton as a graph G=(V, E), where the nodes V include key joint points such as the spine center, hips, and knees, and the edges E represent the natural connections between nodes (such as hips-knees). To solve the difficulty of feature extraction caused by the sparsity of millimeter-wave radar point clouds, an adaptive edge weight mechanism is proposed, and the edge weights are dynamically adjusted according to the Euclidean distance between adjacent nodes. The formula is as follows:

[0042] ω ij = exp(-d ij / σ)

[0043] where d ij is the node spacing, and σ = 0.2 is the scale parameter, making the network pay more attention to the action-sensitive areas;

[0044] S322, Multi-branch spatio-temporal feature extraction: Adopt a partitioning strategy (P = 3) to group the neighbors of each node, aggregate the features of nodes at different distances through learnable weights, capture the joint space collaboration relationship, and use one-dimensional dilated convolution (kernel size 3, dilation rate 2) to slide along the time axis to expand the temporal receptive field and enhance the modeling ability for action continuity;

[0045] S323, Action evaluation metric output: Design multiple independent fully-connected branches (FC_256 → FC_1), corresponding to multiple evaluation tasks respectively, and through experimentally verified unbalanced loss weight allocation, alleviate the model bias problem caused by differences in metric importance. Optionally, the cross-entropy loss function is expressed as:

[0046]

[0047] The mean squared error loss function is expressed as:

[0048]

[0049] The comprehensive loss function is expressed as:

[0050]

[0051] A millimeter-wave radar action evaluation system based on a human body bone model, used to implement the above-mentioned millimeter-wave radar action evaluation method based on a human body bone model, which is characterized by including the following modules:

[0052] Millimeter-wave radar point cloud sampling module: Collect three-dimensional point cloud data of the human body through a millimeter-wave radar device, and adopt an improved DBSCAN clustering algorithm to dynamically adjust the neighborhood radius of the core point based on the hierarchical information of the point cloud and combined with the azimuth angle information, and extract the target points related to the human body bone nodes;

[0053] Bone node estimation module: Based on the extracted target point data, predict the three-dimensional coordinates of each joint of the human body by adjusting the output module of the PointNet neural network model, so as to reconstruct and fit the human body bone structure;

[0054] Human action evaluation module: Construct the three-dimensional coordinates of multiple frames of human body bone nodes into time series data, and use an evaluation system based on the ST-GCN network to analyze the motion trajectory, so as to realize the evaluation and feedback of human actions.

[0055] A millimeter-wave radar action evaluation device based on a human body bone model, including a memory, a processor, and a computer program stored in the memory and running on the processor, which is characterized in that when the processor executes the computer program, it implements the above-mentioned millimeter-wave radar action evaluation method based on a human body bone model.

[0056] Advantages of the present invention:

[0057] In the present invention, non-contact motion monitoring is carried out by using a millimeter-wave radar, which avoids the restraint of the human body by traditional wearable devices, greatly improves the comfort and operation convenience of users. At the same time, the high penetration and anti-interference ability of the millimeter-wave radar enable it to work stably in complex environments and different lighting conditions, with stronger adaptability, and can accurately monitor human motion in a variety of environments, providing more reliable motion assessment results.

[0058] In the present invention, by improving the DBSCAN clustering algorithm and introducing a weighted distance calculation method, the limitations of traditional clustering algorithms in processing non-uniform density point cloud data are effectively overcome, and the accuracy and robustness of bone node extraction are improved. Combined with dynamic behavior analysis based on the ST-GCN network, it can accurately capture the complex dynamic changes and behavior patterns in human motion, providing high-precision real-time assessment and feedback for fields such as motion monitoring, rehabilitation training, and health assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0060] Figure 1 It is a schematic flow chart of the evaluation method according to an embodiment of the present invention;

[0061] Figure 2 It is a schematic diagram of the system function module according to an embodiment of the present invention;

[0062] Figure 3 It is a structural diagram of the improved PointNet network according to an embodiment of the present invention;

[0063] Figure 4 It is a structural diagram of the improved ST-GCN network according to an embodiment of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the accompanying drawings are only for more specifically describing the embodiments, and are not intended to specifically limit the present invention.

[0065] It should be noted that in the specification, the mention of "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc. indicates that the described embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge scope of those skilled in the relevant art.

[0066] Generally, terms can be understood at least in part from their use in context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather can alternatively, at least in part depending on the context, allow for the existence of other factors that are not necessarily explicitly described.

[0067] As Figure 1 shown, a millimeter-wave radar action evaluation method based on a human body bone model includes the following steps:

[0068] I. Millimeter-wave radar point cloud sampling:

[0069] In order to extract human key bone nodes from the three-dimensional point cloud data collected by the millimeter-wave radar, an improved DBSCAN clustering algorithm, called the "millimeter-wave radar point cloud sampling algorithm", is proposed. This algorithm combines the hierarchical characteristics of the millimeter-wave radar point cloud and the resolution characteristics of the millimeter-wave radar, and realizes the sampling of target points related to human bone nodes in the point cloud by optimizing the distance calculation method between target points. The algorithm steps are as follows:

[0070] 1. Two-dimensional projection of the millimeter-wave radar point cloud:

[0071] First, project the millimeter-wave radar point cloud along the front view of the radar. After projection, the coordinates of each target point are converted from three-dimensional to two-dimensional plane coordinates (x, z), while retaining the point cloud hierarchical information (y), spherical serial number (Kn), and horizontal angle (θ h ), pitch angle (θ q ).

[0072] The information of the point is defined as a six-tuple (x, z, y, Kn, θ h , θ q ), where:

[0073] x, z: The position of the target point in the two-dimensional plane.

[0074] y: The distance from the target point to the radar.

[0075] Kn: The spherical number where the target point is located.

[0076] θ h : The horizontal angle of the target point.

[0077] θ q : The pitch angle of the target point.

[0078] 2. Improved target point distance calculation:

[0079] The weighted factor is introduced to optimize the calculation method of the distance between target points, and the target point A(x1, z1, y1, K n1 ,θ h1 ,θ q1 ) and B(x2,z2,y2,K n2 ,θ h2 ,θ q2 ) is:

[0080]

[0081] Δz=α z (z2-z1);

[0082]

[0083] Among them, α z is the weighting factor along the z and x directions, adjusting the weights in different directions, α y is the weighting factor along the y direction, balancing the effect of the sphere spacing, and d is the fixed distance between adjacent spheres.

[0084] 3. Clustering process:

[0085] After completing the improved calculation of the distance between target points, the improved DBSCAN algorithm is used to perform cluster analysis on the millimeter wave radar point cloud data. The main steps are as follows:

[0086] (1) Parameter settings:

[0087] Minimum number of samples (MinPts): According to the density characteristics of the point cloud, set an appropriate minimum number of samples to ensure that there are enough neighborhood points for judging local density.

[0088] Radius parameter (Eps): For each target point P, the neighborhood radius is dynamically modified by introducing the angle-radius coupling model. The neighborhood radius can be expressed by the following formula:

[0089] ε=ε0·(1+k·(|θ h |+|θ q |))

[0090] where k is the angle sensitivity coefficient, which controls the influence weight of the angle on the neighborhood radius. When the incident angle of the radar beam increases, the projection density of the point cloud in the local area decreases, and the neighborhood radius needs to be expanded to ensure the robustness of clustering.

[0091] (2) Point sequence sorting and clustering formation:

[0092] Starting from the core points, recursively merge all density-reachable points.

[0093] Give priority to accessing the point cloud area with a lower spherical serial number and gradually expand the clustering cluster.

[0094] Isolated points with lower density during the clustering process are automatically marked as noise.

[0095] (3) Clustering result analysis:

[0096] Several clustering clusters are formed in the high-density area, and the heights of these clusters may correspond to the key parts of the human body (such as shoulders, elbows, knees, etc.).

[0097] Sparse areas or discrete points are regarded as background noise or irrelevant points, effectively removing interference.

[0098] II. Skeletal node estimation:

[0099] After clustering, target points related to human skeletal nodes are selected from the point cloud. In order to accurately fit these target points into the standard human skeletal model, an improved PointNet neural network model is used for training and prediction.

[0100] (1) Data preprocessing:

[0101] Input data: The point set obtained after clustering, containing the three-dimensional coordinates (x, y, z) of key target points.

[0102] Data standardization: Normalize the point cloud data to avoid biases in network training caused by data in different dimensions.

[0103] Data augmentation: Expand the training samples through methods such as rotation, translation, and noise perturbation to enhance the robustness of the model.

[0104] (2) Network structure design:

[0105] Input layer: Receive the clustered point cloud data (batch size × number of points × 3D coordinates).

[0106] Spatial transformation layer: A learnable parameterized transformation matrix (256×9) implicitly learns the spatial mapping relationship between the millimeter-wave radar point cloud and the human skeleton nodes. A one-dimensional convolution with shared weights (Conv1D, kernel_size=1) is used to extract high-dimensional features point by point, generating an n×1024 intermediate feature tensor. This is compressed into a 1×1024 global feature vector through maximum pooling, and then reduced to 256 dimensions through a fully connected layer. The matrix is then multiplied with the transformation matrix to reconstruct a 3×3 rotation matrix.

[0107] Feature Extraction Layer: This layer progressively extracts high-dimensional spatial features from the normalized point cloud. The input point cloud is first trained using a 1×3 convolution kernel (64 channels) to learn the local spatial associations of individual points, generating an n×64 primary feature tensor. The spatial transformation network further learns a 64×64 rotation matrix to orthogonalize the feature space and enhance geometric consistency. This layer is then gradually expanded to 128 and 1024 channels via two levels of 1×1 convolution, deeply mining the skeletal topology implicit in the point cloud. The final output is an n×1024 high-dimensional feature tensor, which supports the regression of the 3D coordinates of human skeletal nodes.

[0108] Max Pooling Layer: Maps local features to a global representation space. The input is an n×1024-dimensional point-level feature matrix. Maximum sampling is performed channel by channel along the point set dimension to generate a 1×1024-dimensional global feature vector that is independent of the point order. Only the most significant response value of each feature channel is retained.

[0109] The regression output layer takes a 1024-dimensional global feature vector as input, which is then reduced in dimension by a 512-neuron fully connected layer, introducing nonlinear representation using the ReLU activation function. During training, a 30% dropout layer randomly blocks some neuron connections to mitigate overfitting and enhance generalization to dynamic noise in radar point clouds. A normalization layer applies zero-mean normalization to the 512-dimensional features, mitigating gradient fluctuations and accelerating training convergence. Finally, a 33-dimensional fully connected layer linearly maps the features to output 3D spatial coordinate predictions for 11 human skeletal nodes, which are strictly aligned with Kinect data. This design balances model capacity and computational efficiency through feature compression (1024 → 512 → 33). The fully connected weight matrix can be interpreted as a linear decoder from global features to spatial coordinates, directly establishing a mathematical connection between point cloud distribution patterns and human skeletal geometry.

[0110] (3) Training and optimization:

[0111] Loss function: The mean square error (MSE) is used as the loss function to minimize the error between the predicted coordinates and the true coordinates:

[0112]

[0113] Optimization Algorithm: The Adam optimizer is adopted to adaptively adjust the learning rate, improving the training speed and accuracy.

[0114] Early Stopping Mechanism: Monitor the change of loss on the validation set to avoid overfitting.

[0115] (4) Prediction and Fitting Results:

[0116] After the network model is trained, the real-time collected radar point cloud data is input into the model.

[0117] The model outputs the three-dimensional coordinates of 11 key skeletal nodes, realizing the fitting and reconstruction of the human skeletal structure.

[0118] Visually display the superimposed effect of the skeletal nodes and the millimeter-wave radar point cloud to verify the prediction accuracy.

[0119] III. Human Action Evaluation:

[0120] In order to comprehensively evaluate the motion state and behavioral characteristics of the human body, the present invention constructs the three-dimensional coordinate data of multiple frames of human key skeletal nodes obtained into time series data, and uses an evaluation system based on the ST-GCN network to deeply analyze the motion trajectory, realizing the evaluation and feedback of human action behavior.

[0121] (1) Data Construction and Preprocessing:

[0122] Integration of Multiple Frames of Data: Continuously collect the three-dimensional coordinate data of multiple frames of skeletal nodes, set a time window, and arrange each frame of data in chronological order to form a complete motion time series:

[0123] S = {P1, P2,..., P T};

[0124] where Pt represents the three-dimensional coordinates of 11 skeletal nodes of the human body in the t-th frame, and T is the total number of frames within the time window.

[0125] Extraction of Motion Features: Calculate dynamic features such as the relative position, speed, and acceleration between skeletal nodes to enhance the description ability of time series data for motion trends.

[0126] (2) ST-GCN Network Model Design:

[0127] ST-GCN includes a graph convolutional layer and a temporal convolutional layer. After the ST-GCN backbone network, n independent fully-connected classification heads are added, corresponding to the binary classification tasks of n metrics of the evaluated actions respectively. By using a weighted multi-task loss function to solve the problem of importance differences between metrics, it can efficiently capture the long-term dynamic dependencies and context information in time series data. The human motion assessment model is designed based on the ST-GCN network, and the specific structure is as follows:

[0128] Input embedding layer: Map the normalized bone node time series data to a high-dimensional feature space:

[0129] X = Embed(S);

[0130] Input dimension: (T, N, 3), where N is the number of bone nodes.

[0131] Output dimension: (T, d model )), where d model is the model feature dimension.

[0132] Graph structure definition: Model the human skeleton as a graph G = (V, E), where the nodes V include key joint points such as the spine center, hip, and knee, and the edges E represent the natural connections between nodes (such as hip - knee). To solve the difficulty of feature extraction caused by the sparsity of millimeter-wave radar point clouds, an adaptive edge weight mechanism is proposed to dynamically adjust the edge weights according to the Euclidean distance between adjacent nodes:

[0133] ω ij = exp(-d ij / σ)

[0134] where d ij is the node spacing, and σ = 0.2 is the scale parameter to make the network pay more attention to the action-sensitive areas;

[0135] Multi-branch spatio-temporal feature extraction: Adopt a partitioning strategy (P = 3) to group the neighbors of each node, aggregate the features of nodes at different distances through learnable weights to capture the joint space cooperation relationship, and use one-dimensional dilated convolution (kernel size 3, dilation rate 2) to slide along the time axis to expand the temporal receptive field and enhance the modeling ability of action continuity;

[0136] Action evaluation metric output: Design multiple independent fully-connected branches (FC_256 → FC_1), corresponding to multiple evaluation tasks respectively, and use the experimentally verified unbalanced loss weight assignment to alleviate the model bias problem caused by the importance differences of metrics.

[0137] (3) Training and optimization:

[0138] Loss function design:

[0139] For classification tasks, use cross-entropy loss:

[0140]

[0141] The regression task uses the mean squared error:

[0142]

[0143] Comprehensive loss:

[0144]

[0145] Optimization strategy: The AdamW optimizer is adopted, combined with the learning rate warm-up (Warmup) and cosine annealing (CosineAnnealing) strategies to improve the training effect.

[0146] (4) Evaluation results and feedback:

[0147] Action quality score: Output the action evaluation results of the corresponding indicators according to the number of action evaluation indicators.

[0148] Such as Figure 2 As shown, a millimeter-wave radar action evaluation system based on a human body bone model is used to implement the above-mentioned millimeter-wave radar action evaluation method based on a human body bone model, which is characterized by including the following modules:

[0149] Millimeter-wave point cloud sampling module: Collect three-dimensional point cloud data of the human body through a millimeter-wave radar device, and use an improved DBSCAN clustering algorithm to dynamically adjust the neighborhood radius of the core points based on the hierarchical information of the point cloud and combined with the azimuth angle information, and extract the target points related to the human body bone nodes;

[0150] Bone node estimation module: Based on the extracted target point data, predict the three-dimensional coordinates of each joint of the human body by adjusting the output module of the PointNet neural network model, so as to reconstruct and fit the human body bone structure;

[0151] Human action evaluation module: Construct the three-dimensional coordinates of multiple frames of human body bone nodes into time series data, and use an evaluation system based on the ST-GCN network to analyze the motion trajectory, so as to realize the evaluation and feedback of human motion behavior.

[0152] A millimeter-wave radar action evaluation device based on a human body bone model includes a memory, a processor, and a computer program stored in the memory and running on the processor, and is characterized in that when the processor executes the computer program, it implements the above-mentioned millimeter-wave radar action evaluation method based on a human body bone model.

[0153] The present invention encompasses any alternatives, modifications, equivalent methods, and solutions made to the essence and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0154] The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A millimeter-wave radar motion evaluation method based on a human body bone model, characterized in that, It includes the following steps: S1, millimeter-wave radar point cloud sampling: Use a millimeter-wave radar device to collect three-dimensional point cloud data of the human body, and adopt an improved DBSCAN clustering algorithm to dynamically adjust the neighborhood radius of the DBSCAN algorithm based on the hierarchical information of the point cloud and combined with the azimuth angle information, so as to extract the target points related to the human body bone nodes with a clustering thinking; S2, bone node estimation: Based on the target points related to the human body bone nodes extracted, predict the three-dimensional coordinates of each joint of the human body by adjusting the PointNet neural network model, so as to reconstruct and fit the human bone structure; S3, human motion evaluation: Construct the three-dimensional coordinate data of multiple frames of human body bone nodes obtained into time series data, and use the evaluation system of the ST-GCN network to analyze the motion trajectory to realize the evaluation and feedback of human motion behavior.

2. The millimeter-wave radar motion evaluation method based on a human bone model according to claim 1, wherein, The millimeter-wave radar point cloud sampling in S1 includes: S11, 2D Projection of Millimeter-Wave Radar Point Cloud: Project the point cloud along the front view of the radar. After projection, the coordinates of each target point are converted from three-dimensional to two-dimensional plane coordinates, while preserving the hierarchical information, spherical serial number, and horizontal angle. The information of the target point is defined as a six-tuple (x, z, y, Kn, θ h , θ q ), where x and z are the positions of the target point in the two-dimensional plane, y is the distance from the target point to the radar, Kn is the spherical serial number where the target point is located, θ h is the horizontal angle of the target point, and θ q is the pitch angle of the target point; S12, Improvement in target point distance calculation: Introduce a weighting factor to optimize the calculation method of the distance between target points. Define the distance between target points A(x1, z1, y1, K n1 , θ h1 , θ q1 ) and B(x2, z2, y2, K n2 , θ h2 , θ q2 ) as: Δz = α z ·(z2 - z1); Among them, α z is the weighting factor along the z and x directions, which adjusts the weights in different directions. α y is the weighting factor along the y direction to balance the influence of the spherical surface spacing, and d is the fixed distance between adjacent spherical surfaces; S13, clustering process: After completing the improvement of the target point distance calculation, use the improved DBSCAN algorithm to perform clustering analysis on the millimeter-wave radar point cloud data, and form a clustering cluster of points by calculating the neighborhood radius ε of each point.

3. The millimeter-wave radar motion evaluation method based on a human body bone model according to claim 2, wherein The clustering process in S13 includes: S131, parameter setting: Set the minimum number of samples MinPts according to the density characteristics of the point cloud; S132, neighborhood radius: For each target point P, by introducing an angle-radius coupling model, dynamically correct the neighborhood radius, and the neighborhood radius ε can be expressed by the following formula: ε = ε0·(1 + k·(|θ h | + |θ q |)) where k is the angle sensitivity coefficient, which controls the influence weight of the angle on the neighborhood radius. When the incident angle of the radar beam increases, the projection density of the point cloud in the local area decreases, and the neighborhood radius needs to be expanded to ensure the clustering robustness. S133, point sequence sorting and clustering formation: Starting from the core point, recursively merge all density-reachable points to form a clustering cluster Ck; S134, clustering result analysis: The points that do not belong to any cluster are marked as noise points and removed from the output. Calculate the distances from the sample points in all clusters to the hierarchical center and sort them. The priority decreases gradually from far to near.

4. The millimeter-wave radar motion evaluation method based on a human body bone model according to claim 3, wherein The bone node estimation in S2 includes: S21, data preprocessing: The three-dimensional coordinates of the target points obtained after clustering are expressed as (x, y, z). Normalize the target points and perform data augmentation through rotation, translation, and noise perturbation; S22, network structure design: Predict the three-dimensional coordinates of each joint of the human body by adjusting the output structure of the PointNet neural network model, specifically including: Input layer: Receive the target points after clustering; Spatial transformation layer: The learnable parameterized transformation matrix (256×9) implicitly learns the spatial mapping relationship between the millimeter-wave radar point cloud and the human body bone nodes, uses one-dimensional convolution (Conv1D, kernel_size = 1) with shared weights to extract high-dimensional features point by point, generates an intermediate feature tensor of n×1024, compresses it to a global feature vector of 1×1024 through max pooling, and then reduces the dimension to 256 through a fully connected layer and multiplies it with the transformation matrix to reconstruct a 3×3 rotation matrix; Feature extraction layer: This layer progressively extracts high-dimensional spatial features from the normalized point cloud. The input point cloud is first trained using a 1×3 convolution kernel (64 channels) to learn the local spatial associations of individual points, generating an n×64 primary feature tensor. The spatial transformation network further learns a 64×64 rotation matrix to orthogonalize the feature space and enhance geometric consistency. This layer is then gradually expanded to 128 and 1024 channels via two levels of 1×1 convolution, deeply mining the skeletal topology information implicit in the point cloud. The final output is an n×1024 high-dimensional feature tensor, which supports the regression of the 3D coordinates of the human skeletal nodes. Max Pooling Layer: Maps local features to a global representation space. The input is an n×1024-dimensional point-level feature matrix. Maximum sampling is performed channel by channel along the point set dimension to generate a 1×1024-dimensional global feature vector that is independent of the point order. Only the most significant response value of each feature channel is retained. The regression output layer takes a 1024-dimensional global feature vector as input, which is then reduced in dimension by a 512-neuron fully connected layer, introducing nonlinear representation using the ReLU activation function. During training, a 30% dropout layer randomly blocks some neuron connections to mitigate overfitting and enhance generalization to dynamic noise in radar point clouds. A normalization layer applies zero-mean normalization to the 512-dimensional features, mitigating gradient fluctuations and accelerating training convergence. Finally, a 33-dimensional fully connected layer linearly maps the features to output 3D spatial coordinate predictions for 11 human skeletal nodes, which are strictly aligned with Kinect data. This design balances model capacity and computational efficiency through feature compression (1024 → 512 → 33). The fully connected weight matrix can be interpreted as a linear decoder from global features to spatial coordinates, directly establishing a mathematical connection between point cloud distribution patterns and human skeletal geometry. S23, training and optimization: using mean squared error as the loss function to minimize the error between the predicted coordinates and the true coordinates, using the Adam optimizer, adaptively adjusting the learning rate, and monitoring the loss change on the validation set; S24, prediction and fitting results: After the PointNet neural network model training is completed, the millimeter-wave radar point cloud data collected in real time is input into the model, and the three-dimensional coordinates of 11 key bone nodes are output to achieve the fitting and reconstruction of the human skeletal structure. The superposition effect of the bone nodes and point cloud data is visualized to verify the prediction accuracy.

5. The millimeter-wave radar motion evaluation method based on a human body bone model according to claim 4, wherein The human motion assessment in S3 includes: S31, data construction and preprocessing: continuously collect the three-dimensional coordinates of multiple frames of skeletal nodes, set the time window, arrange each frame of data in chronological order to form a complete motion time series, and calculate the relative position, velocity, and acceleration between skeletal nodes; S32, ST-GCN network model design: The ST-GCN network model includes graph convolution layers and temporal convolution layers. After the ST-GCN backbone network, n independent fully connected classification heads are added, corresponding to the binary classification tasks of n indicators of the evaluated action, and the problem of importance differences between indicators is solved through the weighted multi-task loss function; S33, Model Training: The cross-entropy loss function is used to handle classification tasks. At the same time, the classification losses are weighted and combined into a comprehensive loss function to achieve balance. In terms of the optimization strategy, the AdamW optimizer is adopted, combined with the learning rate warm-up and cosine annealing strategies to optimize the training effect; S34, Evaluation Results and Feedback: Output the binary classification results of multiple human motion indicators.

6. The millimeter-wave radar motion evaluation method based on a human body bone model according to claim 5, characterized in that, The ST-GCN network model design in S32 includes: S321, Graph Structure Definition: The human skeleton is modeled as a graph G=(V, E), where the nodes V include key joint points such as the spine center, hips, and knees, and the edges E represent the natural connections between nodes (such as hips-knees). To solve the difficulty of feature extraction caused by the sparsity of millimeter-wave radar point clouds, an adaptive edge weight mechanism is proposed, and the edge weights are dynamically adjusted according to the Euclidean distance between adjacent nodes. The formula is as follows: ω ij = exp(-d ij / σ) where d ij is the node spacing, and σ = 0.2 is the scale parameter to make the network pay more attention to the action-sensitive region; S322, Multi-branch Spatiotemporal Feature Extraction: The partition strategy (P = 3) is used to group the neighbors of each node, and the features of nodes at different distances are aggregated through learnable weights to capture the joint space cooperation relationship. One-dimensional dilated convolution (kernel size 3, dilation rate 2) is used to slide along the time axis to expand the temporal receptive field and enhance the ability to model the continuity of actions; S323, Action Evaluation Index Output: Multiple independent fully connected branches (FC_256→FC_1) are designed, corresponding to multiple evaluation tasks respectively. Through the experimentally verified unbalanced loss weight allocation, the problem of model bias caused by differences in index importance is alleviated.

7. The millimeter-wave radar motion evaluation method based on a human bone model according to claim 6, characterized in that, The cross-entropy loss function is expressed as: The mean square error loss function is expressed as: The comprehensive loss function is expressed as:

8. A millimeter-wave radar motion evaluation system based on a human body bone model, which is used to implement a millimeter-wave radar motion evaluation method based on a human body bone model according to any one of claims 1-7, characterized in that, It includes the following modules: Millimeter-wave Radar Point Cloud Sampling Module: The three-dimensional point cloud data of the human body is collected by a millimeter-wave radar device, and an improved DBSCAN clustering algorithm is adopted. The neighborhood radius of the core point is dynamically adjusted based on the hierarchical information of the point cloud and combined with the azimuth angle information to extract the target points related to the human skeleton nodes; Bone Node Estimation Module: Based on the extracted target point data, the three-dimensional coordinates of each joint of the human body are predicted by adjusting the output module of the PointNet neural network model, so as to reconstruct and fit the human bone structure; Human Action Evaluation Module: The three-dimensional coordinates of the multi-frame human bone nodes obtained are constructed into time series data, and the motion trajectory is analyzed by using an evaluation system based on the ST-GCN network, so as to realize the evaluation and feedback of human actions.

9. A millimeter-wave radar motion evaluation device based on a human body bone model, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements a millimeter-wave radar action evaluation method based on a human bone model as described in any one of claims 1-7.

Citation Information

Cited By

  • Action recognition method based on adaptive skeleton grouping and direction sensitive space-time modeling

    CN121392981A

  • Action recognition method based on adaptive bone grouping and direction-sensitive spatio-temporal modeling

    CN121392981B

  • Radar-based human skeleton estimation method, system and product

    CN121817844A

  • Virtual reality feedback processing method, device and system for cancer postoperative rehabilitation training

    CN122415815A