Dynamic millimeter wave radar point cloud human arm tracking system and method based on joint learning

By introducing joint learning, heuristic and clustering algorithms, spatiotemporal feature coding and multi-task joint learning in the millimeter wave radar point cloud system, the problems of arm motion complexity and semantic information are solved, and fine-grained arm tracking and three-dimensional motion trajectory tracking for privacy protection are realized.

CN120088853APending Publication Date: 2025-06-03SOUTHEAST UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510096037.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing millimeter wave perception system is difficult to achieve fine-grained arm tracking, mainly due to the complexity and dynamic nature of arm movements, as well as the lack of semantic information in the mmWave point cloud.

Method used

Using a dynamic millimeter-wave radar point cloud system based on joint learning, the dynamic millimeter-wave point cloud of arm motion is decoupled through heuristics and clustering algorithms, combining spatiotemporal feature coding and multi-task joint learning, point cloud features are extracted and three-dimensional motion trajectory tracking of arm joints is realized.

Benefits of technology

It realizes the robust tracking of human arm movements with privacy protection, overcomes the complexity and dynamicity of arm movements and the lack of semantic information, and provides a contactless, low-cost three-dimensional arm tracking method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088853A_ABST
    Figure CN120088853A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic millimeter-wave radar point cloud human arm tracking system and method based on joint learning, and the method comprises the steps: obtaining a millimeter-wave point cloud generated by human motion through millimeter-wave radar equipment, and enabling the generated point cloud to serve as the input of an arm point cloud generation module based on heuristic and clustering; decoupling a dynamic millimeter wave point cloud generated by arm movement; taking the dynamic arm point cloud as the input of a spatial-temporal feature encoder, and carrying out feature extraction; and point cloud features are output as input of a multi-task joint learning module, gesture recognition is introduced as an auxiliary task, learning assistance is provided for an arm tracking task, and finally three-dimensional motion trajectory tracking of main joints of the arm is achieved. According to the method, the technical challenges of arm semantic information loss and dynamic change in millimeter wave point cloud are overcome, and the human arm motion robust tracking with privacy protection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless sensing, and mainly relates to a dynamic millimeter-wave radar point cloud human arm tracking system and method based on joint learning. Background Art

[0002] Tracking and understanding human arm movements play an increasingly important role in human-centered intelligent applications. For example, in medical monitoring scenarios, tracking the arm movements of patients is crucial for estimating limb stiffness, movement speed, and postural instability. The arm tracking function also plays an important auxiliary role in fitness. For example, in training such as weightlifting and golf, users can immediately correct their postures based on accurate arm tracking feedback. Arm movements are also integrated into smart home and augmented reality / virtual reality applications to enhance the user experience.

[0003] Currently, in order to track human arm movements, a widely studied direction is to attach motion sensors to the human body to track arm movements, such as IMU sensors. However, this method of attaching multiple sensors to the user's body usually results in a relatively significant intrusive user experience, especially in some medical monitoring and rehabilitation scenarios where patients may have difficulty moving even without attachments. Therefore, vision-based methods are now often used to achieve non-invasive sensing, but these methods can cause privacy issues and are vulnerable to lighting conditions. Especially under lighting conditions, excessive exposure will reduce the arm recognition and tracking accuracy. In the aspect of human non-contact sensing, in recent years, there has been a research boom in using millimeter-wave radar, and its effectiveness in action recognition, pose estimation, etc. has been gradually demonstrated. By placing a millimeter-wave radar system in front of the user to transmit and receive reflected radio frequency spectrum signals and converting them into millimeter-wave point cloud data, continuous capture of the user's arm movements can be achieved. Thanks to the radio frequency sensing characteristics, this non-visual sensing method means that the system will not record any facial information or other sensitive personal data, thus protecting user privacy. However, existing millimeter-wave sensing systems mainly focus on the construction of the user's whole body skeleton and have not modeled the unique movement patterns of the human arm, ultimately resulting in poor performance in the arm tracking task. Therefore, a fine-grained arm tracking scheme is needed.

[0004] Despite the recent success of millimeter-wave radar in various human sensing tasks, achieving fine-grained arm tracking faces the following two unique technical challenges.

[0005] Complexity and Dynamics of Arm Movements: As one of the most expressive parts of the human body, the arm has the greatest degree of joint freedom, resulting in highly dynamic changes in the arm movement manifested in the point cloud of consecutive frames. In addition, the situation of self-occlusion also increases the complexity of the arm tracking process. For example, when the upper body rotates to a specific angle, the torso will occlude the arm, leading to the loss of arm point cloud data.

[0006] Lack of Semantic Information in Millimeter-Wave Point Cloud: Due to the requirement of low cost, the millimeter-wave radar adopted usually only has one chip and a limited number of antenna components (only 3 transmitters and 4 receivers). Combining with the inherent specular reflection characteristics of millimeter-wave signals, it ultimately results in a very sparse output point cloud, making it difficult to understand the corresponding gesture information even with the naked eye, that is, the lack of semantic information in the point cloud data, which further makes it difficult for the depth model to extract features containing arm structure information. Summary of the Invention

[0007] The present invention precisely aims at the problem that it is difficult to achieve fine-grained arm tracking in the prior art, and provides a dynamic millimeter-wave radar point cloud human arm tracking system and method based on joint learning. First, through the millimeter-wave radar device, the millimeter-wave point cloud generated by human movement is obtained, and the generated point cloud is used as the input of the arm point cloud generation module based on heuristics and clustering to decouple the dynamic millimeter-wave point cloud generated by arm movement; then the dynamic arm point cloud is used as the input of the spatio-temporal feature encoder for feature extraction; the output point cloud features are used as the input of the multi-task joint learning module, and gesture recognition is introduced as an auxiliary task to provide learning assistance for the arm tracking task, and finally the three-dimensional motion trajectory tracking of the main joints of the arm is realized. This method overcomes the technical challenges of the lack of arm semantic information and dynamic changes in the millimeter-wave point cloud, and realizes robust tracking of human arm movements with privacy protection.

[0008] To achieve the above object, the technical solution adopted by the present invention is: A dynamic millimeter-wave radar point cloud human arm tracking system based on joint learning, at least including an arm point cloud generation module based on heuristics and clustering, a point cloud feature extraction module based on spatio-temporal feature encoding, and an accurate arm tracking module based on multi-task joint learning;

[0009] The arm point cloud generation module based on heuristics and clustering: Data is collected through the millimeter-wave radar device, and the millimeter-wave point cloud is obtained as the input. Through a heuristic algorithm based on spatial position and a point cloud clustering algorithm based on density, the dynamic millimeter-wave point cloud generated by arm movement is decoupled;

[0010] The point cloud feature extraction module based on spatio-temporal feature encoding: By improving the PointNet network, Transformer network, and bidirectional long short-term memory network Bi-LSTM to construct a spatio-temporal feature encoder, the dynamic millimeter-wave point cloud of the arm is input and mapped into a high-dimensional spatio-temporal feature representation for point cloud feature extraction;

[0011] The precise arm tracking module based on multi-task joint learning: At least includes a gesture recognition branch module and an arm tracking branch module. The high-dimensional spatio-temporal feature representation of the point cloud is used as the input of the multi-task joint learning module. Through the gesture recognition branch module, it guides the arm tracking branch module to output the motion trajectory of the arm joints, realizing the tracking of arm movement.

[0012] As an improvement of the present invention, in the point cloud feature extraction module based on spatio-temporal feature encoding, the PointNet network is constructed by two layers of multi-layer perceptrons MLP with shared weights. The point features pass through the first layer of multi-layer perceptron MLP, and a high-dimensional point feature set is output. A residual link is performed on the arm point cloud set and the high-dimensional point feature set in the feature dimension of each point, and the obtained feature dimension is used as the input of the second layer of multi-layer perceptron MLP, and the final high-dimensional point feature set representation is output;

[0013] The internal multi-head attention mechanism of the Transformer network works in parallel. Each attention head independently calculates the attention weights, multiplies the attention scores of each head with the value matrix, and obtains the weighted features after spatial feature encoding through the attention head. The output multi-head features are weighted through a linear layer to output a feature set.

[0014] As another improvement of the present invention, the gesture recognition branch module at least includes a multi-layer perceptron and a linear layer. After passing through the Softmax function, the global feature vector of each frame in the sequence is mapped to the probability value prediction of each gesture, and the maximum value is selected as the final sequence label;

[0015] The arm tracking branch module at least includes a multi-layer perceptron and a linear layer, and maps the global feature vector of each frame in the sequence to the three-dimensional coordinates of the arm joints.

[0016] To achieve the above object, the technical solution adopted by the present invention is also: A dynamic millimeter-wave radar point cloud human arm tracking method based on joint learning, including the following steps:

[0017] S1. Point cloud generation: The millimeter-wave radar device is used to collect data information on the gesture actions of the user, and continuous millimeter-wave point cloud frames are obtained. Taking the millimeter-wave point cloud frames as the input, they are sequentially passed through the spatial heuristic algorithm and the DBSCAN point cloud clustering algorithm. After filtering out environmental noise, the arm point cloud set is output;

[0018] S2. Point cloud feature extraction: A spatio-temporal feature encoder is constructed through an improved PointNet network, a Transformer network, and a bidirectional long short-term memory network Bi-LSTM, which maps the input of the dynamic arm point cloud into a high-dimensional spatio-temporal feature representation. The improved PointNet network is constructed by two layers of multi-layer perceptrons with shared weights. After the arm point cloud set is input into the improved PointNet network, a high-dimensional point feature set PF is output. t , which is input into the Transformer network to capture the long-range dependencies between dynamic point clouds:

[0019] g t = BiLSTM(f t , f t-1 , f t+1 ; θ b )

[0020] where θ b represents the learnable parameters of Bi-LSTM, and g t is the global feature vector containing context information, and f t is the frame feature;

[0021] S3. Arm tracking: The high-dimensional spatio-temporal feature representation of the point cloud is used as the input of the multi-task joint learning module. The gesture recognition task is introduced to assist in enhancing the neural network's understanding ability of different arm movement patterns, and to guide the arm tracking task to output an accurate arm movement trajectory. In the gesture recognition task, an MLP and a linear layer are used to map the global feature vector of each frame in the sequence to the probability value prediction of each gesture through the Softmax function, and the maximum value is selected as the final sequence label. In the arm tracking task, the global feature vector of each frame in the sequence is mapped to the three-dimensional coordinates of the arm joints through an MLP and a linear layer.

[0022] As an improvement of the present invention, the neighborhood radius of the DBSCAN clustering algorithm is 0.35 meters, the minimum number of points threshold is 3, and the largest sub-cluster is recognized as the millimeter-wave point cloud generated by human arm movement.

[0023] As another improvement of the present invention, in the Transformer network of step S2, the point feature set PF t at any time t is used as the input, and after any attention head head h,t , a query matrix Q h,t , a key matrix K h,t , and a value matrix V h,t are obtained. The attention score of each head is multiplied by the value matrix V h,t , and the weighted feature after spatial feature encoding is obtained through the attention head head h,t :

[0024]

[0025] Where d represents PF t The dimension of each point feature in , H represents the number of attention heads, h∈H;

[0026] The output multi-head features are weighted through the linear layer Linear, and the output feature set PT t It is expressed as:

[0027] PT t =Linear(PT 1,t ,PT 2,t ,…,PT h,t θ l )

[0028] where θ l represents the learnable parameters of the linear layer.

[0029] As another improvement of the present invention, in step S2, the high-dimensional point feature set PF t And the feature set PT after Transformer network encoding t Perform residual connection and output the final point feature set PD′ through the MLP layer t :

[0030] PD′ t =MLP(PF t ,PT t θ r )

[0031] where θ r represents the learnable parameters of MLP;

[0032] The feature set PD′ t The point features pd′ of all points in i,t Aggregated into frame features f through self-attention layer t , get the global information of the frame:

[0033]

[0034] where θ f represents the learnable parameters of the linear layer.

[0035] As another improvement of the present invention, in step S3, for a point cloud frame sequence of length S, a cross entropy loss is used to minimize the error between the predicted category and the true category of each sequence, and the loss function is:

[0036]

[0037] Where G is the number of gesture action classes, is a function with an output of 0 or 1. When is 1, it indicates that the sample s belongs to the class g, is the predicted probability that the sample s belongs to the class g;

[0038] For a point cloud frame sequence with a length of S and K predicted joints per frame, the mean squared error (MSE) loss is used to minimize the error between the prediction and the actual coordinates of the main arm joints. The loss function is:

[0039]

[0040] where ||·|| represents the L2 norm, and J k,t represent the predicted value and the actual value of the three-dimensional coordinates of the arm joints, respectively.

[0041] As a further improvement of the present invention, the following function is used in step S3 to perform end-to-end training on the gesture recognition task and the arm tracking task, calculate the gradient and perform backpropagation, and update the parameters until the network converges:

[0042] L = αL G + βL T

[0043] where α and β are hyperparameters of the model.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] (1) The method and system of the present invention can implement a three-dimensional human arm tracking method based on radio frequency, non-contact, and privacy-protected, which only uses low-cost millimeter-wave radar devices, does not depend on light, and ignores harsh conditions such as smoke and dust, making it easier to implement and avoiding privacy issues.

[0046] (2) In the process of point cloud feature extraction in step S2 of the method of the present invention, based on spatio-temporal feature encoding, it combines local and global perspectives and can effectively extract spatio-temporal features for the dynamic change challenges of sparse millimeter-wave point clouds.

[0047] (3) In the process of accurate human arm tracking based on the multi-task learning framework in step S3 of the method of the present invention, three-dimensional arm tracking (main task) and gesture recognition (auxiliary task) are jointly learned. The gesture recognition task is used to make up for the challenge of missing point cloud semantics and realizes smooth and accurate arm tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a schematic diagram of the overall structure of the system of the present invention;

[0049] Figure 2It is a schematic structural diagram of the point cloud feature extraction module based on spatio-temporal feature encoding and the precise arm tracking module based on multi-task joint learning in the system of the present invention;

[0050] Figure 3 It is a schematic structural diagram of the improved PointNet network structure in the system of the present invention;

[0051] Figure 4 It is a schematic structural diagram of the Transformer global feature extraction network in the system of the present invention;

[0052] Figure 5 It is a graph of the cumulative distribution of the arm joint positioning error probability in the test example of the present invention;

[0053] Figure 6 It is a graph of the cumulative distribution of the arm joint angle error probability in the test example of the present invention. Detailed implementation manners

[0054] The following further clarifies the present invention in conjunction with the accompanying drawings and specific implementation manners. It should be understood that the following specific implementation manners are only used to illustrate the present invention and not to limit the scope of the present invention.

[0055] Example 1

[0056] The dynamic millimeter-wave radar point cloud human arm tracking system based on joint learning at least includes an arm point cloud generation module based on heuristics and clustering, a point cloud feature extraction module based on spatio-temporal feature encoding, and a precise arm tracking module based on multi-task joint learning.

[0057] Among them, the arm point cloud generation module based on heuristics and clustering is as shown in Figure 1 the arm point cloud generation module. Taking the millimeter-wave point cloud obtained by the third perspective millimeter-wave radar device as the input, through the heuristic algorithm based on spatial position and the point cloud clustering algorithm based on density, the environmental noise generated by the multipath effect is filtered out, and the dynamic millimeter-wave point cloud generated by the arm movement is decoupled as the input of the subsequent neural network model;

[0058] The point cloud feature extraction module based on spatio-temporal feature encoding is as shown in Figure 2 the spatio-temporal feature encoding module. Utilizing the movement law and correlation of the dynamic arm point cloud in space and time, by designing the improved PointNet network as shown in Figure 3 , and the Transformer network as shown in Figure 4 , as well as the bidirectional long short-term memory network (Bi-LSTM) to construct a spatio-temporal feature encoder, mapping the input of the dynamic arm point cloud into a high-dimensional spatio-temporal feature representation, effectively solving the challenge of dynamic changes in sparse millimeter-wave point clouds;

[0059] The precise arm tracking module based on multi-task joint learning is as shown inFigure 2 As shown in the multi-task joint learning module, the high-dimensional spatio-temporal feature representation of the point cloud is used as the input of the multi-task joint learning module. The lack of semantic information in the arm point cloud can be compensated by using gesture recognition as an auxiliary task. By utilizing the consistency relationship between gesture recognition and arm tracking, the gesture recognition task is introduced to assist in enhancing the neural network's understanding ability of different arm movement patterns, guiding the arm tracking task to output the accurate main joint movement trajectory of the arm.

[0060] Embodiment 2

[0061] A dynamic millimeter-wave radar point cloud human arm tracking method based on joint learning, as Figure 1 shown, realizes stable and accurate arm trajectory tracking through point cloud spatio-temporal feature encoding and multi-task joint learning, specifically including the following steps:

[0062] Step S1, generation of arm point cloud: Taking the millimeter-wave point cloud obtained by the third-view millimeter-wave radar device as the input, through a heuristic algorithm based on spatial position and a density-based point cloud clustering algorithm, the environmental noise generated by the multipath effect is filtered out, and the dynamic millimeter-wave point cloud generated by the arm movement is decoupled as the input of the subsequent neural network model. The specific steps include:

[0063] S11: The user performs a gesture action. The third-view millimeter-wave radar emits an electromagnetic wave signal and receives the echo to generate an intermediate-frequency signal, which is processed by 3D-FFT to generate a continuous millimeter-wave point cloud frame;

[0064] S12: Taking the millimeter-wave point cloud frame as the input, using a spatial heuristic algorithm to determine the position area of the user's activity, and filtering out the far-point environmental noise generated by the multipath effect according to the spatial range;

[0065] S13: Taking the spatial heuristic algorithm as the input, using the DBSCAN clustering algorithm to filter out the near-point environmental noise, decoupling the millimeter-wave point cloud generated by the arm movement, and the arm point cloud set corresponding to the frame at time t is denoted as PD t 。

[0066] The neighborhood radius (∈) of the DBSCAN clustering algorithm is 0.35 meters, the minimum number of points (MinPts) threshold is 3, and the largest sub-cluster is recognized as the millimeter-wave point cloud generated by human arm movement.

[0067] Step S2, Point Cloud Feature Extraction: It includes a spatial feature encoding module for aggregating the arm point cloud and a temporal feature encoding module for cross-temporal tracking. The spatial feature encoding module consists of an improved Point Cloud Network (PointNet) and a Transformer network. The improved Point Cloud Network can effectively process each point cloud data, while the Transformer network can capture the long-range dependencies between each point. Therefore, this combination enables the model to deeply understand the local and global spatial relationships in the arm point cloud. The specific steps are as follows:

[0068] S21: Construct a PointNet network using a 2-layer shared-weight Multi-Layer Perceptron (MLP) as shown in Figure 3 to map any radar point pd t in the arm point cloud set PD i,t corresponding to any t moment to a higher-dimensional feature representation, i.e., point features.

[0069] After passing through the first-layer MLP with a size of (5, 27), the output is a high-dimensional point feature set PM t represented as:

[0070] PM t = MLP(PD t ; θ m1 )

[0071] where θ 1 represents the learnable parameters of the first-layer MLP.

[0072] S22: Perform a residual connection on PD t and PM t in the feature dimension of each point, and the resulting feature dimension is 32, which is used as the input to the second-layer MLP layer with a size of (32, 64).

[0073] The high-dimensional point feature set PF t output by the second-layer MLP is represented as:

[0074] PF t = MLP(PD t , PM t ; θ m2 )

[0075] where θ 2 represents the learnable parameters of the second-layer MLP, and PF t is the final high-dimensional point feature set representation output by PointNet, and the output feature dimension is 64.

[0076] S23: Use the high-dimensional point feature set PF t as shown in Figure 4The input of the Transformer global feature encoding module shown captures the long-range dependencies between dynamic point clouds.

[0077] Inside the Transformer, the multi-head attention mechanism works in parallel. The number of attention heads H is 4, and the output feature dimension is 64. Each attention head independently calculates the attention weights. Taking the point feature set PF at any time t t as the input, passing through any attention head head h,t , where h ∈ H, to obtain the query matrix Q h,t = PF t × W Q , the key matrix K h,t = PF t × W K , and the value matrix V h,t = PF t × W V , where W Q ∈ R d×1 , W K ∈ R d×1 , W V ∈ R d×d′ are learnable parameters, d represents the dimension of each point feature in PF t , which is 64, and d' represents the dimension of the value matrix, which is 64.

[0078] S24: Multiply the attention scores of each head by the value matrix to obtain the weighted features after spatial feature encoding through the attention head head h,t :

[0079]

[0080] S25: Weight the output multi-head features through a linear layer of size (128, 64), and the output feature set PT t is expressed as:

[0081] PT t = Linear(PT 1,t , PT 2,t , …, PT h,t ; θ l )

[0082] where θ l represents the learnable parameters of the linear layer, and the output feature dimension is 64.

[0083] S26: For the input point feature set PF t and the feature set PT after being encoded by the Transformer tPerform residual connection and output the final more expressive point feature set PD′ through an MLP layer of size (128, 64) t , which can be expressed as:

[0084] PD′ t =MLP(PF t ,PT t θ r )

[0085] where θ r represents the learnable parameters of MLP, and the output feature dimension is 64.

[0086] S27: The feature set PD′ t Point features PD′ of all points in i,t Aggregated into frame features f through a self-attention layer of size (64, 1) t , get the global information of the frame. The formula is:

[0087]

[0088] where θ f represents the learnable parameters of the linear layer, which computes the attention weights for each point feature, allowing the network to dynamically learn the relative contribution of each point.

[0089] S28: A Bi-LSTM network with a hidden layer size of 64 and a layer number of 3 is used to capture the long-term dependencies between consecutive frames. The formula is:

[0090] g t =BiLSTM(f t ,f t-1 ,f t+1 θ b )

[0091] where θ b represents the learnable parameters of Bi-LSTM, and the output is a global feature vector g containing context information t , the corresponding feature dimension is 128, which is used for subsequent multi-task joint learning.

[0092] Step S3: Human arm tracking: Figure 2 As shown in the figure, a multi-task learning framework is proposed to solve the problem of missing semantics in sparse millimeter-wave point clouds. By specifying three-dimensional arm trajectory tracking as the main task and gesture recognition as the auxiliary task, the lack of semantic information in the sparse point cloud data of the arm can be compensated by taking gesture recognition as an auxiliary task. By utilizing the consistent relationship between gesture recognition and arm tracking, the ability of the method of the present invention to track and understand various arm movement patterns can be enhanced.

[0093] Step S3 specifically includes the following steps:

[0094] S31: In the gesture recognition (auxiliary task) branch module, use an MLP of size (128, 64) and a linear layer of size (64, 12), and map the global vector g of each frame in sequence S to the probability value prediction of each gesture through the Softmax function, and select the maximum value as the final sequence label. t Map to the probability value prediction of each gesture, and select the maximum value as the final sequence label.

[0095] S32: For the point cloud frame sequence of length S, the following cross-entropy loss is used to minimize the error between the predicted category and the true category of each sequence. The loss function can be expressed as:

[0096]

[0097] where G is the number of gesture classes, is a function with an output of 0 or 1. When is 1, it indicates that sample s belongs to class g, is the predicted probability that sample s belongs to class g.

[0098] S33: In the arm tracking (main task) branch module, use an MLP of size (128, 64) and a linear layer of size (64, 6 * 3) to map the global vector g of each frame in sequence S t to the three-dimensional coordinates of the arm joints.

[0099] S34: For the point cloud frame sequence of length S with K predicted joints per frame, use the following mean squared error (MSE) loss to minimize the error between the predicted and actual main arm joint coordinates. The loss function can be expressed as:

[0100]

[0101] where ||·|| represents the L2 norm, and J k,t represent the predicted value and the actual value of the three-dimensional coordinates of the arm joints respectively.

[0102] In this embodiment, the number of predicted arm joints K = 6, including the three-dimensional coordinates of six joints: the left and right shoulders, elbows, and wrists.

[0103] S35: Finally, use the following function to train the main task and auxiliary task models end-to-end, calculate the gradient and perform backpropagation, and update the parameters until the network converges:

[0104] L = αL G + βL T

[0105] where α and β are hyperparameters of the model.

[0106] The present invention addresses two major challenges by introducing a new network design framework: the lack of semantic information and the dynamic change problem in point clouds. By implementing a spatio-temporal feature encoder to aggregate dynamic arm point clouds from a global perspective, and through joint learning of 3D arm tracking (main task) and gesture recognition (auxiliary task), using the auxiliary task to guide the network to deeply understand the movement of the arm, a new type of 3D arm trajectory tracking method is finally achieved, providing non-contact and privacy-protected precise sensing for applications such as medical rehabilitation and human-computer interaction.

[0107] Test case

[0108] To verify the effectiveness of the method of the present invention, a human arm movement tracking and estimation experiment was conducted in an indoor environment to demonstrate the performance advantages and impacts of the present solution.

[0109] The experiment was carried out with the following configurations:

[0110] Hardware settings:

[0111] 1. The millimeter-wave radar IWR6843ISK-ODS, used to collect millimeter-wave point cloud data;

[0112] 2. Azure Kinect v2, used to collect the true postures of the human arm and train the network;

[0113] Experimental design:

[0114] 1. An indoor corridor was selected as the test site to simulate a complex indoor environment. Billboards, seats, and metal cabinets were set on both sides to interfere with the millimeter-wave signal, generating multipath effect noise points and increasing the challenge of arm tracking.

[0115] 2. Twelve common gestures in daily life were designed for the experimental actions, including a static sitting posture (aiming to detect the effectiveness of the model in a static state), six arm-only actions such as raising the hand, waving the hand, and making an arm circle, and five actions with significant body displacement accompanied by arm actions such as clapping, nodding, and torso rotation.

[0116] 3. The subject sat about 1.1 meters in front of the device, and the IWR6843ISK-ods millimeter-wave radar and Kinect were placed on specific brackets to collect the millimeter-wave point cloud data of the subject and the true data of the arm joints respectively.

[0117] 4. In each experiment, the subjects performed continuously for 4 minutes and 30 seconds, generating approximately 2,700 frames of point cloud data for each action (radar frame rate is 10 frames per second). ROS melodic was used to collect data from these two different modalities simultaneously. These two modalities have different frame rates (Kinect frame rate is 15 frames per second) and are synchronized based on the ROS timestamp.

[0118] 5. 317,770 frames of point cloud data with joint labels were collected from 12 gesture actions performed by ten screened subjects. The continuous gesture actions were segmented into sequence samples of 20 frames each, and gesture labels (i.e., gestures 1 to 12) were added to each sequence sample for subsequent joint learning.

[0119] 6. The collected data was divided into a dataset in the ratio of 7:2:1, i.e., 70% for model training, 20% for model testing, and 10% for model validation.

[0120] 7. mmPose, Mars, mmMesh, P4Transformer, and mmPose-NLP were selected as comparison methods, where P4Transformer is the current best method.

[0121] The experimental results on the validation set are shown in Tables 1 and 2 below:

[0122] Table 1 Comparison of arm localization errors (cm)

[0123]

[0124] Table 2 Comparison of arm angle estimation errors (°)

[0125]

[0126] As can be seen from the above table, the method of the present invention achieves the best performance on all three arm joints. The average joint position error (angle error) is only 1.67 cm (4.23°), which is 67.6% and 64.5% higher than that of P4Transformer respectively. Different from traditional bone reconstruction methods, the present invention introduces spatial feature encoding and considers the temporal pattern of the gesture sequence, effectively improving the accuracy and coherence of tracking.

[0127] To further analyze the performance of this system on the validation set, we used each group of samples as the input of the model, and then recorded the mean square losses between the output of the model and the true values of the labels collected by Kinect in joint position prediction and joint angle prediction, denoted as position error and angle error. Figure 5 and Figure 6The cumulative distribution function (CDF) of the position error and angle error of the present system is shown. The results show that in 80% of the arm trajectory tracking tasks, the present system achieves position errors of less than 1.5 cm at the shoulder, less than 2 cm at the elbow, and less than 2.5 cm at the wrist, respectively. The angle estimation errors of the humerus and forearm are less than 4° and 6°, respectively, demonstrating the robustness of the present system.

[0128] In summary, the present solution proposes a dynamic millimeter-wave radar point cloud human arm tracking method based on joint learning, which can achieve a radio-frequency-based, non-contact, privacy-protected three-dimensional human arm tracking. The method of the present invention only uses low-cost radar devices, does not rely on light, and ignores harsh conditions such as smoke and dust, making it easier to implement and avoiding privacy issues.

[0129] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. Dynamic millimeter wave radar point cloud human arm tracking system based on joint learning, characterized by ,It at least includes an arm point cloud generation module based on heuristic and clustering, a point cloud feature extraction module based on spatiotemporal feature encoding, and an accurate arm tracking module based on multi-task joint learning; The arm point cloud generation module based on heuristics and clustering: collects data through millimeter wave radar equipment, obtains millimeter wave point cloud as input, and decouples the dynamic millimeter wave point cloud generated by arm movement through a heuristic algorithm based on spatial position and a point cloud clustering algorithm based on density; The point cloud feature extraction module based on spatiotemporal feature encoding: a spatiotemporal feature encoder is constructed through an improved PointNet network, a Transformer network, and a bidirectional long short-term memory network Bi-LSTM, and the dynamic millimeter wave point cloud input of the arm is mapped into a high-dimensional spatiotemporal feature representation to extract point cloud features; The precise arm tracking module based on multi-task joint learning includes at least a gesture recognition branch module and an arm tracking branch module. The high-dimensional spatiotemporal feature representation of the point cloud is used as the input of the multi-task joint learning module. The gesture recognition branch module is used to guide the arm tracking branch module to output the motion trajectory of the arm joint, thereby realizing the tracking of the arm movement.

2. The dynamic millimeter wave radar point cloud human arm tracking system based on joint learning as claimed in claim 1, characterized in that: In the point cloud feature extraction module based on spatiotemporal feature encoding, the PointNet network is constructed by two layers of multi-layer perceptrons MLP with shared weights. The point features are passed through the first layer of multi-layer perceptron MLP, and a high-dimensional point feature set is output. The arm point cloud set and the high-dimensional point feature set are residually linked in the feature dimension of each point. The obtained feature dimension is used as the input of the second layer of multi-layer perceptron MLP, and the final high-dimensional point feature set representation is output; The multi-head attention mechanism inside the Transformer network works in parallel, each attention head independently calculates the attention weight, multiplies the attention score of each head by the value matrix, obtains the weighted features after spatial feature encoding through the attention head, and weights the output multi-head features through the linear layer to output the feature set.

3. The dynamic millimeter wave radar point cloud human arm tracking system based on joint learning as claimed in claim 1, characterized in that: The gesture recognition branch module includes at least a multi-layer perceptron and a linear layer, and maps the global feature vector of each frame in the sequence to the probability value prediction of each gesture through the Softmax function, and selects the maximum value as the final sequence label; The arm tracking branch module includes at least a multi-layer perceptron and a linear layer, which maps the global feature vector of each frame in the sequence into the three-dimensional coordinates of the arm joints.

4. A method for tracking human arms using a dynamic millimeter wave radar point cloud based on joint learning using the system as claimed in claim 1, characterized in that: The steps include: S1. Point cloud generation: The millimeter wave radar device is used to collect data information on the user's gestures, and continuous millimeter wave point cloud frames are obtained. The millimeter wave point cloud frames are used as input, and the spatial heuristic algorithm and DBSCAN point cloud clustering algorithm are used in sequence to filter out environmental noise and output the arm point cloud set; S2. Point cloud feature extraction: A spatiotemporal feature encoder is constructed through an improved PointNet network, a Transformer network, and a bidirectional long short-term memory network Bi-LSTM to map the dynamic arm point cloud input into a high-dimensional spatiotemporal feature representation; the improved PointNet network is constructed by a 2-layer multi-layer perceptron with shared weights. After the arm point cloud set is input into the improved PointNet network, a high-dimensional point feature set PF is output. t , and input it into the Transformer network to capture the long-distance dependencies between dynamic point clouds: g t =BiLSTM(f t ,f t-1 ,f t+1 ;θ b ) where θ b represents the learnable parameters of Bi-LSTM, g t is the global feature vector containing context information, f t is the frame feature; S3. Arm tracking: The high-dimensional spatiotemporal feature representation of the point cloud is used as the input of the multi-task joint learning module, and the gesture recognition task is introduced to assist in enhancing the neural network's ability to understand different arm movement patterns, guiding the arm tracking task to output accurate arm movement trajectories; in the gesture recognition task, the MLP and linear layers are used to map the global feature vector of each frame in the sequence to the probability value prediction of each gesture through the Softmax function, and the maximum value is selected as the final sequence label; in the arm tracking task, the global feature vector of each frame in the sequence is mapped to the three-dimensional coordinates of the arm joints through the MLP and linear layers.

5. The method for tracking human arm using dynamic millimeter wave radar point cloud based on joint learning as claimed in claim 4, characterized in that: The neighborhood radius of the DBSCAN clustering algorithm is 0.35 meters, the minimum point number threshold is 3, and the largest sub-cluster is identified as the millimeter wave point cloud generated by human arm movement.

6. The method for tracking human arm using dynamic millimeter wave radar point cloud based on joint learning as claimed in claim 4, characterized in that: In the Transformer network of step S2, the feature set PF at any time point t is t As input, it passes through any attention head h,t , get the query matrix Q h,t , key matrix K h,t Sum value matrix V h,t , the attention score of each head is combined with the value matrix V h,t Multiply, through the attention head h,t Get the weighted features after spatial feature encoding: Where d represents PF t The dimension of each point feature in , H represents the number of attention heads, h∈H; The output multi-head features are weighted through the linear layer Linear, and the output feature set PT t It is expressed as: PT t =Linear(PT 1,t ,PT 2,t ,…,PT h,t ;θ l ) where θ l represents the learnable parameters of the linear layer.

7. The method for tracking human arm using dynamic millimeter wave radar point cloud based on joint learning as claimed in claim 6, characterized in that: In step S2, the high-dimensional point feature set PF t And the feature set PT after Transformer network encoding t Perform residual connection and output the final point feature set PD′ through the MLP layer t : PD′ t =MLP(PF t ,PT t ;θ r ) where θ r represents the learnable parameters of MLP; The feature set PD′ t The point features pd′ of all points in i,t Aggregated into frame features f through self-attention layer t , get the global information of the frame: where θ f represents the learnable parameters of the linear layer.

8. The method for tracking human arm using dynamic millimeter wave radar point cloud based on joint learning as claimed in claim 4, characterized in that: In step S3, for a point cloud frame sequence of length S, a cross entropy loss is used to minimize the error between the predicted category and the true category of each sequence, and the loss function is: Where G is the number of gesture action categories, is a function whose output is 0 or 1. When it is 1, it indicates that the sample s belongs to category g. is the predicted probability that sample s belongs to category g; For a point cloud frame sequence of length S and K predicted joints per frame, the mean square error MSE loss is used to minimize the error between the predicted and actual arm main joint coordinates. The loss function is: where ||·|| represents the L2 normal form, and J k,t They represent the predicted and actual values ​​of the three-dimensional coordinates of the arm joints respectively.

9. The method for tracking human arm using dynamic millimeter wave radar point cloud based on joint learning as claimed in claim 8, characterized in that: The following function is used in step S3 to train the gesture recognition task and the arm tracking task end-to-end, calculate the gradient and perform back propagation, and update the parameters until the network converges: L=αL G +βL T Where α and β are hyperparameters of the model.

Citation Information

Cited By

  • Daily scene millimeter wave sensing non-human target filtering method and system

    CN121454515A

  • Daily scene millimeter wave perception non-human target filtering method and system

    CN121454515B

  • Millimeter wave sign language data automatic generation and enhancement method based on robot system

    CN121515222A

  • Robotic system-based automated generation and augmentation of millimeter wave sign language data

    CN121515222B