Millimeter wave radar point cloud human body posture estimation method based on speed correction
By using the speed correction technology in the mmWave radar point cloud human posture estimation method, the encoder decoder structure is used to roughly estimate and fine-tune the position of human joints, the uncertainty problem of human posture estimation in the complex environment in the prior art is solved, and higher pose estimation accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411831822.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-12
AI Technical Summary
The existing 3D human pose estimation methods based on images and videos have uncertainties in dealing with complex environments, especially in night environments and privacy scenarios, and the accuracy of human pose estimation based on radar signals is difficult to reach the level of image methods.
A method of human posture estimation based on speed correction is adopted for the human body through the encoder decoder structure, and the joint position of the human body is roughly estimated using the velocity characteristics to further correct the joint position to form a human body posture estimation model with good performance.
It improves the accuracy and robustness of radar-based human posture estimation, and can provide more accurate human posture prediction in a diverse environment, enhancing the application potential of the model.
Smart Images

Figure CN119963640A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human body posture prediction, and in particular to a millimeter wave radar point cloud human body posture estimation method based on velocity correction. Background Art
[0002] Although 3D human pose estimation based on images and videos has made significant progress in many scenarios, they face a fundamental challenge, namely, they are limited to only providing two-dimensional information about the human body. This situation leads to uncertainty in predicting three-dimensional information, especially in complex situations such as occlusion, illumination changes, and image perspective changes. In such cases, relying solely on image and video data may not fully restore the true pose of the human body in three-dimensional space, limiting the accuracy and robustness of applications in actual scenarios.
[0003] In nighttime environments, due to lighting conditions, the image information obtained by the image acquisition device is difficult to estimate human posture due to underexposure, overexposure, and motion blur, and the effect is not ideal. In some privacy scenarios, human posture estimation based on images cannot guarantee privacy.
[0004] Millimeter wave radar is a non-line-of-sight sensing technology that can penetrate optically opaque objects and is less affected by lighting conditions and occlusions. In recent years, human posture estimation based on radar signals has become a new research direction. The basic principle is to accurately estimate the key points and posture information of the human skeleton by analyzing the human body signal reflected by the radar. This method also has the advantage of robustness in diverse environments. However, due to the characteristics of radar signals being weak and susceptible to interference, the accuracy of its posture estimation has been difficult to reach a level that can replace image-based human posture estimation. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides a millimeter wave radar point cloud human posture estimation method based on velocity correction. The present invention mainly uses the encoder-decoder structure to obtain a rough estimate of the human joint position, and further corrects the human joint position using velocity characteristics, thereby obtaining a radar-based human posture estimation model with good performance, so as to accurately predict the human posture from the radar signal.
[0006] The technical means adopted by the present invention are as follows:
[0007] A method for estimating human body posture from millimeter wave radar point cloud based on velocity correction, characterized in that it comprises the following steps:
[0008] Obtain radar point cloud data to be estimated, and perform multi-frame point cloud aggregation on the radar point cloud data in the form of a sliding window, thereby generating radar point cloud input data;
[0009] The radar point cloud input data is respectively input into the trained rough estimation network branch model and the spatial velocity feature fine-tuning network model; the rough estimation network branch model is constructed based on the human sequence codec structure, and is used to roughly estimate the position of human bone points. The rough estimation network branch model includes a point cloud voxelization module, two GRU encoding units, an attention layer and a GRU decoder; the spatial velocity feature fine-tuning network model is constructed based on 3D convolution and long short-term memory network, and is used to extract features of velocity data projected in the point cloud voxelization space to obtain spatial velocity features;
[0010] The rough estimated human skeleton point positions output by the rough estimation network branch model are fine-tuned based on the spatial velocity features output by the spatial velocity feature fine-tuning network model, thereby generating accurate human skeleton point positions.
[0011] Furthermore, the position of human skeleton points is roughly estimated based on the rough estimation network branch model, including:
[0012] Perform voxel mapping encoding on the radar point cloud input data to obtain a sequence corresponding to the point cloud position;
[0013] Learning the encoding representation of the sequence through a GRU encoding unit;
[0014] The attention layer improves the efficiency of the decoder by allowing the learning process to focus on the intermediate encoder states while focusing on the final encoder state.
[0015] The decoder predicts the final human skeleton point sequence encoding representation;
[0016] The human skeleton coding sequence output by the decoder is devoxelized to obtain a rough estimate of the human skeleton point position.
[0017] Furthermore, based on the spatial velocity feature, the network model is fine-tuned to extract features from the radar point cloud data spliced with the velocity feature, including:
[0018] Acquire velocity data based on radar point cloud input data, and project the velocity data into point cloud voxelized space to acquire point cloud velocity data;
[0019] Perform feature extraction on the point cloud velocity data through 3D convolution to obtain spatial velocity features;
[0020] The spatial velocity features are input into the LSTM layer to capture the dependencies between consecutive frames of the time series velocity data.
[0021] Further, based on the spatial velocity feature output by the spatial velocity feature fine-tuning network model, the rough estimate of the human skeleton point position output by the rough estimate network branch model is fine-tuned, including:
[0022] Use a fully connected layer to align the feature channels of the spatial velocity features;
[0023] A fully connected layer is used to fine-tune the rough estimate of the human skeleton point position output by the rough estimate network.
[0024] Furthermore, the radar point cloud data is aggregated into multiple frames in the form of a sliding window to generate radar point cloud input data, including:
[0025] In the form of a sliding window, the previous two frames of radar point cloud data and the current frame data are retained each time to make full use of the continuity characteristics between frames.
[0026] Furthermore, the training steps of the rough estimation network branch model and the spatial velocity feature fine-tuning network model include:
[0027] Acquire human body posture data collected by millimeter wave radar, perform Fourier transform, MTI filtering and CFAR detection on the human body posture data, so as to obtain human body radar point cloud data;
[0028] Acquire the human joint point position data corresponding to the human body radar point cloud data through a Kinect camera;
[0029] The human body radar point cloud data is used as input data and the human body joint point position data is used as output data to train the rough estimation network branch model and the spatial velocity feature fine-tuning network model.
[0030] Compared with the prior art, the present invention has the following advantages:
[0031] The overall framework of the present invention adopts the idea of rough estimation followed by fine tuning. The rough estimation network maps the point cloud into a sequence by voxelization, making full use of the strong learning ability of the serialization relationship of the human body sequence codec, and obtains a rough estimate of the position of the human body joints through the rough estimation network. The fine tuning network uses the radial velocity information obtained by Fourier transform of radar data, effectively extracts the corresponding position velocity features through 3D convolution, and further fine-tunes and optimizes the network prediction results. At the same time, multi-frame aggregate input effectively utilizes the continuity relationship between frames, making the human body posture predicted by the model more coherent, thereby improving the overall performance of the human body posture estimation algorithm based on millimeter wave radar.
[0032] Based on the above reasons, the present invention can be widely promoted in the fields of radar human posture estimation and the like. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0034] Figure 1 The present invention is a flowchart of a method for estimating human body posture from a millimeter-wave radar point cloud based on velocity correction in an embodiment of the present invention.
[0035] Figure 2 This is a model architecture diagram of a millimeter-wave radar point cloud human posture estimation method based on velocity correction in an embodiment of the present invention.
[0036] Figure 3 4 is a flow chart of radar signal processing in an embodiment of the present invention.
[0037] Figure 4 This is a flow chart of point cloud voxelization in an embodiment of the present invention. DETAILED DESCRIPTION
[0038] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0039] like Figure 1 As shown, the present invention provides a method for estimating human body posture from millimeter wave radar point cloud based on velocity correction, comprising:
[0040] S1. Obtain radar point cloud data to be estimated, and perform multi-frame point cloud aggregation on the radar point cloud data in the form of a sliding window to generate radar point cloud input data.
[0041] In this application, TI IWR6843AOP millimeter wave radar is used to collect human body posture radar data. The directly collected human body posture radar data is subjected to Fourier transform, MTI filtering and CFAR detection to obtain the three-dimensional coordinates of the human body radar point data.
[0042] Before using the method of the present invention for posture estimation, the steps of training data acquisition and model training are also included. Specifically, in this application, TI IWR6843AOP millimeter wave radar is used to collect human posture radar data, and a Kinect camera is placed at the same height to obtain the corresponding real human joint point coordinates. The two are aligned according to the time frame, and the coordinates obtained by the camera are used as the supervision data for training the network.
[0043] S2. Input the radar point cloud input data into the trained rough estimation network branch model and the spatial velocity feature fine-tuning network model respectively. The architecture of the rough estimation network branch model and the spatial velocity feature fine-tuning network model is as follows: Figure 2 shown.
[0044] The rough estimation network branch model is built based on the human sequence encoder-decoder structure, and the rough estimation network branch model includes a point cloud voxelization module, two GRU encoding units, an attention layer, and a GRU decoder. The rough estimation network branch based on voxelized data and the encoder-decoder is used to roughly estimate the position of human skeleton points, including:
[0045] a. Perform voxel mapping encoding on the point cloud data to obtain the sequence corresponding to the point cloud position. Figure 4 The figure shows the flow chart of point cloud voxelization. First, the center of the radar point cloud is calculated, and a large cube block (length x width x height, 1.7mx1.4mx2.2m) that can completely cover the human body point cloud is constructed with the center point. Then the large cube block is divided into multiple cube blocks with a side length of 5cm, and sorted to ensure that the serial number of each small cube block is unique. According to the actual position of the point cloud, it is embedded into the corresponding serial number block. The obtained serial number sequence is the encoding sequence obtained after our voxelization. In this way, the spatial information of the point cloud is retained, and the spatial information of the point cloud is reduced from a three-dimensional representation to a one-dimensional sequence.
[0046] b. Learn the encoding representation of the sequence through the GRU encoding unit.
[0047] c. The attention layer improves the efficiency of the decoder by allowing it to focus on the intermediate encoder states while focusing on the final encoder state.
[0048] d. The decoder predicts the final human skeleton point sequence encoding representation.
[0049] e. Perform devoxel mapping on the human skeleton coding sequence output by the decoder to obtain a rough estimate of the position of the human skeleton points.
[0050] Specifically, the rough estimation network model based on the human sequence encoder-decoder, such as Figure 2 As shown, the coarse estimation network contains a voxelization module, an embedding layer, an encoder layer, a decoder layer, and an output layer.
[0051] First, the input radar point cloud data has the shape x∈R N×90×3 , where N represents the number of time frames, 90 is the number of points per frame, and 3 represents the spatial coordinates of each point. In order to make full use of the continuity between time frames, multi-frame aggregation input is adopted. The point cloud data is first voxelized to divide the three-dimensional point cloud data of each frame into voxels. The output shape of voxelization becomes N×90×1. At the same time, each voxel is mapped to a high-dimensional embedding space by looking up the voxel dictionary.
[0052] Next, after the embedding layer, the dimension of the voxel feature data becomes N×90×50, and is further converted to N×4500 through a flattening operation. Then, these data are input into a two-layer GRU encoder module. Each layer of the encoder GRU unit contains 512 hidden units to extract spatial and velocity features in the sequence. The encoding process outputs two types of states: intermediate state sequences (shape is N×512) and final states (shape is 1×512).
[0053] In the decoding stage, the decoder consists of a single-layer GRU and an attention layer, and the hidden state dimension of the GRU unit is also 512. The decoder first receives the final state of the encoder as the initial input to capture global features. At the same time, in each decoding step, the attention layer focuses on the key features of the input data through the intermediate state sequence and outputs the current state of the decoder. The output shape of the decoder is 1×1024, which is concatenated with the prediction result of the previous time step to dynamically adjust the prediction.
[0054] After the decoder output, the concatenated feature map is further processed by the fully connected layer, and the output layer is mapped to the final predicted voxel format. Finally, the voxel dictionary is used to perform a devoxelization operation to restore the predicted voxels to three-dimensional space coordinates, and the skeleton joint positions with a resolution of 25×3 are obtained to achieve a rough human pose estimation.
[0055] The spatial velocity feature fine-tuning network model is built based on 3D convolution and long short-term memory networks, and is used to extract features from velocity data projected in the point cloud voxel space to obtain spatial velocity features. It includes:
[0056] a. Project the velocity into the point cloud voxel space to obtain the point cloud velocity feature;
[0057] b. Extract features of the three-dimensional point cloud spatial data through 3D convolution to obtain spatial velocity features;
[0058] c. Input the spatial velocity features into the LSTM layer to capture the dependencies between consecutive frames of the time series velocity data;
[0059] d. Use a fully connected layer to align the feature channels;
[0060] e. Use a fully connected layer to fine-tune the output of the rough estimate network.
[0061] Build a fine-tuned network model based on convolution and timing models, such as Figure 2 As shown in Figure 1, the fine-tuning network includes a feature extraction module and a fine feature fusion module. This network focuses on fine-tuning the posture information output by the rough estimation network.
[0062] First, the input of the fine-tuning network is the velocity data projected in the point cloud voxel space. The velocity information is associated with the position of each point in the point cloud, that is, the velocity information of the point cloud point. It is obtained through the Doppler effect of the radar. The frequency change of the radar wave echo is proportional to the speed of the object. Therefore, by analyzing the frequency change of the echo signal, that is, the fast Fourier transform (FFT), the speed of the object under the radar perspective can be estimated. The shape is x∈R N×35×29×45 , N is the number of time frames, and 35, 29, and 45 are the spatial dimensions of the point cloud. The input data is processed by the 3D convolution layer to extract its features in the spatial and velocity dimensions. The feature extraction module contains multiple convolution layers, each containing a different number of 3D convolution kernels to extract multi-scale features.
[0063] In the feature extraction module, first, the input data is channel-expanded and transformed into a four-dimensional tensor (35, 29, 45, 1), and then input into the first 3D convolution layer. The first 3D convolution layer has 64 3×3×3 convolution kernels, the activation function is ReLU, and the number of output channels becomes 64; then, it passes through a maximum pooling layer of size 2×2×2; the second 3D convolution layer has 128 3×3×3 convolution kernels, the activation function is also ReLU, and the number of output channels becomes 128; in order to enhance the robustness of the network, a Dropout layer is added after the second 3D convolution layer, keeping the probability at 0.5; finally, it passes through another 2×2×2 maximum pooling layer.
[0064] After the convolution and pooling operations are completed, the feature map is mapped to the final channel dimension through the GlobalAveragePooling3D layer, and then the feature vector enters the long short-term memory network (LSTM) layer. The LSTM unit has 512 hidden nodes to capture the temporal dependencies in the input data. Then, the feature is input into the fully connected layer and mapped to the same dimension (i.e. 2602) as the output of the main branch of the rough estimation network through the ReLU activation function to form a finely adjusted radar speed feature vector.
[0065] S3. Fine-tune the rough estimate of the human skeleton point position output by the rough estimate network branch model based on the spatial velocity feature output by the spatial velocity feature fine-tuning network model, thereby generating an accurate human skeleton point position.
[0066] In this step, the position of the human body posture skeleton points obtained by the rough estimation network is further fine-tuned through the spatial velocity features by using a fine-tuning network model based on 3D convolution and LSTM.
[0067] Radar velocity feature vector F generated by fine-tuning the network v It is adjusted by a scaling factor α (e.g. 0.1) to ensure a moderate impact on the main branch output.
[0068] F v ′ =α·F v
[0069] Then, it is spliced with the output of the main branch, and the coordinate information and feature information are merged column by column to obtain the fused features:
[0070] C fused =[C main ∣F v ′ ]
[0071] The concatenation operation “∣” indicates merging columns in the dimension. The main branch C main The coordinate dimension of the output is (N, D), the feature dimension of the fine-tuning branch output is (N, M), and the fused feature vector C fused The dimension is (N, D+M).
[0072] The final fusion features are directly passed to the output layer, and the fine-tuned human posture prediction results are obtained through the fully connected layer.
[0073] The present invention also discloses that the training steps of the rough estimation network branch model and the spatial velocity feature fine-tuning network model include:
[0074] S01, obtaining human body posture data collected by millimeter wave radar, performing Fourier transform, MTI filtering and CFAR detection on the human body posture data, so as to obtain human body radar point cloud data. Specifically, Figure 3 As shown in the figure, the preprocessing process of millimeter-wave radar data is demonstrated. First, the radar echo signal is processed by Fast-Time FFT to obtain the distance information of the target; then, the frequency shift is analyzed by Slow-Time FFT to extract the radial velocity information of the target. Next, MTI filtering is applied to remove static clutter to ensure the detection of dynamic targets. Target detection is performed through the CFAR algorithm to eliminate background noise and ensure accurate identification of effective targets. Then, Channel FFT is used to obtain the angle information of the target from the multi-channel signal, and coordinate transformation is performed to convert the polar coordinates into three-dimensional space coordinates in the Cartesian coordinate system, and finally generate accurate point cloud data, which provides a basis for subsequent attitude estimation and target analysis.
[0075] S02, obtaining human joint point position data corresponding to the human body radar point cloud data through a Kinect camera;
[0076] S03, using the human body radar point cloud data as input data and the human body joint point position data as output data to train the rough estimation network branch model and the spatial velocity feature fine-tuning network model.
[0077] The present invention improves the accuracy and robustness of posture estimation by combining the advantages of a rough estimation network and a fine-tuning network through a multi-branch network structure. The rough estimation network uses modules such as GRU and attention layer to efficiently extract and model the spatiotemporal features of point cloud data, and achieves a rough estimate of the preliminary coordinates. The fine-tuning network effectively captures spatial velocity features through a feature extraction module composed of modules such as 3D convolutional layers and pooling layers, and further optimizes the output of the rough estimation network. The two network branches work together to effectively improve the accuracy and consistency of posture estimation. At the same time, the multi-frame aggregate input makes full use of the complementarity between frames, and retains the radar point cloud data of the previous two frames and the current frame through a sliding window mechanism to form a multi-frame aggregate data input network. The continuity feature between frames is effectively utilized to improve the coherence of movements when predicting human movements, and further enhances the model's dynamic tracking ability for human posture.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for estimating human body posture from millimeter wave radar point cloud based on velocity correction, characterized in that: The following steps are involved: Obtain radar point cloud data to be estimated, and perform multi-frame point cloud aggregation on the radar point cloud data in the form of a sliding window, thereby generating radar point cloud input data; The radar point cloud input data is respectively input into the trained rough estimation network branch model and the spatial velocity feature fine-tuning network model; the rough estimation network branch model is constructed based on the human sequence codec structure, and is used to roughly estimate the position of human bone points. The rough estimation network branch model includes a point cloud voxelization module, two GRU encoding units, an attention layer and a GRU decoder; the spatial velocity feature fine-tuning network model is constructed based on 3D convolution and long short-term memory network, and is used to extract features of velocity data projected in the point cloud voxelization space to obtain spatial velocity features; The rough estimated human skeleton point positions output by the rough estimation network branch model are fine-tuned based on the spatial velocity features output by the spatial velocity feature fine-tuning network model, thereby generating accurate human skeleton point positions.
2. The method for estimating human body posture from millimeter wave radar point cloud based on velocity correction according to claim 1, characterized in that: The position of human skeleton points is roughly estimated based on the rough estimation network branch model, including: Perform voxel mapping encoding on the radar point cloud input data to obtain a sequence corresponding to the point cloud position; Learning the encoding representation of the sequence through a GRU encoding unit; The attention layer improves the efficiency of the decoder by allowing the learning process to focus on the intermediate encoder states while focusing on the final encoder state. The decoder predicts the final human skeleton point sequence encoding representation; The human skeleton coding sequence output by the decoder is devoxelized to obtain a rough estimate of the human skeleton point position.
3. The method for estimating human body posture from millimeter wave radar point cloud based on velocity correction according to claim 1, characterized in that: Based on the spatial velocity feature, the fine-tuned network model is used to extract features from the radar point cloud data spliced with the velocity feature, including: Acquire velocity data based on radar point cloud input data, and project the velocity data into point cloud voxelized space to acquire point cloud velocity data; Perform feature extraction on the point cloud velocity data through 3D convolution to obtain spatial velocity features; The spatial velocity features are input into the LSTM layer to capture the dependencies between consecutive frames of the time series velocity data.
4. The method for estimating human body posture from millimeter wave radar point cloud based on velocity correction according to claim 1, characterized in that: Fine-tuning the rough estimate of the human skeleton point position output by the rough estimate network branch model based on the spatial velocity feature output by the spatial velocity feature fine-tuning network model includes: Use a fully connected layer to align the feature channels of the spatial velocity features; A fully connected layer is used to fine-tune the rough estimate of the human skeleton point position output by the rough estimate network.
5. The method for estimating human body posture from millimeter wave radar point cloud based on velocity correction according to claim 1, characterized in that: Aggregate multiple frames of radar point cloud data in the form of a sliding window to generate radar point cloud input data, including: In the form of a sliding window, the previous two frames of radar point cloud data and the current frame data are retained each time to make full use of the continuity characteristics between frames.
6. The method for estimating human body posture from millimeter wave radar point cloud based on velocity correction according to claim 1, characterized in that: The training steps of the rough estimation network branch model and the spatial velocity feature fine-tuning network model include: Acquire human body posture data collected by millimeter wave radar, perform Fourier transform, MTI filtering and CFAR detection on the human body posture data, so as to obtain human body radar point cloud data; Acquire the human joint point position data corresponding to the human body radar point cloud data through a Kinect camera; The human body radar point cloud data is used as input data and the human body joint point position data is used as output data to train the rough estimation network branch model and the spatial velocity feature fine-tuning network model.
Citation Information
Patent Citations
Human body posture reconstruction method and system based on millimeter wave radar point cloud
CN118169645A
Posture recognition method and device, equipment and storage medium
CN118212678A
Method, device, computer system for detecting pedestrian based on 3D point clouds
US20240193788A1
Cited By
Segmentation model training method, segmentation method and device for tumor lesion edema area
CN120580430A
Mine vehicle speed estimation method and system based on radar point cloud data, storage medium and program product
CN120928356A