Human posture reconstruction method based on fusion of MIMO millimeter wave radar and infrared camera
By fusing MIMO millimeter-wave radar with an infrared camera, the problems of low resolution and poor stability in human posture reconstruction in existing technologies are solved, achieving high-precision and stable posture reconstruction, which is suitable for indoor applications.
Patent Information
- Application Number
- CN202511631023.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-11-10
AI Technical Summary
Existing human posture reconstruction methods based on radar or WiFi suffer from low resolution, poor stability, and unsatisfactory reconstruction results, which limits their practical application scenarios.
A human posture reconstruction method based on the fusion of MIMO millimeter-wave radar and infrared camera is adopted. By extracting two-dimensional spatial spectrum through radar signal preprocessing and combining it with infrared image features, a fused human posture reconstruction model is constructed to achieve high-precision and stable posture reconstruction.
It improves the resolution and stability of human posture reconstruction, achieving high-precision human posture perception, and is applicable to human-computer interaction, sports and health monitoring, and smart home applications in indoor scenarios.
Smart Images

Figure CN121095381B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of human posture reconstruction technology, specifically providing a human posture reconstruction method that integrates MIMO millimeter-wave radar and infrared camera. Background Technology
[0002] In the field of human posture reconstruction technology, existing methods are mainly based on visual camera perception or wearable perception. They use computer vision algorithms to locate joints from visual cameras or capture human joint positions using wearable devices to achieve human posture reconstruction. However, in everyday human posture reconstruction scenarios, user privacy, comfort, and ease of use need to be considered, making the above posture reconstruction methods unsuitable for direct application in daily life. In contrast, radar and infrared cameras have advantages such as being non-contact, not involving user privacy, and being insensitive to light. Therefore, using radar and infrared cameras to achieve human posture reconstruction has significant application value.
[0003] In recent years, some research has been conducted on non-contact and privacy-preserving methods for human pose reconstruction, and some progress has been made. For example, the paper "Zhao, M., Tian, Y., Zhao, H., et al.: 'RF-based 3D skeletons'. The 2018 Conference of the ACM Special Interest Group on Data Communication, Budapest, Hungary, 2018, pp. 267-281" proposes a three-dimensional human skeleton reconstruction system called RF-Pose3D. This system uses radio frequency signals for human perception and constructs a novel convolutional neural network architecture based on the obtained radio frequency features to estimate the three-dimensional skeletal structure from the radio frequency signals. Another example is the paper "Xue, H., Ju, Y., Miao, C., et al.: 'mmMesh: Towards 3D real-time dynamic human meshconstruction using millimeter-wave', The 19th Annual International Conference on Mobile Systems, Applications, and Services, NY, USA, 2021, pp. The paper “269-282” proposes a method for 3D human body reconstruction using millimeter-wave radar, and discloses a model called mmMesh for reconstructing human body contours from sparse 3D radar point clouds; another example is “Mueller, J., et al.: 'End-to-End Learning for Human Pose Estimation from Raw Millimeter Wave Radar Data', 202458th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove,CA, USA, 2024, pp. 1495-1499”, which proposes an end-to-end learning method using neural networks, combined with a neural network architecture that incorporates an attention mechanism, to directly extract target features from raw radar data, achieving high-precision human pose reconstruction; and the paper “Song, Y., Jin, T., Dai Y., Song, Y., Zhou, X”.The paper "Human pose reconstruction network based on ultra-wideband MIMO radar image with distance-assisted" (JOURNAL OF SIGNAL PROCESSING, 2021, 37, (8), pp:1355-1364) proposes a range-assisted human pose reconstruction network for ultra-wideband MIMO radar images. It uses convolutional neural networks to extract signal intensity and spatial location features of the human target image, and deconvolutional modules to reconstruct the positions of each joint of the human target.
[0004] However, the aforementioned human posture reconstruction methods mainly rely on radar or WiFi for human posture perception and reconstruction, which suffers from problems such as low resolution, poor stability, and unsatisfactory reconstruction results, thus limiting their practical application scenarios. To address these issues, this invention proposes a human posture reconstruction method that integrates MIMO millimeter-wave radar and an infrared camera. Summary of the Invention
[0005] The purpose of this invention is to provide a human posture reconstruction method that fuses MIMO millimeter-wave radar and infrared cameras, addressing the problems of low resolution, poor stability, and unsatisfactory reconstruction results in existing human posture reconstruction methods. This invention extracts a two-dimensional spatial spectrum of the human body using a radar signal preprocessing algorithm, and combines this with synchronously acquired infrared images to obtain two types of human posture feature representations. Then, it constructs a fused human posture reconstruction model based on MIMO millimeter-wave radar and infrared cameras. This fused model enables the effective extraction and fusion of the two types of human posture feature data, achieving high-precision and stable human posture reconstruction, ultimately realizing stable perception of human posture.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A method for human pose reconstruction by fusing MIMO millimeter-wave radar and infrared cameras includes the following steps:
[0008] Step 1. Construct a synchronous acquisition system based on MIMO millimeter-wave radar and infrared camera to synchronously acquire radar data and infrared images, forming a human posture dataset;
[0009] Step 2. Perform data preprocessing on the radar data of the human posture dataset to obtain a preprocessed two-dimensional spatial spectrum of the human body.
[0010] Step 3. Combine the two-dimensional spatial spectrum of the human body with the corresponding infrared image to form the input data, and combine it with the prior human posture to construct a human posture training set;
[0011] Step 4. Construct and train a fusion human pose reconstruction model based on MIMO millimeter-wave radar and infrared camera. The fusion human pose reconstruction model consists of an infrared image feature extraction module, a millimeter-wave radar feature extraction module, an attention-based feature fusion module, and a pose reconstruction module based on fused features. The infrared image and the human body two-dimensional spatial spectrum map are respectively input into the infrared image feature extraction module and the millimeter-wave radar feature extraction module. The output features are fused by the attention-based feature fusion module. The fused features are input into the pose reconstruction module based on fused features to obtain the human pose reconstruction result.
[0012] Step 5. Simultaneously acquire radar data and infrared images of the human target under test using the synchronous acquisition system in Step 1. Input the processed human two-dimensional spatial spectrum map and infrared image into the trained human posture reconstruction model, and output the human posture reconstruction result from the human posture reconstruction model.
[0013] Furthermore, in step 1, the MIMO millimeter-wave radar and infrared camera are set on a support platform at a preset height above the ground. The fields of view of both the MIMO millimeter-wave radar and the infrared camera are perpendicular to the ground and cover the human target. At the same time, the infrared camera and the MIMO millimeter-wave radar are controlled by a frame synchronization signal to achieve synchronous acquisition of radar data and infrared images.
[0014] Furthermore, in step 2, the specific process of data preprocessing is as follows:
[0015] For any frame of radar data, firstly, a fast Fourier transform is performed on the original data matrix to obtain a time-range spectrum; then, Doppler spectrum estimation is performed on the time-range spectrum to obtain a range-Doppler spectrum; next, two-dimensional constant false alarm rate (CFAR) detection is performed on the range-Doppler spectrum to obtain all range-Doppler cells containing human targets, and for each range-Doppler cell containing human targets, a least mean square distortion-free response algorithm is used to perform two-dimensional spatial spectrum estimation to obtain a two-dimensional spatial spectrum; finally, all two-dimensional spatial spectra are accumulated to obtain a human two-dimensional spatial spectrum.
[0016] Furthermore, the time-distance spectrum is represented as follows:
[0017] ,
[0018] in, Indicates the first At what moment is the human target at the [time]? Echo signal strength at each distance cell The first element of the original radar data matrix represents the... line, number Column of complex data, This represents the number of points in the Fast Fourier Transform.
[0019] Furthermore, the distance-Doppler spectrum is represented as follows:
[0020] ,
[0021] in, This represents the echo signal intensity of a human target at the m-th distance cell and the k-th Doppler cell. Represents the window function. Indicates the first At what moment is the human target at the [time]? Echo signal strength at each distance cell This represents the number of points in the Fast Fourier Transform.
[0022] Furthermore, the specific process of two-dimensional spatial spectrum estimation is as follows:
[0023] First, construct the steering vector of the antenna array for the MIMO millimeter-wave radar. :
[0024] ,
[0025] in, Indicates the angle in the horizontal direction. Indicates the angle in the vertical direction; Represents the horizontal guidance vector. ; Represents the guide vector in the vertical direction. ; Indicates matrix transpose. and These represent the number of channels in the horizontal and vertical directions, respectively.
[0026] Then, for the range-Doppler cell containing the target, the autocorrelation matrix is calculated based on the range-Doppler data. :
[0027] ,
[0028] ,
[0029] in, This represents cell data containing the range-Doppler cells of the target. Indicates the first The distance between the channels - Doppler spectrum in the 1st line, number Column data; This indicates the operation of calculating the mean. Indicates conjugate transpose;
[0030] The two-dimensional spatial spectrum is then represented as:
[0031] ,
[0032] in, It represents a two-dimensional spatial spectrum.
[0033] Furthermore, in step 3, the prior human posture acquisition process is as follows: under the same viewpoint of the synchronous acquisition system in step 1, RGB images are synchronously acquired using an RGB camera, and then the prior human posture is extracted through computer vision.
[0034] Furthermore, in step 4, the infrared image feature extraction module and the millimeter-wave radar feature extraction module adopt a symmetrical structure, specifically as follows:
[0035] (Conv1+BatchNorm+Relu)+(Conv2+Relu)+(Conv3+Relu),
[0036] Where Conv represents a convolutional layer, BatchNorm represents batch normalization, and ReLU represents an activation layer.
[0037] Furthermore, in step 4, the feature fusion module based on the attention mechanism adopts a two-dimensional convolutional block attention module.
[0038] Furthermore, in step 4, the specific structure of the pose reconstruction module based on fused features is as follows:
[0039] (Conv4+Relu)+Softargmax,
[0040] Where Conv represents a convolutional layer, ReLU represents an activation layer, and Softargmax represents a Softargmax layer.
[0041] Furthermore, in step 4, the mean squared error loss function is used during the training of the human pose reconstruction model, specifically as follows:
[0042] ,
[0043] in, This represents the mean squared error loss function. Indicates the number of joints. Indicates the first Predicted coordinates of each joint, Indicates the first The label coordinates of each joint.
[0044] Based on the above technical solution, the beneficial effects of the present invention are as follows:
[0045] This invention provides a method for human posture reconstruction by fusing MIMO millimeter-wave radar and infrared camera. The method uses millimeter-wave radar to compensate for the limitation of infrared camera being affected by ambient temperature, and uses infrared camera to compensate for the poor stability of millimeter-wave radar detection results. Finally, by fusing the human posture reconstruction model, the method effectively extracts and fuses the two-dimensional spatial spectral features of millimeter-wave radar and the features of infrared image, thereby achieving stable perception of human posture.
[0046] More specifically, the present invention has the following advantages:
[0047] 1) This invention uses millimeter-wave radar and infrared camera to simultaneously perceive human targets, which can effectively improve perception capabilities;
[0048] 2) This invention extracts the two-dimensional spatial spectrum of the human target from the radar spectrum for attitude characterization, and then combines it with the infrared image captured by the infrared camera, which effectively solves the problem of poor stability of the reconstruction results in the traditional radar-based human attitude reconstruction method and improves the stability of the reconstruction results.
[0049] 3) The fusion human posture reconstruction model based on MIMO millimeter-wave radar and infrared camera proposed in this invention can effectively extract and fuse the two-dimensional spatial spectral features of millimeter-wave radar and infrared image features of human posture, thereby achieving high-precision and stable human posture reconstruction.
[0050] 4) This invention can be applied to human-computer interaction, sports and health monitoring and smart home in indoor scenarios, and has broad application prospects. Attached Figure Description
[0051] Figure 1 This is a flowchart illustrating the human posture reconstruction method based on the fusion of MIMO millimeter-wave radar and infrared camera in an embodiment of the present invention.
[0052] Figure 2 This is a schematic diagram of the radar data preprocessing process in an embodiment of the present invention.
[0053] Figure 3 This is a schematic diagram of the original data cube structure of radar data in an embodiment of the present invention.
[0054] Figure 4 This is a time-range spectrum of radar data in an embodiment of the present invention.
[0055] Figure 5 This is a range-Doppler spectrum of radar data in an embodiment of the present invention.
[0056] Figure 6 This is a comparison diagram of RGB images, human body two-dimensional spatial spectrum and infrared images in an embodiment of the present invention.
[0057] Figure 7 This is a schematic diagram of the structure of the human posture reconstruction model integrated in an embodiment of the present invention.
[0058] Figure 8 This is a simulation test result of the average joint error on the test set in an embodiment of the present invention, where the human posture reconstruction model is integrated. Detailed Implementation
[0059] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0060] This embodiment provides a method for human posture reconstruction by fusing MIMO millimeter-wave radar and infrared camera, the process of which is as follows: Figure 1 As shown, the specific steps include:
[0061] Step 1. Construct a synchronous acquisition system based on MIMO millimeter-wave radar and infrared camera to synchronously acquire radar data and infrared images, forming a human posture dataset;
[0062] To achieve synchronous data acquisition between MIMO millimeter-wave radar and infrared camera, a synchronous acquisition system based on MIMO millimeter-wave radar and infrared camera is constructed. The system uses a three-transmitter, four-receiver millimeter-wave radar as the radar sensing front end, with its antenna array consisting of three transmitting antennas and four receiving antennas. The infrared camera uses an infrared camera with a resolution of 256×192. The infrared camera and MIMO millimeter-wave radar are controlled by a frame synchronization signal to achieve synchronous data acquisition between the radar and the infrared camera.
[0063] Referring to the actual application scenario of human posture reconstruction, the synchronous acquisition system is set in front of the human body and mounted on a tripod at a height of 1.2m above the ground. The fields of view of the MIMO millimeter-wave radar and the infrared camera are perpendicular to the ground. In this scenario, human posture datasets of four human targets performing arbitrary actions at a position 3.5m in front of the synchronous acquisition system are collected. The parameter configuration of the MIMO millimeter-wave radar is shown in Table 1. As can be seen from the table, the data acquisition frame rate of the MIMO millimeter-wave radar and the infrared camera is 10 frames / second. 10 sets of 1-minute data are collected for each target, that is, 6000 (10×60×10) frames of data are collected for each human target, and a total of 24000 frames of human posture data are collected to form a human posture dataset.
[0064] Table 1 MIMO millimeter-wave radar signal parameters
[0065] Parameter type Parameter value Parameter type Parameter value Starting frequency 60GHz FM slope 80MHz / us Chirp cycle 50us Frame period 100ms bandwidth 2.56GHz Sampling rate 4000ksps Number of sampling points 128 Chirp Quantity 64
[0066] Step 2. Perform data preprocessing on the radar data of the human posture dataset to obtain a preprocessed two-dimensional spatial spectrum of the human body.
[0067] The data processing flow of radar data is as follows Figure 2 As shown, for any frame of radar data, its original data cube is as follows: Figure 3 As shown, in order to extract the range information of human targets, a Fast Fourier Transform (FFT) is first performed on the original radar data matrix over a slow time interval to obtain the time-range spectrum, the calculation formula of which is shown below:
[0068] ,
[0069] in, Indicates the first At what moment is the human target at the [time]? Echo signal strength at each distance cell The first element of the original radar data matrix represents the... line, number Column of complex data, The number of points in the Fast Fourier Transform;
[0070] The above calculations yield the human body time-range spectrum, providing distance information between the radar and the human target. This distance information changes over time as the human target moves. After processing, the time-range spectrum of the 12 channels of the MIMO millimeter-wave radar is shown below. Figure 4 As shown;
[0071] To extract the target Doppler, Doppler spectrum estimation is performed on the time-range spectrum to obtain the range-Doppler spectrum, and the calculation formula is shown below:
[0072] ,
[0073] in, This represents the echo signal intensity of a human target at the m-th distance cell and the k-th Doppler cell. Represents the window function. Indicates the first At what moment is the human target at the [time]? Echo signal strength at each distance cell The number of points in the Fast Fourier Transform;
[0074] After processing all range cells in the current frame using the above method, a range-Doppler spectrum is obtained, as shown below. Figure 5 As shown;
[0075] Then, two-dimensional constant false alarm rate (CFAR) detection is performed on the range-Doppler spectrum to obtain range-Doppler cells containing human targets. For range-Doppler cells containing human targets, two-dimensional spatial spectrum estimation is performed using the minimum variance distortionless response (MVDR) algorithm to obtain the two-dimensional spatial spectrum map of the human target. The specific process is as follows:
[0076] First, construct the steering vector of the antenna array for the MIMO millimeter-wave radar. :
[0077] ,
[0078] in, Indicates the angle in the horizontal direction. Indicates the angle in the vertical direction; Represents the horizontal guidance vector. ; Represents the guide vector in the vertical direction. ; Indicates matrix transpose. and These represent the number of channels in the horizontal and vertical directions, respectively.
[0079] Then, for the range-Doppler cells containing the target, the range-Doppler data of the 12 virtual channels formed by the MIMO millimeter-wave radar are extracted, and the autocorrelation matrix is calculated. for:
[0080] ,
[0081] ,
[0082] in, This represents cell data containing the range-Doppler cells of the target. Indicates the first The distance between the channels - Doppler spectrum in the 1st line, number Column data; This indicates the operation of calculating the mean. Indicates conjugate transpose;
[0083] Based on the above calculations, a two-dimensional spatial spectrum estimation is performed, and its power spectrum is obtained. for:
[0084] ,
[0085] This yields a two-dimensional spatial spectrum over the current range-Doppler unit;
[0086] The above estimation is performed on all range-Doppler cells containing human targets to obtain a two-dimensional spatial spectrum of all targets. The two-dimensional spatial spectrum of the human body is obtained by accumulating (summing) all spatial spectra.
[0087] Step 3. Combine the two-dimensional spatial spectrum of the human body with the corresponding infrared image to form the input data, and combine it with the prior human posture to construct a human posture training set;
[0088] The prior human pose acquisition process is as follows: Under the same viewpoint of the synchronous acquisition system in step 1, RGB images are synchronously acquired using an RGB camera, and then the prior human pose is extracted using computer vision. The process of extracting the human pose using computer vision employs existing techniques in the field, which will not be elaborated here. In this embodiment, the synchronously acquired RGB images, human two-dimensional spatial spectrum, and infrared images are as follows: Figure 6 As shown, from left to right, the images are an RGB image, a two-dimensional spatial spectrum of the human body, and an infrared image.
[0089] Step 4. Construct and train a fusion human pose reconstruction model based on MIMO millimeter-wave radar and infrared camera;
[0090] Since human pose reconstruction by fusing MIMO millimeter-wave radar and infrared cameras requires the simultaneous extraction of human two-dimensional spatial spectral maps and infrared image features, and the fusion of features and pose reconstruction based on the detection performance advantages of both, this embodiment constructs a human pose reconstruction model fused with MIMO millimeter-wave radar and infrared cameras, as follows: Figure 7 As shown, it consists of an infrared image feature extraction module, a millimeter-wave radar feature extraction module, an attention-based feature fusion module, and a pose reconstruction module based on fused features;
[0091] Both the infrared image feature extraction module and the millimeter-wave radar feature extraction module are used to extract human pose features from the input image. Therefore, the two modules adopt a symmetrical structure, but the model parameters after training are different; the specific structure is as follows:
[0092] (Conv1+BatchNorm+Relu)+(Conv2+Relu)+(Conv3+Relu), where Conv represents a convolutional layer, BatchNorm represents batch normalization, Relu represents an activation layer, and "+" indicates sequential concatenation. The convolutional layer Conv1 receives the input infrared image or human two-dimensional spatial spectrum, extracts the features, and performs normalization and nonlinear mapping through the batch normalization layer and the activation layer. Then, it goes through the next two convolutional layers to further extract the image features. After extraction, the features are sent to the feature fusion module based on the attention mechanism.
[0093] The feature fusion module based on the attention mechanism is built using a two-dimensional convolutional block attention module (CBAM2D). This module automatically learns the importance of two input features in the channel and spatial dimensions, and then adaptively adjusts different input feature maps to enhance important features and suppress unimportant features.
[0094] The pose reconstruction module based on fused features consists of a convolutional layer, an activation layer, and an output layer, with the following structure:
[0095] (Conv4+Relu)+Softargmax, where Conv represents a convolutional layer, Relu represents an activation layer, and Softargmax represents a softargmax layer; the Conv4 convolutional layer and the Relu activation layer extract features from the received fused features to obtain an output feature map consistent with the number of joints. Each feature map corresponds to the probability distribution of a joint, that is, the point with the highest probability in the output feature map is the position of the corresponding joint point; on this basis, the output layer uses a softargmax layer to extract the corresponding coordinates to obtain the coordinates of each joint point. Finally, after connecting all the joint points, the human pose reconstruction result is obtained.
[0096] The loss function is set as the mean squared error loss function, and the human pose reconstruction model is trained using the human pose training set from step 3. The mean squared error loss (MSE) is calculated as follows:
[0097] ,
[0098] in, Indicates the number of joints. Indicates the first Predicted coordinates of each joint, Indicates the first The label coordinates of each joint;
[0099] Step 5. Simultaneously acquire radar data and infrared images of the human target under test using the synchronous acquisition system in Step 1. Input the processed human two-dimensional spatial spectrum map and infrared image into the trained human posture reconstruction model, and output the human posture reconstruction result from the human posture reconstruction model.
[0100] The beneficial effects of the present invention will be explained in detail below with reference to simulation tests.
[0101] In this embodiment, the human pose training set obtained in step 3 is divided into a training set and a test set in a 3:1 ratio; specifically, 18,000 frames of data from three human targets are used for model training, and 6,000 frames of data from another human target are used for model testing; for example... Figure 8 The figure shows the average joint error of the human pose reconstruction model on the test set in this embodiment. As can be seen from the figure, after the first iteration, the average joint error of the model on the test set is about 6.18 pixels. After two hundred iterations, the model converges to the best effect, and its average joint error is controlled at 3.12 pixels.
[0102] In summary, the human posture reconstruction method proposed in this invention, which integrates MIMO millimeter-wave radar and infrared camera, can achieve stable and high-precision human posture reconstruction.
[0103] The above description is merely a specific embodiment of the present invention. Any feature disclosed in this specification may be replaced by other equivalent or similar features unless otherwise specified. All disclosed features, or steps in all methods or processes, may be combined in any way except for mutually exclusive features and / or steps.
Claims
1. A human posture reconstruction method based on fusion of MIMO millimeter wave radar and infrared camera, characterized in that, Comprising the following steps: Step 1. Synchronous acquisition system is constructed based on MIMO millimeter wave radar and infrared camera, and radar data and infrared image are synchronously acquired to construct human posture data set; Step 2. Data preprocessing is performed on radar data in human posture data set to obtain radar signal preprocessed human two-dimensional spatial spectrum graph; the specific process of data preprocessing is as follows: For any one frame of radar data, firstly, fast Fourier transform is performed on the original data matrix to obtain a time-distance spectrum graph; then, Doppler spectrum estimation is performed on the time-distance spectrum graph to obtain a distance-Doppler spectrum graph; then, two-dimensional constant false alarm rate detection is performed on the distance-Doppler spectrum graph to obtain all distance-Doppler units containing human targets, and two-dimensional spatial spectrum estimation is performed on each distance-Doppler unit containing human targets by using minimum mean square distortionless response algorithm to obtain a two-dimensional spatial spectrum graph; finally, all two-dimensional spatial spectrum graphs are accumulated to obtain a human two-dimensional spatial spectrum graph; The time-distance spectrum graph is represented as: , wherein, represents the echo signal intensity of the human target on the distance unit at the time moment, represents the complex data of the row, the column of the radar raw data matrix, is the number of points of the fast Fourier transform; The distance-Doppler spectrum graph is represented as: , wherein, represents the echo signal intensity of the human target at the mth range cell and the kth Doppler cell, represents a window function, represents the echo signal intensity of the human target at the mth range cell and the kth Doppler cell at the nth time point, represents the echo signal intensity of the human target at the mth range cell at the nth time point, represents the echo signal intensity of the human target at the mth range cell and the kth Doppler cell at the nth time point, is the number of points of the fast Fourier transform; The specific process of two-dimensional spatial spectrum estimation is as follows: First, steering vectors of an antenna array of a MIMO millimeter wave radar are constructed : , wherein, denotes an angle in horizontal direction, denotes an angle in vertical direction; denotes a guiding vector in horizontal direction, denotes a guiding vector in vertical direction, denotes a matrix transpose, and denote the number of channels in horizontal and vertical direction, respectively. Then, for the range-Doppler cell containing the target, the auto-correlation matrix is computed from the range-Doppler data : , , in, This represents cell data containing the range-Doppler cells of the target. Indicates the first The distance between the channels - Doppler spectrum in the 1st line, number Column data; This indicates the operation of calculating the mean. Indicates conjugate transpose; Then, the two-dimensional spatial spectrum is represented as: , wherein represents a two-dimensional spatial spectrum; Step 3. The human two-dimensional spatial spectrum graph and the corresponding infrared image are constructed into input data, and a human posture training set is constructed in combination with prior human posture; Step 4. A fusion human posture reconstruction model based on MIMO millimeter wave radar and infrared camera is constructed and trained, and the fusion human posture reconstruction model comprises an infrared image feature extraction module, a millimeter wave radar feature extraction module, a feature fusion module based on an attention mechanism, and a posture reconstruction module based on fused features; the infrared image and the human two-dimensional spatial spectrum graph are input into the infrared image feature extraction module and the millimeter wave radar feature extraction module, the features output are fused by the feature fusion module based on the attention mechanism, and the fused features are input into the posture reconstruction module based on the fused features to obtain a human posture reconstruction result; Step 5. The radar data and the infrared image of the human target to be measured are synchronously acquired by using the synchronous acquisition system in step 1, the human two-dimensional spatial spectrum graph after data processing and the infrared image are input into the trained fusion human posture reconstruction model, and the human posture reconstruction result is output by the fusion human posture reconstruction model.
2. The method of claim 1, wherein the MIMO millimeter wave radar and infrared camera fusion human pose reconstruction method is characterized by, In step 1, the MIMO millimeter wave radar and the infrared camera are arranged on a support table at a preset height from the ground, the fields of view of the MIMO millimeter wave radar and the infrared camera are perpendicular to the ground and cover the human target; at the same time, the infrared camera and the MIMO millimeter wave radar are controlled by a frame synchronization signal to realize synchronous acquisition of radar data and infrared images.
3. The method of claim 1, wherein the MIMO millimeter wave radar and infrared camera fusion human pose reconstruction method is characterized by, In step 3, the acquisition process of the prior human posture is as follows: the RGB camera is used to synchronously acquire an RGB image under the same view angle of the synchronous acquisition system in step 1, and the prior human posture is obtained by computer vision extraction.
4. The method of claim 1, wherein the MIMO millimeter wave radar and infrared camera fusion human pose reconstruction method is characterized by, In step 4, the infrared image feature extraction module and the millimeter wave radar feature extraction module adopt a symmetrical structure, and the specific structure is as follows: (Conv1+BatchNorm+Relu)+(Conv2+Relu)+(Conv3+Relu), Wherein, Conv represents a convolution layer, BatchNorm represents a batch normalization, and Relu represents an activation layer. The feature fusion module based on the attention mechanism adopts a two-dimensional convolution block attention module. The specific structure of the pose reconstruction module based on the fused features is as follows: (Conv4+Relu)+Softargmax, Wherein, Softargmax represents a Softargmax layer.
5. The method of claim 1, wherein the MIMO millimeter wave radar and infrared camera fusion human pose reconstruction method is characterized by, In step 4, the mean square error loss function is used in the training process of the human body pose reconstruction model, and the specific process is as follows: , wherein, represents a mean squared error loss function, represents the number of joints, represents the predicted coordinates of the j-th joint, represents the label coordinates of the j-th joint.
Citation Information
Patent Citations
Character modeling and fitness action teaching method and system based on non-contact action capture
CN120563731A
Human body posture estimation method based on radar point cloud imaging and multi-dimensional feature fusion
CN120877384A