Through-the-wall radar and binocular camera multi-person posture synchronous acquisition method and system
By combining the method of passing through the wall radar and binocular camera, the problems of high equipment costs, complex deployment and occlusion in the prior art are solved, and high-quality three-dimensional human posture estimation is achieved, which improves positioning accuracy and accuracy of posture estimation.
Patent Information
- Application Number
- CN202510359526.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art has problems in human posture estimation with high equipment costs, complex deployment, insufficient positioning accuracy and occlusion. In particular, the method based on a monocular camera has a large three-dimensional position error in the world coordinate system, and the method based on a binocular camera cannot effectively deal with occlusion of pose key points.
The method of combining the wall-passing radar and binocular camera is adopted to obtain the radar echo signal and camera image, preprocessing, time matching, attitude extraction and fusion are performed, and the three-dimensional attitude is estimated using YOLO and Multi-HMR algorithms, and the accuracy of attitude labels is improved through spatial registration.
The quality of three-dimensional human posture acquisition is improved, the problems of high equipment costs, complex deployment and occlusion are solved, and the positioning accuracy and the accuracy of posture estimation are improved.
Smart Images

Figure CN120385997A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and through-wall radar, and specifically, to a method and system for synchronously collecting multi-person postures of a through-wall radar and a binocular camera. Background Art
[0002] With the development of radar technology, multiple-input multiple-output (MIMO) ultra-wideband (UWB) through-wall radars have high enough range resolution and angular resolution to have the potential to achieve human posture estimation behind walls. In recent years, there have also been studies on human posture estimation behind walls using through-wall radars based on deep learning. To achieve high-precision posture estimation, these methods all use the human postures from vision as labels.
[0003] Currently, there are two main methods for obtaining human posture labels in the academic community: The first is to use a motion capture system composed of 12 cameras. A total of 12 cameras are arranged in various directions at the top of the experimental scene to collect scene photos, and the postures obtained by the 12 cameras are fused through image registration and other methods to obtain the three-dimensional human postures in the scene. The second method is to use a three-dimensional posture estimation method for a monocular camera, and use existing deep learning to estimate the three-dimensional postures of humans from a single image. In addition, there is also a method of estimating two-dimensional postures in the scene using a neural network and restoring the two-dimensional postures to three-dimensional postures through the scene distance information calculated by a binocular camera.
[0004] These three methods all have their own advantages and limitations. The posture acquisition method based on a motion capture system has high precision, but requires high hardware costs and is not easy to deploy. It may also involve complex registration problems among 12 cameras in various scenarios. The posture estimation method based on a monocular camera is simple and easy to deploy, and only relies on a neural network for estimation. Although the existing technology has been able to achieve good posture estimation performance, its monocular nature determines that there are large errors in the three-dimensional positions of humans in the world coordinate system for its posture estimation. The method based on a binocular camera and two-dimensional posture estimation has high positioning accuracy, but cannot well handle problems such as occlusion between posture key points.
[0005] The Chinese invention application with the application number 202310821946.3 discloses "A 3D Human Pose Estimation Method Based on Low-Channel Radar", which includes the following steps: collecting low-channel radar echo signals with human pose information; filtering out clutter; segmenting and performing fast Fourier transform on the radar echo signals after clutter filtering to obtain multiple range-Doppler images and combining them to obtain pose sub-signal samples; generating sample labels with the help of a Kinect depth camera, and further establishing a pose sub-signal sample data set; constructing a pose estimation model based on a convolutional neural network, and training and optimizing the pose estimation model using the data set; using the pose estimation model with optimized parameters to make predictions to obtain the joint coordinates of the human target, and finally generating a 3D human pose estimation image. Summary of the Invention
[0006] In view of the problems existing in the above-mentioned existing methods, the present invention provides a method and system for synchronously collecting multi-person postures of a through-wall radar and a binocular camera. The technical solution adopted by the present invention is as follows:
[0007] In the first aspect of the present invention, a method for synchronously collecting multi-person postures of a through-wall radar and a binocular camera is provided. The method includes:
[0008] Obtaining through-wall radar echo signals and binocular camera images respectively;
[0009] Preprocessing the radar echo signals to obtain a channel-range matrix formed by stacking one-dimensional range images of each transceiver channel of the radar;
[0010] Synthesizing the binocular camera images to obtain a synthesized image, and performing three-dimensional position calculation on the synthesized image to obtain the three-dimensional point cloud of all pixel points in the synthesized image;
[0011] Performing time matching on the channel-range matrix, the synthesized image, and the three-dimensional point cloud to obtain an initial data frame;
[0012] Performing YOLO three-dimensional pose extraction based on the synthesized image and the three-dimensional point cloud in the initial data frame to obtain YOLO three-dimensional pose information;
[0013] Performing Multi-HMR three-dimensional human pose estimation based on the synthesized image in the initial data frame to obtain Multi-HMR three-dimensional pose information;
[0014] Fusing the YOLO three-dimensional pose information and the Multi-HMR three-dimensional pose information to obtain a final pose label;
[0015] Performing spatial registration on the final pose label and the channel-range matrix to obtain a final data frame.
[0016] As a preferred solution, the method for preprocessing the radar echo signal to obtain the channel-distance matrix stacked by the one-dimensional range images of each transceiver channel of the radar includes:
[0017] Perform digital reception, quadrature mixing, windowing, inverse fast Fourier transform, and moving average static clutter suppression on the radar echo signal to obtain the channel-distance matrix stacked by the one-dimensional range images of each transceiver channel of the radar.
[0018] As a preferred solution, the method for time-matching the channel-distance matrix, the synthetic image, and the 3D point cloud to obtain the initial data frame includes:
[0019] Perform nearest neighbor matching on the channel-distance matrix, the synthetic image, and the 3D point cloud according to their respective timestamps during acquisition to obtain the initial data frame.
[0020] As a preferred solution, the method for extracting the YOLO 3D pose from the synthetic image and the 3D point cloud in the initial data frame to obtain the YOLO 3D pose information includes:
[0021] Use the pre-trained YOLO network to estimate the planar pose information of all people in the scene from the synthetic image, then extract the 3D pose positions from the 3D point cloud of the binocular camera according to the pixel positions in the planar pose information, and finally obtain the YOLO 3D pose information.
[0022] As a preferred solution, the method for fusing the YOLO 3D pose information and the Multi-HMR 3D pose information to obtain the final pose label includes:
[0023] Fuse the YOLO 3D pose information and the Multi-HMR 3D pose information through the preset Hungarian algorithm to obtain the final pose label.
[0024] As a preferred solution, the method for spatially registering the final pose label and the channel-distance matrix to obtain the final data frame includes:
[0025] Convert the 3D coordinates of the final pose label to the 3D coordinate system with the through-wall radar as the origin. The specific method is:
[0026] After fixing the positions of the through-wall radar and the binocular camera, place three non-collinear corner reflectors in the scene. Based on the 3D BP imaging algorithm, use the channel-distance matrix of the through-wall radar to obtain the 3D positions p1, p2, p3 of the three corner reflectors, and extract the 3D positions q1, q2, q3 of the three corner reflectors obtained by the binocular camera. Use the least squares method to find a rigid body transformation. The specific formula is q i = R * p i+t, where R is a rotation matrix and t is a translation vector, so as to perform spatial registration on the final pose label and the channel-distance matrix to obtain the final pose label.
[0027] The second aspect of the present invention provides a multi-person pose synchronous acquisition system for a through-wall radar and a binocular camera. The system includes:
[0028] A through-wall radar module, a binocular camera module, a first data preprocessing module, a second data preprocessing module, a time matching module, a first pose extraction module, a second pose extraction module, a pose fusion module, and a spatial registration module;
[0029] The through-wall radar module is used to obtain through-wall radar echo signals;
[0030] The binocular camera module is used to obtain binocular camera images;
[0031] The first data preprocessing module is used to preprocess the radar echo signals to obtain a channel-distance matrix formed by stacking one-dimensional range images of each transceiver channel of the radar;
[0032] The second data preprocessing module is used to synthesize the binocular camera images to obtain a synthesized image, and perform three-dimensional position calculation on the synthesized image to obtain three-dimensional point clouds of all pixel points in the synthesized image;
[0033] The time matching module is used to perform time matching on the channel-distance matrix, the synthesized image, and the three-dimensional point clouds to obtain an initial data frame;
[0034] The first pose extraction module is used to perform YOLO three-dimensional pose extraction based on the synthesized image and the three-dimensional point clouds in the initial data frame to obtain YOLO three-dimensional pose information;
[0035] The second pose extraction module is used to perform Multi-HMR three-dimensional human pose estimation based on the synthesized image in the initial data frame to obtain Multi-HMR three-dimensional pose information;
[0036] The pose fusion module is used to fuse the YOLO three-dimensional pose information and the Multi-HMR three-dimensional pose information to obtain a final pose label
[0037] The spatial registration module is used to perform spatial registration on the final pose label and the channel-distance matrix to obtain a final data frame.
[0038] As a preferred solution, the through-wall radar module is specifically a MIMO UWB low-frequency through-wall radar; the binocular camera module includes two cameras with the same parameters.
[0039] In a third aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for synchronously collecting multi-person postures of a through-wall radar and a binocular camera are implemented.
[0040] In a fourth aspect of the present invention, there is provided a computer device, including a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, the steps of the above-mentioned method for synchronously collecting multi-person postures of a through-wall radar and a binocular camera are implemented.
[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0042] In the present invention, a single binocular camera is used for posture acquisition. Aiming at the problems that the three-dimensional point cloud obtained by the binocular camera cannot handle the occlusion between posture key points, the insufficient kinematic constraints of the posture, and the low positioning accuracy of the three-dimensional posture estimation algorithm based on deep learning, a method of fusing posture extraction and estimation is adopted to improve the quality of three-dimensional human posture acquisition. Description of the Drawings
[0043] Figure 1 It is a flowchart of a method for synchronously collecting multi-person postures of a through-wall radar and a binocular camera provided in this embodiment. Detailed Embodiments
[0044] The drawings are only for illustrative purposes and cannot be construed as limiting the present invention;
[0045] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the embodiments of the present application.
[0046] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present application. The singular forms of "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0047] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0048] In addition, in the description of the present application, unless otherwise specified, "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0049] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0050] Embodiment 1
[0051] Please refer to Figure 1 , this embodiment provides a method for synchronous acquisition of multi-person postures by a through-wall radar and a binocular camera, and the method includes:
[0052] S1: Obtain the through-wall radar echo signal and the binocular camera image respectively;
[0053] S2: Preprocess the radar echo signal to obtain a channel-distance matrix formed by stacking one-dimensional range profiles of each transceiver channel of the radar;
[0054] In a specific embodiment, the method for preprocessing the radar echo signal to obtain a channel-distance matrix formed by stacking one-dimensional range profiles of each transceiver channel of the radar includes:
[0055] Perform digital reception, quadrature mixing, windowing, inverse fast Fourier transform, and sliding average static clutter suppression on the radar echo signal to obtain a channel-distance matrix formed by stacking one-dimensional range profiles of each transceiver channel of the radar.
[0056] S3: Synthesize the binocular camera images to obtain a synthesized image, and perform three-dimensional position calculation on the synthesized image to obtain the three-dimensional point cloud of all pixel points in the synthesized image;
[0057] S4: Perform temporal matching on the channel-distance matrix, synthetic image, and 3D point cloud to obtain an initial data frame;
[0058] In a specific embodiment, the method for performing temporal matching on the channel-distance matrix, synthetic image, and 3D point cloud to obtain an initial data frame includes:
[0059] Perform nearest neighbor matching on the channel-distance matrix, synthetic image, and 3D point cloud according to their respective timestamps during acquisition to obtain an initial data frame.
[0060] S5: Perform YOLO 3D pose extraction based on the synthetic image and 3D point cloud in the initial data frame to obtain YOLO 3D pose information;
[0061] In a specific embodiment, the method for performing YOLO 3D pose extraction based on the synthetic image and 3D point cloud in the initial data frame to obtain YOLO 3D pose information includes:
[0062] Use a pre-trained YOLO network to estimate the planar pose information of all people in the scene from the synthetic image, then extract the 3D pose positions from the 3D point cloud of the binocular camera according to the pixel positions in the planar pose information, and finally obtain the YOLO 3D pose information.
[0063] S6: Perform Multi-HMR 3D human pose estimation based on the synthetic image in the initial data frame to obtain Multi-HMR 3D pose information;
[0064] S7: Fuse the YOLO 3D pose information and the Multi-HMR 3D pose information to obtain a final pose label;
[0065] In a specific embodiment, the method for fusing the YOLO 3D pose information and the Multi-HMR 3D pose information to obtain a final pose label includes:
[0066] Fuse the YOLO 3D pose information and the Multi-HMR 3D pose information through a preset Hungarian algorithm to obtain a final pose label.
[0067] S8: Perform spatial registration on the final pose label and the channel-distance matrix to obtain a final data frame.
[0068] In a specific embodiment, the method for performing spatial registration on the final pose label and the channel-distance matrix to obtain a final data frame includes:
[0069] Convert the 3D coordinates of the final pose label to a 3D coordinate system with the through-wall radar as the origin. The specific method is:
[0070] After fixing the positions of the through-wall radar and the binocular camera, three non-collinear corner reflectors are placed in the scene. Based on the 3D BP imaging algorithm, the 3D positions p1, p2, and p3 of the three corner reflectors are obtained using the channel-distance matrix of the through-wall radar, and the 3D positions q1, q2, and q3 of the three corner reflectors obtained from the binocular camera are extracted. A rigid body transformation is found using the least squares method, and the specific formula is q i = R * p i + t, where R is the rotation matrix and t is the translation vector, thereby performing spatial registration on the final pose label and the channel-distance matrix to obtain the final pose label.
[0071] Embodiment 2
[0072] This embodiment provides a multi-person pose synchronous acquisition system for a through-wall radar and a binocular camera. The system includes:
[0073] A through-wall radar module, a binocular camera module, a first data preprocessing module, a second data preprocessing module, a time matching module, a first pose extraction module, a second pose extraction module, a pose fusion module, and a spatial registration module;
[0074] The through-wall radar module is used to obtain through-wall radar echo signals;
[0075] The binocular camera module is used to obtain binocular camera images;
[0076] The first data preprocessing module is used to preprocess the radar echo signals to obtain a channel-distance matrix formed by stacking one-dimensional range images of each transceiver channel of the radar;
[0077] The second data preprocessing module is used to synthesize the binocular camera images to obtain a synthesized image, and perform 3D position calculation on the synthesized image to obtain the 3D point cloud of all pixel points in the synthesized image;
[0078] The time matching module is used to perform time matching on the channel-distance matrix, the synthesized image, and the 3D point cloud to obtain an initial data frame;
[0079] The first pose extraction module is used to perform YOLO 3D pose extraction based on the synthesized image and the 3D point cloud in the initial data frame to obtain YOLO 3D pose information;
[0080] The second pose extraction module is used to perform Multi-HMR 3D human pose estimation based on the synthesized image in the initial data frame to obtain Multi-HMR 3D pose information;
[0081] The pose fusion module is used to fuse the YOLO three-dimensional pose information and the Multi-HMR three-dimensional pose information to obtain the final pose label
[0082] The spatial registration module is used to perform spatial registration on the final pose label and the channel-distance matrix to obtain the final data frame.
[0083] Embodiment 3
[0084] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for synchronously collecting multi-person poses of a through-wall radar and a binocular camera described in Embodiment 1 are implemented
[0085] Embodiment 4
[0086] This embodiment provides a computer device, including a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, the steps of the method for synchronously collecting multi-person poses of a through-wall radar and a binocular camera described in Embodiment 1 are implemented.
[0087] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for synchronous acquisition of multi-person postures by a through-wall radar and a binocular camera, characterized in that, The method includes: Obtaining through-wall radar echo signals and binocular camera images respectively; Preprocessing the radar echo signals to obtain a channel-distance matrix formed by stacking one-dimensional range images of each transceiver channel of the radar; Synthesizing the binocular camera images to obtain a synthesized image, performing three-dimensional position calculation on the synthesized image to obtain the three-dimensional point cloud of all pixel points in the synthesized image; Performing time matching on the channel-distance matrix, the synthesized image, and the three-dimensional point cloud to obtain an initial data frame; Performing YOLO three-dimensional pose extraction based on the synthesized image and the three-dimensional point cloud in the initial data frame to obtain YOLO three-dimensional pose information; Performing Multi-HMR three-dimensional human pose estimation based on the synthesized image in the initial data frame to obtain Multi-HMR three-dimensional pose information; Fusing the YOLO three-dimensional pose information and the Multi-HMR three-dimensional pose information to obtain a final pose label; Performing spatial registration on the final pose label and the channel-distance matrix to obtain a final data frame.
2. The multi-person pose synchronous acquisition method of a wall-penetrating radar and a binocular camera according to claim 1, characterized in that, The method for preprocessing the radar echo signals to obtain a channel-distance matrix formed by stacking one-dimensional range images of each transceiver channel of the radar includes: Performing digital reception, quadrature mixing, windowing, inverse fast Fourier transform, and sliding average static clutter suppression on the radar echo signals to obtain a channel-distance matrix formed by stacking one-dimensional range images of each transceiver channel of the radar.
3. A method for synchronous acquisition of multi-person postures by a through-wall radar and a binocular camera according to claim 1, characterized in that, The method for performing time matching on the channel-distance matrix, the synthesized image, and the three-dimensional point cloud to obtain an initial data frame includes: Performing nearest neighbor matching on the channel-distance matrix, the synthesized image, and the three-dimensional point cloud according to their respective timestamps during acquisition to obtain an initial data frame.
4. A method for synchronous acquisition of multi-person postures by a wall-penetrating radar and a binocular camera according to claim 1, characterized in that The method for performing YOLO three-dimensional pose extraction based on the synthesized image and the three-dimensional point cloud in the initial data frame to obtain YOLO three-dimensional pose information includes: Using a pre-trained YOLO network to estimate the planar pose information of all people in the scene from the synthesized image, then extracting the three-dimensional pose positions from the three-dimensional point cloud of the binocular camera according to the pixel positions in the planar pose information, and finally obtaining the YOLO three-dimensional pose information.
5. A method for synchronous acquisition of multi-person postures by a wall-penetrating radar and a binocular camera according to claim 1, characterized in that, The method for fusing the YOLO three-dimensional pose information and the Multi-HMR three-dimensional pose information to obtain a final pose label includes: Fusing the YOLO three-dimensional pose information and the Multi-HMR three-dimensional pose information through a preset Hungarian algorithm to obtain a final pose label.
6. A method for synchronous acquisition of multi-person postures by a wall-penetrating radar and a binocular camera according to claim 1, characterized in that The method for performing spatial registration on the final pose label and the channel-distance matrix to obtain a final data frame includes: Converting the three-dimensional coordinates of the final pose label to a three-dimensional coordinate system with the through-wall radar as the origin. The specific method is: After fixing the positions of the fixed through-wall radar and the binocular camera, three non-collinear corner reflectors are placed in the scene. Based on the 3D BP imaging algorithm, the 3D positions p1, p2, and p3 of the three corner reflectors are obtained using the channel-distance matrix of the through-wall radar, and the 3D positions q1, q2, and q3 of the three corner reflectors obtained by the binocular camera are extracted. A rigid body transformation is found using the least squares method, and the specific formula is q i = R * p i + t, where R is the rotation matrix and t is the translation vector, so as to perform spatial registration on the final attitude label and the channel-distance matrix to obtain the final attitude label.
7. A multi-person pose synchronous acquisition system for through-wall radar and binocular cameras, characterized in that, The system includes: A through-wall radar module, a binocular camera module, a first data preprocessing module, a second data preprocessing module, a time matching module, a first pose extraction module, a second pose extraction module, a pose fusion module, and a spatial registration module; The through-wall radar module is used to obtain through-wall radar echo signals; The binocular camera module is used to obtain binocular camera images; The first data preprocessing module is used to preprocess the radar echo signal to obtain a channel-distance matrix formed by stacking the one-dimensional range images of each transceiver channel of the radar; The second data preprocessing module is used to synthesize the binocular camera images to obtain a synthesized image, and perform three-dimensional position calculation on the synthesized image to obtain the three-dimensional point cloud of all pixel points in the synthesized image; The time matching module is used to perform time matching on the channel-distance matrix, the synthesized image, and the three-dimensional point cloud to obtain an initial data frame; The first pose extraction module is used to perform YOLO three-dimensional pose extraction based on the synthesized image and the three-dimensional point cloud in the initial data frame to obtain YOLO three-dimensional pose information; The second pose extraction module is used to perform Multi-HMR three-dimensional human pose estimation based on the synthesized image in the initial data frame to obtain Multi-HMR three-dimensional pose information; The pose fusion module is used to fuse the YOLO three-dimensional pose information and the Multi-HMR three-dimensional pose information to obtain a final pose label; The spatial registration module is used to perform spatial registration on the final pose label and the channel-distance matrix to obtain a final data frame.
8. A multi-person pose synchronous acquisition system for through-wall radar and binocular cameras according to claim 7, characterized in that, The through-wall radar module is specifically a MIMO UWB low-frequency through-wall radar; the binocular camera module includes two cameras with the same parameters.
9. A computer-readable storage medium storing a computer program thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of a method for synchronously collecting multi-person poses of a through-wall radar and a binocular camera according to any one of claims 1 to 6.
10. A computer device, characterized in that: It includes a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, it implements the steps of a method for synchronously collecting multi-person poses of a through-wall radar and a binocular camera according to any one of claims 1 to 6.
Citation Information
Patent Citations
Three-dimensional human body posture estimation method based on low-channel radar
CN117058228A