Vision-millimeter wave based multi-modal non-line-of-sight moving obstacle detection method
By using multimodal fusion technology of millimeter-wave radar and depth camera, the real-time and accuracy problems of obstacle detection in non-line-of-sight areas on mobile platforms are solved, and the presence and speed of people after turning corners are accurately detected. It is applicable to scenarios such as robot patrol, autonomous driving and security monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-05-20
- Publication Date
- 2026-06-19
AI Technical Summary
Existing non-line-of-sight sensing technologies struggle to detect the presence and speed of people behind corners in real time on mobile platforms. In particular, pure radar and pure vision-based solutions suffer from insufficient resolution, complex hardware, high power consumption, and susceptibility to changes in lighting conditions.
By employing a method that integrates millimeter-wave radar and a depth camera, the multipath reflection characteristics of millimeter-wave radar are used to perceive non-line-of-sight areas. Combined with the high-precision wall geometry information provided by the depth camera, point clouds are reconstructed through 3D beamforming and ray tracing. The vehicle speed is then obtained by combining the inertial measurement unit, thereby enabling obstacle detection in non-line-of-sight areas.
It enables real-time detection and precise localization of obstacles behind corners on a mobile platform, reducing hardware complexity and power consumption, and improving detection accuracy and motion speed estimation capabilities in non-line-of-sight areas.
Smart Images

Figure CN122239042A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of millimeter-wave radar sensing and multimodal fusion technology, specifically relating to a vision-millimeter-wave multimodal non-line-of-sight moving obstacle detection method. Background Technology
[0002] With the rapid development of mobile robots and autonomous driving technologies, higher demands are being placed on the perception capabilities of dynamic targets in complex environments. In scenarios with blind spots, such as corridors and intersections, how to detect the movement of people behind corners in advance has become a key issue in ensuring safe robot navigation. Existing non-line-of-sight perception technologies mainly have the following shortcomings:
[0003] While radar-based solutions, such as HoloRadar, can reconstruct the geometry of non-line-of-sight scenes, they are limited by the low resolution of millimeter-wave radar, making it difficult to accurately extract the motion speed information of people and to effectively distinguish between people and static objects. Furthermore, these solutions often employ radar self-rotation scanning, resulting in complex hardware, high power consumption, and difficulty in real-time deployment on mobile platforms.
[0004] Based on a purely visual approach, RGB-D cameras offer high resolution advantages within the line-of-sight range, but they cannot perceive non-line-of-sight areas beyond corners, resulting in inherent blind spots. Furthermore, visual solutions are susceptible to changes in lighting conditions, exhibiting significant performance degradation in low-light or high-light environments.
[0005] Existing multi-sensor fusion solutions often employ post-fusion strategies, failing to fully utilize the high-precision geometric information in the line-of-sight region to physically constrain radar multipath reflections, resulting in insufficient accuracy in localization and velocity measurement in non-line-of-sight regions. More importantly, existing methods primarily focus on target localization and scene reconstruction, lacking accurate estimation of personnel movement speed, making it difficult to meet the needs of dynamic obstacle avoidance and trajectory prediction.
[0006] Therefore, how to design a method suitable for mobile platforms that can detect the presence and speed of people behind corners in real time has become an urgent problem to be solved in the current technology field. Summary of the Invention
[0007] To address the aforementioned technical challenges, this application provides a vision-millimeter wave-based multimodal non-line-of-sight moving obstacle detection method. Using millimeter-wave radar and a depth camera as the core sensing technologies, the method leverages the multipath reflection characteristics of millimeter-wave radar to perceive non-line-of-sight regions and reconstruct point clouds. High-precision wall geometry information provided by the depth camera assists in calculating the reflection path of the millimeter-wave radar signal. An onboard inertial measurement unit acquires the vehicle's own speed, enabling the mobile platform to dynamically perceive the state of moving obstacles after corners, thus meeting the needs of a wide range of applications such as robot patrolling, autonomous driving, and security monitoring.
[0008] To achieve the above objectives, this application employs the following technical solution:
[0009] This application discloses a multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave, which specifically includes the following steps:
[0010] Step 1: Acquire multimodal data: Millimeter-wave radar acquires millimeter-wave radar signals in the corner direction, and depth camera simultaneously acquires RGB-D image data of the wall at line of sight;
[0011] Step 2: Multimodal data preprocessing: Range-Doppler Fast Fourier Transform (RFT) is performed on the millimeter-wave radar signals received by the antenna in the millimeter-wave radar. A radar heatmap in RAE format is then constructed using a 3D beamforming method. The ORB-SLAM3 algorithm is then used to process the RGB-D image data acquired in Step 1 to reconstruct a 3D mesh map of the line-of-sight area. And extract the wall normal vector And through PnP joint calibration, the radar coordinate system is aligned with the camera coordinate system;
[0012] Step 3, Non-line-of-sight path resolution: Radar heatmap based on RAE format 3D grid map of viewing distance area The coordinates of the reflection point of the millimeter-wave radar signal on the wall were calculated using ray tracing and the law of reflection. and the direction of the reflection path And reconstruct the complete point cloud of the non-line-of-sight region. ;
[0013] Step 4, Non-line-of-sight target separation: The motion direction of the moving obstacle is obtained through velocity compensation, and the complete point cloud of the non-line-of-sight region is separated by point cloud classification. Separate the point cloud of moving obstacles and extract the position coordinates of non-line-of-sight moving obstacles. Output the confidence level of the detection. ;
[0014] Step 5, Output: Output the position coordinates of the non-line-of-sight moving obstacle. Direction of motion and confidence level of detection .
[0015] A further improvement to this application is that step 1 specifically includes the following steps:
[0016] Step 1.1: Acquiring millimeter-wave radar data: The millimeter-wave radar transmits frequency-modulated continuous wave (FMCW) signals and receives millimeter-wave radar signals from both line-of-sight and non-line-of-sight areas. The millimeter-wave radar signals are the raw echo data.
[0017] Step 1.2: Acquire depth camera: The depth camera synchronously acquires RGB-D image data of the wall at the line of sight, which is used for subsequent 3D reconstruction and normal vector extraction.
[0018] A further improvement in this application is that the multimodal data preprocessing in step 2 specifically includes the following steps:
[0019] Step 2.1, Range-Doppler Fast Fourier Transform: Perform a range-Doppler Fast Fourier Transform on the millimeter-wave radar signal received by each antenna to resolve the range. The distance of the object on the surface to the signal after the Fast Fourier Transform is represented as:
[0020]
[0021] in, For the received time-domain millimeter-wave radar signal, For frequency, The imaginary unit, It is a natural constant. The integral symbol is used. For time Integral infinitesimal element, For time variables, Let be the kernel function for the Fourier transform, and then perform a Doppler Fast Fourier Transform to separate the object in the velocity dimension. The expression for the Doppler Fast Fourier Transform is:
[0022]
[0023] in, For Doppler frequency shift, The range-Doppler spectrum is obtained from the signal after the range-fast Fourier transform. , For distance, For speed;
[0024] Step 2.2, Constant False Alarm Rate Detection: Analyze the obtained range-Doppler spectrum... A constant false alarm rate (CFAR) detection algorithm is applied to identify echo amplitudes exceeding an adaptive threshold. The criteria for determining the object are:
[0025]
[0026] Adaptive threshold Based on background noise power calculate:
[0027]
[0028] Wherein, γ is the false alarm rate control factor;
[0029] Step 2.3, Doppler velocity extraction: For values exceeding the adaptive threshold Extract the Doppler frequency shift of the object. Calculate radial velocity :
[0030]
[0031] in, For radar wavelength, Radial velocity as measured by radar;
[0032] Simultaneously, the range-Doppler spectrum Micro-Doppler features are extracted by applying short-time Fourier transform. It is used to describe the subtle patterns of echo frequency changes over time caused by moving obstacles (such as the swinging of a person's limbs), and to help distinguish moving obstacles from static objects.
[0033] Step 2.4: Constructing a radar heatmap in RAE format: Using a three-dimensional beamforming method such as the Capon beamformer, the raw echo data from Step 1.1 is used to construct a radar heatmap in RAE format. ,in, For the distance dimension, the number of samples is... For the azimuth dimension, the number of samples is... The number of samples is for the pitch angle dimension;
[0034] Step 2.5, Visual 3D Reconstruction: The ORB-SLAM3 algorithm is used to process the RGB-D image data acquired in Step 1 to reconstruct a 3D mesh map of the viewing distance area in real time. And extract the wall normal vector. ,in For the normal vector in Components on the axis, For the normal vector in Components on the axis, For the normal vector in Components on the axis; Step 2.6, PnP Joint Calibration: The radar coordinate system is aligned with the camera coordinate system through PnP joint calibration.
[0035] A further improvement of this application is that step 2.6 specifically includes the following steps:
[0036] Step 2.6.1: Collect calibration board data: Place a joint calibration board within the common field of view of the radar and camera. The joint calibration board contains both corner reflectors and checkerboard patterns.
[0037] Step 2.6.2, Feature Point Extraction: The radar detects the corner reflector and obtains the three-dimensional point cloud coordinates in the radar coordinate system. The camera detects the corner points of the chessboard grid and obtains the two-dimensional pixel coordinates in the image coordinate system. ,in, The x-coordinate is the pixel coordinate. The vertical coordinate is the pixel coordinate. For the target point in the radar coordinate system Axis coordinates For the target point in the radar coordinate system Axis coordinates For the target point in the radar coordinate system Axis coordinates;
[0038] Step 2.6.3, Coordinate Transformation Solution: By solving the PnP problem, calculate the rotation matrix from the radar coordinate system to the camera coordinate system. Translation vector ,satisfy:
[0039]
[0040] in, These are the coordinates of a three-dimensional point in the radar coordinate system. These are two-dimensional pixel coordinates in the image coordinate system. It is a 3×3 rotation matrix. It is a 3×1 translation vector.
[0041] A further improvement in this application is that the non-line-of-sight path resolution described in step 3 specifically includes the following steps:
[0042] Step 3.1, Radar beam direction definition: For radar heatmaps in RAE format... Each pixel in Corresponding to a radar beam transmission direction ,in, For the first Azimuth angle For the first The elevation angle and the radar's position are determined by three-dimensional point coordinates. Give;
[0043] Step 3.2, Locating the wall reflection point: along the radar beam emission direction Perform ray tracing to calculate the ray and the 3D mesh map of the view distance area. The first intersection point, i.e., the coordinates of the reflection point on the wall. The coordinates of the reflection point satisfy:
[0044]
[0045] in, Radar to reflection point coordinates The distance;
[0046] Step 3.3: Calculate the reflection path direction: According to the law of specular reflection, the reflection path direction... From the direction of radar beam transmission and wall normal vector calculate:
[0047]
[0048] in, Let be the wall normal vector. Represents the dot product of two vectors;
[0049] Step 3.4, Non-line-of-sight range decomposition: Total flight distance measured by radar Includes radar to reflection point coordinates distance and distance from the reflection point to the moving obstacle :
[0050]
[0051] in, The distance from the reflection point to the moving obstacle. The coordinates of the moving obstacle's position;
[0052] Step 3.5: Construct a complete point cloud for the non-line-of-sight region: For each radar beam transmission direction According to radar heatmaps in RAE format The echo intensity is extracted along the radar beam transmission direction. All valid reflection points in the mirror space are mapped to real-world coordinates using the law of reflection, resulting in a complete point cloud of the non-line-of-sight region. :
[0053]
[0054] Each point contains three-dimensional spatial coordinates and corresponding echo intensity information.
[0055] A further improvement of this application is that the non-line-of-sight target separation in step 4 specifically includes the following steps:
[0056] Step 4.1, Vehicle speed acquisition: The vehicle's own speed is acquired by the inertial measurement unit. and direction of movement ;
[0057] Among them, speed The scalar value output by the inertial measurement unit indicates the direction of motion. This is the unit direction vector output by the inertial measurement unit (IMU). An IMU is a mature commercial sensor that integrates an accelerometer, gyroscope, and signal processing circuitry, and can directly output data such as velocity, direction, and angular velocity.
[0058] Step 4.2, Velocity Projection Calculation: Radial Velocity Measured by Radar Includes the projected velocity of the moving obstacle in the radar line of sight and the vehicle's own velocity. The contribution of the two, and their geometric relationship is as follows:
[0059]
[0060] in, Radial velocity measured by radar. The projected velocity of the moving obstacle in the radar line-of-sight direction. The direction of motion of the car relative to radar line of sight The included angle is calculated using the following formula:
[0061] ;
[0062] Step 4.3, Radial velocity extraction of moving obstacles: subtract the vehicle's own velocity. Then, the true radial velocity of the moving obstacle relative to the ground is obtained:
[0063] ;
[0064] Step 4.4, Calculation of the true velocity of the moving obstacle: Combining reflection geometry, the true velocity of the moving obstacle is calculated:
[0065] ;
[0066] in, The actual speed of the moving obstacle. The direction of movement of the moving obstacle and the direction of the reflection path The included angle is calculated using the following formula:
[0067] ;
[0068] Step 4.5, Determining the direction of movement: Based on the actual speed of the moving obstacle. The sign of the sign indicates the direction of motion:
[0069]
[0070] Step 4.6, Point Cloud Classification and Separation: A deep learning point cloud classification network is used to classify the complete point cloud in non-line-of-sight regions. Processing is performed, combined with microDoppler characteristics. Distinguish between moving obstacles and static objects, and separate the point cloud of moving obstacles. ;
[0071] Step 4.7, Location Extraction: From the point cloud of moving obstacles Extract the position coordinates of non-line-of-sight moving obstacles. Extract the location coordinates of the point cloud centroid or cluster center. Used for the output of step 5;
[0072] Step 4.8, Confidence Output: A three-dimensional convolutional neural network is used to output the confidence score of the radar heatmap in RAE format. Feature extraction is performed, and the detection confidence is output through a fully connected layer and a sigmoid function. :
[0073]
[0074] in, This is the logit value output by the fully connected layer in the classification branch.
[0075] A further improvement in this application is that the output in step 5 specifically includes the following steps:
[0076] Step 5.1, Position Output: Output the three-dimensional position coordinates of the non-line-of-sight moving obstacle in the world coordinate system. The unit is meters;
[0077] Step 5.2, Direction Output: Output the direction of movement of non-line-of-sight moving obstacles, with values of "approaching", "moving away" or "stationary";
[0078] Step 5.3, Confidence Output: Output the confidence score for non-line-of-sight detection. The range of values is This indicates the reliability of the test results.
[0079] A further improvement of this application is that the method further includes a model evaluation step, which uses a binary classification evaluation method to calculate the accuracy. Recall rate and Fractions are calculated using the following formula:
[0080]
[0081]
[0082]
[0083] in, To accurately detect the number of samples of moving obstacles, The number of false positives. This represents the number of underreported samples.
[0084] The beneficial effects of this application are:
[0085] This application employs three-dimensional beamforming to construct a RAE heatmap containing distance, azimuth, and elevation angles, providing complete spatial information for subsequent three-dimensional ray tracing and point cloud reconstruction.
[0086] This application transforms radar point clouds to the camera coordinate system through PnP joint calibration, and reconstructs non-line-of-sight point clouds by combining ray tracing and reflection laws, thereby achieving accurate positioning of moving obstacles in the world coordinate system.
[0087] This application uses speed compensation to obtain the true speed of the moving obstacle. Based on the positive or negative sign of the speed, it directly determines whether the moving obstacle is approaching or moving away. The output results are intuitive and the direction of movement is accurately determined.
[0088] The radar and camera in this application are fixedly mounted on a mobile vehicle, eliminating the need for self-rotation scanning. This results in simple hardware and low power consumption. The vehicle speed is directly obtained by the inertial measurement unit, requiring no additional estimation, thus demonstrating adaptability to mobile platforms.
[0089] This application utilizes wall geometry information provided by a depth camera to assist in radar reflection path calculation, thereby improving the accuracy of point cloud reconstruction and positioning in non-line-of-sight areas.
[0090] This application utilizes a 3D convolutional neural network and a point cloud classification network to automatically learn the feature differences between moving obstacles and static objects, eliminating the need for manual feature design and demonstrating strong generalization ability. Attached Figure Description
[0091] Figure 1 This is a framework diagram of this application.
[0092] Figure 2 This is a flowchart of this application.
[0093] Figure 3 This is the ORB-SLAM3 system architecture diagram.
[0094] Figure 4 This is a schematic diagram of PnP joint calibration.
[0095] Figure 5 This is a schematic diagram of the speed compensation geometry in this application.
[0096] Figure 6 It consists of an RGB image and a depth image captured by a camera.
[0097] Figure 7 It is a diagram of the camera's motion trajectory.
[0098] Figure 8 It is a 3D grid map of the viewing area. Detailed Implementation
[0099] The embodiments of the present invention will be disclosed below with reference to the drawings. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details are not intended to limit the present invention. That is, in some embodiments of the present invention, these practical details are not essential. In addition, for the sake of simplicity, some conventional structures and components will be shown in the drawings in a simple schematic manner.
[0100] like Figures 1-2 As shown, this application discloses a multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave, which specifically includes the following steps:
[0101] Step 1: Acquire Multimodal Data: Millimeter-wave radar acquires millimeter-wave radar signals in the corner direction, while a depth camera simultaneously acquires RGB-D image data of the wall at the line of sight. Acquiring multimodal data specifically includes the following steps:
[0102] Step 1.1: Acquiring millimeter-wave radar data: The millimeter-wave radar transmits frequency-modulated continuous wave (FMCW) signals, and the received millimeter-wave radar signals from both line-of-sight and non-line-of-sight areas are the raw echo data.
[0103] Step 1.2: Acquire depth camera: The depth camera synchronously acquires RGB-D image data of the wall at the line of sight, which is used for subsequent 3D reconstruction and normal vector extraction.
[0104] Step 2: Multimodal data preprocessing: Range-Doppler Fast Fourier Transform (RFT) is performed on the millimeter-wave radar signals received by the antenna, and a radar thermal map in RAE format is constructed using a 3D beamforming method. The ORB-SLAM3 algorithm is then used to process the RGB-D image data to reconstruct a 3D mesh map of the line-of-sight area. And extract the wall normal vector The radar coordinate system is then aligned to the camera coordinate system through PnP joint calibration. This includes the following steps:
[0105] Step 2.1, Range-Doppler Fast Fourier Transform: Perform a range-Doppler Fast Fourier Transform on the millimeter-wave radar signal received by each antenna to resolve the range. The distance of the object on the surface to the signal after the Fast Fourier Transform is represented as:
[0106]
[0107] in, For the received time-domain millimeter-wave radar signal, For frequency, The imaginary unit, It is a natural constant. The integral symbol is used. For time Integral infinitesimal element, For time variables, Let be the kernel function for the Fourier transform, and then perform a Doppler Fast Fourier Transform to separate the object in the velocity dimension. The expression for the Doppler Fast Fourier Transform is:
[0108]
[0109] in, For Doppler frequency shift, The range-Doppler spectrum is obtained from the signal after the range-fast Fourier transform. , For distance, For speed;
[0110] Step 2.2, Constant False Alarm Rate Detection: Analyze the obtained range-Doppler spectrum... A constant false alarm rate (CFAR) detection algorithm is applied to identify echo amplitudes exceeding an adaptive threshold. The criteria for determining the object are:
[0111]
[0112] Adaptive threshold Based on background noise power calculate:
[0113]
[0114] in, This is a control factor for the false alarm rate;
[0115] Step 2.3, Doppler velocity extraction: For values exceeding the adaptive threshold Extract the Doppler frequency shift of the object. Calculate radial velocity :
[0116]
[0117] in, For radar wavelength, Radial velocity as measured by radar;
[0118] Simultaneously, the range-Doppler spectrum Micro-Doppler features are extracted by applying short-time Fourier transform. It is used to describe the subtle patterns of echo frequency changes over time caused by moving obstacles (such as the swinging of a person's limbs), helping to distinguish moving obstacles from static objects.
[0119] Step 2.4: Constructing a radar heatmap in RAE format: Using a three-dimensional beamforming method such as the Capon beamformer, the raw radar echo data from Step 1.1 is used to construct a radar heatmap in RAE format. ,in, For the distance dimension, the number of samples is... For the azimuth dimension, the number of samples is... The number of samples is for the pitch angle dimension;
[0120] Step 2.5, Visual 3D Reconstruction: The ORB-SLAM3 algorithm is used to process the RGB-D image data, such as... Figure 3 As shown, the ORB-SLAM3 system processes visual and inertial data through three parallel threads. Input data includes image streams and inertial measurement unit (IMU) data. The image stream is denoted as... ,in, For the first Frame RGB image, For frame index, This represents the total number of frames in the image sequence, used to extract ORB feature points and provide visual observation information. The inertial measurement unit data is denoted as... ,in, For the first The three-axis acceleration vector corresponding to the frame, in meters per second squared (m²). ), For the first The three-axis angular velocity vector corresponding to the frame, in radians per second ( ( ), used to estimate relative motion between frames and assist visual tracking.
[0121] The visual tracking thread receives image streams and inertial measurement unit (IMU) data, extracts ORB feature points, performs initial pose estimation, and outputs the initial pose of the camera in the world coordinate system. ,in It is a 4×4 transformation matrix. This represents the 3D Euclidean transformation group. The local mapping thread performs local bundle adjustment on the keyframes output by the vision tracking thread to construct a local 3D map, outputting a 3D mesh map of the view distance region. The loop closure detection thread identifies visited scenes, performs global bundle adjustment to eliminate cumulative drift, and globally optimizes the map built by the local mapping thread. These three threads work collaboratively to output a real-time 3D mesh map of the viewport area. And the initial pose of the camera in the world coordinate system Simultaneously extract the wall normal vector ,in The normal vectors are respectively in Components in three directions.
[0122] To verify the actual effect of visual 3D reconstruction, RGB-D image sequences were collected in a typical corridor corner scene, and the ORB-SLAM3 algorithm was run.
[0123] Figure 6 RGB images and depth maps captured by the camera.
[0124] Figure 7 The image shows the camera's motion trajectory. The horizontal axis represents the displacement in the X direction (m), and the vertical axis represents the displacement in the Z direction (m). The total trajectory length is approximately 3.79m. The curve is smooth with no obvious jumps, indicating that SLAM tracking is stable and the positioning is drift-free.
[0125] Figure 8 This is a 3D mesh map of the reconstructed viewing area. The green point cloud in the image is concentrated in the plane range of X coordinate -0.9 to 1.0m and Z coordinate 0 to 0.1m. The point cloud has a small thickness and no discrete noise, indicating that the wall reconstruction is flat and has high accuracy.
[0126] The experimental results above show that the ORB-SLAM3 module can stably reconstruct the geometry of the wall at the line of sight and extract reliable normal vectors, meeting the requirements of this method for geometric information of the line of sight region.
[0127] Step 2.6, PnP Joint Calibration: The radar coordinate system is aligned to the camera coordinate system through PnP joint calibration, such as... Figure 4 As shown, the PnP joint calibration specifically includes the following steps:
[0128] Step 2.6.1: Collect calibration board data: Place a joint calibration board within the common field of view of the radar and camera. The joint calibration board contains both corner reflectors and checkerboard patterns.
[0129] Step 2.6.2, Feature Point Extraction: The radar detects the corner reflector and obtains the three-dimensional point cloud coordinates in the radar coordinate system. The camera detects the corner points of the chessboard grid and obtains the two-dimensional pixel coordinates in the image coordinate system. ,in, The x-coordinate is the pixel coordinate. The vertical coordinate is the pixel coordinate.
[0130] Step 2.6.3, Coordinate Transformation Solution: By solving the PnP problem, calculate the rotation matrix from the radar coordinate system to the camera coordinate system. Translation vector ,satisfy:
[0131]
[0132] in, These are the coordinates of a three-dimensional point in the radar coordinate system. These are two-dimensional pixel coordinates in the image coordinate system. It is a 3×3 rotation matrix. It is a 3×1 translation vector.
[0133] Step 3, Non-line-of-sight path resolution: Radar heatmap based on RAE format 3D grid map of viewing distance area The coordinates of the reflection point of the millimeter-wave radar signal on the wall were calculated using ray tracing and the law of reflection. and the direction of the reflection path And reconstruct the complete point cloud of the non-line-of-sight region. Non-line-of-sight path resolution specifically includes the following steps:
[0134] Step 3.1, Radar beam direction definition: For radar heatmaps in RAE format... Each pixel in Corresponding to a radar beam transmission direction ,in, For the first Azimuth angle For the first The elevation angle and the radar's position are determined by three-dimensional point coordinates. Give;
[0135] Step 3.2, Locating the wall reflection point: along the radar beam emission direction Perform ray tracing to calculate the ray and the 3D mesh map of the view distance area. The first intersection point, i.e., the coordinates of the reflection point on the wall. The coordinates of the reflection point satisfy:
[0136]
[0137] in, Radar to reflection point coordinates The distance;
[0138] Step 3.3: Calculate the reflection path direction: According to the law of specular reflection, the reflection path direction... From the direction of radar beam transmission and wall normal vector calculate:
[0139]
[0140] in, Let be the wall normal vector. Represents the dot product of two vectors;
[0141] Step 3.4, Non-line-of-sight range decomposition: Total flight distance measured by radar Includes radar to reflection point coordinates distance and distance from the reflection point to the moving obstacle :
[0142]
[0143] in, The distance from the reflection point to the moving obstacle. The coordinates of the moving obstacle's position;
[0144] Step 3.5: Construct a complete point cloud for the non-line-of-sight region: For each radar beam transmission direction According to radar heatmaps in RAE format The echo intensity is extracted along the radar beam transmission direction. All valid reflection points in the mirror space are mapped to real-world coordinates using the law of reflection, resulting in a complete point cloud of the non-line-of-sight region. :
[0145]
[0146] Each point contains three-dimensional spatial coordinates and corresponding echo intensity information.
[0147] Step 4, Non-line-of-sight target separation: The motion direction of the moving obstacle is obtained through velocity compensation, and the moving obstacle point cloud is separated from the non-line-of-sight point cloud and its position is extracted through point cloud classification. Output detection confidence level . Figure 5 The image shows the radial velocity measured by radar. With the speed of the car itself The actual speed of moving obstacles and included angle , The geometric relationship between them is used to help understand the calculation process of velocity compensation and motion direction determination. Specifically, the non-line-of-sight target separation includes the following steps:
[0148] Step 4.1, Vehicle speed acquisition: The vehicle's own speed is acquired by the inertial measurement unit. and direction of movement Among them, speed The scalar value output by the inertial measurement unit indicates the direction of motion. This is the unit direction vector output by the inertial measurement unit (IMU). An IMU is a mature commercial sensor that integrates an accelerometer, gyroscope, and signal processing circuitry, and can directly output data such as velocity, direction, and angular velocity.
[0149] Step 4.2, Velocity Projection Calculation: Radial Velocity Measured by Radar Includes the projected velocity of the moving obstacle in the radar line of sight and the vehicle's own velocity. The contribution of the two, and their geometric relationship is as follows:
[0150]
[0151] in, Radial velocity measured by radar. The projected velocity of the moving obstacle in the radar line-of-sight direction. The direction of motion of the car relative to radar line of sight The included angle is calculated using the following formula:
[0152] ;
[0153] Step 4.3, Radial velocity extraction of moving obstacles: subtract the vehicle's own velocity. Then, the true radial velocity of the moving obstacle relative to the ground is obtained:
[0154] ;
[0155] Step 4.4, Calculation of the true velocity of the moving obstacle: Combining reflection geometry, the true velocity of the moving obstacle is calculated:
[0156] ;
[0157] in, The actual speed of the moving obstacle. The direction of movement of the moving obstacle and the direction of the reflection path The included angle is calculated using the following formula:
[0158] ;
[0159] Step 4.5, Determining the direction of movement: Based on the actual speed of the moving obstacle. The sign of the sign indicates the direction of motion:
[0160]
[0161] Step 4.6, Point Cloud Classification and Separation: A deep learning point cloud classification network is used to classify the complete point cloud in non-line-of-sight regions. Processing is performed, combined with microDoppler characteristics. Distinguish between moving obstacles and static objects, and separate the point cloud of moving obstacles. ;
[0162] Step 4.7, Location Extraction: From the point cloud of moving obstacles Extract the position coordinates of non-line-of-sight moving obstacles. Extract the location coordinates of the point cloud centroid or cluster center. Used for the result output in step 5.
[0163] Step 4.8, Confidence Output: A three-dimensional convolutional neural network is used to output the confidence score of the radar heatmap in RAE format. Feature extraction is performed, and the detection confidence is output through a fully connected layer and a sigmoid function. :
[0164]
[0165] in, The logit value output by the fully connected layer in the classification branch. This represents the probability of the presence of a moving obstacle.
[0166] Step 5, Output: Output the position coordinates of the non-line-of-sight moving obstacle. Direction of motion and detection confidence Specifically, it includes the following steps:
[0167] Step 5.1, Position Output: Output the three-dimensional position coordinates of the non-line-of-sight moving obstacle in the world coordinate system. The unit is meters;
[0168] Step 5.2, Direction Output: Output the direction of movement of non-line-of-sight moving obstacles, with values of "approaching", "moving away" or "stationary";
[0169] Step 5.3, Confidence Output: Output the confidence level of the detection. The range of values is This indicates the reliability of the test results.
[0170] This application also employs a binary classification evaluation method to calculate the accuracy. Recall rate and Fractions are calculated using the following formula:
[0171]
[0172]
[0173]
[0174] in, To accurately detect the number of samples of moving obstacles, The number of false positives. This represents the number of underreported samples.
[0175] This application achieves real-time presence detection, precise positioning, and direction determination of moving obstacles behind corners on a mobile platform through multimodal fusion and physically guided speed compensation.
[0176] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave, characterized in that: The multimodal non-line-of-sight moving obstacle detection method specifically includes the following steps: Step 1: Acquire multimodal data: Millimeter-wave radar acquires millimeter-wave radar signals in the corner direction, and depth camera simultaneously acquires RGB-D image data of the wall at line of sight; Step 2, Multimodal Data Preprocessing: Perform range-Doppler fast Fourier transform on the millimeter-wave radar signals received by the antenna in the millimeter-wave radar, and construct a radar heat map in RAE format using a three-dimensional beamforming method. The RGB-D image data acquired in step 1 is processed to reconstruct a 3D mesh map of the viewing distance area. And extract the wall normal vector And through PnP joint calibration, the radar coordinate system is aligned with the camera coordinate system; Step 3, Non-line-of-sight path resolution: Radar heatmap based on RAE format 3D grid map of viewing distance area Calculate the coordinates of the reflection point of the millimeter-wave radar signal on the wall. and the direction of the reflection path And reconstruct the complete point cloud of the non-line-of-sight region. ; Step 4, Non-line-of-sight target separation: The motion direction of the moving obstacle is obtained through velocity compensation, and the complete point cloud of the non-line-of-sight region is separated by point cloud classification. Separate the point cloud of moving obstacles and extract the position coordinates of non-line-of-sight moving obstacles. Output the confidence level of the detection. ; Step 5, Output: Output the position coordinates of the non-line-of-sight moving obstacle. Direction of motion and confidence level of detection .
2. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 1, characterized in that: Step 1 specifically includes the following steps: Step 1.1: Acquiring millimeter-wave radar data: The millimeter-wave radar transmits frequency-modulated continuous wave signals and receives millimeter-wave radar signals from both line-of-sight and non-line-of-sight areas, i.e., the raw echo data. Step 1.2: Acquire depth camera: The depth camera synchronously acquires RGB-D image data of the wall at the line of sight, which is used for subsequent 3D reconstruction and normal vector extraction.
3. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 2, characterized in that: Step 2 specifically includes the following steps: Step 2.1, Range-Doppler Fast Fourier Transform: Perform a range-Doppler Fast Fourier Transform on the millimeter-wave radar signal received by each antenna to resolve the range. The distance of the object on the surface to the signal after the Fast Fourier Transform is represented as: in, For the received time-domain millimeter-wave radar signal, For frequency, The imaginary unit, It is a natural constant. The integral symbol is used. For time Integral infinitesimal element, For time variables, Let be the kernel function for the Fourier transform, and then perform a Doppler Fast Fourier Transform to separate the object in the velocity dimension. The expression for the Doppler Fast Fourier Transform is: in, For Doppler frequency shift, The range-Doppler spectrum is obtained from the signal after the range-fast Fourier transform. , For distance, For speed; Step 2.2, Constant False Alarm Rate Detection: Analyze the obtained range-Doppler spectrum... A constant false alarm rate (CFAR) detection algorithm is applied to identify echo amplitudes exceeding an adaptive threshold. The criteria for determining the object are: Adaptive threshold Based on background noise power calculate: in, This is the false alarm rate control factor; Step 2.3, Doppler velocity extraction: For values exceeding the adaptive threshold Extract the Doppler frequency shift of the object. Calculate radial velocity : in, For radar wavelength, Radial velocity as measured by radar; Simultaneously, the range-Doppler spectrum Micro-Doppler features are extracted by applying short-time Fourier transform. ; Step 2.4: Constructing a radar heatmap in RAE format: Using a three-dimensional beamforming method, the raw echo data obtained in Step 1.1 is used to construct a radar heatmap in RAE format. ,in, For the distance dimension, the number of samples is... For the azimuth dimension, the number of samples is... The number of samples is for the pitch angle dimension; Step 2.5, Visual 3D Reconstruction: The ORB-SLAM3 algorithm is used to process the RGB-D image data acquired in Step 1 to reconstruct a 3D mesh map of the viewing distance area in real time. And extract the wall normal vector. , For the normal vector in Components on the axis, For the normal vector in Components on the axis, For the normal vector in Components on the axis; Step 2.6, PnP Joint Calibration: The radar coordinate system is aligned with the camera coordinate system through PnP joint calibration.
4. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 3, characterized in that: Step 2.6 specifically includes the following steps: Step 2.6.1: Collect calibration board data: Place a joint calibration board within the common field of view of the radar and camera. The joint calibration board contains both corner reflectors and checkerboard patterns. Step 2.6.2, Feature Point Extraction: The radar detects the corner reflector and obtains the three-dimensional point cloud coordinates in the radar coordinate system. The camera detects the corner points of the chessboard grid and obtains the two-dimensional pixel coordinates in the image coordinate system. ,in, The x-coordinate is the pixel coordinate. The vertical coordinate is the pixel coordinate. For the target point in the radar coordinate system Axis coordinates For the target point in the radar coordinate system Axis coordinates For the target point in the radar coordinate system Axis coordinates; Step 2.6.3, Coordinate Transformation Solution: By solving the PnP problem, calculate the rotation matrix from the radar coordinate system to the camera coordinate system. Translation vector ,satisfy: in, The coordinates of a three-dimensional point in the radar coordinate system. These are two-dimensional pixel coordinates in the image coordinate system. It is a 3×3 rotation matrix. It is a 3×1 translation vector.
5. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 4, characterized in that: Step 3, the non-line-of-sight path resolution, specifically includes the following steps: Step 3.1, Radar beam direction definition: For radar heatmaps in RAE format... Each pixel in This corresponds to a radar beam transmission direction. ,in, For the first Azimuth angle For the first The elevation angle and the radar's position are determined by three-dimensional point coordinates. Give; Step 3.2, Locating the wall reflection point: along the radar beam emission direction Perform ray tracing to calculate the ray and the 3D mesh map of the view distance area. The first intersection point, i.e., the coordinates of the reflection point on the wall. The coordinates of the reflection point satisfy: in, Radar to reflection point coordinates The distance; Step 3.3: Calculate the reflection path direction: Reflection path direction From the direction of radar beam transmission and wall normal vector calculate: in, Let be the wall normal vector. Represents the dot product of two vectors; Step 3.4, Non-line-of-sight range decomposition: Total flight distance measured by radar Includes radar to reflection point coordinates distance and the distance from the reflection point to the moving obstacle : in, The distance from the reflection point to the moving obstacle. The coordinates of the moving obstacle's position; Step 3.5: Construct a complete point cloud for the non-line-of-sight region: For each radar beam transmission direction According to radar heatmaps in RAE format The echo intensity is extracted along the radar beam transmission direction. All valid reflection points in the mirror space are mapped to real-world coordinates using the law of reflection, resulting in a complete point cloud of the non-line-of-sight region. : Each point contains three-dimensional spatial coordinates and corresponding echo intensity information.
6. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 5, characterized in that: Step 4, the non-line-of-sight target separation, specifically includes the following steps: Step 4.1, Obtaining the vehicle's speed: Obtain the vehicle's own speed. and direction of movement ; Step 4.2, Velocity Projection Calculation: Radial Velocity Measured by Radar Includes the projected velocity of the moving obstacle in the radar line of sight and the vehicle's own velocity. The contribution of the two, and their geometric relationship is as follows: in, Radial velocity measured by radar. The projected velocity of the moving obstacle in the radar line-of-sight direction. The direction of motion of the car relative to radar line of sight The included angle is calculated using the following formula: ; Step 4.3, Radial velocity extraction of moving obstacles: subtract the vehicle's own velocity. Then, the true radial velocity of the moving obstacle relative to the ground is obtained. : ; Step 4.4, Calculation of the true velocity of the moving obstacle: Combining reflection geometry, the true velocity of the moving obstacle is calculated: ; in, The actual speed of the moving obstacle. The direction of movement of the moving obstacle and the direction of the reflection path The included angle is calculated using the following formula: ; Step 4.5, Determining the direction of movement: Based on the actual speed of the moving obstacle. The sign of the sign indicates the direction of motion: Step 4.6, Point Cloud Classification and Separation: A deep learning point cloud classification network is used to classify the complete point cloud in non-line-of-sight regions. The data is processed and combined with the extracted micro-Doppler features. Distinguish between moving obstacles and static objects, and separate the point cloud of moving obstacles. ; Step 4.7, Location Extraction: From the point cloud of moving obstacles Extract the position coordinates of non-line-of-sight moving obstacles. Take the centroid or cluster center of the cloud; Step 4.8, Confidence Output: A three-dimensional convolutional neural network is used to output the confidence score of the radar heatmap in RAE format. Feature extraction is performed, and the detection confidence is output through a fully connected layer and a sigmoid function. : in, This is the logit value output by the fully connected layer in the classification branch.
7. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 1, characterized in that: The output described in step 5 specifically includes the following steps: Step 5.1, Position Output: Output the three-dimensional position coordinates of the non-line-of-sight moving obstacle in the world coordinate system. ; Step 5.2, Direction Output: Output the direction of movement of non-line-of-sight moving obstacles, with values of "approaching", "moving away" or "stationary"; Step 5.3, Confidence Output: Output the confidence level of the detection. The range of values is This indicates the reliability of the test results.