Vision-millimeter wave based multi-modal non-line-of-sight moving obstacle detection method
Patent Information
- Application Number
- CN202610694291.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-05-20
AI Technical Summary
[0003]基于纯雷达的方案,如HoloRadar等系统虽能重建非视距场景几何结构,但受限于毫米波雷达的低分辨率,难以准确提取人员的运动速度信息,且无法有效区分人员与静态物体
[0083] This application employs three-dimensional beamforming to construct a RAE heatmap containing distance, azimuth, and elevation angles, providing complete spatial information for subsequent three-dimensional ray tracing and point cloud reconstruction.
Smart Images

Figure CN122239042B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of millimeter-wave radar sensing and multimodal fusion technology, specifically relating to a vision-millimeter-wave multimodal non-line-of-sight moving obstacle detection method. Background Technology
[0002] With the rapid development of mobile robots and autonomous driving technologies, higher demands are being placed on the perception capabilities of dynamic targets in complex environments. In scenarios with blind spots, such as corridors and intersections, how to detect the movement of people behind corners in advance has become a key issue in ensuring safe robot navigation. Existing non-line-of-sight perception technologies mainly have the following shortcomings:
[0003] While radar-based solutions, such as HoloRadar, can reconstruct the geometry of non-line-of-sight scenes, they are limited by the low resolution of millimeter-wave radar, making it difficult to accurately extract the motion speed information of people and to effectively distinguish between people and static objects. Furthermore, these solutions often employ radar self-rotation scanning, resulting in complex hardware, high power consumption, and difficulty in real-time deployment on mobile platforms.
[0004] Based on a purely visual approach, RGB-D cameras offer high resolution advantages within the line-of-sight range, but they cannot perceive non-line-of-sight areas beyond corners, resulting in inherent blind spots. Furthermore, visual solutions are susceptible to changes in lighting conditions, exhibiting significant performance degradation in low-light or high-light environments.
[0005] Existing multi-sensor fusion solutions often employ post-fusion strategies, failing to fully utilize the high-precision geometric information in the line-of-sight region to physically constrain radar multipath reflections, resulting in insufficient accuracy in localization and velocity measurement in non-line-of-sight regions. More importantly, existing methods primarily focus on target localization and scene reconstruction, lacking accurate estimation of personnel movement speed, making it difficult to meet the needs of dynamic obstacle avoidance and trajectory prediction.
[0006] Therefore, how to design a method suitable for mobile platforms that can detect the presence and speed of people behind corners in real time has become an urgent problem to be solved in the current technology field. Summary of the Invention
[0007] To address the aforementioned technical challenges, this application provides a vision-millimeter wave-based multimodal non-line-of-sight moving obstacle detection method. Using millimeter-wave radar and a depth camera as the core sensing technologies, the method leverages the multipath reflection characteristics of millimeter-wave radar to perceive non-line-of-sight regions and reconstruct point clouds. High-precision wall geometry information provided by the depth camera assists in calculating the reflection path of the millimeter-wave radar signal. Furthermore, the method utilizes visual-inertial fusion of RGB-D images and data from the vehicle's inertial measurement unit to estimate the vehicle's motion state. This enables the mobile platform to dynamically perceive the state of moving obstacles after corners, thus meeting the needs of a wide range of applications such as robot patrolling, autonomous driving, and security monitoring.
[0008] To achieve the above objectives, this application employs the following technical solution:
[0009] This application discloses a multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave, which specifically includes the following steps:
[0010] Step 1: Acquire multimodal data: Millimeter-wave radar acquires millimeter-wave radar signals in the corner direction, and depth camera simultaneously acquires RGB-D image data of the wall at line of sight;
[0011] Step 2: Multimodal data preprocessing: Range-Doppler Fast Fourier Transform (RFT) is performed on the millimeter-wave radar signals received by the antenna in the millimeter-wave radar. A radar heatmap in RAE format is then constructed using a 3D beamforming method. The ORB-SLAM3 algorithm is then used to process the RGB-D image data acquired in Step 1 to reconstruct a 3D mesh map of the line-of-sight area. And extract the wall normal vector And through PnP joint calibration, the radar coordinate system is aligned with the camera coordinate system;
[0012] Step 3, Non-line-of-sight path analysis: Radar heatmap based on RAE format 3D grid map of viewing distance area The coordinates of the reflection point of the millimeter-wave radar signal on the wall were calculated using ray tracing and the law of reflection. and the direction of the reflection path And reconstruct the complete point cloud of the non-line-of-sight region. ;
[0013] Step 4, Non-line-of-sight target separation: The motion direction of the moving obstacle is obtained through velocity compensation, and the complete point cloud of the non-line-of-sight region is separated by point cloud classification. Separate the point cloud of moving obstacles and extract the position coordinates of non-line-of-sight moving obstacles. Output the confidence level of the detection. ;
[0014] Step 5, Output: Output the position coordinates of the non-line-of-sight moving obstacle. Direction of motion and confidence level of detection .
[0015] A further improvement to this application is that step 1 specifically includes the following steps:
[0016] Step 1.1: Acquiring millimeter-wave radar data: The millimeter-wave radar transmits frequency-modulated continuous wave (FMCW) signals and receives millimeter-wave radar signals from both line-of-sight and non-line-of-sight areas. The millimeter-wave radar signals are the raw echo data.
[0017] Step 1.2: Acquire depth camera data: The depth camera simultaneously acquires RGB-D image data of the wall at the line of sight, which will be used for subsequent 3D reconstruction and normal vector extraction.
[0018] A further improvement in this application is that the multimodal data preprocessing in step 2 specifically includes the following steps:
[0019] Step 2.1, Range-Doppler Fast Fourier Transform: Perform a range-Doppler Fast Fourier Transform on the millimeter-wave radar signal received by each antenna to resolve the range. The distance of the object on the surface to the signal after the Fast Fourier Transform is represented as:
[0020]
[0021] in, For the received time-domain millimeter-wave radar signal, For frequency, The imaginary unit, It is a natural constant. The integral symbol is used. For time Integral infinitesimal element, For time variables, Let be the kernel function for the Fourier transform, and then perform a Doppler Fast Fourier Transform to separate the object in the velocity dimension. The expression for the Doppler Fast Fourier Transform is:
[0022]
[0023] in, For Doppler frequency shift, The range-Doppler spectrum is obtained from the signal after the range-fast Fourier transform. , For distance, For speed;
[0024] Step 2.2, Constant False Alarm Rate Detection: Analyze the obtained range-Doppler spectrum... A constant false alarm rate (CFAR) detection algorithm is applied to identify echo amplitudes exceeding an adaptive threshold. The criteria for determining the object are:
[0025]
[0026] Adaptive threshold Based on background noise power calculate:
[0027]
[0028] Wherein, γ is the false alarm rate control factor;
[0029] Step 2.3, Doppler velocity extraction: For values exceeding the adaptive threshold Extract the Doppler frequency shift of the object. Calculate radial velocity :
[0030]
[0031] in, For radar wavelength, Radial velocity as measured by radar;
[0032] Simultaneously, the range-Doppler spectrum Micro-Doppler features are extracted by applying short-time Fourier transform. It is used to describe the subtle patterns of echo frequency changes over time caused by moving obstacles (such as the swinging of a person's limbs), and to help distinguish moving obstacles from static objects.
[0033] Step 2.4: Constructing a radar heatmap in RAE format: Using a three-dimensional beamforming method such as the Capon beamformer, the raw echo data from Step 1.1 is used to construct a radar heatmap in RAE format. ,in, For the distance dimension, the number of samples is... For the azimuth dimension, the number of samples is... The number of samples is for the pitch angle dimension;
[0034] Step 2.5, Visual 3D Reconstruction: The ORB-SLAM3 algorithm is used to process the RGB-D image data acquired in Step 1 to reconstruct a 3D mesh map of the viewing distance area in real time. And extract the wall normal vector. ,in For the normal vector in Components on the axis, For the normal vector in Components on the axis For the normal vector in Components on the axis; Step 2.6, PnP joint calibration: The radar coordinate system is aligned with the camera coordinate system through PnP joint calibration.
[0035] Step 2.6.1: Collect calibration board data: Place a joint calibration board within the common field of view of the radar and camera. The joint calibration board contains both corner reflectors and checkerboard patterns.
[0036] Step 2.6.2, Feature Point Extraction: Pre-calibrate the camera intrinsic parameter matrix. The system calculates lens distortion parameters and corrects distortion at the checkerboard corner points detected by the camera; the radar detects the corner reflectors to obtain the three-dimensional point cloud coordinates in the radar coordinate system. The camera detects the corner points of the chessboard grid and obtains the two-dimensional pixel coordinates in the image coordinate system. ,in The x-coordinate is the pixel coordinate. The vertical coordinate is the pixel coordinate. , and These are the x-axis, y-axis, and z-axis coordinates of the target point in the radar coordinate system, respectively.
[0037] Step 2.6.3, Coordinate Transformation and Projection Solution: By solving the PnP problem, calculate the rotation matrix from the radar coordinate system to the camera coordinate system. Translation vector The radar's 3D points are first transformed into the camera's 3D coordinate system:
[0038]
[0039] Then use the camera intrinsic parameter matrix Mapping perspective projection relationship to two-dimensional pixel coordinates :
[0040]
[0041] in, These are the coordinates of a three-dimensional point in the radar coordinate system. These are the coordinates of a point in the camera's 3D coordinate system after extrinsic parameter transformation. These are two-dimensional pixel coordinates in the image coordinate system. It is a 3×3 rotation matrix. It is a 3×1 translation vector. For the camera intrinsic parameter matrix, This is the homogeneous coordinate scale factor. Thus, the rigid body transformation from radar 3D coordinates to camera 3D coordinates and the perspective projection from camera 3D coordinates to 2D pixel coordinates are completed sequentially.
[0042] A further improvement in this application is that the non-line-of-sight path resolution described in step 3 specifically includes the following steps:
[0043] Step 3.1, Radar beam direction definition: For radar heatmaps in RAE format... Each pixel in This corresponds to a radar beam transmission direction. ,in, For the first Azimuth angle For the first The elevation angle and the radar's position are determined by three-dimensional point coordinates. Give;
[0044] Step 3.2, Locating the wall reflection point: along the radar beam emission direction Perform ray tracing to calculate the ray and the 3D mesh map of the view distance area. The first intersection point, i.e., the coordinates of the reflection point on the wall. The coordinates of the reflection point satisfy:
[0045]
[0046] in, Radar to reflection point coordinates The distance;
[0047] Step 3.3: Calculate the reflection path direction: According to the law of specular reflection, the reflection path direction... From the direction of radar beam transmission and wall normal vector calculate:
[0048]
[0049] in, Let be the wall normal vector. Represents the dot product of two vectors;
[0050] Step 3.4, Non-line-of-sight range decomposition: Total flight distance measured by radar Includes radar to reflection point coordinates distance and distance from the reflection point to the moving obstacle :
[0051]
[0052] in, The distance from the reflection point to the moving obstacle. The coordinates of the moving obstacle's position;
[0053] Step 3.5: Construct a complete point cloud for the non-line-of-sight region: For each radar beam transmission direction According to radar heatmaps in RAE format The echo intensity is extracted along the radar beam transmission direction. All valid reflection points in the mirror space are mapped to real-world coordinates using the law of reflection, resulting in a complete point cloud of the non-line-of-sight region. :
[0054]
[0055] Each point contains three-dimensional spatial coordinates and corresponding echo intensity information.
[0056] A further improvement of this application is that the non-line-of-sight target separation in step 4 specifically includes the following steps:
[0057] Step 4.1, Vehicle Velocity Acquisition: The inertial measurement unit outputs triaxial acceleration and triaxial angular velocity. Combined with the RGB-D image acquired in Step 1 and the camera pose obtained from ORB-SLAM3, visual-inertial fusion is performed to estimate the vehicle's own velocity. and direction of movement ; The unit direction vector represents the integral drift of the vehicle's motion state, which is constrained by RGB-D visual observation and inertial measurement.
[0058] Step 4.2, Velocity Projection Calculation: Radial Velocity Measured by Radar Includes the projected velocity of the moving obstacle in the radar line of sight and the vehicle's own velocity. The contribution of the two, and their geometric relationship is as follows:
[0059]
[0060] in, Radial velocity measured by radar. The projected velocity of the moving obstacle in the radar line-of-sight direction. The direction of motion of the car relative to radar line of sight The included angle is calculated using the following formula:
[0061] ;
[0062] Step 4.3, Radial velocity extraction of moving obstacles: subtract the vehicle's own velocity. Then, the true radial velocity of the moving obstacle relative to the ground is obtained:
[0063] ;
[0064] Step 4.4, Calculation of collision-related directional velocity components: (e.g.) Figure 5 As shown, The velocity component of the moving obstacle in the direction in which it may collide with the car. Direction and radial velocity of the reflection path The angle between the directions. The depth camera obtains the wall orientation and wall normal vector through 3D reconstruction of the line-of-sight area. And the geometry of the corridor corners, combined with the radar echo direction and the direction of the car's movement. Before speed conversion, according to Figure 5 The geometric relationship shown determines the included angle. , Direction of movement of the car The angle between the radar line of sight and the radar line of sight; for Figure 5 The known geometric angles shown are primarily determined by the geometry of the walls and corridor corners recovered by the depth camera, and calculated in conjunction with the radar echo direction. The actual radial velocity used to compensate for the trolley's self-motion is also considered. Convert the velocity components of the moving obstacle in the collision-related direction. :
[0065]
[0066] Step 4.5, Determining the direction of motion: Based on the velocity components of the moving obstacle in the collision-related direction. The sign of the sign indicates the direction of motion:
[0067]
[0068] Step 4.6, Point Cloud Classification and Separation: A deep learning point cloud classification network is used to classify the complete point cloud in non-line-of-sight regions. Processing is performed, combined with microDoppler characteristics. Distinguish between moving obstacles and static objects, and separate the point cloud of moving obstacles. ;
[0069] Step 4.7, Location Extraction: From the point cloud of moving obstacles Extract the position coordinates of non-line-of-sight moving obstacles. Extract the location coordinates of the point cloud centroid or cluster center. Used for the result output in step 5;
[0070] Step 4.8, Confidence Output: A three-dimensional convolutional neural network is used to output the confidence score of the radar heatmap in RAE format. Feature extraction is performed, and the detection confidence is output through a fully connected layer and a sigmoid function. :
[0071]
[0072] in, This is the logit value output by the fully connected layer in the classification branch.
[0073] A further improvement in this application is that the output in step 5 specifically includes the following steps:
[0074] Step 5.1, Position Output: Output the three-dimensional position coordinates of the non-line-of-sight moving obstacle in the world coordinate system. The unit is meters;
[0075] Step 5.2, Direction Output: Output the direction of movement of non-line-of-sight moving obstacles, with values of "approaching", "moving away" or "stationary";
[0076] Step 5.3, Confidence Output: Output the confidence score for non-line-of-sight detection. The range of values is This indicates the reliability of the test results.
[0077] A further improvement of this application is that the method further includes a model evaluation step, which uses a binary classification evaluation method to calculate the accuracy. Recall rate and Fractions are calculated using the following formula:
[0078]
[0079]
[0080]
[0081] in, To accurately detect the number of samples of moving obstacles, The number of false positives. This represents the number of underreported samples.
[0082] The beneficial effects of this application are:
[0083] This application employs three-dimensional beamforming to construct a RAE heatmap containing distance, azimuth, and elevation angles, providing complete spatial information for subsequent three-dimensional ray tracing and point cloud reconstruction.
[0084] This application transforms radar point clouds to the camera coordinate system through PnP joint calibration, and reconstructs non-line-of-sight point clouds by combining ray tracing and reflection laws, thereby achieving accurate positioning of moving obstacles in the world coordinate system.
[0085] This application obtains the velocity component of the moving obstacle in the collision-related direction through velocity compensation, and determines whether the moving obstacle is approaching or moving away based on the positive or negative sign of the velocity component. The output result is intuitive and can reflect the motion state directly related to the collision risk of the vehicle.
[0086] The radar and camera in this application are fixedly mounted on a mobile vehicle, eliminating the need for self-rotation scanning. The hardware is simple and has low power consumption. The vehicle's speed and direction of motion are estimated by visual-inertial fusion of RGB-D images and inertial measurement unit data, demonstrating the adaptability to mobile platforms.
[0087] This application utilizes wall geometry information provided by a depth camera to assist in radar reflection path calculation, thereby improving the accuracy of point cloud reconstruction and positioning in non-line-of-sight areas.
[0088] This application utilizes a 3D convolutional neural network and a point cloud classification network to automatically learn the feature differences between moving obstacles and static objects, eliminating the need for manual feature design and demonstrating strong generalization ability. Attached Figure Description
[0089] Figure 1 This is a framework diagram of this application.
[0090] Figure 2 This is a flowchart of this application.
[0091] Figure 3 This is the ORB-SLAM3 system architecture diagram.
[0092] Figure 4 This is a schematic diagram of PnP joint calibration.
[0093] Figure 5 This is a schematic diagram of the speed compensation geometry in this application.
[0094] Figure 6 It consists of an RGB image and a depth image captured by a camera.
[0095] Figure 7 It is a diagram of the camera's motion trajectory.
[0096] Figure 8 It is a 3D grid map of the viewing area. Detailed Implementation
[0097] The embodiments of the present invention will be disclosed below with reference to the drawings. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details are not intended to limit the present invention. That is, in some embodiments of the present invention, these practical details are not essential. In addition, for the sake of simplicity, some conventional structures and components will be shown in the drawings in a simple schematic manner.
[0098] like Figures 1-2 As shown, this application discloses a multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave, which specifically includes the following steps:
[0099] Step 1: Acquire Multimodal Data: Millimeter-wave radar acquires millimeter-wave radar signals in the corner direction, while a depth camera simultaneously acquires RGB-D image data of the wall at the line of sight. Acquiring multimodal data specifically includes the following steps:
[0100] Step 1.1: Acquiring millimeter-wave radar data: The millimeter-wave radar transmits frequency-modulated continuous wave (FMCW) signals, and the received millimeter-wave radar signals from both line-of-sight and non-line-of-sight areas are the raw echo data.
[0101] Step 1.2: Acquire depth camera data: The depth camera simultaneously acquires RGB-D image data of the wall at the line of sight, which will be used for subsequent 3D reconstruction and normal vector extraction.
[0102] Step 2: Multimodal data preprocessing: The millimeter-wave radar signals received by the antenna undergo range-Doppler fast Fourier transform, and a radar thermal map in RAE format is constructed using a 3D beamforming method. The ORB-SLAM3 algorithm is then used to process the RGB-D image data to reconstruct a 3D mesh map of the line-of-sight area. And extract the wall normal vector The extrinsic parameters of the radar to the camera are obtained through PnP joint calibration, and then combined with the camera's intrinsic parameters to complete the 3D-to-2D projection. Specifically, this includes the following steps:
[0103] Step 2.1, Range-Doppler Fast Fourier Transform: Perform a range-Doppler Fast Fourier Transform on the millimeter-wave radar signal received by each antenna to resolve the range. The distance of the object on the surface to the signal after the Fast Fourier Transform is represented as:
[0104]
[0105] in, For the received time-domain millimeter-wave radar signal, For frequency, The imaginary unit, It is a natural constant. The integral symbol is used. For time Integral infinitesimal element, For time variables, Let be the kernel function for the Fourier transform, and then perform a Doppler Fast Fourier Transform to separate the object in the velocity dimension. The expression for the Doppler Fast Fourier Transform is:
[0106]
[0107] in, For Doppler frequency shift, The range-Doppler spectrum is obtained from the signal after the range-fast Fourier transform. , For distance, For speed;
[0108] Step 2.2, Constant False Alarm Rate Detection: Analyze the obtained range-Doppler spectrum... A constant false alarm rate (CFAR) detection algorithm is applied to identify echo amplitudes exceeding an adaptive threshold. The criteria for determining the object are:
[0109]
[0110] Adaptive threshold Based on background noise power calculate:
[0111]
[0112] in, This is the false alarm rate control factor;
[0113] Step 2.3, Doppler velocity extraction: For values exceeding the adaptive threshold Extract the Doppler frequency shift of the object. Calculate radial velocity :
[0114]
[0115] in, For radar wavelength, Radial velocity as measured by radar;
[0116] Simultaneously, the range-Doppler spectrum Micro-Doppler features are extracted by applying short-time Fourier transform. It is used to describe the subtle patterns of echo frequency changes over time caused by moving obstacles (such as the swinging of a person's limbs), helping to distinguish moving obstacles from static objects.
[0117] Step 2.4: Constructing a radar heatmap in RAE format: Using a three-dimensional beamforming method such as the Capon beamformer, the raw radar echo data from Step 1.1 is used to construct a radar heatmap in RAE format. ,in, For the distance dimension, the number of samples is... For the azimuth dimension, the number of samples is... The number of samples is for the pitch angle dimension;
[0118] Step 2.5, Visual 3D Reconstruction: The ORB-SLAM3 algorithm is used to process the RGB-D image data, such as... Figure 3 As shown, the ORB-SLAM3 system processes visual and inertial data through three parallel threads. Input data includes image streams and inertial measurement unit (IMU) data. The image stream is denoted as... ,in, For the first Frame RGB image, For frame index, This represents the total number of frames in the image sequence, used to extract ORB feature points and provide visual observation information. The inertial measurement unit data is denoted as... ,in, For the first The three-axis acceleration vector corresponding to the frame, in meters per second squared (m²). ), For the first The three-axis angular velocity vector corresponding to the frame, in radians per second ( ( ), used to estimate relative motion between frames and assist visual tracking.
[0119] The visual tracking thread receives image streams and inertial measurement unit (IMU) data, extracts ORB feature points, performs initial pose estimation, and outputs the initial pose of the camera in the world coordinate system. ,in It is a 4×4 transformation matrix. This represents the 3D Euclidean transformation group. The local mapping thread performs local bundle adjustment on the keyframes output by the vision tracking thread to construct a local 3D map, outputting a 3D mesh map of the view distance region. The loop closure detection thread identifies visited scenes, performs global bundle adjustment to eliminate cumulative drift, and globally optimizes the map built by the local mapping thread. These three threads work collaboratively to output a real-time 3D mesh map of the viewport area. And the initial pose of the camera in the world coordinate system Simultaneously extract the wall normal vector ,in The normal vectors are respectively in Components in three directions.
[0120] To verify the actual effect of visual 3D reconstruction, RGB-D image sequences were collected in a typical corridor corner scene, and the ORB-SLAM3 algorithm was run.
[0121] Figure 6 RGB images and depth maps captured by the camera.
[0122] Figure 7 The image shows the camera's motion trajectory. The horizontal axis represents the displacement in the X direction (m), and the vertical axis represents the displacement in the Z direction (m). The total trajectory length is approximately 3.79m. The curve is smooth with no obvious jumps, indicating that SLAM tracking is stable and the positioning is drift-free.
[0123] Figure 8 This is a 3D mesh map of the reconstructed viewing area. The green point cloud in the image is concentrated in the plane range of X coordinate -0.9 to 1.0m and Z coordinate 0 to 0.1m. The point cloud has a small thickness and no discrete noise, indicating that the wall reconstruction is flat and has high accuracy.
[0124] The experimental results above show that the ORB-SLAM3 module can stably reconstruct the geometry of the wall at the line of sight and extract reliable normal vectors, meeting the requirements of this method for geometric information of the line of sight region.
[0125] Step 2.6, PnP Joint Calibration: The radar coordinate system is aligned to the camera coordinate system through PnP joint calibration, such as... Figure 4 As shown, the PnP joint calibration specifically includes the following steps:
[0126] Step 2.6.1: Collect calibration board data: Place a joint calibration board within the common field of view of the radar and camera. The joint calibration board contains both corner reflectors and checkerboard patterns.
[0127] Step 2.6.2, Feature Point Extraction: Pre-calibrate the camera intrinsic parameter matrix. The system calculates lens distortion parameters and corrects distortion at the checkerboard corner points detected by the camera; the radar detects the corner reflectors to obtain the three-dimensional point cloud coordinates in the radar coordinate system. The camera detects the corner points of the chessboard grid and obtains the two-dimensional pixel coordinates in the image coordinate system. ,in The x-coordinate is the pixel coordinate. The vertical coordinate is the pixel coordinate. , and These are the x-axis, y-axis, and z-axis coordinates of the target point in the radar coordinate system, respectively.
[0128] Step 2.6.3, Coordinate Transformation and Projection Solution: By solving the PnP problem, calculate the rotation matrix from the radar coordinate system to the camera coordinate system. Translation vector The radar's 3D points are first transformed into the camera's 3D coordinate system:
[0129]
[0130] Then use the camera intrinsic parameter matrix Mapping perspective projection relationship to two-dimensional pixel coordinates :
[0131]
[0132] in, These are the coordinates of a three-dimensional point in the radar coordinate system. These are the coordinates of a point in the camera's 3D coordinate system after extrinsic parameter transformation. These are two-dimensional pixel coordinates in the image coordinate system. It is a 3×3 rotation matrix. It is a 3×1 translation vector. For the camera intrinsic parameter matrix, This is the homogeneous coordinate scale factor. Thus, the rigid body transformation from radar 3D coordinates to camera 3D coordinates and the perspective projection from camera 3D coordinates to 2D pixel coordinates are completed sequentially.
[0133] Step 3, Non-line-of-sight path analysis: Radar heatmap based on RAE format 3D grid map of viewing distance area The coordinates of the reflection point of the millimeter-wave radar signal on the wall were calculated using ray tracing and the law of reflection. and the direction of the reflection path And reconstruct the complete point cloud of the non-line-of-sight region. Non-line-of-sight path resolution specifically includes the following steps:
[0134] Step 3.1, Radar beam direction definition: For radar heatmaps in RAE format... Each pixel in This corresponds to a radar beam transmission direction. ,in, For the first Azimuth angle For the first The elevation angle and the radar's position are determined by three-dimensional point coordinates. Give;
[0135] Step 3.2, Locating the wall reflection point: along the radar beam emission direction Perform ray tracing to calculate the ray and the 3D mesh map of the view distance area. The first intersection point, i.e., the coordinates of the reflection point on the wall. The coordinates of the reflection point satisfy:
[0136]
[0137] in, Radar to reflection point coordinates The distance;
[0138] Step 3.3: Calculate the reflection path direction: According to the law of specular reflection, the reflection path direction... From the direction of radar beam transmission and wall normal vector calculate:
[0139]
[0140] in, Let be the wall normal vector. Represents the dot product of two vectors;
[0141] Step 3.4, Non-line-of-sight range decomposition: Total flight distance measured by radar Includes radar to reflection point coordinates distance and distance from the reflection point to the moving obstacle :
[0142]
[0143] in, The distance from the reflection point to the moving obstacle. The coordinates of the moving obstacle's position;
[0144] Step 3.5: Construct a complete point cloud for the non-line-of-sight region: For each radar beam transmission direction According to radar heatmaps in RAE format The echo intensity is extracted along the radar beam transmission direction. All valid reflection points in the mirror space are mapped to real-world coordinates using the law of reflection, resulting in a complete point cloud of the non-line-of-sight region. :
[0145]
[0146] Each point contains three-dimensional spatial coordinates and corresponding echo intensity information.
[0147] Step 4, Non-line-of-sight target separation: The motion direction of the moving obstacle is obtained through velocity compensation, and the moving obstacle point cloud is separated from the non-line-of-sight point cloud and its position is extracted through point cloud classification. Output detection confidence level . Figure 5 The image shows the radial velocity measured by radar. The speed of the car's own movement The velocity components of the moving obstacle in the collision-related direction and included angle , The geometric relationship between them is used to aid in understanding the calculation process of velocity compensation and motion direction determination. Specifically, the non-line-of-sight target separation includes the following steps:
[0148] Step 4.1, Vehicle Velocity Acquisition: The inertial measurement unit outputs triaxial acceleration and triaxial angular velocity. Combined with the RGB-D image acquired in Step 1 and the camera pose obtained from ORB-SLAM3, visual-inertial fusion is performed to estimate the vehicle's own velocity. and direction of movement ;in, For velocity scalar, It is a unit direction vector.
[0149] The inertial measurement unit (IMU) does not directly provide the long-term stable absolute velocity of the vehicle. Its three-axis angular velocities are used to estimate the vehicle's attitude changes, and the three-axis accelerations are integrated over a short time after attitude compensation, gravity component subtraction, and sensor bias correction to obtain motion predictions between adjacent image frames. RGB-D images provide visual observations with scale information, and ORB-SLAM3 obtains the keyframe poses. It provides position and attitude constraints for adjacent time steps. Visual-inertial joint optimization uses visual pose constraints to correct for the drift caused by the inertial integral over time, and then obtains the vehicle's velocity based on the displacement and time interval of consecutive time steps. The direction of motion of the trolley is obtained from the direction of displacement. .
[0150] Step 4.2, Velocity Projection Calculation: Radial Velocity Measured by Radar Includes the projected velocity of the moving obstacle in the radar line of sight and the vehicle's own velocity. The contribution of the two, and their geometric relationship is as follows:
[0151]
[0152] in, Radial velocity measured by radar. The projected velocity of the moving obstacle in the radar line-of-sight direction. The direction of motion of the car relative to radar line of sight The included angle is calculated using the following formula:
[0153] ;
[0154] Step 4.3, Radial velocity extraction of moving obstacles: subtract the vehicle's own velocity. Then, the true radial velocity of the moving obstacle relative to the ground is obtained:
[0155] ;
[0156] Step 4.4, Calculation of collision-related directional velocity components: (e.g.) Figure 5 As shown, The velocity component of the moving obstacle in the direction in which it may collide with the car. Direction and radial velocity of the reflection path The angle between the directions. The depth camera obtains the wall orientation and wall normal vector through 3D reconstruction of the line-of-sight area. And the geometry of the corridor corners, combined with the radar echo direction and the direction of the car's movement. Before speed conversion, according to Figure 5 The geometric relationship shown determines the included angle. , Direction of movement of the car The angle between the radar line of sight and the radar line of sight; for Figure 5 The known geometric angles shown are primarily determined by the geometry of the walls and corridor corners recovered by the depth camera, and calculated in conjunction with the radar echo direction. The actual radial velocity used to compensate for the trolley's self-motion is also considered. Convert the velocity components of the moving obstacle in the collision-related direction. :
[0157]
[0158] Step 4.5, Determining the direction of motion: Based on the velocity components of the moving obstacle in the collision-related direction. The sign of the sign indicates the direction of motion:
[0159]
[0160] Step 4.6, Point Cloud Classification and Separation: A deep learning point cloud classification network is used to classify the complete point cloud in non-line-of-sight regions. Processing is performed, combined with microDoppler characteristics. Distinguish between moving obstacles and static objects, and separate the point cloud of moving obstacles. ;
[0161] Step 4.7, Location Extraction: From the point cloud of moving obstacles Extract the position coordinates of non-line-of-sight moving obstacles. Extract the location coordinates of the point cloud centroid or cluster center. Used for the output of step 5.
[0162] Step 4.8, Confidence Output: A three-dimensional convolutional neural network is used to output the confidence score of the radar heatmap in RAE format. Feature extraction is performed, and the detection confidence is output through a fully connected layer and a sigmoid function. :
[0163]
[0164] in, The logit value output by the fully connected layer in the classification branch. This represents the probability that a moving obstacle exists.
[0165] Step 5, Output: Output the position coordinates of the non-line-of-sight moving obstacle. Direction of motion and detection confidence Specifically, it includes the following steps:
[0166] Step 5.1, Position Output: Output the three-dimensional position coordinates of the non-line-of-sight moving obstacle in the world coordinate system. The unit is meters;
[0167] Step 5.2, Direction Output: Output the direction of movement of non-line-of-sight moving obstacles, with values of "approaching", "moving away" or "stationary";
[0168] Step 5.3, Confidence Output: Output the confidence level of the detection. The range of values is This indicates the reliability of the test results.
[0169] This application also employs a binary classification evaluation method to calculate the accuracy. Recall rate and Fractions are calculated using the following formula:
[0170]
[0171]
[0172]
[0173] in, To accurately detect the number of samples of moving obstacles, The number of false positives. This represents the number of underreported samples.
[0174] This application achieves real-time presence detection, precise positioning, and direction determination of moving obstacles behind corners on a mobile platform through multimodal fusion and physically guided speed compensation.
[0175] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave, characterized in that: The multimodal non-line-of-sight moving obstacle detection method specifically includes the following steps: Step 1: Acquire multimodal data: Millimeter-wave radar acquires millimeter-wave radar signals in the corner direction, and depth camera simultaneously acquires RGB-D image data of the wall at line of sight; Step 2, Multimodal Data Preprocessing: Perform range-Doppler fast Fourier transform on the millimeter-wave radar signals received by the antenna in the millimeter-wave radar, and construct a radar heat map in RAE format using a three-dimensional beamforming method. The RGB-D image data acquired in step 1 is processed to reconstruct a 3D mesh map of the viewing distance area. And extract the wall normal vector The external parameters from the radar coordinate system to the camera's three-dimensional coordinate system are obtained through PnP joint calibration, and the projection from the radar's three-dimensional points to the image's two-dimensional pixels is completed by combining the camera's intrinsic parameters. Step 3, Non-line-of-sight path resolution: Radar heatmap based on RAE format 3D grid map of viewing distance area Calculate the coordinates of the reflection point of the millimeter-wave radar signal on the wall. and the direction of the reflection path And reconstruct the complete point cloud of the non-line-of-sight region. ; Step 4, Non-line-of-sight target separation: The motion direction of the moving obstacle is obtained through velocity compensation, and the complete point cloud of the non-line-of-sight region is separated by point cloud classification. Separate the point cloud of moving obstacles and extract the position coordinates of non-line-of-sight moving obstacles. Output the confidence level of the detection. The non-line-of-sight target separation in this step specifically includes the following steps: Step 4.1, Vehicle Speed Acquisition: Visual-inertial fusion is performed using RGB-D images and triaxial acceleration and angular velocity data output by the inertial measurement unit to estimate the vehicle's own speed. and direction of movement ; Step 4.2, Velocity Projection Calculation: Radial Velocity Measured by Radar Includes the projected velocity of the moving obstacle in the radar line of sight and the vehicle's own velocity. The contribution of the two, and their geometric relationship is as follows: in, Radial velocity measured by radar. The projected velocity of the moving obstacle in the radar line-of-sight direction. The direction of motion of the car relative to radar line of sight The included angle is calculated using the following formula: ; Step 4.3, Radial velocity extraction of moving obstacles: subtract the vehicle's own velocity. Then, the true radial velocity of the moving obstacle relative to the ground is obtained. : ; Step 4.4, Calculation of collision-related directional velocity components: The velocity component of the moving obstacle in the direction in which it may collide with the car. Direction and radial velocity of the reflection path The angle between the directions, the depth camera obtains the wall orientation and wall normal vector through 3D reconstruction of the view range. And the geometry of the corridor corners, combined with the radar echo direction and the direction of the car's movement. Before speed conversion, the included angle is determined according to geometric relationships. , in, Direction of movement of the car The angle between the radar line of sight and the radar line of sight; Given the geometric angles, primarily determined by the geometry of the walls and corridor corners recovered by the depth camera, and calculated in conjunction with the radar echo direction, the actual radial velocity used for trolley self-motion compensation is used. Convert the velocity components of the moving obstacle in the collision-related direction. : Step 4.5, Determining the direction of motion: Based on the velocity components of the moving obstacle in the collision-related direction. The sign of the sign indicates the direction of motion: Step 4.6, Point Cloud Classification and Separation: A deep learning point cloud classification network is used to classify the complete point cloud in non-line-of-sight regions. The data is processed and combined with the extracted micro-Doppler features. Distinguish between moving obstacles and static objects, and separate the point cloud of moving obstacles. ; Step 4.7, Location Extraction: From the point cloud of moving obstacles Extract the position coordinates of non-line-of-sight moving obstacles. Take the centroid or cluster center of the cloud; Step 4.8, Confidence Output: A three-dimensional convolutional neural network is used to output the confidence score of the radar heatmap in RAE format. Feature extraction is performed, and the detection confidence is output through a fully connected layer and a sigmoid function. : in, The logit value output by the fully connected layer in the classification branch; Step 5, Output: Output the position coordinates of the non-line-of-sight moving obstacle. Direction of motion and confidence level of detection .
2. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 1, characterized in that: Step 1 specifically includes the following steps: Step 1.1: Acquiring millimeter-wave radar data: The millimeter-wave radar transmits frequency-modulated continuous wave signals and receives millimeter-wave radar signals from both line-of-sight and non-line-of-sight areas, i.e., the raw echo data. Step 1.2: Acquire depth camera data: The depth camera simultaneously acquires RGB-D image data of the wall at the line of sight, which will be used for subsequent 3D reconstruction and normal vector extraction.
3. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 2, characterized in that: Step 2 specifically includes the following steps: Step 2.1, Range-Doppler Fast Fourier Transform: Perform a range-Doppler Fast Fourier Transform on the millimeter-wave radar signal received by each antenna to resolve the range. The distance of the object on the surface to the signal after the Fast Fourier Transform is represented as: in, For the received time-domain millimeter-wave radar signal, For frequency, The imaginary unit, It is a natural constant. The integral symbol is used. For time Integral infinitesimal element, For time variables, Let be the kernel function for the Fourier transform, and then perform a Doppler Fast Fourier Transform to separate the object in the velocity dimension. The expression for the Doppler Fast Fourier Transform is: in, For Doppler frequency shift, The range-Doppler spectrum is obtained from the signal after the range-fast Fourier transform. , For distance, For speed; Step 2.2, Constant False Alarm Rate Detection: Analyze the obtained range-Doppler spectrum... A constant false alarm rate (CFAR) detection algorithm is applied to identify echo amplitudes exceeding an adaptive threshold. The criteria for determining the object are: Adaptive threshold Based on background noise power calculate: in, This is a control factor for the false alarm rate; Step 2.3, Doppler velocity extraction: For values exceeding the adaptive threshold Extract the Doppler frequency shift of the object. Calculate radial velocity : in, For radar wavelength, Radial velocity as measured by radar; Simultaneously, the range-Doppler spectrum Micro-Doppler features are extracted by applying short-time Fourier transform. ; Step 2.4: Constructing a radar heatmap in RAE format: Using a three-dimensional beamforming method, the raw echo data obtained in Step 1.1 is used to construct a radar heatmap in RAE format. ,in, For the distance dimension, the number of samples is... For the azimuth dimension, the number of samples is... The number of samples is for the pitch angle dimension; Step 2.5, Visual 3D Reconstruction: The ORB-SLAM3 algorithm is used to process the RGB-D image data acquired in Step 1 to reconstruct a 3D mesh map of the viewing distance area in real time. And extract the wall normal vector. , For the normal vector in Components on the axis, For the normal vector in Components on the axis, For the normal vector in Components on the axis; Step 2.6, PnP Joint Calibration: The radar coordinate system is aligned with the camera coordinate system through PnP joint calibration.
4. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 3, characterized in that: Step 2.6 specifically includes the following steps: Step 2.6.1: Collect calibration board data: Place a joint calibration board within the common field of view of the radar and camera. The joint calibration board contains both corner reflectors and checkerboard patterns. Step 2.6.2, Feature Point Extraction: Pre-calibrate the camera intrinsic parameter matrix. The system calculates lens distortion parameters and corrects distortion at the checkerboard corner points detected by the camera; the radar detects the corner reflectors to obtain the three-dimensional point cloud coordinates in the radar coordinate system. The camera detects the corner points of the chessboard grid and obtains the two-dimensional pixel coordinates in the image coordinate system. ,in The x-coordinate is the pixel coordinate. The vertical coordinate is the pixel coordinate. , and These are the x-axis, y-axis, and z-axis coordinates of the target point in the radar coordinate system, respectively. Step 2.6.3, Coordinate Transformation and Projection Solution: By solving the PnP problem, calculate the rotation matrix from the radar coordinate system to the camera coordinate system. Translation vector The radar's 3D points are first transformed into the camera's 3D coordinate system: Then use the camera intrinsic parameter matrix Mapping perspective projection relationship to two-dimensional pixel coordinates : in, These are the coordinates of a three-dimensional point in the radar coordinate system. These are the coordinates of a point in the camera's 3D coordinate system after extrinsic parameter transformation. These are two-dimensional pixel coordinates in the image coordinate system. It is a 3×3 rotation matrix. It is a 3×1 translation vector. For the camera intrinsic parameter matrix, The homogeneous coordinate scale factor is used to sequentially perform the rigid body transformation from radar 3D coordinates to camera 3D coordinates and the perspective projection from camera 3D coordinates to 2D pixel coordinates.
5. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 4, characterized in that: Step 3, the non-line-of-sight path resolution, specifically includes the following steps: Step 3.1, Radar beam direction definition: For radar heatmaps in RAE format... Each pixel in Corresponding to a radar beam transmission direction ,in, For the first Azimuth angle For the first The elevation angle and the radar's position are determined by three-dimensional point coordinates. Give; Step 3.2, Locating the wall reflection point: along the radar beam emission direction Perform ray tracing to calculate the ray and the 3D mesh map of the view distance area. The first intersection point, i.e., the coordinates of the reflection point on the wall. The coordinates of the reflection point satisfy: in, Radar to reflection point coordinates The distance; Step 3.3: Calculate the reflection path direction: Reflection path direction From the direction of radar beam transmission and wall normal vector calculate: in, Let be the wall normal vector. Represents the dot product of two vectors; Step 3.4, Non-line-of-sight range decomposition: Total flight distance measured by radar Includes radar to reflection point coordinates distance and the distance from the reflection point to the moving obstacle : in, The distance from the reflection point to the moving obstacle. The coordinates of the moving obstacle's position; Step 3.5: Construct a complete point cloud for the non-line-of-sight region: For each radar beam transmission direction According to radar heatmaps in RAE format The echo intensity is extracted along the radar beam transmission direction. All valid reflection points in the mirror space are mapped to real-world coordinates using the law of reflection, resulting in a complete point cloud of the non-line-of-sight region. : Each point contains three-dimensional spatial coordinates and corresponding echo intensity information. The mirrored space is the spatial region where the virtual image point extends along the radar observation direction in the RAE radar thermal image. The image point in this region is symmetrical to the position of the real object about the wall.
6. The multimodal non-line-of-sight moving obstacle detection method based on vision-millimeter wave as described in claim 1, characterized in that: The output described in step 5 specifically includes the following steps: Step 5.1, Position Output: Output the three-dimensional position coordinates of the non-line-of-sight moving obstacle in the world coordinate system. ; Step 5.2, Direction Output: Output the direction of movement of non-line-of-sight moving obstacles, with values of "approaching", "moving away" or "stationary"; Step 5.3, Confidence Output: Output the confidence level of the detection. The range of values is This indicates the reliability of the test results.