A millimeter wave radar-based target bounding box detection system and method

By using a millimeter-wave radar-based target bounding box detection system, and combining a point cloud data generation and optimization module with a two-dimensional contour prediction neural network, the problems of light limitation and high cost in existing technologies are solved, achieving high-precision and low-cost target bounding box detection, which is suitable for complex scenarios.

CN119511281BActive Publication Date: 2025-12-30NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411587828.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-12-30
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

Existing target detection solutions are susceptible to light limitations, have high deployment and processing costs, and lack accuracy and resolution, especially under low light or extreme weather conditions.

Method used

A target bounding box detection system based on millimeter-wave radar is adopted, including a point cloud data generation module, a data optimization module, and a bounding box detection module. Point cloud data is generated using millimeter-wave radar, and target bounding boxes are detected by removing noise points and enhancing sparse point clouds, combined with a two-dimensional contour prediction neural network.

Benefits of technology

It achieves high-precision, low-cost, and highly applicable target bounding box detection, can work reliably under various lighting conditions, has an error of less than 0.11m, and is suitable for complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119511281B_ABST
    Figure CN119511281B_ABST
Patent Text Reader

Abstract

The application discloses a target bounding box detection system and method based on a millimeter wave radar, which comprises a point cloud data generation module, a data optimization module and a bounding box detection module.The point cloud data generation module is used for collecting millimeter wave intermediate frequency signals by using the millimeter wave radar and generating point cloud data of a target.The data optimization module is used for removing noise in the point cloud data by using a reflection saliency coefficient and enhancing the point cloud by using a cumulative matrix.The bounding box detection module is used for obtaining a target reflection energy spatial distribution map by performing receiving end beamforming on the millimeter wave intermediate frequency signals, fusing the target reflection energy spatial distribution map and the optimized point cloud data to obtain a fused multi-channel feature map, inputting the fused multi-channel feature map into a two-dimensional contour prediction neural network to obtain a two-dimensional contour image of the target, and finally detecting a bounding box of the target in combination with the optimized point cloud data.The application improves the applicability and reliability of target bounding box detection, and can reliably work in various environments with poor light conditions by using a millimeter wave radar related algorithm design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of millimeter-wave wireless sensing technology, specifically relating to a target bounding box detection system and method based on millimeter-wave radar. Background Technology

[0002] Object detection plays a crucial role in numerous applications in production and daily life, enabling machines to perceive the real world and interact further. In intelligent transportation systems, object detection not only aids in vehicle control but is also increasingly applied to emerging applications such as map generation and vehicle communication. Object detection also plays a key role in intelligent unmanned aerial vehicle (UAV) systems, widely used in rescue missions, delivery services, and wildfire suppression. After a target is detected, it is usually necessary to further determine the target's boundary, a process known as bounding box estimation. Reliable and accurate bounding box detection is essential for effective machine-physical interaction, such as ensuring reliable navigation and preventing collisions in autonomous driving. Currently, existing bounding box detection schemes mainly fall into the following categories:

[0003] 1) Computer vision-based detection schemes, including images captured by ordinary cameras, infrared cameras, and depth cameras. These technologies typically perform poorly in scenarios with unfavorable visual conditions, such as low light or extreme weather, which are common in driving scenarios.

[0004] 2) LiDAR-based detection schemes utilize point clouds obtained from LiDAR scanning for detection. However, LiDAR systems are typically very expensive, and the point cloud data is large, requiring significant computing power and resources for processing.

[0005] 3) The detection scheme based on ultrasonic radar obtains the target detection result by analyzing the echo signal of ultrasonic radar. However, ultrasonic radar has low resolution, poor accuracy, and short effective range. Summary of the Invention

[0006] To address the shortcomings of the existing technologies, the present invention aims to provide a target bounding box detection system and method based on millimeter-wave radar, thereby solving the problems of sensor susceptibility to light limitations and high deployment and processing costs in existing target detection scenarios. The present invention utilizes the signal reflected from the target by millimeter-wave radar to achieve the detection of the target bounding box in the scene.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0008] The present invention provides a target bounding box detection system based on millimeter-wave radar, comprising: a point cloud data generation module, a data optimization module, and a bounding box detection module;

[0009] The point cloud data generation module includes a millimeter-wave radar and a point cloud data generation unit;

[0010] Millimeter-wave radar is used to transmit millimeter-wave frequency-modulated continuous wave signals, receive signals reflected from targets, and obtain millimeter-wave intermediate frequency signals through frequency mixing.

[0011] The point cloud data generation unit is used to convert millimeter-wave intermediate frequency signals into the target's distance, velocity, and azimuth angle to generate point cloud data.

[0012] The data optimization module is used to remove noise points and enhance sparse point clouds in point cloud data to obtain optimized point cloud data.

[0013] The bounding box detection module is used to perform beamforming operation on the millimeter-wave intermediate frequency signal to obtain the spatial distribution map of the target reflection energy, and to fuse the spatial distribution map of the target reflection energy with the optimized point cloud data to obtain a fused multi-channel feature map. The fused multi-channel feature map is input into a two-dimensional contour prediction neural network to obtain a predicted two-dimensional contour image of the target. Combined with the optimized point cloud data, the bounding box of the target is finally detected.

[0014] Furthermore, the point cloud data acquisition module transmits a continuously frequency-modulated millimeter-wave signal with a frequency of 60–64 GHz, and receives the echo signal reflected from the target through multiple receiving antennas. The millimeter-wave transmission signal s at time t... Tx (t) and received signal s Rx (t) is expressed by the following formulas:

[0015] s Tx (t)=exp[j(2πf c t+πKt 2 )]

[0016] s Rx (t)=αS Tx [t-2R(t) / c]

[0017] Where j is the imaginary unit, f c Let be the starting frequency of the frequency-modulated continuous wave, K be the rate of change of frequency, R(t) be the distance between the millimeter-wave radar and the target at time t, α be the path loss during reflection, and c be the speed of light in vacuum.

[0018] The intermediate frequency signal s(t) obtained after frequency mixing is expressed as:

[0019] s(t)=s Tx (t)conj[s Rx (t)]≈αexp[j4π(f c +Kt)R(t) / c]

[0020] Among them, conj[s Rx [(t)] represents the received signal s Rx The conjugate of (t);

[0021] The intermediate frequency signal s(t) after mixing contains aliasing of reflection signals from different targets. Performing a Fast Fourier Transform on it to separate the reflection signals from targets at different distances is expressed as:

[0022]

[0023] Where FFT represents Fast Fourier Transform, S i (t) represents the signal component corresponding to a separated target, called the range Fourier signal, R i This indicates the distance between the reflected signal and the millimeter-wave radar.

[0024] Furthermore, the specific steps for the point cloud data acquisition module to generate point cloud data are as follows:

[0025] After the first Fast Fourier Transform, the phase ω of the distance Fourier signal is expressed as:

[0026]

[0027] Where ΔR represents the tiny displacement of the target between adjacent frequency-modulated signals transmitted by the millimeter-wave radar, and λ represents the wavelength corresponding to the center frequency of the millimeter-wave frequency-modulated continuous wave signal;

[0028] Perform a Fast Fourier Transform on the distance Fourier signal to obtain the target's velocity v, as shown in the following formula:

[0029]

[0030] Where Δω is the change in phase, T c The duration of linear frequency modulation;

[0031] After performing the above two fast Fourier transforms on the millimeter-wave intermediate frequency signal, the range-Doppler matrix is ​​obtained. The peak points on the range-Doppler matrix indicate the location of strong reflection, which is the target point corresponding to the target to be detected. The coordinates of the peak points represent the target's distance and velocity, respectively.

[0032] The azimuth angle θ of the target in space is obtained by performing a Fast Fourier Transform on the multiple signals acquired by different receiving antennas of the millimeter-wave radar, as follows:

[0033]

[0034] in, is the phase difference of the signals received by adjacent receiving antennas, and l is the distance between adjacent receiving antennas;

[0035] Based on the distance, velocity, and azimuth of the target point obtained above, the position corresponding to the detected target point is mapped to a Cartesian coordinate system in space to obtain point cloud data.

[0036] Furthermore, the data optimization module performs noise point removal, and the specific steps are as follows:

[0037] The significance coefficient of reflection, RS, is defined as follows:

[0038] RS=S snr ·S v ·S a

[0039] Among them, S snr S represents the significance of the signal-to-noise ratio. v For the significance of speed, S a The angular significance is used as the threshold. The average reflectance coefficient of all points is calculated as the threshold. Points with reflectance coefficient values ​​lower than the threshold are considered noise points and removed to complete the noise point removal of the point cloud data.

[0040] Furthermore, the data optimization module performs sparse point cloud enhancement, with the following specific steps:

[0041] Define a cumulative matrix to enhance sparse point clouds by leveraging the continuity of motion in both time and space dimensions; divide the plane into M×N grids, and define the cumulative matrix BM as follows:

[0042] BM t (m,n)=αBM t-1 (m,n)+count t (m,n) / max(count t )

[0043] Where (m,n) represents the grid coordinates, BM t (m,n) represents the value of the cumulative matrix at coordinates (m,n) at time t, BM t-1 (m,n) represents the value of the cumulative matrix at coordinates (m,n) at time t-1, α represents the time decay factor, and count t (m,n) represents the number of points in the (m,n) grid at time t, and max(count) t ) represents the maximum number of points in the grid; PM t This represents the projection of the denoised point cloud onto the X and Z axes. The cumulative matrix is ​​used to enhance the sparse point cloud projection, as shown in the following expression:

[0044]

[0045] Among them, PM′ t This represents the enhanced point cloud, and γ represents the threshold.

[0046] Furthermore, the bounding box detection module detects the target bounding box based on the acquired millimeter-wave intermediate frequency signal and optimized point cloud data. The specific steps are as follows:

[0047] (1) Extracting the spatial distribution map of target reflected energy based on receiver beamforming;

[0048] (2) The spatial distribution map of the target reflection energy and the optimized point cloud data are fused to obtain a fused multi-channel feature map. The fused multi-channel feature map is used as input to design a two-dimensional contour prediction neural network to predict the target bounding box.

[0049] (3) Based on the prediction results, the bounding box of the target is detected in the three-dimensional coordinate system using the target's depth information and projection rules.

[0050] Furthermore, step (1) specifically includes:

[0051] Multiple intermediate frequency (IF) signals acquired by different receiving antennas of a millimeter-wave radar are aligned onto the same wavefront to achieve beamforming for a specific β-angle direction. The beamformed signal x β The calculation method is as follows:

[0052] x β =[1,e -jπdsinβ ,e -jπ2dsinβ ,…,e -jπ(N-1)dsinβ ]x

[0053] Where β represents the angle between the target and the radial direction of the millimeter-wave radar, N represents the number of millimeter-wave radar receiving antennas, d represents the spacing between adjacent receiving antennas, and x represents the original multi-channel signal vectors obtained from different receiving antennas.

[0054] By using the above method, signals from different directions in space are aggregated to obtain a spatial distribution map of the target's reflected energy.

[0055] Furthermore, step (2) specifically includes:

[0056] The spatial distribution map of the target's reflected energy is convolved with the projection map of the optimized point cloud in the X-axis and Z-axis planes to obtain a fused point cloud contour image. A fused multi-channel feature map containing three different dimensions—distance, signal-to-noise ratio, and contour—is generated as the input to the two-dimensional contour prediction neural network. The depth image of the target captured by the depth camera is used as the expected output of the two-dimensional contour prediction neural network.

[0057] The 2D contour prediction neural network uses a fused multi-channel feature map of three different dimensions to predict the 2D contour of a target. Based on the ResNet architecture, it starts with a convolutional block, passes through a downsampling layer, transitions through a ResNet block, then an upsampling layer, and finally ends with another convolutional block. The Tanh layer introduces L2 regularization to improve the generalization performance of the 2D contour prediction neural network. The loss function during training is expressed as follows:

[0058]

[0059] in, For loss function, Let y be the calculated output value, θ be the desired output value, and L be the L2 regularization parameter. BCE The binary cross-entropy loss is expressed as follows:

[0060]

[0061] Where m and n represent the coordinates of the planar grid;

[0062] The fused multi-channel feature map is used as the input to the 2D contour prediction neural network, and the corresponding depth image is used as the output to train the 2D contour prediction network. After training, the trained 2D contour prediction neural network is obtained, which is used to convert the fused multi-channel feature map into the target 2D contour image.

[0063] Furthermore, step (3) specifically includes:

[0064] The target bounding box is reconstructed in a 3D coordinate system using the distances between the target's 2D contour image and the optimized point cloud. The horizontal and vertical fields of view of the depth camera are denoted as HFOV and VFOV, respectively, and the size of the depth image is I. height ×I weight Let the coordinates of a target point be (x, y, z), and its azi and elevation angles relative to the millimeter-wave radar be azi and ele, respectively. According to the projection principle, the mapping of this point on the depth image is represented as (i, j), as follows:

[0065]

[0066] The point (i,j) on the target 2D contour image is transformed to a 3D coordinate system using the following formula:

[0067]

[0068] Where (x,y,z) are the coordinates of the point in the three-dimensional coordinate system, and range is the distance between the target and the millimeter-wave radar. The above transformation is applied to each vertex of the target's two-dimensional contour image to obtain the target's bounding box in the three-dimensional coordinate system.

[0069] The present invention provides a target bounding box detection method based on millimeter-wave radar, which, based on the above system, includes the following steps:

[0070] 1) Transmit millimeter-wave frequency-modulated continuous wave signals;

[0071] 2) Receive the signal reflected by the target, mix it to obtain a millimeter-wave intermediate frequency signal;

[0072] 3) The range-Doppler matrix is ​​obtained by transforming the millimeter-wave intermediate frequency signal, and the range and velocity of the target are determined based on the coordinates of the peak value of the range-Doppler matrix;

[0073] 4) Perform a Fast Fourier Transform on the millimeter-wave intermediate frequency signals from different receiving antennas to obtain the azimuth angle of the target;

[0074] 5) Map the distance, velocity, and azimuth of the target point to a Cartesian coordinate system in space to obtain point cloud data;

[0075] 6) Calculate the reflectance significance coefficient and remove noise points from the point cloud data based on the threshold;

[0076] 7) Calculate the cumulative matrix and use it to enhance the point cloud;

[0077] 8) Beamforming at different angles on millimeter-wave intermediate frequency signals from different receiving antennas to obtain the spatial distribution map of the target reflected energy;

[0078] 9) Convolve the spatial distribution map of the target's reflected energy with the projection of the optimized point cloud onto the XZ-axis plane to generate a fused multi-channel feature map;

[0079] 10) Use the fused multi-channel feature map as input to train the two-dimensional contour prediction neural network, and obtain the trained two-dimensional contour prediction neural network.

[0080] 11) Input the fused multi-channel feature map into the trained two-dimensional contour prediction neural network to obtain the target two-dimensional contour image;

[0081] 12) Use the distances corresponding to the two-dimensional contour image and the optimized point cloud to obtain the bounding box of the target in the spatial coordinate system.

[0082] The beneficial effects of this invention are:

[0083] 1. This invention can achieve high-precision target bounding box detection: the average error for detecting target bounding boxes is 0.11m.

[0084] 2. This invention improves the applicability and reliability of target bounding box detection: by utilizing millimeter-wave radar-related algorithm design, it can work reliably in various environments with poor lighting conditions.

[0085] 3. This invention reduces the cost of target bounding box detection: it achieves efficient target bounding box detection by using data acquired by low-cost millimeter-wave radar.

[0086] 4. This invention is easy to deploy: it has low environmental requirements, is not easily disturbed, has high robustness, and can be deployed in various complex scenarios such as transportation and production to work normally. Attached Figure Description

[0087] Figure 1 This is an architecture diagram of the system of the present invention;

[0088] Figure 2 This is a schematic diagram illustrating the principle of multiple antenna beamforming in this invention;

[0089] Figure 3a This is a schematic diagram of the spatial distribution of reflected energy according to the present invention;

[0090] Figure 3b This is a schematic diagram of the fused multi-channel feature map of the present invention;

[0091] Figure 4 This is a schematic diagram of the two-dimensional contour prediction neural network structure of the present invention;

[0092] Figure 5 This is a schematic diagram illustrating the principle of two-dimensional contour image projection according to the present invention. Detailed Implementation

[0093] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and accompanying drawings. The content mentioned in the embodiments is not intended to limit the present invention.

[0094] Reference Figures 1 to 5 As shown, the present invention provides a target bounding box detection system based on millimeter-wave radar, comprising: a point cloud data generation module, a data optimization module, and a bounding box detection module;

[0095] The point cloud data generation module includes a millimeter-wave radar and a point cloud data generation unit;

[0096] Millimeter-wave radar is used to transmit millimeter-wave frequency-modulated continuous wave signals, receive signals reflected from targets, and obtain millimeter-wave intermediate frequency signals through frequency mixing.

[0097] The point cloud data generation unit is used to convert millimeter-wave intermediate frequency signals into the target's distance, velocity, and azimuth angle to generate point cloud data.

[0098] The point cloud data acquisition module transmits a continuous frequency modulated millimeter-wave signal with a frequency of 60–64 GHz, and receives the echo signal reflected from the target through multiple receiving antennas. The transmitted millimeter-wave signal s at time t... Tx (t) and received signal s Rx (t) is expressed by the following formulas:

[0099] s Tx (t)=exp[j(2πf c t+πKt 2 )]

[0100] s Rx (t)=αS Tx [t-2R(t) / c]

[0101] Where j is the imaginary unit, f c Let be the starting frequency of the frequency-modulated continuous wave, K be the rate of change of frequency, R(t) be the distance between the millimeter-wave radar and the target at time t, α be the path loss during reflection, and c be the speed of light in vacuum.

[0102] The intermediate frequency signal s(t) obtained after frequency mixing is expressed as:

[0103] s(t)=s Tx (t)conj[s Rx (t)]≈αexp[j4π(f c +Kt)R(t) / c]

[0104] Among them, conj[s Rx [(t)] represents the received signal s Rx The conjugate of (t);

[0105] The intermediate frequency signal s(t) after mixing contains aliasing of reflection signals from different targets. Performing a Fast Fourier Transform on it to separate the reflection signals from targets at different distances is expressed as:

[0106]

[0107] Where FFT represents Fast Fourier Transform, S i (t) represents the signal component corresponding to a separated target, called the range Fourier signal, R i This indicates the distance between the reflected signal and the millimeter-wave radar.

[0108] The specific steps by which the point cloud data acquisition module generates point cloud data are as follows:

[0109] After the first Fast Fourier Transform, the phase ω of the distance Fourier signal is expressed as:

[0110]

[0111] Where ΔR represents the tiny displacement of the target between adjacent frequency-modulated signals transmitted by the millimeter-wave radar, and λ represents the wavelength corresponding to the center frequency of the millimeter-wave frequency-modulated continuous wave signal;

[0112] Perform a Fast Fourier Transform on the distance Fourier signal to obtain the target's velocity v, as shown in the following formula:

[0113]

[0114] Where Δω is the change in phase, T c The duration of linear frequency modulation;

[0115] After performing the above two fast Fourier transforms on the millimeter-wave intermediate frequency signal, the range-Doppler matrix is ​​obtained. The peak points on the range-Doppler matrix indicate the location of strong reflection, which is the target point corresponding to the target to be detected. The coordinates of the peak points represent the target's distance and velocity, respectively.

[0116] The azimuth angle θ of the target in space is obtained by performing a Fast Fourier Transform on the multiple signals acquired by different receiving antennas of the millimeter-wave radar, as follows:

[0117]

[0118] in, is the phase difference of the signals received by adjacent receiving antennas, and l is the distance between adjacent receiving antennas;

[0119] Based on the distance, velocity, and azimuth of the target point obtained above, the position corresponding to the detected target point is mapped to a Cartesian coordinate system in space to obtain point cloud data.

[0120] The data optimization module is used to remove noise points and enhance sparse point clouds in point cloud data to obtain optimized point cloud data.

[0121] Specifically, the data optimization module removes noise points, and the specific steps are as follows:

[0122] The significance coefficient of reflection, RS, is defined as follows:

[0123] RS=S snr ·S v ·S a

[0124] Among them, S snr S represents the significance of the signal-to-noise ratio. v For the significance of speed, S aThe angular significance is used as the threshold. The average reflectance coefficient of all points is calculated as the threshold. Points with reflectance coefficient values ​​lower than the threshold are considered noise points and removed to complete the noise point removal of the point cloud data.

[0125] Specifically, the data optimization module performs sparse point cloud enhancement, and the specific steps are as follows:

[0126] Define a cumulative matrix to enhance sparse point clouds by leveraging the continuity of motion in both time and space dimensions; divide the plane into M×N grids, and define the cumulative matrix BM as follows:

[0127] BM t (m,n)=αBM t-1 (m,n)+count t (m,n) / max(count t )

[0128] Where (m,n) represents the grid coordinates, BM t (m,n) represents the value of the cumulative matrix at coordinates (m,n) at time t, BM t-1 (m,n) represents the value of the cumulative matrix at coordinates (m,n) at time t-1, α represents the time decay factor, and count t (m,n) represents the number of points in the (m,n) grid at time t, and max(count) t ) represents the maximum number of points in the grid; PM t This represents the projection of the denoised point cloud onto the X and Z axes. The cumulative matrix is ​​used to enhance the sparse point cloud projection, as shown in the following expression:

[0129]

[0130] Among them, PM′ t This represents the enhanced point cloud, and γ represents the threshold.

[0131] The bounding box detection module is used to perform beamforming operation on the millimeter-wave intermediate frequency signal to obtain the spatial distribution map of the target reflection energy, and to fuse the spatial distribution map of the target reflection energy with the optimized point cloud data to obtain a fused multi-channel feature map. The fused multi-channel feature map is input into the two-dimensional contour prediction neural network to obtain the predicted two-dimensional contour image of the target. Combined with the optimized point cloud data, the bounding box of the target is finally detected.

[0132] The bounding box detection module detects the target bounding box based on the acquired millimeter-wave intermediate frequency signal and optimized point cloud data. The specific steps are as follows:

[0133] (1) Extracting the spatial distribution map of target reflected energy based on receiver beamforming;

[0134] (2) The spatial distribution map of the target reflection energy and the optimized point cloud data are fused to obtain a fused multi-channel feature map. The fused multi-channel feature map is used as input to design a two-dimensional contour prediction neural network to predict the target bounding box.

[0135] (3) Based on the prediction results, the bounding box of the target is detected in the three-dimensional coordinate system using the target's depth information and projection rules.

[0136] Specifically, step (1) includes:

[0137] Multiple intermediate frequency (IF) signals acquired by different receiving antennas of a millimeter-wave radar are aligned onto the same wavefront to achieve beamforming for a specific β-angle direction. The beamformed signal x β The calculation method is as follows:

[0138] x β =[1,e -jπdsinβ ,e -jπ2dsinβ ,…,e -jπ(N-1)dsinβ ]x

[0139] Where β represents the angle between the target and the radial direction of the millimeter-wave radar, N represents the number of millimeter-wave radar receiving antennas, d represents the spacing between adjacent receiving antennas, and x represents the original multi-channel signal vectors obtained from different receiving antennas.

[0140] By using the above method, signals from different directions in space are aggregated to obtain a spatial distribution map of the target's reflected energy.

[0141] Specifically, step (2) includes:

[0142] The spatial distribution map of the target's reflected energy is convolved with the projection map of the optimized point cloud in the X-axis and Z-axis planes to obtain a fused point cloud contour image. A fused multi-channel feature map containing three different dimensions—distance, signal-to-noise ratio, and contour—is generated as the input to the two-dimensional contour prediction neural network. The depth image of the target captured by the depth camera is used as the expected output of the two-dimensional contour prediction neural network.

[0143] The 2D contour prediction neural network uses a fused multi-channel feature map of three different dimensions to predict the 2D contour of a target. Based on the ResNet architecture, it starts with a convolutional block, passes through a downsampling layer, transitions through a ResNet block, then an upsampling layer, and finally ends with another convolutional block. The Tanh layer introduces L2 regularization to improve the generalization performance of the 2D contour prediction neural network. The loss function during training is expressed as follows:

[0144]

[0145] in, For loss function, Let y be the calculated output value, θ be the desired output value, and L be the L2 regularization parameter. BCE The binary cross-entropy loss is expressed as follows:

[0146]

[0147] Where m and n represent the coordinates of the planar grid;

[0148] The fused multi-channel feature map is used as the input to the 2D contour prediction neural network, and the corresponding depth image is used as the output to train the 2D contour prediction network. After training, the trained 2D contour prediction neural network is obtained, which is used to convert the fused multi-channel feature map into the target 2D contour image.

[0149] Specifically, step (3) includes:

[0150] The target bounding box is reconstructed in a 3D coordinate system using the distances between the target's 2D contour image and the optimized point cloud. The horizontal and vertical fields of view of the depth camera are denoted as HFOV and VFOV, respectively, and the size of the depth image is I. height ×I weight Let the coordinates of a target point be (x, y, z), and its azi and elevation angles relative to the millimeter-wave radar be azi and ele, respectively. According to the projection principle, the mapping of this point on the depth image is represented as (i, j), as follows:

[0151]

[0152] The point (i,j) on the target 2D contour image is transformed to a 3D coordinate system using the following formula:

[0153]

[0154] Where (x,y,z) are the coordinates of the point in the three-dimensional coordinate system, and range is the distance between the target and the millimeter-wave radar. The above transformation is applied to each vertex of the target's two-dimensional contour image to obtain the target's bounding box in the three-dimensional coordinate system.

[0155] The present invention provides a target bounding box detection method based on millimeter-wave radar, which, based on the above system, includes the following steps:

[0156] 1) Transmit millimeter-wave frequency-modulated continuous wave signals;

[0157] 2) Receive the signal reflected by the target, mix it to obtain a millimeter-wave intermediate frequency signal;

[0158] 3) The range-Doppler matrix is ​​obtained by transforming the millimeter-wave intermediate frequency signal, and the range and velocity of the target are determined based on the coordinates of the peak value of the range-Doppler matrix;

[0159] 4) Perform a Fast Fourier Transform on the millimeter-wave intermediate frequency signals from different receiving antennas to obtain the azimuth angle of the target;

[0160] 5) Map the distance, velocity, and azimuth of the target point to a Cartesian coordinate system in space to obtain point cloud data;

[0161] 6) Calculate the reflectance significance coefficient and remove noise points from the point cloud data based on the threshold;

[0162] 7) Calculate the cumulative matrix and use it to enhance the point cloud;

[0163] 8) Beamforming at different angles on millimeter-wave intermediate frequency signals from different receiving antennas to obtain the spatial distribution map of the target reflected energy;

[0164] 9) Convolve the spatial distribution map of the target's reflected energy with the projection of the optimized point cloud onto the XZ-axis plane to generate a fused multi-channel feature map;

[0165] 10) Use the fused multi-channel feature map as input to train the two-dimensional contour prediction neural network, and obtain the trained two-dimensional contour prediction neural network.

[0166] 11) Input the fused multi-channel feature map into the trained two-dimensional contour prediction neural network to obtain the target two-dimensional contour image;

[0167] 12) Use the distances corresponding to the two-dimensional contour image and the optimized point cloud to obtain the bounding box of the target in the spatial coordinate system.

[0168] This invention has many specific applications. The above description is only a preferred embodiment of this invention. It should be noted that for those skilled in the art, several improvements can be made without departing from the principle of this invention, and these improvements should also be considered within the scope of protection of this invention.

Claims

1. A millimeter wave radar based target bounding box detection system, characterized in that, The application relates to a millimeter wave radar target boundary box detection method and device. The millimeter wave radar target boundary box detection method comprises the following steps: The millimeter wave radar target boundary box detection method comprises the following steps: The millimeter wave radar is used for transmitting a millimeter wave frequency-modulated continuous wave signal, receiving a signal reflected from a target, and obtaining a millimeter wave intermediate frequency signal through mixing. The point cloud data generation unit is used for converting the millimeter wave intermediate frequency signal into distance, speed and azimuth angle of the target, and generating point cloud data. The data optimization module is used for removing noise points and enhancing sparse point clouds of the point cloud data to obtain optimized point cloud data. The boundary box detection module is used for performing receiving end beamforming operation on the millimeter wave intermediate frequency signal to obtain a target reflection energy spatial distribution map, fusing the target reflection energy spatial distribution map and the optimized point cloud data to obtain a fused multi-channel feature map, inputting the fused multi-channel feature map into a two-dimensional contour prediction neural network to obtain a predicted target two-dimensional contour image, and finally detecting a boundary box of the target by combining the optimized point cloud data.

2. The millimeter wave radar-based target bounding box detection system of claim 1, wherein, The point cloud data acquisition module transmits a continuous frequency modulation millimeter wave signal with a frequency of 60-64 GHz, and receives the echo signal reflected by the target through multiple receiving antennas. The transmission signal s Tx (t) of the millimeter wave at time t and the received signal s Rx (t) are represented by the following formulas, respectively: s Tx (t) = exp[j(2πf c t + πKt 2 )] s Rx (t) = αs Tx [t - 2R(t) / c] where j is the imaginary unit, f c is the initial frequency of the frequency-modulated continuous wave, K is the rate of change of the frequency, R(t) is the distance between the millimeter wave radar and the target at time t, a is the path loss when reflecting, and c is the speed of light in a vacuum. The intermediate frequency signal s(t) obtained after mixing is expressed as: s(t) = s Tx (t) conj[s Rx (t)] ≈ α exp[j4π(f c + Kt) R(t) / c] wherein conj[s Rx (t)] denotes the conjugate of the received signal s Rx (t). The mixed intermediate frequency signal s(t) contains aliasing of different target reflection signals, and fast Fourier transform is performed on the mixed intermediate frequency signal s(t) to separate reflection signals from different distance targets, and the fast Fourier transform is expressed as: where FFT denotes a fast Fourier transform, S i (t) denotes a signal component corresponding to one target obtained by separation, referred to as a range Fourier signal, R i denotes the distance of the reflection signal from the millimeter wave radar.

3. The millimeter wave radar-based target bounding box detection system of claim 1, wherein, The specific steps of generating point cloud data by the point cloud data acquisition module are as follows: After the first fast Fourier transform, the phase ω of the distance Fourier signal is expressed as: Wherein, ΔR represents a slight displacement of the target between adjacent frequency-modulated signals transmitted by the millimeter wave radar, and lambda represents a wavelength corresponding to a center frequency of the millimeter wave frequency-modulated continuous wave signal. After the two fast Fourier transforms on the millimeter wave intermediate frequency signal, a distance-Doppler matrix is obtained, and the peak point on the distance-Doppler matrix represents a position with strong reflection, which is a target point corresponding to a target to be detected, and the coordinates corresponding to the peak value represent the distance and speed of the target. wherein Δω is the change value of the phase, T c is the time length of the linear frequency modulation; The fast Fourier transform is performed on the multi-channel signals obtained by the different receiving antennas of the millimeter wave radar to obtain the azimuth angle θ of the target in space, and the fast Fourier transform is expressed as: According to the distance, speed and azimuth angle of the target point obtained above, the position corresponding to the detected target point is mapped to the Cartesian coordinate system of space to obtain point cloud data. wherein is the phase difference of the received signal for adjacent receiving antennas, and / is the spacing of the adjacent receiving antennas. The data optimization module removes noise points, and the specific steps are as follows:

4. The millimeter wave radar-based target bounding box detection system of claim 1, wherein, The reflection significance coefficient RS is defined, and the expression is as follows: The data optimization module enhances the sparse point cloud, and the specific steps are as follows: RS = S snr • S v • S a Wherein, S snr is the signal-to-noise ratio significance, S v is the speed significance, S a is the angle significance; the average value of the reflection significance coefficients of all points is calculated as a threshold, and the points with the reflection significance coefficient values lower than the threshold are regarded as noise points and removed, so as to complete the noise point removal of the point cloud data.

5. The millimeter wave radar-based target bounding box detection system of claim 1, wherein, The cumulative matrix is defined, and the continuity of motion in time and space is used to enhance the sparse point cloud; the plane is divided into M*N grids, and the definition of the cumulative matrix BM is as follows: The boundary box detection module realizes the detection of the target boundary box based on the obtained millimeter wave intermediate frequency signal and the optimized point cloud data, and the specific steps are as follows: BM t (m,n) = aBM t-1 (m,n) + count t (m,n) / max(count t ) where (m, n) denotes the coordinates of the grid, BM t (m, n) denotes the value of the cumulative matrix at coordinate (m, n) at time t, BM t-1 (m, n) denotes the value of the cumulative matrix at coordinate (m, n) at time t-1, a denotes the time decay factor, count t (m, n) denotes the number of points in the grid (m, n) at time t, max(count t ) denotes the maximum value of the number of points in the grid; PM t denotes the projection map of the point cloud after denoising in the X and Z axis plane, the cumulative matrix is used to enhance the sparse point cloud projection map, and the expression is as follows: where PM' = PM - P t represents the enhanced point cloud, and γ represents a threshold value.

6. The millimeter wave radar-based target bounding box detection system of claim 1, wherein, (1) Extracting a target reflection energy spatial distribution map based on receiving end beamforming; ​ (2) fuse the target reflected energy spatial distribution map and the optimized point cloud data to obtain a fused multi-channel feature map, take the fused multi-channel feature map as input, and design a two-dimensional contour prediction neural network to predict the target bounding box; (3) based on the prediction result, detect the target bounding box in the three-dimensional coordinate system by using the depth information of the target and the projection rule.

7. The millimeter wave radar-based target bounding box detection system of claim 6, wherein, The step (1) specifically comprises: The multiple intermediate frequency signals obtained by different receiving antennas of the millimeter wave radar are aligned to the same wave front, so that the beam forming effect for a specific β angle direction is achieved, and the signal x after beam forming is obtained β The calculation is as follows: x β = [1, e -jπdsinβ , e -jπ2dsinβ ,..., e -jπ(N-1)dsinβ ]x wherein β represents the included angle between the target and the radial direction of the millimeter wave radar, N represents the number of receiving antennas of the millimeter wave radar, d represents the spacing of adjacent receiving antennas, and x represents the original multi-channel signal vector obtained by different receiving antennas; Through the above method, signals from different directions in space are collected to obtain a target reflected energy spatial distribution map.

8. The millimeter wave radar-based target bounding box detection system of claim 7, wherein, The step (2) specifically comprises: convolve the target reflected energy spatial distribution map and the projection map of the optimized point cloud in the X-axis and Z-axis planes to obtain a fused point cloud contour image, generate a fused multi-channel feature map containing three different dimensions of distance, signal-to-noise ratio and contour as input of the two-dimensional contour prediction neural network, and take the depth image of the target shot by the depth camera as the expected output of the two-dimensional contour prediction neural network; The two-dimensional contour prediction neural network predicts the two-dimensional contour of the target by using the three different dimensions of the fused multi-channel feature map, and the two-dimensional contour prediction neural network is based on the structure of ResNet, starts from a convolution block, passes through a down-sampling layer, transitions through a ResNet block, then passes through an up-sampling layer, and finally ends with a convolution block; a Tanh layer is introduced to L2 regularization to improve the generalization performance of the two-dimensional contour prediction neural network; the loss function of the training process is represented as follows: wherein, is a loss function, is a calculated output value, y is an expected output value, θ is a parameter of L2 regularization, L BCE represents a binary cross-entropy loss, and is represented as follows: wherein m and n represent the coordinates of the plane grid; Take the fused multi-channel feature map as the input of the two-dimensional contour prediction neural network, and take the corresponding depth image as the output to train the two-dimensional contour prediction network; after training, the trained two-dimensional contour prediction neural network is obtained, and the fused multi-channel feature map is converted into a target two-dimensional contour image by the trained two-dimensional contour prediction neural network.

9. The millimeter wave radar-based target bounding box detection system of claim 8, wherein, The step (3) specifically comprises: The target boundary box is reconstructed in a three-dimensional coordinate system by using the target two-dimensional contour image and the distance corresponding to the optimized point cloud. The horizontal field of view and the vertical field of view of the depth camera are denoted as HFOV and VFOV respectively, and the size of the depth image is I height ×I weight The coordinates of a target point are set as (x, y, z), the azimuth and the elevation thereof relative to the millimeter wave radar are denoted as azi and ele respectively, and the mapping of the point on the depth image is denoted as (i, j) according to the projection principle, as follows: convert the point (i, j) on the target two-dimensional contour image into a three-dimensional coordinate system by using the following formula: wherein (x, y, z) is the coordinate of the point in the three-dimensional coordinate system, and range is the distance between the target and the millimeter wave radar; apply the above conversion to each vertex of the target two-dimensional contour image to obtain the bounding box of the target in the three-dimensional coordinate system. 10.A method for target bounding box detection based on millimeter wave radar, based on the system of any one of claims 1-9, characterized in that, The method comprises the following steps: 1) emit a millimeter wave frequency-modulated continuous wave signal; 2) receive the signal reflected by the target, and obtain a millimeter wave intermediate frequency signal after mixing; 3) transform the millimeter wave intermediate frequency signal to obtain a range-Doppler matrix, and determine the distance and speed of the target according to the coordinates of the range-Doppler matrix peak; 4) perform fast Fourier transform on the millimeter wave intermediate frequency signals from different receiving antennas to obtain the azimuth angle of the target; 5) map the distance, speed and azimuth angle of the target point to the Cartesian coordinate system in space to obtain point cloud data; 6) calculate the reflection significance coefficient, and remove the noise points in the point cloud data according to the threshold value; 7) Calculate the cumulative matrix, and enhance the point cloud by using the cumulative matrix; 8) Obtain the spatial distribution map of the target reflection energy by beamforming at different angles for the millimeter wave intermediate frequency signals of different receiving antennas; 9) Convolve the spatial distribution map of the target reflection energy with the projection of the optimized point cloud in the X-Z axis plane to generate a fused multi-channel feature map; 10) Train a two-dimensional contour prediction neural network by taking the fused multi-channel feature map as input, and obtain the trained two-dimensional contour prediction neural network; 11) Input the fused multi-channel feature map into the trained two-dimensional contour prediction neural network to obtain a target two-dimensional contour image; 12) Obtain the bounding box of the target in the spatial coordinate system by using the distance corresponding to the two-dimensional contour image and the optimized point cloud.

Citation Information

Patent Citations

  • Target detection method and device, electronic equipment and storage medium

    CN112766135A

  • Self-learning method and system based on millimeter wave radar target fusion boundary, and vehicle

    CN117115785A