An obstacle detection method based on the fusion of 4D radar and image recognition

Through the obstacle detection method of 4D radar and image recognition fusion, the environmental perception problem of unmanned mine locomotives in complex lighting and dust environments inside the mine is solved, and the identification accuracy and ranging accuracy are improved, ensuring the safe and efficient transportation of mine vehicles.

CN117471463BActive Publication Date: 2025-07-11HEFEI GOCOM INFORMATION &TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311427718.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-07-11
Estimated Expiration
2043-10-30

AI Technical Summary

Technical Problem

Under the complex lighting environment and dust conditions inside the mine, the environmental perception accuracy and recognition rate of unmanned mine locomotives based on cameras and lidars are low, affecting safety and efficiency.

Method used

The obstacle detection method of fusion of 4D radar and image recognition is adopted. By calibrating the parameters of the camera and 4D radar, combining millimeter-wave radar point cloud data and image data, deep convolutional network and feature fusion technology are used to generate more accurate environmental representations.

Benefits of technology

It improves image quality in complex lighting environments and radar recognition accuracy in dust environments, enhances the accuracy of environmental perception and ranging accuracy, and ensures the safe operation and transportation efficiency of mine vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117471463B_ABST
    Figure CN117471463B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent assisted driving, and particularly relates to an obstacle detection method based on the fusion of 4D radar and image recognition, including jointly calibrating a 4D millimeter-wave radar and a camera; collecting millimeter-wave radar point cloud data, filtering out noise points including ground point clouds, voxelizing them into cylinders, and extracting multi-scale BEV features through a radar backbone network; collecting image data through the camera and inputting it into a deep convolutional network to extract multi-scale PV features of the image; predicting a radar 3D occupancy grid using the radar BEV feature map; elevating the image PV features to 3D space based on the radar 3D occupancy grid to obtain a BEV feature map of the image; fusing the radar BEV features and the image BEV features, and predicting 3D bounding boxes through a detection head. The present invention effectively fuses 4D millimeter-wave point cloud and camera image features, and can obtain more accurate 3D target detection results compared with three-dimensional target detection in a single mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent assisted driving, and particularly relates to an obstacle detection method based on the fusion of 4D radar and image recognition. Background Art

[0002] An unmanned mine locomotive is an automated device used in mineral resource development. It can automatically and safely carry out transportation operations inside the mine without a human driver. The unmanned mine locomotive plays an important role in resource extraction, ore transportation, personnel transfer, etc., and is of great significance for improving the efficiency of mineral resource extraction, reducing the probability of mine accidents, and reducing casualties caused by operation errors.

[0003] The technology of unmanned mine locomotives based on cameras mainly relies on computer vision to convert the two-dimensional images captured by the cameras into a three-dimensional understanding of the environment. The cameras capture continuous video frames inside the mine, and each frame contains rich environmental texture information. Then, this image data is transmitted to the on-vehicle computing system for image processing and analysis. In the image processing stage, algorithms use methods such as edge detection and feature point extraction to identify important features in the images, such as road edges, obstacles, etc. Based on these features, machine learning algorithms are used for training and prediction to identify and track dynamic and static obstacles in the environment.

[0004] The technology of unmanned mine locomotives based on lidar mainly relies on the distance measurement and target recognition capabilities of lidar. The lidar emits laser pulses and receives the signals reflected back after these pulses collide with objects. By calculating the time difference between emission and reception, the distance between the radar and the object can be accurately calculated. The lidar can not only obtain the distance information of a single point, but also obtain a panoramic depth map by continuously scanning the surrounding environment. This depth map contains a large amount of three-dimensional space information, which can enable the system to better understand the environment. These depth information are sent to the on-vehicle computing system, and through target detection and classification algorithms, different objects such as obstacles, pedestrians, and vehicles can be identified.

[0005] Since the performance of cameras depends to a large extent on the lighting conditions, the lighting environment inside the mine is complex and variable, which may affect the performance of the cameras. In low-light or strong-reflection situations, the cameras may not be able to clearly capture images, resulting in a decrease in the accuracy of environmental perception. Inside the mine, due to the large amount of dust, the penetration of laser light will be affected, which will reduce the recognition accuracy and ranging accuracy of the radar. Summary of the Invention

[0006] To solve the above problems, the present invention provides an obstacle detection method based on the fusion of 4D radar and image recognition.

[0007] The method includes:

[0008] Step 1, calibrate the internal parameters of the camera and the external parameters of the 4D radar;

[0009] Step 2, collect millimeter-wave radar point cloud data, filter out noise points, retain obstacle point clouds, use the SECOND network to construct the radar backbone and neck. The radar backbone extracts multi-scale BEV features from voxelized columns, capturing the spatial and context information inherent in the radar modality. The radar neck aggregates these multi-scale features into one scale to generate a radar BEV feature map;

[0010] Step 3, collect image data through the camera and input it into the backbone of the constructed obstacle detection model to extract multi-scale PV features of the image, obtaining an image PV feature map; the extracted multi-scale PV features of the image are passed through a 1×1 convolution to convert the channels into the depth network interval, obtaining a multi-scale image depth distribution map;

[0011] Step 4, use the radar BEV feature map to predict the radar 3D occupancy grid;

[0012] Step 5, based on the radar 3D occupancy grid, lift the image PV features to 3D space to obtain the BEV features of the image;

[0013] Step 6, splice the BEV features of the radar and the BEV features of the image, and after fusion through convolution operations, use a detection head to predict the 3D bounding box, and determine the position and orientation of the locomotive based on the predicted bounding box.

[0014] Further, the internal parameters of the camera in Step 1 include the camera principal point coordinates and the focal length.

[0015] Further, the external parameters of the 4D radar in Step 1 include the rotation matrix and the translation matrix relative to the camera and the unmanned platform.

[0016] Further, the obstacle detection model in Step 3 is constructed based on a deep convolutional network, uses the CSPDarknet53 deep convolutional neural network as the backbone to extract image features, and adopts the FPN and PAN structures as the neck.

[0017] Further, in Step 3, the number of discretization intervals of the image depth is set to 512.

[0018] Further, Step 4 specifically includes:

[0019] 3D occupancy grid is:

[0020]

[0021] Among them, it is indicated that the input channel is C P , and the output channel is for the 1×1 convolution, Sigmoid represents the activation function, Z is the height of the predefined 3D occupancy grid, and X and Y represent the dimensions of the radar BEV feature map.

[0022] Furthermore, step five specifically includes:

[0023] Project the coordinates of the predefined 3D voxels onto the image plane, and obtain the 3D voxel features of the image by bilinearly sampling the image PV feature map;

[0024] Perform trilinear sampling on the multi-scale image depth distribution map to obtain the voxel sampling depth probability in the radar coordinate system;

[0025] Multiply the 3D voxel features of the image and the voxel sampling depth probability element-wise to obtain the depth-assisted image 3D features;

[0026] Multiply the depth-assisted image 3D features by the 3D occupancy grid to obtain the radar-assisted image 3D features;

[0027] Concatenate the depth-assisted image 3D features and the radar-assisted image 3D features along the channel dimension and sum them in the scale dimension, and adjust the number of channels through 1×1 convolution to obtain the BEV features of the image.

[0028] Furthermore, the fusion strategy in step six uses attention-based fusion technology.

[0029] Furthermore, the 4D radar is a 4D millimeter-wave radar.

[0030] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0031] The present application adopts advanced image processing technologies, including high dynamic range (HDR) processing and noise suppression technologies, effectively improving the image quality in complex lighting environments, thereby improving the accuracy of environmental perception. For the dust problem of lidar, the present application can distinguish the reflections from dust particles and the reflections from actual obstacles from the received mixed signals by using advanced signal processing and machine learning technologies, thereby improving the recognition accuracy and ranging accuracy of the radar in a dusty environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is the schematic diagram of the obstacle detection method based on the fusion of 4D radar and image recognition provided by the embodiment of the present invention;

[0033] Figure 2Flowchart of the obstacle detection method based on the fusion of 4D radar and image recognition provided by the embodiments of the present invention;

[0034] Figure 3 Schematic diagram provided by the embodiments of the present invention when the obstacle in front of the locomotive is within the safety warning line;

[0035] Figure 4 Schematic diagram provided by the embodiments of the present invention when the obstacle in front of the locomotive is not within the safety warning line. Detailed implementation manners

[0036] The following combines the accompanying drawings and specific embodiments to describe the present invention in detail. Before elaborating on the technical solutions of the embodiments of the present invention, the nouns and terms involved are explained. In this specification, components with the same name or the same reference numeral represent similar or the same structures, and are for illustrative purposes only.

[0037] The obstacle detection method provided by the present invention can be applied to ordinary vehicles, such as obstacle detection of an automobile in a low-illumination environment, and can also be applied to locomotives, such as obstacle detection of a mine locomotive (single-track locomotive, multi-track locomotive) in a low-illumination environment (such as in a roadway). Mine locomotives can be further divided into unmanned (driverless) locomotives and manned (drivered) locomotives. This embodiment will take the obstacle detection method of a mine locomotive, such as an unmanned locomotive running in a roadway, as an example to elaborate on the technical solution of the present invention, and other vehicles such as single-track cranes and rubber-tired vehicles will not be described repeatedly.

[0038] As Figure 1 shown, the millimeter-wave radar and camera installed on the vehicle are connected to the AI processing module through Ethernet, and the detection result of the AI processing module is directly sent to the vehicle controller.

[0039] As Figure 2 shown, the steps of the obstacle detection method provided by the present invention are as follows:

[0040] 1. Calibrate the parameters of the camera and 4D millimeter-wave radar

[0041] The radar and the camera are two different sensing devices, and the data information and representation forms they capture are very different. The radar mainly provides distance information and reflection intensity, generating a set of three-dimensional point cloud data; while the camera mainly provides color and texture information, generating two-dimensional image data. Each of these two types of data has its own advantages and disadvantages. For example, radar data is less affected by lighting and weather conditions, but has a lower spatial resolution; camera data can provide rich color and texture information, has a higher spatial resolution, but is sensitive to lighting and weather conditions.

[0042] Aligning and fusing the data from the radar and the camera can combine the advantages of both and avoid their respective disadvantages, thus obtaining a more comprehensive and accurate environmental representation.

[0043] First, calibrate the internal parameters of the camera, including the coordinates of the camera's principal point and the focal length. These parameters are crucial for determining the geometric relationship between the camera and the object, and with the help of these parameters, the two-dimensional image coordinates in the camera can be converted into three-dimensional world coordinates. In addition, it is also necessary to calibrate the external parameters of the 4D millimeter-wave radar, including the rotation matrix and the translation matrix relative to the camera and the unmanned platform. Based on these parameters, the radar data can be aligned with the camera data, thus realizing data fusion.

[0044] 2. Collect and process the millimeter-wave radar point cloud data

[0045] Collect the point cloud data of the millimeter-wave radar. By filtering out the noise points, including the ground point cloud, and retaining the obstacle point cloud, a clearer environmental representation can be obtained. For example, some points with a height lower than a preset threshold are found in the radar's point cloud data. These points may come from the ground and are filtered out as noise points. Then, use the SECOND network to construct the radar backbone and neck. The radar backbone extracts multi-scale BEV (Bird's Eye View) features from the voxelized columns, capturing the spatial and context information inherent in the radar modality. The radar neck aggregates these multi-scale features into one scale to generate the radar BEV feature map Facilitating subsequent fusion and analysis. where X and Y represent the dimensions of the radar BEV feature map, and C P represents the number of channels.

[0046] 3. Collect and process the image data

[0047] Collect the image data through the camera and input it into the obstacle detection model to extract the multi-scale PV (Perspective View) features of the image, obtaining the image PV feature map. The obstacle detection model is constructed based on a deep convolutional network, using the CSPDarknet53 deep convolutional neural network as the backbone to extract image features, and adopting the FPN (Feature Pyramid Network) and PAN (Path Aggregation Network) structures as the neck to further process and fuse the features extracted by the backbone. The extracted image features are convolved to convert the channels into the depth network interval, obtaining the multi-scale image depth distribution map In this embodiment, the discretization interval number of the image depth is set to 512, corresponding to a radar detection distance of 50 meters, about 0.1 meter per interval.

[0048] 4. Predicting 3D Occupancy Grid Using Radar Data

[0049] Predict the radar 3D occupancy grid using the radar BEV feature map. The height Z of the 3D occupancy grid is predefined. The 3D occupancy grid is:

[0050]

[0051] where represents a 1×1 convolution with input channels C P and output channels , and sigmoid represents the activation function.

[0052] 5. Obtaining the Image BEV Feature Map

[0053] Based on view transformation technology, lift the PV (Per-View) features of the image to 3D space so that they can be compared and fused with the 3D data of the radar. For example, there is an image of a car and the radar return data. The car may only occupy a small part in the image, but may occupy a large 3D space in the radar data. Lift the features of the image to 3D space to compare and fuse these two types of data.

[0054] Project the coordinates of the predefined 3D voxels onto the image plane and obtain the 3D voxel features of the image by bilinear sampling of the image PV feature map to combine the three-dimensional spatial information of the radar with the two-dimensional image information of the camera, so that the image information collected by the camera can have spatiality.

[0055] Perform trilinear sampling (i.e., the 3D version of bilinear sampling) on the multi-scale image depth distribution map to obtain the voxel sampling depth probability in the radar coordinate system to obtain depth information in three-dimensional space, so as to better understand the position and shape of objects in space,

[0056] Multiply the 3D voxel features of the image and the voxel sampling depth probability element-wise to obtain the depth-assisted image 3D features.

[0057] Multiply the depth-assisted image 3D features with the 3D occupancy grid to obtain the radar-assisted image 3D features.

[0058] Concatenate the depth-assisted image 3D features and the radar-assisted image 3D features along the channel dimension and sum them in the scale dimension, and adjust the number of channels by 1×1 convolution, and finally obtain the BEV features of the image.

[0059] 6. Feature Fusion and 3D Bounding Box Prediction

[0060] The BEV features of the radar and the BEV features of the image are concatenated and fused through convolutional operations. Then, a detection head is used to predict the 3D bounding box. This bounding box can help determine the position and orientation of the locomotive, thus assisting the driverless vehicle in navigation and decision-making.

[0061] Note that the fusion strategy here is not limited to a specific method. For example, attention-based fusion techniques can be used to better focus on important features; or an anchor box-based detection head can be used.

[0062] As an example, Figure 3 and Figure 4 respectively show schematic diagrams when the obstacle in front of the locomotive is within and outside the safety warning line. As Figure 3 shown, the line segment OB is the measured distance of the obstacle by the scanning beam, the reference distance line segment is OC, OB is less than the reference distance OC, and the perpendicular distance OA from OB to the track is less than the preset safety distance x0, indicating that the obstacle is within the safety warning line and safety braking and stopping need to be implemented. However, for Figure 4 the shown situation, although the scanning distance OB is also less than the reference distance OC, but the perpendicular distance OA from OB to the track is greater than the preset safety distance x0, indicating that the obstacle is outside the safety warning line. In such cases, braking may not be necessary, but early warning driving can be added.

[0063] The present invention provides a mine obstacle detection method, which makes a fusion judgment based on the point cloud data of the 4D millimeter-wave radar unit and combines with the image data to control the running state of the vehicle. The present invention can be applied to the obstacle detection of mine driverless vehicles in low-illumination environments, reducing the number of underground transportation workers; at the same time, the obstacle detection function of the obstacle detection system of the present invention ensures the driving safety during the vehicle transportation process, generally improving the transportation efficiency of mine vehicles and realizing "personnel reduction and efficiency increase".

[0064] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An obstacle detection method based on the fusion of 4D radar and image recognition, characterized in that, It includes the following steps: Step 1, calibrate the internal parameters of the camera and the external parameters of the 4D radar; Step 2, collect millimeter-wave radar point cloud data, filter out noise points, and retain obstacle point clouds. Use the SECOND network to construct the radar backbone and neck. The radar backbone extracts multi-scale BEV features from voxelized columns, capturing the spatial and context information inherent in the radar modality. The radar neck aggregates these multi-scale features into one scale to generate a radar BEV feature map; Step 3, collect image data through the camera and input it into the backbone of the constructed obstacle detection model to extract multi-scale PV features of the image, obtaining an image PV feature map; the extracted multi-scale PV features of the image pass through a 1×1 convolution to convert the channels into the depth network interval, obtaining a multi-scale image depth distribution map; Step 4, use the radar BEV feature map to predict the radar 3D occupancy grid, where the 3D occupancy grid specifically includes: 3D Occupancy Grid is as follows: Among them, indicates that the input channel is C P , C P represents the number of channels, and the output channel is a 1×1 convolution, Sigmoid represents the activation function, Z is the height of the predefined 3D occupancy grid, and X and Y represent the dimensions of the radar BEV feature map; Step 5, lift the image PV features to 3D space based on the radar 3D occupancy grid to obtain the BEV features of the image, which specifically includes: Project the coordinates of the predefined 3D voxels onto the image plane and obtain the 3D voxel features of the image by bilinear sampling of the image PV feature map; Perform trilinear sampling on the multi-scale image depth distribution map to obtain the voxel sampling depth probability in the radar coordinate system; Multiply the 3D voxel features of the image and the voxel sampling depth probability element-wise to obtain depth-assisted image 3D features; Multiply the depth-assisted image 3D features by the 3D occupancy grid to obtain radar-assisted image 3D features; Concatenate the depth-assisted image 3D features and the radar-assisted image 3D features along the channel dimension and sum them in the scale dimension, and adjust the number of channels through a 1×1 convolution to obtain the BEV features of the image; Step 6, concatenate the BEV features of the radar and the BEV features of the image, fuse them through convolution operations, and then use a detection head to predict the 3D bounding box, and determine the position and orientation of the locomotive based on the predicted bounding box.

2. The obstacle detection method based on the fusion of 4D radar and image recognition according to claim 1, wherein, The internal parameters of the camera described in Step 1 include the camera principal light point coordinates and the focal length.

3. The obstacle detection method based on the fusion of 4D radar and image recognition according to claim 1, characterized in that, The external parameters of the 4D radar described in Step 1 include the rotation matrix and the translation matrix relative to the camera and the unmanned platform.

4. The method for obstacle detection based on the fusion of 4D radar and image recognition according to claim 1, wherein, The obstacle detection model described in Step 3 is constructed based on a deep convolutional network, uses the CSPDarknet53 deep convolutional neural network as the backbone to extract image features, and adopts the FPN and PAN structures as the neck.

5. The obstacle detection method based on the fusion of 4D radar and image recognition according to claim 1, characterized in that, In Step 3, the number of discretization intervals of the image depth is set to 512.

6. The obstacle detection method based on the fusion of 4D radar and image recognition according to claim 1, wherein, The fusion strategy in Step 6 uses an attention-based fusion technique.

7. The obstacle detection method based on the fusion of 4D radar and image recognition as claimed in claim 1, wherein, The 4D radar is a 4D millimeter-wave radar.

Citation Information

Patent Citations

  • Attention-based 4D millimeter wave radar and vision fusion method

    CN116129234A

  • Visual three-dimensional perception method and system based on three-dimensional occupancy prediction and neural rendering

    CN116934977A

  • 3D target detection system and method based on 4D millimeter wave radar and camera fusion

    CN117452396A