A three-dimensional object detection method, system and storage medium for multi-dimensional fusion of close-range and long-range views
Through the three-dimensional object detection method of multi-dimensional data fusion, Centernet and full convolutional network combined with lidar, millimeter-wave radar and cameras, the problems of low long-range detection accuracy and insufficient robustness in autonomous driving are solved, and accurate positioning of vehicles and pedestrians are achieved.
Patent Information
- Application Number
- CN202310711683.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-06-15
AI Technical Summary
The prior art has low accuracy in detection of three-dimensional targets under long-term and adverse weather conditions in autonomous driving, and a single sensor is not robust enough in harsh environments, making it difficult to obtain depth and speed information at the same time.
The center point is detected by the Centernet and features are extracted through a fully convolutional encoding-decoding backbone network, combined with the multi-dimensional data fusion method of lidar, millimeter-wave radar and camera, close-up and long-range detection tasks are separated, and the speed information of millimeter-wave radar is used as a priori feature to perform three-dimensional target detection.
It improves the accuracy and robustness of three-dimensional target detection in autonomous driving, and can accurately identify and locate targets such as vehicles and pedestrians in different environments.
Smart Images

Figure CN116740519B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a three-dimensional object detection method, system and storage medium for multi-dimensional fusion of near view and far view. Background Art
[0002] With the development of autonomous driving technology, many related researches on object detection have emerged in recent years. In traditional three-dimensional object detection, camera-based methods estimate object classification and position through the semantic information of images. However, since it is difficult to obtain depth information from images, additional computing power is required to estimate the depth information of objects. Most lidar-based detection methods use voxel or point cloud projections for detection. Although point clouds retain the geometric information of objects relatively completely, the sparsity and disorder of point clouds reduce the ability to accurately detect distant objects. Mainstream algorithms that use multiple sensors simultaneously use lidar and cameras to utilize their complementary advantages to achieve high-precision three-dimensional object detection. However, the detection accuracy of this method will decline when facing far views and adverse weather conditions. In addition, cameras and lidars cannot directly obtain key speed data to prevent collisions in many cases. In addition to lidars and cameras, millimeter-wave radars are also widely used in assisted driving. Compared with lidars and cameras, millimeter-wave radars have strong penetration ability and show good robustness in harsh environments. In addition, millimeter-wave radars can accurately detect the relative speed of objects. However, the point clouds of millimeter-wave radars are very sparse and can only be used as a source of depth and speed information, resulting in fewer algorithms for detection using millimeter-wave radars. Therefore, in order to reduce the dependence of algorithms on single sensors, enhance the robustness of algorithms, and improve the detection accuracy of near views and far views. Summary of the Invention
[0003] In order to solve the problems in the prior art, the present invention provides a detection method for multi-dimensional data fusion of near view and far view.
[0004] The present invention provides a three-dimensional object detection method for multi-dimensional fusion of near view and far view, including:
[0005] Step 1: Use Centernet to detect the center points of the image and perform regression on the basic attributes, and extract eigenvalue through a fully convolutional encoding-decoding backbone network;
[0006] Step 2: Use the estimated depth to divide the three-dimensional region of interest (ROI) of the object, and then divide the detection task into near view detection and far view detection.
[0007] The present invention also discloses a three-dimensional object detection system for multi-dimensional fusion of near view and far view, including: a memory, a processor, and a computer program stored on the memory, where the computer program is configured to implement the steps of the three-dimensional object detection method of the present invention when called by the processor.
[0008] The present invention also discloses a computer-readable storage medium storing a computer program, which is configured to implement the steps of the three-dimensional object detection method of the present invention when called by a processor.
[0009] The beneficial effects of the present invention are as follows: The three-dimensional object detection method of the present invention combines the advantages of three sensors, namely lidar, millimeter-wave radar, and camera, to implement the technology of 3D object detection in the field of autonomous driving, and can accurately identify and locate targets such as vehicles, pedestrians, and cyclists, and can be applied to actual scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a framework diagram of the three-dimensional object detection method of the present invention;
[0011] Figure 2 is a schematic diagram of coordinate transformation of the three-dimensional object detection method of the present invention;
[0012] Figure 3 is a schematic diagram of a point cloud detection network of the three-dimensional object detection method of the present invention;
[0013] Figure 4 is a schematic diagram of the relationship between the speed of the radar and the key points in the three-dimensional object detection method of the present invention;
[0014] Figure 5 is a schematic diagram of a rotation estimation network of the three-dimensional object detection method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The present invention discloses a detection method for multi-dimensional data fusion of near and far scenes, which adopts different detection schemes for near and far scenes respectively, shares the detection center point and the data detected later, and improves the utilization rate of data. And the speed information of the millimeter-wave radar is used as a priori and features to be incorporated into the detection network to improve the object detection accuracy.
[0016] As Figures 1-5 shown, the present invention discloses a detection method for multi-dimensional data fusion of near and far scenes, including:
[0017] Step 1: Use Centernet (Center Network) to detect the center point of the image and perform regression on the basic attributes, and extract eigenvalue through a fully convolutional encoding-decoding backbone network;
[0018] The specific content of step 1 is as follows:
[0019] Taking I ∈ R W×H×3 as the input image, with width W and height H, the network generates a heat map about the center point R is the output size scaling ratio, C is the type of the center point, Y x,y,c = 1 indicates that the point (x, y) is a key point under type C, while Y x,y,c = 0 indicates that the point (x, y) is a background point under type c. When training the key point detection network, for the ground truth p ∈ R 2 downsample to get By constructing a Gaussian kernel function as shown in the formula:
[0020]
[0021] Project the ground truth onto the heat map, and adaptively adjust the size of the target 2D detection box through σ. σ is only a parameter for adjusting the size of the 2D detection box. The objective function for training:
[0022]
[0023] where N is the number of targets, is the calibrated ground truth heat map, and α and β are hyperparameters of the loss function. For each detected center point, the network predicts a local offset to compensate for the discretization error caused by downsampling in the backbone network. The network regresses the 2D size, 3D size, target depth, and rotation angle of the target. These values are regressed by the main detection head as shown in Figure 1 . Each main detection head consists of a 3×3 convolutional layer and a 1×1 convolutional layer. The former is used as the input and the latter is used as the output. This detection network provides accurate detection of the target center point and 2D detection box for subsequent human detection, and at the same time realizes the preliminary detection of the target 3D information.
[0024] Step 2: Use the estimated depth to divide the three-dimensional region of interest (ROI) of the target, and then divide the detection task into near-view detection and far-view detection.
[0025] The three-dimensional object detection method of the present invention utilizes the different advantages of different data to create complementary features.
[0026] When performing the detection task for the near view, use the RGB-D data within the frustum to construct features about the target, and use the semantic information brought by the 2D detection as prior information. Specifically include:
[0027] Step S1: First, in order to enhance the rotational invariance of the target and reduce the learning pressure of the network, rotate the coordinate axes along the y-axis so that the rotated z-axis passes through the projection of the peak in the center point heat map on the y-axis, and then construct a rotation matrix R y (θ Δy ) ∈ R 3×3 to convert the global coordinates into local coordinates.
[0028] Step S2: In order to filter out point cloud data that is not related to the target, the present invention segments the point cloud data in the viewing cone. First, the point cloud data converted into local coordinates is input into multiple shared Mlp perceptrons for dimensionality upgrading, and each point in the point cloud data is upgraded to a 1024-dimensional feature vector. The global features of the point cloud data are obtained by the Maxpooling layer while maintaining the disorder of the point cloud. Then the global features are connected to each point and a K-dimensional one-hot vector is added to ensure that the segmentation network can make full use of the prior information brought by 2D detection. The same shared Mlp (perceptron) is used to generate n×1 vectors to segment the point cloud data. Finally, the point on the local coordinate y-axis closest to the centroid of the segmented point cloud is used to preliminarily regress the center of the target.
[0029] There is a big difference between the initial target center obtained during point cloud data segmentation and the real center of the target. In order to accurately regress the target center, the present invention uses a special spatial transformation network (T-Net) and uses the relationship between the two-dimensional target center and the depth value regressed by CenterNet to perform dimensionality reduction processing on T-Net, as shown in the formula:
[0030]
[0031] Wherein, d is the difference between the initial target center depth and the actual target center depth value regressed by the t-space transform network (T-Ne), so during training, the present invention constructs a residual-based loss function as shown in the formula:
[0032] L box =C box -C mask -ΔC T-Net (4)
[0033] Among them, C box Represents the prediction box information, C mask represents the mask prediction information, C T-Net Indicates T-net network prediction information;
[0034] After obtaining the center point of the target, the present invention projects the segmented point cloud within the viewing cone onto the XZ axis to form a point cloud map of a BEV (bird's eye view) bird's eye view.
[0035] In the BEV map, the present invention performs rasterization processing on the projected point cloud data to enhance the network's adaptability to the differences caused by different point cloud densities under different sensors. When extracting features from the BEV map, the present invention performs uniform slicing processing on the BEV map at different heights to keep the height information of the point cloud data as much as possible.
[0036] Since the BEV features have poor global description ability for targets, the present invention fuses Point-Wise features with BEV features. In order to combine information from different features, previous work usually used early fusion or late fusion. The present invention performs hierarchical fusion of multi-view features.
[0037] Suppose and Then it can be concluded that: From (a, b), it can be seen that the point cloud segmentation network has strong robustness to noise in the point cloud, and most of the information carried by Point-Wise (point-level) features is about the key points of the target. Due to the lack of description of local information, Point-Wise (point-level) features are more manifested as a kind of weight. Therefore, the present invention uses the fusion method of Element-Wise multiplication for feature fusion, and its fusion method is shown in the formula:
[0038]
[0039] where f is the feature and H represents the perceptron function;
[0040] And in order to prevent the degradation of the perceptron During the fusion of the two perceptron lines, the present invention adds separate auxiliary loss training for the two perceptrons, and the weights are shared between the fusion training and the auxiliary loss training. Finally, during training, due to the certain indivisibility of the regression task, the present invention jointly optimizes multiple tasks and uses the joint loss function shown:
[0041]
[0042] where i, j, k represent variables and P represents the corresponding point.
[0043] When performing the detection task of long-distance vision, the point cloud generated by the lidar will gradually become sparse as the distance increases, and the RGB data provided by the camera also becomes very blurred. While the millimeter-wave radar can maintain high-precision detection at long distances, the point cloud data of the millimeter-wave radar also faces the problem of too sparse point cloud data. Therefore, the present invention uses the method of combining vision with millimeter-wave radar, and fully utilizes the non-visual features of the millimeter-wave radar on the basis of visual detection to create complementary features for the image. For each detection associated with the target within the cone of vision, a single-channel heat map will be generated. The size of the heat map is proportional to the 2D detection box of the target, and its size is controlled by the parameter α. The value of the heat map is the normalized value of the target depth (d): where M is the normalization factor of the target depth, is the center point coordinate of target j, w jand h j are the length and width of the target 2D detection box respectively. In the present invention, the generated feature map is connected in parallel with the image features of the target, and the corresponding relationship between the feature map and the target center point is used to determine the center point of the target, and a frustum of a cone is constructed to divide the ROI. And the feature map is input into the auxiliary detection head to help the main detection task to regress the target depth and rotation information. Since the millimeter-wave radar only returns the relative radial velocity between the target and the main body, in order to regress the absolute velocity of the target, the absolute velocity of the main body needs to be regressed first. By dividing the ROI, the radar point cloud is divided into two parts: target key points and background points, and the relationship between the background points and the main body velocity is:
[0044]
[0045] where V d n is the radial velocity carried by data point n, θ n is the deflection angle carried by data point n, V n is the magnitude of the target main body velocity obtained by regression, and u is the least square error value between the regression value and the true value. After obtaining the main body velocity, the present invention regresses the absolute velocity of the target.
[0046] For a vehicle target with a relatively large volume, multiple radar key points may be included in its detection box. Therefore, the present invention uses the angular difference of the radial velocities between different key points to regress the absolute velocity of the target:
[0047]
[0048] where V d p is the velocity information carried by target key point p, θ pd is the angular information carried by target key point p, θ pt is the velocity direction of the vehicle target obtained by regression, and V T p is the size of the vehicle target obtained by regression.
[0049] For a person target with a relatively small target scale, since it is difficult to find multiple radar key points in the frustum of a cone, the present invention uses a Part-based Person-Reid algorithm to perform reverse tracking on the person target, and divides the ROI through the radial velocity provided by the millimeter-wave radar, as shown in the formula:
[0050]
[0051] where x t and y t are the target positions of the current frame, x t-1 and y t-1is the position of the target to be tracked in the previous frame, C n is the prediction information of target n; V t is the radial velocity provided by the millimeter-wave radar. The present invention regresses the velocity of the target through the position change of the target between frames, uses the direction of the regressed target velocity as the initial value of the target deflection angle, and focuses on training the difference between the target velocity and the target deflection angle during training. And uses the magnitude and direction of the target velocity as global features to access the auxiliary regression network. The present invention constructs a loss function using residuals, as shown in the formula:
[0052]
[0053] where n is the number of targets, θ T accurate deflection angle, θ i predicted deflection angle, Δθ is the difference between the accurate and predicted deflection angles.
[0054] Finally, the present invention uses the obtained target deflection angle as prior information to correct the size of the target, and uses semantic information to construct a prior size. The constructed loss function is as shown in the formula:
[0055]
[0056] where D * represents the true size, represents the predicted size, and δ is a residual value.
[0057] The present invention also discloses a three-dimensional target detection system for multi-dimensional fusion of near and far scenes, including: a memory, a processor, and a computer program stored on the memory. The computer program is configured to implement the steps of the three-dimensional target detection method of the present invention when called by the processor.
[0058] The present invention also discloses a computer-readable storage medium storing a computer program, and the computer program is configured to implement the steps of the three-dimensional target detection method of the present invention when called by a processor.
[0059] The beneficial effects of the present invention are: The three-dimensional target detection method of the present invention integrates the advantages of three sensors, namely lidar, millimeter-wave radar, and camera, realizes the technology of 3D target detection in the field of autonomous driving, can accurately identify and locate targets such as vehicles, pedestrians, and cyclists, and can be applied in actual scenarios.
[0060] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A three-dimensional object detection method for multi-dimensional fusion of close-range and long-range scenes, characterized in that, Including: Step 1: Use Centernet to detect the center point of the image and regress the basic attributes, and extract eigenvalues through a fully convolutional encoding-decoding backbone network; Step 2: Use the estimated depth to divide the three-dimensional region of interest for the target, and then divide the detection task into near-view detection and far-view detection; In Step 2, for the near-view detection task, use the RGB-D data within the frustum to construct features about the target, and use the semantic information brought by 2D detection as prior information. For the far-view detection task, use the method of combining vision and millimeter-wave radar, and fully utilize the non-visual features of the millimeter-wave radar on the basis of vision detection to create complementary features for the image; The specific content of the near-view detection task includes: Step S1: First, rotate the coordinate axes along the y-axis so that the rotated z-axis passes through the projection of the peak in the center point heat map on the y-axis, and then construct a rotation matrix R y (θ Δy ) to convert the global coordinates into local coordinates; Step S2: Segment the point cloud data in the frustum; first, input the point cloud data converted to local coordinates into multiple shared perceptrons for dimensionality increase processing, and raise each point in the point cloud data to a 1024-dimensional feature vector. Obtain the global feature of the point cloud data through the max-pooling layer while maintaining the disorder of the point cloud. Then connect the global feature with each point and add a K-dimensional one-hot vecter to ensure that the segmentation network fully utilizes the prior information brought by 2D detection. Generate an n×1 vector through the same shared perceptron to segment the point cloud data. Finally, use the point on the y-axis of the local coordinates closest to the centroid of the segmented point cloud to initially regress the center of the target; In Step S2, in order to accurately regress the target center, use the relationship between the two-dimensional target center regressed by Centernet and the depth value to perform dimensionality reduction processing on the spatial transformation network. The specific formula is as follows: Among them, d is the difference between the depth of the preliminary target center regressed by the spatial transformation network and the actual target center depth value; Construct a residual-based loss function during training: L box = C box - C mask - ΔC T-Net (4) Among them, C box represents the prediction box information, C mask represents the mask prediction information, C T-Net represents the T-net network prediction information; After obtaining the center point of the target, project the segmented point cloud within the frustum onto the X-Z axis to form a point cloud map in the BEV bird's-eye view, and in the BEV map, rasterize the projected point cloud data; In order to combine information from different features, adopt the Element-Wise multiplication fusion method to fuse the Point-Wise feature and the BEV feature. The formula for the fusion method is as follows: Among them, f is the feature, and H represents the perceptron function; To prevent the degradation of the perceptron while fusing two sensing lines, separate auxiliary loss training for the two perceptrons is added, and the weights are shared between the fusion training and the auxiliary loss training; During training, jointly optimize multiple tasks. The formula for the joint loss function used is as follows: Among them, i, j, k represent variables, and P represents the corresponding point; The specific content of the far-view detection task is: Parallelly connect the generated feature map with the image feature of the target, use the correspondence between the feature map and the target center point to determine the center point of the target, then construct a frustum to divide the ROI, and input the feature map into the auxiliary detection head to help the main detection task perform regression on the target depth and rotation information; In the far-view detection task, it also includes: Step A1, regress the absolute speed of the main body: Divide the radar point cloud into target key points and background points by dividing the ROI. The relationship between the background points and the main body speed is as follows: Among them, V d n is the radial velocity carried by data point n, and θ n is the deflection angle carried by data point n, and V n is the magnitude of the target body velocity obtained by regression, and u is the least square error value between the regression value and the true value; Step A2, regression of the absolute velocity of the target: For a vehicle target with a large volume, the angular difference of the radial velocities between different key points is used to regress the absolute velocity of the target, and the formula is as follows: Among them, V d p is the speed information carried by the target key point p, and θ pd is the angle information carried by the target key point p, and θ pt is the speed direction of the vehicle target obtained by regression, and V T p is the size of the vehicle target obtained by regression; for the person target with a small target scale, the Person-Reid algorithm is used to perform reverse tracking on the person target, and the ROI is divided by the radial velocity provided by the millimeter-wave radar. The formula is as follows: where x t and y t are the target positions of the current frame, x t-1 and y t-1 are the positions of the target to be tracked in the previous frame, C n is the prediction information of target n; V t is the radial velocity provided by the millimeter-wave radar; In the long-range detection task, a loss function is constructed using residuals, and the formula is as follows: where n is the number of targets, θ T is the accurate deflection angle, θ i is the predicted deflection angle, and Δθ is the difference between the accurate and predicted deflection angles; The obtained target deviation angle is used as prior information to correct the size of the target, and semantic information is used to construct a prior size. The formula for the constructed loss function is as follows: Among them, D * represents the actual size, represents the predicted size, and δ is a residual value.
2. The 3D object detection method according to claim 1, wherein The specific content of the first step is as follows: Taking I as the input image, use Centernet to generate a heatmap about the center point. For each detected center point, Centernet predicts local offsets to compensate for the discretization error caused by downsampling in the fully convolutional encoding-decoding backbone network. Centernet regresses the 2D size, 3D size, object depth, and rotation angle of the object.
3. The three-dimensional object detection method according to claim 2, wherein In the heatmap where Y x,y,c = 1 indicates that the point (x, y) is a key point under type C, and Y x,y,c = 0 indicates that the point (x, y) is a background point under type C. When training the key point detection network, the ground truth p is downsampled to obtain By constructing a Gaussian kernel function, the ground truth is projected onto the heatmap, and σ is used to adaptively adjust the size of the target 2D detection box. The formula for the Gaussian kernel function is as follows: The formula for the objective function during training is as follows: The 2D size, 3D size, target depth, and rotation angle are regressed by the main detection head. Each main detection head consists of a 3×3 convolutional layer and a 1×1 convolutional layer. The 3×3 convolutional layer is used as the input, and the 1×1 convolutional layer is used as the output.
4. A three-dimensional object detection system for multi-dimensional fusion of close-range and long-range views, characterized in that, including: A memory, a processor, and a computer program stored on the memory. The computer program is configured to implement the steps of the 3D object detection method according to any one of claims 1-3 when called by the processor.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the 3D object detection method according to any one of claims 1-3 when called by a processor.
Citation Information
Patent Citations
Three-dimensional target detection system and method based on millimeter wave radar and monocular camera
CN113095154A
Radar distance view and image-based multi-sensor fusion detection method, model and model training method
CN114359664A