Object Detection Method, 3D Object Detection Model Training Method and Device

By combining point cloud and image features in machine vision technology and using multiple decoding and update methods, the problem of visual sensors being susceptible to external interference is solved, and the accuracy of object detection is improved.

CN119810820BActive Publication Date: 2025-07-01HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510293256.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-01
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

In machine vision technology, the image collected by vision sensors is susceptible to external interference such as light changes, background color changes and occlusions, resulting in low accuracy of target detection results.

Method used

A target detection method is adopted to obtain the point clouds and images of the space area to be detected, divide and extract the point clouds according to the bird's-eye view BEV grid, and combine the image features to decode and update multiple times until the preset number is reached, so as to improve the accuracy of target detection.

Benefits of technology

By combining point clouds and image features, the impact of external interference on target detection results can be reduced and the accuracy of target detection results can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810820B_ABST
    Figure CN119810820B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an object detection method, a three-dimensional object detection model training method and device, which relate to the field of machine vision technology. The object detection method includes: obtaining a point cloud to be detected and an image to be detected in a spatial region to be detected; extracting an initial point cloud feature to be detected and obtaining an initial image feature to be detected; obtaining a fused point cloud feature to be detected of the current reference three-dimensional point based on the current coordinates and reference features of each reference three-dimensional point; obtaining an intermediate image feature to be detected of the current reference three-dimensional point based on the initial image feature to be detected; updating the current coordinates and reference features of the reference three-dimensional point based on the current fused point cloud feature to be detected and the intermediate image feature to be detected; returning to execute the step of extracting the initial point cloud feature to be detected until a preset number of decoding times; obtaining the position of an object in the spatial region to be detected based on the current coordinates of each reference three-dimensional point. The accuracy of the obtained object detection result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine vision technology, and in particular, to an object detection method, a three-dimensional object detection model training method, and an apparatus. Background Art

[0002] In the field of machine vision technology, a vision sensor can be used to collect images of a space area to be detected, and by performing object detection on the collected images, a detection result of the position of an object in the space area to be detected can be obtained.

[0003] However, the process of collecting images by a vision sensor is vulnerable to external interferences such as illumination changes, background color changes, and occluders, which may lead to low accuracy of the object detection result. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide an object detection method, a three-dimensional object detection model training method, and an apparatus to improve the accuracy of the obtained object detection result. The specific technical solutions are as follows:

[0005] In the first aspect of the embodiments of this application, first, an object detection method is provided. The method includes: obtaining a point cloud and an image of a space area to be detected as a point cloud to be detected and an image to be detected respectively; dividing the point cloud to be detected according to a bird's-eye view (BEV) grid, and extracting the initial point cloud features to be detected of each obtained BEV grid, and performing feature extraction on the image to be detected to obtain the initial image features to be detected; for each reference three-dimensional point, sampling the BEV grids within the neighborhood range of the reference three-dimensional point based on the current coordinates and reference features of the reference three-dimensional point, and fusing the initial point cloud features to be detected of the sampled BEV grids to obtain the current fused point cloud features to be detected of the reference three-dimensional point; obtaining the corresponding part of the reference three-dimensional point in the initial image features to be detected to obtain the current intermediate image features to be detected of the reference three-dimensional point; splicing the current fused point cloud features to be detected and the current intermediate image features to be detected of each reference three-dimensional point to obtain the current spliced features to be detected; decoding the current spliced features to be detected to obtain the current decoding result, and updating the current coordinates and reference features of the reference three-dimensional point based on the current decoding result; returning to execute the step of for each reference three-dimensional point, sampling the BEV grids within the neighborhood range of the reference three-dimensional point based on the current coordinates and reference features of the reference three-dimensional point, and fusing the initial point cloud features to be detected of the sampled BEV grids to obtain the current fused point cloud features to be detected of the reference three-dimensional point until the number of decoding times reaches a preset number; obtaining the positions of objects in the space area to be detected based on the current coordinates of each reference three-dimensional point.

[0006] Optionally, dividing the point cloud to be detected according to the bird's-eye view BEV grid, extracting the initial point cloud features to be detected of each obtained BEV grid, and extracting the features of the image to be detected to obtain the initial image features to be detected, includes: dividing the point cloud to be detected according to the bird's-eye view BEV grid, and determining the three-dimensional points to be detected included in each BEV grid obtained by the division; using the point cloud feature extraction network in the pre-trained three-dimensional object detection model to extract the features of the information to be utilized of the three-dimensional points to be detected included in each BEV grid, to obtain the initial point cloud features to be detected of each BEV grid; and inputting the image to be detected into the image feature extraction network in the three-dimensional object detection model to obtain a plurality of initial image features to be detected; wherein, the information to be utilized of a three-dimensional point includes: the coordinates of the three-dimensional point, and the coordinates of the centroid of the BEV grid to which the three-dimensional point belongs; the three-dimensional object detection model is: trained based on the sample images, sample point clouds in the sample space region, and sample labels indicating the positions of the objects in the sample space region; the three-dimensional object detection model further includes a deformable attention network, a weight attention network, and a decoder; for each reference three-dimensional point, sampling the BEV grids within the neighborhood range of the reference three-dimensional point based on the current coordinates and reference features of the reference three-dimensional point, and fusing the initial point cloud features to be detected of the sampled BEV grids, to obtain the current fused point cloud features to be detected of the reference three-dimensional point, includes: for each reference three-dimensional point, inputting the current reference features of the reference three-dimensional point into the deformable attention network to obtain a preset number of position offsets and the weights corresponding to each position offset; for each position offset, using the position offset to offset the current coordinates of the reference three-dimensional point, and determining the BEV grid to which the offset coordinates belong as the BEV grid corresponding to the position offset; fusing the initial point cloud features to be detected of the BEV grids corresponding to each position offset according to the weights corresponding to each position offset, to obtain the current fused point cloud features to be detected of the reference three-dimensional point; obtaining the corresponding part of the reference three-dimensional point in the initial image features to be detected, to obtain the current intermediate image features to be detected of the reference three-dimensional point, includes: inputting the current reference features of the reference three-dimensional point into the weight attention network to obtain the weights of each initial image feature to be detected; fusing the eigenvalue corresponding to the reference three-dimensional point in each initial image feature to be detected according to the weights of each initial image feature to be detected, to obtain the current intermediate image features to be detected of the reference three-dimensional point; decoding the current concatenated features to be detected to obtain the current decoding result, includes: inputting the current concatenated features to be detected into the decoder to obtain the current decoding result.

[0007] Optionally, the 3D object detection model further includes a multi-layer perceptron; before inputting the current reference feature of each reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset, the method further includes: inputting the initial point cloud features to be detected of each BEV grid and the initial image features to be detected into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix; wherein, the initial transformation matrix is calibrated according to the acquisition methods of the point cloud to be detected and the image to be detected; correcting the initial transformation matrix by using the obtained correction amount; the step of fusing the feature values corresponding to the reference 3D point in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference 3D point includes: converting the current coordinates of the reference 3D point by using the corrected transformation matrix to obtain the pixel coordinates corresponding to the reference 3D point in the image to be detected as the reference coordinates to be detected; fusing the feature values corresponding to the reference coordinates to be detected in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference 3D point.

[0008] Optionally, the step of inputting the initial point cloud features to be detected of each BEV grid and the initial image features to be detected into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix includes: inputting the initial point cloud features to be detected of each BEV grid, the initial image features to be detected, and the current coordinates of each reference 3D point into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix.

[0009] Optionally, before fusing the feature values corresponding to the reference coordinates to be detected in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference 3D point, the method further includes: if the reference coordinates to be detected are not integers, obtaining other pixel coordinates within the neighborhood range of the reference coordinates to be detected; for each initial image feature to be detected, interpolating the feature values corresponding to the obtained other pixel coordinates in the initial image feature to be detected as the feature values corresponding to the reference coordinates to be detected in the initial image feature to be detected.

[0010] Optionally, the image feature extraction network includes a plurality of downsampling layers connected in series; there is one image to be detected, and the plurality of initial image features to be detected include: the image features of the image to be detected extracted by the plurality of downsampling layers; or, the image to be detected includes images of the space area to be detected collected at multiple different shooting angles, and the plurality of initial image features to be detected include the image features of each image to be detected.

[0011] Optionally, determining the three-dimensional points to be detected included in each BEV grid obtained by partitioning includes: for each BEV grid obtained by partitioning, if the number of three-dimensional points in the point cloud to be detected included in this BEV grid is greater than the specified number, then sample the specified number of three-dimensional points from the three-dimensional points included in this BEV grid as the three-dimensional points to be detected included in this BEV grid; if the number of three-dimensional points in the point cloud to be detected included in this BEV grid is equal to the specified number, then use the three-dimensional points included in this BEV grid as the three-dimensional points to be detected included in this BEV grid; if the number of three-dimensional points in the point cloud to be detected included in this BEV grid is less than the specified number, then calculate the difference between the specified number and the number of three-dimensional points in the point cloud to be detected included in this BEV grid, and generate the difference number of three-dimensional points with preset coordinates; combine the generated three-dimensional points and the three-dimensional points in the point cloud to be detected included in this BEV grid to obtain the three-dimensional points to be detected included in this BEV grid.

[0012] Optionally, the information to be utilized of a three-dimensional point further includes at least one of the following: the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, and the feature data at the real position corresponding to the detected three-dimensional point.

[0013] In a second aspect of the embodiments of the present application, a method for training a three-dimensional object detection model is provided. The method includes: obtaining the point cloud and image of a sample space region as a sample point cloud and a sample image respectively, and obtaining a sample label including the position of an object within the sample space region; dividing the sample point cloud according to a bird's-eye view (BEV) grid and determining the sample three-dimensional points included in each BEV grid obtained by the division; using a point cloud feature extraction network in the three-dimensional object detection model with an initial structure to extract features from the information to be utilized of the sample three-dimensional points included in each BEV grid, obtaining the sample initial point cloud features of each BEV grid; and inputting the sample image into an image feature extraction network in the three-dimensional object detection model to obtain a plurality of sample initial image features; wherein, the information to be utilized of a three-dimensional point includes: the coordinates of the three-dimensional point and the coordinates of the centroid of the BEV grid to which the three-dimensional point belongs; the three-dimensional object detection model further includes a deformable attention network, a weight attention network, and a decoder; for each reference three-dimensional point, inputting the current reference feature of the reference three-dimensional point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset; for each position offset, using the position offset to offset the current coordinates of the reference three-dimensional point and determining the BEV grid to which the offset coordinates belong as the BEV grid corresponding to the position offset; fusing the sample initial point cloud features of the BEV grids corresponding to each position offset according to the weights corresponding to each position offset to obtain the current sample fused point cloud feature of the reference three-dimensional point; inputting the current reference feature of the reference three-dimensional point into the weight attention network to obtain the weight of each sample initial image feature; fusing the eigenvalue corresponding to the reference three-dimensional point in each sample initial image feature according to the weights of each sample initial image feature to obtain the current sample intermediate image feature of the reference three-dimensional point; splicing the current sample fused point cloud feature and the current sample intermediate image feature of each reference three-dimensional point to obtain the current sample spliced feature; inputting the current sample spliced feature into the decoder to obtain the current decoding result, and updating the current coordinates and reference feature of the reference three-dimensional point based on the current decoding result; returning to execute the step of for each reference three-dimensional point, inputting the current reference feature of the reference three-dimensional point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset until the number of decoding times reaches a preset number; obtaining a detection result based on the current coordinates of each reference three-dimensional point; wherein, the detection result includes the position of the object in the sample space region; adjusting the model parameters of the three-dimensional object detection model with the initial structure based on the difference between the sample label and the detection result until a preset convergence condition is reached, obtaining the trained three-dimensional object detection model.

[0014] Optionally, the 3D object detection model further includes a multi-layer perceptron; before inputting the current reference feature of each reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset, the method further includes: inputting the sample initial point cloud features of each BEV grid and the sample initial image features into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix; wherein, the initial transformation matrix is calibrated according to the acquisition methods of the sample point cloud and the sample image; correcting the initial transformation matrix by using the obtained correction amount; the step of fusing the eigenvalues corresponding to the reference 3D point in each sample initial image feature according to the weights of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point includes: using the corrected transformation matrix to transform the current coordinates of the reference 3D point to obtain the pixel coordinates corresponding to the reference 3D point in the sample image as the sample reference coordinates; fusing the eigenvalues corresponding to the sample reference coordinates in each sample initial image feature according to the weights of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point.

[0015] Optionally, the step of inputting the sample initial point cloud features of each BEV grid and the sample initial image features into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix includes: inputting the sample initial point cloud features of each BEV grid, the sample initial image features, and the current coordinates of each reference 3D point into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix.

[0016] Optionally, before fusing the eigenvalues corresponding to the sample reference coordinates in each sample initial image feature according to the weights of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point, the method further includes: if the sample reference coordinates are not integers, obtaining other pixel coordinates within the neighborhood range of the sample reference coordinates; for each sample initial image feature, interpolating the eigenvalues corresponding to the obtained other pixel coordinates in the sample initial image feature as the eigenvalues corresponding to the sample reference coordinates in the sample initial image feature.

[0017] Optionally, the image feature extraction network includes a plurality of downsampling layers connected in series; there is one sample image, and the plurality of sample initial image features include: the image features of the sample image extracted by the plurality of downsampling layers; or, the sample image includes images of the sample spatial region collected at multiple different shooting angles, and the plurality of sample initial image features include the image features of each sample image.

[0018] Optionally, determining the sample three-dimensional points included in each BEV grid obtained by partitioning includes: for each BEV grid obtained by partitioning, if the number of three-dimensional points in the sample point cloud included in this BEV grid is greater than a specified number, sampling the specified number of three-dimensional points from the three-dimensional points included in this BEV grid as the sample three-dimensional points included in this BEV grid; if the number of three-dimensional points in the sample point cloud included in this BEV grid is equal to the specified number, taking the three-dimensional points included in this BEV grid as the sample three-dimensional points included in this BEV grid; if the number of three-dimensional points in the sample point cloud included in this BEV grid is less than the specified number, calculating the difference between the specified number and the number of three-dimensional points in the sample point cloud included in this BEV grid, and generating three-dimensional points with the difference number of coordinates being preset values; combining the generated three-dimensional points and the three-dimensional points in the sample point cloud included in this BEV grid to obtain the sample three-dimensional points included in this BEV grid.

[0019] Optionally, the information to be utilized for a three-dimensional point further includes at least one of the following: the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, and the feature data at the true position corresponding to the detected three-dimensional point.

[0020] Optionally, the sample label further includes the category of the object within the sample space region; the position indication of the object within the sample space region: the size of the minimum bounding cube of the object, the coordinates of the center point of the minimum bounding cube of the object, and the orientation of the object; the detection result further includes the category of the object within the sample space region; the differences between the sample label and the detection result include: the position difference of the object within the sample space region, and the category difference of the object within the sample space region.

[0021] In a third aspect of the embodiments of the present application, a target detection device is provided. The device includes: a first acquisition module, configured to acquire a point cloud and an image of a space area to be detected, respectively as a point cloud to be detected and an image to be detected; a first feature extraction module, configured to divide the point cloud to be detected according to a bird's-eye view (BEV) grid, and extract the initial point cloud features to be detected of each obtained BEV grid, and perform feature extraction on the image to be detected to obtain the initial image features to be detected; a first point cloud feature fusion module, configured to, for each reference three-dimensional point, sample the BEV grids within the neighborhood range of the reference three-dimensional point based on the current coordinates and reference features of the reference three-dimensional point, and fuse the initial point cloud features to be detected of the sampled BEV grids to obtain the current fused point cloud features to be detected of the reference three-dimensional point; a first image feature fusion module, configured to obtain the corresponding part of the reference three-dimensional point in the initial image features to be detected to obtain the current intermediate image features to be detected of the reference three-dimensional point; a first splicing module, configured to splice the current fused point cloud features to be detected and the current intermediate image features to be detected of each reference three-dimensional point to obtain the current spliced features to be detected; a first update module, configured to decode the current spliced features to be detected to obtain the current decoding result, and update the current coordinates and reference features of the reference three-dimensional point based on the current decoding result; return to execute the step of, for each reference three-dimensional point, sampling the BEV grids within the neighborhood range of the reference three-dimensional point based on the current coordinates and reference features of the reference three-dimensional point, and fusing the initial point cloud features to be detected of the sampled BEV grids to obtain the current fused point cloud features to be detected of the reference three-dimensional point, until the number of decoding times reaches a preset number; a position acquisition module, configured to obtain the positions of the objects in the space area to be detected based on the current coordinates of each reference three-dimensional point.

[0022] Optionally, the first feature extraction module is specifically configured to divide the to-be-detected point cloud according to the bird's-eye view (BEV) grid, and determine the to-be-detected three-dimensional points included in each BEV grid obtained by the division; use the point cloud feature extraction network in the pre-trained three-dimensional object detection model to extract features from the to-be-used information of the to-be-detected three-dimensional points included in each BEV grid, and obtain the to-be-detected initial point cloud features of each BEV grid; and input the to-be-detected image into the image feature extraction network in the three-dimensional object detection model to obtain a plurality of to-be-detected initial image features; wherein, the to-be-used information of a three-dimensional point includes: the coordinates of the three-dimensional point and the coordinates of the centroid of the BEV grid to which the three-dimensional point belongs; the three-dimensional object detection model is: trained based on the sample images, sample point clouds in the sample space region, and sample labels indicating the positions of the objects in the sample space region; the three-dimensional object detection model further includes a deformable attention network, a weight attention network, and a decoder; the first point cloud feature fusion module is specifically configured to, for each reference three-dimensional point, input the current reference feature of the reference three-dimensional point into the deformable attention network to obtain a preset number of position offsets and the weights corresponding to each position offset; for each position offset, use the position offset to offset the current coordinates of the reference three-dimensional point, and determine the BEV grid to which the offset coordinates belong as the BEV grid corresponding to the position offset; fuse the to-be-detected initial point cloud features of the BEV grids corresponding to each position offset according to the weights corresponding to each position offset to obtain the to-be-detected fused point cloud feature of the reference three-dimensional point at present; the first image feature fusion module is specifically configured to input the current reference feature of the reference three-dimensional point into the weight attention network to obtain the weights of each to-be-detected initial image feature; fuse the feature values corresponding to the reference three-dimensional point in each to-be-detected initial image feature according to the weights of each to-be-detected initial image feature to obtain the to-be-detected intermediate image feature of the reference three-dimensional point at present; the first update module is specifically configured to input the current to-be-detected splicing feature into the decoder to obtain the current decoding result.

[0023] Optionally, the 3D object detection model further includes a multi-layer perceptron; the apparatus further includes: a first correction amount acquisition module, configured to, before the first point cloud feature fusion module performs, for each reference 3D point, inputting the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset, input the initial point cloud features to be detected of each BEV grid and the initial image features to be detected into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix; wherein, the initial transformation matrix is calibrated according to the acquisition manners of the point cloud to be detected and the image to be detected; and correct the initial transformation matrix by using the obtained correction amount; the first image feature fusion module is specifically configured to use the corrected transformation matrix to transform the current coordinates of the reference 3D point to obtain the pixel coordinates corresponding to the reference 3D point in the image to be detected as the reference coordinates to be detected; and fuse the feature values corresponding to the reference coordinates to be detected in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the intermediate image feature to be detected of the reference 3D point currently.

[0024] Optionally, the first correction amount acquisition module is specifically configured to input the initial point cloud features to be detected of each BEV grid, the initial image features to be detected, and the current coordinates of each reference 3D point into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix.

[0025] Optionally, the apparatus further includes: a first interpolation module, configured to, before the first image feature fusion module performs fusing the feature values corresponding to the reference coordinates to be detected in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the intermediate image feature to be detected of the reference 3D point currently, if the reference coordinates to be detected are not integers, obtain other pixel coordinates within the neighborhood range of the reference coordinates to be detected; and for each initial image feature to be detected, perform interpolation on the feature values corresponding to the obtained other pixel coordinates in the initial image feature to be detected as the feature values corresponding to the reference coordinates to be detected in the initial image feature to be detected.

[0026] Optionally, the image feature extraction network includes a plurality of downsampling layers connected in series; there is one image to be detected, and the plurality of initial image features to be detected include: the image features of the image to be detected extracted by the plurality of downsampling layers; or, the image to be detected includes images of the space region to be detected collected at multiple different shooting angles, and the plurality of initial image features to be detected include the image features of each image to be detected.

[0027] Optionally, the first feature extraction module is specifically configured to, for each BEV grid obtained by division, if the number of three-dimensional points in the point cloud to be detected included in the BEV grid is greater than a specified number, sample the specified number of three-dimensional points from the three-dimensional points included in the BEV grid as the three-dimensional points to be detected included in the BEV grid; if the number of three-dimensional points in the point cloud to be detected included in the BEV grid is equal to the specified number, use the three-dimensional points included in the BEV grid as the three-dimensional points to be detected included in the BEV grid; if the number of three-dimensional points in the point cloud to be detected included in the BEV grid is less than the specified number, calculate the difference between the specified number and the number of three-dimensional points in the point cloud to be detected included in the BEV grid, and generate the difference number of three-dimensional points with preset coordinates; combine the generated three-dimensional points and the three-dimensional points in the point cloud to be detected included in the BEV grid to obtain the three-dimensional points to be detected included in the BEV grid.

[0028] Optionally, the information to be utilized of a three-dimensional point further includes at least one of the following: the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, and the feature data at the real position corresponding to the detected three-dimensional point.

[0029] In the fourth aspect of the embodiments of the present application, a three-dimensional object detection model training device is provided, and the device includes:

[0030] A second acquisition module, configured to acquire the point cloud and image of the sample space region as the sample point cloud and sample image respectively, and acquire the sample label including the position of the object within the sample space region; a second feature extraction module, configured to divide the sample point cloud according to the bird's-eye view (BEV) grid and determine the sample 3D points included in each BEV grid obtained by the division; use the point cloud feature extraction network in the 3D object detection model with the initial structure to extract features from the information to be utilized of the sample 3D points included in each BEV grid, so as to obtain the sample initial point cloud feature of each BEV grid; and input the sample image into the image feature extraction network in the 3D object detection model to obtain a plurality of sample initial image features; wherein, the information to be utilized of a 3D point includes: the coordinates of the 3D point and the coordinates of the centroid of the BEV grid to which the 3D point belongs; the 3D object detection model further includes a deformable attention network, a weight attention network and a decoder; a second point cloud feature fusion module, configured to, for each reference 3D point, input the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset; for each position offset, use the position offset to offset the current coordinates of the reference 3D point and determine the BEV grid to which the offset coordinates belong as the BEV grid corresponding to the position offset; fuse the sample initial point cloud features of the BEV grids corresponding to each position offset according to the weights corresponding to each position offset to obtain the current sample fused point cloud feature of the reference 3D point; a second image feature fusion module, configured to input the current reference feature of the reference 3D point into the weight attention network to obtain the weight of each sample initial image feature; fuse the feature values corresponding to the reference 3D point in each sample initial image feature according to the weights of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point; a second splicing module, configured to splice the current sample fused point cloud feature and sample intermediate image feature of each reference 3D point to obtain the current sample spliced feature; a second update module, configured to input the current sample spliced feature into the decoder to obtain the current decoding result, and update the current coordinates and reference feature of the reference 3D point based on the current decoding result; return to execute the step of, for each reference 3D point, inputting the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset until the number of decoding times reaches a preset number; a detection result acquisition module, configured to obtain a detection result based on the current coordinates of each reference 3D point; wherein, the detection result includes the position of the object in the sample space region;A training module, configured to adjust the model parameters of the three-dimensional object detection model with the initial structure based on the difference between the sample label and the detection result until a preset convergence condition is met, thereby obtaining a trained three-dimensional object detection model.

[0031] Optionally, the three-dimensional object detection model further includes a multi-layer perceptron; the apparatus further includes: a second correction amount acquisition module, configured to, before the second point cloud feature fusion module performs, for each reference three-dimensional point, inputting the current reference feature of the reference three-dimensional point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset, input the sample initial point cloud features of each BEV grid and the sample initial image features into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix; wherein, the initial transformation matrix is: calibrated according to the acquisition methods of the sample point cloud and the sample image; using the obtained correction amount to correct the initial transformation matrix; the second image feature fusion module is specifically configured to: use the corrected transformation matrix to transform the current coordinates of the reference three-dimensional point to obtain the pixel coordinates corresponding to the reference three-dimensional point in the sample image as the sample reference coordinates; fuse the eigenvalue corresponding to the sample reference coordinates in each sample initial image feature according to the weight of each sample initial image feature to obtain the current sample intermediate image feature of the reference three-dimensional point.

[0032] Optionally, the second correction amount acquisition module is specifically configured to input the sample initial point cloud features of each BEV grid, the sample initial image features, and the current coordinates of each reference three-dimensional point into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix.

[0033] Optionally, the apparatus further includes: a second interpolation module, configured to, before the second image feature fusion module performs fusing the eigenvalue corresponding to the sample reference coordinates in each sample initial image feature according to the weight of each sample initial image feature to obtain the current sample intermediate image feature of the reference three-dimensional point, if the sample reference coordinates are not integers, obtain other pixel coordinates within the neighborhood range of the sample reference coordinates; for each sample initial image feature, interpolate the eigenvalue corresponding to the obtained other pixel coordinates in the sample initial image feature as the eigenvalue corresponding to the sample reference coordinates in the sample initial image feature.

[0034] Optionally, the image feature extraction network includes a plurality of downsampling layers connected in series; there is one sample image, and the plurality of sample initial image features include: the image features of the sample image extracted by the plurality of downsampling layers; or, the sample image includes images of the sample space region collected at multiple different shooting angles, and the plurality of sample initial image features include the image features of each sample image.

[0035] Optionally, the second feature extraction module is specifically configured to, for each BEV grid obtained by division, if the number of three-dimensional points in the sample point cloud included in the BEV grid is greater than the specified number, sample the specified number of three-dimensional points from the three-dimensional points included in the BEV grid as the sample three-dimensional points included in the BEV grid; if the number of three-dimensional points in the sample point cloud included in the BEV grid is equal to the specified number, use the three-dimensional points included in the BEV grid as the sample three-dimensional points included in the BEV grid; if the number of three-dimensional points in the sample point cloud included in the BEV grid is less than the specified number, calculate the difference between the specified number and the number of three-dimensional points in the sample point cloud included in the BEV grid, and generate the difference number of three-dimensional points with preset coordinates; combine the generated three-dimensional points and the three-dimensional points in the sample point cloud included in the BEV grid to obtain the sample three-dimensional points included in the BEV grid.

[0036] Optionally, the information to be utilized of a three-dimensional point further includes at least one of the following: the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, and the feature data at the real position corresponding to the detected three-dimensional point.

[0037] Optionally, the sample label further includes the category of the object in the sample space region; the position indication of the object in the sample space region: the size of the minimum bounding cube of the object, the coordinates of the center point of the minimum bounding cube of the object, and the orientation of the object; the detection result further includes the category of the object in the sample space region; the differences between the sample label and the detection result include: the position difference of the object in the sample space region, and the category difference of the object in the sample space region.

[0038] In a fifth aspect of the embodiments of the present application, an electronic device is provided, including: a memory for storing a computer program; a processor for implementing any one of the object detection methods described in the first aspect above, or implementing any one of the three-dimensional object detection model training methods described in the second aspect above when executing the program stored in the memory.

[0039] In the sixth aspect of the embodiments of the present application, a computer-readable storage medium is provided. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, it implements the object detection method described in any one of the first aspects above, or implements the three-dimensional object detection model training method described in any one of the second aspects above.

[0040] The embodiments of the present application also provide a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the object detection method described in any one of the first aspects above, or execute the three-dimensional object detection model training method described in any one of the second aspects above.

[0041] For an object detection method provided by the embodiments of the present application, an electronic device acquires a point cloud and an image of a space area to be detected, as a point cloud to be detected and an image to be detected respectively; divides the point cloud to be detected according to a BEV grid, extracts the initial point cloud features to be detected of each obtained BEV grid, and extracts features of the image to be detected to obtain the initial image features to be detected; for each reference three-dimensional point, samples the BEV grids within the neighborhood range of the reference three-dimensional point based on the current coordinates and reference features of the reference three-dimensional point, and fuses the initial point cloud features to be detected of the sampled BEV grids to obtain the current fused point cloud features to be detected of the reference three-dimensional point; acquires the corresponding part of the reference three-dimensional point in the initial image features to be detected to obtain the current intermediate image features to be detected of the reference three-dimensional point; splices the current fused point cloud features to be detected and the current intermediate image features to be detected of each reference three-dimensional point to obtain the current spliced features to be detected; decodes the current spliced features to be detected to obtain the current decoding result, and updates the current coordinates and reference features of the reference three-dimensional point based on the current decoding result; returns to execute the step of, for each reference three-dimensional point, sampling the BEV grids within the neighborhood range of the reference three-dimensional point based on the current coordinates and reference features of the reference three-dimensional point, and fusing the initial point cloud features to be detected of the sampled BEV grids to obtain the current fused point cloud features to be detected of the reference three-dimensional point, until the number of decoding times reaches a preset number; based on the current coordinates of each reference three-dimensional point, obtains the positions of the objects in the space area to be detected.

[0042] Based on the above processing, the electronic device can obtain the current feature of the to-be-detected fused point cloud based on the to-be-detected point cloud, and obtain the current feature of the to-be-detected intermediate image based on the to-be-detected image. Based on the current feature of the to-be-detected fused point cloud and the current feature of the to-be-detected intermediate image, update the current coordinates and reference features of each reference 3D point. After a preset number of updates, determine the position of the object in the to-be-detected space region according to the current coordinates of each reference 3D point. That is, perform object detection on the to-be-detected space region based on the point cloud and the image jointly. In this way, since the process of obtaining the point cloud is not easily affected by external interferences such as changes in illumination, background color, and occlusions, combining the point cloud and the image for object detection can make up for the defect that the accuracy of the object detection result obtained only based on the image is not high, reduce the influence of external interferences on the accuracy of the object detection result, and improve the accuracy of the obtained object detection result.

[0043] Of course, it is not necessary for any product or method implementing the present application to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.

[0045] Figure 1 It is the first flowchart of the object detection method provided by the embodiment of the present application;

[0046] Figure 2 It is the second flowchart of the object detection method provided by the embodiment of the present application;

[0047] Figure 3 It is a schematic diagram of the object detection method provided by the embodiment of the present application;

[0048] Figure 4 It is a schematic diagram of the learnable radar-vision matrix mapping sampling module 302 provided by the embodiment of the present application;

[0049] Figure 5 It is a flowchart of the 3D object detection model training method provided by the embodiment of the present application;

[0050] Figure 6 It is a structural diagram of the object detection device provided by the embodiment of the present application;

[0051] Figure 7 It is a structural diagram of the 3D object detection model training device provided by the embodiment of the present application;

[0052] Figure 8 A structural diagram of the electronic device provided by the embodiment of the present application. Specific implementation manners

[0053] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.

[0054] In the field of machine vision technology, a vision sensor can be used to collect an image of a space area to be detected. By performing object detection on the collected image, a detection result of the position of an object in the space area to be detected can be obtained. However, the process of collecting an image by a vision sensor is susceptible to external interferences such as changes in illumination, changes in background color, and occlusions, which may result in low accuracy of the object detection result.

[0055] To solve the above problems, the present application provides an object detection method applied to an electronic device. The electronic device can be a terminal or a server. The electronic device can obtain a point cloud to be detected and an image to be detected in a space area to be detected; divide the point cloud to be detected according to a BEV (Bird’s Eye View) grid, and extract the initial point cloud feature to be detected of each obtained BEV grid, and perform feature extraction on the image to be detected to obtain the initial image feature to be detected; for each reference 3D point, sample the BEV grids within the neighborhood range of the reference 3D point based on the current coordinates and reference features of the reference 3D point, and fuse the initial point cloud features to be detected of the sampled BEV grids to obtain the current fused point cloud feature to be detected of the reference 3D point; obtain the corresponding part of the reference 3D point in the initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference 3D point; splice the current fused point cloud features to be detected and the current intermediate image features to be detected of each reference 3D point to obtain the current spliced feature to be detected; decode the current spliced feature to be detected to obtain the current decoding result, and update the current coordinates and reference features of the reference 3D point based on the current decoding result; return to execute the step of, for each reference 3D point, sampling the BEV grids within the neighborhood range of the reference 3D point based on the current coordinates and reference features of the reference 3D point, and fusing the initial point cloud features to be detected of the sampled BEV grids to obtain the current fused point cloud feature to be detected of the reference 3D point until the number of decoding times reaches a preset number; based on the current coordinates of each reference 3D point, obtain the position of an object in the space area to be detected. Combining the point cloud and the image to jointly perform object detection to improve the accuracy of the obtained object detection result.

[0056] See Figure 1 , Figure 1 which is the first flowchart of the object detection method provided by the embodiment of the present application. The method may include the following steps:

[0057] S101: Obtain the point cloud and image of the space area to be detected, and use them as the point cloud to be detected and the image to be detected respectively.

[0058] S102: Divide the point cloud to be detected according to the bird's-eye view grid, extract the initial point cloud features to be detected of each obtained bird's-eye view grid, and perform feature extraction on the image to be detected to obtain the initial image features to be detected.

[0059] S103: For each reference 3D point, sample the bird's-eye view grids within the neighborhood range of the reference 3D point based on the current coordinates and reference features of the reference 3D point, and fuse the initial point cloud features to be detected of the sampled bird's-eye view grids to obtain the current fused point cloud features to be detected of the reference 3D point.

[0060] S104: Obtain the corresponding part of the reference 3D point in the initial image features to be detected to obtain the current intermediate image features to be detected of the reference 3D point.

[0061] S105: Concatenate the current fused point cloud features to be detected and the intermediate image features to be detected of each reference 3D point to obtain the current concatenated features to be detected.

[0062] S106: Decode the current concatenated features to be detected to obtain the current decoding result, and update the current coordinates and reference features of the reference 3D point based on the current decoding result; return to execute the step of sampling the bird's-eye view grids within the neighborhood range of the reference 3D point based on the current coordinates and reference features of the reference 3D point, and fusing the initial point cloud features to be detected of the sampled bird's-eye view grids to obtain the current fused point cloud features to be detected of the reference 3D point for each reference 3D point until the number of decoding times reaches the preset number of times.

[0063] S107: Based on the current coordinates of each reference 3D point, obtain the positions of the objects in the space area to be detected.

[0064] Based on the object detection method provided in the embodiments of the present application, the electronic device can obtain the current to-be-detected fused point cloud feature based on the to-be-detected point cloud, and obtain the current to-be-detected intermediate image feature based on the to-be-detected image. Based on the current to-be-detected fused point cloud feature and the current to-be-detected intermediate image feature, update the current coordinates and reference features of each reference 3D point. After a preset number of updates, determine the positions of the objects in the to-be-detected space region according to the current coordinates of each reference 3D point. That is, perform object detection on the to-be-detected space region based on the point cloud and the image together. In this way, since the process of obtaining the point cloud is not easily affected by external interferences such as changes in illumination, background color, and occlusions, combining the point cloud and the image for object detection together can make up for the defect of low accuracy of the object detection result obtained only based on the image, reduce the influence of external interferences on the accuracy of the object detection result, and improve the accuracy of the obtained object detection result.

[0065] Regarding step S101, the to-be-detected space region and the categories of objects that may exist therein are related to the business requirements of the application scenario. For example, in the digital intersection scenario of modern traffic management that integrates multiple sensors and intelligent technologies to achieve comprehensive perception, analysis, and control of intersection traffic information, the to-be-detected space region may include the roads around the traffic intersection, and the objects that may exist may include vehicles, roadblocks, etc. In the intelligent loading scenario, the to-be-detected space region may be the cargo placement area inside the carriage, and the objects that may exist may include baffles, bolts, air conditioner indoor units, etc.

[0066] By scanning the to-be-detected space region with a vision sensor, an image of the to-be-detected space region can be collected; by scanning the to-be-detected space region with a radar sensor, a point cloud of the to-be-detected space region can be collected. The electronic device can obtain the image of the to-be-detected space region from the vision sensor as the to-be-detected image, and obtain the point cloud of the to-be-detected space region from the radar sensor as the to-be-detected point cloud. For example, the radar sensor can be a Millimeter-wave RADAR (millimeter-wave radar). The millimeter-wave radar can detect based on the electromagnetic wave in the millimeter-wave band. The frequency domain of the electromagnetic wave in the millimeter-wave band is usually 30 GHz (gigahertz) to 300 GHz; the wavelength usually ranges from 1 millimeter to 10 millimeters. Alternatively, the radar sensor can also be a lidar. The present application does not limit this.

[0067] It can be understood that there can be multiple vision sensors. For example, in the aforementioned digital intersection scenario, vision sensors can be set at different positions around the road. By adjusting the shooting angles of these vision sensors, the acquisition areas of these vision sensors can all include the to-be-detected space region. In this way, the images of the to-be-detected space region can be collected at different shooting angles by the vision sensors set at different positions. Correspondingly, the to-be-detected image is multiple at this time.

[0068] By setting visual sensors at different positions, the images to be detected in the space area to be detected are collected at different shooting angles respectively. Then, the information of the space area to be detected included in the images to be detected is more comprehensive, which can reduce the probability of problems such as missed detection and false detection caused by occlusion between objects, and improve the accuracy of the subsequent target detection results.

[0069] Regarding step S102, BEV refers to: a three-dimensional view drawn by looking down at the ground undulation from a high point. The electronic device can divide the space area to be detected into multiple BEV grids. One BEV grid corresponds to a cubic space area, and one BEV grid can be called a Pillar. The number of grids of the divided BEV grids can be jointly determined by considering the service requirements and device performance. The more the number of grids, the greater the computational complexity during target detection, and the higher the accuracy of the target detection results; the fewer the number of grids, the less the computational complexity during target detection, and the shorter the time taken to obtain the target detection results. Therefore, when the service requirements require a shorter time to obtain the target detection results, a smaller number of grids can be set; when the device performance is good and a higher accuracy of the target detection results is required, a larger number of grids can be set. For example, the number of divided BEV grids can be denoted as H×W. That is, on the BEV plane, the space area to be detected is divided into H rows of BEV grids, and each row of BEV grids includes W columns of BEV grids.

[0070] The point cloud to be detected includes multiple three-dimensional points. For each three-dimensional point, the electronic device can determine the Pillar to which the three-dimensional point belongs, that is, determine the BEV grid to which the three-dimensional point belongs. In this way, the electronic device also divides the point cloud to be detected according to the BEV grids, and determines the three-dimensional points included in each BEV grid.

[0071] The coordinates of the three-dimensional points can be coordinates in the world coordinate system. For example, the XOY plane of the world coordinate system can be the horizontal ground, the X-axis can be a regional boundary of the space area to be detected on the horizontal ground; the positive direction of the Z-axis can be the direction perpendicular to the horizontal ground and upward. According to the right-hand rule of the three-dimensional coordinate system, the world coordinate system can be constructed. Obviously, the projections of the three-dimensional points belonging to the same BEV grid on the XOY plane belong to the same area. Therefore, for each three-dimensional point included in the point cloud to be detected, the electronic device can determine the area to which the projection of the three-dimensional point belongs only according to the coordinate value of the three-dimensional point on the X-axis and the coordinate value on the Y-axis, and the BEV grid corresponding to the area is the BEV grid to which the three-dimensional point belongs. It can reduce the computational complexity when dividing the point cloud to be detected, reduce the division time, and further reduce the target detection time, improving the target detection efficiency.

[0072] In some embodiments, in the above world coordinate system, if the radar sensor can collect the coordinate value of a three-dimensional point on the Z-axis, the coordinates of each three-dimensional point included in the point cloud to be detected may include: the coordinate value on the X-axis, the coordinate value on the Y-axis, and the coordinate value on the Z-axis.

[0073] In some embodiments, if the radar sensor fails to collect the coordinate value of a three-dimensional point on the Z-axis, the coordinates of each three-dimensional point included in the point cloud to be detected may include: the coordinate value on the X-axis, and the coordinate value on the Y-axis.

[0074] Alternatively, a preset virtual coordinate value may also be used as the coordinate value of the three-dimensional point on the Z-axis, such as set to 0 or 1. At this time, the coordinates of each three-dimensional point include: the coordinate value on the X-axis, the coordinate value on the Y-axis, and the preset virtual coordinate value.

[0075] For each BEV grid, the electronic device may perform feature extraction on the three-dimensional points included in the BEV grid to obtain the initial point cloud feature to be detected of the BEV grid. Exemplarily, the electronic device may process the coordinates of the three-dimensional points included in the BEV grid based on PointNN (a deep learning algorithm for point cloud processing) to obtain the initial point cloud feature to be detected of the BEV grid. The initial point cloud feature to be detected of the BEV grid may also be referred to as the Pillar feature.

[0076] Moreover, the electronic device may also perform feature extraction on the image to be detected to obtain the initial image feature to be detected. Exemplarily, the electronic device may extract the initial image feature to be detected based on a convolutional neural network. For example, it may perform feature extraction based on ResNet (residual connection) and based on VGG (Visual Geometry Group Network), but not limited thereto.

[0077] Regarding step S103, when performing target detection, the electronic device may pre-determine a plurality of reference three-dimensional points. Subsequently, by adjusting the coordinates of the reference three-dimensional points, the coordinates of the reference three-dimensional points gradually approach the position of the object in the space area to be detected. Therefore, in order to reduce the probability of missed detection, the number of reference three-dimensional points is usually much larger than the number of objects that may exist in the space area to be detected. And it can be understood that the more the number of reference three-dimensional points, the greater the computational amount when the electronic device performs target detection, and the longer the time consumed for target detection. Therefore, according to the business requirements of the application scenario, the number of reference three-dimensional points is usually set within a certain numerical range. For example, in the aforementioned digital intersection scenario, the number of reference three-dimensional points may be set to not less than 600 and not more than 1000.

[0078] For each reference 3D point, the electronic device can obtain the current coordinates and reference features of the reference 3D point. The reference features can be a 256-dimensional vector. For example, the electronic device can randomly generate 256 numbers within a preset data range as the reference features of the reference 3D point. For example, the preset data range can be from 0 to 100.

[0079] In one implementation, according to the size of the space region to be detected, the electronic device can directly generate coordinates belonging to the space region to be detected as the coordinates of the reference 3D point. For example, in the aforementioned world coordinate system, one vertex of the space region to be detected is used as the coordinate origin; the length of the space region to be detected in the X-axis direction is 100 meters, the length in the Y-axis direction is 200 meters, and the length in the Z-axis direction is 50 meters. Then, the electronic device can generate random numbers within the interval [0, 100] as the coordinates of the reference 3D point on the X-axis; generate random numbers within the interval [0, 200] as the coordinates of the reference 3D point on the Y-axis; generate random numbers within the interval [0, 50] as the coordinates of the reference 3D point on the Z-axis.

[0080] In one implementation, the electronic device can perform object detection on the space region to be detected based on a 3D object detection model. During the training process of the 3D object detection model, the coordinates and reference features of the reference 3D point can be adjusted. The electronic device can directly obtain the coordinates and reference features of the reference 3D point obtained through the training of the 3D object detection model. For example, when training the 3D object detection model, the electronic device can randomly determine three numbers within the interval [0, 1] as the coordinate values of the reference 3D point on the X-axis, Y-axis, and Z-axis respectively at present. The reference 3D point determined at this time can be called the initial 3D point. Then, based on the current coordinates of the reference 3D point, the 3D object detection model is trained, and the current coordinates of the reference 3D point are updated. After the 3D object detection model is trained, the current coordinates and reference features of the reference 3D point, that is, the current coordinates and reference features of the reference 3D point directly obtained by the electronic device. Subsequently, the electronic device can scale the current coordinates of the reference 3D point according to the size of the space region to be detected. For example, in the above example, multiply the coordinate value of the reference 3D point on the X-axis at present by 100; multiply the coordinate value of the reference 3D point on the Y-axis at present by 200; multiply the coordinate value of the reference 3D point on the Z-axis at present by 50. That is, map the reference 3D point into the space region to be detected. The coordinates of the initial 3D point do not depend on the characteristics of any information collected by sensors. Therefore, the 3D object detection model trained based on the initial 3D point can adapt to different application scenarios and has high versatility. The specific method for the electronic device to perform object detection on the space region to be detected based on the 3D object detection model is described in detail in the subsequent embodiments.

[0081] For each reference 3D point, according to the current coordinates of the reference 3D point, the electronic device can determine the BEV grid to which the reference 3D point currently belongs. According to the current reference feature of the reference 3D point, sample BEV grids from the BEV grid to which the reference 3D point currently belongs and other surrounding BEV grids, and then fuse the initial point cloud features to be detected of the sampled BEV grids to obtain the fused point cloud features to be detected of the reference 3D point currently.

[0082] Regarding step S104, since the point cloud to be detected and the image to be detected are both obtained for the space region to be detected, there is a mapping relationship between the 3D points in the space region to be detected and the pixel coordinates in the image to be detected. By calibrating according to the acquisition methods of the point cloud to be detected and the image to be detected, this mapping relationship can be obtained. For example, a technician can place a calibration board in the space region to be detected. The calibration point cloud of the calibration board is collected by the acquisition method of the point cloud to be detected, and the calibration image of the calibration board is collected by the acquisition method of the image to be detected. For each fiducial point on the calibration board, according to the coordinates of the 3D point corresponding to the fiducial point in the calibration point cloud and the pixel coordinates of the fiducial point in the calibration image, the mapping relationship between the 3D points in the space region to be detected and the pixel coordinates in the image to be detected can be solved.

[0083] When there are multiple images to be detected, for each image to be detected, by calibrating according to the acquisition methods of the point cloud to be detected and the image to be detected, the mapping relationship between the pixel coordinates in the image to be detected and the 3D points in the space region to be detected can be obtained. Moreover, since the initial image features to be detected are obtained by feature extraction of the image to be detected, there is also a mapping relationship between the feature values in the initial image features to be detected and the pixel coordinates in the image to be detected.

[0084] That is, a 3D point in the space region to be detected corresponds to a pixel coordinate. The part corresponding to this pixel coordinate in the initial image features to be detected is the part corresponding to this 3D point in the initial image features to be detected.

[0085] Exemplarily, a feature value in the initial image features to be detected is obtained by performing feature extraction on the pixel points in an image region in the image to be detected. Then, the part corresponding to a 3D point in the initial image to be detected is the feature value obtained by performing feature extraction based on the pixel point corresponding to the 3D point in the image to be detected.

[0086] In this way, for each reference 3D point, the electronic device can obtain the part corresponding to the reference 3D point in the initial image features to be detected, and obtain the intermediate image features to be detected of the reference 3D point currently.

[0087] For steps S105 and S106, for each reference 3D point, the electronic device can obtain the current to-be-detected fused point cloud feature and the current to-be-detected intermediate image feature of the reference 3D point. Then, the eigenvalues of the current to-be-detected fused point cloud features of each reference 3D point and the eigenvalues of the current to-be-detected intermediate image features of each reference 3D point are concatenated to obtain the current to-be-detected concatenated feature.

[0088] Specifically, if the current to-be-detected fused point cloud feature of reference 3D point i is denoted as Ri, the current to-be-detected intermediate image feature is denoted as Ci, and the total number of reference 3D points is denoted as I, then the electronic device can perform concat (concatenation) on the eigenvalues of the current to-be-detected fused point cloud features of the reference 3D points and the eigenvalues of the current to-be-detected intermediate image features of each reference 3D point, and the obtained feature can be denoted as: {R1, R2, R3, …, RI, C1, C2, C3, ……, CI}.

[0089] Then, the feature obtained by concat is encoded to obtain the current to-be-detected concatenated feature. Exemplarily, the electronic device can use a convolutional neural network to jointly encode the feature obtained by concat to obtain the current to-be-detected concatenated feature.

[0090] By decoding the current to-be-detected concatenated feature, the current coordinate adjustment amount of each reference 3D point and the current reference feature adjustment amount of each reference 3D point can be obtained. For each reference 3D point, the sum of the current coordinate and the coordinate adjustment amount of the reference 3D point is calculated to obtain the current coordinate of the reference 3D point; the sum of the current reference feature and the reference feature adjustment amount of the reference 3D point is calculated to obtain the current reference feature of the reference 3D point. For example, the electronic device can decode the current to-be-detected concatenated feature based on the Decoder in the deep learning model. For example, the decoder can be the decoding module in the Transformer (transformation network) for performing decoding processing. In this way, the electronic device also performs one decoding and updates the coordinates and reference features of each reference 3D point once.

[0091] Then, the electronic device can use the current coordinates and reference features of each reference 3D point updated by this decoding, and return the steps of performing, for each reference 3D point, sampling the BEV grid within the neighborhood range of the reference 3D point based on the current coordinates and reference features of the reference 3D point, and fusing the initial point cloud features to be detected of the sampled BEV grid to obtain the current point cloud features to be detected and fused of the reference 3D point. That is, based on the current coordinates and reference features of each reference 3D point, the current stitching features to be detected are determined again, so as to perform the next decoding based on the current stitching features to be detected, and update the coordinates and reference features of each reference 3D point again until the number of decoding times reaches the preset number. The preset number can be determined based on the type of decoder used. For example, when using the decoding module in Transformer as the decoder, since Transformer usually includes 6 decoding modules, the preset number can be 6.

[0092] Regarding step S107, after the number of decoding times reaches the preset number, for each reference 3D point, the current coordinates of the reference 3D point are the positions of the objects existing in the to-be-detected space area to be detected (hereinafter referred to as the to-be-screened positions). It can be understood that different reference 3D points may indicate the same object, that is, an object may correspond to multiple to-be-screened positions. Therefore, the electronic device can post-process the to-be-screened positions to obtain the positions of the objects in the to-be-detected space area.

[0093] In some embodiments, the decoding result may further include the confidence of the reference 3D point, and the confidence can indicate the probability that there is an object at the position indicated by the reference 3D point. When the confidence of a reference 3D point is greater than the confidence threshold, it means that the probability of an object existing at the reference 3D point is relatively high. Then, the coordinates of the reference 3D point can be used as the position of the object existing in the to-be-detected space area detected, as the target detection result of the to-be-detected space area. The confidence threshold can be determined according to the detection accuracy requirements in the actual scenario. The higher the confidence threshold, the higher the accuracy of the detected object; the lower the confidence threshold, the lower the probability of problems such as missed detection. When high accuracy of the detected object is required, the confidence threshold can be set to a relatively large value, such as 0.95; when the probability of problems such as missed detection is required to be reduced, the confidence threshold can be set to a relatively small value, such as 0.7.

[0094] In some embodiments, the position of the detected object may indicate: the coordinates of the center point of the minimum bounding cube of the object, the size of the minimum bounding cube of the object, and the orientation of the object. If the position of the detected object includes the above information, the subsequent electronic device may calculate the vertex coordinates of the minimum bounding cube of the object based on the above information, that is, determine the position of the object. Alternatively, the position of the detected object may be the vertex coordinates of the minimum bounding cube of the object, that is, the electronic device may directly obtain the position of the object.

[0095] In some embodiments, the decoding result may further include: the category of the object existing at the reference three-dimensional point. Then, according to the decoding result, the electronic device may also obtain the category of the detected object, meeting the service requirements in the actual scenario.

[0096] In some embodiments, the electronic device may perform object detection on the space region to be detected based on an end-to-end three-dimensional object detection model. On Figure 1 this basis, referring to Figure 2 , the above step S102 may include the following steps:

[0097] S1021: Divide the point cloud to be detected according to the bird's-eye view grid, and determine the three-dimensional points to be detected included in each bird's-eye view grid obtained by the division; use the point cloud feature extraction network in the pre-trained three-dimensional object detection model to extract features from the information to be used of the three-dimensional points to be detected included in each bird's-eye view grid, and obtain the initial point cloud features to be detected of each bird's-eye view grid; and input the image to be detected into the image feature extraction network in the three-dimensional object detection model to obtain a plurality of initial image features to be detected.

[0098] Among them, the information to be used of a three-dimensional point includes: the coordinates of the three-dimensional point and the coordinates of the centroid of the BEV grid to which the three-dimensional point belongs; the three-dimensional object detection model is trained based on the sample images, sample point clouds in the sample space region, and sample labels indicating the positions of the objects in the sample space region; the three-dimensional object detection model further includes a deformable attention network, a weight attention network, and a decoder.

[0099] Correspondingly, the above step S103 may include the following steps: S1031: For each reference 3D point, input the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset; for each position offset, use the position offset to offset the current coordinates of the reference 3D point, and determine the bird's-eye view grid to which the offset coordinates belong as the bird's-eye view grid corresponding to the position offset; fuse the initial point cloud features to be detected of the bird's-eye view grids corresponding to each position offset according to the weights corresponding to each position offset to obtain the current initial point cloud features to be detected and fused point cloud features of the reference 3D point.

[0100] Correspondingly, the above step S104 may include the following steps:

[0101] S1041: Input the current reference feature of the reference 3D point into the weight attention network to obtain the weight of each initial image feature to be detected; fuse the feature values corresponding to the reference 3D point in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the current intermediate image features to be detected of the reference 3D point.

[0102] Correspondingly, the above step S106 may include the following steps:

[0103] S1061: Input the current initial image features to be detected and fused point cloud features into the decoder to obtain the current decoding result, and update the current coordinates and reference features of the reference 3D point based on the current decoding result; return to execute the step of sampling the bird's-eye view grids within the neighborhood range of the reference 3D point based on the current coordinates and reference features of the reference 3D point and fusing the initial point cloud features to be detected of the sampled bird's-eye view grids to obtain the current initial point cloud features to be detected and fused point cloud features of the reference 3D point until the number of decoding times reaches the preset number of times.

[0104] Since the electronic device will subsequently use the point cloud feature extraction network in the pre-trained 3D object detection model to extract the features of the information to be utilized of the 3D points to be detected included in each BEV grid. It can be understood that the format of the input data of the point cloud feature extraction network is set. For example, the input data of the point cloud feature extraction network is a tensor of (D, N, P), where D represents the number of information dimensions of the information to be utilized of a 3D point. When the information to be utilized of a 3D point includes the coordinates of the 3D point and the coordinates of the centroid of the BEV grid to which the 3D point belongs, D can be 6. N represents the number of 3D points to be detected in each BEV grid (i.e., the specified number in the subsequent embodiments). P represents the number of BEV grids obtained by division, that is, the product of the aforementioned H and W.

[0105] For each BEV grid, after the electronic device determines the three-dimensional points included in the BEV grid, it can determine the minimum circumscribed cube of the three-dimensional points included in the BEV grid, and then use the centroid of the minimum circumscribed cube as the centroid of the BEV grid.

[0106] Since the number of three-dimensional points included in each BEV grid may be different, in order to obtain the initial point cloud features to be detected for each BEV grid through the point cloud feature extraction network, the electronic device can determine the three-dimensional points to be detected included in the BEV grid according to the number of three-dimensional points included in the BEV grid.

[0107] In some embodiments, the foregoing determination of the three-dimensional points to be detected included in each BEV grid obtained by partitioning may include the following steps: for each BEV grid obtained by partitioning, if the number of three-dimensional points in the point cloud to be detected included in the BEV grid is greater than the specified number, sample the specified number of three-dimensional points from the three-dimensional points included in the BEV grid as the three-dimensional points to be detected included in the BEV grid; if the number of three-dimensional points in the point cloud to be detected included in the BEV grid is equal to the specified number, use the three-dimensional points included in the BEV grid as the three-dimensional points to be detected included in the BEV grid; if the number of three-dimensional points in the point cloud to be detected included in the BEV grid is less than the specified number, calculate the difference between the specified number and the number of three-dimensional points in the point cloud to be detected included in the BEV grid, and generate the difference number of three-dimensional points with preset coordinates; combine the generated three-dimensional points and the three-dimensional points in the point cloud to be detected included in the BEV grid to obtain the three-dimensional points to be detected included in the BEV grid.

[0108] For each BEV grid obtained by partitioning, when the number of three-dimensional points in the point cloud to be detected included in the BEV grid is greater than the specified number, that is, not all the information to be utilized of the three-dimensional points included in the BEV grid needs to be input into the point cloud feature extraction network. At this time, the electronic device can sample the three-dimensional points included in the BEV grid and use the specified number of sampled three-dimensional points as the three-dimensional points to be detected included in the BEV grid. Correspondingly, in the input data of the point cloud feature extraction network, the data corresponding to the BEV grid is the information to be utilized of the sampled three-dimensional points to be detected.

[0109] When the number of three-dimensional points in the point cloud to be detected included in the BEV grid is equal to the specified number, that is, all the information to be utilized of the three-dimensional points included in the BEV grid is input into the point cloud feature extraction network. At this time, the electronic device can directly determine all the three-dimensional points included in the BEV grid as the three-dimensional points to be detected. Correspondingly, in the input data of the point cloud feature extraction network, the data corresponding to the BEV grid is the information to be utilized of all the three-dimensional points in the BEV grid.

[0110] When the number of three-dimensional points in the point cloud to be detected included in the BEV grid is less than the specified number, even if all the three-dimensional points included in the BEV grid are determined as the three-dimensional points to be detected, the data corresponding to the BEV grid in the above-set data format cannot be obtained. At this time, the electronic device can perform supplementary processing on the data corresponding to the BEV grid. Specifically, the electronic device can calculate the difference between the specified number and the number of three-dimensional points in the point cloud to be detected included in the BEV grid (hereinafter referred to as the number of groups to be supplemented), and the number of groups to be supplemented is the number of data groups that need to be supplemented to the data corresponding to the BEV grid. Then, generate three-dimensional points with preset values for the number of groups to be supplemented. When each value in the information to be utilized of a three-dimensional point is a preset value, it can be indicated that the three-dimensional point is used to fill the input data so that the input data has the set format. For example, the preset value can be 0 or 1.

[0111] For example, if the information to be utilized of a three-dimensional point is represented as: (x, y, z, xc, yc, zc), where (x, y, z) represents the coordinates of the three-dimensional point; (xc, yc, zc) represents the coordinates of the centroid of the BEV grid to which the three-dimensional point belongs. When the preset value is 0, the information to be utilized of the three-dimensional point generated by the electronic device based on the preset value can be (0, 0, 0, 0, 0, 0).

[0112] Furthermore, for the BEV grid, the electronic device can use the generated three-dimensional points and the three-dimensional points in the point cloud to be detected included in the BEV grid as the three-dimensional points to be detected included in the BEV grid. Correspondingly, in the input data of the point cloud feature extraction network, the data corresponding to the BEV grid is: the information to be utilized of all three-dimensional points in the BEV grid and the information to be utilized of the three-dimensional points generated based on the preset value.

[0113] Based on the above processing, when the number of three-dimensional points included in a BEV grid is not the specified number, the electronic device can determine the three-dimensional points to be detected included in the BEV grid based on the number of three-dimensional points included in the BEV grid in a sampling or supplementary processing manner. Subsequently, the point cloud feature extraction network can be used to extract the features of the information to be utilized of the three-dimensional points to be detected included in each BEV grid, reducing the probability of problems such as low accuracy of the initial point cloud features to be detected of each BEV grid extracted due to the inconsistent number of three-dimensional points included in each BEV grid, and improving the accuracy of the target detection results obtained based on the initial point cloud features to be detected of each BEV grid.

[0114] In some embodiments, the information to be utilized of a three-dimensional point further includes at least one of the following: the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, and the feature data at the real position corresponding to the detected three-dimensional point.

[0115] When the information to be utilized for a three-dimensional point further includes the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, that is, the input data of the point cloud feature extraction network carries data indicating the relationship between some of the input data, the difficulty for the point cloud feature extraction network to learn the relationship between the input data will be lower, and the relationship between the learned input data will be more accurate. In this case, the initial point cloud features to be detected for each BEV grid obtained will be more accurate.

[0116] When the information to be utilized for a three-dimensional point further includes the feature data at the true position corresponding to the detected three-dimensional point, the point cloud feature extraction network performs feature extraction based on the feature data at the actual detected true positions. Then, the point cloud features to be detected for each BEV grid extracted will be more in line with the actual scenario, that is, more accurate. Subsequently, the accuracy of the object detection results obtained based on the more accurate point cloud features to be detected for each BEV grid will be higher. For example, the feature data can be the signal-to-noise ratio, reflection intensity, etc. When the radar sensor can detect multiple types of feature data at the true position corresponding to the three-dimensional point, the information to be utilized for the three-dimensional point can also include the multiple types of feature data to further improve the accuracy of the initial point cloud features to be detected for each BEV grid extracted.

[0117] Exemplarily, when the information to be utilized for a three-dimensional point further includes the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, and the feature data at the true position corresponding to the detected three-dimensional point, the data to be utilized for a three-dimensional point can be expressed as: (x, y, z, r, xc, yc, zc, xp, yp). At this time, the number of information dimensions of the information to be utilized for the three-dimensional point is 9. That is, when the input data of the point cloud feature extraction network is a tensor of (D, N, P), D is 9. Among them, (x, y, z) represents the coordinates of the three-dimensional point; r represents the feature data at the true position corresponding to the three-dimensional point; (xc, yc, zc) represents the coordinates of the centroid of the BEV grid to which the three-dimensional point belongs; (xp, yp) represents the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs.

[0118] In some embodiments, when the radar sensor can collect the coordinate value of the three-dimensional point on the Z-axis, the coordinate value of the three-dimensional point included in the information to be utilized for the three-dimensional point can be: the true coordinate value collected by the radar sensor. When the radar sensor fails to collect the coordinate value of the three-dimensional point on the Z-axis, a preset virtual coordinate value can be used as the coordinate value of each three-dimensional point on the Z-axis. For example, the preset virtual coordinate value can be set to 0 or 1.

[0119] After the electronic device obtains the image to be detected, it can use an image feature extraction network to extract features from the image to be detected, obtaining multiple initial image features to be detected. For example, the image feature extraction network may include ResNet, VGG, etc. in the foregoing embodiments, but is not limited thereto.

[0120] In some embodiments, the image feature extraction network includes a plurality of downsampling layers connected in series; there is one image to be detected, and the multiple initial image features to be detected include: the image features of the image to be detected extracted by the plurality of downsampling layers; or, the image to be detected includes images of the space area to be detected collected at multiple different shooting angles, and the multiple initial image features to be detected include the image features of each image to be detected.

[0121] After inputting the image to be detected into the first downsampling layer in the series connection, the plurality of downsampling layers can sequentially perform downsampling processing on their own input data. Then, the output data of different downsampling layers, that is, the image features of the image to be detected extracted by different downsampling layers, are different initial image features to be detected. Obviously, the number of feature values included in different initial image features to be detected is different. Taking the downsampling layer as a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1 as an example, for an image 1 with 5×5 pixel points, by performing one convolution on image 1 through the first downsampling layer, an image feature a with a size of 3×3 can be obtained, that is, the output data of the first downsampling layer; by performing one convolution on image feature a through the second downsampling layer, an image feature b with a size of 1 can be obtained, that is, the output data of the second downsampling layer. The above image feature a and image feature b are the image features of image 1 extracted by the plurality of downsampling layers.

[0122] Therefore, in the case where there is one image to be detected, multiple initial image features to be detected can be extracted through the plurality of downsampling layers connected in series included in the image feature extraction network.

[0123] In the case where there are multiple images to be detected, that is, in the case where multiple vision sensors respectively collect images of the space area to be detected at different shooting angles, for each image to be detected, through the image feature extraction network, one image feature of the image to be detected can be extracted; or, in the manner of extracting multiple initial image features of one image to be detected as described above, multiple initial image features of each image to be detected can be extracted respectively.

[0124] Based on the above processing, multiple initial image features to be detected of the image to be detected are extracted through the image feature extraction network, that is, more abundant and information-rich initial image features to be detected can be obtained. Then, subsequent obtaining the object detection result based on the more information-rich initial image features to be detected means obtaining the object detection result by integrating more information, which can improve the accuracy of the obtained object detection result.

[0125] For each reference 3D point, the electronic device can input the current reference feature of the reference 3D point into the DAT (Deformable Attention Transformer), and obtain the number of position offsets to be obtained preset (i.e., the preset sampling number). The deformable attention network can process the current reference feature of the reference 3D point based on the deformable attention mechanism to obtain a preset sampling number of position offsets and the weight corresponding to each position offset. For each position offset, the electronic device can use the position offset to offset the current coordinate of the reference 3D point.

[0126] In the case where the electronic device directly generates the coordinates within the space region to be detected as the coordinates of the initial 3D point, the electronic device can directly calculate the sum of the current coordinate of the reference 3D point and the position offset, and the calculated coordinate is the offset coordinate. In the case where the coordinate of the reference 3D point is obtained by the electronic device based on a number randomly determined from the interval [0, 1], the electronic device can first scale the coordinate of the reference 3D point according to the size of the space region to be detected, and then calculate the sum of the scaled coordinate and the position offset, and the calculated coordinate is the offset coordinate.

[0127] The BEV grid to which the offset coordinate belongs, that is, the BEV grid corresponding to the position offset, is also the BEV grid sampled based on the position offset. According to the weights corresponding to the respective position offsets, calculate the weighted sum of the initial point cloud features to be detected of the BEV grids corresponding to the position offsets, and obtain the current initial fused point cloud feature to be detected of the reference 3D point.

[0128] Exemplarily, the electronic device can calculate the current initial fused point cloud feature to be detected of the reference 3D point based on the following formula (1):

[0129] ; (1)

[0130] where represents the current initial fused point cloud feature to be detected of the i-th reference 3D point; represents the initial point cloud feature to be detected of the BEV grid corresponding to the k-th position offset of the i-th reference 3D point; Represents the weight corresponding to the k-th position offset of the i-th reference 3D point currently; Represents the current coordinates of the i-th reference 3D point; Represents the coordinates obtained by scaling the current coordinates of the i-th reference 3D point according to the size of the space region to be detected; Represents the k-th position offset of the i-th reference 3D point currently; K represents the preset sampling number.

[0131] For each reference 3D point, the electronic device can input the current reference feature of the reference 3D point into the weight attention network. The weight attention network processes the current reference feature of the reference 3D point based on the Attention mechanism to obtain the weights of each initial image feature to be detected. After the electronic device determines the part of the reference 3D point corresponding to each initial image feature to be detected, that is, after determining the eigenvalue corresponding to the reference 3D point in each initial image feature to be detected, it can calculate the weighted sum of the determined eigenvalues according to the weights of each initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference 3D point.

[0132] In some embodiments, the 3D object detection model further includes an MLP (Multilayer Perceptron).

[0133] Before the above steps, for each reference 3D point, inputting the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset, the method may further include the following steps: Step 1: Input the initial point cloud features to be detected and the initial image features to be detected of each BEV grid into the multilayer perceptron to obtain a correction amount for correcting the initial transformation matrix. The initial transformation matrix is calibrated according to the acquisition methods of the point cloud to be detected and the image to be detected.

[0134] Step 2: Use the obtained correction amount to correct the initial transformation matrix.

[0135] Correspondingly, the above steps of fusing the eigenvalues corresponding to the reference 3D point in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference 3D point may include the following steps: Step 3: Use the corrected transformation matrix to transform the current coordinates of the reference 3D point to obtain the pixel coordinates corresponding to the reference 3D point in the image to be detected as the reference coordinates to be detected. Step 4: Fuse the eigenvalues corresponding to the reference coordinates to be detected in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference 3D point.

[0136] The mapping relationship between the three-dimensional points in the space region to be detected and the pixel coordinates in the image to be detected obtained by the above-mentioned calibration according to the acquisition methods of the point cloud to be detected and the image to be detected, that is, the initial transformation matrix. The initial transformation matrix can be denoted as . That is, the initial transformation matrix is a 3×3 matrix.

[0137] It can be understood that the initial point cloud features to be detected of a BEV grid include multiple eigenvalues, and the initial image features to be detected also include multiple eigenvalues. The electronic device can perform concat on the initial point cloud features to be detected of each BEV grid and the initial image features to be detected. Then, the features obtained by concat are input into the MLP, and the 3×3 matrix output by the MLP is used as the correction amount for correcting the initial transformation matrix.

[0138] In some embodiments, the above-mentioned step of inputting the initial point cloud features to be detected of each BEV grid and the initial image features to be detected into the multi-layer perceptron to obtain the correction amount for correcting the initial transformation matrix may include the following steps: inputting the initial point cloud features to be detected of each BEV grid, the initial image features to be detected, and the current coordinates of each reference three-dimensional point into the multi-layer perceptron to obtain the correction amount for correcting the initial transformation matrix.

[0139] It can be understood that at this time, decoding has not been performed based on the reference three-dimensional points, and the coordinates of the reference three-dimensional points have not been updated. When the electronic device generates the coordinates of the reference three-dimensional points, the current coordinates of each reference three-dimensional point are the coordinates generated by the electronic device; when the electronic device performs object detection on the space region to be detected based on the three-dimensional object detection model, the current coordinates of each reference three-dimensional point are the current coordinates of each reference three-dimensional point after the three-dimensional object detection model is trained.

[0140] The electronic device can perform concat on the initial point cloud features to be detected of each BEV grid, multiple initial image features to be detected, and the current coordinates of each reference three-dimensional point. Then, the features obtained by concat are input into the MLP, and the 3×3 matrix output by the MLP is used as the correction amount for correcting the initial transformation matrix. In this way, that is, the current coordinates of the reference three-dimensional points that comprehensively indicate the positions of the objects in the space region to be detected are jointly used to determine the correction amount, and the correction amount can indicate the positions where there are objects in the space region to be detected. After correcting the initial transformation matrix based on the correction amount, the position of the object indicated by the reference position to be detected obtained based on the corrected transformation matrix is more accurate, thereby improving the accuracy of the subsequent obtained object detection results.

[0141] Then, the electronic device can correct the initial transformation matrix using the obtained correction amount. That is, for each element in the initial transformation matrix, calculate the sum value of the element and the element at the same position in the obtained correction amount, and update the element. Updating the elements in the initial transformation matrix means correcting the initial transformation matrix using the obtained correction amount. Furthermore, for each reference three-dimensional point, the electronic device can use the corrected transformation matrix to transform the coordinates of the reference three-dimensional point to obtain the pixel coordinates in the pixel coordinate system, that is, obtain the corresponding reference coordinates to be detected in the image to be detected. The eigenvalue corresponding to the reference coordinate to be detected in each initial image feature to be detected is the eigenvalue corresponding to the reference three-dimensional point in each initial image feature to be detected.

[0142] As mentioned above, there is a mapping relationship between the eigenvalues in the initial image features to be detected and the pixel points in the image to be detected. After the electronic device obtains the corresponding reference coordinates to be detected in the image to be detected, it can determine the eigenvalues corresponding to the reference coordinates to be detected in the initial image features to be detected according to the above mapping relationship.

[0143] For example, if the number of pixel points included in the image to be detected is 1000×1000, the initial image features to be detected include: the initial image feature to be detected 1 with the number of eigenvalues being 200×200; the initial image feature to be detected 2 with the number of eigenvalues being 100×100; the initial image feature to be detected 3 with the number of eigenvalues being 20×20, etc. Correspondingly, at this time, if the reference coordinate to be detected is (100, 100), then the eigenvalues corresponding to the reference coordinate to be detected in the initial image features to be detected include: the eigenvalue at the coordinate (20, 20) in the initial image feature to be detected 1; the eigenvalue at the coordinate (10, 10) in the initial image feature to be detected 2; the eigenvalue at the coordinate (2, 2) in the initial image feature to be detected 3.

[0144] Furthermore, according to the weights of each initial image feature to be detected, fuse the eigenvalues corresponding to the reference coordinate to be detected in each initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference three-dimensional point.

[0145] Based on the above processing, the electronic device can correct the initially calibrated transformation matrix based on the initial point cloud features to be detected in each BEV grid and the initial image features to be detected. That is, the initially calibrated transformation matrix is corrected based on the actual application scenario, and the corrected transformation matrix is more suitable for the actual application scenario compared to the initially calibrated transformation matrix. The current intermediate image features to be detected obtained based on the corrected transformation matrix are more accurate compared to the current intermediate image features to be detected obtained based on the initially calibrated transformation matrix, enabling subsequent object detection based on the more accurate current intermediate image features to be detected, which can improve the accuracy of the obtained object detection results.

[0146] Exemplarily, the electronic device can obtain the corrected calibration matrix based on the following formula (2):

[0147] ; (2)

[0148] where, represents the corrected calibration matrix; , represents the initially calibrated transformation matrix; represents the correction amount obtained through a multi-layer perceptron; , represents the coordinates of each reference 3D point, and M represents the number of reference 3D points; , represents the initial image features to be detected, and C1 represents the number of feature channels of the initial image features to be detected; , represents the initial point cloud features to be detected in each BEV grid, and C2 represents the number of feature dimensions of the initial point cloud features to be detected in each BEV grid.

[0149] Correspondingly, the electronic device can calculate the current intermediate image features to be detected of the reference 3D point based on the following formula (3):

[0150] ; (3)

[0151] where, represents the current intermediate image features to be detected of the i-th reference 3D point; represents the current coordinates of the i-th reference 3D point; represents the reference coordinates to be detected corresponding to the current coordinates of the i-th reference 3D point in the n-th image to be detected; represents: the reference coordinates to be detected extracted through the j-th downsampling layer corresponding image features; represents the weight of the image features extracted through the j-th downsampling layer in the n-th image to be detected of the i-th reference 3D point, m represents the number of serially connected downsampling layers included in the image feature extraction network; N represents the number of images to be detected.

[0152] In some embodiments, based on the weighted attention network, the electronic device can obtain the weight of each image to be detected. Further, for each image to be detected, the electronic device can calculate the quotient of the weight of the image to be detected and the number of downsampling layers, and use the calculation result as the weight of each of the multiple image features of the image to be detected extracted through the multiple downsampling layers.

[0153] In some embodiments, since the initial image features to be detected and the initial point cloud features to be detected of each BEV grid need to be input into the MLP to obtain the correction amount for correcting the initial transformation matrix, that is, the initial point cloud features to be detected of each BEV grid also need to have a specified format, and this format is related to the format of the initial point cloud to be detected. For example, the electronic device can increase the dimension of the features output by the point cloud feature extraction network to obtain a tensor with dimensions (C, N, P); perform pooling on the tensor after dimension increase in the N dimension to obtain a tensor with dimensions (C, P). And as mentioned above, P is the product of H and W, then the tensor with dimensions (C, P) can be transformed into a tensor with dimensions (C, H, W) as the initial point cloud features to be detected of each BEV grid. Wherein, C is the number of output channels of the image feature extraction network.

[0154] For each reference 3D point, the detected reference coordinates obtained by performing coordinate transformation on the current coordinates of the reference 3D point may not be integers, that is, the detected reference coordinates may not correspond to pixel points in the image to be detected.

[0155] In some embodiments, when the detected reference coordinates are not integers, the electronic device can round the non-integer coordinate values of the detected reference coordinates, and use the feature value corresponding to the rounded coordinates in each of the initial image features to be detected as the feature value corresponding to the detected reference coordinates.

[0156] In some embodiments, before fusing the feature values corresponding to the detected reference coordinates in each of the initial image features to be detected according to the weights of the initial image features to be detected to obtain the intermediate image features to be detected of the reference 3D point at present, the method may further include the following steps: if the detected reference coordinates are not integers, obtain other pixel coordinates within the neighborhood range of the detected reference coordinates; for each initial image feature to be detected, interpolate the feature values corresponding to the obtained other pixel coordinates in the initial image feature to be detected as the feature value corresponding to the detected reference coordinates in the initial image feature to be detected.

[0157] The electronic device can obtain other pixel coordinates within the neighborhood range of the reference coordinate to be detected. For example, for the coordinate values in the reference coordinate to be detected that are not integers, they are respectively rounded up and rounded down. The pixel coordinates at the coordinate values determined by rounding are determined as other pixel coordinates within the neighborhood range of the reference coordinate to be detected. For example, when the reference coordinate to be detected is (1.5, 3.5), rounding up 1.5 gives 2; rounding down 1.5 gives 1; rounding up 3.5 gives 4; rounding down 3.5 gives 3. Then the other pixel coordinates within the neighborhood range of the reference coordinate to be detected include: (1, 3), (2, 3), (1, 4), (2, 4).

[0158] At this time, the other pixel coordinates correspond to feature values in each initial image feature to be detected. For each initial image feature to be detected, the electronic device can interpolate the feature values corresponding to the obtained other pixel coordinates in this initial image feature to be detected as the feature value corresponding to the reference coordinate to be detected in this initial image feature to be detected. For example, the electronic device can use bilinear interpolation to interpolate the obtained feature values.

[0159] Based on the above processing, when the reference coordinate to be detected is not an integer, the electronic device comprehensively determines the feature value corresponding to the reference coordinate to be detected by combining all other pixel coordinates within the neighborhood range of the reference coordinate to be detected. Then the feature value corresponding to the reference coordinate to be detected determined in this way is more accurate, and the current intermediate image feature to be detected of this reference 3D point calculated based on the more accurate feature value is more accurate, which can improve the accuracy of the target detection result obtained based on the current intermediate image feature to be detected.

[0160] After the electronic device stitches the current intermediate image feature to be detected and the current fused point cloud feature to be detected of each reference 3D point to obtain the current stitched feature to be detected, it can input the current stitched feature to be detected into the decoder to obtain the current decoding result. The decoder can be the Decoder based on the deep learning model in the foregoing embodiments, such as the decoding module in Transformer. The decoding result can include the coordinate offset of each reference 3D point, the reference feature offset of each reference 3D point, etc. Based on the offsets obtained by decoding, the electronic device can update the current coordinates and reference features of each reference 3D point.

[0161] In some embodiments, the decoding result can also include the size offset of the minimum circumscribed cube of the object existing at this reference 3D point, the category offset of the category of the object existing at this reference 3D point, etc. The electronic device can use the decoding result to update the size of the minimum circumscribed cube of the object existing at this reference 3D point, the category of the object existing at this reference 3D point, etc.

[0162] Based on the above processing, the electronic device can perform object detection based on the end-to-end trained 3D object detection model, reducing the maintenance cost and improving the versatility of the object detection method.

[0163] In some embodiments, referring to Figure 3 , Figure 3 is a schematic diagram of an object detection method provided by an embodiment of the present application. In the embodiment of the present application, according to the functions of the 3D object detection model, the 3D object detection model is divided into 4 modules. Among them, the multi-sensor feature extraction module 301 includes an image backbone network (i.e., the image feature extraction network in the foregoing embodiment) and a radar backbone network (i.e., the point cloud feature extraction network in the foregoing embodiment). The learnable radar-vision matrix mapping sampling module 302 includes the MLP in the foregoing embodiment. The learnable bird's-eye view feature extraction module 303 based on the query vector includes the DAT and the weight attention network in the foregoing embodiment. The 3D detection decoder module 304 includes the decoder in the foregoing embodiment. The reference 3D points can also be referred to as anchor points; the reference features of each reference 3D point can be denoted as . Among them, M represents the number of reference 3D points, represents the current reference feature of the i-th reference 3D point, and C represents the number of output channels of the image feature extraction network. The coordinates and reference features of a reference 3D point can be called a 3D Query (3D query vector). According to the correlation between the Query (query vector) and the Key (key), the corresponding Value (value) can be found, and the weights of each Value can be obtained. Subsequently, based on the weights of each Value, the fusion result can be calculated. In the process of obtaining the weights of each position offset based on DAT in the present application, each position offset is the Value, and the weight of each position offset is the weight of the Value obtained based on the correlation between the Query and the Key; in the process of obtaining the weights of each initial image feature to be detected based on the Attention mechanism, each initial image feature to be detected is the Value, and the weight of each initial image feature to be detected is the weight of the Value obtained based on the correlation between the Query and the Key.

[0164] After the electronic device obtains the single camera / multi-camera image (i.e., one or more images to be detected in the foregoing embodiment, hereinafter referred to as the camera image for short) and the radar point cloud (i.e., the point cloud to be detected in the foregoing embodiment), it can input the camera image and the radar point cloud into the multi-sensor feature extraction module 301. Through the image backbone network included in the multi-sensor feature extraction module 301, the image features of the camera image (i.e., the initial image features to be detected in the foregoing embodiment) are extracted; through the radar backbone network, the grid features of the radar point cloud (i.e., the initial point cloud features to be detected of each BEV grid in the foregoing embodiment) are extracted.

[0165] The above-mentioned method of extracting the point cloud features to be detected for each BEV grid based on the point cloud feature extraction network can also be referred to as the method of extracting the Pillar features of each grid by using the method of constructing the Pillar grid. Specifically, the electronic device can divide the radar point cloud to determine the three-dimensional points included in each of the H×W BEV grids; for each BEV grid, determine the three-dimensional points to be detected in the BEV grid according to the number of three-dimensional points included in the BEV grid; use the radar backbone network to extract features from the information to be utilized of the three-dimensional points to be detected in the BEV grid. The extracted features are dimensionally increased to obtain a tensor with dimensions (C, N, P); pooling is performed in the dimension of N to obtain a tensor with dimensions (C, P); the tensor with dimensions (C, P) is deformed into a tensor with dimensions (C, H, W) as the grid feature of each BEV grid.

[0166] The electronic device uses the learnable radar-vision matrix mapping sampling module 302 to concat the extracted image features, grid features, and each three-dimensional coordinate (i.e., the coordinates of each reference three-dimensional point in the foregoing embodiments), and then uses the MLP to extract features from the concat features to obtain a correction amount for correcting the initial transformation matrix. Furthermore, the obtained correction amount is used to correct the RV (Radar-View) matrix (i.e., the initial transformation matrix in the foregoing embodiments). The transformation matrix is a homography matrix that can label the mapping relationship between the BEV space and the image pixel positions.

[0167] Exemplarily, refer to Figure 4 , Figure 4 which is a schematic diagram of the learnable radar-vision matrix mapping sampling module 302 provided in the embodiments of the present application. The radar features (i.e., the initial point cloud features to be detected of each BEV grid in the foregoing embodiments), three-dimensional coordinates, and image features are input into the learnable multi-layer perceptron module, and the output data of the learnable multi-layer perceptron module is used to correct the radar-vision matrix to obtain an optimized radar-vision matrix (i.e., the corrected transformation matrix in the foregoing embodiments).

[0168] Furthermore, through the learnable bird's-eye view feature extraction module 303 based on the query vector, for each reference three-dimensional point, the current reference feature of the reference three-dimensional point is input into the DAT included in the learnable bird's-eye view feature extraction module 303 based on the query vector to obtain a preset number of position offsets and the weight corresponding to each position offset; for each position offset, the current coordinates of the reference three-dimensional point are offset by the position offset, and the BEV grid to which the offset coordinates belong is determined as the BEV grid corresponding to the position offset; the initial point cloud features to be detected of the BEV grids corresponding to each position offset are feature-fused according to the weights corresponding to the position offsets to obtain the current point cloud feature to be detected and fused for the reference three-dimensional point.

[0169] Using the weight attention network included in the learnable bird's-eye view feature extraction module 303 based on the query vector, for each reference 3D point, input the current reference feature of the reference 3D point into the weight attention network to obtain the weight of each initial image feature to be detected; according to the weights of the initial image features to be detected, perform feature fusion on the feature values corresponding to the coordinates to be detected in each initial image feature to be detected, and obtain the intermediate image feature to be detected of the reference 3D point currently.

[0170] Then, the electronic device splices the current intermediate image feature to be detected and the current fused point cloud feature to be detected of each reference 3D point to obtain the current spliced feature to be detected, that is, the query vector feature of each reference 3D point.

[0171] Furthermore, the electronic device inputs the current spliced feature to be detected into the 3D detection decoder module 304, and decodes the current spliced feature to be detected through the decoder included in the 3D detection decoder module 304 to obtain the current decoding result; update the current coordinates and reference features of each reference 3D point based on the decoding result. Then send the current coordinates and reference features of each reference 3D point to the learnable bird's-eye view feature extraction module 303 based on the query vector. Through the learnable bird's-eye view feature extraction module 303 based on the query vector, obtain the current spliced feature to be detected based on the current coordinates and reference features of each reference 3D point, and then the 3D detection decoder module 304 decodes the current spliced feature to be detected. Repeat this process until the number of decoding times reaches the preset number of times to obtain the current coordinates of each reference 3D point. Subsequently, the electronic device can obtain the positions of the objects in the space region to be detected based on the current coordinates of each reference 3D point.

[0172] Based on the above processing, the electronic device can obtain the current fused point cloud feature to be detected based on the point cloud to be detected; obtain the current intermediate image feature to be detected based on the image to be detected. Update the current coordinates and reference features of each reference 3D point based on the current fused point cloud feature to be detected and the current intermediate image feature to be detected. After a preset number of updates, determine the positions of the objects in the space region to be detected according to the current coordinates of each reference 3D point. That is, perform object detection on the space region to be detected based on the point cloud and the image together. In this way, since the process of obtaining the point cloud is not easily affected by external interferences such as illumination changes, background color changes, and occlusions, combining the point cloud and the image for object detection can make up for the defect that the accuracy of the object detection result obtained only based on the image is not high, reduce the influence of external interferences on the accuracy of the object detection result, and improve the accuracy of the obtained object detection result.

[0173] According to the object detection method provided by the embodiments of the present application, in the radar-vision traffic scenario, the electronic device can jointly perform object detection of three-dimensional objects based on millimeter-wave radar and images (hereinafter referred to as radar-vision fusion three-dimensional detection). That is, in the digital intersection scenario, based on the radar-vision fusion end-to-end model framework of BEV Query, and based on the learnable RV matrix mapping sampling, a more robust "BEV-image" mapping is realized, making the radar-vision fusion three-dimensional detection more general and having stronger generalization ability. By performing detection through an end-to-end three-dimensional object detection model, the customization and maintenance cost of the post-fusion logic can be reduced, and the information of multiple sensors can be fully utilized.

[0174] Due to the differences in the performance of different sensors, for example, the images collected by visual sensors contain rich semantic information, so usually, the accuracy of the detection results obtained based on visual sensors is relatively high and has been widely used in perception tasks such as object detection and object segmentation. However, when affected by external interferences such as changes in illumination, background color, and occlusions, and in situations where the object movement route is complex and the background color is messy, the accuracy of the detection results obtained based on visual sensors will decrease significantly, that is, the robustness of object detection based on visual sensors is not high. Radar sensors can measure information such as the movement speed, relative distance, and azimuth angle of objects and are not easily affected by the above external interferences. However, the detection results obtained based on radar sensors are relatively sparse and noisy, that is, the robustness of object detection based on radar sensors is relatively high, but the accuracy of the obtained detection results is relatively not high. Therefore, by fusing the information of multiple sensors such as radar and vision, the all-weather perception ability of the algorithm can be improved, that is, the object detection method has high robustness and high accuracy.

[0175] Compared with the method of separately performing object detection based on visual information (i.e., images) and radar information (i.e., point clouds) and then performing post-fusion of radar and vision in the BEV space, since the post-fusion processing method has problems such as insufficient use of sensor information, many logical rules, and high maintenance costs, based on the object detection method provided in this application, by fusing visual information and radar information and then performing object detection based on the fused information, it has obvious generalization performance advantages. Compared with the post-fusion processing method, it has better versatility and stronger generalization ability. It is compatible with multiple sensors, such as pluggably supporting multi-modal data such as radar, vision, and lidar, and has extremely high multi-scenario adaptability. Moreover, compared with the post-fusion processing method in the image perspective, since the object scales are unified in the BEV space, the features in the BEV space can more intuitively represent information such as the spatial position, size, orientation, and distance of objects, and can provide a unified physical space to realize the information fusion of different sensors. Therefore, performing fusion in the BEV space, that is, subsequently detecting the position of objects in the BEV space, has better algorithm performance advantages compared with performing post-fusion in the image perspective.

[0176] Compared with the method of transforming an image into a pseudo-point cloud based on a display transformation, or, processing an image based on a depth prediction method to obtain BEV dense features and then extracting local point cloud features based on an attention mechanism, these methods have problems such as feature redundancy and a large amount of ineffective computational overhead. Based on the object detection method provided in this application, a learnable Query is used to represent the position of an object, image features are extracted based on the RV matrix mapping relationship, and feature fusion of different sensors is performed in the BEV space, thereby improving the overall 3D detection performance and algorithm generalization of radar-vision fusion. There is no need to map the entire image to be detected into the BEV space, which can reduce ineffective computational overhead.

[0177] This application also provides a model training method, which is also applied to an electronic device. To distinguish it from the aforementioned electronic device for object detection, the electronic device for executing the model training method is hereinafter referred to as the training device. The training device and the electronic device can be the same device, or they can also be different devices, and this application does not limit this. In the case where the training device and the electronic device are different devices, after the training device obtains a trained three-dimensional object detection model by executing the model training method provided in this application, it can send the trained three-dimensional object detection model to the electronic device, so that after the electronic device deploys the received three-dimensional object detection model, it performs object detection based on the deployed three-dimensional object detection model.

[0178] See Figure 5 , Figure 5 is a flowchart of a three-dimensional object detection model training method provided by an embodiment of this application. The method may include the following steps:

[0179] S501: Obtain the point cloud and image of the sample space region as the sample point cloud and sample image respectively, and obtain the sample label containing the position of the object within the sample space region.

[0180] S502: Divide the sample point cloud according to the bird's-eye view grid, and determine the sample 3D points included in each divided bird's-eye view grid; use the point cloud feature extraction network in the 3D object detection model with the initial structure to extract features from the information to be utilized of the sample 3D points included in each bird's-eye view grid, and obtain the sample initial point cloud features of each bird's-eye view grid; and input the sample image into the image feature extraction network in the 3D object detection model to obtain multiple sample initial image features.

[0181] Among them, the information to be utilized of a 3D point includes: the coordinates of the 3D point and the coordinates of the centroid of the BEV grid to which the 3D point belongs; the 3D object detection model also includes a deformable attention network, a weight attention network and a decoder.

[0182] S503: For each reference 3D point, input the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset; for each position offset, use the position offset to offset the current coordinates of the reference 3D point, and determine the bird's-eye view grid to which the offset coordinates belong as the bird's-eye view grid corresponding to the position offset; fuse the sample initial point cloud features of the bird's-eye view grids corresponding to each position offset according to the weights corresponding to each position offset to obtain the current sample fused point cloud feature of the reference 3D point.

[0183] S504: Input the current reference feature of the reference 3D point into the weight attention network to obtain the weight of each sample initial image feature; fuse the eigenvalue corresponding to the reference 3D point in each sample initial image feature according to the weights of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point.

[0184] S505: Concatenate the current sample fused point cloud feature and sample intermediate image feature of each reference 3D point to obtain the current sample concatenated feature.

[0185] S506: Input the current sample concatenated feature into the decoder to obtain the current decoding result, and update the current coordinates and reference features of the reference 3D point based on the current decoding result; return to execute the step of inputting the current reference feature of each reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset until the number of decoding times reaches the preset number of times.

[0186] S507: Obtain a detection result based on the current coordinates of each reference 3D point.

[0187] Among them, the detection result includes the position of the object in the sample space region.

[0188] S508: Based on the difference between the sample label and the detection result, adjust the model parameters of the 3D object detection model of the initial structure until the preset convergence condition is reached, and obtain a trained 3D object detection model.

[0189] Based on the 3D object detection model training method provided by this application, the training device can obtain a trained 3D object detection model, and then send the trained 3D object detection model to the electronic device, so that after the electronic device deploys the received 3D object detection model, it performs object detection based on the deployed 3D object detection model. When the electronic device performs object detection based on the end-to-end 3D object detection model, it combines the point cloud and the image to jointly perform object detection, improving the accuracy of the obtained object detection result, reducing the maintenance cost, and improving the versatility of the object detection method.

[0190] Regarding step S501, the sample space region and the categories of objects that may exist therein are related to the business requirements of the application scenario to which the trained 3D object detection model is to be applied. If the trained 3D object detection model needs to be applied in a digital intersection scenario, the sample space region may include the roads around the traffic intersection, and the objects that may exist therein may include vehicles, roadblocks, etc.

[0191] By scanning the sample space region with a vision sensor, an image of the sample space region can be collected; by scanning the sample space region with a radar sensor, a point cloud of the sample space region can be collected. The training device can obtain the image of the sample space region from the vision sensor as a sample image, and obtain the point cloud of the sample space region from the radar sensor as a sample point cloud.

[0192] Regarding step S502, the training device can divide the sample space region into multiple BEV grids, and one grid can be called a Pillar. The sample point cloud includes multiple 3D points. For each 3D point, the training device can determine the Pillar to which the 3D point belongs, that is, determine the BEV grid to which the 3D point belongs. In this way, the training device also divides the sample point cloud according to the BEV grid, and determines the 3D points included in each BEV grid.

[0193] The coordinates of the three-dimensional points can be the coordinates in the world coordinate system. When dividing the sample point cloud, for each three-dimensional point included in the sample point cloud, the training device can determine the area to which the three-dimensional point belongs only according to the coordinate value of the three-dimensional point on the X-axis and the coordinate value on the Y-axis, and the BEV grid corresponding to the area is the BEV grid to which the three-dimensional point belongs. It can reduce the computational complexity when dividing the sample point cloud, reduce the division time-consuming, and further reduce the model training time-consuming, improving the model training efficiency.

[0194] For each BEV grid, after the training device determines the three-dimensional points included in the BEV grid, it can determine the minimum circumscribed cube of the three-dimensional points included in the BEV grid, and then use the centroid of the minimum circumscribed cube as the centroid of the BEV grid.

[0195] Since the format of the input data of the point cloud feature extraction network is set, and the number of three-dimensional points included in each BEV grid may be different, therefore, in order to obtain the sample initial point cloud features of each BEV grid through the point cloud feature extraction network, the training device can determine the sample three-dimensional points included in the BEV grid according to the number of three-dimensional points included in each BEV grid.

[0196] Furthermore, for each BEV grid, the electronic device can input the sample three-dimensional points included in the BEV grid into the point cloud feature extraction network in the three-dimensional object detection model with the initial structure, and the point cloud feature extraction network extracts the features of the information to be utilized of the sample three-dimensional points included in the BEV grid to obtain the sample initial point cloud features of the BEV grid.

[0197] In some embodiments, when the radar sensor can collect the coordinate value of the three-dimensional point on the Z-axis, the coordinate value of the three-dimensional point included in the information to be utilized of a three-dimensional point can be: the real coordinate value collected by the radar sensor. When the radar sensor fails to collect the coordinate value of the three-dimensional point on the Z-axis, a preset virtual coordinate value can be used as the coordinate value of each three-dimensional point on the Z-axis. For example, the preset virtual coordinate value can be set to 0 or 1.

[0198] After the training device obtains the sample image, it can use the image feature extraction network to extract the features of the sample image to obtain multiple sample initial image features. For example, the image feature extraction network can include ResNet, VGG, etc. in the foregoing embodiments, but is not limited thereto.

[0199] For step S503, for each reference 3D point, the training device can input the current reference feature of the reference 3D point into the DAT, and obtain the number of position offsets that need to be obtained and are preset. Based on the deformable attention mechanism, the DAT processes the current reference feature of the reference 3D point, and can output the set number of position offsets and the weight corresponding to each position offset. For each position offset, the training device can use the position offset to offset the current coordinates of the reference 3D point.

[0200] In the case where the training device directly generates the coordinates within the sample space region as the coordinates of the reference 3D point, the training device can directly calculate the sum of the current coordinates of the reference 3D point and the position offset, and the calculated coordinates are the offset coordinates. In the case where the coordinates of the reference 3D point are obtained by the training device based on a number randomly determined from the interval [0, 1], the training device can first scale the coordinates of the reference 3D point according to the size of the sample space region, and then calculate the sum of the scaled coordinates and the position offset, and the calculated coordinates are the offset coordinates.

[0201] The BEV grid to which the offset coordinates belong, that is, the BEV grid corresponding to the position offset, is also the BEV grid sampled based on the position offset. According to the weights corresponding to the respective position offsets, calculate the weighted sum of the sample initial point cloud features of the BEV grids corresponding to the respective position offsets, and obtain the current sample fusion point cloud feature of the reference 3D point.

[0202] For step S504, for each reference 3D point, the training device can input the current reference feature of the reference 3D point into the weight attention network. Based on the Attention mechanism, the weight attention network processes the current reference feature of the reference 3D point, and can output the weights of the respective sample initial image features. After the training device determines the part corresponding to the reference 3D point in the respective sample initial image features, that is, after determining the eigenvalue corresponding to the reference 3D point in the respective sample initial image features, it can calculate the weighted sum of the determined eigenvalues according to the weights of the respective sample initial image features, and obtain the current sample intermediate image feature of the reference 3D point.

[0203] For steps S505 and S506, after the training device splices the current sample fusion point cloud features and sample intermediate image features of each reference 3D point to obtain the current sample splicing features, it can input the current sample splicing features into the decoder to obtain the current decoding result. The decoder can be the Decoder based on the deep learning model in the foregoing embodiments, such as the decoding module in Transformer. The decoding result can include the position offset of each reference 3D point, the reference feature offset of each reference 3D point, etc. Based on the offsets obtained by decoding, the training device can update the current coordinates, reference features, etc. of each reference 3D point.

[0204] Then, the training device can use the current coordinates and reference features of each reference 3D point updated by this decoding, and return to execute the step of, for each reference 3D point, inputting the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset. That is, based on the current coordinates and reference features of each reference 3D point, the current sample splicing features are determined again, so as to perform the next decoding based on the current sample splicing features and update the coordinates and reference features of each reference 3D point again until the number of decoding times reaches the preset number. The preset number can be determined based on the type of decoder used.

[0205] For steps S507 and S508, after the number of decoding times reaches the preset number, for each reference 3D point, the current coordinate of the reference 3D point is the position of the object existing in the predicted sample space region, that is, the obtained detection result. Then the training device can obtain the difference between the sample label and the detection result. For example, the training device can calculate the loss function value representing the difference between the sample label and the detection result based on a preset loss function. The preset loss function can be a cross-entropy loss function, a mean square error loss function, etc.

[0206] In some embodiments, the decoding result can also include the confidence of the reference 3D point, that is, the probability that there is an object at the position indicated by the reference 3D point. When the confidence of a reference 3D point is greater than the confidence threshold, it means that the probability of an object existing at the reference 3D point is relatively high. Then the coordinates of the reference 3D point can be used as the position of the object existing in the detected sample space region, as the target detection result of the sample space region. Subsequently, the training device can determine the difference between the detection result determined based on the confidence and the sample label.

[0207] Then, the training device can adjust the model parameters of the three-dimensional object detection model with the initial structure based on the obtained differences until the preset convergence condition is reached, and obtain the trained three-dimensional object detection model. That is, the parameters of the point cloud feature extraction network, image feature extraction network, deformable attention network, weight attention network, and decoder included in the three-dimensional object detection model with the initial structure are adjusted. The preset convergence condition can include any one of the following: the number of training times reaches the preset number of training times, and the difference between the loss function value calculated this time and the loss function value calculated last time is less than the preset difference threshold.

[0208] In some embodiments, when the trained three-dimensional object detection model is obtained, the current coordinates and reference features of each reference three-dimensional point can be the coordinates and reference features of the reference three-dimensional points directly obtained by the electronic device in the foregoing object detection method. Then, based on the foregoing object detection method, the electronic device will adjust the obtained coordinates and reference features of the reference three-dimensional points, that is, adjust the coordinates and reference features of the reference three-dimensional points obtained by training, which can improve the accuracy of the position of the object in the to-be-detected spatial region obtained by detection compared with adjusting the coordinates and reference features of the reference three-dimensional points generated by the electronic device.

[0209] In some embodiments, determining the sample three-dimensional points included in each BEV grid obtained by division may include the following steps: for each BEV grid obtained by division, if the number of three-dimensional points in the sample point cloud included in the BEV grid is greater than the specified number, then sample the specified number of three-dimensional points from the three-dimensional points included in the BEV grid as the sample three-dimensional points included in the BEV grid; if the number of three-dimensional points in the sample point cloud included in the BEV grid is equal to the specified number, then use the three-dimensional points included in the BEV grid as the sample three-dimensional points included in the BEV grid; if the number of three-dimensional points in the sample point cloud included in the BEV grid is less than the specified number, then calculate the difference between the specified number and the number of three-dimensional points in the sample point cloud included in the BEV grid, and generate the difference number of three-dimensional points with preset coordinates; combine the generated three-dimensional points and the three-dimensional points in the sample point cloud included in the BEV grid to obtain the sample three-dimensional points included in the BEV grid.

[0210] For each BEV grid obtained by division, when the number of three-dimensional points in the sample point cloud included in the BEV grid is greater than the specified number, that is, not all the to-be-exploited information of the three-dimensional points included in the BEV grid needs to be input into the point cloud feature extraction network. At this time, the training device can sample the three-dimensional points included in the BEV grid and use the sampled specified number of three-dimensional points as the sample three-dimensional points included in the BEV grid. Correspondingly, in the input data of the point cloud feature extraction network, the data corresponding to the BEV grid is the to-be-exploited information of the sampled sample three-dimensional points.

[0211] When the number of three-dimensional points in the sample point cloud included in the BEV grid is equal to the specified number, that is, all the information to be utilized of the three-dimensional points included in the BEV grid is input into the point cloud feature extraction network. At this time, the training device can directly determine all the three-dimensional points included in the BEV grid as sample three-dimensional points. Correspondingly, in the input data of the point cloud feature extraction network, the data corresponding to the BEV grid is the information to be utilized of all the three-dimensional points in the BEV grid.

[0212] When the number of three-dimensional points in the sample point cloud included in the BEV grid is less than the specified number, even if all the three-dimensional points included in the BEV grid are determined as sample three-dimensional points, the data corresponding to the BEV grid in the above-mentioned set data format cannot be obtained. At this time, the training device can perform supplementary processing on the data corresponding to the BEV grid. Specifically, the training device can calculate the difference between the specified number and the number of three-dimensional points in the sample point cloud included in the BEV grid, and generate three-dimensional points with the difference number of coordinates being preset values. When each value in the information to be utilized of a three-dimensional point is a preset value, it can indicate that the three-dimensional point is used to fill the input data to make the input data have a fixed format, rather than the three-dimensional points included in the point cloud of the sample space region. For example, the preset value can be 0 or 1.

[0213] Based on the above processing, when the number of three-dimensional points included in a BEV grid is not the specified number, the training device can determine the sample three-dimensional points included in the BEV grid based on the number of three-dimensional points included in the BEV grid, in a sampling or supplementary processing manner. Subsequently, the point cloud feature extraction network can be used to extract the features of the information to be utilized of the sample three-dimensional points included in each BEV grid, reducing the probability of problems such as the low accuracy of the sample initial point cloud features extracted from each BEV grid due to the inconsistent number of three-dimensional points included in each BEV grid, and thus improving the accuracy of the target detection results obtained based on the sample initial point cloud features of each BEV grid.

[0214] In some embodiments, the information to be utilized of a three-dimensional point further includes at least one of the following: the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, and the feature data at the real position corresponding to the detected three-dimensional point.

[0215] When the information to be utilized for a three-dimensional point further includes the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, that is, in the input data of the point cloud feature extraction network, data indicating the relationship between some input data has been carried, the difficulty for the point cloud feature extraction network to learn the relationship between the input data will be lower, and the relationship between the learned input data will be more accurate. In this case, the initial point cloud features of each BEV grid obtained will be more accurate.

[0216] When the information to be utilized for a three-dimensional point further includes the feature data at the true position corresponding to the detected three-dimensional point, the point cloud feature extraction network performs feature extraction based on the feature data at each actually detected true position. Then, the sample point cloud features of each BEV grid extracted are more in line with the actual scenario, that is, more accurate. Consequently, the accuracy of the object detection results obtained based on the more accurate sample point cloud features of each BEV grid will be higher. For example, the feature data can be signal-to-noise ratio, reflection intensity, etc. When the radar sensor can detect multiple types of feature data at the true position corresponding to the three-dimensional point, the information to be utilized for the three-dimensional point can also include the multiple types of feature data to further improve the accuracy of the initial point cloud features of each BEV grid extracted.

[0217] In some embodiments, the image feature extraction network includes a plurality of downsampling layers connected in series; there is one sample image, and the multiple initial sample image features include: the image features of the sample image extracted by the plurality of downsampling layers; or, the sample image includes images of the sample space region collected at multiple different shooting angles, and the multiple initial sample image features include the image features of each sample image.

[0218] After inputting the sample image into the first downsampling layer in the series connection, the plurality of downsampling layers can sequentially perform downsampling processing on their input data. Then, the output data of different downsampling layers, that is, the image features of the sample image extracted by the plurality of downsampling layers, are different initial sample image features.

[0219] Therefore, when there is one sample image, multiple initial sample image features can be extracted through the plurality of downsampling layers connected in series included in the image feature extraction network.

[0220] When there are multiple sample images, that is, when multiple vision sensors collect images of the sample space region from different shooting angles, for each sample image, through the image feature extraction network, one image feature of the sample image can be extracted; or, in the same way as extracting multiple initial sample image features of one sample image described above, multiple initial sample image features of each sample image can be extracted respectively.

[0221] Based on the above processing, multiple sample initial image features of the sample image are extracted through the image feature extraction network, that is, more abundant and information-rich sample initial image features can be obtained. Then, obtaining the prediction label based on the more information-rich sample initial image features is to obtain the prediction label by integrating more information, which can improve the accuracy of the obtained prediction label.

[0222] In some embodiments, the 3D object detection model further includes an MLP (Multilayer Perceptron).

[0223] Before inputting the current reference feature of each reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset, the method may further include the following steps: Step 1: Input the sample initial point cloud features and sample initial image features of each BEV grid into the multilayer perceptron to obtain a correction amount for correcting the initial transformation matrix. The initial transformation matrix is calibrated according to the acquisition methods of the sample point cloud and the sample image. Step 2: Use the obtained correction amount to correct the initial transformation matrix.

[0224] Correspondingly, the above steps of fusing the eigenvalue corresponding to the reference 3D point in each sample initial image feature according to the weight of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point may include the following steps: Step 3: Use the corrected transformation matrix to transform the current coordinates of the reference 3D point to obtain the pixel coordinates corresponding to the reference 3D point in the sample image as the sample reference coordinates. Step 4: Fuse the eigenvalue corresponding to the sample reference coordinates in each sample initial image feature according to the weight of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point.

[0225] The initial transformation matrix can be calibrated according to the acquisition methods of the sample point cloud and the sample image. The method of calibrating according to the acquisition methods of the sample point cloud and the sample image is similar to the method of calibrating according to the acquisition methods of the point cloud to be detected and the image to be detected, and the relevant introduction of the foregoing embodiments can be referred to.

[0226] The training device can concat the sample initial point cloud features and sample initial image features of each BEV grid. Then, input the features obtained by concat into the MLP to obtain a 3×3 matrix output by the MLP as the correction amount for correcting the initial transformation matrix.

[0227] In some embodiments, inputting the sample initial point cloud features and the sample initial image features of each BEV grid into a multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix may include the following steps: Inputting the sample initial point cloud features, the sample initial image features of each BEV grid, and the current coordinates of each reference 3D point into a multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix.

[0228] It can be understood that at this time, decoding has not been performed based on the reference 3D points, nor has the coordinates of the reference 3D points been updated. When the training device generates the coordinates of the reference 3D points, the current coordinates of each reference 3D point are the coordinates generated by the training device; when the training device performs object detection on the sample space region based on the 3D object detection model, the current coordinates of each reference 3D point are the current coordinates of each reference 3D point after the 3D object detection model is trained.

[0229] The training device can concat the sample initial point cloud features of each BEV grid, multiple sample initial image features, and the current coordinates of each reference 3D point, and then input the features obtained by concat into the MLP. The MLP performs feature extraction on the features obtained by concat to obtain a 3×3 matrix output by the MLP as the correction amount for correcting the initial transformation matrix. In this way, that is, the current coordinates of the reference 3D points that comprehensively indicate the positions of the objects in the sample space region are jointly used to determine the correction amount, and the correction amount can indicate the positions where there are objects in the sample space region. After correcting the initial transformation matrix based on the correction amount, the positions of the objects indicated by the sample reference positions obtained based on the corrected transformation matrix are more accurate, thereby improving the accuracy of the trained 3D object detection model and improving the accuracy of the subsequent object detection results obtained based on the 3D object detection model.

[0230] Then, the training device can use the obtained correction amount to correct the initial transformation matrix. That is, for each element in the initial transformation matrix, calculate the sum value of the element and the element at the same position in the obtained correction amount, and update the element. Updating the elements in the initial transformation matrix means using the obtained correction amount to correct the initial transformation matrix. Furthermore, for each reference 3D point, the training device can use the corrected transformation matrix to transform the coordinates of the reference 3D point to obtain the pixel coordinates in the pixel coordinate system, that is, obtain the corresponding sample reference coordinates in the sample image. The feature values corresponding to the sample reference coordinates in each sample initial image feature are the feature values corresponding to the reference 3D point in each sample initial image feature.

[0231] Based on the above processing, the training device can correct the initially calibrated transformation matrix based on the sample initial point cloud features and sample initial image features of each BEV grid. That is, the initially calibrated transformation matrix is corrected based on the actual application scenario, and the corrected transformation matrix is more suitable for the actual application scenario compared to the initial transformation matrix. The current sample intermediate image features obtained based on the corrected transformation matrix are more accurate than the current sample intermediate image features obtained based on the initial transformation matrix. Then, the accuracy of the 3D object detection model trained based on the more accurate current sample intermediate image features is higher, improving the accuracy of the object detection results obtained based on the 3D object detection model.

[0232] In some embodiments, since the training device needs to generate a correction amount for correcting the initially calibrated transformation matrix based on the sample initial point cloud features of each BEV grid, that is, the sample initial point cloud features of each BEV grid also need to have a specified format. After the training device obtains the features output by the point cloud feature extraction network, it can also increase the dimension of the output features to obtain a tensor with dimensions (C, N, P); then, for the tensor with increased dimensions of each BEV grid, pooling is performed in the N dimension to obtain a tensor with dimensions (C, P). As described above, P is the product of H and W, so the tensor with dimensions (C, P) can be transformed into a tensor with dimensions (C, H, W) as the sample initial point cloud features of each BEV grid. Here, C is the number of output channels of the image feature extraction network.

[0233] For each reference 3D point, the sample reference coordinates obtained by performing coordinate transformation on the current coordinates of the reference 3D point may not be integers, that is, the sample reference coordinates may not correspond to pixel points in the sample image.

[0234] In some embodiments, when the sample reference coordinates are not integers, the training device can perform rounding processing on the non-integer coordinate values of the sample reference coordinates, and use the feature value corresponding to the rounded coordinates in each sample initial image feature as the feature value corresponding to the sample reference coordinates.

[0235] In some embodiments, before fusing the feature values corresponding to the sample reference coordinates in each sample initial image feature according to the weights of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point, the method may further include the following steps: if the sample reference coordinates are not integers, obtain other pixel coordinates within the neighborhood range of the sample reference coordinates; for each sample initial image feature, interpolate the feature values corresponding to the obtained other pixel coordinates in the sample initial image feature as the feature values corresponding to the sample reference coordinates in the sample initial image feature.

[0236] The training device can obtain other pixel coordinates within the neighborhood of the sample reference coordinates. For example, for the coordinate values in the sample reference coordinates that are not integers, rounding up and rounding down are respectively performed, and the pixel coordinates at the coordinate values determined by rounding are determined as the other pixel coordinates within the neighborhood of the sample reference coordinates.

[0237] At this time, there are corresponding feature values for the other pixel coordinates in each sample initial image feature. For each sample initial image feature, the training device can interpolate the feature values corresponding to the obtained other pixel coordinates in this sample initial image feature as the feature value corresponding to the sample reference coordinate in this sample initial image feature.

[0238] Based on the above processing, when the sample reference coordinate is not an integer, the training device can still comprehensively determine the feature value corresponding to the sample reference coordinate by combining each other pixel coordinate within the neighborhood of the sample reference coordinate. Then, the feature value corresponding to the sample reference coordinate determined in this way is more accurate, and the current sample intermediate image feature of this reference three-dimensional point calculated based on the more accurate feature value is more accurate, which can improve the accuracy of the target detection result obtained based on the current sample intermediate image feature.

[0239] In some embodiments, the sample label further includes: for each object within the sample space region, the category of the object; the detection result further includes: for each reference three-dimensional point, the category of the object at the position indicated by this reference three-dimensional point based on the decoding result; the differences between the sample label and the detection result include: the position difference between the position of the object included in the sample label and the position of the object included in the detection result; and the category difference between the category of the object included in the sample label and the category of the object included in the detection result.

[0240] For each object within the sample space region, the position of the object included in the sample label can indicate: the size of the minimum bounding cube of the object, the coordinates of the center point of the minimum bounding cube of the object, and the orientation of the object.

[0241] Correspondingly, for each reference three-dimensional point, the position of the object at the position indicated by this reference three-dimensional point included in the detection result can indicate: the size of the minimum bounding cube of the object, the coordinates of the center point of the minimum bounding cube of the object, and the orientation of the object.

[0242] It can be understood that there may be multiple objects within the sample space region; there may be multiple reference 3D points indicating the positions of the objects. Therefore, for each object within the sample space region, according to the position of the object included in the sample label, the position of the minimum bounding cube of the object (hereinafter referred to as the sample position) can be determined. Furthermore, the training device can calculate the intersection over union (IoU) between the sample position and the positions indicated by the respective reference 3D points (hereinafter referred to as the detection positions), and determine the detection positions with an IoU greater than the preset IoU threshold as the detection results corresponding to the object. Obviously, the detection results may be one or may include multiple ones.

[0243] As mentioned above, the position of an object can include the size of the minimum bounding cube of the object, the coordinates of the center point of the minimum bounding cube of the object, and the orientation of the object; alternatively, the position of an object can also include the coordinates of the respective vertices of the minimum bounding cube of the object.

[0244] Then, in the case where the position of an object includes the size of the minimum bounding cube of the object, the coordinates of the center point of the minimum bounding cube of the object, and the orientation of the object, for each detection result corresponding to the object, the training device can calculate the size difference between the size of the minimum bounding cube of the object included in the sample label and the size of the minimum bounding cube of the object included in the detection result; calculate the coordinate difference between the coordinates of the center point of the minimum bounding cube of the object included in the sample label and the coordinates of the center point of the minimum bounding cube of the object included in the detection result; calculate the orientation difference between the orientation of the minimum bounding cube of the object included in the sample label and the orientation of the minimum bounding cube of the object included in the detection result. The size difference, coordinate difference, and orientation difference are the position differences between the object and the detection result. Subsequently, the training device can train the 3D object detection model with the initial structure based on the position differences between each object and its respective detection results.

[0245] Then, in the case where the position of an object includes the coordinates of the respective vertices of the minimum bounding cube of the object, for each detection result corresponding to the object, according to the coordinates of the respective vertices of the minimum bounding cube of the object included in the sample label (hereinafter referred to as the sample vertices) and the coordinates of the respective vertices included in the detection result (hereinafter referred to as the detection vertices), the sample vertex and the detection vertex (hereinafter referred to as the vertex pair) indicating the same vertex of the minimum bounding cube of the object can be determined. Furthermore, for each vertex pair, the training device can calculate the difference between the coordinates of the sample vertex and the detection vertex included in the vertex pair. The respective differences of each vertex pair are the position differences between the object and the detection result. Subsequently, the training device can train the 3D object detection model with the initial structure based on the position differences between each object and its respective detection results.

[0246] Moreover, when performing object detection based on machine vision, for each detected object, the obtained object detection result can also indicate the category of the object. Correspondingly, when training a three-dimensional object detection model for object detection, for each object in the sample space region, the sample label can also include the category of the object; for each reference three-dimensional point, the detection result further includes: the category of the object at the position indicated by the reference three-dimensional point obtained based on the decoding result.

[0247] It can be understood that the category of an object can be represented in the form of a vector. For example, the first element in the vector represents the probability that the category of the object is category A; the second element in the vector represents the probability that the category of the object is category B. If the category of the object in the sample space region is category A, the sample label can include vector 1 (1, 0). If the detection result includes vector 2 (0.7, 0.3), the training device can calculate the loss function value representing the difference between vector 1 and vector 2 based on the loss function, that is, obtain the category difference of the object in the sample space region between the sample label and the detection result, and then adjust the model parameters of the three-dimensional object detection model based on the loss function value.

[0248] Correspondingly, in the embodiments of the present application, for each reference three-dimensional point, the decoding result further includes: the coordinate adjustment amount of the center point of the object (hereinafter referred to as the first object) existing at the reference three-dimensional point, the size adjustment amount of the minimum circumscribed cube of the first object, the orientation adjustment amount of the minimum circumscribed cube of the first object, and the category adjustment amount of the first object. The training device can use the decoding result to update the current coordinates of the reference three-dimensional point, the size of the minimum circumscribed cube of the first object, the orientation of the first object, the category of the first object, etc.

[0249] Based on the above processing, training the three-dimensional object detection model according to the positions and categories of the objects in the sample space region, then through the trained three-dimensional object detection model, the categories of the objects in the to-be-detected space region can also be detected, meeting the business requirements.

[0250] Based on the same inventive concept as the above object detection method, the present application further provides an object detection device. Refer to Figure 6 , Figure 6 which is a structural diagram of the object detection device provided by the embodiments of the present application. The device includes:

[0251] A first acquisition module 601, configured to acquire the point cloud and image of the to-be-detected space region, respectively, as the to-be-detected point cloud and the to-be-detected image;

[0252] The first feature extraction module 602 is configured to divide the to-be-detected point cloud according to the bird's-eye view (BEV) grid, extract the to-be-detected initial point cloud features of each obtained BEV grid, and extract features from the to-be-detected image to obtain the to-be-detected initial image features;

[0253] The first point cloud feature fusion module 603 is configured to, for each reference 3D point, sample the BEV grids within the neighborhood range of the reference 3D point based on the current coordinates and reference features of the reference 3D point, and fuse the to-be-detected initial point cloud features of the sampled BEV grids to obtain the to-be-detected fused point cloud features of the reference 3D point at present;

[0254] The first image feature fusion module 604 is configured to obtain the corresponding part of the reference 3D point in the to-be-detected initial image features to obtain the to-be-detected intermediate image features of the reference 3D point at present;

[0255] The first splicing module 605 is configured to splice the to-be-detected fused point cloud features and the to-be-detected intermediate image features of each reference 3D point at present to obtain the to-be-detected splicing features at present;

[0256] The first update module 606 is configured to decode the to-be-detected splicing features at present to obtain the decoding result at present, and update the current coordinates and reference features of the reference 3D point based on the decoding result at present; return to execute the step of, for each reference 3D point, sampling the BEV grids within the neighborhood range of the reference 3D point based on the current coordinates and reference features of the reference 3D point, and fusing the to-be-detected initial point cloud features of the sampled BEV grids to obtain the to-be-detected fused point cloud features of the reference 3D point at present, until the number of decoding times reaches a preset number;

[0257] The position acquisition module 607 is configured to obtain the positions of the objects in the to-be-detected spatial region based on the current coordinates of each reference 3D point.

[0258] Optionally, the first feature extraction module 602 is specifically configured to divide the to-be-detected point cloud according to a bird's-eye view (BEV) grid, and determine the to-be-detected three-dimensional points included in each BEV grid obtained by the division; use a point cloud feature extraction network in a pre-trained three-dimensional object detection model to extract features from the to-be-used information of the to-be-detected three-dimensional points included in each BEV grid, so as to obtain the to-be-detected initial point cloud features of each BEV grid; and input the to-be-detected image into an image feature extraction network in the three-dimensional object detection model to obtain a plurality of to-be-detected initial image features; wherein, the to-be-used information of a three-dimensional point includes: the coordinates of the three-dimensional point and the coordinates of the centroid of the BEV grid to which the three-dimensional point belongs; the three-dimensional object detection model is: trained based on sample images, sample point clouds in a sample space region, and sample labels indicating the positions of objects in the sample space region; the three-dimensional object detection model further includes a deformable attention network, a weight attention network, and a decoder; the first point cloud feature fusion module 603 is specifically configured to, for each reference three-dimensional point, input the current reference feature of the reference three-dimensional point into the deformable attention network to obtain a preset number of position offsets and the weights corresponding to each position offset; for each position offset, use the position offset to offset the current coordinates of the reference three-dimensional point, and determine the BEV grid to which the coordinates after the offset belong as the BEV grid corresponding to the position offset; fuse the to-be-detected initial point cloud features of the BEV grids corresponding to each position offset according to the weights corresponding to each position offset to obtain the to-be-detected fused point cloud feature of the reference three-dimensional point at present; the first image feature fusion module 604 is specifically configured to input the current reference feature of the reference three-dimensional point into the weight attention network to obtain the weights of each to-be-detected initial image feature; fuse the feature values corresponding to the reference three-dimensional point in each to-be-detected initial image feature according to the weights of each to-be-detected initial image feature to obtain the to-be-detected intermediate image feature of the reference three-dimensional point at present; the first update module 606 is specifically configured to input the current to-be-detected splicing feature into the decoder to obtain the current decoding result.

[0259] Optionally, the 3D object detection model further includes a multi-layer perceptron; the apparatus further includes: a first correction amount acquisition module, configured to, before the first point cloud feature fusion module 603 performs, for each reference 3D point, inputting the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset, input the initial point cloud features to be detected of each BEV grid and the initial image features to be detected into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix; wherein, the initial transformation matrix is calibrated according to the acquisition manners of the point cloud to be detected and the image to be detected; correcting the initial transformation matrix by using the obtained correction amount; the first image feature fusion module 604 is specifically configured to use the corrected transformation matrix to transform the current coordinates of the reference 3D point to obtain the pixel coordinates corresponding to the reference 3D point in the image to be detected as the reference coordinates to be detected; fusing the feature values corresponding to the reference coordinates to be detected in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the intermediate image feature to be detected of the reference 3D point currently.

[0260] Optionally, the first correction amount acquisition module is specifically configured to input the initial point cloud features to be detected of each BEV grid, the initial image features to be detected, and the current coordinates of each reference 3D point into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix.

[0261] Optionally, the apparatus further includes: a first interpolation module, configured to, before the first image feature fusion module 604 performs fusing the feature values corresponding to the reference coordinates to be detected in each initial image feature to be detected according to the weights of each initial image feature to be detected to obtain the intermediate image feature to be detected of the reference 3D point currently, if the reference coordinates to be detected are not integers, obtain other pixel coordinates within the neighborhood range of the reference coordinates to be detected; for each initial image feature to be detected, interpolate the feature values corresponding to the obtained other pixel coordinates in the initial image feature to be detected as the feature values corresponding to the reference coordinates to be detected in the initial image feature to be detected.

[0262] Optionally, the image feature extraction network includes a plurality of downsampling layers connected in series; there is one image to be detected, and the plurality of initial image features to be detected include: the image features of the image to be detected extracted by the plurality of downsampling layers; or, the image to be detected includes images of the space region to be detected collected at multiple different shooting angles, and the plurality of initial image features to be detected include the image features of each image to be detected.

[0263] Optionally, the first feature extraction module 602 is specifically configured to, for each BEV grid obtained by partitioning, if the number of three-dimensional points in the point cloud to be detected included in the BEV grid is greater than the specified number, sample the specified number of three-dimensional points from the three-dimensional points included in the BEV grid as the three-dimensional points to be detected included in the BEV grid; if the number of three-dimensional points in the point cloud to be detected included in the BEV grid is equal to the specified number, use the three-dimensional points included in the BEV grid as the three-dimensional points to be detected included in the BEV grid; if the number of three-dimensional points in the point cloud to be detected included in the BEV grid is less than the specified number, calculate the difference between the specified number and the number of three-dimensional points in the point cloud to be detected included in the BEV grid, and generate the difference number of three-dimensional points with preset coordinates; combine the generated three-dimensional points and the three-dimensional points in the point cloud to be detected included in the BEV grid to obtain the three-dimensional points to be detected included in the BEV grid.

[0264] Optionally, the information to be utilized for a three-dimensional point further includes at least one of the following: the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, and the feature data at the true position corresponding to the detected three-dimensional point.

[0265] Based on the target detection device provided in the embodiments of the present application, an electronic device can obtain the current feature of the point cloud to be detected for fusion based on the point cloud to be detected, and obtain the current intermediate image feature to be detected based on the image to be detected. Based on the current feature of the point cloud to be detected for fusion and the current intermediate image feature to be detected, update the current coordinates and reference features of each reference three-dimensional point. After a preset number of updates, determine the positions of the objects in the space region to be detected according to the current coordinates of each reference three-dimensional point. That is, perform target detection on the space region to be detected based on the point cloud and the image jointly. In this way, since the process of obtaining the point cloud is not easily affected by external interferences such as changes in illumination, background color, and occlusions, and target detection is performed jointly by combining the point cloud and the image, it can make up for the defect that the accuracy of the target detection result obtained only based on the image is not high, reduce the influence of external interferences on the accuracy of the target detection result, and improve the accuracy of the obtained target detection result.

[0266] In the technical solution of the present application, operations such as acquisition, storage, use, processing, transmission, provision, and disclosure of the point cloud and the image are all carried out under the condition of obtaining authorization. It should be noted that the three-dimensional target detection model in this embodiment is not a three-dimensional target detection model for a specific scenario and cannot reflect the scenario information of a specific scenario. It should be noted that the point cloud and the image in this embodiment come from a public data set.

[0267] Based on the same inventive concept as the above three-dimensional object detection model training method, this application also provides a three-dimensional object detection model training device. Refer to Figure 7 , Figure 7 which is a structural diagram of the three-dimensional object detection model training device provided by the embodiments of this application. The device includes:

[0268] A second acquisition module 701, configured to acquire the point cloud and image of the sample space region as the sample point cloud and sample image respectively, and acquire the sample label including the position of the object within the sample space region;

[0269] A second feature extraction module 702, configured to divide the sample point cloud according to the bird's-eye view (BEV) grid, and determine the sample three-dimensional points included in each BEV grid obtained by the division; use the point cloud feature extraction network in the three-dimensional object detection model with the initial structure to extract features from the information to be utilized of the sample three-dimensional points included in each BEV grid, so as to obtain the sample initial point cloud feature of each BEV grid; and input the sample image into the image feature extraction network in the three-dimensional object detection model to obtain a plurality of sample initial image features;

[0270] Wherein, the information to be utilized of a three-dimensional point includes: the coordinates of the three-dimensional point and the coordinates of the centroid of the BEV grid to which the three-dimensional point belongs; the three-dimensional object detection model further includes a deformable attention network, a weight attention network and a decoder;

[0271] A second point cloud feature fusion module 703, configured to input the current reference feature of each reference three-dimensional point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset; for each position offset, use the position offset to offset the current coordinates of the reference three-dimensional point, and determine the BEV grid to which the offset coordinates belong as the BEV grid corresponding to the position offset; fuse the sample initial point cloud features of the BEV grids corresponding to each position offset according to the weights corresponding to each position offset to obtain the current sample fusion point cloud feature of the reference three-dimensional point;

[0272] A second image feature fusion module 704, configured to input the current reference feature of the reference three-dimensional point into the weight attention network to obtain the weight of each sample initial image feature; fuse the feature values corresponding to the reference three-dimensional point in each sample initial image feature according to the weights of each sample initial image feature to obtain the current sample intermediate image feature of the reference three-dimensional point;

[0273] A second splicing module 705, configured to splice the current sample fusion point cloud feature and sample intermediate image feature of each reference three-dimensional point to obtain the current sample splicing feature;

[0274] A second update module 706, configured to input the current sample stitching feature into the decoder to obtain the current decoding result, and update the current coordinates and reference features of the reference 3D point based on the current decoding result; return to execute the step of, for each reference 3D point, inputting the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset, until the number of decoding times reaches the preset number;

[0275] A detection result acquisition module 707, configured to obtain a detection result based on the current coordinates of each reference 3D point; wherein, the detection result includes the position of an object in the sample space region;

[0276] A training module 708, configured to adjust the model parameters of the 3D object detection model with the initial structure based on the difference between the sample label and the detection result until the preset convergence condition is reached, so as to obtain a trained 3D object detection model.

[0277] Optionally, the 3D object detection model further includes a multi-layer perceptron; the apparatus further includes:

[0278] A second correction amount acquisition module, configured to, before the second point cloud feature fusion module 703 executes the step of, for each reference 3D point, inputting the current reference feature of the reference 3D point into the deformable attention network to obtain a preset number of position offsets and the weight corresponding to each position offset, input the sample initial point cloud features of each BEV grid and the sample initial image feature into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix; wherein, the initial transformation matrix is calibrated according to the acquisition methods of the sample point cloud and the sample image; use the obtained correction amount to correct the initial transformation matrix; the second image feature fusion module 704 is specifically configured to: use the corrected transformation matrix to transform the current coordinates of the reference 3D point to obtain the pixel coordinates corresponding to the reference 3D point in the sample image as the sample reference coordinates; fuse the eigenvalue corresponding to the sample reference coordinates in each sample initial image feature according to the weight of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point.

[0279] Optionally, the second correction amount acquisition module is specifically configured to input the sample initial point cloud features of each BEV grid, the sample initial image feature, and the current coordinates of each reference 3D point into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix.

[0280] Optionally, the device further includes: a second interpolation module, configured to, before the second image feature fusion module 704 performs fusing, according to the weights of the initial image features of each sample, the feature values corresponding to the sample reference coordinates in the initial image features of each sample to obtain the current sample intermediate image features of the reference three-dimensional point, if the sample reference coordinates are not integers, obtain other pixel coordinates within the neighborhood range of the sample reference coordinates; for each initial image feature of each sample, perform interpolation on the feature values corresponding to the obtained other pixel coordinates in the initial image feature of each sample as the feature values corresponding to the sample reference coordinates in the initial image feature of each sample.

[0281] Optionally, the image feature extraction network includes a plurality of downsampling layers connected in series; there is one sample image, and the plurality of initial image features of the sample include: the image features of the sample image extracted by the plurality of downsampling layers; or, the sample images include images of the sample space region collected at a plurality of different shooting angles, and the plurality of initial image features of the sample include the image features of each sample image.

[0282] Optionally, the second feature extraction module 702 is specifically configured to, for each BEV grid obtained by division, if the number of three-dimensional points in the sample point cloud included in the BEV grid is greater than a specified number, sample the specified number of three-dimensional points from the three-dimensional points included in the BEV grid as the sample three-dimensional points included in the BEV grid; if the number of three-dimensional points in the sample point cloud included in the BEV grid is equal to the specified number, use the three-dimensional points included in the BEV grid as the sample three-dimensional points included in the BEV grid; if the number of three-dimensional points in the sample point cloud included in the BEV grid is less than the specified number, calculate the difference between the specified number and the number of three-dimensional points in the sample point cloud included in the BEV grid, and generate the difference number of three-dimensional points with preset coordinates; combine the generated three-dimensional points and the three-dimensional points in the sample point cloud included in the BEV grid to obtain the sample three-dimensional points included in the BEV grid.

[0283] Optionally, the information to be utilized of a three-dimensional point further includes at least one of the following: the coordinate offset of the three-dimensional point relative to the centroid of the BEV grid to which the three-dimensional point belongs, and the feature data at the true position corresponding to the detected three-dimensional point.

[0284] Optionally, the sample label further includes the category of the object within the sample space region; the position indication of the object within the sample space region: the size of the minimum bounding cube of the object, the coordinates of the center point of the minimum bounding cube of the object, and the orientation of the object; the detection result further includes the category of the object within the sample space region; the differences between the sample label and the detection result include: the position difference of the object within the sample space region, and the category difference of the object within the sample space region.

[0285] Based on the 3D object detection model training device provided in this application, the training device can obtain a trained 3D object detection model, and then send the trained 3D object detection model to the electronic device, so that after the electronic device deploys the received 3D object detection model, it can perform object detection based on the deployed 3D object detection model. When the electronic device performs object detection based on the end-to-end 3D object detection model, it combines the point cloud and the image to jointly perform object detection, improving the accuracy of the obtained object detection result, reducing the maintenance cost, and improving the versatility of the object detection method.

[0286] An embodiment of this application further provides an electronic device, as Figure 8 shown, including: a memory 801 for storing a computer program; a processor 802 for implementing the steps of any of the object detection methods in the above embodiments, or implementing the steps of any of the 3D object detection model training methods in the above embodiments when executing the program stored on the memory 801.

[0287] And the above electronic device may further include a communication bus and / or a communication interface, and the processor 802, the communication interface, and the memory 801 complete communication with each other through the communication bus.

[0288] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0289] The communication interface is used for communication between the above electronic device and other devices.

[0290] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0291] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0292] In another embodiment provided by the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of any of the above target detection methods, or implements the steps of any of the above three-dimensional target detection model training methods.

[0293] In another embodiment provided by the present application, a computer program product containing instructions is further provided. When it runs on a computer, it causes the computer to execute any of the target detection methods in the above embodiments, or execute any of the three-dimensional target detection model training methods in the above.

[0294] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.

[0295] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0296] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the apparatus, electronic device, computer-readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0297] The foregoing are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.

Claims

1. A target detection method, characterized in that: The method comprises: Obtaining a point cloud and an image of the space area to be detected as the point cloud to be detected and the image to be detected respectively; Dividing the point cloud to be detected according to the bird's-eye view BEV grid, extracting the initial point cloud features to be detected of each BEV grid, and performing feature extraction on the image to be detected to obtain the initial image features to be detected; For each reference 3D point, the current reference feature of the reference 3D point is input into the deformable attention network in the pre-trained 3D target detection model to obtain a preset number of sampling position offsets and a weight corresponding to each position offset; for each position offset, the current coordinates of the reference 3D point are offset by using the position offset, and the BEV grid to which the offset coordinates belong is determined as the BEV grid corresponding to the position offset; the initial point cloud features to be detected of the BEV grids corresponding to each position offset are fused according to the weights corresponding to each position offset to obtain the current fused point cloud features to be detected of the reference 3D point; Acquire a portion of the reference three-dimensional point corresponding to the initial image feature to be detected, and obtain a current intermediate image feature to be detected of the reference three-dimensional point; The current fused point cloud features to be detected and the intermediate image features to be detected of each reference 3D point are spliced ​​to obtain the current spliced ​​features to be detected; Decode the current splicing feature to be detected to obtain the current decoding result, and calculate the sum of the current coordinates of the reference 3D point and the current coordinate adjustment amount to obtain the current coordinates of the reference 3D point; calculate the sum of the current reference feature of the reference 3D point and the current reference feature adjustment amount to obtain the current reference feature of the reference 3D point; return to execute the steps of sampling the BEV grid within the neighborhood of the reference 3D point based on the current coordinates and reference features of the reference 3D point for each reference 3D point, and fusing the initial point cloud features to be detected of the sampled BEV grid to obtain the current fused point cloud features to be detected of the reference 3D point, until the number of decoding times reaches a preset number; wherein the current decoding result includes: the current coordinate adjustment amount of the reference 3D point, and the current reference feature adjustment amount of the reference 3D point; Based on the current coordinates of each reference three-dimensional point, the position of the object in the spatial area to be detected is obtained.

2. The method according to claim 1, characterized in that The step of dividing the point cloud to be detected according to the bird's-eye view BEV grid, extracting the initial point cloud features to be detected of each BEV grid, and performing feature extraction on the image to be detected to obtain the initial image features to be detected includes: The point cloud to be detected is divided according to the bird's-eye view BEV grid, and the three-dimensional points to be detected contained in each BEV grid obtained by the division are determined; the feature extraction of the information to be used of the three-dimensional points to be detected contained in each BEV grid is performed using the point cloud feature extraction network in the pre-trained three-dimensional target detection model to obtain the initial point cloud features to be detected of each BEV grid; and the image to be detected is input into the image feature extraction network in the three-dimensional target detection model to obtain a plurality of initial image features to be detected; The information to be used of a three-dimensional point includes: the coordinates of the three-dimensional point and the coordinates of the centroid of the BEV grid to which the three-dimensional point belongs; the three-dimensional object detection model is obtained by training based on a sample image of a sample space region, a sample point cloud and a sample label representing the position of an object in the sample space region; the three-dimensional object detection model also includes a weighted attention network and a decoder; The obtaining of the portion of the reference 3D point corresponding to the initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference 3D point includes: inputting the current reference feature of the reference 3D point into the weighted attention network to obtain the weight of each initial image feature to be detected; fusing the feature values ​​corresponding to the reference 3D point in each initial image feature to be detected according to the weight of each initial image feature to be detected to obtain the current intermediate image feature to be detected of the reference 3D point; The decoding of the current splicing feature to be detected to obtain the current decoding result includes: inputting the current splicing feature to be detected into the decoder to obtain the current decoding result.

3. The method according to claim 2, characterized in that The three-dimensional target detection model also includes a multi-layer perceptron; Before inputting the current reference feature of each reference 3D point into a deformable attention network in a pre-trained 3D object detection model to obtain a preset number of sampling position offsets and a weight corresponding to each position offset, the method further includes: Input the initial point cloud features to be detected of each BEV grid and the initial image features to be detected into the multi-layer perceptron to obtain a correction amount for correcting the initial conversion matrix; wherein the initial conversion matrix is ​​obtained by calibration according to the acquisition method of the point cloud to be detected and the image to be detected; Correcting the initial transformation matrix using the obtained correction amount; According to the weights of the initial image features to be detected, the feature values ​​corresponding to the reference three-dimensional point in the initial image features to be detected are fused to obtain the current intermediate image feature to be detected of the reference three-dimensional point, including: The corrected transformation matrix is ​​used to transform the current coordinates of the reference three-dimensional point to obtain the pixel coordinates corresponding to the reference three-dimensional point in the image to be detected as the reference coordinates to be detected; According to the weight of each initial image feature to be detected, the feature values ​​corresponding to the reference coordinate to be detected in each initial image feature to be detected are fused to obtain the current intermediate image feature to be detected of the reference three-dimensional point.

4. The method according to claim 3, characterized in that The initial point cloud features to be detected of each BEV grid and the initial image features to be detected are input into the multi-layer perceptron to obtain a correction amount for correcting the initial conversion matrix, including: The initial point cloud features to be detected of each BEV grid, the initial image features to be detected, and the current coordinates of each reference three-dimensional point are input into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix.

5. The method according to claim 3, characterized in that: Before fusing the feature values ​​corresponding to the reference coordinate to be detected in each of the initial image features to be detected according to the weights of each of the initial image features to be detected to obtain the current intermediate image feature to be detected of the reference three-dimensional point, the method further includes: If the reference coordinate to be detected is not an integer, obtaining coordinates of other pixels within the neighborhood of the reference coordinate to be detected; For each initial image feature to be detected, feature values ​​corresponding to other acquired pixel coordinates in the initial image feature to be detected are interpolated as feature values ​​corresponding to the reference coordinates to be detected in the initial image feature to be detected.

6. The method according to claim 2, characterized in that The image feature extraction network includes a plurality of downsampling layers connected in series; the image to be detected is one, and the plurality of initial image features to be detected include: image features of the image to be detected extracted by the plurality of downsampling layers; or, The image to be detected includes images of the spatial area to be detected collected at multiple different shooting angles, and the multiple initial image features to be detected include image features of each image to be detected.

7. The method according to claim 2, characterized in that The three-dimensional points to be detected contained in each BEV grid obtained by the determination of the division include: For each BEV grid obtained by division, if the number of three-dimensional points in the point cloud to be detected contained in the BEV grid is greater than a specified number, sampling the specified number of three-dimensional points from the three-dimensional points contained in the BEV grid as the three-dimensional points to be detected contained in the BEV grid; If the number of the three-dimensional points in the to-be-detected point cloud contained in the BEV grid is equal to the specified number, the three-dimensional points contained in the BEV grid are used as the to-be-detected three-dimensional points contained in the BEV grid; If the number of three-dimensional points in the point cloud to be detected contained in the BEV grid is less than the specified number, the difference between the specified number and the number of three-dimensional points in the point cloud to be detected contained in the BEV grid is calculated to generate three-dimensional points whose coordinates are preset values ​​of the difference; the generated three-dimensional points are combined with the three-dimensional points in the point cloud to be detected contained in the BEV grid to obtain the three-dimensional points to be detected contained in the BEV grid.

8. A three-dimensional object detection model training method, characterized in that: The method comprises: Acquire a point cloud and an image of a sample space region as a sample point cloud and a sample image, respectively, and acquire a sample label including a position of an object in the sample space region; Divide the sample point cloud according to the bird's-eye view BEV grid, and determine the sample three-dimensional points contained in each BEV grid obtained by the division; use the point cloud feature extraction network in the three-dimensional target detection model of the initial structure including a deformable attention network, a weighted attention network, a decoder and an image feature extraction network to extract features from the coordinates of the sample three-dimensional points contained in each BEV grid and the coordinates of the centroid of the BEV grid to which the sample three-dimensional points belong, and obtain the sample initial point cloud features of each BEV grid; input the sample image into the image feature extraction network to obtain multiple sample initial image features; Input the current reference feature of each reference 3D point into the deformable attention network to obtain a preset number of sampling position offsets and a weight corresponding to each position offset; use each position offset to offset the current coordinates of the reference 3D point, and determine that the BEV grid to which the offset coordinates belong is the BEV grid corresponding to the position offset; fuse the sample initial point cloud features of the BEV grid corresponding to each position offset according to the weight corresponding to each position offset, and obtain the current sample fused point cloud features of the reference 3D point; Inputting the current reference feature of the reference 3D point into the weighted attention network to obtain the weight of each sample initial image feature; fusing the feature values ​​corresponding to the reference 3D point in each sample initial image feature according to the weight of each sample initial image feature to obtain the current sample intermediate image feature of the reference 3D point; The current sample fusion point cloud features and the sample intermediate image features of each reference 3D point are spliced ​​to obtain the current sample splicing features; Input the current sample splicing feature into the decoder to obtain the current position offset and the current reference feature offset of the reference 3D point, respectively calculate the sum of the current coordinates of the reference 3D point and the current position offset, and the sum of the current reference feature of the reference 3D point and the current reference feature offset, to obtain the current coordinates and the current reference feature of the reference 3D point; return to execute the step of inputting the current reference feature of each reference 3D point into the deformable attention network to obtain a preset number of sampling position offsets and a weight corresponding to each position offset, until the number of decoding times reaches the preset number; Obtaining a detection result including the position of the object in the sample space region based on the current coordinates of each reference three-dimensional point; Based on the difference between the sample label and the detection result, the model parameters of the three-dimensional object detection model of the initial structure are adjusted until a preset convergence condition is reached to obtain a trained three-dimensional object detection model.

9. The method according to claim 8, characterized in that The three-dimensional target detection model also includes a multi-layer perceptron; Before inputting the current reference feature of each reference 3D point into the deformable attention network to obtain a preset number of sampling position offsets and a weight corresponding to each position offset, the method further includes: Inputting the sample initial point cloud features of each BEV grid and the sample initial image features into the multi-layer perceptron to obtain a correction amount for correcting the initial conversion matrix; wherein the initial conversion matrix is ​​obtained by calibration according to the acquisition method of the sample point cloud and the sample image; Correcting the initial transformation matrix using the obtained correction amount; The step of fusing the feature values ​​corresponding to the reference three-dimensional point in each sample initial image feature according to the weight of each sample initial image feature to obtain the current sample intermediate image feature of the reference three-dimensional point includes: The corrected transformation matrix is ​​used to transform the current coordinates of the reference three-dimensional point to obtain the pixel coordinates corresponding to the reference three-dimensional point in the sample image as the sample reference coordinates; According to the weight of each sample initial image feature, the feature values ​​corresponding to the sample reference coordinates in each sample initial image feature are fused to obtain the current sample intermediate image feature of the reference three-dimensional point.

10. The method according to claim 9, characterized in that The sample initial point cloud features of each BEV grid and the sample initial image features are input into the multi-layer perceptron to obtain a correction amount for correcting the initial conversion matrix, including: The sample initial point cloud features of each BEV grid, the sample initial image features, and the current coordinates of each reference three-dimensional point are input into the multi-layer perceptron to obtain a correction amount for correcting the initial transformation matrix.

11. The method according to claim 9, characterized in that Before fusing the feature values ​​corresponding to the sample reference coordinates in the sample initial image features according to the weights of the sample initial image features to obtain the current sample intermediate image feature of the reference three-dimensional point, the method further includes: If the sample reference coordinate is not an integer, obtaining coordinates of other pixels within the neighborhood of the sample reference coordinate; For each sample initial image feature, feature values ​​corresponding to other acquired pixel coordinates in the sample initial image feature are interpolated as feature values ​​corresponding to the sample reference coordinates in the sample initial image feature.

12. The method according to claim 8, characterized in that The image feature extraction network includes a plurality of downsampling layers connected in series; the sample image is one, and the plurality of sample initial image features include: image features of the sample image extracted by the plurality of downsampling layers; or, The sample image includes images of the sample space area collected at multiple different shooting angles, and the multiple sample initial image features include image features of each sample image.

13. The method according to claim 8, characterized in that The sample three-dimensional points contained in each BEV grid obtained by the determination of the division include: For each BEV grid obtained by division, if the number of three-dimensional points in the sample point cloud contained in the BEV grid is greater than a specified number, sampling the specified number of three-dimensional points from the three-dimensional points contained in the BEV grid as the sample three-dimensional points contained in the BEV grid; If the number of three-dimensional points in the sample point cloud contained in the BEV grid is equal to the specified number, the three-dimensional points contained in the BEV grid are used as the sample three-dimensional points contained in the BEV grid; If the number of three-dimensional points in the sample point cloud contained in the BEV grid is less than the specified number, the difference between the specified number and the number of three-dimensional points in the sample point cloud contained in the BEV grid is calculated to generate three-dimensional points whose coordinates are preset values. The generated three-dimensional points are combined with the three-dimensional points in the sample point cloud contained in the BEV grid to obtain the sample three-dimensional points contained in the BEV grid.

14. A target detection device, characterized in that: The device comprises: A first acquisition module is used to acquire a point cloud and an image of a spatial area to be detected as a point cloud to be detected and an image to be detected, respectively; A first feature extraction module is used to divide the point cloud to be detected according to the bird's-eye view BEV grid, extract the initial point cloud features to be detected of each BEV grid, and perform feature extraction on the image to be detected to obtain the initial image features to be detected; The first point cloud feature fusion module is used to input the current reference feature of each reference 3D point into a deformable attention network in a pre-trained 3D target detection model to obtain a preset number of sampling position offsets and a weight corresponding to each position offset; for each position offset, use the position offset to offset the current coordinates of the reference 3D point, and determine the BEV grid to which the offset coordinates belong as the BEV grid corresponding to the position offset; fuse the initial point cloud features to be detected of the BEV grids corresponding to each position offset according to the weights corresponding to each position offset, to obtain the current fused point cloud features to be detected of the reference 3D point; A first image feature fusion module is used to obtain a portion of the reference three-dimensional point corresponding to the initial image feature to be detected, and obtain a current intermediate image feature to be detected of the reference three-dimensional point; The first stitching module is used to stitch the current fused point cloud features to be detected and the intermediate image features to be detected of each reference three-dimensional point to obtain the current stitching features to be detected; The first updating module is used to decode the current splicing feature to be detected to obtain the current decoding result, and calculate the sum of the current coordinates of the reference 3D point and the current coordinate adjustment amount to obtain the current coordinates of the reference 3D point; calculate the sum of the current reference feature of the reference 3D point and the current reference feature adjustment amount to obtain the current reference feature of the reference 3D point; return to execute the steps of sampling the BEV grid within the neighborhood of the reference 3D point based on the current coordinates and reference features of the reference 3D point for each reference 3D point, and fusing the initial point cloud features to be detected of the sampled BEV grid to obtain the current fused point cloud features to be detected of the reference 3D point, until the number of decoding times reaches a preset number; wherein the current decoding result includes: the current coordinate adjustment amount of the reference 3D point, and the current reference feature adjustment amount of the reference 3D point; The position acquisition module is used to obtain the position of the object in the spatial area to be detected based on the current coordinates of each reference three-dimensional point.

15. A three-dimensional object detection model training device, characterized in that: The device comprises: A second acquisition module is used to acquire a point cloud and an image of a sample space region as a sample point cloud and a sample image respectively, and acquire a sample label including a position of an object in the sample space region; The second feature extraction module is used to divide the sample point cloud according to the bird's-eye view BEV grid, determine the sample three-dimensional points contained in each BEV grid obtained by the division; use the point cloud feature extraction network in the three-dimensional target detection model of the initial structure including the deformable attention network, the weighted attention network, the decoder and the image feature extraction network to extract features of the coordinates of the sample three-dimensional points contained in each BEV grid and the coordinates of the centroid of the BEV grid to which the sample three-dimensional points belong, and obtain the sample initial point cloud features of each BEV grid; input the sample image into the image feature extraction network to obtain multiple sample initial image features; The second point cloud feature fusion module is used to input the current reference feature of each reference 3D point into the deformable attention network to obtain a preset number of sampling position offsets and a weight corresponding to each position offset; use each position offset to offset the current coordinates of the reference 3D point, and determine that the BEV grid to which the offset coordinates belong is the BEV grid corresponding to the position offset; according to the weight corresponding to each position offset, the sample initial point cloud features of the BEV grid corresponding to each position offset are fused to obtain the current sample fused point cloud features of the reference 3D point; The second image feature fusion module is used to input the current reference feature of the reference 3D point into the weighted attention network to obtain the weight of each sample initial image feature; according to the weight of each sample initial image feature, the corresponding feature value of the reference 3D point in each sample initial image feature is fused to obtain the current sample intermediate image feature of the reference 3D point; The second stitching module is used to stitch the current sample fusion point cloud features and the sample intermediate image features of each reference three-dimensional point to obtain the current sample stitching features; The second updating module is used to input the current sample splicing feature into the decoder to obtain the current position offset and the current reference feature offset of the reference 3D point, respectively calculate the sum of the current coordinates of the reference 3D point and the current position offset, and the sum of the current reference feature of the reference 3D point and the current reference feature offset, to obtain the current coordinates and the current reference feature of the reference 3D point; return to execute the step of inputting the current reference feature of each reference 3D point into the deformable attention network to obtain a preset number of sampling position offsets and a weight corresponding to each position offset, until the number of decoding times reaches the preset number; A detection result acquisition module, used to obtain a detection result including the position of the object in the sample space area based on the current coordinates of each reference three-dimensional point; The training module is used to adjust the model parameters of the three-dimensional object detection model of the initial structure based on the difference between the sample label and the detection result, until a preset convergence condition is reached to obtain a trained three-dimensional object detection model.

Citation Information

Patent Citations

  • Complex road target detection method based on multi-modal fusion aerial view

    CN117058646A

  • Target detection method based on image and point cloud fusion, vehicle and storage medium

    CN118115844A