A target detection method, device, equipment and storage medium
By dividing point clouds into sub-point clouds and weighting the fusion feature, the point dislocation problem in the prior art is solved, and high-accurate object detection is achieved in motion scenarios.
Patent Information
- Application Number
- CN202210577183.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-25
AI Technical Summary
The existing object detection scheme has point misalignment problems when integrating the 2-dimensional features of the image and the 3-dimensional features of the point cloud, resulting in poor target detection results in motion scenarios.
By dividing the point cloud into multiple sub-point clouds, determining the attention coefficient based on the radar characteristics and image characteristics of each sub-point cloud, performing feature weighting, and finally projecting the fusion feature to the bird's eye view for target detection.
It effectively reduces the performance losses caused by misalignment of local projection points in motion scenarios, and improves the accuracy and effectiveness of target detection.
Smart Images

Figure CN114998610B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular, to an object detection method, apparatus, device, and storage medium. Background Art
[0002] With the gradual entry of the fields of autonomous driving and intelligent transportation into the public eye, accurate object detection is particularly important. In current object detection solutions, generally, based on the images collected by a camera and the point cloud collected by a lidar, the 2D features of the image are fused with the 3D features of the point cloud to achieve object detection.
[0003] Since the 2D features of the image and the 3D features of the point cloud come from different sensors and are not from the same source, it is not appropriate to directly fuse the two. Additionally, when fusing features, it is necessary to project the point cloud onto the image. In a moving scenario, there will be a deviation between the projected point and the actual pixel point, resulting in point misalignment and poor object detection effects.
[0004] In summary, the current object detection solutions have certain defects and cannot accurately and effectively perform object detection work. Summary of the Invention
[0005] The present application provides an object detection method, apparatus, device, and storage medium, which can accurately and effectively perform object detection work.
[0006] In a first aspect, the present application provides an object detection method, which includes: obtaining the radar features and image features of each point in the point cloud corresponding to the radar signal; the image feature of a point in the point cloud is the image feature of the point corresponding to the point in the two-dimensional image; dividing the point cloud into N sub-point clouds; N is greater than or equal to 1; determining the attention coefficient of each sub-point cloud according to the radar features and image features of each point in the sub-point cloud; the attention coefficient is used to indicate the mutual attention of the radar information between the points in the sub-point cloud and the mutual attention of the image information; respectively determining the fusion feature of each sub-point cloud; the fusion feature of a sub-point cloud is obtained by weighting the radar features and image features of the sub-point cloud by the attention coefficient of the sub-point cloud; projecting the fusion features of the N sub-point clouds onto a bird's-eye view for object detection.
[0007] In one possible implementation, for the first sub-point cloud, determining the attention coefficient of the first sub-point cloud includes: using a multi-layer perceptron to map the radar features and image features of each point in the first sub-point cloud to obtain a vector of radar information and a vector of image information of the first sub-point cloud; determining a first coefficient and a second coefficient respectively according to the vector of radar information and the vector of image information as the attention coefficient of the first sub-point cloud; the first coefficient is the mutual attention matrix of the radar information among each point in the first sub-point cloud; the second coefficient is the mutual attention matrix of the image information among each point in the first sub-point cloud.
[0008] In another possible implementation, determining a first coefficient and a second coefficient respectively according to the vector of radar information and the vector of image information includes: multiplying the query key value in the vector of radar information by the dictionary key value in the vector of radar information as the first coefficient; multiplying the query key value in the vector of image information by the dictionary key value in the vector of image information as the second coefficient.
[0009] In yet another possible implementation, the radar feature of the first sub-point cloud is the eigenvalue in the vector of radar information of the first sub-point cloud, and the image feature of the first sub-point cloud is the eigenvalue in the vector of image information of the first sub-point cloud; determining the fusion feature of the first sub-point cloud includes: using the first coefficient and the second coefficient to weight the radar features in the vector of radar information to obtain the radar fusion feature of the first sub-point cloud; using the first coefficient and the second coefficient to weight the image features in the vector of image information to obtain the image fusion feature of the first sub-point cloud; connecting the radar fusion feature and the image fusion feature of the first sub-point cloud to obtain the fusion feature of the first sub-point cloud.
[0010] In yet another possible implementation, dividing the point cloud into N sub-point clouds includes: using the farthest point sampling algorithm and the k-nearest neighbor algorithm to divide the point cloud into N sub-point clouds.
[0011] In yet another possible implementation, projecting the fusion features of the N sub-point clouds onto a bird's-eye view for object detection includes: using a deep learning network to extract the features of the bird's-eye view to generate 3D detection boxes; using a deep learning network to predict the category and scale of the 3D detection boxes.
[0012] In yet another possible implementation, obtaining the radar features and image features of each point in the point cloud corresponding to the radar signal includes: performing feature extraction on the radar signal to obtain the radar features of each point in the point cloud; performing feature extraction on the image signal to obtain the features of the image signal; processing the features of the image signal through the intrinsic matrix and the extrinsic matrix to obtain the image features of each point in the point cloud.
[0013] In another possible implementation, the radar signal is a distance view signal, and the method further includes: obtaining the features of the distance view signal; obtaining the radar features of each point in the point cloud corresponding to the radar signal, including: corresponding the features of the distance view signal to the point cloud to obtain the radar features of each point in the point cloud corresponding to the radar signal.
[0014] The target detection method provided by the embodiments of the present application divides the point cloud into multiple sub-point clouds. For each sub-point cloud, according to the radar features and image features of each point in the sub-point cloud, the attention coefficient of the sub-point cloud is determined. Further, according to the attention coefficient, the fusion feature of the sub-point cloud is determined to be projected onto the bird's-eye view for target detection. By dividing the sub-point clouds and weighting the features of the sub-point clouds using the attention coefficients of the sub-point clouds, the features of each point in the local area are aggregated from the features of the surrounding points according to the attention coefficients, increasing the weight relationship between each point and the surrounding points in the local area, effectively compensating for the performance loss caused by the misalignment of local projection points in the moving scene, and performing the target detection work more accurately and effectively. Moreover, compared with the traditional method of directly fusing radar features and image features, this solution determines the attention coefficients through radar features and image features, and uses the attention coefficients to weight the radar features and image features respectively, correlating the radar features and image features to a certain extent. Therefore, the method of fusing the weighted radar features and image features in this solution is smoother, with better fusion effect and superior performance.
[0015] In a second aspect, the present application provides a target detection device, which includes: an acquisition module, a division module, a determination module, and a detection module; the acquisition module is used to acquire the radar features and image features of each point in the point cloud corresponding to the radar signal; the image feature of a point in the point cloud is the image feature of the point corresponding to the point in the two-dimensional image; the division module is used to divide the point cloud into N sub-point clouds; N is greater than or equal to 1; the determination module is used to respectively determine the fusion feature of each sub-point cloud according to the radar features and image features of each point in the sub-point cloud; the fusion feature is used to reflect the radar information and image information of the sub-point cloud; the detection module is used to project the fusion features of the N sub-point clouds onto the bird's-eye view for target detection.
[0016] In a possible implementation, the determination module is specifically used to respectively determine the attention coefficient of each sub-point cloud according to the radar features and image features of each point in the sub-point cloud; the attention coefficient is used to indicate the mutual attention of the radar information and the mutual attention of the image information among the points in the sub-point cloud; respectively determine the fusion feature of each sub-point cloud; the fusion feature of a sub-point cloud is obtained by weighting the radar features and image features of the sub-point cloud by the attention coefficient of the sub-point cloud.
[0017] In another possible implementation manner, for the first sub-point cloud, the determination module is specifically configured to use a multi-layer perceptron to map the radar features and image features of each point in the first sub-point cloud to obtain a vector of radar information and a vector of image information of the first sub-point cloud; according to the vector of radar information and the vector of image information, determine a first coefficient and a second coefficient respectively as the attention coefficients of the first sub-point cloud; the first coefficient is the mutual attention matrix of the radar information among each point in the first sub-point cloud; the second coefficient is the mutual attention matrix of the image information among each point in the first sub-point cloud.
[0018] In yet another possible implementation manner, the determination module is specifically configured to multiply the query key value in the vector of radar information by the dictionary key value in the vector of radar information as the first coefficient; multiply the query key value in the vector of image information by the dictionary key value in the vector of image information as the second coefficient.
[0019] In yet another possible implementation manner, the radar feature of the first sub-point cloud is the eigenvalue in the vector of radar information of the first sub-point cloud, and the image feature of the first sub-point cloud is the eigenvalue in the vector of image information of the first sub-point cloud; the determination module is specifically configured to use the first coefficient and the second coefficient to weight the radar features in the vector of radar information to obtain the radar fusion feature of the first sub-point cloud; use the first coefficient and the second coefficient to weight the image features in the vector of image information to obtain the image fusion feature of the first sub-point cloud; connect the radar fusion feature and the image fusion feature of the first sub-point cloud to obtain the fusion feature of the first sub-point cloud.
[0020] In yet another possible implementation manner, the partitioning module is specifically configured to use the farthest point sampling algorithm and the k-nearest neighbor algorithm to partition the point cloud into N sub-point clouds.
[0021] In yet another possible implementation manner, the detection module is specifically configured to use a deep learning network to extract the features of the bird's-eye view image and generate a 3D detection box; use a deep learning network to predict the category and scale of the 3D detection box.
[0022] In yet another possible implementation manner, the acquisition module is specifically configured to extract the features of the radar signal to obtain the radar features of each point in the point cloud; extract the features of the image signal to obtain the features of the image signal; process the features of the image signal through the intrinsic matrix and the extrinsic matrix to obtain the image features of each point in the point cloud.
[0023] In yet another possible implementation manner, the radar signal is a distance view signal, and the acquisition module is further configured to acquire the features of the distance view signal; the acquisition module is specifically configured to correspond the features of the distance view signal to the point cloud to obtain the radar features of each point in the point cloud corresponding to the radar signal.
[0024] In a third aspect, the present application provides a server, which includes: a processor and a memory; the memory stores instructions executable by the processor; when the processor is configured to execute the instructions, the server implements the method of the first aspect described above.
[0025] In a fourth aspect, the present application provides a computer-readable storage medium, which includes: computer software instructions; when the computer software instructions run on an electronic device, the electronic device implements the method of the first aspect described above.
[0026] In a fifth aspect, the present application provides a computer program product, which when run on a computer, causes the computer to execute the steps of the related method described in the first aspect to implement the method of the first aspect.
[0027] The beneficial effects of the second to fifth aspects described above can refer to the corresponding descriptions of the first aspect and will not be elaborated here. Description of the Drawings
[0028] Figure 1 It is a schematic diagram of the application environment of a target detection method provided by the present application;
[0029] Figure 2 It is a schematic diagram of the process of a target detection method provided by the present application;
[0030] Figure 3 It is a schematic diagram of the process of a method for obtaining features of each point provided by the present application;
[0031] Figure 4 It is a schematic diagram of the structure of a fully convolutional network with an encoder-decoder structure provided by the present application;
[0032] Figure 5 It is a schematic diagram of the process of determining attention coefficients provided by the present application;
[0033] Figure 6 It is a schematic diagram of the process of determining fused features provided by the present application;
[0034] Figure 7 It is a schematic diagram of the process of a target detection process provided by the present application;
[0035] Figure 8 It is a schematic diagram of the network structure for generating detection boxes provided by the present application;
[0036] Figure 9 It is a schematic diagram of the process of an LV fusion method based on self-attention provided by the present application;
[0037] Figure 10 It is a schematic diagram of the composition of a target detection device provided by the present application;
[0038] Figure 11 A structural schematic diagram of a server provided for this application. Specific implementation manners
[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0040] It should be noted that in the embodiments of the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.
[0041] To facilitate the clear description of the technical solutions in the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the words such as "first" and "second" do not limit the quantity and execution order. As described in the background art, perception is the basis and core of the autonomous driving system. 2D perception is difficult to support high-level autonomous driving, and 3D perception needs to be further adopted. Currently, there are several mainstream 3D object detection methods based on multi-sensor fusion: a 3D detector for multi-modal fusion, a 3D detector for cross-view spatial feature fusion, a 3D detector for enhancing point cloud features using image semantic information, etc. In the above lidar vision (LV) fusion framework, the high-dimensional features of the images output by the 2D network are mainly connected with the high-dimensional features of the laser point cloud output by the 3D network, so as to realize 3D object detection. However, this method has the problem that the direct fusion of features from different sources is inappropriate, and there is a point misalignment in the motion scene, which affects the accuracy of the object detection work.
[0042] Based on this, the embodiments of the present application provide an object detection method, which can divide the point cloud into multiple sub-point clouds, and use the sub-attention coefficients of each sub-point cloud to correct the radar features and image features of the sub-point cloud, so as to accurately and effectively perform the object detection work.
[0043] The object detection method provided by the present application can be applied to an application environment as Figure 1 shown. As Figure 1As shown in the figure, the application environment may include: a target detection device 101, a camera 102, and a lidar 103. The target detection device 101 is respectively connected to the camera 102 and the lidar 103.
[0044] Among them, the target detection device 101 can be applied to a server. Here, the server mentioned can be a server cluster composed of multiple servers, or a single server, or a computer. Specifically, the target detection device 101 can be a processor or a processing chip in the server, etc. The embodiments of the present application do not limit the specific device form of the above server. Figure 1 Taking the target detection device 101 applied to a single server as an example is shown in the figure.
[0045] The above camera 102 is mainly used to collect red green blue (RGB) images required for LV fusion. The embodiments of the present application do not limit the specific device form of the camera 102, Figure 1 Taking the camera 102 as a camera as an example is shown in the figure. The above lidar 103 is mainly used to collect point clouds required for LV fusion. The embodiments of the present application do not limit the specific device form of the lidar 103.
[0046] Figure 2 It is a schematic flowchart of a target detection method provided by an embodiment of the present application. Exemplarily, the target detection method provided by the present application can be applied to Figure 1 the application environment shown in the figure.
[0047] As Figure 2 shown in the figure, the target detection method provided by the present application can specifically include the following steps:
[0048] S201. The target detection device obtains the radar feature and the image feature of each point in the point cloud corresponding to the radar signal.
[0049] Among them, the image feature of a point in the point cloud is the image feature of the point corresponding to the point in the two-dimensional image.
[0050] In some embodiments, the target detection device can receive the point cloud collected by the lidar and the two-dimensional image collected by the camera. Then, it obtains the radar feature of each point in the point cloud and the image feature of the point corresponding to each point in the point cloud in the two-dimensional image.
[0051] Specifically, the steps of obtaining the radar feature and the image feature of each point can be as Figure 3 shown in the figure, including the following steps S201a - S201c.
[0052] S201a. The target detection device performs feature extraction on the radar signal to obtain the radar feature of each point in the point cloud.
[0053] In some embodiments, the target detection device may extract features from radar signals to obtain the radar features of each point in the point cloud. Specifically, the target detection device may use different 3D backbone networks to extract different radar signals. Exemplarily, the following describes two mainstream radar signals, the point cloud signal and the range view signal, respectively.
[0054] The point cloud signal is an irregular signal composed of n points, and each point has 4 channels, namely the {x, y, z} position information and the reflectivity information r. Therefore, the point cloud signal can be expressed as p l ∈R n×4 , where R represents a matrix. The target detection device may use a deep learning network to extract features from the point cloud signal, and the type and internal structure of the deep learning network are not limited in this application.
[0055] Considering the disorder and irregularity of the point cloud, exemplarily, the embodiments of this application use a multi-level point cloud feature extraction network (PointNet++) to extract the radar features of the point cloud signal. PointNet++ mainly consists of a downsampling module and an upsampling module. Among them, the downsampling module consists of a multi-layer perceptron and a downsampling layer, and the upsampling module consists of a multi-layer perceptron and an upsampling layer. The point cloud signal can be processed by the downsampling module and the upsampling module to obtain the features of the point cloud, such as expressed as f l ∈R n×128 , and then the radar features of each point in the point cloud can be obtained, with a dimension of 128.
[0056] In the case where the radar signal is a range view signal, the target detection device obtains the features of the range view signal. Further, the features of the range view signal are corresponded to the point cloud to obtain the radar features of each point in the point cloud corresponding to the radar signal. The range view signal x d usually has a regular shape, that is, the width w and the height h. There are 4 channels in the range view signal, namely the {x, y, z} position information and the reflectivity information r. Therefore, the range view signal can be expressed as x d ∈R w×h×4 .
[0057] Among them, the target detection device may also use a deep learning network to extract features from the range view signal, and the type and internal structure of the deep learning network are not limited herein.
[0058] Exemplarily, the embodiments of this application may use a fully convolutional network with an encoder-decoder structure to extract features from the range view signal. The structure of the fully convolutional network with an encoder-decoder structure is as Figure 4As shown in the figure. The encoder consists of a convolutional layer and a pooling layer. The convolutional layer performs convolution on the input, and the pooling layer performs pooling on the output of the convolutional layer to achieve downsampling of features. Different from the encoder, the decoder consists of a convolutional layer and an upsampling layer, and the upsampling layer performs upsampling on the output of the convolutional layer. After the distance view signal is processed by the encoder and the decoder, a feature map with the same size as the input distance view (i.e., the feature of the above-mentioned distance view signal) is output, which can be denoted as f d ∈R w×h×128 . Further, the features of the distance view are processed through the corresponding formula with the point cloud to obtain the radar features of each point in the point cloud.
[0059] S201b. The target detection device extracts features from the image signal to obtain the features of the image signal;
[0060] In some embodiments, the target detection device can extract features from the image signal collected by the camera to obtain the features of the image signal.
[0061] Specifically, the target detection device can pre-train an image detection network using an image detection task, and then use this network to extract image features. Exemplarily, in the embodiments of the present application, a one-stage 2D detector (Yolo v3) of a deep convolutional neural network (Darknet-53) is used to process the image signal to obtain a feature map that is 32 times downsampled from the original image (i.e., the feature of the above-mentioned image signal).
[0062] S201c. The target detection device processes the features of the image signal through the intrinsic matrix and the extrinsic matrix to obtain the image features of each point in the point cloud.
[0063] In some embodiments, after obtaining the features of the image signal, the target detection device can process the features of the image signal through the intrinsic matrix of the camera and the extrinsic matrix between the camera and the lidar to achieve the correspondence between the point cloud and the pixel points in the image, so as to obtain the image features of each point in the point cloud.
[0064] In order to optimize the situation of point misalignment when projecting the point cloud onto the image in a moving scene, in this embodiment, the point cloud is divided into multiple local regions, and then each local region is corrected to solve the local projection error caused by moving objects. That is, the target detection device executes the following S202 - S205.
[0065] S202. The target detection device divides the point cloud into N sub-point clouds.
[0066] Where N is greater than or equal to 1.
[0067] In some exemplary embodiments, the target detection device may divide the point cloud into N sub-point clouds. Specifically, the target detection device may use the farthest point sampling algorithm and the k-nearest neighbor algorithm to divide the point cloud into N sub-point clouds, with a maximum of K points in each sub-point cloud. For the specific usage of the algorithm, please refer to the relevant technical documents and will not be elaborated here. It should be noted that the specific size of N can be adjusted according to the actual scenario. The finer the division granularity (i.e., the larger the value of N), the better the optimization effect on point misalignment.
[0068] S203. The target detection device determines the attention coefficient of each sub-point cloud according to the radar feature and image feature of each point in the sub-point cloud.
[0069] Among them, the attention coefficient is used to indicate the mutual attention of radar information and the mutual attention of image information among the points in the sub-point cloud. The radar information is the data information corresponding to the radar signal, and the image information is the data information corresponding to the image signal.
[0070] In the embodiments of the present application, the attention coefficient of each sub-point cloud is used to correct the radar feature and image feature of the sub-point cloud to optimize the problem of point misalignment. Therefore, in some embodiments, after dividing the point cloud into multiple sub-point clouds, for each sub-point cloud, the target detection device may determine the attention coefficient of each sub-point cloud according to the radar feature and image feature of each point in the sub-point cloud.
[0071] Specifically, taking the first sub-point cloud among the following N sub-point clouds as an example, the specific process of determining the attention coefficient of the sub-point cloud is described. This determination process is as Figure 5 shown, including the following S203a - S203b.
[0072] S203a. The target detection device uses a multi-layer perceptron to map the radar feature and image feature of each point in the first sub-point cloud to obtain the vector of radar information and the vector of image information of the first sub-point cloud.
[0073] In some embodiments, the target detection device may use a multi-layer perceptron to map the radar feature of each point in the first sub-point cloud to obtain the vector of radar information of the first sub-point cloud. Similarly, the target detection device may use a multi-layer perceptron to map the image feature of each point in the first sub-point cloud to obtain the vector of image information of the first sub-point cloud.
[0074] Specifically, the above vector of radar information can be expressed as {Q L , K L , V L}, and the above vector of image information can be expressed as {Q C , K C , V C}. Q, K, V ∈ R K×128They are three feature vectors, which are obtained by processing through perceptrons at different levels. Among them, Q is called the query key, K is called the dictionary key, and V is called the feature value. A multi-layer perceptron is a feedforward artificial neural network model that maps multiple input data sets to a single output data set. Among them, a neural network is a machine learning model, which is a machine learning technology that simulates the neural network of the human brain to achieve artificial intelligence similar to humans. The input and output of the neural network can be configured according to actual needs, and the neural network is trained with sample data to minimize the error between its output and the true output corresponding to the sample data. The specific implementation of the multi-layer perceptron in the embodiments of the present application will not be elaborated in detail.
[0075] S203b. The target detection device determines a first coefficient and a second coefficient respectively according to the vector of the radar information and the vector of the image information, and uses them as the attention coefficients of the first sub-point cloud.
[0076] Among them, the first coefficient is the mutual attention matrix of the radar information among each point in the first sub-point cloud. The second coefficient is the mutual attention matrix of the image information among each point in the first sub-point cloud.
[0077] In some embodiments, the target detection device can determine the mutual attention matrix of the radar information among each point in the first sub-point cloud as the first coefficient according to the vector of the radar information. Similarly, the target detection device can determine the mutual attention matrix of the image information among each point in the first sub-point cloud as the second coefficient according to the vector of the image information. The first coefficient and the second coefficient are used as the attention coefficients of the first sub-point cloud.
[0078] Specifically, according to the vector of the radar information and the vector of the image information, the first coefficient and the second coefficient are determined respectively by using the following formula:
[0079] A mod =Q mod ·K mod , mod ∈ {L, C}
[0080] Among them, Q is the above-mentioned query key, and K is the above-mentioned dictionary key.
[0081] Multiply the query key in the vector of the radar information by the dictionary key in the vector of the radar information as the first coefficient. That is, multiply Q L by K L to perform matrix multiplication to obtain the mutual attention matrix A L as the first coefficient.
[0082] Multiply the query key in the vector of the image information by the dictionary key in the vector of the image information as the second coefficient. That is, multiply Q C by K CPerform a matrix multiplication to obtain the mutual attention matrix A C As the second coefficient.
[0083] S204. The target detection device respectively determines the fusion features of each sub-point cloud.
[0084] Among them, the fusion features are used to reflect the radar information and image information of the sub-point cloud. The fusion features of a sub-point cloud are obtained by weighting the radar features and image features of the sub-point cloud with the attention coefficients of the sub-point cloud.
[0085] In some embodiments, after determining the attention coefficients of the sub-point cloud, the target detection device can respectively determine the fusion features of each sub-point cloud. Taking the first sub-point cloud as an example, the specific steps for determining the fusion features are as Figure 6 shown, including the following S204a - S204c. Among them, the radar features of the first sub-point cloud are the eigenvalues in the vector of the radar information of the first sub-point cloud. The image features of the first sub-point cloud are the eigenvalues in the vector of the image information of the first sub-point cloud.
[0086] S204a. The target detection device uses the first coefficient and the second coefficient to weight the eigenvalues in the vector of the radar information to obtain the radar fusion features of the first sub-point cloud.
[0087] The target detection device uses the first coefficient A L and the second coefficient A C , to weight the eigenvalues V L in the vector of the radar information to obtain the radar fusion features F L of the first sub-point cloud. The specific weighting formula is as follows:
[0088]
[0089] Among them, softmax is a normalization function in the field of deep learning, and d k is the dimension size (128 in the embodiments of the present application).
[0090] S204b. The target detection device uses the first coefficient and the second coefficient to weight the eigenvalues in the vector of the image information to obtain the image fusion features of the first sub-point cloud.
[0091] The target detection device uses the first coefficient A L and the second coefficient A C , to weight the eigenvalues V C in the vector of the image information to obtain the radar fusion features F C of the first sub-point cloud. The specific weighting formula is as follows:
[0092]
[0093] Among them, softmax and d k have the same meaning as above and will not be elaborated here.
[0094] S204c. The target detection device connects the radar fusion feature and the image fusion feature of the first sub-point cloud to obtain the fusion feature of the first sub-point cloud.
[0095] In some embodiments, after determining the radar fusion feature F L and the image fusion feature F C of the first sub-point cloud, the target detection device can connect the radar fusion feature F L and the image fusion feature F C of the first sub-point cloud to perform vector connection to obtain the final fusion feature of the first sub-point cloud.
[0096] It can be understood that in the embodiments of the present application, the attention coefficients of the sub-point clouds are used to weight the radar features and image features of the sub-point clouds. The features of each point in the local area are aggregated from the features of the surrounding points according to the attention coefficients, increasing the correlation between each point and the surrounding points in the local area. In this way, the performance loss caused by the misalignment of the projection points can be effectively reduced, which affects the work of target detection.
[0097] S205. The target detection device projects the fusion features of N sub-point clouds onto the bird's-eye view for target detection.
[0098] In some embodiments, after determining the fusion feature of each sub-point cloud, the target detection device can project the fusion features of N sub-point clouds onto the bird's-eye view in sequence according to the divided positions, thereby realizing target detection.
[0099] Specifically, the process of target detection is as Figure 7 shown, specifically including the following S205a - S205b.
[0100] S205a. The target detection device uses a deep learning network to extract the features of the bird's-eye view and generate a 3D detection box.
[0101] In some embodiments, after projecting the fusion features of N sub-point clouds onto the bird's-eye view, the target detection device can use a deep learning network to extract the features of the bird's-eye view and generate a 3D detection box.
[0102] Exemplarily, the target detection device may use a convolutional neural network in a deep learning network to extract features of the bird's-eye view. Among them, the convolutional neural network is a deep neural network with a convolutional structure. The convolutional neural network includes a feature extractor composed of convolutional layers and subsampling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as convolving a trainable filter with an input image or a convolutional feature plane (feature map). The convolutional layer refers to the neuron layer in the convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of the convolutional neural network, a neuron can only be connected to some neighboring layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weight here is the convolutional kernel. Sharing weights can be understood as a way of extracting image information that is independent of position. The underlying principle here is that the statistical information of a certain part of the image is the same as that of other parts. That is to say, the image information learned in one part can also be used in another part. Therefore, for all positions on the image, the same learned image information can be used. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels, the richer the image information reflected by the convolution operation.
[0103] The convolutional kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolutional kernel can learn to obtain reasonable weights. Additionally, the direct benefit brought by sharing weights is to reduce the connections between the layers of the convolutional neural network and at the same time reduce the risk of overfitting.
[0104] For example, as Figure 8 shown in a schematic diagram of a network structure for generating detection boxes. The convolutional neural network sequentially performs 3 3×3 convolutions (conv) on the input bird's-eye view features (denoted as (C, W, H) in the figure), and the stride of each convolution is 2, obtaining three different scales of features, denoted as (C, W / 2, H / 2), (2C, W / 4, H / 4), and (4C, W / 8, H / 8) respectively. Further, perform deconvolution (deconv) on these three scales of features to restore them to the same scale, then concatenate the three features (concat) to obtain the concatenated features and generate 3D detection boxes.
[0105] S205b. The target detection device uses a deep learning network to predict the category and scale of the 3D detection box.
[0106] In some embodiments, after generating the 3D detection box, the target detection device may use a deep learning network to predict the category and scale of the 3D detection box to achieve the work of target detection.
[0107] Exemplarily, the target detection device may use a convolutional neural network in a deep learning network to predict the category and scale of the 3D detection box. Specifically, the target detection device may use 1×1 convolution to extract the features of the 3D detection box to achieve the prediction of the category and scale of the detection box.
[0108] It can be understood that the embodiments of the present application provide a target detection method. After obtaining the bird's-eye view, a convolutional neural network can be used to generate a 3D detection box in the bird's-eye view. The 3D detection box can enclose the targets (such as pedestrians, vehicles, trees, etc.) in the bird's-eye view. Further, a 1×1 convolutional neural network is used to predict the category and scale of the targets enclosed by the 3D detection box. While realizing the detection of the category of the target, the three-dimensional information such as the spatial position, size, and orientation of the target is estimated.
[0109] The technical solutions provided in the above embodiments at least bring the following beneficial effects. The target detection method provided in the embodiments of the present application divides the point cloud into multiple sub-point clouds. For each sub-point cloud, according to the radar features and image features of each point in the sub-point cloud, the attention coefficient of the sub-point cloud is determined. Further, the fusion feature of the sub-point cloud is determined according to the attention coefficient and projected onto the bird's-eye view for target detection. This solution divides the sub-point cloud and weights the features of the sub-point cloud using the attention coefficient of the sub-point cloud, so that the features of each point in the local area are aggregated from the features of the surrounding points according to the attention coefficient, increasing the weight relationship between each point and the surrounding points in the local area, effectively compensating for the performance loss caused by the misalignment of local projection points in the motion scene, and performing the target detection work more accurately and effectively. Moreover, compared with the traditional method of directly fusing radar features and image features, this solution determines the attention coefficient through radar features and image features, and uses the attention coefficient to weight the radar features and image features respectively, correlating the radar features and image features to a certain extent. Therefore, the method of fusing and weighting the radar features and image features in this solution is relatively smoother, the fusion effect is better, and the performance is superior.
[0110] Further, the target detection method provided in this solution takes the point cloud as the intermediate perspective, re-weights the point cloud using image features (such as image texture features, color features, etc.), adds 2D image features to the 3D position and size features, and realizes 3D target detection to meet the growing perception requirements of high-level autonomous driving.
[0111] Figure 9A flowchart of an LV fusion method based on self-attention provided by an embodiment of the present application. LiDAR signals such as point clouds and distance views are input into the 3D backbone in the feature extraction module to extract the point-by-point features of the LiDAR information. The RGB image is input into the 2D backbone in the feature extraction module to extract the point-by-point features of the image information (equivalent to S201 above). Further, the point-by-point features of the LiDAR information and the point-by-point features of the image information are input into the fusion module (or called the Transformer fusion module) for point-by-point fusion. During the fusion process, the fusion module extracts the Q, K, and V feature vectors of the LiDAR information and the Q, K, and V feature vectors of the image information respectively to obtain the attention matrix of the LiDAR information and the attention matrix of the image information. The V features of the LiDAR information and the V features of the image information are weighted and spliced across modalities using the attention matrix to obtain the fusion vector of each local fusion subset (equivalent to the sub-point cloud above) (equivalent to S202 - S204 above). Then, it is projected onto the bird's-eye view and the 3D box prediction network is used for target prediction (equivalent to S205 above).
[0112] In an exemplary embodiment, the present application also provides a target detection device. The target detection device may include one or more functional modules for implementing the target detection method in the above method embodiments.
[0113] For example, Figure 10 A schematic diagram of the composition of a target detection device provided by an embodiment of the present application. As Figure 7 shown, the target detection device includes: an acquisition module 1001, a division module 1002, a determination module 1003, and a detection module 1004. The acquisition module 1001, the division module 1002, the determination module 1003, and the detection module 1004 are connected to each other.
[0114] The acquisition module 1001 is used to acquire the LiDAR features and image features of each point in the point cloud corresponding to the LiDAR signal, and the image feature of a point in the point cloud is the image feature of the point corresponding to it in the two-dimensional image.
[0115] The division module 1002 is used to divide the point cloud into N sub-point clouds. N is greater than or equal to 1.
[0116] The determination module 1003 is used to respectively determine the fusion feature of each sub-point cloud according to the LiDAR features and image features of each point in the sub-point cloud; the fusion feature is used to reflect the LiDAR information and image information of the sub-point cloud.
[0117] The detection module 1004 is used to project the fusion features of the N sub-point clouds onto the bird's-eye view for target detection.
[0118] In some embodiments, the determination module 1003 is specifically configured to determine the attention coefficient of each sub-point cloud according to the radar features and image features of each point in the sub-point cloud; the attention coefficient is used to indicate the mutual attention of radar information and the mutual attention of image information among the points in the sub-point cloud; determine the fusion feature of each sub-point cloud respectively; the fusion feature of a sub-point cloud is obtained by weighting the radar features and image features of the sub-point cloud with the attention coefficient of the sub-point cloud.
[0119] In some embodiments, for the first sub-point cloud, the determination module 1003 is specifically configured to map the radar features and image features of each point in the first sub-point cloud by using a multi-layer perceptron to obtain the vector of radar information and the vector of image information of the first sub-point cloud; determine the first coefficient and the second coefficient respectively according to the vector of radar information and the vector of image information as the attention coefficient of the first sub-point cloud; the first coefficient is the mutual attention matrix of radar information among each point in the first sub-point cloud; the second coefficient is the mutual attention matrix of image information among each point in the first sub-point cloud.
[0120] In some embodiments, the determination module 1003 is specifically configured to multiply the query key value in the vector of radar information by the dictionary key value in the vector of radar information as the first coefficient; multiply the query key value in the vector of image information by the dictionary key value in the vector of image information as the second coefficient.
[0121] In some embodiments, the radar feature of the first sub-point cloud is the eigenvalue in the vector of radar information of the first sub-point cloud, and the image feature of the first sub-point cloud is the eigenvalue in the vector of image information of the first sub-point cloud.
[0122] The determination module 1003 is specifically configured to weight the radar features in the vector of radar information by using the first coefficient and the second coefficient to obtain the radar fusion feature of the first sub-point cloud; weight the image features in the vector of image information by using the first coefficient and the second coefficient to obtain the image fusion feature of the first sub-point cloud; connect the radar fusion feature and the image fusion feature of the first sub-point cloud to obtain the fusion feature of the first sub-point cloud.
[0123] In some embodiments, the partitioning module 1002 is specifically configured to partition the point cloud into N sub-point clouds by using the farthest point sampling algorithm and the k-nearest neighbor algorithm.
[0124] In some embodiments, the detection module 1004 is specifically configured to extract the features of the bird's-eye view by using a deep learning network to generate a 3D detection box; predict the category and scale of the 3D detection box by using a deep learning network.
[0125] In some embodiments, the acquisition module 1001 is specifically configured to extract features from the radar signal to obtain the radar features of each point in the point cloud; extract features from the image signal to obtain the features of the image signal; and process the features of the image signal through the intrinsic matrix and the extrinsic matrix to obtain the image features of each point in the point cloud.
[0126] In some embodiments, the radar signal is a distance view signal, and the acquisition module 1001 is further configured to obtain the features of the distance view signal.
[0127] The acquisition module 1001 is specifically configured to correspond the features of the distance view signal to the point cloud to obtain the radar features of each point in the point cloud corresponding to the radar signal.
[0128] When the functions of the above integrated modules are implemented in the form of hardware, an exemplary structural diagram of a server is provided in an embodiment of the present application. The server may be the target detection device in the above embodiment. As Figure 11 shown, the server 1100 includes: a processor 1102, a communication interface 1103, and a bus 1104. Optionally, the server may further include a memory 1101.
[0129] The processor 1102 may be configured to implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 1102 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 1102 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0130] The communication interface 1103 is used to connect to other devices through a communication network. The communication network may be an Ethernet, a wireless access network, a wireless local area network (WLAN), etc.
[0131] The memory 1101 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0132] As a possible implementation, the memory 1101 can exist independently of the processor 1102. The memory 1101 can be connected to the processor 1102 through the bus 11011 and is used to store instructions or program code. When the processor 1102 calls and executes the instructions or program code stored in the memory 1101, the object detection method provided by the embodiments of the present application can be implemented.
[0133] In another possible implementation, the memory 1101 can also be integrated with the processor 1102.
[0134] The bus 1104 can be an extended industry standard architecture (EISA) bus, etc. The bus 1104 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the object detection device is divided into different functional modules to complete all or part of the functions described above.
[0136] The embodiments of the present application also provide a computer-readable storage medium. All or part of the processes in the above method embodiments can be completed by computer instructions instructing relevant hardware. This program can be stored in the above computer-readable storage medium. When this program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be the memory of any of the foregoing embodiments. The above computer-readable storage medium can also be an external storage device of the above object detection device, such as a plug-in hard disk equipped on the above object detection device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the above computer-readable storage medium can also include both the internal storage unit of the above object detection device and the external storage device. The above computer-readable storage medium is used to store the above computer program and other programs and data required by the above object detection device. The above computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0137] The embodiments of the present application also provide a computer program product. This computer product includes a computer program. When this computer program product runs on a computer, it enables the computer to execute any one of the object detection methods provided in the above embodiments.
[0138] Although the present application has been described in combination with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the drawings, the disclosed content, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0139] Although the present application has been described in combination with specific features and their embodiments, obviously, various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are only exemplary descriptions of the present application defined by the appended claims, and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these changes and modifications.
[0140] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A target detection method, characterized in that, the method includes: Obtain the radar features and image features of each point in the point cloud corresponding to the radar signal; the image feature of a point in the point cloud is the image feature of the point corresponding to it in the two-dimensional image; Divide the point cloud into N sub-point clouds; N is greater than or equal to 1; According to the radar features and image features of each point in the sub-point cloud, determine the attention coefficient of each sub-point cloud respectively; the attention coefficient is used to indicate the mutual attention of radar information and the mutual attention of image information among the points in the sub-point cloud; Determine the fusion features of each of the sub-point clouds respectively; the fusion feature of a sub-point cloud is obtained by weighting the radar features and image features of the sub-point cloud with the attention coefficient of the sub-point cloud; Project the fusion features of the N sub-point clouds onto a bird's-eye view for target detection.
2. The method according to claim 1, characterized in that, For the first sub-point cloud, determining the attention coefficient of the first sub-point cloud includes: Using a multi-layer perceptron to map the radar features and image features of each point in the first sub-point cloud to obtain the vector of radar information and the vector of image information of the first sub-point cloud; According to the vector of radar information and the vector of image information, determine a first coefficient and a second coefficient respectively as the attention coefficient of the first sub-point cloud; the first coefficient is the mutual attention matrix of radar information among each point in the first sub-point cloud; the second coefficient is the mutual attention matrix of image information among each point in the first sub-point cloud.
3. The method according to claim 2, characterized in that, The determining the first coefficient and the second coefficient respectively according to the vector of radar information and the vector of image information includes: Multiply the query key value in the vector of radar information by the dictionary key value in the vector of radar information as the first coefficient; Multiply the query key value in the vector of image information by the dictionary key value in the vector of image information as the second coefficient.
4. The method according to claim 3, characterized in that, The radar feature of the first sub-point cloud is the eigenvalue in the vector of radar information of the first sub-point cloud, and the image feature of the first sub-point cloud is the eigenvalue in the vector of image information of the first sub-point cloud; Determining the fusion feature of the first sub-point cloud includes: Using the first coefficient and the second coefficient to weight the eigenvalues in the vector of radar information to obtain the radar fusion feature of the first sub-point cloud; Using the first coefficient and the second coefficient to weight the eigenvalues in the vector of image information to obtain the image fusion feature of the first sub-point cloud; Connect the radar fusion feature and the image fusion feature of the first sub-point cloud to obtain the fusion feature of the first sub-point cloud.
5. The method according to any one of claims 1-4, characterized in that, The dividing the point cloud into N sub-point clouds includes: Using the farthest point sampling algorithm and the k-nearest neighbor algorithm to divide the point cloud into N sub-point clouds.
6. The method according to any one of claims 1-4, characterized in that, the projecting the fusion features of the N sub-point clouds onto a bird's-eye view for object detection includes: using a deep learning network to extract features of the bird's-eye view to generate 3D detection boxes; using a deep learning network to predict the class and scale of the 3D detection boxes.
7. The method according to any one of claims 1-4, characterized in that, the obtaining the radar features and image features of each point in the point cloud corresponding to the radar signal includes: performing feature extraction on the radar signal to obtain the radar features of each point in the point cloud; performing feature extraction on the image signal to obtain the features of the image signal; processing the features of the image signal through an intrinsic matrix and an extrinsic matrix to obtain the image features of each point in the point cloud.
8. The method according to claim 7, characterized in that, the radar signal is a distance view signal, and the method further includes: obtaining the features of the distance view signal; the obtaining the radar features of each point in the point cloud corresponding to the radar signal includes: corresponding the features of the distance view signal to the point cloud to obtain the radar features of each point in the point cloud corresponding to the radar signal.
9. An object detection device, characterized in that, the device includes: an acquisition module, a division module, a determination module, and a detection module; the acquisition module is configured to acquire the radar features and image features of each point in the point cloud corresponding to the radar signal; the image feature of a point in the point cloud is the image feature of the point corresponding to the point in the two-dimensional image; the division module is configured to divide the point cloud into N sub-point clouds; N is greater than or equal to 1; the determination module is configured to respectively determine the attention coefficients of each sub-point cloud according to the radar features and image features of each point in the sub-point cloud; the attention coefficient is used to indicate the mutual attention of the radar information and the mutual attention of the image information among the points in the sub-point cloud; the determination module is further configured to respectively determine the fusion features of each of the sub-point clouds; the fusion feature of a sub-point cloud is obtained by weighting the radar features and image features of the sub-point cloud by the attention coefficient of the sub-point cloud; the detection module is configured to project the fusion features of the N sub-point clouds onto a bird's-eye view for object detection.
10. The device according to claim 9, characterized in that, for the first sub-point cloud, the determination module is specifically configured to use a multi-layer perceptron to map the radar features and image features of each point in the first sub-point cloud to obtain the vector of the radar information and the vector of the image information of the first sub-point cloud; according to the vector of the radar information and the vector of the image information, respectively determine a first coefficient and a second coefficient as the attention coefficient of the first sub-point cloud; the first coefficient is the mutual attention matrix of the radar information among the points in the first sub-point cloud; the second coefficient is the mutual attention matrix of the image information among the points in the first sub-point cloud; The determining module is specifically configured to multiply the query key value in the vector of the radar information by the dictionary key value in the vector of the radar information as the first coefficient; multiply the query key value in the vector of the image information by the dictionary key value in the vector of the image information as the second coefficient; The radar feature of the first sub-point cloud is the eigenvalue in the vector of the radar information of the first sub-point cloud, and the image feature of the first sub-point cloud is the eigenvalue in the vector of the image information of the first sub-point cloud; the determining module is specifically configured to use the first coefficient and the second coefficient to weight the eigenvalue in the vector of the radar information to obtain the radar fusion feature of the first sub-point cloud; use the first coefficient and the second coefficient to weight the eigenvalue in the vector of the image information to obtain the image fusion feature of the first sub-point cloud; Connect the radar fusion feature of the first sub-point cloud and the image fusion feature of the first sub-point cloud to obtain the fusion feature of the first sub-point cloud; The partitioning module is specifically configured to partition the point cloud into N sub-point clouds by using the farthest point sampling algorithm and the k-nearest neighbor algorithm; The detection module is specifically configured to use a deep learning network to extract the features of the bird's-eye view and generate a 3D detection box; use a deep learning network to predict the category and scale of the 3D detection box; The obtaining module is specifically configured to extract the features of the radar signal to obtain the radar feature of each point in the point cloud; extract the features of the image signal to obtain the features of the image signal; process the features of the image signal through the intrinsic matrix and the extrinsic matrix to obtain the image feature of each point in the point cloud; The radar signal is a distance view signal; the obtaining module is further configured to obtain the features of the distance view signal; The obtaining module is specifically configured to correspond the features of the distance view signal to the point cloud to obtain the radar feature of each point in the point cloud corresponding to the radar signal.
11. A server, characterized in that, the server includes: a processor and a memory; the memory stores instructions executable by the processor; when the processor is configured to execute the instructions, the server implements the method according to any one of claims 1-8.
12. A computer-readable storage medium, characterized in that, the computer-readable storage medium includes: computer software instructions; when the computer software instructions run in an electronic device, the electronic device implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Three-dimensional target detection method and device based on multi-sensor information fusion
CN110929692A
Target detection method, device, equipment and system
CN111856445A