Three-dimensional target detection method, system and electronic equipment

By performing depth mapping and feature fusion of multiple camera images and point cloud data, the problem of depth information gap and computational complexity in laser point cloud and camera data fusion is solved, and high-precision and efficient three-dimensional object detection is achieved.

CN116805410BActive Publication Date: 2025-06-06INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310775040.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-28
Publication Date
2025-06-06
Estimated Expiration
2043-06-28

AI Technical Summary

Technical Problem

When the prior art uses laser point cloud and camera multimodal data for three-dimensional object detection, it is difficult to effectively solve the depth information gap and computational complexity between modes, resulting in low detection accuracy and efficiency.

Method used

By deeply mapping the collected multiple camera images and point cloud data to generate images to be processed, the trained image feature model and point cloud feature model are used to extract and fuse three-dimensional features respectively, and the three-dimensional object detection results are output.

Benefits of technology

It realizes efficient detection of three-dimensional targets by multimodal data based on image and point cloud data, which significantly improves detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116805410B_ABST
    Figure CN116805410B_ABST
Patent Text Reader

Abstract

The present application provides a three-dimensional target detection method, system and electronic device, including depth mapping of multiple collected camera images and point cloud data to generate an image to be processed; according to the image to be processed and the trained image feature model, the three-dimensional features of the image to be processed in the image coordinate system are determined and the three-dimensional features in the image coordinate system are converted into three-dimensional features in the reference coordinate system to determine the first bird's-eye view feature; according to the image to be processed and the trained point cloud feature model, the point cloud features of the image to be processed are extracted and multi-scale fused to determine the second bird's-eye view feature; the first bird's-eye view feature and the second bird's-eye view feature are fused and the three-dimensional target detection result is determined based on the fused first bird's-eye view feature and the second bird's-eye view feature. Through a more effective model, the detection of three-dimensional targets based on multi-modal data of image and point cloud data is realized, and the detection accuracy is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, in particular to three-dimensional target detection methods, systems and electronic equipment. Background Art

[0002] The rapid development of artificial intelligence technology has brought about a leap in autonomous driving technology. The autonomous driving system is like the human eye, which provides more than 70% of the information to the brain. The autonomous driving system includes three parts: perception, planning, and control. The perception part provides a large amount of information input for the autonomous driving car. The perception system perceives the surrounding environment through various sensors such as cameras and radars. It not only needs to accurately identify vehicles, pedestrians, obstacles, traffic signs, etc. in the surrounding environment, but also needs to accurately locate and predict their speed. Therefore, the quality of an autonomous driving system often depends on the quality of its perception system. The fusion of LiDAR and camera can provide crucial information input for the realization of highly robust and high-precision three-dimensional target detection, and has always been a key research direction in the industry. On the one hand, because LiDAR can effectively capture spatial information, point cloud data has a natural 3D advantage, which maximizes the ranging accuracy, speed and direction of target detection. On the other hand, the camera has rich texture information, powerful semantics and image context understanding capabilities to ensure the effective recognition of concrete road information such as pedestrians and traffic signs.

[0003] In the existing 3D target detection technology based on laser point cloud and camera multimodal data fusion, one part adopts the latest Transformer (a neural network architecture) fusion architecture, searches on image features through point cloud features, obtains fusion features, and finally predicts targets such as vehicles, pedestrians and bicycles; the other part realizes the merging of image features and point cloud features through simple feature splicing, and then connects to a conventional 3D detection head for target prediction. Although these two methods can generally complete 3D target detection, due to the differences and complexity of laser point cloud and camera multimodal information, it is still a huge challenge to combine the two completely different modal geometric and semantic features in one representation space: on the one hand, since point cloud radar is much sparser than camera data, it is difficult for existing model fusion methods to solve the depth information gap between inherent modes, and there are still great challenges in accuracy; on the other hand, in cross-modal fusion interaction, point cloud radar involves fine division of voxels and a large number of 3D convolution calculations, which are complex and time-consuming; and images bring considerable computational pressure due to multiple cameras, high resolution, and complex feature extraction networks.

[0004] Therefore, there is an urgent need for a three-dimensional target detection method that improves the accuracy of three-dimensional target detection to solve the above technical problems. Summary of the invention

[0005] Based on this, it is necessary to provide a three-dimensional target detection method to address the above technical issues in order to achieve three-dimensional detection of the target.

[0006] In a first aspect, the present application provides a three-dimensional object detection method, the method comprising:

[0007] Depth mapping is performed on the collected multiple camera images and point cloud data to generate an image to be processed;

[0008] According to the image to be processed and the trained image feature model, determine the three-dimensional features of the image to be processed in the image coordinate system and convert the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system to output a first bird's-eye view feature, so as to determine the first bird's-eye view feature;

[0009] According to the image to be processed and the trained point cloud feature model, point cloud features of the image to be processed are extracted and multi-scale fusion is performed to output a second bird's-eye view feature to determine the second bird's-eye view feature;

[0010] The first bird's-eye view feature and the second bird's-eye view feature are fused, and a three-dimensional target detection result is determined based on the fused first bird's-eye view feature and the second bird's-eye view feature.

[0011] In some embodiments, the image feature model includes:

[0012] Receive the input image to be processed, and extract image features of the image to be processed based on a deep neural network;

[0013] Performing enhancement processing on the image features based on a multi-layer neural network to generate an enhanced image feature map;

[0014] Performing depth prediction based on the image features of the depth prediction module to generate a corresponding depth map;

[0015] Determine the three-dimensional features in the image coordinate system based on the enhanced image feature map and the corresponding depth map;

[0016] The three-dimensional features in the image coordinate system are converted into three-dimensional features in a reference coordinate system for output.

[0017] In some embodiments, the point cloud feature model includes:

[0018] Receiving the input image to be processed, and extracting point cloud features of the point cloud data pasted in the image to be processed;

[0019] Based on a preset point cloud feature extraction network, the point cloud features are multi-scale fused and the multi-scale fused point cloud features are output.

[0020] In some embodiments, the step of performing multi-scale fusion on the point cloud features based on a preset point cloud feature extraction network and outputting the multi-scale fused point cloud features includes:

[0021] Use multi-layer 3D convolutional networks to convolve point cloud features in sequence and record the intermediate features output by each layer of 3D convolutional network;

[0022] All the intermediate features are combined to determine the multi-scale fused point cloud features.

[0023] In some embodiments, the first bird's-eye view feature is a three-dimensional feature in a reference coordinate system output by an image feature model;

[0024] The second bird's-eye view feature is a multi-scale fused point cloud feature output by the point cloud feature model.

[0025] In some embodiments, the method further comprises:

[0026] After the image to be processed is input into the point cloud feature model, the point cloud feature model projects the point cloud data in the image to be processed onto the camera image to generate target depth information and outputs it to the image feature model;

[0027] The image feature model receives the target depth information output by the point cloud feature model, and performs depth correction on the generated depth map;

[0028] The image feature model determines the three-dimensional features in the image coordinate system based on the depth map after depth correction and the enhanced image feature map.

[0029] In some embodiments, the training process of the image feature model and the point cloud feature model includes:

[0030] Extract the point cloud data and image data of each sample in the training data set to build a sample database;

[0031] Randomly selecting candidate samples from the sample database to determine a candidate sample set;

[0032] Get the existing sample set at the current moment;

[0033] Performing depth mapping processing on samples in the existing sample set and the spare sample set according to the three-dimensional true value frame corresponding to each sample in the existing sample set and the spare sample set;

[0034] Paste the point cloud data into the corresponding depth map processed sample to generate an enhanced dataset;

[0035] The image feature model and the point cloud feature model are trained using the enhanced data set.

[0036] In some embodiments, the training process of the image feature model and the point cloud feature model further includes:

[0037] BEV encoding the first bird's eye view feature to determine a first detection feature;

[0038] decoding the first detection feature;

[0039] determining a prediction loss and a classification loss in response to the decoded first detection feature;

[0040] Determine a target loss function based on the prediction loss and the classification loss;

[0041] The image feature model and the point cloud feature model are modified based on the target loss function.

[0042] In a second aspect, the present application provides a three-dimensional object detection system, the system comprising:

[0043] A preprocessing module, used for performing depth mapping on the collected multiple camera images and point cloud data to generate an image to be processed;

[0044] A model prediction module, used to determine the three-dimensional features of the image to be processed in the image coordinate system according to the image to be processed and the trained image feature model, and convert the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system to determine the first bird's-eye view features;

[0045] The model prediction module is further used to extract the point cloud features of the image to be processed and perform multi-scale fusion according to the image to be processed and the trained point cloud feature model to determine the second bird's-eye view features;

[0046] The data processing module is used to fuse the first bird's-eye view feature and the second bird's-eye view feature and determine a three-dimensional target detection result based on the fused first bird's-eye view feature and the second bird's-eye view feature.

[0047] In a third aspect, the present application provides an electronic device, the electronic device comprising:

[0048] one or more processors;

[0049] and a memory associated with the one or more processors, the memory being used to store program instructions, wherein when the program instructions are read and executed by the one or more processors, the program instructions perform the following operations:

[0050] The beneficial effects achieved by this application are:

[0051] The present application provides a three-dimensional target detection method, including performing depth mapping on multiple collected camera images and point cloud data to generate an image to be processed; determining the three-dimensional features of the image to be processed in the image coordinate system according to the image to be processed and the trained image feature model and converting the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system to determine the first bird's-eye view feature; extracting the point cloud features of the image to be processed and performing multi-scale fusion according to the image to be processed and the trained point cloud feature model to determine the second bird's-eye view feature; fusing the first bird's-eye view feature and the second bird's-eye view feature and determining the three-dimensional target detection result based on the fused first bird's-eye view feature and the second bird's-eye view feature. Through the newly constructed model, the detection of three-dimensional targets based on multi-modal data of image and point cloud data is realized, which greatly improves the detection accuracy.

[0052] Furthermore, the present application also proposes to introduce point cloud data to generate a depth map in the image space when performing depth prediction, and then use the generated depth map to correct the depth map predicted by the depth estimation module to further improve the accuracy.

[0053] Furthermore, the present application also proposes to fuse point cloud features through multi-scale fusion technology to further enhance the feature expression capability of point clouds.

[0054] Furthermore, the present application also discloses a method for enhancing sample data during the model training process, by generating a data sample library by combining the target's laser point cloud information and image information in advance. During the training process, samples are randomly selected from the data sample library, and sample screening and mapping are performed according to their three-dimensional spatial coordinate information to complete data enhancement.

[0055] Furthermore, the present application also discloses an auxiliary supervision method during the model training process to continuously correct the model and further enhance the advantages of image texture features. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work, among which:

[0057] Figure 1 is a diagram of a three-dimensional target detection architecture provided in an embodiment of the present application;

[0058] Figure 2 is a schematic diagram of a data enhancement method provided in an embodiment of the present application;

[0059] Figure 3is a schematic diagram of a depth prediction method provided in an embodiment of the present application;

[0060] Figure 4 is a schematic diagram of a multi-scale fusion method provided in an embodiment of the present application;

[0061] Figure 5 is a flow chart of a three-dimensional target detection method provided in an embodiment of the present application;

[0062] Figure 6 is a schematic diagram of a three-dimensional target detection system provided in an embodiment of the present application;

[0063] Figure 7 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0065] It should be understood that in the description of the present application, unless the context clearly requires otherwise, words such as "include", "comprises", and the like in the entire specification and claims should be interpreted as inclusive rather than exclusive or exhaustive; that is, the meaning of "including but not limited to".

[0066] It should also be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.

[0067] It should be noted that the terms "S1", "S2", etc. are only used for the purpose of describing the steps, and do not specifically refer to the order or sequence, nor are they used to limit the present application. They are only for the convenience of describing the method of the present application, and cannot be understood as indicating the order of the steps. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0068] As described in the background technology, the perception module in the automatic driving system outputs the position, length, width, height and speed of the detection target in three-dimensional space by reading the laser point cloud information and the new machine image information, so as to facilitate the subsequent decision-making and planning of the driving vehicle. Although the existing three-dimensional target detection method based on the fusion of multimodal data such as laser point cloud and camera can realize the detection of three-dimensional targets, it is difficult to combine the two completely different modal geometry and semantic features in one representation space due to the differences and complexity of the multimodal information of laser point cloud and camera. On the one hand, since the point cloud radar is much sparser than the camera data, the existing model fusion method is difficult to solve the depth information gap between the inherent modes, and there are still great challenges in accuracy. On the other hand, in the cross-modal fusion interaction, the point cloud radar involves the fine division of voxels, a large number of three-dimensional convolution calculations, and the calculation is complex and time-consuming; and the image brings considerable computational pressure due to the multi-camera, high resolution, and complex feature extraction network.

[0069] Therefore, the present application provides a three-dimensional target detection method based on multimodal data, which realizes efficient feature extraction and fusion optimization of lidar and camera by constructing a more effective model and providing a more effective model training architecture, through a more sophisticated feature extraction network and more powerful data preprocessing.

[0070] Embodiment 1

[0071] The embodiment of the present application provides a three-dimensional target detection architecture, which can be applied to the field of autonomous driving and can also be applied to roads to detect road information. This embodiment is described when it is applied to the field of autonomous driving. Figure 1 As shown in the architecture diagram, the architecture provided by the present application mainly includes a synchronous mapping module 110, an image processing module 120, a point cloud processing module 130, an auxiliary supervision module 140 and a fusion detection module 150. First, by extracting the image features of the camera image corresponding to each camera and projecting the extracted image features back into a ray in the three-dimensional space, a dense first bird's eye view (BEV) feature F is generated. B At the same time, the point cloud features are extracted from the laser point cloud data and projected to generate the second bird's-eye view feature F B2 , the first bird's-eye view feature and the second bird's-eye view feature are fused and decoded to obtain the final feature output F F2 , and then use the set 3D target detection head to decode the above feature output to determine the target detection result. It provides a more effective detection architecture, which realizes efficient feature extraction and fusion optimization of lidar and camera through a more sophisticated feature extraction network and more powerful data preprocessing.

[0072] Before implementing 3D object detection using the 3D object detection architecture disclosed in the embodiment of the present application, it is necessary to train the models included in the architecture, namely, the image feature model in the image processing module 120 and the point cloud feature model in the point cloud processing module 130. The training process includes:

[0073] The data set used for training is enhanced to generate an enhanced data set, and the enhanced data set is used to input the point cloud feature model and the image feature model for training.

[0074] Specifically, due to the large number of model parameters in the training process of deep neural networks, overfitting often occurs, which affects the generalization performance of the model. At the same time, due to the uneven distribution of the number of target type samples in the data set, the detection performance of this embodiment for a small number of sample categories will be affected. In order to solve the above problems and improve the generalization performance and detection accuracy of the model, in addition to using conventional image data enhancement methods such as scaling, cropping, rotating, and flipping, we also use Figure 2 As shown in the figure, we also provide a data enhancement method, and train the image feature model and point cloud feature model through the generated data enhancement set. The specific data enhancement method includes:

[0075] a1. Build a sample database.

[0076] Specifically, from the entire training data set, according to the 3D label G of each sample S S Extract the corresponding point cloud P S (Point cloud data) and image block I S (Image data). All extracted sample data and corresponding related information (camera parameters, radar parameters and vehicle position corresponding to the sample time, etc.) are organized into a sample database.

[0077] a2. Random sample selection.

[0078] During the training process, the candidate sample set is randomly sampled from the sample database. According to the camera intrinsic parameters at the current moment, the external parameter relationship between the camera and the radar, and the three-dimensional true value frame of the sample, the camera index and image true value frame of the sample are obtained.

[0079] a3. Get the current frame information.

[0080] Get the image truth frame, 3D truth frame, corresponding image data and point cloud data corresponding to all existing samples at the current moment.

[0081] a4. Synchronize the texture in the current frame.

[0082] Depth mapping is performed on all samples in the candidate sample set and the existing sample set according to the depth value corresponding to the center point of the three-dimensional truth box under each camera from far to near, and the corresponding point cloud data is pasted into the depth-mapped samples according to the camera index pasted by the obtained example sample to generate an enhanced data set.

[0083] The above training process also includes:

[0084] The auxiliary supervision module 140 performs the first bird's-eye view feature F B Perform BEV encoding to determine the first detection feature F F , where F F ∈R X×Y×C , F B ∈R X×Y×C ; X, Y represent feature dimensions, C represents the number of channels of image features, F B and F F X, Y and C are basically consistent. Then the detection head is set to detect F F Decode it into [x,y,z,w,h,l,yaw,v x ,v y ], the above are the coordinates, length, width, height, rotation angle and speed of the center point of the three-dimensional truth box; according to [x, y, z, w, h, l, yaw, v x ,v y ] performs prediction and classification to determine the prediction loss L cls and classification loss L reg The specific determination method is a conventional method in the art and will not be described in detail here. cls and classification loss L reg Determine the target loss function L, where λ cls ,λ reg is the weight corresponding to each loss. The auxiliary supervision module converts and modifies the image feature model and the point cloud feature model according to the generated target loss function until the value of the target loss function reaches a satisfactory value and stops iterating. At this time, the image feature model and the point cloud feature model have completed training. The auxiliary supervision module 140 participates in training but not in reasoning, does not bring additional calculations, and optimizes and improves the prediction ability of the target.

[0085] Using the three-dimensional object detection architecture disclosed in the embodiment of this application to realize three-dimensional object detection specifically includes the following contents:

[0086] The synchronous mapping module 110 is used to perform depth mapping on the collected multiple camera images and point cloud data to generate an image to be processed. Specifically, mapping is performed from far to near according to the depth values ​​corresponding to the center points of the three-dimensional true value frames of the multiple camera images under each camera, and the image data is enhanced. After the depth mapping is completed, the corresponding point cloud data is pasted into the mapped camera image according to the camera index pasted by the example sample to generate an image to be processed.

[0087] The image processing module 120 is used to determine the first bird's-eye view feature according to the image to be processed and the trained image feature model. Specifically, the image processing module 120 inputs the image to be processed into the trained image feature model, and the image model extracts the image features of the image to be processed based on the deep neural network after receiving the input image to be processed. Where n represents the number of cameras, H F ,W F represents the scale of the image feature, and C represents the number of channels of the image feature.

[0088] The image processing module 120 also includes an image feature enhancement module 121 and a depth prediction module 122; wherein the image feature enhancement module 121 uses a multi-layer neural network to enhance the image feature F 0 Perform enhancement processing to generate an enhanced image feature map Further improve the image feature expression ability; the depth prediction module 122 is a multi-layer neural network that inputs image features The corresponding depth map generated Where D is the number of depth quantization, that is, the specified depth [depth min ,depth max ] is divided into D units, and the value in the i-th unit represents the probability that the depth value of the current feature point is within the depth range of the i-th unit. Preferably, in order to further improve the accuracy of depth estimation, the embodiment of the present application designs a two-layer cascaded neural network, such as Figure 3 As shown, the first depth prediction map estimated by the first layer of the neural network is used as a new feature, which is cascaded with the feature output by the first layer of the neural network in the feature channel, and then input into the second layer of the cascaded neural network. Finally, the second depth prediction map is output. The second depth map generated at this time is the depth map generated by the depth prediction module 122 in the aforementioned content.

[0089] Furthermore, the depth prediction module 122 is also affected by the point cloud feature model when generating a depth map. The point cloud feature model projects the point cloud data in the image to be processed onto the camera image to generate target depth information and outputs it to the image feature model; after the image feature model receives the target depth information output by the point cloud feature model, it uses the received target depth information to perform depth correction on the generated depth map.

[0090] The image processing module 120 enhances the image feature map F and the predicted corresponding depth map Depth through image feature mapping, and generates a three-dimensional feature map in the image coordinate system through existing technology. Specifically, F I Each point feature F I (u, v, i) = Depth(u, v, i) × F(u, v). Using existing technology, the features of each point in the three-dimensional feature F in the image coordinate system are converted to the autonomous driving vehicle coordinate system through the camera intrinsic parameters and the rotation and translation relationship between the camera and the autonomous driving vehicle. The points converted to the autonomous driving vehicle coordinate system are voxelized to form three-dimensional features in the autonomous driving vehicle coordinate system, where the features falling into the same voxel grid are accumulated, and the features of the voxel grid where no feature points fall are set to all 0. Finally, along the height dimension, the features corresponding to the voxels at all heights are accumulated to output the final three-dimensional feature F in the autonomous driving vehicle coordinate system. B ∈R X×Y×C , at this time the three-dimensional feature F in the coordinate system of the autonomous driving vehicle B ∈R X×Y×C That is the first bird's-eye view feature, where X and Y represent the dimensions of the bird's-eye view feature.

[0091] The point cloud processing module 130 is used to determine the second bird's-eye view features according to the image to be processed and the trained point cloud feature model. Specifically, the point cloud processing module 130 inputs the image to be processed into the point cloud feature model, extracts point cloud features using the point cloud feature model, and uses an existing voxelization method, which is not described in detail in this application. The point cloud features are multi-scale fused using a preset point cloud feature extraction network in the point cloud feature model. Specifically, Figure 4 As shown in FIG. 1 , after the input point cloud features are voxelized, the point cloud features are convolved in turn using a multi-layer three-dimensional convolutional network and the intermediate features output by each layer of the three-dimensional convolutional network are recorded; all the intermediate features are combined to determine the multi-scale fused point cloud features F B2 It is understandable that the present embodiment can also fuse the point cloud features according to a conventional point cloud feature extraction network, that is, after the input point cloud features are voxelized into three dimensions, they are directly reduced to a two-dimensional bird's-eye view space through multiple layers of three-dimensional convolution.

[0092] The fusion detection module 150 is used to detect the first bird's-eye view feature F generated by the image processing module 120. B And the second bird's-eye view feature F generated by the point cloud feature module B2 Then based on the fused first bird's-eye view feature F B and the second bird's-eye view feature F B2Determine the three-dimensional target detection result. Specifically, the conventional BEV encoding is used to implement it, that is, the fused F B and F B2 Perform BEV encoding to obtain the final feature output F F2 ∈R X×Y×C , and then connect the 3D target detection head to F F2 Decode into [x, y, z, w, h, l, yaw, v x ,v y ], and according to [x, y, z, w, h, l, yaw, v x , v y ] performs prediction and classification to determine the three-dimensional target detection results, where the three-dimensional target detection results include the target's three-dimensional position, length, width, height, speed and classification score.

[0093] Embodiment 2

[0094] Corresponding to the above-mentioned embodiment 1, the present application also provides a three-dimensional target detection method in an embodiment, such as Figure 5 The flowchart shown specifically includes:

[0095] 5100, performing depth mapping on the collected multiple camera images and point cloud data to generate an image to be processed;

[0096] 5200. Determine, according to the image to be processed and the trained image feature model, a three-dimensional feature of the image to be processed in an image coordinate system and convert the three-dimensional feature in the image coordinate system into a three-dimensional feature in a reference coordinate system to determine a first bird's-eye view feature;

[0097] Wherein, the image feature model includes:

[0098] Receive the input image to be processed, and extract image features of the image to be processed based on a deep neural network;

[0099] Performing enhancement processing on the image features based on a multi-layer neural network to generate an enhanced image feature map;

[0100] Performing depth prediction based on the image features of the depth prediction module to generate a corresponding depth map;

[0101] Determine the three-dimensional features in the image coordinate system based on the enhanced image feature map and the corresponding depth map;

[0102] The three-dimensional features in the image coordinate system are converted into three-dimensional features in a reference coordinate system for output.

[0103] It should be noted that the reference coordinate system disclosed in the embodiment of the present application is a coordinate system established based on the observation position. By setting the reference coordinate system, the applicable scenarios of the three-dimensional target detection method provided in the embodiment of the present application are expanded to meet the needs of different scenarios. For example, in the field of autonomous driving, it is a coordinate system constructed with the vehicle itself as the origin, and in the field of road detection, it is a coordinate system constructed with the origin of the road detection point.

[0104] 5300. Extract point cloud features of the image to be processed and perform multi-scale fusion according to the image to be processed and the trained point cloud feature model to determine a second bird's-eye view feature;

[0105] Wherein, the point cloud feature model includes:

[0106] Receiving the input image to be processed, and extracting point cloud features of the point cloud data pasted in the image to be processed;

[0107] Based on a preset point cloud feature extraction network, the point cloud features are multi-scale fused and the multi-scale fused point cloud features are output.

[0108] The step of performing multi-scale fusion on the point cloud features based on a preset point cloud feature extraction network and outputting the multi-scale fused point cloud features comprises:

[0109] Use multi-layer 3D convolutional networks to convolve point cloud features in sequence and record the intermediate features output by each layer of 3D convolutional network;

[0110] All the intermediate features are combined to determine the multi-scale fused point cloud features.

[0111] Wherein, the first bird's-eye view feature is a three-dimensional feature in a reference coordinate system output by the image feature model;

[0112] The second bird's-eye view feature is a multi-scale fused point cloud feature output by the point cloud feature model.

[0113] The method further comprises:

[0114] After the image to be processed is input into the point cloud feature model, the point cloud feature model projects the point cloud data in the image to be processed onto the camera image to generate target depth information and outputs it to the image feature model;

[0115] The image feature model receives the target depth information output by the point cloud feature model, and performs depth correction on the generated depth map;

[0116] The image feature model determines the three-dimensional features in the image coordinate system based on the depth map after depth correction and the enhanced image feature map.

[0117] The training process of the image feature model and the point cloud feature model includes:

[0118] Extract the point cloud data and image data of each sample in the training data set to build a sample database;

[0119] Randomly selecting candidate samples from the sample database to determine a candidate sample set;

[0120] Get the existing sample set at the current moment;

[0121] Performing depth mapping processing on samples in the existing sample set and the spare sample set according to the three-dimensional true value frame corresponding to each sample in the existing sample set and the spare sample set;

[0122] Paste the point cloud data into the corresponding depth map processed sample to generate an enhanced dataset;

[0123] The image feature model and the point cloud feature model are trained using the enhanced data set.

[0124] The training process of the image feature model and the point cloud feature model further includes:

[0125] BEV encoding the first bird's eye view feature to determine a first detection feature;

[0126] decoding the first detection feature;

[0127] determining a prediction loss and a classification loss in response to the decoded first detection feature;

[0128] Determine a target loss function based on the prediction loss and the classification loss;

[0129] The image feature model and the point cloud feature model are modified based on the target loss function.

[0130] 540. Fuse the first bird's-eye view feature and the second bird's-eye view feature and determine a three-dimensional target detection result based on the fused first bird's-eye view feature and the second bird's-eye view feature.

[0131] There is no sequence for step 520 and step 530. Step 520 may be performed first and then step 530; step 530 may be performed first and then step 520; or step 520 and step 530 may be performed simultaneously.

[0132] Embodiment 3

[0133] Corresponding to the first and second embodiments, the present application also provides a three-dimensional target detection system, such as Figure 6 As shown, including:

[0134] A pre-processing module 610 is used to perform depth mapping on the collected multiple camera images and point cloud data to generate an image to be processed;

[0135] A model prediction module 620 is used to determine the three-dimensional features of the image to be processed in the image coordinate system according to the image to be processed and the trained image feature model, and to convert the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system, so as to determine the first bird's-eye view features;

[0136] The model prediction module 620 is further used to extract the point cloud features of the image to be processed and perform multi-scale fusion according to the image to be processed and the trained point cloud feature model to determine the second bird's-eye view features;

[0137] The data processing module 630 is used to fuse the first bird's-eye view feature and the second bird's-eye view feature and determine a three-dimensional target detection result based on the fused first bird's-eye view feature and the second bird's-eye view feature.

[0138] In some embodiments, the model prediction module 620 is also used to perform the following operations using the image feature model: receiving the input image to be processed, extracting image features of the image to be processed based on a deep neural network; enhancing the image features based on a multi-layer neural network to generate an enhanced image feature map; performing depth prediction on the image features based on a depth prediction module to generate a corresponding depth map; determining the three-dimensional features in the image coordinate system based on the enhanced image feature map and the corresponding depth map; and converting the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system for output.

[0139] In some embodiments, the model prediction module 620 is also used to perform the following operations using the point cloud feature model: receiving the input image to be processed, extracting the point cloud features of the point cloud data pasted in the image to be processed; performing multi-scale fusion of the point cloud features based on a preset point cloud feature extraction network and outputting the multi-scale fused point cloud features.

[0140] Among them, the first bird's-eye view feature is a three-dimensional feature in a reference coordinate system output by an image feature model; and the second bird's-eye view feature is a multi-scale fused point cloud feature output by a point cloud feature model.

[0141] In some embodiments, the model prediction module 620 is also used to use the point cloud feature model to convolve the point cloud features in sequence through multiple layers of three-dimensional convolutional networks and record the intermediate features output by each layer of the three-dimensional convolutional network; merge all the intermediate features to determine the point cloud features after multi-scale fusion.

[0142] In some embodiments, the model prediction module 620 is also used to, after the image to be processed is input into the point cloud feature model, use the point cloud feature model to project the point cloud data in the image to be processed onto the camera image to generate target depth information and output it to the image feature model; use the image feature model to receive the target depth information output by the point cloud feature model, and perform depth correction on the generated depth map; and use the image feature model to determine the three-dimensional features in the image coordinate system based on the depth map after depth correction and the enhanced image feature map.

[0143] In some embodiments, the three-dimensional target detection system also includes a model training module 640 (not shown in the figure), which is used to extract point cloud data and image data of each sample in the training data set to construct a sample database; randomly select alternative samples from the sample database to determine an alternative sample set; obtain the existing sample set at the current moment; perform depth mapping on the samples in the existing sample set and the spare sample set according to the three-dimensional true value box corresponding to each sample in the existing sample set and the spare sample set; paste the point cloud data into the corresponding depth map processed sample to generate an enhanced data set; and use the enhanced data set to train the image feature model and the point cloud feature model.

[0144] In some embodiments, the model training module 640 is also used to perform BEV encoding on the first bird's-eye view feature to determine a first detection feature; decode the first detection feature; determine a prediction loss and a classification loss in response to the decoded first detection feature; determine a target loss function based on the prediction loss and the classification loss; and correct the image feature model and the point cloud feature model based on the target loss function.

[0145] Embodiment 4

[0146] Corresponding to all the above embodiments, an embodiment of the present application provides an electronic device, including:

[0147] One or more processors; and a memory associated with the one or more processors, the memory being used to store program instructions, the program instructions when read and executed by the one or more processors, performing the following operations:

[0148] Depth mapping is performed on the collected multiple camera images and point cloud data to generate an image to be processed;

[0149] According to the image to be processed and the trained image feature model, determine the three-dimensional features of the image to be processed in the image coordinate system and convert the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system to determine the first bird's-eye view features;

[0150] According to the image to be processed and the trained point cloud feature model, point cloud features of the image to be processed are extracted and multi-scale fusion is performed to determine a second bird's-eye view feature;

[0151] The first bird's-eye view feature and the second bird's-eye view feature are fused, and a three-dimensional target detection result is determined based on the fused first bird's-eye view feature and the second bird's-eye view feature.

[0152] in, Figure 7 The architecture of the electronic device is shown as an example, which may include a processor 710, a video display adapter 711, a disk drive 712, an input / output interface 713, a network interface 714, and a memory 720. The processor 710, the video display adapter 711, the disk drive 712, the input / output interface 713, the network interface 714, and the memory 720 may be communicatively connected via a bus 730.

[0153] Among them, the processor 710 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided in this application.

[0154] The memory 720 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 720 can store an operating system 721 for controlling the execution of the electronic device 700, and a basic input and output system (BIOS) 722 for controlling the low-level operations of the electronic device 700. In addition, a web browser 723, a data storage management system 724, and an icon font processing system 727, etc. can also be stored. The above-mentioned icon font processing system 727 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided in the present application is implemented by software or firmware, the relevant program code is stored in the memory 720 and is called and executed by the processor 710.

[0155] The input / output interface 713 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0156] The network interface 714 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0157] The bus 730 comprises a pathway for transmitting information between the various components of the device (eg, the processor 710, the video display adapter 711, the disk drive 712, the input / output interface 713, the network interface 714, and the memory 720).

[0158] In addition, the electronic device 700 can also obtain information on specific collection conditions from the virtual resource object collection condition information database for use in condition judgment, etc.

[0159] It should be noted that, although the above device only shows a processor 710, a video display adapter 711, a disk drive 712, an input / output interface 713, a network interface 714, a memory 720, a bus 730, etc., in the specific implementation process, the device may also include other components necessary for normal execution. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.

[0160] It can be seen from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a disk, an optical disk, etc., including several instructions for enabling a computer device (which can be a personal computer, a cloud service end, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0161] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0162] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A three-dimensional object detection method, It is characterized in that The method comprises: Depth mapping is performed on the collected multiple camera images and point cloud data to generate an image to be processed; According to the image to be processed and the trained image feature model, determine the three-dimensional features of the image to be processed in the image coordinate system and convert the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system to determine the first bird's-eye view features; According to the image to be processed and the trained point cloud feature model, point cloud features of the image to be processed are extracted and multi-scale fusion is performed to determine a second bird's-eye view feature; fusing the first bird's-eye view feature and the second bird's-eye view feature and determining a three-dimensional target detection result based on the fused first bird's-eye view feature and the second bird's-eye view feature; The image feature model comprises: Receive the input image to be processed, and extract image features of the image to be processed based on a deep neural network; Performing enhancement processing on the image features based on a multi-layer neural network to generate an enhanced image feature map; Performing depth prediction based on the image features of the depth prediction module to generate a corresponding depth map; Determine the three-dimensional features in the image coordinate system based on the enhanced image feature map and the corresponding depth map; Convert the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system for output; The point cloud feature model includes: Receiving the input image to be processed, and extracting point cloud features of the point cloud data pasted in the image to be processed; Based on a preset point cloud feature extraction network, the point cloud features are multi-scale fused and the multi-scale fused point cloud features are output; After the image to be processed is input into the point cloud feature model, the point cloud feature model projects the point cloud data in the image to be processed onto the camera image to generate target depth information and outputs it to the image feature model; The image feature model receives the target depth information output by the point cloud feature model, and performs depth correction on the generated depth map; The image feature model determines the three-dimensional features in the image coordinate system based on the depth map after depth correction and the enhanced image feature map.

2. The method according to claim 1, It is characterized in that The step of performing multi-scale fusion on the point cloud features based on a preset point cloud feature extraction network and outputting the multi-scale fused point cloud features comprises: Use multi-layer 3D convolutional networks to convolve point cloud features in sequence and record the intermediate features output by each layer of 3D convolutional network; All the intermediate features are combined to determine the multi-scale fused point cloud features.

3. The method according to claim 1, It is characterized in that The first bird's-eye view feature is a three-dimensional feature in a reference coordinate system output by the image feature model; The second bird's-eye view feature is a multi-scale fused point cloud feature output by the point cloud feature model.

4. The method according to any one of claims 1 to 3, It is characterized in that The training process of the image feature model and the point cloud feature model includes: Extract the point cloud data and image data of each sample in the training data set to build a sample database; Randomly selecting candidate samples from the sample database to determine a candidate sample set; Get the existing sample set at the current moment; Performing depth mapping processing on samples in the existing sample set and the spare sample set according to the three-dimensional true value frame corresponding to each sample in the existing sample set and the spare sample set; Paste the point cloud data into the corresponding depth map processed sample to generate an enhanced dataset; The image feature model and the point cloud feature model are trained using the enhanced data set.

5. The method according to any one of claims 1 to 3, It is characterized in that The training process of the image feature model and the point cloud feature model also includes: BEV encoding the first bird's eye view feature to determine a first detection feature; decoding the first detection feature; determining a prediction loss and a classification loss in response to the decoded first detection feature; Determine a target loss function based on the prediction loss and the classification loss; The image feature model and the point cloud feature model are modified based on the target loss function.

6. A three-dimensional object detection system, It is characterized in that The system comprises: A preprocessing module, used for performing depth mapping on the collected multiple camera images and point cloud data to generate an image to be processed; A model prediction module, used to determine the three-dimensional features of the image to be processed in the image coordinate system according to the image to be processed and the trained image feature model, and convert the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system to determine the first bird's-eye view features; The model prediction module is further used to extract the point cloud features of the image to be processed and perform multi-scale fusion according to the image to be processed and the trained point cloud feature model to determine the second bird's-eye view features; A data processing module, configured to fuse the first bird's-eye view feature and the second bird's-eye view feature and determine a three-dimensional target detection result based on the fused first bird's-eye view feature and the second bird's-eye view feature; Wherein, the image feature model includes: Receive the input image to be processed, and extract image features of the image to be processed based on a deep neural network; Performing enhancement processing on the image features based on a multi-layer neural network to generate an enhanced image feature map; Performing depth prediction based on the image features of the depth prediction module to generate a corresponding depth map; Determine the three-dimensional features in the image coordinate system based on the enhanced image feature map and the corresponding depth map; Convert the three-dimensional features in the image coordinate system into three-dimensional features in the reference coordinate system for output; The point cloud feature model includes: Receiving the input image to be processed, and extracting point cloud features of the point cloud data pasted in the image to be processed; Based on a preset point cloud feature extraction network, the point cloud features are multi-scale fused and the multi-scale fused point cloud features are output; Wherein, the model prediction module is further used for, after the image to be processed is input into the point cloud feature model, the point cloud feature model projects the point cloud data in the image to be processed onto the camera image to generate target depth information and output it to the image feature model; The image feature model receives the target depth information output by the point cloud feature model, and performs depth correction on the generated depth map; The image feature model determines the three-dimensional features in the image coordinate system based on the depth map after depth correction and the enhanced image feature map.

7. An electronic device, It is characterized in that The electronic device comprises: one or more processors; And a memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the method according to any one of claims 1-5 is executed.

Citation Information

Patent Citations

  • Three-dimensional target detection method and device based on multi-sensor information fusion

    CN110929692A

  • Three-dimensional target detection method and device and computer storage medium

    CN115909269A