3D Object Detection Method, Encoder, and Decoder

By interacting and fusion of images and three-dimensional point cloud features in three-dimensional object detection, the problem of not being able to fully utilize the information complementarity advantages of multi-sensor data in the prior art is solved, and high-quality three-dimensional object detection is achieved.

CN115063768BActive Publication Date: 2025-06-17ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210810402.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-11
Publication Date
2025-06-17
Estimated Expiration
2042-07-11

AI Technical Summary

Technical Problem

The existing three-dimensional object detection method based on multi-sensors cannot fully utilize the information complementarity of images and three-dimensional point cloud data, resulting in the inability to obtain high-quality three-dimensional detection results.

Method used

By acquiring the images and three-dimensional point clouds of the three-dimensional environment, extracting the image feature map and point cloud feature map, and performing feature interaction, we obtain an enhanced image feature map that integrates point cloud features and an enhanced point cloud feature map that integrates image features, and then performing three-dimensional object detection.

Benefits of technology

The effective fusion of images and three-dimensional point cloud data is achieved, and the richness and comprehensiveness of feature information is enhanced, thereby improving the accuracy and efficiency of three-dimensional object detection and obtaining high-quality detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063768B_ABST
    Figure CN115063768B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a three-dimensional object detection method, an electronic device, and a computer storage medium. The three-dimensional object detection method includes: obtaining an image and a three-dimensional point cloud of a three-dimensional environment to be detected, as well as an image feature map corresponding to the image and a point cloud feature map corresponding to the three-dimensional point cloud; performing feature interaction between the image feature map and the point cloud feature map, and obtaining an enhanced image feature map integrating point cloud features and an enhanced point cloud feature map integrating image features according to the feature interaction result; and performing three-dimensional object detection in the three-dimensional environment based on the enhanced image feature map and the enhanced point cloud feature map. Through the embodiments of the present application, high-quality three-dimensional object detection results can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of object detection, and in particular, to a three-dimensional object detection method, an electronic device, and a computer storage medium. Background Art

[0002] In fields such as driverless and robotics, the ability to perceive the environment is an important ability to ensure their normal operation. To achieve correct perception of the environment, three-dimensional object detection is a basic function of environmental perception. Through three-dimensional object detection, the position and category of objects in a three-dimensional scene can be predicted, and the size and orientation of the objects can be determined. Furthermore, based on this, operation tasks such as trajectory prediction and path planning can be realized.

[0003] In order to achieve the perception of the surrounding environment, vehicles or robots with autonomous driving functions are often equipped with a variety of sensors, such as lidar, surround-view cameras, etc., in order to make the different modalities of data collected by various sensors complementary in information. An existing multi-sensor based three-dimensional object detection method uses the images collected by the surround-view camera and the point clouds collected by the lidar independently for detection, and integrates the two at the detection result level. Although this method can directly use the existing detection algorithms for a single modality and the fusion difficulty is relatively low, however, this method generates initial detection results independently using the data of each modality, resulting in the inability to fully utilize the information complementary advantages of the two modalities of data. Therefore, high-quality three-dimensional detection results cannot be obtained. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a three-dimensional object detection solution to at least partially solve the above problems.

[0005] According to the first aspect of the embodiments of the present application, a three-dimensional object detection method is provided, including: obtaining an image and a three-dimensional point cloud of a three-dimensional environment to be detected, as well as an image feature map corresponding to the image and a point cloud feature map corresponding to the three-dimensional point cloud; performing feature interaction between the image feature map and the point cloud feature map, and obtaining an enhanced image feature map integrating point cloud features and an enhanced point cloud feature map integrating image features according to the feature interaction result; based on the enhanced image feature map and the enhanced point cloud feature map, performing three-dimensional object detection in the three-dimensional environment.

[0006] According to the second aspect of the embodiments of the present application, an encoder is provided. The encoder includes a plurality of encoding layers. Each encoding layer includes: a feature input part for receiving the image feature map and the point cloud feature map output by the previous layer; a feature interaction part for respectively performing a first feature interaction between image features and a second feature interaction between point cloud features and image features on the image feature map, and performing a third feature interaction between point cloud features and a fourth feature interaction between image features and point cloud features on the point cloud feature map; and a feature fusion part for performing feature fusion based on the first feature interaction result and the second feature interaction result to obtain an enhanced image feature map, and performing feature fusion based on the third feature interaction result and the fourth feature interaction result to obtain an enhanced point cloud feature map.

[0007] According to the third aspect of the embodiments of the present application, a decoder is provided. The decoder includes a plurality of decoding layers. Adjacent decoding layers are used to decode different types of feature maps, where the different types of feature maps include an enhanced point cloud feature map and an enhanced image feature map. The enhanced point cloud feature map and the enhanced image feature map are obtained by respectively performing feature interaction between the point cloud feature map and the image feature map corresponding to the three-dimensional environment to be detected, and obtaining the corresponding enhanced point cloud feature map fused with image features and the enhanced image feature map fused with point cloud features according to the feature interaction result. The point cloud feature map and the image feature map are respectively obtained by performing feature extraction on the image and the three-dimensional point cloud corresponding to the three-dimensional environment.

[0008] According to the fourth aspect of the embodiments of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the method described in the first aspect.

[0009] According to the fifth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the method described in the first aspect.

[0010] According to the solution provided by the embodiments of the present application, when performing 3D object detection, detection is carried out based on two aspects of modal information of the 3D environment to be detected, namely images and 3D point clouds. Moreover, during the detection process, feature interaction is performed between the image feature map corresponding to the image and the point cloud feature map corresponding to the 3D point cloud, so that the modal information of one aspect can be effectively integrated into the modal information of the other aspect, enabling the modal information of the two aspects to complement each other, thereby achieving feature enhancement of the integrated modal information, and thus obtaining the corresponding enhanced image feature map and enhanced point cloud feature map. Based on the enhanced image feature map and the enhanced point cloud feature map, 3D object detection is then carried out. Since the corresponding information is more comprehensive, rich, and characteristic, the 3D object detection is more accurate and efficient, and high-quality 3D object detection results can be obtained. Description of the Drawings

[0011] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0012] Figure 1A Schematic diagram of an exemplary system applicable to the solution of the embodiments of the present application;

[0013] Figure 1B Schematic diagram of the structure of an exemplary 3D object detection model applicable to the solution of the embodiments of the present application;

[0014] Figure 2A Flowchart of the steps of a 3D object detection method according to an embodiment of the present application;

[0015] Figure 2B For Figure 2A Schematic diagram of the structure of an encoding layer in the illustrated embodiment;

[0016] Figure 2C Based on Figure 2B Schematic diagram of feature interaction from point cloud features to image features based on the illustrated encoder;

[0017] Figure 2D Based on Figure 2B Schematic diagram of feature mapping from image features to point cloud features based on the illustrated encoder;

[0018] Figure 2E For Figure 2A Schematic diagram of the structure of a decoder in the illustrated embodiment;

[0019] Figure 2F For Figure 2ESchematic diagram of the structure of a decoding layer in the decoder shown in

[0020] Figure 2G is Figure 2A Schematic diagram of an example scenario in the illustrated embodiment;

[0021] Figure 3 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0022] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.

[0023] The specific implementation of the embodiments of the present application will be further described below in conjunction with the accompanying drawings of the embodiments of the present application.

[0024] Figure 1A An exemplary system applicable to the solution of the embodiment of the present application is shown. As shown in FIG. 1, the system 100 may include a cloud server 102, a communication network 104, and / or one or more user devices 106. In FIG. 1, multiple user devices are exemplified.

[0025] The cloud server 102 may be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 102 may perform any suitable function. For example, in some embodiments, the cloud server 102 may be used for three-dimensional object detection based on the images and three-dimensional point clouds of the three-dimensional environment to be detected. As an optional example, in some embodiments, the cloud server 102 may be used for feature interaction based on the image feature map corresponding to the image of the three-dimensional environment to be detected and the point cloud feature map corresponding to the three-dimensional point cloud, and obtaining an enhanced image feature map and an enhanced point cloud feature map after feature fusion according to the feature interaction result; further, performing three-dimensional object detection based on the enhanced image feature map and the enhanced point cloud feature map. As another example, in some embodiments, the cloud server 102 may be used to send the three-dimensional object detection result to the user device, or send the result after processing the downstream operation task based on the three-dimensional object detection result to the user device.

[0026] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user device 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud server 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user device 106 and the cloud server 102, such as a network link, a dial-up link, a wireless link, a hardwired link, any other suitable communication link, or any suitable combination of such links.

[0027] The user device 106 can include any one or more user devices suitable for interacting with a user and capable of collecting images and point cloud data of a three-dimensional environment. In some embodiments, the user device 106 can include any suitable type of device. For example, in some embodiments, the user device 106 can include a vehicle with an autonomous driving function, an aircraft, a robot, and / or any other suitable type of user device.

[0028] In an alternative embodiment, the three-dimensional object detection provided by the embodiments of the present application can be implemented based on an exemplary three-dimensional object detection model structure, such as Figure 1B shown.

[0029] As can be seen from the figure, the three-dimensional object detection model structure includes a feature extractor, an encoder, and a decoder.

[0030] Among them, the feature extractor includes two parts. One part is used to extract features from the three-dimensional point cloud to obtain a corresponding point cloud feature map. Figure 1B An example in the figure is to extract features from the collected three-dimensional point cloud data through a point cloud feature extractor and generate a corresponding bird's-eye view point cloud feature map. However, those skilled in the art should understand that other non-bird's-eye view forms of point cloud feature maps are also applicable to the solution of the embodiments of the present application. The other part is used to extract features from the image to obtain a corresponding image feature map. Figure 1BIn the example, an image feature extractor extracts features from the panoramic images collected by the panoramic cameras and generates corresponding panoramic image feature maps. However, those skilled in the art should understand that images collected by other non-panoramic cameras or non-panoramic images collected are applicable to the solution of the embodiments of the present application. Among them, the above two parts can be implemented by any appropriate structure with the function of feature extraction for the corresponding modal data, including but not limited to encoder structures, convolutional network structures, etc.

[0031] The encoder in the 3D object detection model has a multi-modal feature interaction function, and constructs a deep structure for the feature fusion stage, so that the features of the two modalities can perform intensive interaction to achieve the mutual fusion and enhancement of the features of the two modalities. The specific structure and specific function implementation of this encoder will be described in detail in the corresponding parts hereinafter. Through this encoder, feature interaction can be performed between the image features corresponding to the images of the 3D environment and the point cloud features corresponding to the 3D point cloud, obtaining enhanced image features fused with point cloud features and enhanced point cloud features fused with image features.

[0032] The decoder in the 3D object detection model can perform 3D object detection based on the enhanced image feature map and the enhanced point cloud feature map after the encoder performs feature interaction enhancement. The specific structure and specific function implementation of this decoder will be described in detail in the corresponding parts hereinafter. In a feasible manner, the decoder can alternately converge the enhanced features of the two modalities to iteratively refine the prediction boxes and generate the final prediction. Subsequently, the prediction result, that is, the 3D object detection result, is output.

[0033] Hereinafter, based on the above system and model structure description, the 3D object detection solution provided by the embodiments of the present application will be described through embodiments.

[0034] Refer to Figure 2A , which shows a flowchart of the steps of a 3D object detection method according to an embodiment of the present application.

[0035] The 3D object detection method includes the following steps:

[0036] Step S202: Obtain the image and 3D point cloud of the 3D environment to be detected, as well as the image feature map corresponding to the image and the point cloud feature map corresponding to the 3D point cloud.

[0037] The three-dimensional environment to be detected is usually the physical environment where certain devices are located, especially vehicles or aircraft or robots with autonomous driving functions. These devices use various sensors on their own, such as lidar, cameras (which can be panoramic or non-panoramic cameras), etc. to collect data on their own physical environment, that is, the three-dimensional environment. In the embodiments of the present application, it mainly includes image acquisition and three-dimensional point cloud acquisition. These two types of data form two types of modal data for the three-dimensional environment, so that devices with autonomous driving functions can perceive the surrounding environment based on these two types of collected modal data.

[0038] After obtaining the two types of modal data of the three-dimensional environment, i.e., images and three-dimensional point clouds, feature extraction can be performed separately to obtain an image feature map corresponding to the image and a point cloud feature map corresponding to the three-dimensional point cloud. When using a three-dimensional object detection model as shown in Figure 1B For the image of the three-dimensional environment, the feature extractor for the image in its feature extractor can be used to perform feature extraction on the image of the three-dimensional environment to obtain an image feature map; for the three-dimensional point cloud of the three-dimensional environment, the feature extractor for the three-dimensional point cloud in its feature extractor can be used to perform feature extraction on the three-dimensional point cloud of the three-dimensional environment to obtain a point cloud feature map. Optionally, this point cloud feature map can be a bird's-eye view feature map of the point cloud to more effectively capture the information in the three-dimensional point cloud, especially the depth information. However, as mentioned above, other methods of feature extraction are also applicable to the solutions of the embodiments of the present application.

[0039] Step S204: Perform feature interaction between the image feature map and the point cloud feature map, and obtain an enhanced image feature map that fuses point cloud features and an enhanced point cloud feature map that fuses image features according to the feature interaction result.

[0040] The image and the three-dimensional point cloud represent the three-dimensional environment from different perspectives. However, due to the different acquisition methods and presentation methods, the information carried by the two is also different. In the traditional method, these two parts of data are processed separately and only integrated at the result level. However, since the two parts of data are processed independently during the processing, the obtained results are also biased, resulting in inaccurate results even when integrated at the result level. Therefore, the solution provided by the embodiments of the present application performs feature interaction on the features corresponding to these two parts of data at the feature stage to make them complementary, so that the generated feature map after interaction can carry more abundant and comprehensive information, providing high-quality feature data for subsequent three-dimensional object detection.

[0041] Based on this, in a feasible manner, this step can be implemented as follows: for the image feature map, perform a first feature interaction from image features to image features and a second feature interaction from point cloud features to image features; according to the results of the first feature interaction and the second feature interaction, obtain an enhanced image feature map that incorporates point cloud features. Also, for the point cloud feature map, perform a third feature interaction from point cloud features to point cloud features and a fourth feature interaction from image features to point cloud features; according to the results of the third feature interaction and the fourth feature interaction, obtain an enhanced point cloud feature map that incorporates image features. Through this method, not only can the two modal features be effectively fused, but also the feature representation of their own modal features can be further enhanced.

[0042] Among them, the first feature interaction from image features to image features and the third feature interaction from point cloud features to point cloud features can both be obtained by performing feature enhancement processing based on the feature data of their own modality. For example, local attention calculation is performed. Specifically, local attention calculation is performed on the image features to achieve the first feature interaction, obtaining an enhanced first part of the image features; local attention calculation is performed on the point cloud features to achieve the third feature interaction, obtaining an enhanced first part of the point cloud features.

[0043] For the second feature interaction from point cloud features to image features, in a feasible manner, the point cloud features corresponding to the position points in the image feature map can be projected and transformed into image features. Among them, when projecting and transforming the point cloud features corresponding to the position points in the image feature map into image features, the three-dimensional point cloud can be projected and transformed into a bird's-eye view with depth information; according to the positions of the pixel points in the bird's-eye view, the point cloud features at the corresponding positions in the point cloud feature map are obtained; the obtained point cloud features are projected and transformed into image features. Through the above method, feature alignment and fusion can be accurately performed.

[0044] For each pixel point in the image feature map, it corresponds to a position (coordinate). Based on this position, the corresponding position in the point cloud feature map can be determined (such as through projection or coordinate transformation), and this position corresponds to corresponding point cloud features. Then, after obtaining the point cloud features, they can be fused with the image features at the position of this pixel point to obtain the fused image features at the position of this pixel point.

[0045] For the fourth feature interaction from image features to point cloud features, in a feasible manner, for a certain pixel point in the point cloud feature map, multiple reference pixel points of this pixel point in the point cloud feature map can be obtained; based on the positions of this pixel point and the multiple reference pixel points, multiple image features at corresponding positions can be obtained from the image feature map; the point cloud feature corresponding to this pixel point and the obtained multiple image features are fused. Among them, when obtaining multiple image features at corresponding positions from the image feature map based on the positions of this pixel point and the multiple reference pixel points, projection transformation can be performed based on the positions of this pixel point and the multiple reference pixel points to obtain corresponding two-dimensional coordinates; based on the two-dimensional coordinates, feature acquisition at corresponding positions is performed on the image feature map to obtain corresponding multiple image features. In this way, the deviation that may be caused by collecting the image features of a single pixel point is avoided, making the collected image features more objective and effective.

[0046] When adopting a 3D object detection model as shown in Figure 1B , the above-mentioned first, second, third, and fourth feature interactions can all be implemented by the encoder of the 3D object detection model.

[0047] In one example, the encoder may include multiple encoding layers, and each encoding layer includes: a feature input part, a feature interaction part, and a feature fusion part. It should be noted that, unless otherwise specified, in the embodiments of the present application, the quantities related to "multiple", such as "multiple" and "multiple types", mean two or more.

[0048] Among them, the feature input part is used to receive the image feature map and the point cloud feature map output by the previous layer (the feature extraction layer of the feature extractor before the previous encoding layer or the first encoding layer); the feature interaction part is used to respectively perform the first feature interaction between image features to image features and the second feature interaction between point cloud features to image features on the image feature map, and perform the third feature interaction between point cloud features to point cloud features and the fourth feature interaction between image features to point cloud features on the point cloud feature map; the feature fusion part is used to perform feature fusion based on the first feature interaction result and the second feature interaction result to obtain an enhanced image feature map, and perform feature fusion based on the third feature interaction result and the fourth feature interaction result to obtain an enhanced point cloud feature map.

[0049] An exemplary structure of an encoding layer can be as shown in Figure 2B . The input of the first encoding layer of the encoder is the features h p and h c, each of the other encoding layers takes the point cloud feature maps enhanced by the previous encoding layer, such as the bird's-eye view feature map of the point cloud and the image feature map, as inputs, and outputs the multi-modal feature map of the same shape enhanced by this layer, that is, the enhanced point cloud feature map (such as the enhanced bird's-eye view feature map of the point cloud) h' p and the enhanced image feature map h' c .

[0050] Specifically, as Figure 2B shown, each encoding layer contains four feature interactions, which can be denoted as φ p→c , φ p→p , φ c→p , φ c→c . Among them, φ x→y (x, y ∈ {p, c}) represents the feature interaction from modality x to modality y, and its inputs are h x and h y , and the output is the enhanced y with the same form as h Subsequently, the features enhanced by each interaction are further fused through two MLPs (Multi-Layer Perceptrons). This process can be formalized as:

[0051]

[0052]

[0053] The following is an explanation of the above four feature interactions:

[0054] (1) The feature interactions φ p→p and φ c→c within both modalities perform local self-attention operations with the feature maps in their own modalities as inputs to achieve the interaction and enhancement of the features in their own modalities.

[0055] (2) For the feature interaction φ p→c from the point cloud feature to the image feature, in the embodiments of the present application, first use the Warp operation, specifically the BEVWarp operation for the bird's-eye view feature map, to organize h p into the same form as h c , denoted as Subsequently, use h c as the parameter Q for self-attention calculation, use as K and V, and perform local self-attention calculation. The self-attention calculation result obtained is used as the feature interaction result.

[0056] The Warp operation is a transformation operation that can transform data from one type to another through certain transformation algorithms, such as Euclidean transformation, similarity transformation, deviation transformation, projective transformation, etc. Based on this principle, in practical applications, those skilled in the art can improve and process it according to actual needs to meet their own requirements.

[0057] In this embodiment, taking the BEV Warp operation as an example, the process of organizing h p into the same form as h c will be described. As Figure 2C shown.

[0058] As can be seen from Figure 2C when organizing h p into the same form as h c firstly, the point cloud feature map h p needs to be projected onto the image to obtain a sparse depth map; then, depth completion is performed on the sparse depth map to obtain a pixel-by-pixel dense depth map; each pixel in the dense depth map is then lifted to a unique 3D point corresponding in 3D space according to its depth to become a pseudo point cloud; the pseudo point cloud is then projected onto the bird's-eye view feature map, i.e., the BEV feature map, to obtain the corresponding features; then, the obtained features are filled into the corresponding positions of the image feature map to form BEV features in the form of image features. This part of the features is subsequently fused with another part of the image features (i.e., the features after φ c→c ) that have undergone feature interaction at the corresponding position, and then an enhanced image feature map fused with point cloud features can be obtained.

[0059] (3) For the feature interaction flow φ c→p from image features to point cloud features, a correspondence can first be established between each pixel in the point cloud feature map, such as the bird's-eye view BEV feature map, and a set of features in the image feature map. Exemplarily, as Figure 2D shown, for each pixel in the BEV feature map, a preset number of feature points, such as randomly selecting 20 feature points, are selected from the spatial range where the pixel is located and projected onto the image feature map to sample 20 corresponding image features. The specific setting of the preset number can be flexibly set by those skilled in the art according to actual needs.

[0060] Based on this, when performing the multi-head attention calculation in φ c→p , taking h p as the source of the parameter Q for the multi-head attention calculation, the corresponding K and V when each pixel on the BEV feature map serves as Q are several image features determined by the above method, and local attention calculation is performed. The obtained attention calculation result is used as the feature interaction result. This part of the features is subsequently fused with another part of the point cloud features (i.e., φ p→pAfter fusing the features), an enhanced point cloud feature map incorporating image features can be obtained.

[0061] Through the above process, during the feature processing stage of the model, the interaction and fusion between different modality features are achieved. The obtained enhanced image feature map and enhanced point cloud feature map will provide richer, more comprehensive information and more accurate data basis for subsequent 3D object detection.

[0062] It should be noted that the above encoder can be implemented not only by code but also as an encoder chip or other hardware forms such as FPGA that implement the functions of the above encoder through logic circuits.

[0063] Step S206: Perform 3D object detection in a 3D environment based on the enhanced image feature map and the enhanced point cloud feature map.

[0064] Based on the enhanced image feature map and the enhanced point cloud feature map, 3D object detection can be performed in a conventional manner. However, in order to obtain better detection accuracy, in the embodiments of the present application, a method of alternately extracting features from the enhanced image feature map and the enhanced point cloud feature map and performing 3D object detection in a 3D environment based on the extracted features is adopted. In this way, by alternately interacting with the information expressed by the features of different modalities, information can be more effectively utilized to obtain more accurate detection results.

[0065] In a feasible manner, before alternately extracting features from the enhanced image feature map and the enhanced point cloud feature map, 3D object position prediction can also be performed respectively based on the point cloud feature map before the aforementioned feature interaction and the enhanced point cloud feature map after the feature interaction to obtain corresponding first prediction result and second prediction result; according to the first prediction result and the second prediction result, a preset number of candidate prediction boxes are obtained, where the probability that a candidate prediction box belongs to a certain type of target is greater than a preset probability. Among them, the preset number and the preset probability can both be appropriately set by those skilled in the art according to actual needs, and the embodiments of the present application do not limit this. Based on this, alternately extracting features from the enhanced image feature map and the enhanced point cloud feature map can be implemented as: alternately extracting features from the enhanced image feature map and the enhanced point cloud feature map based on the information of the candidate prediction boxes and performing 3D object detection in a 3D environment based on the extracted features. In this way, 3D object detection can be performed based on the candidate prediction boxes to accelerate the convergence speed of the 3D object detection model. Among them, the first prediction result and the second prediction result are the first prediction heatmap and the second prediction heatmap respectively, and the k-th channel in the first prediction heatmap and the second prediction heatmap represents the probability that the k-th type of target exists at the current position. The information of the candidate prediction boxes includes the point cloud features corresponding to the candidate prediction boxes and the information of the target category vectors corresponding to the candidate prediction boxes, so as to provide an initial effective reference for subsequent further detection and improve the detection efficiency.

[0066] Based on this, during subsequent detection, when extracting features from the enhanced image feature map, obtain the point cloud candidate prediction box corresponding to the previous enhanced point cloud feature map of the enhanced image feature map; project the point cloud candidate prediction box onto the enhanced image feature map to obtain the corresponding image candidate prediction box; when extracting features from the enhanced point cloud feature map, obtain the image candidate prediction box corresponding to the previous enhanced image feature map of the enhanced point cloud feature map, and convert the image candidate prediction box to the enhanced point cloud feature map to obtain the corresponding point cloud candidate prediction box.

[0067] Among them, when projecting the point cloud candidate prediction box onto the enhanced image feature map to obtain the corresponding image candidate prediction box, it can be adopted: magnify the point cloud candidate prediction box by a preset multiple and then project it onto the enhanced image feature map; obtain the corresponding region of the projection result on the enhanced image feature map; extract the features of the region, and obtain the corresponding image candidate prediction box according to the feature extraction result. Thus, the prediction deviation caused by projection can be avoided, and a more accurate image candidate prediction box can be obtained.

[0068] Furthermore, in order to further improve the detection efficiency and exclude data that is invalid for detection, in a feasible manner, before extracting features from the enhanced image feature map, perform multi-head self-attention calculation on the features corresponding to the point cloud candidate prediction box corresponding to the previous enhanced point cloud feature map to obtain the relative position relationship between the candidate targets corresponding to each point cloud candidate prediction box, and / or remove duplicate point cloud candidate prediction boxes; or, before extracting features from the enhanced point cloud feature map, perform multi-head self-attention calculation on the features corresponding to the image candidate prediction box corresponding to the previous enhanced image feature map to obtain the relative position relationship between the candidate targets corresponding to each image candidate prediction box, and / or remove duplicate image candidate prediction boxes.

[0069] When adopting the 3D object detection model as shown in Figure 1B , the above process can be implemented by the decoder of the 3D object detection model.

[0070] In one example, the decoder may include multiple decoding layers, and adjacent decoding layers are used to decode different types of feature maps, where different types of feature maps include enhanced point cloud feature maps and enhanced image feature maps. The enhanced point cloud feature map and the enhanced image feature map are feature maps obtained by the foregoing method.

[0071] Exemplarily, the structure of the decoder is as shown in Figure 2E , and its decoding layers are respectively denoted as θ (0) , θ (1) , …… θ (t) . Among them, the first layer θ (0)The structure of the decoding layer in the standard Transformer decoder is adopted. To better focus on the local regions where the targets may exist, for the cross-attention calculation used in this decoding layer, in the embodiments of this application, the point cloud feature map before the feature interaction of the aforementioned encoder, that is, before the feature enhancement, such as the bird's-eye view BEV feature map h before the feature enhancement, is used p Each pixel in is used as K and V in the cross-attention calculation.

[0072] To accelerate convergence, the embodiments of this application also use a query initialization strategy that depends on the input. Specifically, the point cloud feature maps h before and after enhancement are used p and h' p are used to predict two heatmaps respectively, where the k-th channel at each position represents the probability that the center of the k-th type of object exists at the current position. After adding the two heatmaps, the m positions with the largest values are taken as the initial Q (query) coordinates. Q is encoded as the sum of the feature and the class vector at this position, which is shown as q in the figure p . Among them, m is an integer greater than or equal to 2. init .

[0073] The subsequent 2×N decoding layers of this decoder can be regarded as composed of N cascaded basic units, where each basic unit contains two decoding layers for query (Q)-feature dynamic interaction (that is Figure 2E the θ in (i) ), therefore, the decoding layer in this example is also called the query-feature dynamic interaction layer. Each query-feature dynamic interaction layer sequentially extracts features from h' p and h' c to improve the expression of the Q vector and generate the prediction box b i for each layer.

[0074] Among them, the structure of each query-feature dynamic interaction layer (that is θ (i) ) is as shown in Figure 2F . Its input is h', which is the feature map after feature enhancement of the previous layer. If the previous layer is used to decode the image feature map, then h' is h' c , if the previous layer is used to decode the point cloud feature map, then h' is h' p . b i-1 is the prediction box output by the previous layer, and q i-1 is the q (query) vector output by the previous layer.

[0075] Based on this, the specific feature interaction process of each query-feature dynamic interaction layer is as follows:

[0076] The first step: First, before performing feature interaction, each query-feature dynamic interaction layer first processes a group of query vectors output by the previous layer, that is, qi-1 Perform multi-head self-attention calculation, Figure 2F denoted as MHSA in the figure, to infer the relative positional relationship between the corresponding targets of each predicted bounding box b i-1 and eliminate duplicate predicted bounding boxes. The output of MHSA is added to q i-1 and normalized using the LayerNorm method (denoted as Add&Norm after MHSA in the figure) to obtain the updated query vector q i-1 .

[0077] Second step: Subsequently, for the updated query vector q i-1 , project the predicted bounding box b i-1 of the three-dimensional target decoded from the previous layer onto the image feature map h′ of the corresponding modality at this layer to obtain a two-dimensional predicted bounding box on this feature map, and use the RoIAlign method (a method for mapping the generated predicted bounding box into a feature map of a fixed size) to extract the RoI (region of interest) feature R of size 7×7 i . Specifically, use the maximum circumscribed rectangle of the two-dimensional convex polygon obtained by projecting the predicted bounding box of the three-dimensional target as the RoI. Since in scenarios such as autonomous driving, the target usually has a relatively small scale on the point cloud feature map, such as the BEV feature map, when processing point cloud features, the predicted bounding box of the three-dimensional target is enlarged by a factor of two before projection.

[0078] Third step: Subsequently, perform DynConv (dynamic convolution) processing on q i-1 updated in the first step and R i obtained in the second step. Specifically, q i-1 updated in the first step is mapped to two groups of 1×1 convolutional kernels, and these two groups of convolutional kernels are successively used to perform convolution on the RoI feature R i . After the convolution process, the RoI feature R i is flattened and dimensionally reduced to obtain an output with the same shape as q i-1 . Subsequently, this output is added to q i-1 updated in the first step and normalized using the LayerNorm method (denoted as Add&Norm after DynConv in the figure) to form the query vector q i-1 fused with the RoI feature.

[0079] Fourth step: Use a two-layer feed-forward neural network to update the query vector q i-1 fused with the RoI feature. The updated q i-1 and the feed-forward neural network ( Figure 2FThe inputs shown as FFN) are added together and normalized using the LayerNorm method (shown as Add&Norm in the lower right corner of the figure), and then the q output by the query-feature dynamic interaction layer is obtained. i 。

[0080] Step 5: At the end of each query-feature dynamic interaction layer, a two-layer feed-forward network ( Figure 2F shown as CLS® in the figure) independently decodes the output q i to obtain the improved prediction box b of this layer i 。

[0081] The prediction box decoded by the feed-forward network after the last query-feature dynamic interaction layer will be output as the final detection box of the final three-dimensional target, that is, the three-dimensional target detection result.

[0082] It should be noted that the above decoder can be implemented not only by code, but also as a decoder chip or other hardware forms (such as FPGA, etc.) that implement the functions of the above decoder through logic circuits. And, the above encoder-decoder architecture is similar, which can be implemented not only by code, but also in the form of an encoder chip - decoder chip or other hardware forms such as FPGA.

[0083] Through the above encoder-decoder architecture, the solution of the embodiment of the present application can perform multi-modal feature fusion and use the fused features to accurately detect three-dimensional targets.

[0084] In addition, during training, for the decoder, the Hungarian loss calculated based on the set of actual prediction boxes and the set of prediction boxes predicted by the model can be optimized for each decoding layer. And, in addition to this, the heatmap prediction loss (such as using Gaussian focal loss) can also be optimized to supervise query initialization. The final loss is the sum of the above Hungarian loss and heatmap loss. In addition, the Adam optimizer can be used to optimize the overall training of the three-dimensional target detection model, and the one-cycle learning rate strategy can be adopted.

[0085] Next, taking a specific scenario as an example, the above process will be described exemplarily, as Figure 2G shown.

[0086] Figure 2G In this example, taking the autonomous driving scenario as an example, it is assumed that a lidar and a camera are installed in an autonomous driving vehicle. The lidar collects the three-dimensional point cloud of the three-dimensional environment where the autonomous driving vehicle is located, and the camera collects the image of the environment where the autonomous driving vehicle is located.

[0087] The collected three-dimensional point cloud and image are input as Figure 1BThe three-dimensional object detection model shown extracts three-dimensional point cloud features through a feature extractor for three-dimensional point clouds to generate a point cloud feature map, and extracts image features through a feature extractor for images to generate an image feature map. Then, both the point cloud feature map and the image feature map are input into an encoder, and the encoder performs feature interaction to obtain an enhanced image feature map and an enhanced point cloud feature map. Next, the enhanced image feature map and the enhanced point cloud feature map output by the encoder are input into a decoder, and the decoder alternately decodes the enhanced image feature map and the enhanced point cloud feature map, and finally outputs the detection results of each three-dimensional object, that is, the detection boxes of each three-dimensional object and the information of the corresponding three-dimensional object, such as the category, position, size, orientation, etc. of the three-dimensional object. Further, based on the three-dimensional object detection results, driving planning can be performed for an autonomous driving vehicle.

[0088] In this embodiment, when performing three-dimensional object detection, detection is performed based on two aspects of modal information of the three-dimensional environment to be detected, namely images and three-dimensional point clouds. And during the detection process, feature interaction is performed between the image feature map corresponding to the image and the point cloud feature map corresponding to the three-dimensional point cloud, so that the modal information of one aspect can be effectively fused into the modal information of the other aspect, enabling the modal information of the two aspects to complement each other, thereby realizing feature enhancement of the incorporated modal information, and thus obtaining the corresponding enhanced image feature map and enhanced point cloud feature map. Based on the enhanced image feature map and the enhanced point cloud feature map, three-dimensional object detection is then performed. Since the corresponding information is more comprehensive, rich, and characteristic, the three-dimensional object detection is more accurate and efficient, and high-quality three-dimensional object detection results can be obtained.

[0089] Refer to Figure 3 , which shows a schematic structural diagram of an electronic device according to an embodiment of the present application. The specific implementation of the electronic device in the specific embodiment of the present application is not limited.

[0090] As Figure 3 shown, the electronic device may include: a processor 302, a communication interface 304, a memory 306, and a communication bus 308.

[0091] Among them:

[0092] The processor 302, the communication interface 304, and the memory 306 complete communication with each other through the communication bus 308.

[0093] The communication interface 304 is used to communicate with other electronic devices or servers.

[0094] A processor 302 is configured to execute a program 310, and specifically can execute the relevant steps in the embodiments of the above three-dimensional object detection method.

[0095] Specifically, the program 310 may include program code, and the program code includes computer operation instructions.

[0096] The processor 302 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0097] A memory 306 is configured to store the program 310. The memory 306 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0098] The program 310 is specifically configured to cause the processor 302 to execute the operations corresponding to the three-dimensional object detection method described in the foregoing method embodiments.

[0099] For the specific implementation of each step in the program 310, reference may be made to the corresponding steps and descriptions in the corresponding units in the foregoing method embodiments, and they have corresponding beneficial effects, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above may refer to the corresponding process descriptions in the foregoing method embodiments, and will not be elaborated here.

[0100] The embodiments of the present application also provide a computer program product, including computer instructions, and the computer instructions instruct a computing device to execute the operations corresponding to the three-dimensional object detection method in the foregoing method embodiments.

[0101] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present application may be split into more components / steps, or two or more components / steps or partial operations of components / steps may be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0102] The method according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored on such a recording medium and processed by software using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0103] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such an implementation should not be considered to exceed the scope of the embodiments of the present application.

[0104] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application, and the patent protection scope of the embodiments of the present application shall be defined by the claims.

Claims

1. A three-dimensional object detection method, comprising: Obtain an image and a 3D point cloud of the 3D environment to be detected, as well as an image feature map corresponding to the image and a point cloud feature map corresponding to the 3D point cloud; Perform feature interaction between the image feature map and the point cloud feature map, and obtain an enhanced image feature map fused with point cloud features and an enhanced point cloud feature map fused with image features according to the feature interaction result; Based on the enhanced image feature map and the enhanced point cloud feature map, perform 3D object detection in the 3D environment; The performing feature interaction between the image feature map and the point cloud feature map, and obtaining an enhanced image feature map fused with point cloud features and an enhanced point cloud feature map fused with image features according to the feature interaction result includes: For the image feature map, perform a first feature interaction from image features to image features and a second feature interaction from point cloud features to image features; according to the results of the first feature interaction and the second feature interaction, obtain an enhanced image feature map fused with point cloud features; And, For the point cloud feature map, perform a third feature interaction from point cloud features to point cloud features and a fourth feature interaction from image features to point cloud features; according to the results of the third feature interaction and the fourth feature interaction, obtain an enhanced point cloud feature map fused with image features.

2. The method according to claim 1, wherein, The second feature interaction from point cloud features to image features includes: Perform a projection transformation on the point cloud features corresponding to the position points in the image feature map to convert them into image features.

3. The method according to claim 2, wherein, The performing a projection transformation on the point cloud features corresponding to the position points in the image feature map to convert them into image features includes: Project the 3D point cloud into a bird's-eye view with depth information; According to the positions of the pixel points in the bird's-eye view, obtain the point cloud features at the corresponding positions in the point cloud feature map; Perform a projection transformation on the obtained point cloud features to convert them into image features.

4. The method according to claim 1, wherein, The fourth feature interaction from image features to point cloud features includes: For a certain pixel point in the point cloud feature map, obtain multiple reference pixel points of the pixel point in the point cloud feature map; Based on the positions of the pixel point and the multiple reference pixel points, obtain multiple image features at the corresponding positions from the image feature map; Fuse the point cloud features corresponding to the pixel point and the obtained multiple image features.

5. The method according to claim 4, wherein, The obtaining multiple image features at the corresponding positions from the image feature map based on the positions of the pixel point and the multiple reference pixel points includes: Based on the positions of the pixel point and the multiple reference pixel points, perform a projection transformation to obtain corresponding two-dimensional coordinates; Based on the two-dimensional coordinates, perform feature acquisition at the corresponding positions on the image feature map to obtain corresponding multiple image features.

6. The method according to claim 1, wherein, The performing 3D object detection in the 3D environment based on the enhanced image feature map and the enhanced point cloud feature map includes: Perform feature extraction on the enhanced image feature map and the enhanced point cloud feature map alternately in sequence, and perform 3D object detection in the 3D environment based on the extracted features.

7. The method according to claim 6, wherein, Before performing feature extraction on the enhanced image feature map and the enhanced point cloud feature map alternately in sequence, the method further includes: respectively performing three-dimensional target position prediction based on the point cloud feature map and the enhanced point cloud feature map to obtain corresponding first prediction result and second prediction result; obtaining a preset number of candidate prediction boxes according to the first prediction result and the second prediction result, where the probability that a candidate prediction box belongs to a certain type of target is greater than a preset probability. Performing feature extraction on the enhanced image feature map and the enhanced point cloud feature map alternately in sequence, and performing three-dimensional target detection in the three-dimensional environment based on the extracted features, includes: performing feature extraction on the enhanced image feature map and the enhanced point cloud feature map alternately in sequence based on the information of the candidate prediction boxes, and performing three-dimensional target detection in the three-dimensional environment based on the extracted features.

8. The method according to claim 7, wherein, The first prediction result and the second prediction result are respectively a first prediction heat map and a second prediction heat map; the k-th channel in the first prediction heat map and the second prediction heat map represents the probability that a k-th type of target exists at the current position. The information of the candidate prediction boxes includes the point cloud features corresponding to the candidate prediction boxes and the information of the target category vectors corresponding to the candidate prediction boxes.

9. The method according to claim 7 or 8, wherein, Performing feature extraction on the enhanced image feature map and the enhanced point cloud feature map alternately in sequence, and performing three-dimensional target detection in the three-dimensional environment based on the extracted features, includes: When performing feature extraction on the enhanced image feature map, obtaining the point cloud candidate prediction boxes corresponding to the previous enhanced point cloud feature map of this enhanced image feature map; projecting the point cloud candidate prediction boxes onto this enhanced image feature map to obtain corresponding image candidate prediction boxes. When performing feature extraction on the enhanced point cloud feature map, obtaining the image candidate prediction boxes corresponding to the previous enhanced image feature map of this enhanced point cloud feature map, and converting the image candidate prediction boxes to this enhanced point cloud feature map to obtain corresponding point cloud candidate prediction boxes.

10. The method according to claim 9, wherein, Before performing feature extraction on the enhanced image feature map, performing multi-head self-attention calculation on the features corresponding to the point cloud candidate prediction boxes corresponding to the previous enhanced point cloud feature map to obtain the relative position relationship between the candidate targets corresponding to each point cloud candidate prediction box, and / or removing duplicate point cloud candidate prediction boxes. Or, Before performing feature extraction on the enhanced point cloud feature map, performing multi-head self-attention calculation on the features corresponding to the image candidate prediction boxes corresponding to the previous enhanced image feature map to obtain the relative position relationship between the candidate targets corresponding to each image candidate prediction box, and / or removing duplicate image candidate prediction boxes.

11. According to the method of claim 9, wherein, The projecting the point cloud candidate prediction boxes onto this enhanced image feature map to obtain corresponding image candidate prediction boxes includes: Projecting the point cloud candidate prediction boxes onto this enhanced image feature map after magnifying them by a preset multiple. Obtaining the region corresponding to the projection result on this enhanced image feature map. Performing feature extraction on the features of the region, and obtaining corresponding image candidate prediction boxes according to the feature extraction result.

12. An encoder, wherein, The encoder includes multiple encoding layers. Each encoding layer includes: A feature input part for receiving the image feature map and the point cloud feature map output by the previous layer; A feature interaction part for respectively performing a first feature interaction between image features and a second feature interaction between point cloud features and image features on the image feature map, and performing a third feature interaction between point cloud features and a fourth feature interaction between image features and point cloud features on the point cloud feature map; A feature fusion part for performing feature fusion based on the first feature interaction result and the second feature interaction result to obtain an enhanced image feature map, and performing feature fusion based on the third feature interaction result and the fourth feature interaction result to obtain an enhanced point cloud feature map.

13. A decoder, wherein, The decoder includes a plurality of decoding layers; Adjacent decoding layers are used to decode different types of feature maps, where the different types of feature maps include the enhanced point cloud feature map and the enhanced image feature map; The enhanced image feature map is: for the image feature map corresponding to the three-dimensional environment to be detected, performing a first feature interaction from image features to image features and a second feature interaction from point cloud features to image features; according to the results of the first feature interaction and the second feature interaction, obtaining an enhanced image feature map that fuses point cloud features; the image feature map is obtained by performing feature extraction on the image corresponding to the three-dimensional environment; and, The enhanced point cloud feature map is: for the point cloud feature map corresponding to the three-dimensional environment to be detected, performing a third feature interaction from point cloud features to point cloud features and a fourth feature interaction from image features to point cloud features; according to the results of the third feature interaction and the fourth feature interaction, obtaining an enhanced point cloud feature map that fuses image features; the point cloud feature map is obtained by performing feature extraction on the three-dimensional point cloud corresponding to the three-dimensional environment.

Citation Information

Patent Citations

  • Target detection method and device, electronic equipment and storage medium

    CN113554643A

  • Learning method and learning device for integrating image acquired by camera and point-cloud map acquired by radar or LiDAR corresponding to image at each of convolution stages in neural network and testing method and testing device using the same

    US10408939B1