Obstacle detection method and device, equipment, storage medium and program product
By extracting and fusion of image data and point cloud data, the problem of insufficient information fusion in the prior art is solved, and more comprehensive obstacle detection and higher detection accuracy are achieved.
Patent Information
- Application Number
- CN202510091803.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, only a single representation space is selected for information fusion, resulting in insufficient information exploration or interference from irrelevant information.
By acquiring image data and point cloud data, the first bird's eye view feature map and the second bird's eye view feature map are extracted respectively, and the global and local features are characterized by fusion to determine the location range and category of obstacles.
It realizes the integration of long-range information in dense space and short-range information in sparse space to avoid the problems of insufficient information exploration or interference from irrelevant information, and improves the understanding and detection accuracy of complex scenarios.
Smart Images

Figure CN120088758A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to an obstacle detection method, device, equipment, storage medium and program product. Background Art
[0002] 3D object detection aims to accurately predict key information such as the position, size, category, direction, and speed of obstacles around the ego vehicle, and it plays a crucial role in the fields of autonomous driving and robotics. In recent years, multi-modal 3D object detection has attracted increasing attention due to its excellent robustness and high performance.
[0003] Among numerous sensors, cameras and lidars are two of the most common types. Cameras can provide rich and detailed semantic information, endowing in-depth connotations for scene understanding; while lidars can accurately capture the structural details of objects. Their advantages complement each other. However, since input data of different modalities usually have heterogeneous data representation forms, an effective fusion strategy has become the core problem in the field of multi-modal 3D object detection.
[0004] In the prior art, most methods project features of different modalities into the same representation space, then explore the corresponding relationships of cross-modal features in this space, and finally fuse the aligned features with an encoder. However, only selecting a single representation space for information fusion is likely to cause problems such as insufficient information exploration or interference from irrelevant information. Summary of the Invention
[0005] The present invention provides an obstacle detection method, device, equipment, storage medium and program product to solve the problems in the prior art that only selecting a single representation space for information fusion is likely to cause problems such as insufficient information exploration or interference from irrelevant information.
[0006] The present invention provides an obstacle detection method, including: obtaining a first bird's-eye view feature map of image data, and obtaining a second bird's-eye view feature map of point cloud data; determining a first global feature and a first local feature of the first bird's-eye view feature map, and determining a second global feature and a second local feature of the second bird's-eye view feature map; fusing the first global feature and the second global feature to obtain a third global feature, and fusing the first local feature and the second local feature to obtain a third local feature; determining obstacle information according to the third global feature and the third local feature; wherein, the obstacle information includes the position range of the obstacle and the category of the obstacle.
[0007] A method for obstacle detection provided by the present invention, the obtaining of the first bird's-eye view feature map of the image data includes: inputting the image data into a two-dimensional backbone network to obtain the first bird's-eye view feature map; wherein, the two-dimensional backbone network includes an image encoder and an image-bird's-eye view transformation module, and the image-bird's-eye view transformation module includes a depth prediction network and a perspective transformation module; the image encoder is used to extract the image features of the image data; the image-bird's-eye view transformation module is used to extract the depth information and semantic features of the image features through the depth prediction network, combine the depth information with the semantic features to form a frustum feature, and convert the frustum feature into the first bird's-eye view feature map through the perspective transformation module.
[0008] A method for obstacle detection provided by the present invention, the obtaining of the second bird's-eye view feature map of the point cloud data includes: inputting the point cloud data into a three-dimensional backbone network to obtain a voxel feature; performing a compression operation on the voxel feature along the Z axis to obtain the second bird's-eye view feature map.
[0009] A method for obstacle detection provided by the present invention, the determining of the first global feature and the first local feature of the first bird's-eye view feature map includes: inputting the first bird's-eye view feature map into a distance-based feature decomposition network to obtain the first global feature and the first local feature; wherein, the distance-based feature decomposition network includes a long-distance encoder and a short-distance encoder, and the distance-based feature decomposition network is used to decompose and extract features from different distance scales.
[0010] A method for obstacle detection provided by the present invention, the feature fusion of the first global feature and the second global feature to obtain a third global feature includes: inputting the first global feature and the second global feature into a distance-based fusion network for feature fusion to obtain the third global feature; wherein, the distance-based fusion network includes a long-distance fusion network and a short-distance fusion network, and the distance-based fusion network is used to fuse multi-modal features from different distance scales.
[0011] According to an obstacle detection method provided by the present invention, determining obstacle information based on the third global feature and the third local feature includes: inputting the third global feature and the third local feature into a split prediction network to obtain the obstacle information; wherein, the split prediction network includes a class prediction branch and a detection box prediction branch; the class prediction branch is used to fuse the local features for classification prediction in the third global feature and the third local feature, and input the fused features into a classification prediction head to obtain the class of the obstacle; the detection box prediction branch is used to fuse the local features for detection box prediction in the third global feature and the third local feature, and input the fused features into a detection box prediction head to obtain the position range of the obstacle.
[0012] The present invention also provides an obstacle detection device, including the following modules: an acquisition module and a processing module; the acquisition module is used to acquire a first bird's-eye view feature map of image data and a second bird's-eye view feature map of point cloud data; the processing module is used to determine a first global feature and a first local feature of the first bird's-eye view feature map, and determine a second global feature and a second local feature of the second bird's-eye view feature map; fuse the first global feature and the second global feature to obtain a third global feature, and fuse the first local feature and the second local feature to obtain a third local feature; determine obstacle information based on the third global feature and the third local feature; wherein, the obstacle information includes the position range of the obstacle and the class of the obstacle.
[0013] According to an obstacle detection device provided by the present invention, the acquisition module is used to input the image data into a two-dimensional backbone network to obtain the first bird's-eye view feature map; wherein, the two-dimensional backbone network includes an image encoder and an image-bird's-eye view transformation module, and the image-bird's-eye view transformation module includes a depth prediction network and a perspective transformation module; the image encoder is used to extract the image features of the image data; the image-bird's-eye view transformation module is used to extract the depth information and semantic features of the image features through the depth prediction network, combine the depth information with the semantic features to form a frustum feature, and convert the frustum feature into the first bird's-eye view feature map through the perspective transformation module.
[0014] According to an obstacle detection device provided by the present invention, the acquisition module is used to input the point cloud data into a three-dimensional backbone network to obtain a voxel feature; perform a compression operation on the voxel feature along the Z axis to obtain the second bird's-eye view feature map.
[0015] An obstacle detection device provided according to the present invention, the processing module is configured to input the first bird's-eye view feature map into a distance-based feature decomposition network to obtain the first global feature and the first local feature; wherein, the distance-based feature decomposition network includes a long-distance encoder and a short-distance encoder, and the distance-based feature decomposition network is configured to decompose and extract features from different distance scales.
[0016] An obstacle detection device provided according to the present invention, the processing module is configured to input the first global feature and the second global feature into a distance-based fusion network for feature fusion to obtain a third global feature; wherein, the distance-based fusion network includes a long-distance fusion network and a short-distance fusion network, and the distance-based fusion network is configured to fuse multi-modal features from different distance scales.
[0017] An obstacle detection device provided according to the present invention, the processing module is configured to input the third global feature and the third local feature into a separate prediction network to obtain the obstacle information; wherein, the separate prediction network includes a class prediction branch and a detection box prediction branch; the class prediction branch is configured to fuse the local features for classification prediction in the third global feature and the third local feature, and input the fused features into a classification prediction head to obtain the class of the obstacle; the detection box prediction branch is configured to fuse the local features for detection box prediction in the third global feature and the third local feature, and input the fused features into a detection box prediction head to obtain the position range of the obstacle.
[0018] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the obstacle detection method as described in any one of the above.
[0019] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the obstacle detection method as described in any one of the above.
[0020] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the obstacle detection method as described in any one of the above.
[0021] The obstacle detection method, device, equipment, storage medium and program product provided by the present invention can perform feature fusion on two different modalities of global features to obtain a third global feature, perform feature fusion on two different modalities of local features to obtain a third local feature, and determine the position range and category of the obstacle according to the third global feature and the third local feature. Since long-range information can be fused in a dense space and short-range information can be fused in a sparse space, on the one hand, problems such as insufficient information exploration or interference from irrelevant information can be avoided, so that the perception of the environment is more comprehensive, thereby improving the understanding ability of complex scenes; on the other hand, a receptive field matching the prediction task can be designed, so as to achieve targeted feature fusion, thereby improving the fusion effect and detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0023] Figure 1 is one of the flow diagrams of the obstacle detection method provided by the present invention; Figure 2 is the second flow diagram of the obstacle detection method provided by the present invention; Figure 3 is the structural diagram of the distance-based fusion network provided by the present invention; Figure 4 is the structural diagram of the obstacle detection device provided by the present invention; Figure 5 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To make the objectives, technical solutions and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the present application with reference to the drawings in the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0025] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0026] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device including that element. In addition, it should be noted that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0027] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order.
[0028] Some exemplary embodiments are described in the embodiments of the present application for the purpose of illustration. It should be understood that the present application can be implemented in other ways not specifically shown in the drawings.
[0029] As Figure 1 shown, the embodiments of the present application provide an obstacle detection method, which can be applied to an obstacle detection device. The obstacle detection method may include S101 - S104: S101. The obstacle detection device acquires a first bird's-eye view feature map of the image data and a second bird's-eye view feature map of the point cloud data.
[0030] Optionally, the above-mentioned image data may be multi-view images, and the above-mentioned point cloud data may be the point cloud of a lidar.
[0031] Optionally, as Figure 2As shown in the figure, the obstacle detection device obtains the first bird's-eye view feature map of the image data, including: inputting the image data into a two-dimensional backbone network to obtain the first bird's-eye view feature map; wherein, the two-dimensional backbone network includes an image encoder and an image-bird's-eye view transformation module, and the image-bird's-eye view transformation module includes a depth prediction network and a perspective transformation module; the image encoder is used to extract the image features of the image data; the image-bird's-eye view transformation module is used to extract the depth information and semantic features of the image features through the depth prediction network, combine the depth information with the semantic features to form a frustum feature, and convert the frustum feature into the first bird's-eye view feature map through the perspective transformation module.
[0032] Specifically, the image data can be expressed as , where B represents the batch size, H represents the image height, that is, the number of pixels in the vertical direction of the image, which determines the vertical size of the image. W represents the image width, that is, the number of pixels in the horizontal direction of the image, which determines the horizontal size of the image. C represents the color channel of the image.
[0033] The two-dimensional backbone network includes an image encoder and an image-bird's-eye view transformation module. The function of the image encoder is to encode the input image data and extract some feature information in the image data. The image-bird's-eye view transformation module includes a depth prediction network and a perspective transformation module. The depth prediction network is composed of multiple layers of convolutional neural networks, and it can predict the depth according to the input information. Depth is very important in 3D scene understanding, as it helps us know the distance of objects in the scene from the observation point. Moreover, the depth predicted by this network will perform a dot product operation with the semantic features to obtain the frustum feature. The semantic features contain the category information of different objects or regions in the image. By performing the dot product operation, the depth information and semantic information are combined to form the frustum feature for better description of the 3D scene. The perspective transformation module can, together with BEVPool sampling, convert the input frustum feature into the bird's-eye view feature map of the image. The bird's-eye view feature map is a feature representation of the scene from a top-down perspective.
[0034] The first bird's-eye view feature map is expressed as . Here, represents the height of the feature map, which may be different from the height H of the image data because, after a series of transformations of the network, its size may have changed. is the width of the feature map. Similarly, it may also be different from the width W of the image data. is the number of channels of the feature map. The number of channels reflects the richness of the feature map. Different channels can represent different types of features, such as different texture features, color features, or different aspects of semantic features, etc.
[0035] Optionally, as Figure 2 shown, the obstacle detection device obtains a second bird's-eye view feature map of the point cloud data, including: inputting the point cloud data into a three-dimensional backbone network to obtain voxel features; performing a compression operation on the voxel features along the Z axis to obtain the second bird's-eye view feature map; wherein, the three-dimensional backbone network includes a point cloud encoder for extracting features from the point cloud data.
[0036] Specifically, the point cloud data can be represented as , where N represents the number of points, that is, the number of points detected by the lidar. K represents the feature dimension of the points, which means that each point has K features.
[0037] The three-dimensional backbone network is mainly composed of a point cloud encoder. The role of the point cloud encoder is to encode the point cloud data of the lidar. Since the point cloud data is a discrete and irregular data structure, the point cloud encoder will convert these discrete point cloud data into more regular and easier-to-process voxel features. After obtaining the voxel features, the voxel features will be compressed along the Z axis to obtain the second bird's-eye view feature map. Compressing along the Z axis means compressing the three-dimensional voxel features in the vertical direction to obtain a two-dimensional bird's-eye view feature map.
[0038] The second bird's-eye view feature map can be represented as , where represents the height of the feature map, which reflects the size of the bird's-eye view feature map in the vertical direction. represents the width of the feature map, corresponding to the size in the horizontal direction. And represents the number of channels of the feature map.
[0039] S102. The obstacle detection device determines the first global feature and the first local feature of the first bird's-eye view feature map, and determines the second global feature and the second local feature of the second bird's-eye view feature map.
[0040] Optionally, as Figure 2As shown, the obstacle detection device determines the first global feature and the first local feature of the first bird's-eye view feature map, including: inputting the first bird's-eye view feature map into a distance-based feature decomposition network to obtain the first global feature and the first local feature; wherein, the distance-based feature decomposition network includes a long-distance encoder and a short-distance encoder, and the distance-based feature decomposition network is used to decompose and extract features from different distance scales.
[0041] Specifically, the distance-based feature decomposition network includes a long-distance encoder and a short-distance encoder. Among them, the long-distance encoder consists of multiple deformable attention layers and can extract long-distance context information. The deformable attention layer is an advanced attention mechanism. Different from the traditional attention mechanism, it can dynamically adjust the attention position according to the input features, so as to capture long-distance information more flexibly. For example, when processing the bird's-eye view feature map, the long-distance encoder can associate feature elements that are far apart and consider the relationship between them. The short-distance encoder consists of multiple convolutional layers and is mainly used to extract local feature information. The convolutional layer can capture local features by sliding the window. For the bird's-eye view feature map, the short-distance encoder can focus on the features within a local area, such as the texture, color, or shape features within a small area. Through the distance-based feature decomposition network, the feature information at different distances can be processed more carefully to achieve more accurate and effective feature fusion.
[0042] It should be noted that when processing the first bird's-eye view feature map and the second bird's-eye view feature map, the distance-based feature decomposition network has the same structure but different parameters.
[0043] It should be noted that the global feature is a comprehensive description of the bird's-eye view feature map, including the characteristics of the image or point cloud data at the overall level, and can reflect the overall trend and main information of the data. For example, for image data, the first global feature can be shape features, texture features, color distribution features, and for point cloud data, the second global feature can be spatial distribution features and density features. The local feature focuses on describing the characteristics of local regions in the image or point cloud data and can capture the detailed information in the data. For example, for image data, the first local feature can be edge features, local texture features, local color features, and for point cloud data, the second local feature can be local geometric shape features and local point cloud density change features.
[0044] S103. The obstacle detection device fuses the first global feature and the second global feature to obtain a third global feature, and fuses the first local feature and the second local feature to obtain a third local feature.
[0045] Optionally, as Figure 2As shown, the obstacle detection device performs feature fusion on the first global feature and the second global feature to obtain a third global feature, including: inputting the first global feature and the second global feature into a distance-based fusion network for feature fusion to obtain a third global feature; wherein, the distance-based fusion network includes a long-distance fusion network and a short-distance fusion network, and the distance-based fusion network is used to fuse multi-modal features at different distance scales.
[0046] Specifically, as Figure 3 shown, the distance-based fusion network includes a long-distance fusion network and a short-distance fusion network: The input of the long-distance fusion network is the global feature of multiple modalities. In the embodiment of the present application, the input of the long-distance fusion network is the first global feature and the second global feature, and the output is the global feature of the obstacle.
[0047] The long-distance fusion network occurs in a dense feature space (bird's-eye view space) and consists of multi-scale information transmission, three-path feature fusion, a candidate extraction module, and a cross-attention mechanism.
[0048] The multi-scale information transmission can first sample the original feature map to multiple different sizes for spatial feature fusion, and then sample it back to the original size for channel feature fusion. This operation can capture feature information at different scales. By sampling and fusing the feature map at different scales, it is possible to better utilize the information at different resolutions, which helps to extract more comprehensive features.
[0049] The three-path feature fusion consists of parallel Squeeze-and-Excitation (SE) Bloc, a convolutional module, and Atrous Spatial Pyramid Pooling. The Squeeze-and-Excitation (SE) Block can recalibrate the relationship between channels. By learning the correlation between channels, different weights are assigned to different channels, enabling the network to pay more attention to important channels, thereby enhancing the feature representation ability. The convolutional module slides the convolutional kernel on the feature map to extract local feature information and further transform and process the features. Atrous Spatial PyramidPooling can perform feature extraction at different sampling rates, expand the receptive field without reducing the spatial resolution, and capture context information of different sizes. Through the parallel processing and fusion of these three modules, a fused bird's-eye view feature map can be obtained.
[0050] The candidate extraction module can predict the heat distribution of the appearance of obstacles based on the fused bird's-eye view feature map. The heat distribution can reflect which areas are more likely to have obstacles.
[0051] In the cross-attention mechanism, the query is the obstacle candidate feature, the Key and Value are the fused bird's-eye view feature maps, and the output feature is the general obstacle feature. Through this cross-attention mechanism, the network can screen and focus on more valuable information from the fused bird's-eye view feature maps according to the obstacle candidate features, and further extract the general features related to the obstacles to help better understand and describe the obstacles.
[0052] The input of the short-distance fusion network is the multi-modal local features. In the embodiment of the present application, the input of the short-distance fusion network is the first local feature and the second local feature, and the output is the local feature of the obstacle. The local feature of the obstacle is composed of the local feature for classification prediction and the local feature for detection box prediction.
[0053] The short-distance fusion network occurs in the sparse space (candidate space), including a candidate extraction module and a self-attention module. The candidate extraction module directly reuses the indexes of the obstacle candidates in the long-distance fusion network. This can ensure the consistency of the two networks when processing obstacle candidates, and at the same time reduce the computational amount. By reusing the indexes of the obstacle candidates, the obstacle candidate features of the image and the obstacle candidate objects of the lidar can be obtained quickly.
[0054] After that, the obstacle candidate features of the image and the obstacle candidate objects of the lidar are fused through the self-attention module to obtain the local feature for classification prediction and the local feature for detection box prediction. The self-attention module can perform weighted processing on different features according to the correlation between the features, highlight the more important features, and make the finally obtained local features more suitable for classification and detection box prediction.
[0055] S104. The obstacle detection device determines obstacle information according to the third global feature and the third local feature.
[0056] Wherein, the obstacle information includes the position range of the obstacle and the category of the obstacle.
[0057] Optionally, as Figure 2As shown, the obstacle detection device determines obstacle information based on the third global feature and the third local feature, including: inputting the third global feature and the third local feature into a separate prediction network to obtain the obstacle information; wherein, the separate prediction network includes a class prediction branch and a detection box prediction branch; the class prediction branch is used to fuse the local features for classification prediction in the third global feature and the third local feature, and input the fused features into a classification prediction head to obtain the class of the obstacle; the detection box prediction branch is used to fuse the local features for detection box prediction in the third global feature and the third local feature, and input the fused features into a detection box prediction head to obtain the position range of the obstacle.
[0058] Specifically, the separate prediction network includes a class prediction branch and a detection box prediction branch. This separate design helps to separately process different prediction tasks (class prediction and detection box prediction), improving the scalability and pertinence of the network.
[0059] The class prediction branch can fuse the global feature of the obstacle with the local features for classification prediction, and then input the fused features into the classification prediction head to predict the class of the obstacle.
[0060] The detection box prediction branch can fuse the global feature of the obstacle with the local features for detection box prediction, and input the fused features into the detection box prediction head to predict the detection box of the obstacle.
[0061] Finally, by combining the prediction results of the two branches, information such as the position, size, orientation, speed, and type of the obstacle can be obtained.
[0062] In the embodiment of the present application, two different modalities of global features can be feature-fused to obtain the third global feature, and two different modalities of local features can be feature-fused to obtain the third local feature, and the position range and the class of the obstacle can be determined based on the third global feature and the third local feature. Since long-range information can be fused in the dense space and short-range information can be fused in the sparse space, on the one hand, problems such as insufficient information exploration or interference from irrelevant information can be avoided, so that the perception of the environment is more comprehensive, and thus the understanding ability of complex scenes is improved; on the other hand, a receptive field can be designed to match the prediction task, so as to achieve targeted feature fusion, and thus the fusion effect and detection accuracy are improved.
[0063] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0064] For the obstacle detection method provided by the embodiments of the present application, the execution subject can be an obstacle detection device, or a control module for obstacle detection in the obstacle detection device. In the embodiments of the present application, taking the obstacle detection device executing the obstacle detection method as an example, the obstacle detection device provided by the embodiments of the present application is described.
[0065] It should be noted that the embodiments of the present application can divide the functional modules of the obstacle detection device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. Optionally, the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0066] As Figure 4 shown, the embodiments of the present application provide an obstacle detection device 400. The obstacle detection device 400 includes: an acquisition module 401 and a processing module 402. The acquisition module 401 is used to acquire a first bird's-eye view feature map of image data and a second bird's-eye view feature map of point cloud data; the processing module 402 is used to determine a first global feature and a first local feature of the first bird's-eye view feature map, and determine a second global feature and a second local feature of the second bird's-eye view feature map; perform feature fusion on the first global feature and the second global feature to obtain a third global feature, and perform feature fusion on the first local feature and the second local feature to obtain a third local feature; determine obstacle information according to the third global feature and the third local feature; wherein, the obstacle information includes the position range of the obstacle and the category of the obstacle.
[0067] Optionally, the obtaining module 401 is configured to input the image data into a two-dimensional backbone network to obtain the first bird's-eye view feature map; wherein, the two-dimensional backbone network includes an image encoder and an image-bird's-eye view transformation module, and the image-bird's-eye view transformation module includes a depth prediction network and a perspective transformation module; the image encoder is configured to extract image features of the image data; the image-bird's-eye view transformation module is configured to extract depth information and semantic features of the image features through the depth prediction network, combine the depth information with the semantic features to form a frustum feature, and convert the frustum feature into the first bird's-eye view feature map through the perspective transformation module.
[0068] Optionally, the obtaining module 401 is configured to input the point cloud data into a three-dimensional backbone network to obtain voxel features; and perform a compression operation on the voxel features along the Z axis to obtain the second bird's-eye view feature map.
[0069] Optionally, the processing module 402 is configured to input the first bird's-eye view feature map into a distance-based feature decomposition network to obtain the first global feature and the first local feature; wherein, the distance-based feature decomposition network includes a long-distance encoder and a short-distance encoder, and the distance-based feature decomposition network is configured to decompose and extract features from different distance scales.
[0070] Optionally, the processing module 402 is configured to input the first global feature and the second global feature into a distance-based fusion network for feature fusion to obtain a third global feature; wherein, the distance-based fusion network includes a long-distance fusion network and a short-distance fusion network, and the distance-based fusion network is configured to fuse multimodal features from different distance scales.
[0071] Optionally, the processing module 402 is configured to input the third global feature and the third local feature into a separated prediction network to obtain the obstacle information; wherein, the separated prediction network includes a class prediction branch and a detection box prediction branch; the class prediction branch is configured to fuse local features for classification prediction in the third global feature and the third local feature, and input the fused features into a classification prediction head to obtain the class of the obstacle; the detection box prediction branch is configured to fuse local features for detection box prediction in the third global feature and the third local feature, and input the fused features into a detection box prediction head to obtain the position range of the obstacle.
[0072] In the embodiments of the present application, the global features of two different modalities can be fused to obtain a third global feature, and the local features of two different modalities can be fused to obtain a third local feature. Then, the position range and category of the obstacle can be determined based on the third global feature and the third local feature. Since long-range information can be fused in the dense space and short-range information can be fused in the sparse space, on the one hand, problems such as insufficient information exploration or interference from irrelevant information can be avoided, so that the perception environment is more comprehensive, and thus the ability to understand complex scenes is improved; on the other hand, a receptive field matching the prediction task can be designed, so as to achieve targeted feature fusion, and thus the fusion effect and detection accuracy are improved.
[0073] Figure 5 An example of the physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute an obstacle detection method, and the method includes: obtaining a first bird's-eye view feature map of image data, and obtaining a second bird's-eye view feature map of point cloud data; determining a first global feature and a first local feature of the first bird's-eye view feature map, and determining a second global feature and a second local feature of the second bird's-eye view feature map; fusing the first global feature and the second global feature to obtain a third global feature, and fusing the first local feature and the second local feature to obtain a third local feature; determining obstacle information based on the third global feature and the third local feature; where the obstacle information includes the position range and category of the obstacle.
[0074] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0075] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the obstacle detection method provided by the above-mentioned various methods. The method includes: obtaining a first bird's-eye view feature map of image data and obtaining a second bird's-eye view feature map of point cloud data; determining a first global feature and a first local feature of the first bird's-eye view feature map, and determining a second global feature and a second local feature of the second bird's-eye view feature map; performing feature fusion on the first global feature and the second global feature to obtain a third global feature, and performing feature fusion on the first local feature and the second local feature to obtain a third local feature; determining obstacle information according to the third global feature and the third local feature; wherein the obstacle information includes the position range of the obstacle and the category of the obstacle.
[0076] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the obstacle detection method provided by the above-mentioned various methods. The method includes: obtaining a first bird's-eye view feature map of image data and obtaining a second bird's-eye view feature map of point cloud data; determining a first global feature and a first local feature of the first bird's-eye view feature map, and determining a second global feature and a second local feature of the second bird's-eye view feature map; performing feature fusion on the first global feature and the second global feature to obtain a third global feature, and performing feature fusion on the first local feature and the second local feature to obtain a third local feature; determining obstacle information according to the third global feature and the third local feature; wherein the obstacle information includes the position range of the obstacle and the category of the obstacle.
[0077] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0078] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An obstacle detection method, characterized in that: include: Acquire a first bird's-eye view feature map of the image data, and acquire a second bird's-eye view feature map of the point cloud data; Determine a first global feature and a first local feature of the first bird's-eye view feature map, and determine a second global feature and a second local feature of the second bird's-eye view feature map; Performing feature fusion on the first global feature and the second global feature to obtain a third global feature, and performing feature fusion on the first local feature and the second local feature to obtain a third local feature; Determine obstacle information according to the third global feature and the third local feature; The obstacle information includes the location range of the obstacle and the category of the obstacle.
2. The obstacle detection method according to claim 1, characterized in that: The step of acquiring a first bird's-eye view feature map of the image data comprises: Inputting the image data into a two-dimensional backbone network to obtain the first bird's-eye view feature map; The two-dimensional backbone network includes an image encoder and an image-bird's-eye view transformation module, and the image-bird's-eye view transformation module includes a depth prediction network and a perspective transformation module; The image encoder is used to extract image features of the image data; The image-bird's-eye view transformation module is used to extract the depth information and semantic features of the image features through the depth prediction network, combine the depth information with the semantic features to form a cone feature, and convert the cone feature into the first bird's-eye view feature map through the perspective transformation module.
3. The obstacle detection method according to claim 1, characterized in that: The step of obtaining a second bird's-eye view feature map of the point cloud data comprises: Inputting the point cloud data into a three-dimensional backbone network to obtain voxel features; The voxel features are compressed along the Z axis to obtain the second bird's-eye view feature map.
4. The obstacle detection method according to claim 1, characterized in that: The determining of the first global feature and the first local feature of the first bird's-eye view feature map comprises: Inputting the first bird's-eye view feature map into a distance-based feature decomposition network to obtain the first global feature and the first local feature; The distance-based feature decomposition network includes a long-distance encoder and a short-distance encoder, and the distance-based feature decomposition network is used to decompose and extract features from different distance scales.
5. The obstacle detection method according to claim 1, characterized in that: The step of fusing the first global feature with the second global feature to obtain a third global feature includes: Inputting the first global feature and the second global feature into a distance-based fusion network for feature fusion to obtain a third global feature; The distance-based fusion network includes a long-distance fusion network and a short-distance fusion network, and the distance-based fusion network is used to fuse multimodal features from different distance scales.
6. The obstacle detection method according to claim 1, characterized in that: The determining the obstacle information according to the third global feature and the third local feature includes: Inputting the third global feature and the third local feature into a separate prediction network to obtain the obstacle information; Wherein, the separated prediction network includes a category prediction branch and a detection box prediction branch; The category prediction branch is used to fuse the third global feature and the local feature used for classification prediction in the third local feature, and input the fused feature into the classification prediction head to obtain the category of the obstacle; The detection box prediction branch is used to fuse the third global feature and the local feature used for detection box prediction in the third local feature, and input the fused feature into the detection box prediction head to obtain the position range of the obstacle.
7. An obstacle detection device, characterized in that: include: Acquisition module and processing module; The acquisition module is used to acquire a first bird's-eye view feature map of the image data and a second bird's-eye view feature map of the point cloud data; The processing module is used to determine a first global feature and a first local feature of the first bird's-eye view feature map, determine a second global feature and a second local feature of the second bird's-eye view feature map; perform feature fusion on the first global feature and the second global feature to obtain a third global feature, perform feature fusion on the first local feature and the second local feature to obtain a third local feature; and determine obstacle information according to the third global feature and the third local feature; The obstacle information includes the location range of the obstacle and the category of the obstacle.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the obstacle detection method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the obstacle detection method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the obstacle detection method according to any one of claims 1 to 6 is implemented.