Obstacle detection method and device, vehicle and storage medium

By multimodal fusion of the image and point cloud data of the vehicle environment, extracting and encoding features, the problem of inaccurate obstacle detection is solved, and more efficient obstacle detection and automatic driving system performance improvement is achieved.

CN120147997APending Publication Date: 2025-06-13HAOMO TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311714483.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional obstacle detection methods have the problem of inaccurate detection, which is difficult to provide more comprehensive and accurate perception capabilities, affecting the performance of autonomous driving systems.

Method used

By obtaining the image and point cloud data of the vehicle environment, image features are extracted and three-dimensional position encoding is performed, and feature extraction is achieved by combining bird's-eye point cloud features to achieve multimodal fusion for obstacle detection.

Benefits of technology

The accuracy of obstacle detection is improved, so that image features and point cloud features are in the same dimension, enhancing the perception and performance of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147997A_ABST
    Figure CN120147997A_ABST
Patent Text Reader

Abstract

The invention provides an obstacle detection method and device, a vehicle and a storage medium, and belongs to the field of automobiles, and the method comprises the steps: obtaining an image of a target environment where the vehicle is located, extracting image features from the image of the target environment, and obtaining a feature map of the target environment, the target environment comprising a to-be-detected obstacle; performing three-dimensional position coding on feature pixel points in the feature map; obtaining aerial view point cloud features of the target environment; according to the aerial view point cloud features, feature extraction is carried out on the feature map of the three-dimensional position coding, and aerial view image features are obtained; and performing obstacle detection on the target environment based on the aerial view point cloud features and the aerial view image features. Therefore, the three-dimensional position codes and the two-dimensional image features are fused, so that the image features and the point cloud features are in the same dimension, the accuracy of extracting the image features is further improved, and obstacles can be detected more accurately when obstacle detection is carried out according to the aerial view image features and the aerial view point cloud features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of automobiles, and particularly relates to an obstacle detection method, device, vehicle, and storage medium. Background Art

[0002] In the automatic driving system of a vehicle, it is necessary to control the driving state of the vehicle according to the obstacles in the driving path of the vehicle. Therefore, obstacle detection is required.

[0003] In the process of obstacle detection, two types of sensors, namely a camera and a lidar, are commonly used to obtain environmental information, so as to perform obstacle detection on the environmental information and obtain an obstacle detection result. Among them, the environmental information obtained by these two sensors has different characteristics and advantages. The camera can provide high-resolution image data and can capture details such as the texture and color of an object, which is beneficial to the recognition and classification of the object; while the lidar can provide geometric information such as the distance and shape of an object, which is beneficial to achieving accurate spatial perception ability.

[0004] Therefore, there is an urgent need for an obstacle detection method that combines multimodal fusion to provide more comprehensive and accurate perception ability, thereby improving the performance of the automatic driving system. Summary of the Invention

[0005] The purpose of this application is to provide an obstacle detection method, device, vehicle, and storage medium, aiming to solve the problem of inaccurate obstacle detection in traditional obstacle detection.

[0006] The first aspect of the embodiment of this application provides an obstacle detection method, and the method includes:

[0007] Obtain an image of the target environment where the vehicle is located, extract image features from the image of the target environment to obtain a feature map of the target environment, and the target environment includes obstacles to be detected;

[0008] Perform three-dimensional position encoding on the feature pixel points in the feature map;

[0009] Obtain the bird's-eye view point cloud feature of the target environment;

[0010] According to the bird's-eye view point cloud feature, perform feature extraction on the feature map encoded by the three-dimensional position to obtain a bird's-eye view image feature;

[0011] Based on the bird's-eye view point cloud feature and the bird's-eye view image feature, perform obstacle detection on the target environment.

[0012] In some embodiments, the performing three-dimensional position encoding on the feature pixel points in the feature map includes:

[0013] Establish a target coordinate system with the position of the vehicle as the origin;

[0014] Determine the depth value of the feature pixel points in the feature map;

[0015] Based on the depth information of the pixel points, project the feature pixel points into the target coordinate system;

[0016] Determine the coordinates of the feature pixel points in the target coordinate system as the three-dimensional position encoding of the feature pixel points.

[0017] In some embodiments, before determining the depth value of the feature pixel points in the feature map, the method further includes:

[0018] Obtain training samples, where the training samples include sample images and sample point clouds of the same scene;

[0019] Determine the depth information of the pixel points in the sample image through a depth determination model;

[0020] Determine the error of the depth information according to the correspondence between the sample point cloud and the pixel points in the sample image;

[0021] Adjust the model parameters of the depth determination model according to the error, and determine the depth information of the pixel points in the sample image according to the depth determination model with adjusted parameters until the depth determination model converges;

[0022] The determination of the depth value of the feature pixel points in the feature map includes:

[0023] Determine the depth value of the feature pixel points in the feature map through the depth determination model.

[0024] In some embodiments, the feature extraction of the feature map of the three-dimensional position encoding according to the bird's-eye view point cloud feature to obtain the bird's-eye view image feature includes:

[0025] Use the bird's-eye view point cloud feature as the query, use the feature map of the three-dimensional position encoding as the key and value, and perform feature extraction on the feature map based on the cross-attention model to obtain the bird's-eye view image feature.

[0026] In some embodiments, after performing feature extraction on the feature map based on the cross-attention model to obtain the bird's-eye view image feature, the method further includes:

[0027] Project the bird's-eye view image features into a three-dimensional detection grid to obtain the occupancy prediction result of the feature pixels for the three-dimensional detection grid. The occupancy prediction result represents the occupancy of the feature pixels of the bird's-eye view image features in each cell of the three-dimensional detection grid;

[0028] Project the bird's-eye view point cloud features aligned with the bird's-eye view image features into the three-dimensional detection grid to obtain the occupancy ground-truth result of the feature pixels for the three-dimensional detection grid. The occupancy ground-truth result represents the occupancy of the feature pixels of the bird's-eye view point cloud features in each cell of the three-dimensional detection grid;

[0029] Adjust the parameters of the cross-attention model according to the occupancy ground-truth result and the occupancy prediction result.

[0030] In some embodiments, the adjusting the parameters of the cross-attention model according to the occupancy ground-truth result and the occupancy prediction result includes:

[0031] Verify the occupancy prediction result according to the occupancy ground-truth result to obtain a verification result. The verification result represents whether the occupancy of the feature pixels of the bird's-eye view image features and the feature pixels of the bird's-eye view point cloud features in the same cell is the same;

[0032] Use the cells with the same occupancy as positive samples and the cells with different occupancy as negative samples to adjust the parameters of the cross-attention model.

[0033] In some embodiments, the detecting obstacles in the target environment based on the bird's-eye view point cloud features and the bird's-eye view image features includes:

[0034] Fuse the bird's-eye view point cloud features and the bird's-eye view image features to obtain fused features;

[0035] Detect obstacles based on the fused features.

[0036] A second aspect of the embodiments of the present application provides an obstacle detection device, the device includes:

[0037] A first extraction unit, configured to extract an image of the target environment where the vehicle is located, and extract image features from the image of the target environment to obtain a feature map of the target environment. The target environment includes obstacles to be detected;

[0038] A position encoding unit, configured to perform three-dimensional position encoding on the feature pixels in the feature map;

[0039] An acquisition unit, configured to acquire the bird's-eye view point cloud features of the target environment;

[0040] A second extraction unit, configured to extract features from the feature map encoded with three-dimensional positions according to the bird's-eye view point cloud features, so as to obtain bird's-eye view image features;

[0041] A detection unit, configured to perform obstacle detection on the target environment based on the bird's-eye view point cloud features and the bird's-eye view image features.

[0042] In some embodiments, the position encoding unit is configured to establish a target coordinate system with the position where the vehicle is located as the origin; determine the depth value of the feature pixel points in the feature map; project the feature pixel points into the target coordinate system based on the depth information of the pixel points; and determine the coordinates of the feature pixel points in the target coordinate system as the three-dimensional position encoding of the feature pixel points.

[0043] In some embodiments, the apparatus further includes:

[0044] An acquisition unit, configured to acquire training samples, where the training samples include sample images and sample point clouds of the same scene;

[0045] A first determination unit, configured to determine the depth information of the pixel points in the sample image through a depth determination model;

[0046] A second determination unit, configured to determine the error of the depth information according to the correspondence between the sample point cloud and the pixel points in the sample image;

[0047] A first adjustment unit, configured to adjust the model parameters of the depth determination model according to the error, and determine the depth information of the pixel points in the sample image according to the depth determination model with adjusted parameters until the depth determination model converges;

[0048] The position encoding unit is configured to determine the depth value of the feature pixel points in the feature map through the depth determination model.

[0049] In some embodiments, the second extraction unit is configured to use the bird's-eye view point cloud features as a query, use the feature map encoded with three-dimensional positions as a key and a value, and extract features from the feature map based on a cross-attention model to obtain the bird's-eye view image features.

[0050] In some embodiments, the apparatus further includes:

[0051] A first projection unit, configured to project the bird's-eye view image features into a three-dimensional detection grid to obtain an occupancy prediction result of the feature pixel points on the three-dimensional detection grid, where the occupancy prediction result represents the occupancy situation of the feature pixel points of the bird's-eye view image features on each cell in the three-dimensional detection grid;

[0052] A second projection unit, configured to project the bird's-eye view point cloud features aligned with the bird's-eye view image features into the three-dimensional detection grid, to obtain an occupancy ground truth result of the feature pixels on the three-dimensional detection grid, where the occupancy ground truth result represents the occupancy of the feature pixels of the bird's-eye view point cloud features in each cell of the three-dimensional detection grid;

[0053] A second adjustment unit, configured to adjust the parameters of the cross-attention model according to the occupancy ground truth result and the occupancy prediction result.

[0054] In some embodiments, the second adjustment unit is configured to verify the occupancy prediction result according to the occupancy ground truth result, to obtain a verification result, where the verification result represents whether the occupancy of the feature pixels of the bird's-eye view image features and the feature pixels of the bird's-eye view point cloud features in the same cell is the same; cells with the same occupancy are used as positive samples, and cells with different occupancy are used as negative samples, to adjust the parameters of the cross-attention model.

[0055] In some embodiments, the detection unit is configured to perform feature fusion on the bird's-eye view point cloud features and the bird's-eye view image features, to obtain a fused feature; and perform obstacle detection based on the fused feature.

[0056] A third aspect of the embodiments of the present application provides a vehicle, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the above-mentioned obstacle detection method is implemented.

[0057] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned obstacle detection method is implemented.

[0058] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:

[0059] In the embodiments of the present application, in the target environment of the obstacle to be detected, the bird's-eye view point cloud features and the image features are collected, and the image features are three-dimensionally encoded, so as to perform bird's-eye view feature extraction based on the three-dimensionally encoded features. When extracting the bird's-eye view image features in this way, the image features are first three-dimensionally position-encoded, and the three-dimensional position encoding is fused with the two-dimensional image features, so that the two-dimensional image features have three-dimensional spatial position information, enabling the image features and the point cloud features to be in the same dimension, thereby improving the accuracy of extracting the image features, and thus enabling more accurate obstacle detection when detecting obstacles based on the bird's-eye view image features and the bird's-eye view point cloud features. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1Shows a schematic diagram of an obstacle detection system involved in an obstacle detection method provided by an exemplary embodiment;

[0061] Figure 2 Shows a schematic flowchart of an obstacle detection method provided by an exemplary embodiment;

[0062] Figure 3 Shows a schematic flowchart of an obstacle detection method provided by an exemplary embodiment;

[0063] Figure 4 Shows a schematic flowchart of an obstacle detection method provided by an exemplary embodiment;

[0064] Figure 5 Shows a schematic diagram of the structure of an obstacle detection device provided by an exemplary embodiment;

[0065] Figure 6 Is a schematic diagram of the structure of a vehicle provided by an embodiment of the present invention. Detailed implementation manners

[0066] In order to make the technical problems, technical solutions, and beneficial effects to be solved by the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0067] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.

[0068] The following explains the terms used in the present application.

[0069] Bird's Eye View feature BEV: That is, Bird's Eye View, briefly described as BEV. In the field of autonomous driving, it is a vehicle environment perception technology that provides a more comprehensive scene perception from a top-down perspective.

[0070] Please refer to Figure 1, which shows an obstacle detection system involved in an obstacle detection method provided by an exemplary embodiment. The obstacle detection system includes: a camera 10 (camera), a lidar 20 (lidar), and a controller 30. The camera 10 and the lidar 20 are disposed on a vehicle. Among them, the number and the positions on the vehicle of the camera 10 and the lidar 20 can both be set as needed. Moreover, the number of the camera 10 and the lidar 20 can be the same or different, and the positions of the camera 10 and the lidar 20 on the vehicle can also be the same or different. In the embodiments of the present application, no specific limitations are made thereto.

[0071] The camera 10 is configured to collect an image of the target environment where the vehicle is located and send the collected image to the controller 30. In some embodiments, the number of the cameras 10 is determined according to the viewing angle of each camera 10, and the sum of the viewing angles of the multiple cameras 10 is not less than the angle of the environment to be monitored. For example, to monitor obstacles in the 360-degree environment around the vehicle, when the viewing angles of the cameras 10 are the same and each viewing angle is 120 degrees, the number of the cameras 10 is at least 3. The camera 10 can be any device having an image collection function. For example, the camera 10 is an in-vehicle camera, etc.

[0072] The lidar 20 is configured to collect point cloud data of the target environment where the vehicle is located and send the collected point cloud data to the controller 30. The number and the setting manner of the lidar 20 are similar to those of the camera 10, and will not be elaborated herein. The lidar 20 can be any device having a point cloud data collection function. For example, the lidar 20 is an in-vehicle radar, etc.

[0073] The controller 30 receives the image and the point cloud data, and fuses the image and the point cloud data through multimodal fusion, so as to perform obstacle detection on the target environment where the vehicle is located according to the image and the point cloud data.

[0074] In related technologies, multi-modal fusion is often performed through multi-modal fusion methods such as Bevfusion and transfuion. Among them, Bevfuion uses the LSS (Lift, Splat, Shoot) scheme to predict the depth of each feature pixel, and then projects the features of the image into the BEV space based on the camera 10 parameters and stitches them together with the features of the point cloud data for the next detection task. This scheme has problems of inaccurate depth estimation and poor feature stitching and fusion effects. Transfuion is an order fusion method that uses the Transformers decoder layer as the detection head. First, a sparse set of object queries is used to generate initial bounding boxes according to the radar characteristics. Then, the Transformers decoder adaptively fuses the spatial and context relationships related to the object queries and useful image features to complete object detection. This scheme has a problem of high computational complexity and the speed does not meet the requirements for vehicle deployment. In summary, there is an urgent need for an obstacle detection method that combines multi-modal fusion to provide more comprehensive and accurate perception capabilities, thereby improving the performance of the autonomous driving system.

[0075] To solve the above technical problems, an obstacle detection method is proposed. Refer to Figure 2 , which shows a flowchart of an obstacle detection method provided by an embodiment of the present application. By way of example and not limitation, this method is applied to a vehicle provided with the above obstacle detection system.

[0076] S201, the vehicle acquires an image of the target environment where the vehicle is located, extracts image features from the image of the target environment, and obtains a feature map of the target environment, where the target environment includes obstacles to be detected.

[0077] The vehicle acquires an image of the target environment through a camera. In some embodiments, the image can be images from multiple perspectives. The vehicle aligns the multiple images according to the acquisition times of the multiple images, and forms the image of the target environment with the images having the same acquisition time. For example, the multiple images can be images acquired by a 7view camera, and correspondingly, the vehicle forms the image of the target environment with the images acquired by the 7view camera.

[0078] Refer to Figure 3 , after the vehicle acquires an image of the target environment, it extracts the image features of the image of the target environment. The vehicle can extract the image features through any feature extraction method. For example, the vehicle extracts the image features of the image of the target environment by using a residual network (resnet34) as the feature extraction network.

[0079] After the vehicle extracts image features, it performs multi-scale feature fusion on the image features to obtain a feature map of the image features. Among them, the vehicle can perform multi-scale feature fusion on the image features through any multi-scale feature fusion method. For example, the vehicle uses a Feature Pyramid Network (FPN) to perform multi-scale feature fusion on the extracted image features to obtain the feature map.

[0080] S202, the vehicle performs three-dimensional position encoding on the feature pixel points in the feature map.

[0081] The position encoding of the feature map uses three-dimensional position encoding. In some embodiments, the vehicle constructs a target coordinate system according to the position of the vehicle, and the target coordinate system can be a world coordinate system. The vehicle also constructs a camera coordinate system based on the camera parameters of the camera. The vehicle determines the mapping relationship between the target coordinate system and the camera coordinate system. In this step, the vehicle determines the 3D point of the feature pixel point in the feature map in the camera coordinate system according to the position of the feature pixel point in the feature map, and projects the 3D point in the camera coordinate system into the target coordinate system based on the mapping relationship between the target coordinate system and the camera coordinate system, and determines the coordinate position of the feature pixel point in the world coordinate system as the three-dimensional position encoding of the feature pixel point.

[0082] It should be noted that when the image of the target environment is an image collected by multiple cameras, the vehicle can respectively construct the camera coordinate systems of each camera, and based on each camera coordinate system, respectively determine the mapping relationship between each camera coordinate system and the target coordinate system, so as to obtain the ego-vehicle surround-view panoramic three-dimensional world coordinate system.

[0083] In some embodiments, the vehicle processes the three-dimensional coordinates of the feature points through a Multi-Layer Perceptron (MLP) to obtain the same dimension as the feature map, and then fuses the two-dimensional feature map with the three-dimensional position encoding information through pixel addition (ADD operation) to obtain a feature map with three-dimensional position encoding.

[0084] S203, the vehicle obtains the bird's-eye view point cloud feature of the target environment.

[0085] Please continue to refer to Figure 3. The vehicle collects point cloud data of the target environment through a lidar, and performs point cloud processing on the point cloud data to obtain the bird's-eye view point cloud feature of the point cloud data. For example, the vehicle uses the pointpillar method to process the point cloud data to obtain the point cloud data in the pillar mode, so as to obtain three-dimensional matrix data, and performs feature extraction on the three-dimensional matrix data through a Convolutional Neural Networks (CNN) to obtain the bird's-eye view point cloud feature.

[0086] It should be noted that in the embodiments of the present application, the execution order of steps S203 and steps S201 - S202 is not specifically limited. For example, the vehicle may first execute steps S201 - S202 and then execute step S203; the vehicle may also first execute step S203 and then execute steps S201 - S202; the vehicle may also execute steps S201 - S202 and step S203 simultaneously. In the embodiments of the present application, no specific limitation is made in this regard.

[0087] S204. The vehicle extracts features from the feature map encoded by the three - dimensional position according to the bird's - eye view point cloud feature to obtain the bird's - eye view image feature.

[0088] When the vehicle extracts features from the feature map encoded by the three - dimensional position, in order to make the obtained image feature be the bird's - eye view image feature, in this step, the vehicle extracts features from the feature map encoded by the three - dimensional position according to the bird's - eye view point cloud feature.

[0089] In some embodiments, the vehicle uses the bird's - eye view point cloud feature as the query, and uses the feature map encoded by the three - dimensional position as the key and value. Based on the cross - attention model, the vehicle extracts features from the feature map to obtain the bird's - eye view image feature. In this way, the cross - attention algorithm is used between the bird's - eye view point cloud feature and the feature map to determine the bird's - eye view image feature, and the more accurate bird's - eye view feature is generated by the interaction between the point cloud feature and the image feature.

[0090] In some embodiments, after the vehicle extracts the bird's - eye view image feature, it can also perform self - supervised adjustment on the cross - attention model according to the bird's - eye view image feature. This process includes:

[0091] (1) The vehicle projects the bird's - eye view image feature into a three - dimensional detection grid to obtain the occupancy prediction result of the feature pixel points on the three - dimensional detection grid. This occupancy prediction result represents the occupancy situation of the feature pixel points of the bird's - eye view image feature on each cell in the three - dimensional detection grid.

[0092] The vehicle constructs a three - dimensional detection grid, which is used to detect the feature pixel points in the bird's - eye view feature. In this step, the vehicle projects the bird's - eye view image feature into the three - dimensional detection grid and performs binary classification on each cell in the three - dimensional detection grid, that is, determines whether each cell is occupied by the feature pixel points.

[0093] (2) The vehicle projects the bird's - eye view point cloud feature aligned with the bird's - eye view image feature into the three - dimensional detection grid to obtain the ground - truth occupancy result of the feature pixel points on the three - dimensional detection grid. This ground - truth occupancy result represents the occupancy situation of the feature pixel points of the bird's - eye view point cloud feature on each cell in the three - dimensional detection grid.

[0094] This step is based on the same principle as step (1) and will not be elaborated here.

[0095] (3) The vehicle adjusts the parameters of the cross-attention model according to the occupancy ground truth result and the occupancy prediction result.

[0096] Check the occupancy prediction result according to the occupancy ground truth result to obtain a verification result, which indicates whether the occupancy of the feature pixel points of the bird's-eye view image feature and the feature pixel points of the bird's-eye view point cloud feature are the same in the same cell; use the cells with the same occupancy as positive samples and the cells with different occupancy as negative samples to adjust the parameters of the cross-attention model.

[0097] In this implementation, the occupancy of the bird's-eye view point cloud feature and the bird's-eye view image feature in the 3D detection grid is used for detection, so as to compare the extracted network features to obtain positive and negative samples, and then adjust the parameters of the cross-attention model according to the positive and negative samples, thus realizing cross-modal self-supervised training and improving the accuracy of the cross-attention model.

[0098] S205. The vehicle performs obstacle detection on the target environment based on the bird's-eye view point cloud feature and the bird's-eye view image feature.

[0099] In this step, the vehicle fuses the bird's-eye view point cloud feature and the bird's-eye view image feature to obtain a fused feature; and performs obstacle detection based on the fused feature.

[0100] Among them, the vehicle splices the bird's-eye view point cloud feature and the bird's-eye view image feature, adaptively fuses and further extracts and fuses the spliced features, and finally generates a bird's-eye view feature for detecting obstacles. Among them, the vehicle can adaptively fuse the spliced features through a channel self-attention mechanism. In this process, by performing global average pooling and global max pooling on the spliced bird's-eye view features (H*W*C) respectively, two unit feature vectors (1*1*C) are obtained, and the two unit feature vectors are summed after passing through two fully connected layers, and then the weights for determining the bird's-eye view feature are generated through an activation layer (sigmoid layer). Through the above fusion method, the accuracy of the feature fusion network in determining the importance of feature channels is higher. And, through adaptive fusion by the channel attention mechanism, the channels of different modalities can be further fused, improving the fusion effect.

[0101] In some embodiments, when the vehicle generates a bird's-eye view feature for detecting obstacles, it can further fuse the bird's-eye view features. For example, the vehicle further extracts and fuses the fused features through a bird's-eye view feature decoder (BEV decoder) composed of multiple convolutional layers, thus improving the fusion effect.

[0102] After obtaining the fused bird's-eye view features, the vehicle performs obstacle detection on the fused bird's-eye view features. In the embodiments of the present application, the vehicle can use any obstacle detection method for obstacle detection. For example, the vehicle uses a key point detector to detect the center (CenterPoint) of an object. That is, first, the center of the obstacle is detected on the bird's-eye view features, and then the size, direction, and speed of the detection frame of the obstacle are determined through a regression network to obtain the obstacle detection result.

[0103] In the embodiments of the present application, in the target environment of the obstacle to be detected, the bird's-eye view point cloud features and image features are collected, and the image features are three-dimensionally encoded, so as to extract bird's-eye view features based on the three-dimensionally encoded features. When extracting the bird's-eye view image features in this way, the image features are first three-dimensionally position-encoded, and the three-dimensionally position-encoded features are fused with the two-dimensional image features, so that the two-dimensional image features have three-dimensional spatial position information, enabling the image features and the point cloud features to be in the same dimension, thereby improving the accuracy of extracting the image features. Therefore, when detecting obstacles based on the bird's-eye view image features and the bird's-eye view point cloud features, the obstacles can be detected more accurately.

[0104] See Figure 4 , which shows a flowchart of an obstacle detection method provided by an embodiment of the present application. By way of example and not limitation, this method is applied to a vehicle equipped with the above-mentioned obstacle detection system.

[0105] S401, the vehicle extracts an image of the target environment where the vehicle is located, extracts image features from the image of the target environment to obtain a feature map of the target environment, and the target environment includes the obstacle to be detected.

[0106] The principle of this step is the same as that of step S201, and will not be elaborated here.

[0107] S402, the vehicle establishes a target coordinate system with the position where the vehicle is located as the origin.

[0108] In some embodiments, the vehicle constructs a target coordinate system according to the position where the vehicle is located, and the target coordinate system can be a world coordinate system. In some embodiments, when the image of the target environment is an image collected by multiple cameras, the vehicle can respectively construct the camera coordinate systems of each camera, and based on each camera coordinate system, respectively determine the mapping relationship between each camera coordinate system and the target coordinate system, so as to obtain the ego-vehicle surround panoramic three-dimensional world coordinate system.

[0109] S403, the vehicle determines the depth values of the feature pixel points in the feature map.

[0110] In some embodiments, the vehicle determines the depth value of the feature pixel points in the feature map through a depth determination model. Accordingly, before this step, the vehicle trains the depth determination model. The process can be as follows: obtaining training samples, where the training samples include sample images and sample point clouds of the same scene; determining the depth information of the pixel points in the sample image through the depth determination model; determining the error of the depth information according to the correspondence between the sample point cloud and the pixel points in the sample image; adjusting the model parameters of the depth determination model according to the error, and determining the depth information of the pixel points in the sample image according to the depth determination model with adjusted parameters until the depth determination model converges.

[0111] During the training process, the point cloud is back-projected onto the feature map through the internal and external parameters of the camera, providing depth supervision information for each pixel point.

[0112] S404. The vehicle projects the feature pixel point into the target coordinate system based on the depth information of the pixel point.

[0113] The vehicle also constructs a camera coordinate system based on the camera parameters of the camera. The vehicle determines the mapping relationship between the target coordinate system and the camera coordinate system. In this step, the vehicle determines the 3D point of the feature pixel point in the feature map in the camera coordinate system according to the position of the feature pixel point in the feature map, projects the 3D point in the camera coordinate system into the target coordinate system based on the mapping relationship between the target coordinate system and the camera coordinate system, and determines the coordinate position of the feature pixel point in the world coordinate system as the three-dimensional position encoding of the feature pixel point.

[0114] S405. The vehicle determines the coordinate of the feature pixel point in the target coordinate system as the three-dimensional position encoding of the feature pixel point.

[0115] In some embodiments, the vehicle processes the three-dimensional coordinates of the feature points through a Multi-Layer Perceptron (MLP) to obtain the same dimension as the feature map, and then fuses the two-dimensional feature map and the three-dimensional position encoding information through pixel addition (ADD operation) to obtain a feature map of three-dimensional position encoding.

[0116] S406. The vehicle obtains the bird's-eye view point cloud feature of the target environment.

[0117] The principle of this step is the same as that of step S203 and will not be elaborated here.

[0118] S407. The vehicle extracts features from the feature map of three-dimensional position encoding according to the bird's-eye view point cloud feature to obtain the bird's-eye view image feature.

[0119] The principle of this step is the same as that of step S204 and will not be elaborated here.

[0120] S408. The vehicle performs obstacle detection on the target environment based on the bird's-eye view point cloud feature and the bird's-eye view image feature.

[0121] This step is the same as that of step S205 in principle and will not be elaborated here.

[0122] In the embodiment of the present application, in the target environment of the obstacle to be detected, the bird's-eye view point cloud feature and the image feature are collected, and the image feature is three-dimensionally encoded, so as to extract the bird's-eye view feature based on the three-dimensional encoded feature. When extracting the bird's-eye view image feature in this way, the image feature is first three-dimensionally position-encoded, and the three-dimensional position encoding is fused with the two-dimensional image feature, so that the two-dimensional image feature has three-dimensional spatial position information, making the image feature and the point cloud feature in the same dimension, thereby improving the accuracy of extracting the image feature. Therefore, when detecting obstacles based on the bird's-eye view image feature and the bird's-eye view point cloud feature, obstacles can be detected more accurately.

[0123] Further, a depth value is predicted for each pixel point in the feature map, so that each feature pixel point can be projected into the target coordinate system through the internal and external parameters of the camera. Furthermore, the bird's-eye view spatial coordinates are used as the position encoding of the image pixels, so that the position encodings of the image and the point cloud in the bird's-eye view space can be represented in the same form. On this basis, the cross-attention mechanism is executed.

[0124] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0125] See Figure 5 , which shows a schematic structural diagram of an obstacle detection device provided by the present application. Each unit included is used to execute each step in the above embodiment. See Figure 5 This obstacle detection device includes:

[0126] The first extraction unit 501 is used to extract an image of the target environment where the vehicle is located, extract image features from the image of the target environment, and obtain a feature map of the target environment. The target environment includes obstacles to be detected;

[0127] The position encoding unit 502 is used to perform three-dimensional position encoding on the feature pixel points in the feature map;

[0128] The acquisition unit 503 is used to acquire the bird's-eye view point cloud feature of the target environment;

[0129] The second extraction unit 504 is used to extract features from the feature map with the three-dimensional position encoding according to the bird's-eye view point cloud feature to obtain a bird's-eye view image feature;

[0130] The detection unit 505 is configured to perform obstacle detection on the target environment based on the bird's-eye view point cloud feature and the bird's-eye view image feature.

[0131] In some embodiments, the position encoding unit 502 is configured to establish a target coordinate system with the position where the vehicle is located as the origin; determine the depth value of the feature pixel points in the feature map; project the feature pixel points into the target coordinate system based on the depth information of the pixel points; and determine the coordinates of the feature pixel points in the target coordinate system as the three-dimensional position encoding of the feature pixel points.

[0132] In some embodiments, the apparatus further includes:

[0133] An acquisition unit 503, configured to acquire training samples, where the training samples include sample images and sample point clouds of the same scene;

[0134] A first determination unit, configured to determine the depth information of the pixel points in the sample image through a depth determination model;

[0135] A second determination unit, configured to determine the error of the depth information according to the correspondence between the sample point cloud and the pixel points in the sample image;

[0136] A first adjustment unit, configured to adjust the model parameters of the depth determination model according to the error, and determine the depth information of the pixel points in the sample image according to the depth determination model after the parameters are adjusted until the depth determination model converges;

[0137] The position encoding unit 502 is configured to determine the depth value of the feature pixel points in the feature map through the depth determination model.

[0138] In some embodiments, the second extraction unit 504 is configured to use the bird's-eye view point cloud feature as a query, use the feature map of the three-dimensional position encoding as a key and a value, and perform feature extraction on the feature map based on a cross-attention model to obtain the bird's-eye view image feature.

[0139] In some embodiments, the apparatus further includes:

[0140] A first projection unit, configured to project the bird's-eye view image feature into a three-dimensional detection grid to obtain an occupancy prediction result of the feature pixel points on the three-dimensional detection grid, where the occupancy prediction result represents the occupancy situation of the feature pixel points of the bird's-eye view image feature on each cell in the three-dimensional detection grid;

[0141] A second projection unit, configured to project the bird's-eye view point cloud features aligned with the bird's-eye view image features into the three-dimensional detection grid, to obtain an occupancy ground truth result of the feature pixels on the three-dimensional detection grid, where the occupancy ground truth result represents the occupancy of the feature pixels of the bird's-eye view point cloud features in each cell of the three-dimensional detection grid;

[0142] A second adjustment unit, configured to adjust the parameters of the cross-attention model according to the occupancy ground truth result and the occupancy prediction result.

[0143] In some embodiments, the second adjustment unit is configured to verify the occupancy prediction result according to the occupancy ground truth result, to obtain a verification result, where the verification result represents whether the occupancy of the feature pixels of the bird's-eye view image features and the feature pixels of the bird's-eye view point cloud features in the same cell is the same; cells with the same occupancy are used as positive samples, and cells with different occupancy are used as negative samples, to adjust the parameters of the cross-attention model.

[0144] In some embodiments, the detection unit 505 is configured to perform feature fusion on the bird's-eye view point cloud features and the bird's-eye view image features, to obtain fused features; and perform obstacle detection based on the fused features.

[0145] Figure 6 is a schematic diagram of a vehicle provided by an exemplary embodiment of the present application. As Figure 6 shown, the vehicle 6 in this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60, such as a program for obstacle detection. When the processor 60 executes the computer program 62, the steps in the above-mentioned various embodiments of the tailgate control method are implemented, such as Figure 2 the steps S201 to S205 shown. Alternatively, when the processor 60 executes the computer program 62, the functions of each unit in the above-mentioned various device embodiments are implemented, such as Figure 5 the functions of the units 501 to 505 shown.

[0146] Exemplarily, the computer program 62 can be divided into one or more units, and the one or more units are stored in the memory 61 and executed by the processor 60 to complete the present application. The one or more units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 62 in the vehicle 6. For example, the computer program 62 can be divided into a first extraction unit, a position encoding unit, an acquisition unit, a second extraction unit, and a detection unit, and the specific functions of each module are as follows:

[0147] The first extraction unit 501 is configured to extract an image of the target environment where the vehicle is located, extract image features from the image of the target environment, and obtain a feature map of the target environment, where the target environment includes obstacles to be detected;

[0148] The position encoding unit 502 is configured to perform three-dimensional position encoding on the feature pixel points in the feature map;

[0149] The acquisition unit 503 is configured to acquire the bird's-eye view point cloud feature of the target environment;

[0150] The second extraction unit 504 is configured to perform feature extraction on the feature map encoded in three dimensions according to the bird's-eye view point cloud feature to obtain a bird's-eye view image feature;

[0151] The detection unit 505 is configured to perform obstacle detection on the target environment based on the bird's-eye view point cloud feature and the bird's-eye view image feature.

[0152] The vehicle 6 may be any vehicle with an obstacle detection function. The vehicle 6 may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 6 merely examples of the vehicle 6 do not constitute a limitation on the vehicle 6, and may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the vehicle 6 may further include an input / output device, a network access device, a bus, etc.

[0153] The so-called processor 60 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0154] The memory 61 may be an internal storage unit of the vehicle 6, such as a hard disk or memory of the vehicle 6. The memory 61 may also be an external storage device of the vehicle 6, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the vehicle 6. Further, the memory 61 may also include both an internal storage unit of the vehicle 6 and an external storage device. The memory 61 is used to store the computer program and other programs and data required by the terminal device. The memory 61 may also be used to temporarily store the data that has been output or will be output.

[0155] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.

[0156] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0157] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0158] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0159] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0160] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0161] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0162] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0163] The embodiment of the present application also provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned method embodiments when executed.

[0164] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. An obstacle detection method, characterized in that, the method includes: Obtain an image of the target environment where the vehicle is located, extract image features from the image of the target environment to obtain a feature map of the target environment, and the target environment includes obstacles to be detected; Perform three-dimensional position encoding on the feature pixel points in the feature map; Obtain the bird's-eye view point cloud feature of the target environment; According to the bird's-eye view point cloud feature, perform feature extraction on the feature map with three-dimensional position encoding to obtain a bird's-eye view image feature; Based on the bird's-eye view point cloud feature and the bird's-eye view image feature, perform obstacle detection on the target environment.

2. The method according to claim 1, characterized in that, the performing three-dimensional position encoding on the feature pixel points in the feature map includes: Establish a target coordinate system with the position where the vehicle is located as the origin; Determine the depth value of the feature pixel points in the feature map; Based on the depth information of the pixel points, project the feature pixel points into the target coordinate system; Determine the coordinates of the feature pixel points in the target coordinate system as the three-dimensional position encoding of the feature pixel points.

3. The method according to claim 2, characterized in that, before determining the depth value of the feature pixel points in the feature map, the method further includes: Obtain training samples, where the training samples include sample images and sample point clouds of the same scene; Determine the depth information of the pixel points in the sample image through a depth determination model; According to the correspondence between the sample point cloud and the pixel points in the sample image, determine the error of the depth information; According to the error, adjust the model parameters of the depth determination model, and according to the depth determination model with adjusted parameters, determine the depth information of the pixel points in the sample image until the depth determination model converges; the determining the depth value of the feature pixel points in the feature map includes: Determine the depth value of the feature pixel points in the feature map through the depth determination model.

4. The method according to claim 1, characterized in that, the performing feature extraction on the feature map with three-dimensional position encoding according to the bird's-eye view point cloud feature to obtain a bird's-eye view image feature includes: Use the bird's-eye view point cloud feature as a query, use the feature map with three-dimensional position encoding as a key and a value, and based on a cross-attention model, perform feature extraction on the feature map to obtain the bird's-eye view image feature.

5. The method according to claim 4, characterized in that, after performing feature extraction on the feature map based on the cross-attention model to obtain the bird's-eye view image feature, the method further includes: Project the bird's-eye view image feature into a three-dimensional detection grid to obtain an occupancy prediction result of the feature pixel points for the three-dimensional detection grid, and the occupancy prediction result represents the occupancy situation of the feature pixel points of the bird's-eye view image feature for each cell in the three-dimensional detection grid; Project the bird's-eye view point cloud features aligned with the bird's-eye view image features into the three-dimensional detection grid to obtain the occupancy ground truth results of the feature pixels for the three-dimensional detection grid. The occupancy ground truth results represent the occupancy of the feature pixels of the bird's-eye view point cloud features in each cell of the three-dimensional detection grid. Adjust the parameters of the cross-attention model according to the occupancy ground truth results and the occupancy prediction results.

6. The method according to claim 5, wherein, the adjusting the parameters of the cross-attention model according to the occupancy ground truth results and the occupancy prediction results includes: Verifying the occupancy prediction results according to the occupancy ground truth results to obtain a verification result, where the verification result indicates whether the occupancy of the feature pixels of the bird's-eye view image features and the feature pixels of the bird's-eye view point cloud features in the same cell is the same; Using the cells with the same occupancy as positive samples and the cells with different occupancy as negative samples to adjust the parameters of the cross-attention model.

7. The method according to claim 1, wherein, the performing obstacle detection on the target environment based on the bird's-eye view point cloud features and the bird's-eye view image features includes: Performing feature fusion on the bird's-eye view point cloud features and the bird's-eye view image features to obtain fused features; Performing obstacle detection based on the fused features.

8. An obstacle detection device, wherein, the device includes: A first extraction unit for extracting an image of the target environment where the vehicle is located, and extracting image features from the image of the target environment to obtain a feature map of the target environment, where the target environment includes obstacles to be detected; A position encoding unit for performing three-dimensional position encoding on the feature pixels in the feature map; An acquisition unit for acquiring the bird's-eye view point cloud features of the target environment; A second extraction unit for extracting features from the feature map encoded in three-dimensional position according to the bird's-eye view point cloud features to obtain bird's-eye view image features; A detection unit for performing obstacle detection on the target environment based on the bird's-eye view point cloud features and the bird's-eye view image features.

9. A vehicle, wherein, the vehicle includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the obstacle detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program. The computer program, when executed by a processor, implements the obstacle detection method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Obstacle detection method, obstacle detection model training method and electronic equipment

    CN120726610A

  • Height detection method and device, vehicle detection method and device, vehicle and storage medium

    CN121214393A