Method, device, and vehicle for detecting obstacles

By combining the visual layer information of the prior map and the vehicle camera image, the change detection model is used to extract the difference information of road areas, the problem of low-short obstacles and pit hole detection in the prior art is solved, efficient and low-cost obstacle recognition is achieved, and the safety of autonomous driving is improved.

CN114842453BActive Publication Date: 2025-07-22HANGZHOU ZHIHUI MANTU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210652249.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2025-07-22
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

In the prior art, it is difficult to efficiently detect low obstacles and pits in road driving areas in autonomous driving. The camera sensor solution has low recall rate, while the lidar solution is costly and has a risk of ground estimation errors, making it difficult to meet safe driving requirements.

Method used

The visual layer information of the prior map is combined with the vehicle camera image, and the difference information of road areas is obtained through the change detection model, and areas of dangerous roadblocks such as low obstacles and pits are extracted, so as to reduce the pressure of the vehicle-end real-time algorithm and improve recognition performance.

Benefits of technology

On the basis of not increasing the cost of autonomous driving vehicles, the detection accuracy and recall rate of low obstacles and pits are improved, and the safety and identification performance of autonomous driving are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842453B_ABST
    Figure CN114842453B_ABST
Patent Text Reader

Abstract

A method, device and vehicle for detecting obstacles, the method comprising: obtaining a real-time camera image captured by a vehicle camera, and a road area mask corresponding to the road area in the camera image; obtaining visual layer information of a prior map; inputting the camera image, the visual layer information of the prior map and the road area mask into a change detection model to output a change area mask; and performing a fusion post-processing operation according to the change area mask to output obstacle information. This method can improve the detection accuracy of obstacles in the road driving area without increasing the cost of autonomous vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular, to a method, device, and vehicle for detecting obstacles. Background Art

[0002] With the increasing development and gradual maturity of the field of autonomous driving, the industry's requirements for the overall performance of autonomous driving perception are also getting higher and higher, especially for safety requirements. In the road driving area, the existence of low obstacles and potholes and other dangerous roadblocks has become one of the safety problems that the autonomous driving perception system must face.

[0003] Currently, the detection schemes for low obstacles and potholes in the road driving area include the scheme combining camera sensors and the scheme combining lidar. However, both of the above schemes have deficiencies when applied to vehicle mass production. For example, if the scheme of camera sensors is adopted, since low obstacles and potholes do not have unified visual features, unlike conventional obstacles that conform to certain rules, such as people and vehicles, it is difficult to accurately detect low obstacles and potholes based on conventional deep learning neural network algorithms, and the recall rate of samples is relatively low. If the lidar scheme is adopted, it is necessary to rely on lidar for ground height estimation algorithms to calculate the ground height near the vehicle during real-time operation, and then use the comparison between the lidar point cloud height and the ground height to identify dangerous roadblocks such as low obstacles and potholes. This scheme has high requirements for the beam of the lidar, so the cost of the lidar is relatively high, which is contradictory to the original intention of low-cost mass production of autonomous driving. In addition, due to the limited diameter of the pothole, there is a risk that the laser line cannot be scanned and the ground estimation is incorrect. Therefore, the current ground estimation algorithm relying on lidar is difficult to meet the requirements of high-recall safe driving for autonomous driving.

[0004] Based on the above problems, there is an urgent need in the industry for a method to improve the detection accuracy of obstacles in the road driving area on the premise of controlling production costs. Summary of the Invention

[0005] This application provides a method, device, and vehicle for detecting obstacles to improve the detection accuracy of obstacles in the road driving area.

[0006] In a first aspect, a method for detecting obstacles is provided, including: obtaining a real-time camera image captured by a vehicle camera, and a road area mask corresponding to the road area in the camera image; obtaining visual layer information of a prior map; inputting the camera image, the visual layer information of the prior map, and the road area mask into a change detection model to output a change area mask; performing a fusion post-processing operation according to the change area mask to output obstacle information.

[0007] In combination with the first aspect, in a possible implementation manner, the step of inputting the camera image, the visual layer information of the prior map, and the road area mask into a change detection model to output a change area mask includes: respectively performing feature extraction on the camera image, the visual layer information of the prior map, and the road area mask through a feature encoder to obtain a feature map of the camera image, a feature map of the visual layer information of the prior map, and a feature map of the road area mask; extracting difference comparison information between the feature map of the camera image and the feature map of the visual layer information of the prior map; performing correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output high-level semantic features; and performing decoding processing on the high-level semantic features through a feature decoder to output the change area mask.

[0008] In combination with the first aspect, in a possible implementation manner, the step of performing correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output high-level semantic features includes: performing dimensionality reduction convolution on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output a dimensionality-reduced feature map; and inputting the dimensionality-reduced feature map and the difference comparison information into a correlation filtering module to output the high-level semantic features.

[0009] In combination with the first aspect, in a possible implementation manner, the road area mask corresponding to the road area in the camera image is obtained by the following method: inputting the camera image into a panoramic segmentation model to output the road area mask, and the panoramic segmentation model is used for region division and category annotation of the camera image.

[0010] In combination with the first aspect, in a possible implementation manner, the step of performing post-fusion processing operations according to the change area mask to output obstacle information includes: extracting a dangerous roadblock area according to the change area mask; performing visual tracking filtering processing on the dangerous roadblock area; and converting the position information of the dangerous roadblock area from two-dimensional coordinates to three-dimensional coordinates.

[0011] In a second aspect, a training method for obstacle detection is provided, and the method includes: obtaining training data, where the training data includes a camera image captured by a camera in a vehicle, visual layer information of a prior map, and a road area mask corresponding to the road area in the camera image; and inputting the training data into a change detection model to train parameters of the change detection model, where the change detection model is used to output a change area mask.

[0012] In combination with the second aspect, in a possible implementation, the training data is input into the change detection model to train the parameters of the change detection model, including: respectively extracting features from the camera image, the visual layer information of the prior map, and the road area mask through a feature encoder to obtain the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask; extracting the difference comparison information between the feature map of the camera image and the feature map of the visual layer information of the prior map; performing correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output high-level semantic features; and decoding the high-level semantic features through a feature decoder to output the change area mask.

[0013] In combination with the second aspect, in a possible implementation, the performing correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output high-level semantic features includes: performing dimensionality reduction convolution on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output the dimensionality-reduced feature map; and inputting the dimensionality-reduced feature map and the difference comparison information into a correlation filtering module to output the high-level semantic features.

[0014] In combination with the second aspect, in a possible implementation, the training method further includes: performing a fusion post-processing operation according to the change area mask to output obstacle information.

[0015] In combination with the second aspect, in a possible implementation, the performing a fusion post-processing operation according to the change area mask to output obstacle information includes: extracting a dangerous roadblock area according to the change area mask; performing visual tracking filtering processing on the dangerous roadblock area; and converting the position information of the dangerous roadblock area from two-dimensional coordinates to three-dimensional coordinates.

[0016] In a third aspect, a device for detecting obstacles is provided, including: an acquisition module, configured to acquire a real-time camera image captured by a vehicle camera, and a road area mask corresponding to the road area in the camera image; the acquisition module is further configured to acquire the visual layer information of a prior map; a processing module, configured to input the camera image, the visual layer information of the prior map, and the road area mask into a change detection model to output a change area mask; and the processing module is further configured to perform a fusion post-processing operation according to the change area mask to output obstacle information.

[0017] In combination with the third aspect, in a possible implementation manner, the processing module is specifically configured to: respectively perform feature extraction on the camera image, the visual layer information of the prior map, and the road area mask through a feature encoder to obtain a feature map of the camera image, a feature map of the visual layer information of the prior map, and a feature map of the road area mask; extract difference comparison information between the feature map of the camera image and the feature map of the visual layer information of the prior map; perform correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output high-level semantic features; and perform decoding processing on the high-level semantic features through a feature decoder to output the change area mask.

[0018] In combination with the third aspect, in a possible implementation manner, the processing module is specifically configured to: perform dimensionality reduction convolution on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output a dimensionality-reduced feature map; and input the dimensionality-reduced feature map and the difference comparison information into a correlation filtering module to output the high-level semantic features.

[0019] In combination with the third aspect, in a possible implementation manner, the acquisition module is specifically configured to: input the camera image into a panoramic segmentation model to output the road area mask, and the panoramic segmentation model is used for region division and category annotation of the camera image.

[0020] In combination with the third aspect, in a possible implementation manner, the processing module is specifically configured to: extract a dangerous roadblock area according to the change area mask; perform visual tracking filtering processing on the dangerous roadblock area; and convert the position information of the dangerous roadblock area from two-dimensional coordinates to three-dimensional coordinates.

[0021] Fourth aspect, a device for detecting obstacles is provided, and the device includes: an acquisition module, configured to acquire training data, where the training data includes a camera image captured by a camera in a vehicle, visual layer information of a prior map, and a road area mask corresponding to a road area in the camera image; and a processing module, configured to input the training data into a change detection model to train parameters of the change detection model, and the change detection model is used for outputting a change area mask.

[0022] In combination with the fourth aspect, in a possible implementation manner, the processing module is specifically configured to: respectively perform feature extraction on the camera image, the visual layer information of the prior map, and the road area mask through a feature encoder to obtain a feature map of the camera image, a feature map of the visual layer information of the prior map, and a feature map of the road area mask; extract the difference comparison information between the feature map of the camera image and the feature map of the visual layer information of the prior map; perform correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output high-level semantic features; and perform decoding processing on the high-level semantic features through a feature decoder to output the changed area mask.

[0023] In combination with the fourth aspect, in a possible implementation manner, the processing module is specifically configured to: perform dimensionality reduction convolution on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output a dimensionality-reduced feature map; and input the dimensionality-reduced feature map and the difference comparison information into a correlation filtering module to output the high-level semantic features.

[0024] In combination with the fourth aspect, in a possible implementation manner, the processing module is further configured to perform a post-fusion processing operation according to the changed area mask to output obstacle information.

[0025] In combination with the fourth aspect, in a possible implementation manner, the processing module is specifically configured to: extract a dangerous roadblock area according to the changed area mask; perform visual tracking filtering processing on the dangerous roadblock area; and convert the position information of the dangerous roadblock area from two-dimensional coordinates to three-dimensional coordinates.

[0026] In a fifth aspect, a computer device is provided, including a processor, where the processor is configured to call a computer program from a memory, and when the computer program is executed, the processor is configured to execute the method in the first aspect or any possible implementation manner in the first aspect, or is configured to execute the method in the second aspect or any possible implementation manner in the second aspect.

[0027] In a sixth aspect, a computer-readable storage medium is provided for storing a computer program, where the computer program includes code for executing the method in the first aspect or any possible implementation manner in the first aspect, or code for executing the method in the second aspect or any possible implementation manner in the second aspect.

[0028] In a seventh aspect, there is provided a computer program product including a computer program which comprises code for performing the method in the above-mentioned first aspect or any possible implementation manner of the first aspect, or code for performing the method in the above-mentioned second aspect or any possible implementation manner of the second aspect.

[0029] In an eighth aspect, there is provided a vehicle which includes a computing device for performing the method in the above-mentioned first aspect or any possible implementation manner of the first aspect, or for performing the method in the above-mentioned second aspect or any possible implementation manner of the second aspect.

[0030] In the embodiments of the present application, a method for detecting obstacles is proposed. This method utilizes the visual layer information of the existing prior map, combines it with the camera image acquired by the vehicle, and detects the changed area in the road driving area through the difference change between the two, so as to obtain the area where obstacles exist, thereby improving the safety of ensuring autonomous driving. Compared with the generative algorithm model of the neural network algorithm based on label learning, the discriminative algorithm model for judging the changed area based on the layer information of the prior map is more in line with the human sensory principle and can improve the recall rate and obstacle recognition performance. Moreover, through the algorithm scheme of combining the visual layer information of the prior map, part of the computational pressure of the recognition algorithm is transferred to the offline layer information, reducing the real-time algorithm recognition pressure at the vehicle end, and then more optimal algorithm performance can be obtained by using the computing resources. Thus, the obstacle recognition performance of the vehicle is improved without increasing the cost of the autonomous driving vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0032] Figure 1 is a schematic diagram of an application scenario of an embodiment of the present application;

[0033] Figure 2 is a framework schematic diagram of a computing device 200 according to an embodiment of the present application;

[0034] Figure 3 is a structural schematic diagram of a change detection model 210 according to an embodiment of the present application;

[0035] Figure 4 is a schematic diagram of the working process of a fusion module 212 according to an embodiment of the present application;

[0036] Figure 5 is a structural schematic diagram of a post-fusion processing module 220 according to an embodiment of the present application;

[0037] Figure 6 is a schematic flowchart of a method for detecting obstacles according to an embodiment of the present application;

[0038] Figure 7 is a schematic flowchart of a training method for detecting obstacles according to an embodiment of the present application;

[0039] Figure 8 is a schematic structural diagram of device 800 according to an embodiment of the present application;

[0040] Figure 9 is a schematic structural diagram of device 900 according to an embodiment of the present application;

[0041] Figure 10 is a schematic structural diagram of device 1000 according to an embodiment of the present application. Detailed implementation manners

[0042] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0043] First, the nouns involved in the present application are explained:

[0044] Prior map: refers to the map information collected and made in advance. The prior map may include multiple layer information. For example, it may include visual layer information and lidar layer information. Among them, the visual layer information may include RGB image information, and the lidar layer information may include the height information of the constructed 3D map.

[0045] Change detection: By performing differential detection on the visual information obtained by the vehicle-mounted camera and the layer information of the prior map, the area that has changed relative to the layer information of the prior map is found.

[0046] Neural network algorithm of deep learning: refers to an algorithm based on data driving. It can learn the internal laws and representation levels of sample data, and then obtain a model that can analyze data such as images, texts, and sounds.

[0047] Mask: is a template for image filters, which can be used to filter pixels of an image through a matrix to extract the required objects or regions.

[0048] Semantic features: The semantic features of an image include a visual layer, an object layer, and a concept layer. The visual layer, which is also the low-level semantic feature, includes features such as color, texture, and shape in the image. The object layer is the intermediate-level semantic feature and can include the attribute features of the image. For example, the state of an object at a certain moment. The concept layer is the high-level semantic feature, which refers to the information closest to human understanding. The higher the level of the semantic feature, the stronger the semanticity and the stronger the discrimination ability.

[0049] The method for detecting obstacles provided by this application aims to solve the above technical problems in the prior art.

[0050] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.

[0051] To solve the problems described above, an embodiment of this application proposes a method and device for detecting obstacles. This method utilizes the layer information of an existing prior map, combines it with the sensor information obtained by the vehicle, and through a change detection algorithm between the two, obtains the changed areas in the road driving area, thereby obtaining the areas with low obstacles and potholes and other dangerous roadblocks, and further improving the safety of ensuring autonomous driving. Compared with the generative algorithm model of the neural network algorithm based on label learning, the discriminative algorithm model for judging the changed areas based on the prior layer information is more in line with the human sensory principle and can improve the recall rate and obstacle recognition performance.

[0052] Among them, the recall rate, also known as the completeness rate, refers to the ratio of the number of positive examples detected in the sample to the number of original positive examples in the sample. The higher the recall rate of the neural network model, the better the recognition performance.

[0053] Figure 1 It is a schematic diagram of the application scenario of an embodiment of this application. As Figure 1 shown, a camera 110 and a computing device 200 are installed in the vehicle 100. Among them, the camera 110 can be used to obtain image or video data, and send the image or video data to the computing device 200. The computing device 200 can analyze and process the obtained image or video data to obtain an analysis result for use in obstacle detection, collision avoidance, and navigation installation and other solutions.

[0054] Optionally, the above computing device 200 may include a vehicle head unit, which may be other processors, processing chips or devices installed inside the vehicle. Alternatively, the functions implemented by the computing device 200 may be executed by multiple processors distributed in a decentralized manner inside the vehicle. Or, the above computing device 200 is not necessarily installed inside the vehicle and may also be other devices outside the vehicle. For example, the computing device 200 may also be replaced by a network device. The above network device includes but is not limited to a cloud server. For example, the vehicle 100 may upload image or video data to the cloud server, and the cloud server analyzes and processes the data and returns the processing result to the vehicle 100.

[0055] For the sake of simplicity, Figure 1 only the modules related to the embodiments of the present application are shown. Optionally, the vehicle 100 may further include more functional modules or units. For example, devices such as lidar and millimeter-wave radar may also be installed in the vehicle 100.

[0056] It should be understood that Figure 1 the description of the application scenarios is only an example and not a limitation. In practice, appropriate deformations and additions or subtractions can be made on the basis of the above scenarios, and the solutions of the embodiments of the present application are still applicable.

[0057] Figure 2 is a framework schematic diagram of the computing device 200 according to an embodiment of the present application. As Figure 2 shown, the computing device 200 includes a change detection model 210, a fusion post-processing module 220, and a panoramic segmentation model 230. Among them, the change detection model 210 is a neural network model based on deep learning. The inputs of the change detection model 210 include: camera images, visual layer information of the prior map, and road area masks.

[0058] Among them, the camera image refers to the real-time image information acquired by the camera 110 installed in the vehicle 100. The prior map may refer to the layer information of the vehicle driving area acquired in advance, which is usually used to make a high-precision map and can be made from the data collected by a survey vehicle equipped with a camera and lidar in advance in the vehicle driving area. The prior map includes visual layer information and / or laser layer information. Among them, the visual layer information may include RGB image information, and the laser layer information may include the height information of the constructed 3D map. It should be noted that in some examples of the present application, only the visual layer information in the prior map may be used for obstacle detection without the laser layer information. In other words, in some examples, the prior map may only include visual layer information.

[0059] The road area mask refers to a template used to extract the road area in an image. The road area mask can be obtained through the panoramic segmentation model 230. The input of the panoramic segmentation model 230 includes a camera image, and the output includes a road area mask, which is used to divide the input image into regions and label the categories. As an example, the panoramic segmentation model 230 can assign semantic labels and instance identifiers (IDs) to each pixel point in the image. Among them, the semantic label is used to indicate the category of the object, and the instance ID is used to indicate different labels of the same type of object. The panoramic segmentation model 230 is a neural learning network. The embodiments of the present application do not limit the specific implementation manner of the panoramic segmentation model 230, as long as it can implement the function of panoramic segmentation. As an example, the panoramic segmentation model 230 can adopt the SegNet network structure.

[0060] The output of the change detection model 210 is a change area mask. The change detection model 210 is a neural network model based on deep learning. It can compare the visual layer information of the camera image and the prior map, and determine the area where the camera image has changed relative to the prior map through a differential change detection algorithm. Then, combined with the road area mask, it obtains the areas with low obstacles, potholes and other dangerous roadblocks on the road driving area. Compared with the generative algorithm model of the neural network algorithm based on label learning, the discriminative algorithm model of the change detection model 210 based on the visual layer information of the prior map to judge the changed area is more in line with the human sensory principle and can improve the recall rate and obstacle recognition performance.

[0061] The fusion post-processing module 220 is used to post-process the change area mask output by the change detection model 210 to obtain obstacle information. Among them, the obstacle information may include obstacle coordinate information. As an example, the fusion post-processing module 220 can perform processing such as dangerous roadblock area extraction, fusion processing to remove false detections, and three-dimensional coordinate conversion on the change area mask, and finally obtain the three-dimensional coordinate information of the obstacle.

[0062] Optionally, the fusion post-processing module 220 can also combine the sensor information obtained by the lidar to post-process the change area mask to improve the accuracy of obstacle detection.

[0063] Traditional autonomous driving algorithms impose all recognition pressure on the in-vehicle algorithms of autonomous vehicles. Limited by the problem of limited computing resources on the vehicle side, they cannot fully utilize all their recognition capabilities. In the embodiments of the present application, while ensuring a high recall rate for low obstacles and potholes and other dangerous roadblocks in the road driving area, through an algorithmic solution that combines the visual layer information of the prior map, part of the computational pressure of the recognition algorithm is transferred to the offline layer information, reducing the real-time algorithm recognition pressure on the vehicle side. Furthermore, better algorithm performance can be obtained by using the computing resources. Thus, without increasing the cost of autonomous vehicles, the obstacle recognition performance of the vehicle is improved.

[0064] In addition, the algorithm principle of the solution in the embodiments of the present application determines the obstacles by judging the principle of road area change. Different low obstacles and potholes on the road surface can be classified and labeled and treated as the changed road surface area, improving the recognition scope of the current algorithm.

[0065] Moreover, the solution in the embodiments of the present application uses offline information for discriminative algorithm modeling, which is more in line with the human sensory principle and improves the overall algorithm effect.

[0066] Figure 3 It is a schematic structural diagram of a change detection model 210 according to an embodiment of the present application. As Figure 3 shown, the change detection model 210 includes a feature encoder 211, a fusion module 212, and a feature decoder 213. Among them, the feature encoder 211 is used to extract features from the input data respectively to obtain the feature maps of the input data. For example, the feature encoder 211 can be used to extract features from the camera image, the visual layer information of the prior map, and the road area mask respectively, and output the feature maps of the camera image, the visual layer information of the prior map, and the road area mask respectively.

[0067] The fusion module 212 is used to compare the camera image and the visual layer information of the prior map according to the feature maps extracted by the feature encoder 211 to determine the changed area in the camera image, and perform relevant filtering in combination with the road area mask to output high-level semantic features. For example, the fusion module 212 includes a differential contrast feature attention module and a relevant filtering module. The differential feature attention module can be used to extract the differential contrast information between the feature map of the camera image and the feature map of the visual layer information of the prior map, and use the differential contrast information and each feature map as the input of the relevant filtering module, and process them through the relevant filtering module to obtain high-level semantic features. The relevant filtering module is used to further strengthen the extracted differential contrast information.

[0068] The feature decoder 213 is used to receive the high-level semantic features and the road area mask, and perform decoding to output the changed area mask. The feature decoder 213 is used to convert the high-level semantic features into an image.

[0069] Figure 4 is a schematic flowchart of the working process of the fusion module 212 according to an embodiment of the present application. As Figure 4 shown, as an example, the difference contrast feature attention module in the fusion module 212 is used to receive the feature map of the camera image and the feature map of the visual layer information of the prior map, and extract the difference contrast information between the two based on the differential change detection algorithm. At the same time, the fusion module 212 performs dimensionality reduction convolution on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to obtain the feature map after dimensionality reduction. The difference contrast information and the feature map after dimensionality reduction are used as inputs to the correlation filtering module. The correlation filtering module uses the correlation in signal processing to train the correlation filter by extracting the target features to strengthen the difference contrast information and output the high-level semantic features.

[0070] Figure 5 is a schematic structural diagram of the post-fusion processing module 220 according to an embodiment of the present application. The post-fusion processing module 220 is used to receive the changed area mask and perform post-processing to output obstacle information. As Figure 5 shown, as an example, the post-fusion processing module 220 mainly includes three modules, namely the dangerous roadblock area extraction module 221, the fusion processing de-misdetection module 222, and the three-dimensional coordinate conversion module 223. The dangerous roadblock area extraction module 221 can use the morphological processing method to obtain the concerned dangerous roadblock area. The fusion processing de-misdetection module 222 can use visual tracking filtering processing to obtain a stable detection result to improve the accuracy. The three-dimensional coordinate conversion module can be used to obtain the three-dimensional coordinate information by projecting the obtained two-dimensional dangerous roadblock area information through point cloud projection.

[0071] Figure 6 is a schematic flowchart of the method for detecting obstacles according to an embodiment of the present application. This method can be executed by any computing device set inside the vehicle, or can be executed by a device outside the vehicle, such as a cloud server. As Figure 6 shown, this method includes the following contents.

[0072] S601. Obtain the real-time camera image captured by the vehicle camera, and the road area mask corresponding to the road area in the camera image.

[0073] Among them, the camera image refers to the real-time image information obtained by the camera installed in the vehicle.

[0074] Optionally, the road area mask can be obtained through a panoramic segmentation model. For example, input the camera image into the panoramic segmentation model to output the road area mask. The panoramic segmentation model is used to divide the regions of the camera image and label the categories. As an example, the panoramic segmentation model 230 can assign semantic labels and instance IDs to each pixel point in the image. Among them, the semantic label is used to indicate the category of the object, and the instance ID is used to indicate different labels of the same type of object. The panoramic segmentation model is a neural learning network. The embodiments of the present application do not limit the specific implementation manner of the panoramic segmentation model, as long as it can implement the function of panoramic segmentation. As an example, the panoramic segmentation model can adopt the SegNet network structure.

[0075] S602. Obtain the visual layer information of the prior map.

[0076] Among them, the prior map is pre-collected map information. The prior map can refer to the visual layer information of the vehicle driving area obtained in advance. The prior map can be made from the data collected in advance by a map surveying and mapping team using a survey vehicle equipped with a camera and a lidar in the vehicle driving area. The prior map can be used to make a high-precision map. The prior map may include visual layer information and / or lidar layer information. Among them, the visual layer information may include RGB image information, and the lidar layer information may include the height information of the constructed 3D map. It should be noted that in some examples of the present application, only the visual layer information in the prior map can be used for obstacle detection, without the need for lidar layer information. In other words, in some examples, the prior map may only include visual layer information.

[0077] S603. Input the camera image, the visual layer information of the prior map, and the road area mask into the change detection model to output the change area mask.

[0078] Among them, the change area mask is used to indicate the area where the camera image has changed relative to the visual layer information of the prior map. The change detection model is a neural network model based on deep learning. It can compare the camera image and the visual layer information of the prior map, and determine the area where the camera image has changed relative to the prior map through a differential change detection algorithm. Then, combined with the road area mask, the areas with low obstacles, potholes and other dangerous roadblocks on the road driving area can be obtained.

[0079] In some examples, the process by which the change detection model obtains the change region mask based on the input information includes: respectively extracting features from the camera image, the visual layer information of the prior map, and the road region mask through a feature encoder to obtain the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road region mask; extracting the difference comparison information between the feature maps of the camera image and the visual layer information of the prior map; performing correlation filtering on the difference comparison information based on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road region mask to output high-level semantic features; and decoding the high-level semantic features through a feature decoder to output the change region mask.

[0080] In some examples, performing correlation filtering on the difference comparison information based on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road region mask to output high-level semantic features includes: performing dimensionality reduction convolution on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road region mask to output the dimensionality-reduced feature map; and inputting the dimensionality-reduced feature map and the difference comparison information into a correlation filtering module to output high-level semantic features.

[0081] In some examples, decoding the high-level semantic features through a feature decoder includes: the feature decoder receiving the high-level semantic features and the road region mask and performing decoding to output the change region mask. The feature decoder is used to convert the high-level semantic features into an image.

[0082] S604. Perform post-processing operations for fusion based on the change region mask to output obstacle information.

[0083] Among them, the above obstacle information may include the coordinate information of the obstacle. As an example, the above post-processing operations for fusion may include: extracting the dangerous roadblock area according to the change region mask; performing visual tracking filtering processing on the dangerous roadblock area; and converting the position information of the dangerous roadblock area from two-dimensional coordinates to three-dimensional coordinates.

[0084] In the embodiments of the present application, a method for detecting obstacles is proposed. This method utilizes the visual layer information of the existing prior map, combines it with the camera image obtained by the vehicle, and through difference change detection between the two, obtains the changed area in the road driving area, thereby obtaining the area where there are obstacles, and further improving the safety of ensuring autonomous driving. Compared with the generative algorithm model of the neural network algorithm based on label learning, the discriminative algorithm model for judging the changed area based on the layer information of the prior map is more in line with the human sensory principle and can improve the recall rate and obstacle recognition performance.

[0085] Figure 7It is a schematic flowchart of a training method for detecting obstacles according to an embodiment of the present application. This method can be executed by any computing device. For example, it can be executed by a computing device in the vehicle, or by a computing device outside the vehicle, or can be executed in a server. As Figure 7 shown, this method includes the following contents.

[0086] S701. Obtain training data, where the training data includes camera images captured by a camera in the vehicle, visual layer information of a prior map, and a road area mask corresponding to the road area in the camera image.

[0087] Among them, the prior map is pre-collected map information, and the road area mask is used to extract the road area in the camera image.

[0088] Optionally, the road area mask can be obtained through a panoramic segmentation model. For example, input the camera image into the panoramic segmentation model to output the road area mask, and the panoramic segmentation model is used to divide and label categories of the camera image.

[0089] S702. Input the training data into a change detection model to train the parameters of the change detection model, and the change detection model is used to output a change area mask.

[0090] Among them, the change area mask is used to indicate the area where the camera image changes relative to the visual layer information of the prior map.

[0091] In S702, inputting the training data into the change detection model to train the parameters of the change detection model includes: respectively performing feature extraction on the camera image, the visual layer information of the prior map, and the road area mask through a feature encoder to obtain a feature map of the camera image, a feature map of the visual layer information of the prior map, and a feature map of the road area mask; extracting the difference comparison information between the feature map of the camera image and the feature map of the visual layer information of the prior map; performing correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output high-level semantic features; and performing decoding processing on the high-level semantic features through a feature decoder to output a change area mask.

[0092] In some examples, performing correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output high-level semantic features includes: performing dimensionality reduction convolution on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output a dimensionality-reduced feature map; and inputting the dimensionality-reduced feature map and the difference comparison information into a correlation filtering module to output high-level semantic features.

[0093] In some examples, Figure 7 the method further includes: performing a fusion post - processing operation according to the change region mask to output obstacle information.

[0094] As an example, the above - mentioned fusion post - processing operation may include: extracting a dangerous roadblock area according to the change region mask; performing visual tracking filtering processing on the dangerous roadblock area; and converting the position information of the dangerous roadblock area from two - dimensional coordinates to three - dimensional coordinates.

[0095] Figure 8 is a schematic structural diagram of a device 800 according to an embodiment of the present application. The device 800 can be used to execute Figure 6 the method in, or execute the method executed by the computing device 200 in the above text.

[0096] The device 800 includes: an acquisition module 810, configured to acquire a real - time camera image captured by a vehicle camera, and a road region mask corresponding to the road region in the camera image; the acquisition module 810 is further configured to acquire visual layer information of a prior map; a processing module 820, configured to input the camera image, the visual layer information of the prior map, and the road region mask into a change detection model to output a change region mask; the processing module 820 is further configured to perform a fusion post - processing operation according to the change region mask to output obstacle information.

[0097] Figure 9 is a schematic structural diagram of a device 900 according to an embodiment of the present application. The device 900 can be used to execute Figure 7 the method in, or execute the method executed by the computing device 200 in the above text.

[0098] The device 900 includes: an acquisition module 910, configured to acquire training data, where the training data includes a camera image captured by a camera in a vehicle, visual layer information of a prior map, and a road region mask corresponding to the road region in the camera image, the prior map is pre - collected map information, and the road region mask is used to extract the road region in the camera image; a processing module 920, configured to input the training data into a change detection model to train the parameters of the change detection model, and the change detection model is used to output a change region mask.

[0099] Figure 10 is a schematic structural diagram of a device 1000 according to an embodiment of the present application. The device 1000 is used to execute Figure 6 or Figure 7 the method in, or execute the method executed by the computing device 200 in the above text.

[0100] The device 1000 includes a processor 1010, which is used to execute the computer programs or instructions stored in the memory 1020, or read the data stored in the memory 1020, so as to execute the methods in the above method embodiments. Optionally, the processor 1010 is one or more.

[0101] Optionally, as Figure 10 shown, the device 1000 further includes a memory 1020, which is used to store computer programs or instructions and / or data. The memory 1020 can be integrated with the processor 1010, or can be separately provided. Optionally, the memory 1020 is one or more.

[0102] Optionally, as Figure 10 shown, the device 1000 further includes a communication interface 1030, which is used for receiving and / or sending signals. For example, the processor 1010 is used to control the communication interface 1030 to receive and / or send signals.

[0103] Optionally, the device 1000 is used to implement the operations performed by the computing device in the above method embodiments.

[0104] For example, the processor 1010 is used to execute the computer programs or instructions stored in the memory 1020 to implement the related operations of the computing device in the above method embodiments. For example, the processor 1010 is used to: obtain a real-time camera image captured by a vehicle camera, and a road area mask corresponding to the road area in the camera image; obtain visual layer information of a prior map; input the camera image, the visual layer information of the prior map, and the road area mask into a change detection model to output a change area mask; perform a fusion post-processing operation according to the change area mask to output obstacle information.

[0105] For another example, the processor 1010 is used to: obtain training data, where the training data includes a camera image captured by a camera in a vehicle, visual layer information of a prior map, and a road area mask corresponding to the road area in the camera image; input the training data into a change detection model to train the parameters of the change detection model, and the change detection model is used to output a change area mask.

[0106] It should be noted that Figure 10 the device 1000 in

[0107] In the embodiments of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor can be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a GPU (which can be understood as a type of microprocessor), or a DSP, etc.; in another implementation, the processor can achieve certain functions through the logical relationship of a hardware circuit, and the logical relationship of this hardware circuit is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a type of ASIC, such as an NPU, a TPU, a DPU, etc.

[0108] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method. For example: a CPU, a GPU, an NPU, a TPU, a DPU, a microprocessor, a DSP, an ASIC, an FPGA, or a combination of at least two of these processor forms.

[0109] In addition, each unit in the above device can be integrated in whole or in part, or can be independently implemented. In one implementation, these units are integrated together and implemented in the form of a system-on-a-chip (SOC). The SOC can include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The types of the at least one processor can be different. For example, it includes a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0110] Correspondingly, the embodiments of the present application also provide a computer-readable storage medium storing a computer program. When the computer program / instructions are executed by a processor, the processor is caused to implement Figure 6 or Figure 7 the steps in the method in, or execute the method executed by the computing device 200 in the above text.

[0111] Correspondingly, the embodiments of the present application also provide a computer program product including computer program / instructions. When the computer program / instructions are executed by a processor, the processor is caused to implement Figures 6 to 7 the steps in the method in, or execute the method executed by the computing device 200 in the above text.

[0112] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0113] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0114] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0116] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0117] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0118] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0119] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the above elements.

[0120] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for detecting obstacles, characterized in that, including: obtaining a real-time camera image captured by a vehicle camera, and a road region mask corresponding to a road region in the camera image; obtaining visual layer information of a prior map; inputting the camera image, the visual layer information of the prior map, and the road region mask into a change detection model to output a change region mask; performing a fusion post-processing operation according to the change region mask to output obstacle information; the inputting the camera image, the visual layer information of the prior map, and the road region mask into a change detection model to output a change region mask includes: respectively performing feature extraction on the camera image, the visual layer information of the prior map, and the road region mask through a feature encoder to obtain a feature map of the camera image, a feature map of the visual layer information of the prior map, and a feature map of the road region mask; extracting difference comparison information between the feature map of the camera image and the feature map of the visual layer information of the prior map; performing correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road region mask to output high-level semantic features; performing decoding processing on the high-level semantic features through a feature decoder to output the change region mask.

2. The method according to claim 1, wherein the performing correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road region mask to output high-level semantic features includes: performing dimensionality reduction convolution on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road region mask to output a dimensionality-reduced feature map; inputting the dimensionality-reduced feature map and the difference comparison information into a correlation filtering module to output the high-level semantic features.

3. The method according to claim 1, characterized in that, the road region mask corresponding to the road region in the camera image is obtained by the following method: inputting the camera image into a panoramic segmentation model to output the road region mask, where the panoramic segmentation model is used for region division and category annotation of the camera image.

4. The method according to claim 1, characterized in that, the performing a fusion post-processing operation according to the change region mask to output obstacle information includes: extracting a dangerous roadblock region according to the change region mask; performing visual tracking filtering processing on the dangerous roadblock region; converting the position information of the dangerous roadblock region from two-dimensional coordinates to three-dimensional coordinates.

5. A device for detecting obstacles, characterized in that, including: an acquisition module for obtaining a real-time camera image captured by a vehicle camera, and a road region mask corresponding to a road region in the camera image; the acquisition module is further configured to obtain visual layer information of a prior map; a processing module for inputting the camera image, the visual layer information of the prior map, and the road region mask into a change detection model to output a change region mask; the processing module is further configured to perform a fusion post-processing operation according to the change region mask to output obstacle information; The processing module is specifically configured to: respectively perform feature extraction on the camera image, the visual layer information of the prior map, and the road area mask through a feature encoder to obtain a feature map of the camera image, a feature map of the visual layer information of the prior map, and a feature map of the road area mask; extract the difference comparison information between the feature map of the camera image and the feature map of the visual layer information of the prior map; perform correlation filtering on the difference comparison information according to the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output high-level semantic features; decode the high-level semantic features through a feature decoder to output the change area mask.

6. The device according to claim 5, wherein The processing module is specifically configured to: perform dimensionality reduction convolution on the feature map of the camera image, the feature map of the visual layer information of the prior map, and the feature map of the road area mask to output a dimensionality-reduced feature map; input the dimensionality-reduced feature map and the difference comparison information into a correlation filtering module to output the high-level semantic features.

7. The device according to claim 5, characterized in that, The acquisition module is specifically configured to: input the camera image into a panoramic segmentation model to output the road area mask, and the panoramic segmentation model is used to perform region division and label categories on the camera image.

8. The device according to claim 5, wherein The processing module is specifically configured to: extract a dangerous roadblock area according to the change area mask; perform visual tracking filtering processing on the dangerous roadblock area; convert the position information of the dangerous roadblock area from two-dimensional coordinates to three-dimensional coordinates.

9. An electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method for detecting obstacles according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by the processor, they are used to implement the method for detecting obstacles according to any one of claims 1 to 4.

11. A vehicle, characterized in that, A computing device is provided in the vehicle, and the computing device is used to execute the method for detecting obstacles according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Vehicle control method and device, computer equipment and computer readable storage medium

    CN111666921A

  • Obstacle detection method and device

    CN112149460A