Vehicle drivable area detection method, device, electronic device and storage medium
By using a fisheye camera to acquire images and combine shared feature extraction networks and pre-trained models for processing, the problems of poor adaptability and high cost of vehicle driving area detection in the autonomous driving system are solved, and more efficient and accurate detection effects are achieved.
Patent Information
- Application Number
- CN202410213523.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-02-27
AI Technical Summary
The prior art has problems such as poor adaptability, low perceptual accuracy and low cost efficiency in the detection of vehicle travelable areas in autonomous driving systems.
The fisheye camera is used to collect the vehicle's surrounding environment images, extract the general feature map through the shared feature extraction network, and input it into the pre-trained semantic segmentation model and obstacle detection model for processing, and determine the final feasible area with the fusion strategy.
It improves the accuracy of vehicle driving area detection, reduces detection costs, improves detection efficiency, and makes the autonomous driving system more reliable and economical in various driving environments.
Smart Images

Figure CN118038409B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of new energy vehicles, and particularly to a method, device, electronic device and storage medium for detecting the drivable area of a vehicle. Background Art
[0002] The development of autonomous driving technology aims to achieve autonomous navigation and safe driving of vehicles in different environments. In this process, accurately detecting the freespace around the vehicle - that is, the drivable area without obstacles or other obstructions - is crucial. This requirement involves path planning, obstacle avoidance, and safe navigation in various traffic environments, including roads, lanes, intersections, and other scenarios.
[0003] In the prior art, freespace detection mainly includes the following methods:
[0004] Geometry-based method: This method relies on camera parameters and the geometric information of the image to identify freespace, but it has obvious deficiencies in dealing with complex scenes, unstructured information, and in-depth semantic understanding, especially when it is necessary to understand the semantic information of various objects (such as vehicles, pedestrians) in the environment.
[0005] Motion-based method: Such methods identify freespace by analyzing the optical flow or motion patterns captured by the camera, but they are insufficient in adapting to dynamic scenes, especially prone to misjudgment when encountering occlusion situations.
[0006] Semantic segmentation model-based method: Using deep learning models for semantic segmentation of images, although it improves the detection accuracy to a certain extent, it faces problems such as strong dependence on a large amount of labeled data, insufficient robustness in complex scenes, and misjudgment caused by obstacle occlusion.
[0007] Camera and lidar fusion-based method: This method attempts to improve the detection accuracy by fusing visual information and lidar data. However, the challenges it faces include high sensor costs and the complexity of the data fusion process.
[0008] Generally speaking, the prior art still faces problems such as poor adaptability, low perception accuracy, and low cost efficiency when dealing with freespace detection in complex environments. Therefore, there is an urgent need for an improved freespace detection method to improve the reliability and accuracy of the autonomous driving system in various driving environments. Summary of the Invention
[0009] In view of this, embodiments of the present application provide a method, device, electronic device, and storage medium for detecting a drivable area of a vehicle, so as to solve the problems of poor adaptability and perception accuracy, high detection cost, and low detection efficiency existing in the prior art for the drivable area detection method.
[0010] In the first aspect of the embodiments of the present application, a method for detecting a drivable area of a vehicle is provided, including: using a fisheye camera installed on the vehicle to collect image data of the surrounding environment of the vehicle, correcting and stitching the collected images to obtain a panoramic view; using a shared feature extraction network to extract features from the panoramic view to obtain a general feature map for semantic segmentation and obstacle detection; inputting the general feature map into the head network of a pre-trained semantic segmentation model for processing to generate a semantic segmentation result including a drivable area and a non-drivable area; inputting the general feature map into the head network of a pre-trained obstacle detection model for processing to obtain the grounding points corresponding to the obstacles in the panoramic view; using a predetermined fusion strategy to determine the final drivable area in the panoramic view according to the semantic segmentation result and the grounding points corresponding to the obstacles in the panoramic view, where the fusion strategy includes marking the non-drivable area after the grounding points as a drivable area.
[0011] In the second aspect of the embodiments of the present application, a device for detecting a drivable area of a vehicle is provided, including: an image processing module configured to use a fisheye camera installed on the vehicle to collect image data of the surrounding environment of the vehicle, correct and stitch the collected images to obtain a panoramic view; a feature extraction module configured to use a shared feature extraction network to extract features from the panoramic view to obtain a general feature map for semantic segmentation and obstacle detection; a semantic segmentation module configured to input the general feature map into the head network of a pre-trained semantic segmentation model for processing to generate a semantic segmentation result including a drivable area and a non-drivable area; an obstacle detection module configured to input the general feature map into the head network of a pre-trained obstacle detection model for processing to obtain the grounding points corresponding to the obstacles in the panoramic view; a fusion determination module configured to use a predetermined fusion strategy to determine the final drivable area in the panoramic view according to the semantic segmentation result and the grounding points corresponding to the obstacles in the panoramic view, where the fusion strategy includes marking the non-drivable area after the grounding points as a drivable area.
[0012] In the third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the steps of the above method when executing the computer program.
[0013] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0014] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects:
[0015] By using a fisheye camera installed on a vehicle, image data of the vehicle surrounding environment is collected, the collected images are corrected and stitched to obtain a panoramic view; a shared feature extraction network is used to extract features from the panoramic view to obtain a general feature map for semantic segmentation and obstacle detection; the general feature map is input into the head network of a pre-trained semantic segmentation model for processing to generate a semantic segmentation result including drivable areas and non-drivable areas; the general feature map is input into the head network of a pre-trained obstacle detection model for processing to obtain the grounding points corresponding to obstacles in the panoramic view; a predetermined fusion strategy is used to determine the final drivable area in the panoramic view according to the semantic segmentation result and the grounding points corresponding to obstacles in the panoramic view, where the fusion strategy includes marking the non-drivable area after the grounding points as a drivable area. The present application improves the accuracy of the drivable area detection result, reduces the detection cost, and improves the detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0017] Figure 1 is a schematic flowchart of a method for detecting a drivable area of a vehicle provided by an embodiment of the present application;
[0018] Figure 2 is a schematic structural diagram of a device for detecting a drivable area of a vehicle provided by an embodiment of the present application;
[0019] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.
[0021] Freespace refers to the drivable area around a vehicle where there are no obstacles or other obstructions. This concept includes roads, lanes, intersections, etc., areas where vehicles can drive freely and safely. Having good freespace detection and understanding is crucial for autonomous driving systems and can be used for path planning and navigation, obstacle avoidance and safety, parking and parking lot navigation, etc. The following are the existing technologies and disadvantages related to freespace detection schemes:
[0022] 1. Geometry-based methods: Use camera parameters and image geometry information for freespace detection. Disadvantages include poor adaptability to complex scenes, difficulty in processing unstructured information, perception inaccuracy, and lack of semantic understanding (unable to deeply understand the semantic information of objects in the environment, such as vehicles, pedestrians, bicycles, etc.).
[0023] 2. Motion-based methods: Detect freespace by analyzing the optical flow or motion patterns of the camera. Disadvantages include sensitivity to dynamic scenes and occlusions.
[0024] 3. Semantic segmentation models: Use deep learning models such as U-Net, deeplabv3+ etc. for image semantic segmentation. Disadvantages include the need for a large amount of labeled data, robustness challenges in complex scenes, and the problem of object occlusion. Obstacles may occlude the drivable area and turn it into an undrivable area, resulting in misjudgment by the system.
[0025] 4. Camera and lidar fusion: Combine visual information and lidar data to improve the accuracy of freespace detection. Disadvantages include sensor cost and the complexity of data fusion.
[0026] In view of the problems existing in the above-mentioned existing technologies, the present application realizes efficient acquisition of the freespace area by adopting the fusion of pure vision semantic segmentation and obstacle ground contact point detection, stitching BEV images through multiple fisheye cameras, sharing the model backbone, reducing the sensor cost and chip computing power requirements, and improving the robustness to obstacle occlusion, making the autonomous driving system more economical, lightweight and reliable.
[0027] To solve the problems of high sensor cost, the requirements of multi-sensor fusion for chip computing power and operation time, and the occlusion of obstacles in pure vision, this application adopts a method that combines pure vision semantic segmentation and obstacle grounding point detection to efficiently achieve freespace detection, with the following advantages:
[0028] 1. Reduce the sensor cost. Only four fisheye cameras are needed to stitch together a BEV image;
[0029] 2. The model is more lightweight. The semantic segmentation model of deeplabv3+ and the detection model of yolov5 share a backbone (i.e., a shared feature extraction network), and different heads (i.e., head networks) are used to implement a multi-task model. Therefore, the computing power and operation time of the chip are reduced, making the acquisition of the freespace area more efficient;
[0030] 3. For the problem of obstacle occlusion, first, the head of the semantic segmentation model of deeplabv3+ outputs the drivable area and non-drivable area of the BEV image, and the head of the detection model of yolov5 outputs the grounding points of the obstacles. By fusing the two heads, the true drivable area can be obtained, and the drivable area will not become a non-drivable area due to the occlusion of obstacles.
[0031] The following combines the accompanying drawings and specific embodiments to describe the content of the technical solution of this application in detail.
[0032] Figure 1 It is a schematic flowchart of the vehicle drivable area detection method provided by the embodiment of this application. Figure 1 The vehicle drivable area detection method can be executed by the vehicle-end autonomous driving system. As Figure 1 shown, the vehicle drivable area detection method can specifically include:
[0033] S101, Use the fisheye cameras installed on the vehicle to collect image data of the vehicle's surrounding environment, and perform correction and stitching processing on the collected images to obtain a panoramic view;
[0034] S102, Use the shared feature extraction network to extract features from the panoramic view to obtain a general feature map for semantic segmentation and obstacle detection;
[0035] S103, Input the general feature map into the head network of the pre-trained semantic segmentation model for processing to generate a semantic segmentation result including the drivable area and non-drivable area;
[0036] S104, Input the general feature map into the head network of the pre-trained obstacle detection model for processing to obtain the grounding points corresponding to the obstacles in the panoramic view;
[0037] S105. Using a predetermined fusion strategy, determine the final drivable area in the panoramic view according to the semantic segmentation result and the grounding points corresponding to the obstacles in the panoramic view, where the fusion strategy includes marking the non-drivable area after the grounding points as a drivable area.
[0038] In some embodiments, use a fisheye camera installed on the vehicle to collect image data of the vehicle's surrounding environment, and perform correction and stitching processing on the collected images to obtain a panoramic view, including: using four fisheye cameras installed in the front, rear, left, and right of the vehicle to collect image data of the vehicle's surrounding environment, correct the collected images, and use geometric transformation and image fusion techniques to stitch the corrected images into a panoramic view.
[0039] Specifically, the embodiments of the present application use a method of four fisheye cameras to acquire and process images of the vehicle's surrounding environment to form a panoramic view for vehicle drivable area detection. The following describes in detail how to collect images through cameras installed at key positions of the vehicle and how to efficiently process these images to support the decision-making process of the autonomous driving system.
[0040] Further, first, install fisheye cameras at four key positions in the front, rear, left, and right of the vehicle. This layout ensures that the vehicle can obtain a 360-degree environmental view, capturing detailed information in all directions, including different types of obstacles and road features. Synchronously collect the surrounding environment images of the vehicle through the above four fisheye cameras. These cameras should be able to provide high-quality image data under various lighting and weather conditions to ensure the accuracy and reliability of subsequent processing.
[0041] Further, correct the images obtained from each camera to eliminate the distortion effect of the fisheye lens. This step is crucial for improving the geometric accuracy of the images, ensuring that the subsequent stitched BEV (Bird's Eye View) image can maintain consistency during geometric transformation. At the same time, apply geometric transformation and image fusion techniques to stitch the four corrected images into a continuous panoramic BEV image. In this process, it is necessary to process the seams between the images to ensure that there is no obvious transition area at the stitching point, thereby generating a high-quality panoramic view.
[0042] In one example, in the formed BEV image, various types of obstacles can be recognized and labeled, such as buildings, pedestrians, other vehicles, cones, and water horses, etc., while clearly distinguishing the remaining drivable areas. This analysis is crucial for subsequent drivable area detection and path planning. In the autonomous driving system, this BEV image will be used to perform various tasks, such as path planning, obstacle avoidance, and navigation, etc. For example, in a parking lot environment, by analyzing the BEV image, the system can identify available parking spaces and provide a path for the vehicle to safely drive into the parking space while avoiding obstacles.
[0043] Through the method of the above embodiments, the embodiments of the present application not only provide an effective method for acquiring and processing images of the vehicle surrounding environment, but also ensure that this information can effectively support the decision-making and operation of autonomous driving vehicles, thereby improving the overall performance and safety of the autonomous driving system.
[0044] In some embodiments, the pre-training method of the semantic segmentation model includes:
[0045] Using a fisheye camera to collect environmental images of the vehicle in different scenarios, and performing correction and stitching processing on the environmental images to obtain an original panoramic image;
[0046] Using a preset data annotation tool to classify and label each pixel in the original panoramic image to obtain the semantic category corresponding to each pixel in the original panoramic image;
[0047] After the data annotation is completed, convert the annotation data from the original format to the target format, where the pixel value of each pixel in the target format represents the category to which the pixel belongs;
[0048] Divide the dataset composed of the annotation data in the target format into a training set, a validation set, and a test set, use the training set to train the semantic segmentation model, and use the validation set and the test set to verify and test the functions of the trained semantic segmentation model.
[0049] Specifically, the embodiments of the present application also provide a pre-training method for the semantic segmentation model. The semantic segmentation model of the embodiments of the present application can adopt the semantic segmentation model of DeepLabV3+. The following combines specific embodiments to detail the pre-training process of the semantic segmentation model, which can specifically include the following content:
[0050] Data preparation: Use a calibrated fisheye camera to collect pictures in a variety of different scenarios, and stitch these pictures into an AVM (All-in-View Monitoring System) image (the same as the BEV image in the above embodiments). This process involves image collection under various environments and conditions to ensure the diversity and comprehensiveness of the data.
[0051] Annotation tool selection: Public data annotation tools, such as labelme, are selected to perform semantic segmentation tasks. These tools can help annotators assign the correct semantic categories to each pixel in the image.
[0052] Data annotation process: Using the selected annotation tool, each pixel in the image is classified and labeled, and these categories may include the ego vehicle, pedestrians, other vehicles, obstacles, and buildings, etc. This step is the basis for achieving accurate semantic segmentation and requires precise identification and differentiation of different elements in the image.
[0053] Data format conversion: After annotation, the data is saved in an appropriate format. For the convenience of model training, the annotated data is converted from json format to png format, where the value of each pixel represents the category it belongs to.
[0054] Data splitting: The entire dataset is divided into a training set, a validation set, and a test set. The purpose of doing this is that during the model development process, the model can be effectively trained, its performance can be verified, and the final test can be carried out to ensure that the model has good generalization ability and accuracy.
[0055] In some embodiments, the general feature map is input into the head network of a pre-trained semantic segmentation model for processing to generate a semantic segmentation result including a drivable area and a non-drivable area, including: inputting the general feature map into the head network of the pre-trained semantic segmentation model, and using the head network of the pre-trained semantic segmentation model to classify each pixel in the general feature map into the corresponding semantic category to obtain a semantic segmentation feature map, where the semantic segmentation feature map contains a drivable area and a non-drivable area.
[0056] Specifically, the embodiments of the present application adopt a shared deep learning model Backbone, which simultaneously provides feature extraction for the semantic segmentation model (DeepLabV3+) and the obstacle detection model (YOLOv5). Using the shared Backbone reduces the requirements for computing power and operation time and improves the system efficiency. By using the shared Backbone, the Head of the semantic segmentation model and the Head of the detection model are respectively connected to form a multi-task model. The semantic segmentation Head is used to generate the drivable area and the non-drivable area of the BEV image, and the Head of the detection model is used to output the grounding points of the obstacles.
[0057] At the beginning of the model, the embodiment of the present application introduces a shared convolutional part (i.e., a shared feature extraction network) responsible for extracting general features (i.e., general feature maps). The shared feature extraction network part may include convolutional layers, activation functions, and pooling layers to learn low-level and mid-level features from the input image. It should be noted that this shared feature extraction network can adopt any effective feature extraction architecture, such as ResNet, etc.
[0058] In one example, the extracted general feature maps are input into a pre-trained DeepLabV3+ semantic segmentation head network. This head network uses advanced semantic analysis techniques, such as the ASPP module and upsampling techniques, to accurately classify each pixel while maintaining high resolution. The head network of the semantic segmentation model classifies each pixel in the general feature map into corresponding semantic categories, such as drivable areas, non-drivable areas, road markings, obstacles, etc.
[0059] Furthermore, after being processed by the head network, the resulting semantic segmentation feature maps depict the drivable and non-drivable areas in detail, providing an in-depth understanding and analysis of the vehicle's surrounding environment. The generation of these feature maps ensures accurate discrimination of various areas even in an environment with dense obstacles or complexity, providing a reliable basis for subsequent driving decisions.
[0060] For example, the image data of the vehicle's surrounding environment is first processed by the shared convolutional network to extract general feature maps describing the environment. Then, these feature maps are input into the pre-trained DeepLabV3+ model, which has been trained on a large number of highway and urban driving scenarios and can accurately identify and classify various road elements. The output of the DeepLabV3+ model is a high-resolution semantic segmentation map that clearly marks the drivable and non-drivable areas, as well as other relevant road information. In practical applications, this method can identify and distinguish roads, sidewalks, intersections, static obstacles, and dynamic obstacles, etc., providing detailed road and surrounding environment information for the autonomous driving system.
[0061] Through the method of the above embodiment, the pre-trained semantic segmentation model can be effectively utilized to process the general feature maps of the vehicle environment, generating accurate semantic segmentation feature maps, thereby improving the environmental perception ability and driving safety of the vehicle autonomous driving system. In addition, this embodiment also provides important information support for subsequent navigation decisions and path planning.
[0062] In some embodiments, the general feature map is input into the head network of a pre-trained obstacle detection model for processing to obtain the grounding points corresponding to the obstacles in the panoramic view, including: inputting the general feature map into the head network of the pre-trained obstacle detection model, and using the head network of the pre-trained obstacle detection model to identify the obstacles in the general feature map to determine the category of the obstacles and the position of the grounding points corresponding to the obstacles, where the grounding point represents the position where the obstacle contacts the ground.
[0063] Specifically, in the embodiments of the present application, the extracted general feature map is input into the head network of a pre-trained obstacle detection model (such as YOLOv5). This head network uses a multi-layer convolutional structure, activation functions, and pooling layers, as well as a specific YOLO detection layer, to specifically identify and classify the obstacles in the image. This head network is responsible for identifying the obstacles in the general feature map and determining the category of each obstacle and its grounding point, that is, the specific position of the intersection of the actual obstacle and the ground.
[0064] Furthermore, after being processed by the obstacle detection head network, an output containing the obstacle grounding point information is obtained. This information is crucial for understanding the spatial layout of the obstacles and their potential impact on the vehicle's driving path. Ensure that this output information accurately reflects the positions and categories of the various obstacles in the panoramic view, providing key data for subsequent navigation and obstacle avoidance decisions.
[0065] For example, in an example, the general feature map of the vehicle environment is processed by the obstacle detection network of YOLOv5, which has been trained with a large amount of traffic environment data to optimize its obstacle detection performance. The network output includes the positions (grounding points) and categories of the various obstacles, such as identifying and distinguishing pedestrians, vehicles, roadblocks, etc., providing detailed obstacle spatial information. In practical applications, this method can accurately identify vehicles parked by the roadside, pedestrians walking, or any object occupying the vehicle's driving path in complex traffic scenarios and accurately mark their grounding points.
[0066] Through the method of the above embodiments, high-precision obstacle detection and grounding point positioning can be effectively achieved, thereby improving the environmental perception ability and driving safety of the vehicle's autonomous driving system. In addition, this embodiment provides accurate obstacle information for the vehicle's path planning and obstacle avoidance strategies, further enhancing the reliability and efficiency of the autonomous driving system.
[0067] In some embodiments, determining the final drivable area in the panoramic view according to the semantic segmentation result and the grounding points corresponding to the obstacles in the panoramic view includes: marking the non-drivable area after the grounding points of the obstacles in the semantic segmentation feature map as a drivable area according to the positions of the grounding points corresponding to the obstacles, so as to obtain the final drivable area detection result, where the drivable area detection result includes the positions of the final drivable area and the obstacles.
[0068] Specifically, first, use a semantic segmentation model to process the images of the vehicle's surrounding environment to identify the semantic categories of each pixel in the image, such as drivable areas, road markings, obstacles, etc., so as to generate a semantic segmentation feature map, which includes drivable areas, non-drivable areas, and various obstacles. Secondly, identify the grounding points of each obstacle in the BEV image through an obstacle detection model. These grounding points are the positions where the obstacles contact the ground and provide key information for determining the actual occupied space of the obstacles.
[0069] Furthermore, the embodiment of the present application also provides a fusion layer, which is used to fuse the outputs of the object detection branch of YOLOv5 and the semantic segmentation branch of DeepLabV3+. For example, different fusion strategies (weighted fusion, cascaded fusion, etc.) can be adopted. The goal of fusion is to integrate the information of the two tasks to better support the overall performance of the model. In addition, the output layer is responsible for generating the final multi-task result. The output layer can process the object detection boxes and semantic segmentation results simultaneously, so that the final model output includes the information of the two tasks.
[0070] The following combines specific embodiments to detail the specific content of the fusion strategy provided by the present application, which can specifically include the following content:
[0071] Adjust the corresponding area markings in the semantic segmentation feature map according to the positions of the obstacle grounding points. Specifically, change the area after the obstacle grounding point from a non-drivable area to a drivable area to correct the misjudgment that may be caused by the obstacle projection. This process ensures that the finally determined drivable area excludes the wrong expansion area of the obstacle and more accurately reflects the space where the vehicle can safely drive.
[0072] In addition, comprehensive analysis and evaluation can also be performed on the fused drivable area detection result. Standard evaluation metrics and tests in actual driving scenarios can be used to ensure that this fusion strategy can accurately identify the drivable area in different road and traffic conditions.
[0073] In one example, the vehicle captures surrounding images through its environmental perception system and generates a BEV image. This image undergoes semantic segmentation processing to clearly distinguish roads, sidewalks, vehicles, and other obstacles. The obstacle detection model identifies the grounding points of each obstacle in the BEV image and adjusts the region markings in the semantic segmentation map accordingly. For example, if the grounding point of a stop sign is determined, the area after its vertical projection in the BEV map will be relabeled as a drivable area. The application of this fusion strategy significantly improves the accuracy of drivable area determination, especially in urban environments and complex traffic scenarios, and can effectively avoid misguidance caused by obstacle occlusion.
[0074] Through the method of the above embodiments, the embodiments of the present application provide a technology for accurately determining the drivable area around a vehicle, particularly strengthening the judgment ability in the case of obstacle occlusion, thereby enhancing the safety and reliability of the autonomous driving system. The application of this method not only improves the environmental perception accuracy of autonomous driving vehicles but also provides more accurate and reliable information for vehicle navigation and path planning.
[0075] In some embodiments, the semantic segmentation model includes a backbone network, an ASPP module, upsampling, and skip connections. The semantic segmentation model is used to output high-resolution semantic segmentation results. The obstacle detection model includes a series of convolutional layers, activation functions, pooling layers, and YOLO layers. The obstacle detection model is used to generate object detection boxes and obstacle category information.
[0076] Specifically, the semantic segmentation model of the embodiments of the present application adopts the DeepLabV3+ architecture, which includes an efficient backbone network, such as ResNet or MobileNetV2, for basic feature extraction. The ASPP (Atrous Spatial Pyramid Pooling) module is then applied to process image features of different scales, enhancing the model's ability to recognize objects of different sizes. Upsampling and skip connections are used to restore the image resolution, ensuring that the output semantic segmentation results have high resolution and can finely depict different semantic categories.
[0077] Furthermore, the obstacle detection model adopts the YOLOv5 architecture, which contains multiple layers of convolutional networks for extracting deep features in the image. These features are further processed through activation functions and pooling layers, and then the object detection boxes and obstacle category information are output through the YOLO layers. This model is specifically designed to achieve fast and efficient obstacle detection, and the generated detection boxes and category information are crucial for subsequent fusion analysis.
[0078] It should be noted that the embodiments of the present application can also adopt optimization algorithms and hardware acceleration technologies to ensure that the system can efficiently perform inference of the multi-task model in real-time scenarios to obtain the drivable area: using model fusion technology to connect the semantic segmentation model and the object detection model to a shared Backbone. Reducing the redundant parameters of the model through pruning technology to reduce the computational complexity. Selecting a dedicated deep learning chip (such as the BPU of J5) to improve the speed of model inference. Utilizing the parallelism of the hardware to perform inferences of semantic segmentation and object detection simultaneously. Adopting multi-task scheduling to ensure effective cooperation between the two tasks to improve the overall efficiency of the system. Using deep learning libraries and frameworks for hardware, such as CUDA, TensorRT, etc., to fully utilize the performance of the hardware acceleration technology. Optimizing the inference process of the model to reduce computational and memory overhead. Reducing the model size and improving the processing speed by model lightweighting, sharing the Backbone, and converting the model to int8.
[0079] By comprehensively adopting these strategies, it can be ensured that in real-time scenarios, the system can efficiently perform inference of the multi-task model, generate the results of the drivable area and detect obstacles. This is crucial for the real-time performance and accuracy in applications such as autonomous driving.
[0080] According to the technical solutions provided by the embodiments of the present application, the technical solutions of the present application have at least the following beneficial effects:
[0081] 1. Reduce sensor costs
[0082] By adopting the solution of only four fisheye cameras stitched into a BEV image, the cost of sensors has been successfully reduced. Compared with traditional multi-sensor fusion systems, this design not only meets the functional requirements, but also reduces the number and complexity of hardware devices, thus bringing significant cost savings.
[0083] 2. Model lightweighting
[0084] By sharing a single backbone and using different heads of the semantic segmentation model of deeplabv3+ and the detection model of yolov5, a lightweight multi-task model is achieved. This reduces the computing power and operation time requirements of the chip, and improves the efficient acquisition of the freespace area by the system in real-time scenarios.
[0085] 3. Improve robustness
[0086] For the problem of obstacle occlusion, the solution of the present application obtains the true drivable area by fusing the outputs of the semantic segmentation model and the obstacle detection model. This fusion method enhances the robustness to obstacle occlusion, ensures that the autonomous driving system can more accurately understand the environment, and reduces the risk of misjudgment.
[0087] 4. Efficient acquisition of the freespace area
[0088] Through the fusion of pure visual semantic segmentation and obstacle ground contact point detection, the solution of this application shows high efficiency in acquiring the freespace area. This method combines the advantages of two key models, enabling the autonomous driving system to understand the surrounding environment more accurately and improving the ability to quickly detect the drivable area of the road.
[0089] 5. Economical, lightweight and reliable
[0090] The solution of this application makes the autonomous driving system more economical because it reduces costs, chip computing power requirements, realizes the economy of sensors and the lightweight of the system. At the same time, by improving robustness and efficiently acquiring the freespace area, the reliability of the system is ensured, making it more suitable for various driving scenarios.
[0091] Generally speaking, the technical solution of this application has achieved remarkable beneficial effects in terms of accuracy, economy, lightweight, robustness, etc., providing strong support for the further development of autonomous driving technology.
[0092] The following is an embodiment of the device of this application, which can be used to execute the method embodiment of this application. For the details not disclosed in the device embodiment of this application, please refer to the method embodiment of this application.
[0093] Figure 2 It is a schematic structural diagram of a vehicle drivable area detection device provided by an embodiment of this application. As Figure 2 shown, the vehicle drivable area detection device includes:
[0094] An image processing module 201, configured to collect image data of the surrounding environment of the vehicle by using a fisheye camera installed on the vehicle, and perform correction and stitching processing on the collected images to obtain a panoramic view;
[0095] A feature extraction module 202, configured to extract features from the panoramic view by using a shared feature extraction network to obtain a general feature map for semantic segmentation and obstacle detection;
[0096] A semantic segmentation module 203, configured to input the general feature map into the head network of a pre-trained semantic segmentation model for processing to generate a semantic segmentation result including a drivable area and a non-drivable area;
[0097] An obstacle detection module 204, configured to input the general feature map into the head network of a pre-trained obstacle detection model for processing to obtain the ground contact points corresponding to the obstacles in the panoramic view;
[0098] The fusion determination module 205 is configured to use a predetermined fusion strategy to determine the final drivable area in the panoramic view according to the semantic segmentation result and the grounding point corresponding to the obstacle in the panoramic view, where the fusion strategy includes marking the non-drivable area after the grounding point as a drivable area.
[0099] In some embodiments, Figure 2 The image processing module 201 collects image data of the vehicle surrounding environment by using four fisheye cameras installed in the front, rear, left, and right of the vehicle, corrects the collected images, and stitches the corrected images into a panoramic view by using geometric transformation and image fusion techniques.
[0100] In some embodiments, Figure 2 The model pre-training module 206 collects environmental images of the vehicle in different scenarios by using fisheye cameras, corrects and stitches the environmental images to obtain an original panoramic image; classifies and labels each pixel in the original panoramic image by using a preset data annotation tool to obtain the semantic category corresponding to each pixel in the original panoramic image; after the data annotation is completed, converts the annotation data from the original format to a target format, where the pixel value of each pixel in the target format represents the category to which the pixel belongs; divides the dataset composed of the annotation data in the target format into a training set, a validation set, and a test set, trains the semantic segmentation model by using the training set, and validates and tests the function of the trained semantic segmentation model by using the validation set and the test set.
[0101] In some embodiments, Figure 2 The semantic segmentation module 203 inputs the general feature map into the head network of the pre-trained semantic segmentation model, and uses the head network of the pre-trained semantic segmentation model to classify each pixel in the general feature map into the corresponding semantic category to obtain a semantic segmentation feature map, where the semantic segmentation feature map includes a drivable area and a non-drivable area.
[0102] In some embodiments, Figure 2 The obstacle detection module 204 inputs the general feature map into the head network of the pre-trained obstacle detection model, and uses the head network of the pre-trained obstacle detection model to identify the obstacles in the general feature map to determine the category of the obstacles and the position of the grounding point corresponding to the obstacles, where the grounding point represents the position where the obstacle contacts the ground.
[0103] In some embodiments, Figure 2 The fusion determination module 205 marks the non-drivable area after the grounding point position of the obstacle in the semantic segmentation feature map as a drivable area according to the grounding point position corresponding to the obstacle to obtain the final drivable area detection result, where the drivable area detection result includes the final drivable area and the position of the obstacle.
[0104] In some embodiments, the semantic segmentation model includes a backbone network, an ASPP module, upsampling, and skip connections. The semantic segmentation model is used to output high-resolution semantic segmentation results. The obstacle detection model includes a series of convolutional layers, activation functions, pooling layers, and YOLO layers. The obstacle detection model is used to generate object detection boxes and obstacle category information.
[0105] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0106] Figure 3 It is a schematic structural diagram of the electronic device 3 provided by the embodiment of the present application. As Figure 3 shown, the electronic device 3 in this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 301 executes the computer program 303, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0107] Exemplarily, the computer program 303 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 302 and executed by the processor 301 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 303 in the electronic device 3.
[0108] The electronic device 3 can be a desktop computer, a notebook, a palm computer, a cloud server, and other electronic devices. The electronic device 3 can include, but is not limited to, the processor 301 and the memory 302. Those skilled in the art can understand that Figure 3 this is only an example of the electronic device 3, and does not constitute a limitation to the electronic device 3. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0109] The processor 301 may be a Central Processing Unit (CPU), or may be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0110] The memory 302 may be an internal storage unit of the electronic device 3. For example, it may be the hard disk or memory of the electronic device 3. The memory 302 may also be an external storage device of the electronic device 3. For example, it may be a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 3. Further, the memory 302 may also include both the internal storage unit and the external storage device of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device. The memory 302 may also be used to temporarily store data that has been output or is to be output.
[0111] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0112] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0113] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0114] In the embodiments provided in this application, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0115] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0116] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0117] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0118] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for detecting a vehicle drivable area, characterized in that: include: The fisheye camera installed on the vehicle is used to collect image data of the vehicle's surrounding environment, and the collected images are corrected and stitched to obtain a panoramic view; Extracting features from the panoramic view using a shared feature extraction network to obtain a common feature map for semantic segmentation and obstacle detection, wherein the shared feature extraction network is respectively connected to a semantic segmentation model and an obstacle detection model to form a multi-task model; Inputting the universal feature map into the head network of the pre-trained semantic segmentation model for processing to generate a semantic segmentation result including a drivable area and a non-drivable area; Inputting the universal feature map into the head network of the pre-trained obstacle detection model for processing to obtain the ground point corresponding to the obstacle in the panoramic view; Determine a final drivable area in the panoramic view according to the semantic segmentation result and the grounding point corresponding to the obstacle in the panoramic view by using a predetermined fusion strategy, wherein the fusion strategy includes marking a non-drivable area after the grounding point as a drivable area; The step of determining a final drivable area in the panoramic view according to the semantic segmentation result and the grounding point corresponding to the obstacle in the panoramic view includes: According to the grounding point position corresponding to the obstacle, the non-drivable area after the obstacle grounding point position in the semantic segmentation feature map is marked as a drivable area to obtain a final drivable area detection result, wherein the drivable area detection result includes the position of the final drivable area and the obstacle.
2. The method according to claim 1, characterized in that The fisheye camera installed on the vehicle is used to collect image data of the vehicle's surrounding environment, and the collected images are corrected and spliced to obtain a panoramic view, including: Four fisheye cameras installed on the front, rear, left and right sides of the vehicle are used to collect image data of the vehicle's surrounding environment, the collected images are corrected, and the corrected images are spliced into the panoramic view using geometric transformation and image fusion technology.
3. The method according to claim 1, characterized in that The pre-training method of the semantic segmentation model includes: Using a fisheye camera to collect environmental images of the vehicle in different scenes, and correcting and stitching the environmental images to obtain an original panoramic image; Using a preset data annotation tool to classify and annotate each pixel in the original panoramic image, so as to obtain a semantic category corresponding to each pixel in the original panoramic image; After the data annotation is completed, the annotated data is converted from an original format to a target format, wherein the pixel value of each pixel in the target format indicates the category to which the pixel belongs; The data set consisting of the annotated data in the target format is divided into a training set, a validation set and a test set, the semantic segmentation model is trained using the training set, and the function of the trained semantic segmentation model is verified and tested using the validation set and the test set.
4. The method according to claim 1, characterized in that: The step of inputting the general feature map into the head network of the pre-trained semantic segmentation model for processing to generate a semantic segmentation result including a drivable area and a non-drivable area includes: The universal feature map is input into the head network of a pre-trained semantic segmentation model, and each pixel in the universal feature map is classified into a corresponding semantic category using the head network of the pre-trained semantic segmentation model to obtain a semantic segmentation feature map, wherein the semantic segmentation feature map contains a drivable area and a non-drivable area.
5. The method according to claim 1, characterized in that The step of inputting the general feature map into a head network of a pre-trained obstacle detection model for processing to obtain a ground point corresponding to an obstacle in the panoramic view includes: The universal feature map is input into the head network of a pre-trained obstacle detection model, and the head network of the pre-trained obstacle detection model is used to identify obstacles in the universal feature map to determine the category of the obstacle and the grounding point position corresponding to the obstacle, wherein the grounding point represents the position where the obstacle contacts the ground.
6. The method according to claim 1, characterized in that The semantic segmentation model includes a backbone network, an ASPP module, upsampling and skip connections, and is used to output high-resolution semantic segmentation results. The obstacle detection model includes a series of convolutional layers, activation functions, pooling layers and YOLO layers, and is used to generate target detection frames and obstacle category information.
7. A vehicle drivable area detection device, characterized in that: include: The image processing module is configured to collect image data of the vehicle's surrounding environment using a fisheye camera installed on the vehicle, and to correct and stitch the collected images to obtain a panoramic view; a feature extraction module configured to extract features from the panoramic view using a shared feature extraction network to obtain a common feature map for semantic segmentation and obstacle detection, wherein the shared feature extraction network is respectively connected to a semantic segmentation model and an obstacle detection model to form a multi-task model; A semantic segmentation module is configured to input the universal feature map into a head network of a pre-trained semantic segmentation model for processing, and generate a semantic segmentation result including a drivable area and a non-drivable area; An obstacle detection module is configured to input the universal feature map into a head network of a pre-trained obstacle detection model for processing, so as to obtain a ground point corresponding to an obstacle in the panoramic view; a fusion determination module configured to determine a final drivable area in the panoramic view according to the semantic segmentation result and a grounding point corresponding to the obstacle in the panoramic view using a predetermined fusion strategy, wherein the fusion strategy includes marking a non-drivable area after the grounding point as a drivable area; Among them, the fusion determination module is used to mark the non-drivable area after the obstacle grounding point position in the semantic segmentation feature map as a drivable area according to the grounding point position corresponding to the obstacle, so as to obtain the final drivable area detection result, wherein the drivable area detection result includes the position of the final drivable area and the obstacle.
8. An electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Obstacle detection method and device
CN113841154A
Driving area detection method and device, electronic equipment and storage medium
CN114596544A
Automatic driving multi-task visual perception method
CN116824537A
Method and system for optimizing drivable area based on target detection results
CN117496468A