A drivable area detection method based on target detection
By adopting a target detection method based on driving area detection, a pavement target detection model is constructed to extract pavement target information, and the problems of high computational complexity and insufficient detection accuracy in the prior art are solved, and more efficient and accurate detection effects are achieved.
Patent Information
- Application Number
- CN202210267631.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-03-17
AI Technical Summary
The existing driving area detection methods have high computational complexity and are difficult to run on general embedded devices, and the model training effect and detection accuracy are insufficient.
Using a feasible area detection method based on target detection, a pavement target detection model with several pavement target detection frames is constructed, and the model is trained to extract pavement target information and reduce the calculation complexity.
It realizes that while reducing the computational complexity, the model training effect and detection accuracy are ensured, and the accuracy and effect of driving area detection are improved.
Smart Images

Figure CN114581655B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road drivable area detection, and in particular to a drivable area detection method based on target detection. Background Art
[0002] Transportation vehicles are closely related to people's lives, and cars are the main personal transportation vehicles. Among them, drivable area detection is a key component of modern driver assistance systems. It is one of the key technologies of modern assistance and automatic driving systems. There are two main methods. One is to detect obstacles first and then estimate the drivable area. There are many types of obstacles and complex environments on actual roads, which makes it difficult to effectively detect obstacles. The second is to directly detect the drivable area. The direct detection method of the drivable area usually adopts algorithms based on binocular cameras, lidar sensors, and algorithms based on semantic segmentation networks.
[0003] In view of the problems in existing drivable area detection methods that laser radar is expensive and difficult to popularize, and binocular camera calibration accuracy is required to be high, the Chinese patent with publication number CN112418186A discloses "A drivable area detection method and device", which includes: using a monocular camera to obtain road traffic images, and constructing an initial tensor containing color channels based on the obtained road traffic images; for each road traffic image, obtaining the coordinate information of each pixel point in the road traffic image, adding the coordinate information of the pixel point to the color channel corresponding to the pixel point, and obtaining a tensor to be processed corresponding to the road traffic image; inputting the tensor to be processed into a pre-trained neural network containing an attention module to obtain a feature tensor representing the drivable area and the non-drivable area; based on the pixel points representing the drivable area in the feature tensor, constructing the outline of the drivable area, and controlling the vehicle to travel within the outline of the drivable area.
[0004] The drivable area detection method in the above existing scheme uses a semantic segmentation network to classify images at the pixel level, and then detects the drivable area of the road. However, the applicant found that the complexity of the existing semantic segmentation network is relatively high and it is difficult to run on general embedded devices. Therefore, how to design a detection method with lower computational complexity and higher efficiency, while ensuring the model training effect and detection accuracy, is particularly important and prominent. Summary of the invention
[0005] In view of the above-mentioned deficiencies in the prior art, the technical problem to be solved by the present invention is: how to provide a drivable area detection method based on target detection, which can reduce the computational complexity while ensuring the model training effect and detection accuracy, thereby improving the accuracy and effect of drivable area detection, and providing another feasible and efficient implementation method for drivable area detection.
[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A method for detecting a drivable area based on target detection comprises the following steps:
[0008] S1: Construct a road object detection model with several road object detection frames;
[0009] S2: construct a training data set including training images and corresponding training annotated images;
[0010] S3: During training, the training image is input into the road target detection model, and the road target information in the training image is extracted through each road target detection frame; then the road target detection model is trained based on the road target information extraction results of each road target detection frame combined with the training annotated image corresponding to the training image;
[0011] S4: During detection, the road surface image to be tested is input into a trained road surface target detection model, and the drivable area is determined by extracting road surface target information from each output road surface target detection frame.
[0012] Preferably, in step S1, the road object detection model uses a lightweight network as a backbone network, and the feature layer includes a plurality of feature layers of corresponding sizes corresponding to each road object detection frame; by performing 1x1 convolution, 3x3 convolution and maxpooling operations on the last feature layer of the backbone network, a plurality of corresponding feature layers are generated;
[0013] Among them, relatively large road target information is obtained on a feature layer with a relatively small size, and relatively small road target information is obtained on a feature layer with a relatively large size.
[0014] Preferably, in step S1, the center of the road surface object detection frame of the road surface object detection model is designed to be in an area one third of the distance from the bottom of the image.
[0015] Preferably, in step S2, a corresponding training annotated image is generated by annotating a plurality of rectangular frames on the training image.
[0016] Preferably, in step S2, each road object detection frame is serially numbered in order from left to right to generate a corresponding sequential relationship including a spatial structure.
[0017] Preferably, in step S3, the road target detection frame of the road target detection model is nearly overlapped with the marked ground target, that is, the extraction of road target information that can be used for training is completed when the top edge of the road target detection frame contacts the first non-road target.
[0018] Preferably, in step S3, a training loss function for training a road target detection model is generated based on the road target information extraction results of each road target detection frame and the training annotation image of the corresponding training image; the training loss function includes a regression loss function, a confidence loss function and a sequence loss function;
[0019] Among them, the regression loss function is used to locate the position of the road target, the confidence loss function is used to determine the road category, and the sequence loss function is used to characterize the spatial structural relationship between each road target.
[0020] Preferably, the training loss function is expressed as Loss = λ r ·L regression +λ c ·L confidence +λ s ·L sep ;
[0021] in,
[0022]
[0023]
[0024] Where: L regression represents the regression loss; L confidence represents the confidence loss; L sep represents sequence loss; λ r , c , s They represent the weight hyperparameters of regression loss, confidence loss, and sequence loss respectively; K represents the size of the feature map; M represents the number of road target detection boxes; x, y, w, h, and Represent the parameters of the road target detection frame corresponding to the true value and the predicted value respectively; the ijth road target detection frame extracts the road target information, then is 1, otherwise is 0; the ijth road target detection frame does not extract the road target information and is a negative sample, then is 1, otherwise is 0; represents the probability of the ijth road target detection box extracting the road target information; s represents the sequence number; SEQset represents the set of all sequence numbers; P i j (s) represents the probability that the serial number value of the ijth road target detection box is s; when the value in {·} is true, 1{·} takes 1, otherwise it takes 0; a noobj Represents the hyperparameters of the settings.
[0025] Preferably, in step S4, the road surface image to be tested is input into a trained road surface target detection model, and the corresponding road surface target detection frames are output; then, non-maximum suppression is performed on each road surface target detection frame to obtain a preliminary result frame; finally, adjacent preliminary result frames are merged according to a set adaptation method, and corresponding drivable road surface information is generated based on the corresponding road surface target information to determine the drivable area.
[0026] Preferably, in step S4, the trained road object detection model is converted into a corresponding tflite model, and the tflite model is run in a tensorflow lite environment built on the embedded device.
[0027] Compared with the prior art, the drivable area detection method of the present invention has the following beneficial effects:
[0028] The present invention trains the road target detection model through training images and corresponding training annotated images, so that the road target information that does not contain any obstacles can be extracted through the road target detection frame and the drivable area can be determined, thereby converting the road obstacle detection of the prior art into a simpler road target detection. On the one hand, compared with the existing method of first detecting obstacles and then estimating the drivable area, the present invention can greatly reduce the requirements and restrictions on training data, so that under the premise of limited number and type of training data, the training effect and detection accuracy of the detection model can be effectively guaranteed, thereby improving the accuracy and effect of drivable area detection; on the other hand, compared with the existing image full-pixel semantic segmentation network and other methods, the present invention can greatly reduce the computational complexity and at the same time guarantee the model training effect and detection accuracy. The present invention provides a new idea for road drivable area detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to make the purpose, technical solution and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:
[0030] Figure 1 It is a logic block diagram of the drivable area detection method;
[0031] Figure 2 Diagram of annotating images for training. DETAILED DESCRIPTION
[0032] The following is a further detailed description through specific implementation methods:
[0033] Example:
[0034] This embodiment discloses a drivable area detection method based on target detection.
[0035] like Figure 1 As shown, the drivable area detection method based on target detection includes the following steps:
[0036] S1: Construct a road object detection model with several road object detection frames;
[0037] S2: construct a training data set including training images and corresponding training annotated images;
[0038] S3: During training, the training image is input into the road target detection model, and the road target information in the training image is extracted through each road target detection frame; then the road target detection model is trained based on the road target information extraction results of each road target detection frame combined with the training annotated image corresponding to the training image;
[0039] S4: During detection, the road surface image to be tested is input into a trained road surface target detection model, and the drivable area is determined by extracting road surface target information from each output road surface target detection frame.
[0040] The present invention trains the road target detection model through training images and corresponding training annotated images, so that the road target information that does not contain any obstacles can be extracted through the road target detection frame and the drivable area can be determined, thereby converting the road obstacle detection of the prior art into a simpler road target detection. On the one hand, compared with the existing method of first detecting obstacles and then estimating the drivable area, the present invention can greatly reduce the requirements and restrictions on training data, so that under the premise of limited number and type of training data, the training effect and detection accuracy of the detection model can be effectively guaranteed, thereby improving the accuracy and effect of drivable area detection; on the other hand, compared with the existing image full-pixel semantic segmentation network and other methods, the present invention can greatly reduce the computational complexity and at the same time guarantee the model training effect and detection accuracy. The present invention provides a new idea for road drivable area detection.
[0041] In the specific implementation process, the road target detection model uses a lightweight network (mobilenet) as the backbone network, and the feature layer includes multiple feature maps of corresponding sizes corresponding to each road target detection frame (anchors); by performing 1x1 convolution, 3x3 convolution and maxpooling operations on the last feature layer of the backbone network, several corresponding feature layers are generated; among them, relatively large road target information is obtained on the feature layer with relatively small size, and relatively small road target information is obtained on the feature layer with relatively large size.
[0042] The road target detection model of the present invention can effectively extract road target information that does not contain any obstacles through the road target detection frame, thereby ensuring the accuracy of drivable area detection.
[0043] In the specific implementation process, the center of the road target detection frame of the road target detection model is designed to be one-third of the distance from the bottom of the image. Since the road surface is generally located below the image obtained by the vehicle camera, the road target detection frame (anchors) does not need to traverse the entire image.
[0044] The present invention designs the road target detection frame (anchors) in the area one-third of the distance from the bottom of the image, so that the calculation amount of the road target detection model when detecting in the drivable area can be reduced while ensuring the accuracy of road target information extraction, thereby improving the detection efficiency of the road target detection model.
[0045] In the specific implementation process, Figure 2 As shown, the corresponding training annotated image is generated by marking several rectangular boxes on the training image. Regarding the annotation of images in the training data, there are two types of image sources. First, a camera is installed on the car to collect a large number of images for road testing and manually annotate them. The annotation method is to annotate the road surface from left to right as some rectangular boxes similar to a histogram. The bottom of the rectangular box reaches the bottom of the road surface in the image, and the top reaches the top of the road surface close to the non-road target. Second, based on the pixel-level road annotation data set purchased from the open source or market, a conversion program is designed to convert the pixel-level annotation into the above marking method. Based on this method, a large amount of annotation data can be obtained. In order to increase the proportion of training data of real roads, the existing road detection algorithm can be used to infer the road images collected by itself to obtain a pixel-level road annotation data set. According to the above method 2, these data can also be converted into training annotated images for training.
[0046] In the specific implementation process, each road object detection frame is marked with a serial number from left to right to generate a corresponding sequential relationship including a spatial structure. In this embodiment, continuous reinforcement based on the loss function during the training process can enable the network to learn a deeper spatial structure relationship.
[0047] In the specific implementation process, the road target detection frame of the road target detection model is close to the marked ground target, that is, the road target information that can be used for training is extracted when the top edge of the road target detection frame contacts the first non-road target. In this way, the accuracy of the drivable area detection can be improved as much as possible.
[0048] In the specific implementation process, a training loss function for training a road target detection model is generated based on the road target information extraction results of each road target detection frame and the training annotation image of the corresponding training image; the training loss function includes a regression loss function, a confidence loss function and a sequence loss function;
[0049] Among them, the regression loss function is used to locate the position of the road target, the confidence loss function is used to determine the road category, and the sequence loss function is used to characterize the spatial structural relationship between each road target.
[0050] The training loss function is expressed as Loss = λ r ·L regression +λ c ·L confidence +λ s ·L sep ;
[0051] in,
[0052]
[0053]
[0054] Where: L regression represents the regression loss; L confidence represents the confidence loss; L sep represents sequence loss; λ r , c , s They represent the weight hyperparameters of regression loss, confidence loss, and sequence loss respectively; K represents the size of the feature map; M represents the number of road target detection boxes; x, y, w, h, and Represent the parameters of the road target detection frame corresponding to the true value and the predicted value respectively; the ijth road target detection frame extracts the road target information, then is 1, otherwise is 0; the ijth road target detection frame does not extract the road target information and is a negative sample, then is 1, otherwise is 0; represents the probability of the ijth road target detection box extracting the road target information; s represents the sequence number; SEQset represents the set of all sequence numbers; P i j (s) represents the probability that the serial number value of the ijth road target detection box is s; when the value in {·} is true, 1{·} takes 1, otherwise it takes 0; a noobj Represents the hyperparameters of the settings.
[0055] The present invention can effectively improve the training effect of the road target detection model through the above-mentioned training data set and training loss function, so that the road target detection model can accurately extract road target information, thereby effectively improving the accuracy of drivable area detection.
[0056] During the specific implementation process, during detection, the road surface image to be tested is input into the trained road surface target detection model, and the corresponding road surface target detection frames are output; then non-maximum suppression (NMS) is performed on each road surface target detection frame to obtain a preliminary result frame; finally, adjacent preliminary result frames are merged according to the set adaptation method, and the corresponding drivable road surface information is generated based on the corresponding road surface target information to determine the drivable area.
[0057] In this embodiment, the trained road object detection model is converted into a corresponding tflite model, and the tflite model is run in the tensorflow lite environment built on the embedded device.
[0058] Some existing methods, such as the semantic segmentation network method, will extract all road target information in the image when detecting the drivable area, that is, all road targets without obstacles. However, in some scenarios, it is not necessary to detect all road information. For example, when there is an obstacle (such as other vehicles) between the current vehicle position and the road directly in front, the road information ahead of the obstacle may not need to be detected. The road detection algorithm of the present invention can effectively reduce the complexity of detection by controlling the position and size of the road target detection frame so that the road target detection frame only extracts the road targets from the bottom of the image to the first obstacle, while ensuring the accuracy of the extraction of road target information, thereby further improving the accuracy of drivable area detection.
[0059] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described with reference to the preferred embodiments of the present invention, it should be understood by those skilled in the art that various changes can be made to it in form and detail without departing from the spirit and scope of the present invention as defined in the appended claims. At the same time, common sense such as the well-known specific structures and characteristics in the embodiments are not described in detail here. Finally, the scope of protection claimed by the present invention shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. A drivable area detection method based on target detection, characterized in that: The following steps are involved: S1: Construct a road object detection model with several road object detection frames; In step S1, the road object detection model uses a lightweight network as a backbone network, and the feature layer includes a plurality of feature layers of corresponding sizes corresponding to each road object detection frame; by performing 1x1 convolution, 3x3 convolution and maxpooling operations on the last feature layer of the backbone network, a plurality of corresponding feature layers are generated; Among them, relatively large road target information is obtained on a feature layer with a relatively small size, and relatively small road target information is obtained on a feature layer with a relatively large size; The center of the road object detection frame of the road object detection model is designed to be one-third of the distance from the bottom of the image; S2: construct a training data set including training images and corresponding training annotated images; S3: During training, the training image is input into the road target detection model, and the road target information in the training image is extracted through each road target detection frame; then the road target detection model is trained based on the road target information extraction results of each road target detection frame combined with the training annotated image corresponding to the training image; In step S3, the road target detection frame of the road target detection model is nearly overlapped with the marked ground target, that is, the extraction of road target information that can be used for training is completed when the top edge of the road target detection frame contacts the first non-road target; S4: During detection, the road surface image to be tested is input into a trained road surface target detection model, and the drivable area is determined by extracting road surface target information from each output road surface target detection frame.
2. The method for detecting a drivable area based on target detection according to claim 1, characterized in that: In step S2, a corresponding training annotated image is generated by annotating a number of rectangular boxes on the training image.
3. The method for detecting a drivable area based on target detection according to claim 1, characterized in that: In step S2, each road object detection frame is serially labeled from left to right to generate a corresponding sequential relationship including a spatial structure.
4. The method for detecting a drivable area based on target detection according to claim 1, characterized in that: In step S3, a training loss function for training a road target detection model is generated based on the road target information extraction results of each road target detection frame and the training annotation image of the corresponding training image; the training loss function includes a regression loss function, a confidence loss function and a sequence loss function; Among them, the regression loss function is used to locate the position of the road target, the confidence loss function is used to determine the road category, and the sequence loss function is used to characterize the spatial structural relationship between each road target.
5. The method for detecting a drivable area based on target detection according to claim 4, characterized in that: The training loss function is expressed as Loss = λ r ·L regression +λ c ·L confidence +λ s ·L sep ; in, Where: L regression represents the regression loss; L confidence represents the confidence loss; L sep represents sequence loss; λ r , c , s They represent the weight hyperparameters of regression loss, confidence loss, and sequence loss respectively; K represents the size of the feature map; M represents the number of road target detection boxes; x, y, w, h, and Represent the parameters of the road target detection frame corresponding to the true value and the predicted value respectively; the ijth road target detection frame extracts the road target information, then is 1, otherwise is 0; the ijth road target detection frame does not extract the road target information and is a negative sample, then is 1, otherwise is 0; represents the probability of the ijth road target detection box extracting the road target information; s represents the sequence number; SEQset represents the set of all sequence numbers; P i j (s) represents the probability that the serial number value of the ijth road target detection box is s; when the value in {·} is true, 1{·} takes 1, otherwise it takes 0; a noobj Represents the hyperparameters of the settings.
6. The method for detecting a drivable area based on target detection according to claim 1, characterized in that: In step S4, the road surface image to be tested is input into the trained road surface target detection model, and the corresponding road surface target detection frames are output; then, non-maximum suppression is performed on each road surface target detection frame to obtain a preliminary result frame; finally, adjacent preliminary result frames are merged according to the set adaptation method, and corresponding drivable road surface information is generated based on the corresponding road surface target information to determine the drivable area.
7. The method for detecting a drivable area based on target detection according to claim 1, characterized in that: In step S4, the trained road object detection model is converted into a corresponding tflite model, and the tflite model is run in the tensorflow lite environment built on the embedded device.
Citation Information
Patent Citations
Drivable area detection method and device
CN112418186A
Automobile drivable area planning method based on multi-task neural network
CN112418236A
Pavement disease detection method and device, terminal equipment and storage medium
CN113255605A