Training method, detection method and equipment for detection model of drivable area

By labeling sample points in the sample image and training the neural network, combining position labels and category weights, and optimizing the detection model, the problem of time-consuming and low accuracy in detecting travelable areas in the prior art is solved, and more efficient and accurate detection is achieved.

CN120279515APending Publication Date: 2025-07-08BEIJING YINWO AUTOMOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510346353.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art calculates the process of detecting a travelable area and relies on image segmentation and post-processing, affecting the detection accuracy.

Method used

By acquiring sample images and labeling sample points, the detection model is trained using the initial neural network, combining position labels, category weights and loss functions, the neural network is optimized to improve detection accuracy.

Benefits of technology

It improves the accuracy and efficiency of detecting travelable areas, reduces dependence on image segmentation, and enhances the recognition ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279515A_ABST
    Figure CN120279515A_ABST
Patent Text Reader

Abstract

The invention provides a training method, a detection method and equipment for a detection model of a drivable area, and the method comprises the steps: obtaining a sample image shot by a fisheye camera on a vehicle, the sample image comprises a plurality of sample points, and the labels of the sample points comprise a sample position label and a sample type; the sample image is input into the initial neural network, a plurality of prediction points and prediction results corresponding to the prediction points are obtained, and the prediction results comprise prediction position labels, prediction down-sampling error values and prediction categories. Determining a position label loss function based on the predicted position label and a category weight corresponding to the predicted category; determining a sampling error loss function based on the predicted downsampling error value; a category loss function is determined based on the sample category and the predicted category. Training the initial neural network by using a position label loss function, a sampling error loss function and a category loss function to obtain a trained detection model;
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and specifically relates to a training method, a detection method, and a device for a detection model of a drivable area. Background Art

[0002] With the development of autonomous driving technology, more and more vehicles are equipped with systems that can implement autonomous driving functions or assisted driving functions. To achieve autonomous driving, it is necessary to identify the drivable area from the road environment around the vehicle, so as to control the vehicle to drive in the drivable area.

[0003] Currently, when detecting the drivable area, a full-image segmentation method is usually adopted. The road surface, vehicles, sidewalks, and other areas are segmented in the image, and then the contour points of the drivable area are found through post-processing. The entire calculation process takes a long time, and the final detection result depends not only on the effect of image segmentation but also on the post-processing method, which affects the accuracy of detecting the drivable area. Summary of the Invention

[0004] In view of this, the present application is committed to providing a training method, a detection method, and a device for a detection model of a drivable area, so as to improve the accuracy of detecting the drivable area.

[0005] In a first aspect, the present application provides a training method for a detection model of a drivable area, and the method includes:

[0006] Obtain a sample image, where multiple sample points are marked in the sample image, and the labels of the sample points include: a sample position label and a sample category. The sample position label is used to indicate whether the sample point is a positive sample point or a negative sample point, and multiple positive sample points form the boundary of the drivable area;

[0007] Input the sample image into an initial neural network to obtain multiple prediction points and the prediction results corresponding to the prediction points. The prediction results include: a prediction position label, a predicted downsampling error value, and a prediction category. The predicted downsampling error value is determined based on the position of the prediction point and the position of the corresponding sample point of the prediction point in the sample image;

[0008] Based on the prediction position label and the category weight corresponding to the prediction category, determine a position label loss function; based on the predicted downsampling error value, determine a sampling error loss function; based on the sample category and the prediction category, determine a category loss function;

[0009] Train the initial neural network based on the position label loss function, the sampling error loss function, and the category loss function until the training cutoff condition is met, and obtain a trained detection model.

[0010] In one possible implementation, determining the position label loss function based on the predicted position label and the class weight corresponding to the predicted class includes:

[0011] Determine a plurality of first prediction points corresponding to the positions of the positive sample points among the plurality of prediction points;

[0012] Based on the predicted position label of the first prediction point, the class weight corresponding to the predicted class of the first prediction point, the hyperparameters of the balance factor and the modulation factor, determine the first positive loss function;

[0013] Based on the plurality of first positive loss functions corresponding to the plurality of first prediction points, determine the second positive loss function;

[0014] Based on the predicted position label of the second prediction point, the balance factor, and the hyperparameters of the modulation factor, determine the first negative loss function, where the second prediction point represents any other prediction point among the prediction points except the plurality of first prediction points;

[0015] Based on the plurality of first negative loss functions corresponding to the plurality of second prediction points, determine the second negative loss function;

[0016] Based on the average value of the second positive loss function and the second negative loss function, determine the position label loss function.

[0017] In one possible implementation, determining the second positive loss function based on the plurality of first positive loss functions corresponding to the plurality of first prediction points includes:

[0018] Calculate the sum of the plurality of first positive loss functions to obtain the second positive loss function;

[0019] The determining the second negative loss function based on the plurality of first negative loss functions corresponding to the plurality of second prediction points includes:

[0020] Calculate the average value of the plurality of first negative loss functions to obtain the second negative loss function.

[0021] In one possible implementation, obtaining the predicted downsampling error value corresponding to the prediction point includes:

[0022] Determine the product between the pixel position of the prediction point and the downsampling multiple of the initial neural network;

[0023] Based on the error between the position of the sample point corresponding to the prediction point in the sample image and the product, obtain the predicted downsampling error value.

[0024] In a possible implementation, the positive sample points are obtained by equally spacing and sampling the boundaries of the drivable area in the sample image according to a preset number of pixel columns.

[0025] In a second aspect, the present application provides a method for detecting a drivable area, the method comprising:

[0026] Obtaining a to-be-detected image captured by a camera on a vehicle;

[0027] Inputting the to-be-detected image into the detection model according to any one of the implementations in the first aspect above to obtain a plurality of detection points and detection results corresponding to the detection points;

[0028] Determining boundary points of the drivable area based on the detection results, so as to obtain the drivable area of the vehicle based on the boundary points.

[0029] In a possible implementation, the detection results include position tags, and determining boundary points of the drivable area based on the detection results includes:

[0030] When the confidence corresponding to the position tag is greater than or equal to a preset threshold, determining the detection point corresponding to the position tag as a boundary point of the drivable area.

[0031] In a possible implementation, the method further comprises:

[0032] Calculating, based on the pixel positions of the boundary points and the downsampling multiple of the detection model, the original pixel positions corresponding to the boundary points in the to-be-detected image;

[0033] Using an inverse perspective transformation method to determine the coordinates corresponding to the original pixel positions in a bird's-eye view coordinate system;

[0034] Obtaining the drivable area of the vehicle in the bird's-eye view coordinate system based on the coordinates.

[0035] In a possible implementation, the camera comprises four fisheye cameras, the detection results further include categories, and when the boundary points are located in the fusion area of two adjacent fisheye cameras, the method further comprises:

[0036] For a first fisheye camera, tracking the number of consecutive frames in which the boundary point appears in the image captured by the first fisheye camera, and determining a tracking score based on the number of frames, where the first fisheye camera represents any one of the two adjacent fisheye cameras;

[0037] Determining a category score based on the category of the boundary point;

[0038] Determining a confidence score based on the confidence corresponding to the position tag of the boundary point;

[0039] Based on the tracking score, the category score, and the confidence score, determine a total score;

[0040] In the fusion region, retain the boundary points determined by the fisheye camera with a higher total score among two adjacent fisheye cameras.

[0041] In a third aspect, the present application provides a training device for a drivable area detection model, the device including:

[0042] A sample acquisition unit, configured to obtain sample images captured by a camera on a vehicle, where multiple sample points are marked in the sample images, and the labels of the sample points include: a sample position label and a sample category, the sample position label is used to indicate that the label corresponding to the sample point position is a positive sample point or a negative sample point, and multiple positive sample points form the boundary of the drivable area;

[0043] A prediction unit, configured to input the sample image into an initial neural network to obtain multiple prediction points and prediction results corresponding to the prediction points, the prediction results including: a predicted position label, a predicted downsampling error value, and a predicted category, the predicted downsampling error value being determined based on the position of the prediction point and the position of the sample point corresponding to the prediction point in the sample image;

[0044] A loss function determination unit, configured to determine a position label loss function based on the predicted position label and the category weight corresponding to the predicted category; determine a sampling error loss function based on the predicted downsampling error value; determine a category loss function based on the sample category and the predicted category;

[0045] A training unit, configured to train the initial neural network based on the position label loss function, the sampling error loss function, and the category loss function until a training cutoff condition is met, to obtain a trained detection model.

[0046] In a fourth aspect, the present application provides a drivable area detection device, the device including:

[0047] An image acquisition unit, configured to obtain a to-be-detected image captured by a camera on a vehicle;

[0048] A detection unit, configured to input the to-be-detected image into the detection model according to any one of the implementation manners of the first aspect above to obtain multiple detection points and detection results corresponding to the detection points;

[0049] A determination unit, configured to determine boundary points of the drivable area based on the detection results, so as to obtain the drivable area of the vehicle based on the boundary points.

[0050] Fifth aspect, the present application provides an electronic device, the device comprising: a memory and a processor;

[0051] The memory is used for storing relevant program codes;

[0052] The processor is used for calling the program codes to execute the training method of the drivable area detection model according to any implementation manner of the first aspect or the detection method of the drivable area according to any implementation manner of the second aspect.

[0053] Sixth aspect, the present application provides a computer-readable storage medium, the computer-readable storage medium is used for storing a computer program, and the computer program is used for executing the training method of the drivable area detection model according to any implementation manner of the first aspect or the detection method of the drivable area according to any implementation manner of the second aspect.

[0054] Seventh aspect, the present application provides a computer program product, the computer program product comprises computer programs / instructions, and when the computer programs / instructions are executed by a processor, the training method of the drivable area detection model according to any implementation manner of the first aspect or the detection method of the drivable area according to any implementation manner of the second aspect is implemented.

[0055] In the above implementation manner of the present application, in order to train a detection model for detecting the drivable area of a vehicle, it is necessary to prepare training samples in advance, that is, obtain sample images captured by a camera on the vehicle, and multiple sample points are marked on the sample images. The label of each sample point includes: a sample position label and a sample category. Among them, the sample position label is used to indicate whether the sample point is a positive sample point or a negative sample point, and can identify the position of the sample point. The boundary of the drivable area is composed of multiple positive sample points. That is, the positive sample points are the boundary points of the drivable area. Then, the sample images are input into the initial neural network, and multiple prediction points output by the initial neural network and the prediction result corresponding to each prediction point are obtained. Among them, the prediction result includes: a prediction position label, a predicted downsampling error value, and a prediction category. The prediction position label is used to indicate whether the label corresponding to the prediction point position is a positive sample point or a negative sample point. The predicted downsampling error value is determined based on the position of the prediction point and the position of the sample point corresponding to the prediction point in the sample image. Among them, the sample category of the positive sample points includes multiple categories, and each category corresponds to a category weight. That is, the prediction category also corresponds to a category weight. The position label loss function can be determined based on the prediction position label and the category weight corresponding to the prediction category; the sampling error loss function is determined based on the predicted downsampling error value; the category loss function is determined based on the sample category and the prediction category. Then, the initial neural network is trained using the position label loss function, the sampling error loss function, and the category loss function until the training cutoff condition is met, and a trained detection model is obtained. Subsequently, the to-be-detected image around the vehicle can be input into the detection model to obtain the drivable area of the vehicle. Through the method provided by the present application, a neural network can be trained using pre-annotated sample images to obtain a detection model, and the loss function is determined by combining positive and negative samples to improve the accuracy of training the neural network, thereby improving the accuracy of detecting the drivable area using the detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments provided in the present application. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a flowchart of a training method for a detection model of a drivable area provided by an embodiment of the present application.

[0058] Figure 2 It is a schematic diagram for determining positive sample points provided by an embodiment of the present application.

[0059] Figure 3aSchematic diagram of a structure of an initial neural network provided by an embodiment of the present application.

[0060] Figure 3b Schematic diagram of a structure of a residual block provided by an embodiment of the present application.

[0061] Figure 4 Flowchart of a method for detecting a drivable area provided by an embodiment of the present application.

[0062] Figure 5 Schematic diagram of a shooting range of a four-way fisheye camera provided by an embodiment of the present application.

[0063] Figure 6 Schematic diagram of a training device of a detection model for a drivable area provided by an embodiment of the present application.

[0064] Figure 7 Schematic diagram of a detection device for a drivable area provided by an embodiment of the present application.

[0065] Figure 8 Schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0066] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. The described embodiments are only exemplary embodiments of the present application and not all implementation manners. Those skilled in the art can obtain other embodiments without creative work in combination with the embodiments of the present application, and these embodiments are also within the protection scope of the present application.

[0067] In order to achieve autonomous driving, it is necessary to identify the drivable area from the road environment around the vehicle, so as to control the vehicle to drive in the drivable area. Currently, when detecting the drivable area, a full-image segmentation method is usually adopted, where areas such as the road surface, vehicles, and sidewalks are segmented in the image, and then the contour points of the drivable area are found through post-processing. The entire calculation process takes a long time, and the final detection result depends not only on the effect of image segmentation but also on the post-processing method, which affects the accuracy of detecting the drivable area.

[0068] Based on this, the embodiment of the present application provides a training method, a detection method and a device for a detection model of a drivable area, so as to improve the accuracy of detecting the drivable area. In the specific implementation, a sample image taken by a camera on a vehicle is obtained, and a plurality of sample points are marked in the sample image, and the label of each sample point includes: a sample position label and a sample category as the sample true value. Among them, the sample position label is used to indicate that the sample point is a positive sample point or a negative sample point, and the position of the sample point can be identified, and the boundary of the drivable area is composed of a plurality of positive sample points. That is, the boundary points of the drivable area of ​​the vehicle are marked in the sample image as positive sample points. Then the sample image is input into the initial neural network, and a plurality of prediction points output by the initial neural network and the prediction results corresponding to each prediction point are obtained, wherein the prediction results include: a prediction position label, a prediction downsampling error value and a prediction category, and the prediction position label is used to indicate that the label corresponding to the prediction point position is a positive sample point or a negative sample point, and the prediction downsampling error value is determined based on the position of the prediction point and the position of the sample point corresponding to the prediction point in the sample image. Among them, the sample category of the positive sample point includes multiple categories, and a category weight is pre-set for each category. For each prediction point, the prediction category corresponding to the prediction point will be output, and the category weight corresponding to the prediction category can be obtained from the correspondence between the category weights corresponding to the above-mentioned pre-set different categories. The position label loss function can be determined based on the predicted position label and the category weight corresponding to the prediction category; the sampling error loss function can be determined based on the predicted downsampling error value; the category loss function can be determined based on the sample category and the prediction category. Then, the initial neural network is trained using the position label loss function, the sampling error loss function, and the category loss function until the training cutoff condition is met to obtain a trained detection model. Thereby, the image to be detected around the vehicle can be subsequently input into the detection model to obtain the drivable area of ​​the vehicle. Through the method provided in the present application, a neural network can be trained using pre-labeled sample images to obtain a detection model, and the loss function can be determined in combination with positive and negative samples to improve the accuracy of the trained neural network, thereby improving the accuracy of detecting the drivable area using the detection model.

[0069] In order to facilitate the understanding of the technical solutions provided by the embodiments of the present application, a detailed introduction will be given below in conjunction with the drawings in the specification.

[0070] In order to improve the accuracy of detecting the drivable area, it is necessary to obtain a detection model for detecting the drivable area. The following first introduces the training process of obtaining the detection model.

[0071] See also Figure 1 As shown, it is a flowchart of a method for training a detection model of a drivable area provided in an embodiment of the present application.

[0072] This method can be executed by an image processing device and mainly includes the following steps:

[0073] S101: Obtain a sample image.

[0074] The sample image can be made from an image captured by a camera on a vehicle. Among them, multiple sample points are marked on the sample image, and the label of each sample point includes: a sample position label and a sample category.

[0075] That is, after obtaining the sample image captured by the camera, the sample image can be marked as the sample ground truth for training. Since the training model has the ability to detect the drivable area, when marking the sample image, it is necessary to mark the drivable area in the sample image. For example, the boundary points that make up the drivable area can be marked as positive sample points. The other pixel positions in the sample image except the positive sample points can be marked as negative sample points.

[0076] Among them, the label of the sample point includes a sample position label and a sample category as the sample ground truth. The sample position label is used to indicate that the label corresponding to the position where different sample points are located is a positive sample point or a negative sample point. For example, the sample position label corresponding to the position where the positive sample point is located can be 1, and the sample position label corresponding to the position where the negative sample point is located can be 0. That is, the sample position label 1 is used to represent the positive sample point, and the sample position label 0 is used to represent the negative sample point.

[0077] The sample category is used to represent the category corresponding to the sample point. In a possible implementation manner, in the embodiments of the present application, the positive sample points can be divided into multiple categories, that is, the sample category includes multiple categories of positive sample points and negative sample points. For example, the categories of positive sample points can include curbs, vehicles, pedestrians, no-parking signs, and meaningless areas, etc. Among them, the no-parking signs can include no-parking signs, cones, etc. The meaningless area can represent the boundary of the drivable area that cannot be accurately determined. The embodiments of the present application do not distinguish specific categories for negative sample points, and other sample points except the above-mentioned several categories of positive sample points all belong to negative sample points. So as to train the initial neural network based on the sample category of the sample point and the predicted category of the prediction point.

[0078] It should be noted that the above method for determining the sample category of the sample point is only an exemplary description and is not limited to the above form.

[0079] In a possible implementation, when labeling positive sample points, the drivable area can be first divided in the sample image, and then the boundary of the drivable area can be sampled at equal intervals according to a preset number of pixel columns to obtain positive sample points. For example, every 4 pixel columns can be used as an interval to obtain a vertical line, and the intersection of the vertical line and the boundary of the drivable area is used as a positive sample point. For a specific schematic diagram, please refer to Figure 2 as shown in Figure 2 FIG. Figure 2 is a schematic diagram of determining positive sample points provided by an embodiment of the present application. Figure 2 The interval between two adjacent lines in FIG. Figure 2 is 4 pixel columns, and the intersection of the vertical line and the boundary of the drivable area is used as a positive sample point, so as to obtain 8 positive sample points.

[0080] S102: Input the sample image into the initial neural network to obtain a plurality of prediction points and the prediction results corresponding to the prediction points.

[0081] In the embodiment of the present application, the initial neural network can be used as the basic structure of the detection model. The initial neural network can be used to extract the features of the image and realize the detection and classification of the targets in the image, etc. By inputting the sample image into the initial neural network, the sample image can be continuously downsampled to complete feature extraction, and regression prediction can be performed based on the finally obtained image features to obtain a plurality of prediction points and the prediction results corresponding to each prediction point. The prediction results of the prediction points include: prediction position labels, prediction downsampling error values, and prediction categories.

[0082] Among them, the prediction position label is used to indicate whether the prediction point at this position is a positive sample point or a negative sample point, and the prediction position label can be determined according to the confidence level. That is, the initial neural network can output the confidence level corresponding to the prediction point. The confidence level can be a value in the range of 0 to 1, which is used to indicate the accuracy of the prediction position label of the prediction point being a positive sample point, so as to determine whether the prediction position label of the prediction point represents a positive sample point or a negative sample point according to this value. For example, when the confidence level is greater than a certain threshold, it indicates that the prediction position label of the prediction point is a positive sample point; otherwise, it indicates that the prediction position label of the prediction point is a positive sample point. That is, a preset threshold can be set in advance, and by comparing the confidence level corresponding to the prediction point with the preset threshold, it is determined whether the prediction position label of the prediction point is a positive sample point or a negative sample point.

[0083] Among them, the prediction category can indicate which category the prediction point corresponds to among the positive sample points or whether it is a negative sample point. Since the sample category includes positive sample points and negative sample points, and the positive sample points can include multiple categories, that is, the sample category includes negative sample points and multiple categories of positive sample points, when the initial neural network predicts the category of each prediction point, it can output the probability values ​​corresponding to the multiple categories of positive sample points and negative sample points, each probability value can be a value in the range of [0,1], and then determine that the category corresponding to the largest value among the multiple probability values ​​is the prediction category of the prediction point.

[0084] Among them, the predicted downsampling error value is determined based on the position of the predicted point and the position of the sample point corresponding to the predicted point in the sample image. After the sample image is input into the initial neural network, the sample image can be feature extracted by operations such as convolution, that is, the sample image is continuously downsampled to obtain the deepest image features, and regression prediction is performed through the image features. That is, the deepest image features are obtained by downsampling the sample image according to a certain multiple, wherein the downsampling multiple can be obtained by predetermining the structure of the initial neural network. The pixel position of the predicted point obtained based on the deepest image feature is an integer, but the pixel position obtained by directly downsampling the pixel position of the sample point corresponding to the predicted point is not necessarily an integer, so there is an error between the pixel position of the predicted point and the pixel position obtained by downsampling the pixel position of the sample point. Specifically, the pixel position of the predicted point can be multiplied by the downsampling multiple to obtain the upsampled pixel position, and then the difference between the upsampled pixel position and the pixel position of the sample point corresponding to the predicted point is calculated to obtain the predicted downsampling error value.

[0085] In a possible implementation, when the sample image is large, the processing of the sample image by the initial neural network will occupy more computing resources. Therefore, after obtaining the sample image, the sample image can be preprocessed to adjust the size of the sample image to meet the requirements of the initial neural network input. For example, the sample image can be preprocessed by random cropping, proportional reduction, etc. to adjust it to the required size. In the embodiment of the present application, the image size input to the initial neural network can be set to 320*320.

[0086] In one possible implementation, the initial neural network can be composed of convolution blocks and residual blocks to extract image features of different sizes to obtain richer feature information in the sample image. Figure 3a and Figure 3b As shown, Figure 3a A schematic diagram of the structure of an initial neural network provided in an embodiment of the present application, Figure 3b A schematic diagram of the structure of a residual block provided in an embodiment of the present application.

[0087] The initial neural network can be composed of residual blocks, convolutional blocks, and transposed convolutional blocks. The main functions of the residual blocks and convolutional blocks are to extract image features of different sizes, so as to increase the acquisition of feature information in the sample image. The transposed convolutional blocks can increase the size of the image features, so that the image features of the same size can be fused. The residual blocks can be composed of convolutional blocks and basic residual blocks, where the basic residual blocks include convolutional blocks and activation functions. Based on the structure of the initial neural network provided by the embodiments of the present application, the downsampling multiple corresponding to the deepest image features is 16 times.

[0088] S103: Determine the position label loss function based on the predicted position label and the class weight corresponding to the predicted class; determine the sampling error loss function based on the predicted downsampling error value; determine the class loss function based on the sample class and the predicted class.

[0089] After obtaining the prediction results of the prediction points, the loss function can be determined based on the prediction results and the labels of the sample points, so as to train the initial neural network. In specific implementation, since the positive sample points can include multiple categories, such as curbs, vehicles, pedestrians, no-parking signs, and meaningless regions, etc., in order to enable the detection model to accurately identify the boundary points of different categories, class weights can be assigned to the positive sample points of different categories in advance, so that the sum of the class weights corresponding to the positive sample points of each category is 1.

[0090] In a possible implementation manner, it can be determined that the class weights corresponding to multiple different categories of positive sample points are equal values, so that the model can accurately identify each category. Or, in combination with the actual application scenario, the class weights corresponding to the positive sample points of the categories with lower occurrence frequencies can be set to larger values, increasing the contribution of the positive sample points of the minority categories in the loss function, so that the model can pay more attention to these minority categories during the training process to improve the accuracy of model training.

[0091] Based on this, when the initial neural network outputs the predicted position labels, predicted downsampling error values, and predicted classes of multiple prediction points, the position label loss function can be determined based on the predicted position labels and the class weights corresponding to the predicted classes. Among them, the class weight corresponding to the predicted class represents the class weight corresponding to the corresponding sample class. The sampling error loss function can be determined based on the predicted downsampling error value. The class loss function is determined based on the sample class and the predicted class.

[0092] In a possible implementation, the position label loss function can be determined in the following manner: Determine a plurality of first prediction points corresponding to the positions of the positive sample points among the plurality of prediction points. Based on the predicted position labels of the first prediction points, the class weights corresponding to the predicted classes of the first prediction points, the balance factor, and the hyperparameters of the modulation factor, determine the first positive loss function. Based on the plurality of first positive loss functions corresponding to the plurality of first prediction points, determine the second positive loss function. Based on the predicted position labels of the second prediction points, the balance factor, and the hyperparameters of the modulation factor, determine the first negative loss function, where the second prediction points represent any other prediction points among the prediction points except the plurality of first prediction points. Based on the plurality of first negative loss functions corresponding to the plurality of second prediction points, determine the second negative loss function. Based on the average value of the second positive loss function and the second negative loss function, determine the position label loss function.

[0093] Among the multiple prediction points output by the initial neural network, determine multiple first prediction points corresponding to the positions of the positive sample points. Since multiple positive sample points are annotated in the sample image, for the pixel positions of each positive sample point in the sample image, the first prediction points corresponding to each positive sample point can be found based on the pixel positions of the multiple prediction points. That is, the pixel positions of the first prediction points correspond to the pixel positions of the positive sample points.

[0094] Then, based on the confidence corresponding to the predicted position label of the first prediction point, the class weight corresponding to the predicted class of the first prediction point, the balance factor, and the hyperparameters of the modulation factor, the first positive loss function can be determined. Among them, the hyperparameters of the balance factor and the modulation factor can be preset. The balance factor can be used to balance the ratio of positive and negative sample points, and the hyperparameters of the modulation factor are mainly used to adjust the contribution degree of the samples to the loss. The larger the hyperparameters of the modulation factor, the greater the degree of adjusting the loss.

[0095] Since the prediction points include multiple first prediction points, the final second positive loss function can be determined based on the multiple first positive loss functions corresponding to the multiple first prediction points. Optionally, since the first positive loss function is determined based on the class weights corresponding to the first prediction points, the sum of the multiple first positive loss functions can be calculated to obtain the second positive loss function.

[0096] Among multiple prediction points output by the initial neural network, except for multiple first prediction points corresponding to positive sample points, the remaining prediction points can be determined as second prediction points, that is, the prediction points corresponding to negative sample points are second prediction points. The first negative loss function can be determined based on the confidence, balance factor, and hyperparameter of the modulation factor corresponding to the prediction position label of the second prediction point. Since there are multiple second prediction points among the prediction points, the final second negative loss function can be determined based on multiple first negative loss functions corresponding to the multiple second prediction points. Optionally, the average value of multiple first negative loss functions can be calculated as the second negative loss function.

[0097] After obtaining the second positive loss function and the second negative loss function, the average value of the second positive loss function and the second negative loss function can be calculated to determine the position label loss function.

[0098] By classifying positive sample points of the drivable area, presetting different class weights for different classes, and combining the loss function of positive sample points and the loss function of negative sample points to determine the position label loss function, the model can more accurately detect boundary points and improve the accuracy of detecting the drivable area.

[0099] The process of determining the position label loss function will be introduced below in combination with a specific embodiment.

[0100] In specific implementation, the balance factor can be preset as α, the hyperparameter of the modulation factor can be expressed as γ, and the class weight of the first prediction point (corresponding to the positive sample point) can be expressed as ω i , where i = 1, 2,..., n, and n represents the number of multiple first prediction points included in the prediction points, that is, the number of positive sample points. The calculation formula of the first positive loss function can be expressed in the following form:

[0101] p_loss1(i) = ―α × ω i × log(prediction) × (1 ― prediction) γ ;

[0102] Among them, p_loss1(i) represents the first positive loss function, prediction represents the confidence corresponding to the prediction position label, and ω i represents the class weight corresponding to the first prediction point. Then, the second positive loss function can be expressed as

[0103] The calculation formula of the first negative loss function can be expressed in the following form:

[0104] n_loss1(j) = -(1 - α) × log(1 - prediction) × prediction γ ;

[0105] Where j = 1, 2, …, m, and m represents the number of second prediction points (corresponding to negative sample points). Then, the second negative loss function can be expressed as: Location label loss

[0106] In a possible implementation, when determining the sampling error loss function based on the predicted downsampling error value, Smooth L1 Loss can be used to determine the sampling error loss function corresponding to the predicted downsampling error value. Since there are multiple sample points in the sample image, a predicted downsampling error value can be determined for each pair of sample points and prediction points, thus obtaining multiple predicted downsampling error values. When determining the sampling error loss function based on the predicted downsampling error value, the first sampling error loss function corresponding to each predicted downsampling error value can be determined first, and then the average value of the multiple first sampling error loss functions corresponding to the multiple predicted downsampling error values can be calculated as the sampling error loss function.

[0107] Specifically, when the predicted downsampling error value is less than 1, the first sampling error loss function can be determined based on the squared error of the predicted downsampling error value. For example, 1 / 2 of the sum of the squares of the predicted downsampling error values can be determined as the first sampling error loss function. When the predicted downsampling error value is greater than or equal to 1, the first sampling error loss function can be determined based on the linear error of the predicted downsampling error value. For example, the difference between the predicted downsampling error value and 1 / 2 can be determined as the first sampling error loss function.

[0108] In a possible implementation, when determining the class loss function based on the sample class and the predicted class, binary cross-entropy loss can be used to determine the class loss function. According to the above embodiments, it can be known that the probability value corresponding to the sample class of the sample points can be determined in advance, and the initial neural network can output the probability values corresponding to multiple predicted classes. Therefore, based on the probability value corresponding to each sample class and the probability value of the corresponding output predicted class, the binary cross-entropy loss can be calculated as the class loss function.

[0109] S104: Train the initial neural network based on the location label loss function, the sampling error loss function, and the class loss function until the training termination condition is met, and obtain the trained detection model.

[0110] After determining the position label loss function, the sampling error loss function, and the category loss function in the above manner, the parameters of the initial neural network can be adjusted by methods such as gradient descent to train the initial neural network. Continuously adjust the parameters of the neural network until the training termination condition is met. For example, the training termination condition may include that the position label loss function, the sampling error loss function, and the category loss function are less than a preset value, or the number of training iterations reaches a preset number, etc., so as to obtain a trained detection model.

[0111] Through the method provided by the embodiments of the present application, a neural network can be trained using pre-annotated sample images to obtain a detection model, and the category weights of boundary points of different categories can be preset, so that the model pays more attention to learning to identify boundary points of minority categories. By combining positive and negative samples to determine the loss function, the accuracy of training the neural network can be improved to obtain a more accurate detection model.

[0112] Based on the above method embodiments, the embodiments of the present application also provide a method for detecting a drivable area. Refer to Figure 4 As shown, it is a flowchart of a method for detecting a drivable area provided by the embodiments of the present application.

[0113] Optionally, this method can be executed by a detection device. Among them, the detection device and the image processing device can be the same device or different devices. When the detection device and the image processing device belong to different devices, the detection device can obtain the trained detection model from the image processing device and store it in the detection device, and then the detection model can be directly called from the storage area.

[0114] Optionally, this method can be applied to an automatic parking scenario. For example, the vehicle's Home-zone Parking Assist (HPA) can help the vehicle complete parking after detecting the drivable area of the vehicle.

[0115] This method includes the following steps:

[0116] S401: Obtain a to-be-detected image captured by a camera on the vehicle.

[0117] In order to detect the drivable area around the vehicle, the road environment around the vehicle can be photographed by a camera on the vehicle to obtain a to-be-detected image. By identifying the to-be-detected image, the drivable area of the vehicle can be obtained.

[0118] Optionally, in the embodiments of the present application, in order to more accurately detect the drivable area of the vehicle, four-way fish-eye cameras can be installed in four directions of the vehicle to more comprehensively photograph the road environment around the vehicle.

[0119] S402: Input the image to be detected into the detection model to obtain multiple detection points and the detection results corresponding to each detection point.

[0120] After obtaining the image to be detected, the image to be detected can be input into the detection model, and multiple detection points and the detection results corresponding to each detection point can be obtained through the detection model. Among them, the detection result of each detection point can include a position label and a category. Among them, the position label is used to indicate whether the detection point is a boundary point of the drivable area, and the detection model can output the confidence corresponding to the detection point, and determine whether the detection point is a boundary point according to the confidence. The category can indicate which category of boundary point the detection point belongs to or is a non-boundary point. For example, the boundary points can be categories such as curbs, vehicles, pedestrians, no-parking signs, or meaningless areas.

[0121] S403: Determine the boundary points of the drivable area based on the detection results, so as to obtain the drivable area of the vehicle based on the boundary points.

[0122] After the detection model outputs the detection results of the detection points, the detection results can be used to determine multiple boundary points of the drivable area, and the boundaries of the drivable area are formed based on the multiple boundary points.

[0123] Since the position label is included in the detection result of the detection point and is determined according to the confidence, in one possible implementation, a preset threshold corresponding to the position label can be determined in advance, and the preset threshold is used to represent the lower limit value of the current detection point belonging to the boundary point. That is, when the confidence corresponding to the detection point is greater than or equal to the preset threshold, it can be determined that the position label of the detection point is a boundary point of the drivable area. When the confidence corresponding to the detection point is less than the preset threshold, it can be determined that the position label of the detection point is not a boundary point of the drivable area.

[0124] It should be noted that the embodiments of the present application do not limit the specific value of the preset threshold, and can be determined in combination with different application scenarios. According to the above embodiments, for example, the preset threshold can be set to 0.8. That is, when the confidence corresponding to the detection point is greater than or equal to 0.8, it can be determined that the position label of the detection point is a positive sample point, for example, the position label is 1, which is a boundary point of the drivable area. When the confidence is less than 0.8, it can be determined that the position label of the detection point is a negative sample point, for example, the position label is 0, which is not a boundary point of the drivable area.

[0125] After determining the boundary points of the drivable area, since the detection results are obtained by downsampling the image to be detected to extract features, the original pixel position corresponding to the boundary point in the image to be detected can also be restored based on the pixel position of the boundary point at this time and the downsampling ratio. For example, the original pixel position corresponding to the boundary point can be calculated by multiplying the pixel position of the boundary point by the downsampling ratio.

[0126] After determining the original pixel positions of each boundary point, the position of the drivable area in the image coordinate system is determined. In an actual application scenario, the drivable area around the vehicle is usually recognized by the vehicle's own intelligent driving system, and the vehicle is controlled to drive in the drivable area. Currently, the intelligent driving system of the vehicle usually realizes control based on the Bird's Eye View (BEV) coordinate system. Therefore, in order to facilitate the intelligent driving system of the vehicle to make control decisions, the inverse perspective mapping (IPM) transformation method can also be used to realize the conversion from the image coordinate system to the BEV coordinate system, determine the coordinates corresponding to the original pixel positions of the boundary points of the drivable area in the image to be detected in the BEV coordinate system, and determine the drivable area of the vehicle in the BEV coordinate system based on the coordinates of each boundary point, so as to more conveniently enable the intelligent driving system to make control decisions and control the vehicle to drive in the drivable area.

[0127] In a possible implementation manner, in order to reduce the jump problem caused by the unstable output of the detection model and reduce the false detection rate of the boundary points, during the process of tracking the boundary points of the drivable area, the Kalman filtering method can also be used to filter the coordinates of the boundary points to increase the stability of the predicted boundary points.

[0128] In an actual application scenario, four fisheye cameras can usually be installed in four directions of the vehicle to facilitate comprehensively photographing the surrounding environment of the vehicle and accurately determining the drivable area of the vehicle. There may be an overlapping area, that is, a fusion area, between the shooting ranges of two adjacent fisheye cameras. For details, see Figure 5 as shown Figure 5 is a schematic diagram of the shooting ranges of four fisheye cameras provided by an embodiment of the present application.

[0129] In this application scenario, fisheye cameras can be installed in the front, rear, left, and right of the vehicle respectively. Among them, the shooting range of the front fisheye camera consists of a front-left fusion area, a front area, and a front-right fusion area. The front-left fusion area represents the overlapping shooting range of the front fisheye camera and the left fisheye camera. The front area represents the range that can be photographed only by the front fisheye camera. The front-right fusion area represents the overlapping shooting range of the front fisheye camera and the right fisheye camera. The shooting range of the left fisheye camera consists of a front-left fusion area, a left area, and a rear-left fusion area. The left area represents the range that can be photographed only by the left fisheye camera. The rear-left fusion area represents the overlapping shooting range of the rear fisheye camera and the left fisheye camera. Similarly, the shooting ranges of the other fisheye cameras can be determined.

[0130] In a possible implementation, for the fusion region of two adjacent fisheye cameras, that is, the images to be detected in the fusion region can be captured by both adjacent fisheye cameras. Based on the images to be detected in the fusion region captured by the two fisheye cameras respectively, the detection results of the boundary points of the drivable region in the fusion region can be determined. Moreover, the detection results of the boundary points determined based on the two fisheye cameras may be different. Embodiments of the present application can determine the final detection result of the boundary points in the fusion region through the following methods:

[0131] For the detection result of the boundary points obtained by the first fisheye camera (any one of the two adjacent fisheye cameras), since the detection result output by the detection model includes the category, the category score can be determined based on the category corresponding to the boundary point. That is, for different categories that the boundary point may correspond to, different category scores corresponding to the categories can be preset in advance. Thus, after the category of the boundary point is output, the category score corresponding to this category can be determined.

[0132] Optionally, a higher category score can be set for the category with a higher occurrence frequency or a higher risk. For example, when the category of the boundary point is a curb, a vehicle, or a pedestrian, the risk is higher, and the category score is set to 2 points; when the category of the boundary point is a no-parking sign or a meaningless area, the risk is lower, and the category score is set to 1 point.

[0133] According to the above embodiments, the detection result output by the detection model includes the confidence corresponding to the position label. Based on this confidence, it is determined whether the detection point is a boundary point. Therefore, the confidence score can be determined based on the confidence corresponding to the position label. The higher the confidence, the higher the accuracy of predicting the boundary point. Therefore, for different confidences, different confidence scores can be preset in advance. When the confidence is higher, the corresponding confidence score is also higher. For example, when the confidence is greater than or equal to 0.8, the corresponding confidence score can be set to 3 points; when the confidence is greater than 0.6 and less than 0.8, the corresponding confidence score can be set to 2 points; when the confidence is less than or equal to 0.6, the corresponding confidence score can be set to 1 point.

[0134] When the same boundary point exists in multiple consecutive frames of images captured by the fisheye camera, it indicates that the boundary point has high stability and accuracy. Therefore, the boundary point can also be tracked to determine a boundary point with higher stability, improving the accuracy of determining the drivable region. Specifically, for the boundary points determined by any one of the two adjacent fisheye cameras (i.e., the first fisheye camera), the number of consecutive frames in which the boundary point appears in the images captured by the first fisheye camera is tracked, and the tracking score is determined according to this number of frames.

[0135] Similar to the category score and the confidence score, similarly, a mapping relationship between the number of consecutive frames in which a boundary point appears and the tracking score can be established in advance. It can be set that the more consecutive frames in which the boundary point appears, the higher the corresponding tracking score. For example, when the number of consecutive frames in which the boundary point appears is greater than or equal to 5 frames, the corresponding tracking score can be set to 3 points; when the number of consecutive frames in which the boundary point appears is 2 frames, 3 frames or 4 frames, the corresponding tracking score can be set to 2 points; when the boundary point appears in only 1 frame, the corresponding tracking score can be set to 1 point.

[0136] It should be noted that the category scores corresponding to different categories, the confidence scores corresponding to different confidence levels, and the tracking scores corresponding to different numbers of consecutive frames provided in the above embodiments are only an exemplary illustration and do not impose any formal limitation on this application. The specific values of the category score, the confidence score, the number of frames, and the tracking score can be set in combination with actual requirements, and none of them will affect the implementation of the embodiments of this application.

[0137] Then, based on the tracking score, the category score, and the confidence score, the total score can be determined. For example, the tracking score, the category score, and the confidence score can be added together to obtain the total score.

[0138] Based on the above process, for the boundary points of the fusion regions respectively determined for the two fisheye cameras, two total scores corresponding to the boundary points respectively determined by the two fisheye cameras can be obtained. Then, compare the magnitudes of the two total scores, and only retain the boundary points determined by the fisheye camera with the higher total score in the fusion region, and delete the boundary points determined by the fisheye camera with the lower total score, so as to determine the drivable region of the vehicle according to the detection results of the retained boundary points.

[0139] Based on the above method embodiments, the embodiments of this application further provide a training device for a drivable region detection model. Refer to Figure 6 As shown, it is a schematic diagram of a training device for a drivable region detection model provided by the embodiments of this application.

[0140] The device 600 may include:

[0141] A sample acquisition unit 601, configured to acquire a sample image, where multiple sample points are marked in the sample image, and the labels of the sample points include: a sample position label and a sample category, the sample position label is used to indicate whether the sample point is a positive sample point or a negative sample point, and multiple positive sample points form the boundary of the drivable region;

[0142] A prediction unit 602, configured to input the sample image into an initial neural network to obtain a plurality of prediction points and prediction results corresponding to the prediction points, where the prediction results include: a predicted position label, a predicted downsampling error value, and a predicted category, and the predicted downsampling error value is determined based on the position of the prediction point and the position of the sample point corresponding to the prediction point in the sample image;

[0143] A loss function determination unit 603, configured to determine a position label loss function based on the predicted position label and the class weight corresponding to the predicted category; determine a sampling error loss function based on the predicted downsampling error value; determine a class loss function based on the sample class and the predicted class;

[0144] A training unit 604, configured to train the initial neural network based on the position label loss function, the sampling error loss function, and the class loss function until a training termination condition is met, to obtain a trained detection model.

[0145] In a possible implementation manner, the loss function determination unit 603 is specifically configured to determine a plurality of first prediction points corresponding to the positions of the positive sample points among the prediction points; determine a first positive loss function based on the predicted position label of the first prediction points, the class weight corresponding to the predicted category of the first prediction points, hyperparameters of a balance factor, and a modulation factor; determine a second positive loss function based on a plurality of the first positive loss functions corresponding to the plurality of the first prediction points; determine a first negative loss function based on the predicted position label of the second prediction points, the balance factor, and the hyperparameters of the modulation factor, where the second prediction points represent any other prediction points except the plurality of the first prediction points among the prediction points; determine a second negative loss function based on a plurality of the first negative loss functions corresponding to the plurality of the second prediction points; and determine the position label loss function based on an average value of the second positive loss function and the second negative loss function.

[0146] In a possible implementation manner, the loss function determination unit 603 is specifically configured to calculate a sum of the plurality of first positive loss functions to obtain the second positive loss function;

[0147] The loss function determination unit 603 is specifically configured to calculate an average value of the plurality of first negative loss functions to obtain the second negative loss function.

[0148] In a possible implementation manner, the prediction unit 602 is specifically configured to determine a product of the pixel position of the prediction point and the downsampling multiple of the initial neural network; and obtain the predicted downsampling error value based on an error between the position of the sample point corresponding to the prediction point in the sample image and the product.

[0149] In a possible implementation, the positive sample points are obtained by equally spacing and sampling the boundaries of the drivable areas in the sample image according to a preset number of pixel columns.

[0150] In addition, the present application also provides a detection device for a drivable area. Refer to Figure 7 As shown, it is a schematic diagram of a detection device for a drivable area provided by an embodiment of the present application.

[0151] The device 700 includes:

[0152] An image acquisition unit 701, configured to acquire a to-be-detected image captured by a camera on the vehicle;

[0153] A detection unit 702, configured to input the to-be-detected image into a detection model to obtain a plurality of detection points and detection results corresponding to the detection points;

[0154] A determination unit 703, configured to determine boundary points of the drivable area based on the detection results, so as to obtain the drivable area of the vehicle based on the boundary points.

[0155] In a possible implementation, the detection result includes a position label, and the determination unit 703 is specifically configured to determine the detection point corresponding to the position label as a boundary point of the drivable area when the confidence level corresponding to the position label is greater than or equal to a preset threshold.

[0156] In a possible implementation, the device further includes: a conversion unit, configured to calculate, based on the pixel position of the boundary point and the downsampling multiple of the detection model, the original pixel position corresponding to the boundary point in the to-be-detected image; use an inverse perspective transformation method to determine the coordinates corresponding to the original pixel position in a bird's-eye view coordinate system; and obtain the drivable area of the vehicle in the bird's-eye view coordinate system based on the coordinates.

[0157] In a possible implementation, the camera includes four fisheye cameras, and the detection result further includes a category. When the boundary point is located in the fusion area of two adjacent fisheye cameras, the determination unit 703 is further configured to, for a first fisheye camera, track the number of consecutive frames in which the boundary point appears in the image captured by the first fisheye camera, and determine a tracking score based on the number of frames, where the first fisheye camera represents any one of the two adjacent fisheye cameras; determine a category score based on the category of the boundary point; determine a confidence score based on the confidence level corresponding to the position label of the boundary point; determine a total score based on the tracking score, the category score, and the confidence score; and retain, in the fusion area, the boundary points determined by the fisheye camera with a higher total score among the two adjacent fisheye cameras.

[0158] Based on the above method embodiments and apparatus embodiments, the embodiments of the present application further provide an electronic device. The following will be introduced with reference to the accompanying drawings.

[0159] See Figure 8 , Figure 8 which is a schematic diagram of an electronic device provided by an embodiment of the present application.

[0160] The device 800 includes: a memory 801 and a processor 802;

[0161] The memory 801 is used to store relevant program codes;

[0162] The processor 802 is used to call the program codes to execute the training method of the drivable area detection model or the drivable area detection method described in the above method embodiments.

[0163] In addition, the embodiments of the present application further provide a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the training method of the drivable area detection model or the drivable area detection method described in the above method embodiments.

[0164] The embodiments of the present application further provide a computer program product, which includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the training method of the drivable area detection model or the drivable area detection method described in the above method embodiments is implemented.

[0165] It should be noted that the above computer-readable medium of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0166] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0167] It should be noted that the various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other. In particular, for system or apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, reference may be made to the partial description of the method embodiments. The apparatus embodiments described above are merely illustrative. The units or modules described as separate components may or may not be physically separated. The components shown as units or modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network units. Some or all of the units or modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods, apparatuses, and devices according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0169] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0170] It should also be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0171] The steps of the methods or algorithms described in connection with the embodiments disclosed in this application can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0172] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined in this application can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments shown in this application, but will conform to the widest scope consistent with the principles and novel features disclosed in this application.

Claims

1. A training method for a detection model of a drivable area, characterized in that, The method includes: Obtain a sample image, in which a plurality of sample points are labeled. The labels of the sample points include: a sample position label and a sample category. The sample position label is used to indicate whether the sample point is a positive sample point or a negative sample point. The plurality of positive sample points form the boundary of the drivable area; Input the sample image into an initial neural network to obtain a plurality of prediction points and the prediction results corresponding to the prediction points. The prediction results include: a predicted position label, a predicted downsampling error value, and a predicted category. The predicted downsampling error value is determined based on the position of the prediction point and the position of the sample point corresponding to the prediction point in the sample image; Based on the predicted position label and the category weight corresponding to the predicted category, determine a position label loss function; based on the predicted downsampling error value, determine a sampling error loss function; based on the sample category and the predicted category, determine a category loss function; Train the initial neural network based on the position label loss function, the sampling error loss function, and the category loss function until the training termination condition is met to obtain a trained detection model.

2. The method according to claim 1, wherein The determining the position label loss function based on the predicted position label and the category weight corresponding to the predicted category includes: Determine a plurality of first prediction points among the plurality of prediction points that correspond to the positions of the positive sample points; Based on the predicted position label of the first prediction point, the category weight corresponding to the predicted category of the first prediction point, and the hyperparameters of the balance factor and the modulation factor, determine a first positive loss function; Based on the plurality of first positive loss functions corresponding to the plurality of first prediction points, determine a second positive loss function; Based on the predicted position label of the second prediction point, the balance factor, and the hyperparameters of the modulation factor, determine a first negative loss function, where the second prediction point represents any other prediction point among the prediction points except the plurality of first prediction points; Based on the plurality of first negative loss functions corresponding to the plurality of second prediction points, determine a second negative loss function; Based on the average value of the second positive loss function and the second negative loss function, determine the position label loss function.

3. The method according to claim 2, wherein The determining the second positive loss function based on the plurality of first positive loss functions corresponding to the plurality of first prediction points includes: Calculate the sum of the plurality of first positive loss functions to obtain the second positive loss function; The determining the second negative loss function based on the plurality of first negative loss functions corresponding to the plurality of second prediction points includes: Calculate the average value of the plurality of first negative loss functions to obtain the second negative loss function.

4. The method according to claim 1, wherein Obtaining the predicted downsampling error value corresponding to the prediction point includes: Determine the product between the pixel position of the prediction point and the downsampling multiple of the initial neural network; Based on the error between the position of the sample point corresponding to the prediction point in the sample image and the product, obtain the predicted downsampling error value.

5. The method according to any one of claims 1 to 4, characterized in that The positive sample points are obtained by equally spacing sampling the boundaries of the drivable regions in the sample image according to a preset number of pixel columns.

6. A detection method for a drivable area, characterized in that, The method includes: Obtaining a to-be-detected image captured by a camera on the vehicle; Inputting the to-be-detected image into the detection model according to any one of claims 1 to 5, and obtaining a plurality of detection points and the detection results corresponding to the detection points; Determining boundary points of the drivable region based on the detection results, so as to obtain the drivable region of the vehicle based on the boundary points.

7. The method according to claim 6, wherein The detection results include position tags, and determining the boundary points of the drivable region based on the detection results includes: When the confidence level corresponding to the position tag is greater than or equal to a preset threshold, determining the detection point corresponding to the position tag as a boundary point of the drivable region.

8. The method according to claim 7, characterized in that, The method further includes: Calculating, based on the pixel positions of the boundary points and the downsampling multiple of the detection model, the original pixel positions corresponding to the boundary points in the to-be-detected image; Using an inverse perspective transformation method to determine the coordinates corresponding to the original pixel positions in the bird's-eye view coordinate system; Obtaining the drivable region of the vehicle in the bird's-eye view coordinate system based on the coordinates.

9. The method according to claim 7, wherein The camera includes four fisheye cameras, the detection results further include categories, and when the boundary points are located in the fusion region of two adjacent fisheye cameras, the method further includes: For the first fisheye camera, tracking the number of consecutive frames in which the boundary points appear in the image captured by the first fisheye camera, and determining a tracking score based on the number of frames, where the first fisheye camera represents any one of the two adjacent fisheye cameras; Determining a category score based on the category of the boundary points; Determining a confidence score based on the confidence level corresponding to the position tag of the boundary points; Determining a total score based on the tracking score, the category score, and the confidence score; Retaining the boundary points determined by the fisheye camera with a higher total score among the two adjacent fisheye cameras in the fusion region.

10. An electronic device, characterized in that, The device includes: a memory and a processor; The memory is used for storing relevant program codes; The processor is used for calling the program codes to execute the training method of the detection model for the drivable region according to any one of claims 1 to 5 or the detection method of the drivable region according to any one of claims 6 to 9.