A Visual Multi-Task Processing Method Based on Deep Learning
Through the deep learning-based visual multitasking neural network model and the design of shared feature encoder and decoder, the speed and resource requirements of multitasking in autonomous driving systems are solved, and efficient multitasking visual processing effect is achieved.
Patent Information
- Application Number
- CN202211515937.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-11-30
AI Technical Summary
In the prior art, target detection, feasible area detection and lane line detection tasks need to be handled separately, resulting in the autonomous driving system being unable to meet the requirements in terms of speed and resource requirements, especially when deployed on embedded devices, the power consumption and computing power are strictly limited.
The visual multitasking neural network model based on deep learning is adopted, and the combined processing of object detection, feasible area detection and lane line detection is realized through shared feature encoder and multiple decoders. The bounding box positioning loss function is used to optimize bounding box positioning, and the data set preprocessing and enhancement technology are combined to improve the generalization ability of the model.
It realizes the simultaneous multitasking processing in a single neural network, reduces the amount of model parameters and inference time, improves processing speed and accuracy, adapts to harsh driving environments, and meets the multitasking visual processing requirements of the autonomous driving system.
Smart Images

Figure CN115909245B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-task processing in autonomous driving, and particularly to a visual multi-task processing method based on deep learning. Background Art
[0002] The driving environment perception system for autonomous driving is very important because it can obtain visual information from devices such as RGB cameras, depth cameras, and infrared cameras, and can be used as reference input information for vehicle autonomous driving decisions. In order for the vehicle to have the ability of intelligent driving, the visual perception system is required to be able to obtain external information and understand the scene, and the information provided to the decision-making system includes: detection of pedestrians, vehicles, and obstacles, determination of drivable areas, lane lines, etc. The driving perception system includes target detection to help the vehicle identify pedestrians, vehicles, and obstacles, drive safely, and comply with traffic rules. Drivable area segmentation and lane detection also need to be carried out because they are the keys to planning the driving route of the vehicle.
[0003] In the field of autonomous driving, significant progress has been made in object detection algorithms for detecting vehicles, lanes, pedestrians, etc. Moreover, lane detection and lane lines in autonomous driving systems have also developed rapidly. Through these technologies mentioned above, the positions between vehicles can be located, and the feasible areas around the vehicles can be determined to ensure the safety of the moving vehicles. In the past, these tasks were usually processed separately, that is, multiple tasks such as object detection, drivable area detection, and lane line detection were not related to each other. The most classic method to complete these three tasks simultaneously is to use the Faster R-CNN and YOLOv4 algorithms to handle the object detection task, use the ENet and PSPNet algorithms to handle the drivable area segmentation detection task, and use the SCNN and SADENet algorithms to handle the lane line detection task. Although the above methods perform very well in the processing of single tasks, including speed and accuracy. However, in the actual application of autonomous driving systems, it is not only necessary to process single vision tasks. If it is necessary to complete the three tasks of object detection, drivable area detection, and lane line detection, then three tasks need to be executed, and three times are required to process the three tasks. Obviously, this is unacceptable for autonomous driving tasks with extremely high speed requirements. When deploying a driving perception system on an embedded device in autonomous driving, limited power consumption, computing power, and latency need to be considered. In autonomous driving environment perception, there is often a lot of interrelated information among different tasks, and this part of the information can be shared. Visual multi-task neural networks are suitable for this situation because they can implement multiple task processing with one network body without the need to separately process multiple neural network tasks in series or in parallel. And because visual multi-task neural networks can share the same feature extraction backbone, information can be shared among multiple tasks, and multiple models are merged into one visual multi-task neural network model, which greatly reduces redundant parameters, greatly reduces the computing power requirements, and can achieve excellent results in multiple tasks. Summary of the Invention
[0004] The object of the present invention is to overcome the deficiencies of the prior art and propose a visual multi-task processing method based on deep learning, which can simultaneously complete the object detection task, the drivable area detection task, and the lane line detection task.
[0005] To achieve the above object, the technical solution provided by the present invention is: a vision multi-task processing method based on deep learning, which uses a vision multi-task processing neural network model based on deep learning to simultaneously complete multi-task vision processing required in vehicle autonomous driving, including object detection tasks, drivable area detection tasks, and lane line detection tasks. Among them, the vision multi-task processing neural network model consists of an input layer, a shared feature encoder, a bottleneck module, and three decoders for different tasks. The decoders share feature maps with similar tasks to achieve joint semantic understanding, and the decoder used for the object detection task uses CIoU to measure the loss value;
[0006] The specific implementation of this vision multi-task processing method includes the following steps:
[0007] S1. Obtain the data set and perform preprocessing, including: performing a scaling operation on the data set to meet the input requirements of the vision multi-task processing neural network model, performing an enhancement operation on the data set, and performing a style conversion on the data set to better simulate the actual harsh driving weather environment; dividing the preprocessed data set into a training set and a test set;
[0008] S2. Adjust the training parameters, construct a data set generator, and train the vision multi-task neural network model step by step: first train the shared encoder of the vision multi-task neural network model, and then train the three decoders of the vision multi-task neural network model for different tasks respectively;
[0009] S3. Collect the RGB image data in the test set, input it into the trained vision multi-task neural network model for prediction, obtain the object detection prediction result, the drivable area prediction result, and the lane line prediction result, and draw all the prediction results on the test RGB image for display output.
[0010] Further, in step S1, the BDD100K dataset is used. The original RGB image format of the BDD100K dataset is in jpg format with a resolution of 1280×720. The label format for object detection is in json format. During the training process, the json format labels need to be converted into {x, y, w, h, class}, where (x, y) represents the coordinates of the bounding box, (w, h) represents the width and height of the bounding box, and class represents the category of the object. The drivable area labels and lane line labels in the BDD100K dataset are in png format with a resolution of 1280×720. The sizes of the drivable area labels and lane line labels need to be converted into the sizes of the outputs of the two decoders corresponding to the drivable area detection task and the lane line detection task of the visual multi-task neural network model. A color transformation enhancement operation is performed on the BDD100K dataset. Using the histogram equalization algorithm, calculate the grayscale histogram of the image, find the total number of image pixels, normalize the histogram distribution, calculate the cumulative score of the grayscale levels of the image, and find the grayscale value of the enhanced image to obtain an image after histogram equalization. A scene transformation operation is performed on the dataset. The CycleGAN algorithm is used for scene transformation. CycleGAN is a style conversion neural network. Using CycleGAN, the BDD100K dataset is subjected to weather transformation, including converting sunny days to thunderstorm days, sunny days to snowy days, sunny days to hazy days, and sunny days to rainy days, expanding the number of the autonomous driving dataset, and enabling the visual multi-task processing neural network model to learn more data images in harsh environments, making the model more generalizable.
[0011] Further, in step S2, the specific situation of the visual multi-task processing neural network model is as follows:
[0012] a. Construct the input layer of the visual multi-task processing neural network model, with the requirements: input an RGB image, and obtain an RGB image with a size of N×M through scaling or cropping. N is the horizontal resolution size of the RGB image after scaling or cropping, and M is the vertical resolution size of the RGB image after scaling or cropping. Then convert it into a tensor with a dimension of N×M×3.
[0013] b. Construct the shared feature encoder of the visual multi-task processing neural network model. The modules used in the shared feature encoder include the CBM module, the CSPx module, the CBLx module, and the SPP module.
[0014] The CBM module consists of a Conv operation, a BatchNorm operation, and a Mish activation function in sequence.
[0015] The CSPx module has two branches, namely the backbone and the shortcut. The CSPx module starts with a CBM module, and this starting CBM module is connected to the backbone and the shortcut of the CSPx module respectively. The backbone consists of a CBM module, x ResUnit modules, and a CBM module in sequence. The shortcut consists of a CBM module. The backbone and the shortcut of the CSPx module perform a Concat operation, and there is another CBM module at the end; The ResUnit module consists of a backbone and a shortcut. The backbone of the ResUnit consists of two CBM modules. The shortcut of the ResUnit directly adds the input feature to the output feature of the backbone;
[0016] The CBLx module consists of x CBL modules, and each CBL module consists of a Conv operation, a BatchNorm operation, and a LeakyReLU activation function in sequence;
[0017] The SPP module consists of 4 branches, namely Maxpool operations with sizes of 3×3, 5×5, and 9×9, and an empty operation shortcut. Then, these 4 branches perform a Concat operation;
[0018] The shared feature encoder is composed of a CBM module, a CSP1 module, a CSP2 module, a first CSP8 module, a second CSP8 module, a CSP4 module, a CBL3 module, and an SPP module in sequence;
[0019] c. Build the bottleneck module of the visual multi-task processing neural network model, which consists of three CBL6 modules, namely the first CBL6 module, the second CBL6 module, and the third CBL6 module. Requirements: Input the features extracted by the SPP module into the first CBL6 module. After performing an UpSample operation on the output features of the first CBL6 module, then perform a Concat operation with the output features of the second CSP8 module to obtain the feature output to the second CBL6 module. After performing an UpSample operation on the output features of the second CBL6 module, then perform a Concat operation with the output features of the first CSP8 module to obtain the feature output to the third CBL6 module;
[0020] d. Define three decoders for different tasks in the visual multi-task processing neural network model, namely: the object detection decoder, the drivable area detection decoder, and the lane line detection decoder;
[0021] The constructed object detection decoder has object detection heads at three scales. Each object detection head consists of a CBL1 module. The object detection heads at the three scales are named Y1, Y2, and Y3 respectively; Y3 obtains the features output by the first CBL6 module, Y2 obtains the features output by the second CBL6 module, and Y1 obtains the features output by the third CBL6 module;
[0022] The constructed drivable area detection decoder consists of two CBL3 modules, namely the first CBL3 module and the second CBL3 module. Requirements: The features extracted by the third CBL6 module are subjected to an UpSample operation and the output features are sent to the first CBL3 module. The output features of the first CBL3 module are subjected to an UpSample operation and the output features are sent to the second CBL3 module;
[0023] The constructed lane line detection decoder consists of two CBL3 modules, namely the third CBL3 module and the fourth CBL3 module. Requirements: The features extracted by the third CBL6 module are subjected to an UpSample operation and the output features are sent to the third CBL3 module. The output features of the third CBL3 module are subjected to an UpSample operation and the output features are sent to the fourth CBL3 module.
[0024] Furthermore, the visual multi-task processing neural network model designs loss functions for the object detection task, the drivable area detection task, and the lane line detection task, namely the object detection loss, the drivable area detection loss, and the lane line detection loss, as shown in formula (1);
[0025] L all = αL det + βL da + γL ll (1)
[0026] In the formula, L all is the total loss value, α, β, and γ are the weight parameters of the object detection loss, the drivable area detection loss, and the lane line detection loss respectively, L det is the object detection loss value, L da is the drivable area detection loss value, L ll is the lane line detection loss value;
[0027] The object detection loss function consists of a localization loss, an object confidence loss, and a class loss, as shown in formula (2);
[0028] L det = λ1L ciou + λ2L obj + λ3L cla (2)
[0029]
[0030]
[0031]
[0032] Wherein, λ1 represents the weight parameter of the positioning loss, λ2 represents the weight parameter of the target confidence loss, and λ3 represents the weight parameter of the classification loss, and L ciou is the positioning loss value, and L obj is the target confidence loss value, and L cla is the classification loss value; C i represents the predicted confidence of the i-th grid bounding box, represents the target confidence of the i-th grid bounding box, represents that there is no target in the j-th bounding box of the i-th grid; c represents the category, and c belongs to the total category cl asses , p i (c) represents the probability that the prediction of the i-th grid is the category c, represents the probability that the target value of the i-th grid is the category c; S represents dividing the image into S×S grids, and S 2 represents the total number of grids S×S after image segmentation, and the range of i is from 0 to S 2 , B represents the number of bounding boxes in the i-th grid, and the range of j is from 0 to B, represents that there is a target in the j-th bounding box of the i-th grid. The CIoU regression positioning loss is used, as shown in formulas (6), (7), and (8). CIoU is an improved algorithm of IoU. IoU represents the intersection over union of two bounding boxes. CIoU takes into account the overlapping area, center distance, and aspect ratio of the bounding boxes. The CIoU regression positioning loss used makes the neural network model more capable of reflecting the accuracy of bounding box positioning during training;
[0033]
[0034]
[0035]
[0036] Among them, formula (6) is the calculation method of CIoU. IoU is the intersection over union, that is, the intersection area of two bounding boxes divided by the union area, and ρ 2 (b, b gt ) represents the Euclidean distance between the predicted bounding box b and the true bounding box b gt . diag represents the length of the diagonal of the minimum circumscribed rectangle of the two bounding boxes. v is a parameter used to measure the consistency of the aspect ratio, as shown in formula (7). a is a positive trade-off parameter, as shown in formula (8). w, h, and w gt , h gt represent the height and width of the predicted bounding box and the height and width of the true bounding box, respectively;
[0037] Define L da as the drivable area detection loss value, as shown in formula (9). Input the drivable area detection feature
[0038] In the figure, p is the pixel value in the feature map, and Ω is the set of feature maps. is the true pixel value in the target feature map, and the cross-entropy method is used to calculate the detection loss value of the drivable area.
[0039]
[0040] Define L ll as the lane detection loss value, as shown in Equation (10), where ω1, ω2, and ω3 are weight parameters, and L ll_seg is the lane detection segmentation loss function, which is calculated using cross-entropy. The input is the lane detection feature map and is calculated in the same way as L da L ll_iou is the intersection over union (IoU) of the predicted lane and the true lane, and L ll_exist is the loss for the existence of lanes.
[0041] L ll = ω1L ll_seg + ω2L ll_iou + ω3L ll_exist (10).
[0042] Furthermore, in step S2, the adam optimizer is used, the batch_size value is set, the total number of training epochs is set, and the training loss value and validation loss value obtained in each epoch are printed. If the training loss value decreases and the validation loss value also decreases, it proves that the model has not been trained yet, and training continues to improve the model's performance. When the training loss value decreases but the validation loss value increases, it indicates that the model has started to overfit and training needs to be stopped; different learning rates learningrate are set according to different epochs, and as the epoch value increases, the learning rate is dynamically adjusted to decrease.
[0043] A dataset generator is created, and the BDD100K dataset is imported in batches into the visual multi-task processing neural network model for training; the training set and validation set divided from the BDD100K dataset are imported; the training set is used to train the visual multi-task processing neural network model. When training, the content imported by the dataset generator includes the input RGB image, target detection bounding box, drivable area detection segmentation image, and lane detection segmentation image, and the number of imports per batch is controlled by the batch size; the validation set is used to verify the effect of training the visual multi-task processing neural network model, and the mean average precision (mAP), intersection over union (IoU), and recall rate predicted by the target detection decoder are output, and the precision and IoU predicted by the drivable area detection decoder and lane detection decoder are output.
[0044] Initialize the visual multi-task processing neural network model, the dataset generator, and the training parameters, start training the visual multi-task processing neural network model, train the model to the set epoch value, freeze the shared feature encoder of the visual multi-task processing neural network model, and then separately train the object detection decoder, the drivable area detection decoder, and the lane line detection decoder of the visual multi-task processing neural network model in sequence;
[0045] Further, the specific process of step S3 is as follows:
[0046] S31. Use the test set images in the BDD100k dataset, use a GPIO camera and a USB camera to capture an RGB image as the input image, scale the size of the captured RGB image to the input size N×M of the visual multi-task processing neural network model, and convert it into a tensor of N×M×3;
[0047] S32. Input the obtained tensor of N×M×3 into the visual multi-task processing neural network model for prediction. The model will output the results of three parts, namely the object detection result, the drivable area detection result, and the lane line detection result; post-process these three results respectively. The object detection obtains multiple results in the format of {x, y, w, h, conf, class}; (x, y) are the coordinates of the bounding box; (w, h) are the width and height of the bounding box; conf is the confidence of the bounding box; class is the category of the object. Set the IoU value and the confidence value, use the NMS non-maximum suppression algorithm to select the appropriate bounding box, and draw the bounding box on the image; the drivable area detection and the lane line detection are in the data format of instance segmentation. It is necessary to perform a smoothing operation and a binarization operation on the output matrix to obtain the result, and draw it on the image to obtain an RGB image with the object detection result, the drivable area detection result, and the lane line detection result drawn on it.
[0048] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0049] 1. The visual multi-task processing neural network model based on deep learning proposed by the present invention only needs one encoder and three decoders for different tasks respectively to simultaneously complete the three tasks of object detection, drivable area detection, and lane line detection, greatly reducing the number of model parameters and the inference time. The decoders share feature maps with similar tasks to achieve joint semantic understanding, which is more in line with the multi-task visual processing required in autonomous driving technology.
[0050] 2. The visual multi-task neural network model uses a relatively lightweight neural network main structure, which is very fast both in training and inference, and the number of parameters is sufficient to bear the feature information required by the three decoders, meeting both the speed requirements and the accuracy requirements.
[0051] 3. The present invention enhances the BDD100K dataset in multiple modes, which can better improve the generalization performance of the model, better conform to the harsh environments encountered in actual driving, and can adapt to different driving scenarios and various weather conditions, which also reflects the potential market and application value of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is an effect diagram of histogram equalization enhancement for the dataset.
[0053] Figure 2 It is a schematic diagram of the visual multi-task processing neural network model.
[0054] Figure 3 It is an architecture diagram of the visual multi-task processing neural network model.
[0055] Figure 4 It is a display diagram of the input predicted RGB image effect (a grayscale image is shown here).
[0056] Figure 5 It is an effect diagram of the input RGB image entering the visual multi-task processing neural network model; among them, the left figure shows the target detection result, and the right figure shows the drivable area detection result and the lane line detection result.
[0057] Figure 6 It is a display effect diagram of the target detection result, the drivable area detection result and the lane line detection result drawn on the RGB image (a grayscale image is shown here). DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0059] This embodiment is implemented under the Pytorch deep learning framework. The computer is configured with an Intel Core I7-12800H processor, 64GB of memory, an nVIDIA GeForce RTX3090 24GB graphics card, and a Windows11 22H2 operating system. This embodiment discloses a visual multi-task processing method based on deep learning. This method uses a visual multi-task processing neural network model based on deep learning to simultaneously complete multi-task visual processing required in vehicle autonomous driving, including object detection tasks, drivable area detection tasks, and lane line detection tasks. Among them, the visual multi-task processing neural network model consists of an input layer, a shared feature encoder, a bottleneck module, and three decoders for different tasks. The decoders share feature maps with similar tasks to achieve joint semantic understanding, and the decoder used for the object detection task uses CIoU to measure the loss value. The specific process is as follows:
[0060] 1) Use the BDD100K dataset to train the visual multi-task neural network model. The original RGB image format of the BDD100K dataset is in jpg format, with a resolution of 1280×720. The label format for object detection is in json format. During training, the json format label needs to be converted into {x, y, w, h, class}, that is, the coordinates (x, y) of the bounding box; the width and height (w, h) of the bounding box; the class of the object class. The drivable area labels and lane line labels in the BDD100K dataset are in png format, with a resolution of 1280×720. The sizes of the drivable area labels and lane line labels need to be converted into the sizes output by the two decoders corresponding to the drivable area detection task and the lane line detection task of the visual multi-task neural network model;
[0061] Perform color transformation enhancement operations on the dataset. Use the histogram equalization algorithm to calculate the grayscale histogram of the image, find the total number of image pixels, normalize the histogram distribution, calculate the cumulative score of the grayscale levels of the image, and find the grayscale value of the enhanced image to obtain an image after histogram equalization, as shown in Figure 1 shown. Perform scene transformation operations on the dataset. Use the CycleGAN algorithm for scene conversion. CycleGAN is a style conversion neural network model. In this invention, CycleGAN is used to perform weather conversion on the dataset in BDD100K, including converting sunny days to thunderstorm days, sunny days to snowy days, sunny days to hazy days, and sunny days to rainy days in bad weather, expanding the number of autonomous driving datasets, and enabling the visual multi-task processing neural network model to learn more data images in harsh environments, making the model more generalizable.
[0062] 2) Construct a visual multi-task processing neural network model, as shown in Figure 2As shown in the figure, the vision multi-task processing neural network model is divided into four parts, namely the input layer, the shared feature encoder, the bottleneck module, and three decoders for different tasks. The three decoders for different tasks are defined as the object detection decoder, the drivable area detection decoder, and the lane line detection decoder;
[0063] Build the input layer of the vision multi-task processing neural network model with the following requirements: Input an RGB image, and obtain an RGB image with the size of N×M through scaling or cropping. N is the horizontal resolution size of the RGB image after scaling or cropping, and M is the vertical resolution size of the RGB image after scaling or cropping. Then convert it into a tensor with the dimension of N×M×3;
[0064] Build the shared feature encoder of the vision multi-task processing neural network model. The modules used in the shared feature encoder include the CBM module, the CSPx module, the CBLx module, and the SPP module, as shown in Figure 3 the figure;
[0065] The CBM module consists of a Conv operation, a BatchNorm operation, and a Mish activation function in sequence;
[0066] The CSPx module has two branches, the main branch and the shortcut. The CSPx module starts with a CBM module, and this starting CBM module is respectively connected to the main branch and the shortcut of the CSPx module. The main branch consists of a CBM module, x ResUnit modules, and a CBM module in sequence. The shortcut consists of a CBM module. The main branch and the shortcut of the CSPx module perform a Concat operation, and finally there is another CBM module;
[0067] The ResUnit module consists of a main branch and a shortcut. The main branch of the ResUnit consists of two CBM modules, and the shortcut of the ResUnit directly adds the input feature and the output feature of the main branch;
[0068] The CBLx module consists of x CBL modules, and the CBL module consists of a Conv operation, a BatchNorm operation, and a LeakyReLU activation function in sequence;
[0069] The SPP module consists of 4 branches, namely the Maxpool operations with the sizes of 3×3, 5×5, and 9×9, and an empty operation shortcut. Then perform a Concat operation on these 4 branches;
[0070] The shared feature encoder is composed of ① CBM module, ② CSP1 module, ③ CSP2 module, ④ CSP8 module, ⑤ CSP8 module, ⑥ CSP4 module, ⑦ CBL3 module, ⑧ SPP module in sequence;
[0071] Construct the bottleneck module of the visual multi-task processing neural network model, which consists of the ⑨CBL6 module, the ⑩CBL6 module, and the CBL6 module. Requirements: Input the features extracted by the ⑧SPP module into the ⑨CBL6 module, perform an UpSample operation on the output features of the ⑨CBL6 module and perform a Concat operation with the output features of the ⑤CSP8 module to obtain the feature output to the ⑩CBL6 module. Perform an UpSample operation on the output features of the ⑩CBL6 module and perform a Concat operation with the output features of the ④CSP8 module to obtain the feature output to the CBL6 module. Construct the bottleneck module of the visual multi-task processing neural network model in the above order;
[0072] Construct three decoders for different tasks of the visual multi-task processing neural network model, and design three decoders: the object detection decoder, the drivable area detection decoder, and the lane line detection decoder;
[0073] Construct three scales of object detection heads for the constructed object detection decoder. Each object detection head consists of a CBL1 module. The three scales of object detection heads are named Y1, Y2, and Y3 respectively. Y3 obtains the features output by the ⑨CBL6 module, Y2 obtains the features output by the ⑩CBL6 module, and Y1 obtains the features output by the CBL6 module;
[0074] The constructed drivable area detection decoder consists of the CBL3 module and the CBL3 module. Requirements: Input Perform an UpSample operation on the features extracted by the CBL6 module and output the features to the CBL3 module. Perform an UpSample operation on the output features of the CBL3 module and output the features to the CBL3 module;
[0075] The constructed lane line detection decoder consists of the CBL3 module and the CBL3 module. Requirements: Input Perform an UpSample operation on the features extracted by the CBL6 module and output the features to the CBL3 module. Perform an UpSample operation on the output features of the CBL3 module and output the features to the CBL3 module.
[0076] 3) Design of the total loss function. The total loss function is divided into three parts, namely the object detection loss, the drivable area detection loss, and the lane line detection loss, as shown in Equation (1);
[0077] L all = αL det + βL da + γL ll (1)
[0078] In the formula, L all is the total loss value, α, β, and γ are the weight parameters of the object detection loss, the drivable area detection loss, and the lane line detection loss respectively, L det is the object detection loss value, L da is the drivable area detection loss value, L ll is the lane line detection loss value;
[0079] The object detection loss function consists of the localization loss, the object confidence loss, and the classification loss, as shown in Equation (2);
[0080] L det = λ1L ciou + λ2L obj + λ3L cla (2)
[0081]
[0082]
[0083]
[0084] In the formula, λ1, λ2, and λ3 are the weight parameters of the localization loss, the object confidence loss, and the classification loss respectively, L ciou is the localization loss value, L obj is the object confidence loss value, L cla is the classification loss value; C i represents the predicted confidence of the i-th grid bounding box, represents the object confidence of the i-th grid bounding box, represents that there is no object in the j-th bounding box of the i-th grid; c represents the class, c belongs to the total class cl asses , p i (c) represents the probability that the prediction of the i-th grid is class c, represents the probability that the object value of the i-th grid is class c; S represents dividing the image into S×S grids, S 2 represents the total number of grids S×S after image segmentation, the range of i is from 0 to S 2 , B represents the number of bounding boxes in the i-th grid, the range of j is from 0 to B, It is indicated that there is a target in the j-th bounding box of the i-th grid. The CIoU regression localization loss is used, as shown in Formulas (6), (7), and (8). CIoU is an improved algorithm of IoU. IoU represents the intersection over union of two bounding boxes. CIoU takes into account the overlapping area of the bounding boxes, the center distance, and the aspect ratio factors. The CIoU regression localization loss used enables the neural network model to better reflect the accuracy of bounding box localization during training;
[0085]
[0086]
[0087]
[0088] Among them, Formula (6) is the calculation method of CIoU. IoU is the intersection over union, that is, the area of the intersection of two bounding boxes divided by the area of the union, ρ 2 (b, b gt ) represents the Euclidean distance between the predicted bounding box b and the ground truth bounding box b gt . diag represents the length of the diagonal of the minimum bounding rectangle of the two bounding boxes. v is a parameter used to measure the consistency of the aspect ratio, as shown in Formula (7). a is a positive trade-off parameter, as shown in Formula (8). w, h and w gt , h gt represent the height and width of the predicted bounding box and the height and width of the ground truth bounding box respectively;
[0089] Among them, L obj is the target confidence loss value. C i represents the predicted confidence of the bounding box of the i-th grid, represents the target confidence of the bounding box of the i-th grid, represents that there is no target in the j-th bounding box of the i-th grid;
[0090] Among them, L cla is the class loss value. c represents the class. c belongs to classes the total number of classes. p i (c) represents the probability that the prediction of the i-th grid is class c, represents the probability that the target value of the i-th grid is class c;
[0091] Among them, L da is the drivable area detection loss value, as shown in Formula (9). The input is the drivable area detection feature map. p is the pixel value in the feature map. Ω is the set of feature maps, is the true value of the pixel in the target feature map. The cross-entropy method is used to calculate the drivable area detection loss value;
[0092]
[0093] Among them, L ll is the lane line detection loss value as shown in formula (10), where ω1, ω2, ω3 are weight parameters, and L ll_seg is the lane line detection segmentation loss function, calculated using cross-entropy, inputting the lane line detection feature map and having the same calculation method as L da ; L ll_iou is the intersection over union of the predicted lane line and the true lane line, and L ll_exist is the loss for the existence of lane lines;
[0094] L ll = ω1L ll_seg + ω2L ll_iou + ω3L ll_exist (10)
[0095] 4) Use the adam optimizer, set the batch_size value, set the total number of training epochs, print the training loss value and validation loss value obtained in each epoch. If the training loss value decreases and the validation loss value also decreases, it proves that the model has not been trained yet, and continuous training can improve the performance of the model. When the training loss value decreases but the validation loss value increases, it indicates that the model has started to overfit and training needs to be stopped; set different learning rates according to different epochs, and dynamically adjust the learning rate to decrease as the epoch value increases;
[0096] Create a dataset generator, and batch import the BDD100K dataset into the visual multi-task processing neural network model for training; import the training set and validation set divided from the BDD100K dataset; the training set is used to train the visual multi-task processing neural network model. When training, the content imported by the dataset generator includes the input RGB image, object detection bounding box, drivable area detection segmentation image, and lane line detection segmentation image, and the number of imports per batch is controlled by the batch size; the validation set is used to verify the effect of training the visual multi-task processing neural network model, and output the mAP average precision, IoU intersection over union, and Recall recall rate predicted by the object detection decoder, and output the precise accuracy and IoU intersection over union predicted by the drivable area detection decoder and lane line detection decoder;
[0097] Initialize the constructed visual multi-task processing neural network model, dataset generator, and training parameters, start training the visual multi-task processing neural network model, train the model to the set epoch value, freeze the shared feature encoder of the visual multi-task processing neural network model, and then separately train the object detection decoder, drivable area detection decoder, and lane line detection decoder of the visual multi-task processing neural network model in sequence;
[0098] 5) Collect test data. Use the test set images in the BDD100k dataset. Use a GPIO camera or a USB camera to capture an RGB image as the input image. Scale the size of the captured RGB image to the input size N×M of the visual multi-task processing neural network model, and convert it into a tensor of N×M×3. The sample data used in this embodiment is shown in Figure 4 the figure. The original picture is an RGB picture, shown in grayscale;
[0099] Input the obtained tensor of N×M×3 into the visual multi-task processing neural network model for prediction. The model will output the results of three parts, namely the object detection result, the drivable area detection result, and the lane line detection result. Post-process these three results respectively. The object detection obtains multiple results in the format of {x, y, w, h, conf, class}; (x, y) are the coordinates of the bounding box; (w, h) are the width and height of the bounding box; conf is the confidence of the bounding box; class is the category of the object. Set the IoU value and the confidence value, and use the NMS (Non-Maximum Suppression) algorithm to select the suitable bounding box. The effect is shown in Figure 5 the left figure, and draw the bounding box on the image; the drivable area detection and the lane line detection are in the data format of instance segmentation. It is necessary to perform a smoothing operation and a binarization operation on the output matrix to obtain the result. The effect is shown in Figure 5 the right figure, and draw it on the image to obtain an RGB image with the object detection result, the drivable area detection result, and the lane line detection result drawn on it. The effect is shown in Figure 6 the figure. The original picture is an RGB image, shown in grayscale.
[0100] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and shall be included in the protection scope of the present invention.
Claims
1. A vision multi-task processing method based on deep learning, characterized in that This method uses a vision multi-task processing neural network model based on deep learning to simultaneously complete multi-task vision processing required in vehicle autonomous driving, including object detection tasks, drivable area detection tasks, and lane line detection tasks. Among them, the specific situation of this vision multi-task processing neural network model is as follows: a. Construct the input layer of the vision multi-task processing neural network model; b. Construct the shared feature encoder of the vision multi-task processing neural network model. The modules used in the shared feature encoder are CBM module, CSPx module, CBLx module, and SPP module; The CBM module consists of Conv operation, BatchNorm operation, and Mish activation function in sequence; The CSPx module has two branches: a backbone and a shortcut. The CSPx module starts with a CBM module, and this starting CBM module is respectively connected to the backbone and the shortcut of the CSPx module. The backbone consists of a CBM module, x ResUnit modules, and a CBM module in sequence. The shortcut consists of a CBM module. The backbone and the shortcut of the CSPx module perform a Concat operation, and finally, there is another CBM module. The ResUnit module consists of a backbone and a shortcut. The backbone of the ResUnit consists of two CBM modules, and the shortcut of the ResUnit directly adds the input feature to the output feature of the backbone; The CBLx module consists of x CBL modules, and the CBL module consists of Conv operation, BatchNorm operation, and LeakyReLU activation function in sequence; The SPP module consists of 4 branches, namely Maxpool operations with sizes of 3×3, 5×5, and 9×9, and an empty operation shortcut, and then these 4 branches perform a Concat operation; The shared feature encoder is composed of a CBM module, a CSP1 module, a CSP2 module, a first CSP8 module, a second CSP8 module, a CSP4 module, a CBL3 module, and an SPP module in sequence; c. Construct the bottleneck module of the vision multi-task processing neural network model, which consists of three CBL6 modules; d. Define three decoders for different tasks in the vision multi-task processing neural network model, namely: object detection decoder, drivable area detection decoder, and lane line detection decoder; The decoders share feature maps for similar tasks to achieve joint semantic understanding, and the decoder used for the object detection task uses CIoU to measure the loss value; The specific implementation of this vision multi-task processing method includes the following steps: S1. Obtain the dataset and perform preprocessing, including: performing a scaling operation on the dataset to meet the input requirements of the vision multi-task processing neural network model, performing an enhancement operation on the dataset, and performing a style conversion on the dataset to better simulate the actual harsh driving weather environment; dividing the preprocessed dataset into a training set and a test set; S2. Adjust the training parameters, construct a dataset generator, and train the visual multi-task neural network model step by step: first train the shared encoder of the visual multi-task neural network model, and then train the three decoders for different tasks of the visual multi-task neural network model respectively; S3. Collect the RGB image data in the test set, input it into the trained visual multi-task neural network model for prediction, obtain the target detection prediction results, drivable area prediction results, and lane line prediction results, and draw all the prediction results on the test RGB image for display output.
2. The visual multi-task processing method based on deep learning according to claim 1, wherein, In step S1, the BDD100K dataset is used. The original RGB image format of the BDD100K dataset is in jpg format with a resolution of 1280×720. The label format for target detection is in json format. During the training process, the json format labels need to be converted into {x, y, w, h, class}, where (x, y) represents the coordinates of the bounding box, (w, h) represents the width and height of the bounding box, and class represents the category of the target. The drivable area labels and lane line labels in the BDD100K dataset are in png format with a resolution of 1280×720. The sizes of the drivable area labels and lane line labels need to be converted into the sizes of the outputs of the two decoders corresponding to the drivable area detection task and lane line detection task of the visual multi-task neural network model. Perform color transformation enhancement operations on the BDD100K dataset. Use the histogram equalization algorithm to calculate the grayscale histogram of the image, find the total number of image pixels, normalize the histogram distribution, calculate the cumulative score of the grayscale levels of the image, and find the grayscale value of the enhanced image to obtain an image after histogram equalization. Perform scene transformation operations on the dataset. Use the CycleGAN algorithm for scene conversion. CycleGAN is a style conversion neural network. Use CycleGAN to perform weather conversion on the BDD100K dataset, including converting sunny days to thunderstorm days, sunny days to snowy days, sunny days to foggy days, and sunny days to rainy days, to expand the number of autonomous driving datasets and enable the visual multi-task processing neural network model to learn more data images in harsh environments, making the model more generalizable.
3. A visual multi-task processing method based on deep learning according to claim 2, characterized in that, Construct the input layer of the visual multi-task processing neural network model, with the requirements: input an RGB image, and obtain an RGB image with a size of N×M through scaling or cropping. N is the horizontal resolution size of the RGB image after scaling or cropping, and M is the vertical resolution size of the RGB image after scaling or cropping. Then convert it into a tensor with a dimension of N×M×3; Build a bottleneck module for the visual multi-task processing neural network model, which consists of three CBL6 modules, namely the first CBL6 module, the second CBL6 module, and the third CBL6 module. Requirements: Input the features extracted by the SPP module into the first CBL6 module. After performing the UpSample operation on the output features of the first CBL6 module, perform the Concat operation with the output features of the second CSP8 module to obtain the feature output to the second CBL6 module. After performing the UpSample operation on the output features of the second CBL6 module, perform the Concat operation with the output features of the first CSP8 module to obtain the feature output to the third CBL6 module; The constructed object detection decoder has object detection heads at three scales, and each object detection head consists of a CBL1 module. The object detection heads at the three scales are named Y1, Y2, and Y3 respectively; Y3 obtains the features output by the first CBL6 module, Y2 obtains the features output by the second CBL6 module, and Y1 obtains the features output by the third CBL6 module; The constructed drivable area detection decoder consists of two CBL3 modules, namely the first CBL3 module and the second CBL3 module. Requirements: Input the features extracted by the third CBL6 module, perform the UpSample operation and output the features to the first CBL3 module, and perform the UpSample operation on the output features of the first CBL3 module and output the features to the second CBL3 module; The constructed lane line detection decoder consists of two CBL3 modules, namely the third CBL3 module and the fourth CBL3 module. Requirements: Input the features extracted by the third CBL6 module, perform the UpSample operation and output the features to the third CBL3 module, and perform the UpSample operation on the output features of the third CBL3 module and output the features to the fourth CBL3 module.
4. A visual multi-task processing method based on deep learning according to claim 3, characterized in that The visual multi-task processing neural network model designs loss functions for the object detection task, the drivable area detection task, and the lane line detection task respectively, namely the object detection loss, the drivable area detection loss, and the lane line detection loss, as shown in Formula (1); L all = αL det + βL da + γL ll (1) Where L all is the total loss value, α, β, and γ are the weight parameters of the object detection loss, the drivable area detection loss, and the lane line detection loss respectively, and L det is the object detection loss value, L da is the drivable area detection loss value, and L ll is the lane line detection loss value; The object detection loss function consists of a localization loss, an object confidence loss, and a class loss, as shown in Formula (2); L det = λ1L ciou + λ2L obj + λ3L cla (2) Where, λ1 represents the weight parameter of the localization loss, λ2 represents the weight parameter of the object confidence loss, and λ3 represents the weight parameter of the class loss, L ciou is the localization loss value, L obj is the object confidence loss value, L cla is the class loss value; C i represents the predicted confidence of the bounding box of the i-th grid, represents the object confidence of the bounding box of the i-th grid, represents that there is no object in the j-th bounding box of the i-th grid; c represents the class, c belongs to the total class classes, p i (c) represents the probability that the prediction of the i-th grid is the class c, represents the probability that the object value of the i-th grid is the class c; S represents dividing the image into S×S grids, S 2 represents the total number of grids S×S after image segmentation, and the range of i is from 0 to S 2 , B represents the number of bounding boxes in the i-th grid, and the range of j is from 0 to B, represents that there is an object in the j-th bounding box of the i-th grid. The CIoU regression localization loss is used, see Formulas (6), (7), (8). CIoU is an improved algorithm of IoU. IoU represents the intersection-over-union ratio of two bounding boxes. CIoU takes into account the overlapping area, center distance, and aspect ratio factors of the bounding boxes. The CIoU regression localization loss used makes the neural network model more capable of reflecting the localization accuracy of the bounding boxes during training; Among them, formula (6) is the calculation method of CIoU, where IoU is the intersection over union, that is, the area of the intersection of two bounding boxes divided by the area of the union, and ρ 2 (b, b gt ) represents the Euclidean distance between the predicted bounding box b and the ground-truth bounding box b gt , diag represents the length of the diagonal of the minimum bounding rectangle of the two bounding boxes, v is a parameter used to measure the consistency of the aspect ratio, as shown in formula (7), a is a positive trade-off parameter, as shown in formula (8), w, h and w gt 、h gt represent the height and width of the predicted bounding box and the height and width of the ground-truth bounding box, respectively; Define L da is the drivable area detection loss value, as shown in Equation (9). The input is the drivable area detection feature map, p is the pixel value in the feature map, and Ω is the set of feature maps. is the true pixel value in the target feature map. The cross-entropy method is used to calculate the drivable area detection loss value. Define L ll is the lane line detection loss value, see Equation (10), where ω1, ω2, ω3 are weight parameters, and L ll_seg is the lane line detection segmentation loss function, calculated using cross-entropy, inputting the lane line detection feature map and having the same calculation method as L da L ll_iou is the intersection over union of the predicted lane line and the ground truth lane line, and L ll_exist is the loss for the existence of lane lines; L ll = ω1L ll_seg + ω2L ll_iou + ω3L ll_exist (10).
5. A visual multi-task processing method based on deep learning according to claim 4, characterized in that, In step S2, use the adam optimizer, set the batch_size value, set the total number of training epochs, and print the training loss value and the validation loss value obtained in each epoch. If the training loss value decreases and the validation loss value also decreases, it proves that the model has not been trained yet, and continue to train to improve the performance of the model. When the training loss value decreases but the validation loss value increases, it means that the model has started to overfit, and training needs to be stopped; Set different learning rates according to different epochs. As the epoch value increases, dynamically adjust the learning rate to decrease; Create a dataset generator and batch import the BDD100K dataset into the vision multi-task processing neural network model for training; import the training set and validation set divided from the BDD100K dataset. The training set is used to train the vision multi-task processing neural network model. The content imported by the dataset generator during training includes the input RGB image, object detection bounding boxes, drivable area detection segmentation image, and lane line detection segmentation image. The number of imports per batch is controlled by the batch size. The validation set is used to verify the effect of training the vision multi-task processing neural network model, and outputs the mAP average precision, IoU intersection over union, and Recall recall rate predicted by the object detection decoder, and outputs the precise accuracy and IoU intersection over union predicted by the drivable area detection decoder and lane line detection decoder. Initialize the vision multi-task processing neural network model, dataset generator, and training parameters, and start training the vision multi-task processing neural network model. Train the model to the set epoch value, freeze the shared feature encoder of the vision multi-task processing neural network model, and then separately train the object detection decoder, drivable area detection decoder, and lane line detection decoder of the vision multi-task processing neural network model in sequence.
6. A visual multi-task processing method based on deep learning according to claim 5, characterized in that The specific process of step S3 is as follows: S31. Use the test set images in the BDD100k dataset, use a GPIO camera and a USB camera to capture an RGB image as the input image, scale the size of the captured RGB image to the input size N×M of the vision multi-task processing neural network model, and convert it into a tensor of N×M×3. S32. Input the obtained tensor of N×M×3 into the vision multi-task processing neural network model for prediction. The model will output three parts of results, namely object detection results, drivable area detection results, and lane line detection results. Post-process these three results respectively. The object detection obtains multiple results in the format of {x, y, w, h, conf, class}; (x, y) are the coordinates of the bounding box; (w, h) are the width and height of the bounding box; conf is the confidence of the bounding box; class is the category of the object. Set the IoU value and confidence value, and use the NMS non-maximum suppression algorithm to select appropriate bounding boxes and draw the bounding boxes on the image. The drivable area detection and lane line detection are in the data format of instance segmentation. It is necessary to perform a smoothing operation and a binarization operation on the output matrix to obtain the results and draw them on the image to obtain an RGB image with object detection results, drivable area detection results, and lane line detection results drawn on it.
Citation Information
Patent Citations
Pedestrian intention multi-task identification and trajectory prediction method under view angle of intelligent automobile
CN114120439A
Automatic driving automobile road scene understanding method based on multi-task learning
CN115294550A