Multi-task automatic driving perception method and device based on deep learning and storage medium
By building a deep learning network model that can complete multiple autonomous driving perception tasks at the same time, the problem that perception tasks in the prior art are difficult to meet practical application needs, and efficient and accurate multi-task perception effects are achieved.
Patent Information
- Application Number
- CN202510185590.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, multi-task learning networks in the field of autonomous driving perception are difficult to meet practical application needs, especially in object detection, feasible area detection and lane line detection.
Using a multi-task autonomous driving perception method based on deep learning, a network model that can simultaneously complete object detection, feasible area detection, lane line detection and height and width detection is constructed. The method includes steps such as data acquisition and labeling, construction and training of network models, and optimization of loss functions.
It realizes the efficiency and accuracy of multi-task perception in the field of autonomous driving perception, meets the actual application needs, and further improves the system's comprehensive perception ability through the addition of height and width detection tasks.
Smart Images

Figure CN120147985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of assisted driving, and in particular to a multi-task autonomous driving perception method, device and storage medium based on deep learning. Background Art
[0002] In the field of deep learning, the technology of applying multi-task learning networks to autonomous driving perception tasks has a certain history. The main purpose of multi-task learning networks is to simultaneously learn and optimize multiple related tasks to improve the generalization ability, efficiency and performance of the model. Regarding multi-task learning algorithms in the field of autonomous driving perception, there are such as Multi-Net, DLT-Net, YOLOP, and these models all use a single network to complete multiple perception tasks at the same time. For the models mentioned above, they all adopt an encoder-decoder network structure. Different decoders are usually used to complete different specific tasks, and different decoders share an encoder to extract features. Such an approach usually enhances the comprehensive perception ability of the model, reduces network overhead, and improves the operation efficiency of the model. In the field of multi-task learning, network models can be roughly divided into two structures: soft parameter sharing and hard parameter sharing. Among the models mentioned above, Multi-Net and YOLOP both adopt the hard parameter sharing network structure, while DLT-Net has interactions between features in the encoder part.
[0003] With the development and progress of autonomous driving technology, object detection, drivable area detection and lane line detection in the perception field, as traditional tasks that often need to be completed in multi-task networks, can no longer meet the needs of actual applications.
[0004] In view of this, the present invention is proposed. Summary of the Invention
[0005] The main purpose of the present invention is to disclose a multi-task autonomous driving perception method, device and storage medium based on deep learning, which is used to solve the problem that object detection, drivable area detection and lane line detection in the perception field in the prior art, as traditional tasks that often need to be completed in multi-task networks, can no longer meet the needs of actual applications.
[0006] To achieve the above object, according to one aspect of the present invention, a multi-task autonomous driving perception method based on deep learning is provided, and the following technical solutions are adopted:
[0007] A multi-task autonomous driving perception method based on deep learning, characterized in that it includes: S101: collecting real scene data under various conditions, annotating part of the collected data, and dividing it into a training set and a verification set according to a certain ratio; S103: constructing a network model that can solve the problem and meet the actual application needs according to the needs; S105: after data preprocessing, the annotated training set and verification set are input into the constructed network model, and a preset number of rounds of training are performed according to a preset method; S107: comparing actual indicators, and selecting an unlabeled data set as a test set to evaluate whether the model performance release meets the application requirements. If so, obtain the optimal model and its corresponding optimal parameters, continuously optimize the model, and deploy the application; if not, return to execute S103.
[0008] Furthermore, the training for a preset number of rounds according to a preset method includes: using a data set to jointly train the three tasks of target detection, drivable area detection, and lane line detection to obtain relatively optimal network weights; loading the relatively optimal network weights into the network model, freezing the weights of the encoder and the weights of the decoder responsible for the three tasks of target detection, drivable area detection, and lane line detection, and using a data set with height and width restrictions to jointly train the height and width restriction detection tasks.
[0009] Furthermore, for the target detection task, the loss function used in training is:
[0010]
[0011]
[0012]
[0013]
[0014] in is the focal loss, is the L1 smoothing loss. represents the true value, Represents the predicted value. Indicates a pre-selected box. is the total number of pre-selected boxes. represents the real frame, is the total number of ground-truth boxes.
[0015] Furthermore, regarding the drivable area detection and lane line detection tasks, the loss function used in training is:
[0016]
[0017] in is the number of samples correctly predicted as positive, is the number of samples mispredicted as the positive class, is the number of samples mispredicted as the negative class, and are hyperparameters of the loss.
[0018] Furthermore, for the height and width limit detection task, the loss function used during training is:
[0019]
[0020]
[0021]
[0022]
[0023]
[0024] Among them, is the height limit loss, is the width limit loss. is the confidence loss, is the offset loss, is the class loss, , , correspond to the constant parameters of the three losses respectively. is the confidence annotation, is the predicted confidence; is the coordinate offset annotation, is the predicted coordinate offset; is the class annotation, is the predicted class.
[0025] Furthermore, the training for a preset number of rounds according to the preset method further includes: The object detection and semantic segmentation tasks use the fused features, and these two tasks share the same feature fusion layer. For the height and width limit detection task, the features extracted by the encoder are directly used.
[0026] According to another aspect of the present invention, a multi-task autonomous driving perception device based on deep learning is provided, and the following technical solutions are adopted:
[0027] The multi-task autonomous driving perception device based on deep learning includes: an annotation module, which is used to collect real-scene data under various conditions, annotate some of the collected data, and divide it into a training set and a validation set according to a certain ratio; a construction module, which is used to construct a network model that can solve problems and meet the actual application requirements according to the needs; an input module, which is used to preprocess the annotated training set and validation set and then input them into the built network model, and perform training for a preset number of rounds according to a preset method; an output module, which is used to compare the actual indicators, select the unannotated data set as the test set to evaluate the model performance. If it meets the application requirements, the optimal model and its corresponding best parameters are obtained, the model is continuously optimized, and the application is deployed; if not, return to start the construction module.
[0028] Further, the input module includes: a training module, which is used to jointly train the data set for three tasks of object detection, drivable area detection, and lane line detection to obtain relatively optimal network weights; a loading module, which is used to load the relatively optimal network weights into the network model, freeze the weights of the encoder and the weights of the decoder responsible for the three tasks of object detection, drivable area detection, and lane line detection, and use the data set with height and width limits to jointly train the height and width limit detection task.
[0029] According to another aspect of the present invention, a storage medium is provided, and the following technical solutions are adopted:
[0030] Based on the existing tasks, the present invention innovatively adds a height and width limit detection task. For the newly added height and width limit detection task, the present invention uses key point detection technology to effectively detect the edge key points of the height and width limit device; for the object detection task, the present invention uses one-stage object detection technology to effectively detect road traffic objects; for the drivable area detection and lane line detection tasks, the present invention uses semantic segmentation technology to effectively detect the drivable area and lane lines. Generally speaking, the present invention innovatively builds an autonomous driving multi-task perception network model that can simultaneously complete object detection, drivable area detection, lane line detection, and height and width limit detection. Brief Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.
[0032] Figure 1 It is a flowchart of a multi-task autonomous driving perception method based on deep learning according to an embodiment of the present invention;
[0033] Figure 2 It is a design optimization flowchart of a multi-task autonomous driving perception method based on deep learning according to an embodiment of the present invention;
[0034] Figure 3 It is a network flowchart of the model according to an embodiment of the present invention;
[0035] Figure 4 It is a diagram showing the invention effect of the present invention;
[0036] Figure 5 It is another diagram showing the invention effect of the present invention;
[0037] Figure 6 The invention model flowchart according to an embodiment of the present invention; and
[0038] Figure 7 It is a schematic structural diagram of a multi-task autonomous driving perception device based on deep learning according to an embodiment of the present invention. Specific embodiments
[0039] The following will describe the embodiments of the present invention in detail with reference to the accompanying drawings. However, the present invention can be implemented in many different ways defined and covered by the claims.
[0040] Figure 1 It is a flowchart of a multi-task autonomous driving perception method based on deep learning according to an embodiment of the present invention.
[0041] Figure 2 It is a design optimization flowchart of the model according to an embodiment of the present invention.
[0042] See Figure 1 As shown, a multi-task autonomous driving perception method based on deep learning includes:
[0043] A multi-task autonomous driving perception method based on deep learning, characterized in that it includes: S101: Collect real-scene data under various conditions, label some of the collected data, and divide it into a training set and a validation set according to a certain ratio; S103: Construct a network model that can solve problems and meet the actual application requirements according to the needs; S105: After preprocessing the labeled training set and validation set, input them into the built network model and train them for a preset number of rounds according to a preset method; S107: Compare the actual indicators, and select the unlabeled data set as the test set to evaluate the model performance. If it meets the application requirements, obtain the optimal model and its corresponding best parameters, continuously optimize the model, and deploy and apply it; if not, return to execute S103.
[0044] More specifically, see Figure 2 As shown, a multi-task autonomous driving perception method based on deep learning includes:
[0045] Step 20: Collect real-scene data under various conditions, label some of the collected data, and divide it into a training set and a validation set according to a certain ratio;
[0046] Step 21: Design and build a network model that can solve problems and meet the actual application requirements according to the needs;
[0047] Step 22: After preprocessing the labeled training set and validation set, input them into the built network model and train for a certain number of rounds according to a certain method;
[0048] Step 23: Compare the experimental indicators and select the unlabeled data set as the test set to evaluate whether the model performance meets the application requirements? If so, execute Step 24; if not, return to execute Step 21;
[0049] Step 24: Obtain the optimal model and its corresponding best training parameters, continuously optimize the model and deploy it for use.
[0050] Figure 3 This is the network flow chart of the model described in the embodiment of the present invention.
[0051] Furthermore, for the training strategy of the network model, the present invention first uses the BDD100K data set to perform joint training on the three tasks of object detection, drivable area detection, and lane line detection to obtain relatively optimal network weights; load the previously obtained optimal network weights into the network model again, freeze the weights of the encoder and the weights of the decoder responsible for the above three tasks (object detection, drivable area detection, lane line detection), and use the data set with height and width limits to perform joint training on the height and width limit detection task. As Figure 3 shown, the labeled data set 30 is input to the encoder 31, and the decoder 32 outputs respectively, the object detection result 33, the drivable area detection task result 34, the lane line detection task result 35, and the height and width limit detection task result 36. For the specific effect, see Figures 4 - 5 shown. In the legend, the white area on the ground is the drivable area detection result, the black area is the lane line detection result; the white wire frame in the figure is the traffic object detection result; the white dots in the figure are the height and width limit detection results.
[0052] EfficientNets has a smaller model size and fewer parameters while maintaining accuracy. This means that EfficientNets performs better when computing resources are limited, such as when deployed on mobile devices or edge devices. EfficientNets can be scaled according to the needs of the task, thereby achieving a flexible balance of performance and resources in different application scenarios. This scalability makes it suitable for a variety of computing resources and application scenarios. Due to the excellent performance of EfficientNets, this paper uses EfficientNets_b3 as the encoder of the proposed network model to extract the features of the input information. For details, see Figure 6 shown.
[0053] The model of the present invention uses Bilateral Feature Pyramid Network (BiFPN) to fuse the features extracted by the encoder. BiFPN is a feature fusion network structure used to improve the performance of network models, especially in some advanced visual tasks such as target detection and semantic segmentation. It is improved on the basis of Feature Pyramid Network (FPN) by introducing bilateral connections and feature weighting to enhance the expressiveness of features and the transmission of semantic information, thereby improving the expressiveness of features and the performance of the network while maintaining the advantages of computational efficiency and light weight. In the model of the present invention, the two types of tasks, target detection and semantic segmentation, use features fused by BiFPN, and these two types of tasks share the same feature fusion layer. As for the height and width limit detection task, the features extracted by the encoder are directly used. For details, see Figure 6 shown.
[0054] Regarding the target detection task, the loss function used in its training, the present invention uses Focal loss and L1 smoothing loss to train the model for convergence. Its mathematical formula is as follows:
[0055]
[0056]
[0057]
[0058]
[0059] in is the focal loss, is the L1 smoothing loss. represents the true value, Represents the predicted value. Indicates a pre-selected box. is the total number of pre-selected boxes. Denotes the ground truth box, is the total number of ground truth boxes.
[0060] Regarding the drivable area detection and lane line detection tasks, for the loss functions used during training, the present invention adopts the Focal loss and Tvershy loss to train the model for convergence. The mathematical formulas are as follows:
[0061]
[0062] Where is the number of samples correctly predicted as positive classes, is the number of samples wrongly predicted as positive classes, is the number of samples wrongly predicted as negative classes, and are hyperparameters of the loss.
[0063] Regarding the height and width limit detection task, for the loss functions used during training, the present invention adopts the mean squared error loss and cross-entropy loss to train the model for convergence. The mathematical formulas are as follows:
[0064]
[0065]
[0066]
[0067]
[0068]
[0069] Where, is the height limit loss, is the width limit loss. is the confidence loss, is the offset loss, is the class loss, 、 、 correspond to the constant parameters of the three losses respectively. is the confidence annotation, is the predicted confidence; is the coordinate offset annotation, is the predicted coordinate offset; is the class annotation, is the predicted class.
[0070] In Table 1, the index results of the height and width limit detection task of the present invention are shown; and when only the height limit detection task is performed, compared with the previously proposed height limit detection network MF-KNet, the F1 index result of the present invention is improved by 2.4 percentage points. Whether it is the height and width limit joint detection task or the single height limit detection task, the F1 index result of the model of the present invention is better than that of MF-KNet.
[0071] In Table 2, the index comparison results of the object detection task of the present invention are shown. Compared with several other network models, the recall rate of the model of the present invention reaches the highest 94.1%, but the mAP50 result value of the model of the present invention still needs to be improved.
[0072] In Table 3, the index result comparison of the drivable area detection task of the present invention is shown. By observing and comparing the experimental index mIoU in Table 3, it can be found that the drivable area detection comparison of the network model proposed by the present invention does not reach the optimal result, but also reaches the sub-optimal result.
[0073] Table 1 Comparison of index results of the height and width limit detection task and the single height limit detection task of the present invention
[0074] Table 2 Comparison of index results of the object detection task of the present invention
[0075] Table 3 Comparison of index results of the drivable area detection task of the present invention
[0076] In Table 4, the index result comparison of the lane line detection task of the present invention is shown. By observing and comparing Table 4, it can be found that the comparison of the model in this article reaches the optimal in both of the two experimental indexes of Accuracy and IoU.
[0077] Table 4 Comparison of index results of the lane line detection task of the present invention
[0078] Figure 7 It is a schematic structural diagram of a multi-task automatic driving perception device based on deep learning according to an embodiment of the present invention.
[0079] See Figure 7As shown in the figure, the multi-task autonomous driving perception device based on deep learning includes: an annotation module 70, which is used to collect real-scene data under various conditions, annotate some of the collected data, and divide it into a training set and a validation set according to a certain ratio; a construction module 72, which is used to construct a network model that can solve problems and meet the actual application requirements according to the needs; an input module 74, which is used to preprocess the annotated training set and validation set and then input them into the built network model, and train them for a preset number of rounds according to a preset method; an output module 76, which is used to compare the actual indicators, select the unannotated data set as the test set to evaluate the model performance to release and meet the application requirements. If so, obtain the optimal model and its corresponding best parameters, continuously optimize the model, and deploy the application; if not, return to start the construction module.
[0080] Preferably, the input module 74 includes: a training module (not shown in the figure), which is used to jointly train the data set for three tasks: object detection, drivable area detection, and lane line detection, to obtain relatively optimal network weights; a loading module (not shown in the figure), which is used to load the relatively optimal network weights into the network model, freeze the weights of the encoder and the weights of the decoder responsible for the three tasks of object detection, drivable area detection, and lane line detection, and jointly train the height and width limit detection task with the data set of height and width limits.
[0081] The storage medium provided by the present invention includes the above-mentioned multi-task autonomous driving perception method based on deep learning.
[0082] The present invention innovatively adds a height and width limit detection task on the basis of the existing tasks. For the newly added height and width limit detection task, the present invention uses the key point detection technology to effectively detect the edge key points of the height and width limit device; for the object detection task, the present invention uses the one-stage object detection technology to effectively detect the road traffic objects; for the drivable area detection and lane line detection tasks, the present invention uses the semantic segmentation technology to effectively detect the drivable area and lane lines. Generally speaking, the present invention innovatively builds an autonomous driving multi-task perception network model that can simultaneously complete object detection, drivable area detection, lane line detection, and height and width limit detection.
[0083] Only some exemplary embodiments of the present embodiment are described by way of illustration above. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the protection scope of the claims of the present invention.
Claims
1. A multi-task autonomous driving perception method based on deep learning, characterized in that: include: S101: Collect real scene data under various conditions, annotate part of the collected data, and divide it into a training set and a verification set according to a certain ratio; S103: construct a network model that can solve the problem and meet the actual application requirements according to the requirements; S105: After data preprocessing, the labeled training set and validation set are input into the constructed network model, and a preset number of rounds of training are performed according to a preset method; S107: Compare the actual indicators and select an unlabeled data set as a test set to evaluate whether the model performance release meets the application requirements. If so, obtain the optimal model and its corresponding optimal parameters, continue to optimize the model, and deploy the application; if not, return to execute S103.
2. The multi-task autonomous driving perception method according to claim 1, characterized in that: The training of a preset number of rounds according to a preset method includes: The dataset is used for joint training of three tasks: target detection, drivable area detection, and lane line detection, to obtain the relatively optimal network weights. The relatively optimal network weights are loaded into the network model, and the weights of the encoder and the decoder responsible for the three tasks of target detection, drivable area detection, and lane line detection are frozen. The height and width limit detection tasks are jointly trained using the height and width limit data set.
3. The multi-task autonomous driving perception method according to claim 2, characterized in that: Regarding the target detection task, the loss function used in training is: in is the focal loss, is the L1 smoothing loss, represents the true value, represents the predicted value, Indicates a pre-selected box. is the total number of pre-selected boxes, represents the real frame, is the total number of ground-truth boxes.
4. The multi-task autonomous driving perception method according to claim 3, characterized in that: Regarding the drivable area detection and lane line detection tasks, the loss function used in training is: in is the number of samples correctly predicted as positive, is the number of samples incorrectly predicted as positive, is the number of samples incorrectly predicted as negative, and is the hyperparameter of the loss.
5. The multi-task autonomous driving perception method according to claim 4, characterized in that: Regarding the height and width limit detection task, the loss function used in training is: in, It is the maximum loss limit. is the width limitation loss, is the confidence loss, is the offset loss, is the category loss, , , They correspond to the constant parameters of the three losses respectively. is the confidence mark, is the prediction confidence; is the coordinate offset label, To predict the coordinate offset; Label the category. is the predicted category.
6. The multi-task autonomous driving perception method according to claim 5, characterized in that: The training of a preset number of rounds according to a preset method also includes: The two types of tasks, object detection and semantic segmentation, use the fused features, and these two types of tasks share the same feature fusion layer. The height and width limit detection task directly uses the features extracted by the encoder.
7. A multi-task autonomous driving perception device based on deep learning, characterized in that: include: The annotation module is used to collect real scene data under various conditions, annotate part of the collected data, and divide it into training set and verification set according to a certain ratio; Construction module, used to construct a network model that can solve the problem and meet the actual application requirements according to the requirements; The input module is used to input the labeled training set and validation set into the constructed network model after data preprocessing, and perform a preset number of rounds of training according to the preset method; The output module is used to compare actual indicators and select unlabeled data sets as test sets to evaluate whether the model performance release meets the application requirements. If so, the optimal model and its corresponding optimal parameters are obtained, the model is continuously optimized, and the application is deployed; if not, it returns to the startup construction module.
8. The multi-task autonomous driving perception device according to claim 7, characterized in that: The input module comprises: The training module is used to perform joint training on the three tasks of target detection, drivable area detection, and lane line detection on the dataset to obtain the relatively optimal network weights; The loading module is used to load the relatively optimal network weights into the network model, freeze the weights of the encoder and the weights of the decoder responsible for the three tasks of target detection, drivable area detection, and lane line detection, and use the height and width limit data set to jointly train the height and width limit detection tasks.
9. A storage medium, characterized in that: A multi-task autonomous driving perception method comprising the method described in any one of claims 1-6.