ADAS perception optimization method based on multi-task learning

By building an ADAS-aware optimization method for multi-task learning, a multi-task network model that shares Backbone, Neck and Head parts is adopted to solve the problems of large amount of computing and information loss in the ADAS system, and improve detection accuracy and operation efficiency.

CN120388273APending Publication Date: 2025-07-29WU HAN XUAN YUAN ZHI JIA KE JI YOU XIAN GONG SI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510511149.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the existing advanced assisted driving ADAS system, the network models of multiple perception tasks are independent and serially operated, resulting in large calculations, low efficiency, loss of information, and complex calculations of lane line decoding parts, affecting the system operation efficiency and accuracy.

Method used

A ADAS-aware optimization method for multi-task learning is constructed, and a multi-task network model that shares Backbone, Neck and Head parts is adopted. Through the improved CSP-Darknet53 structure and 3×3 convolution, combined with multi-task loss function, feature fusion and decoding are realized, and model parameters are optimized.

Benefits of technology

It significantly improves the detection accuracy and operation efficiency of the ADAS perception system, reduces the calculation amount and resource consumption, solves the problems of inefficiency and information loss in traditional methods, and realizes efficient collaborative perception in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388273A_ABST
    Figure CN120388273A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of auxiliary driving systems, and discloses an ADAS perception optimization method based on multi-task learning, and the method comprises the steps: constructing a multi-task network model which comprises a Backbone part, a Neck part and a Head part which are shared; in the Backbone part, an initial C3 module is replaced by a CSP-Darknet53 structure of a C2F module, and initial 6 * 6 convolution is replaced by 3 * 3 convolution; extracting features through the improved Backbone part, and inputting the features into the Neck part of each task for fusion; the features output by the Neck part of each task are decoded through the Head part, and prediction results of target detection, lane line detection, distance estimation and scene classification are generated; by using the multi-task loss function to optimize the model parameters, the detection effect and precision of advanced auxiliary driving can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of assisted driving systems, and in particular to an ADAS perception optimization method based on multi-task learning. Background Art

[0002] In advanced driver assistance systems (ADAS), various perception information is required for subsequent decision-making and planning. The output of object detection information on images, the judgment of object distance values, and the detection information of lane lines all play a crucial role in subsequent FCW, PCW, and LDW. For the perception input part of the camera, the impact of various weather conditions determines the quality of the perception output. Therefore, classifying and judging the camera scene before the perception module input is very important for subsequent object detection, lane line detection, and ranging. Scene classification includes elevated roads, tunnels, basements, highways, etc. These pre-judgment conditions can assist the detection effect and accuracy of the perception module.

[0003] Currently, the infrared ADAS algorithm network method is based on a serial multiple network model approach. There is no connection or association between multiple network tasks, and each task is independent and separate. At the same time, performing image inference on multiple models to obtain multiple network task results will cause information loss.

[0004] For multiple algorithm tasks, deploying four network models for object detection algorithm, lane line detection, ranging, and scene classification has a high computational cost. The network computational amount and parameter quantity are bound to increase, resulting in a relatively large occupancy of the end-side CPU memory and causing the end-side to run smoothly.

[0005] Based on the multi-task network model of the object detection algorithm and the lane line segmentation algorithm, although it solves the problem of the running efficiency of multiple network models, the lane line decoding part using the semantic segmentation upsampling network will also cause a large number of parameter calculations. At the same time, judging and extracting each pixel point will also cause the problem of low running efficiency, and extracting and fitting curves for each segmented pixel point will also introduce a lot of parameter calculations.

[0006] Therefore, the present invention provides an ADAS perception optimization method based on multi-task learning. Summary of the Invention

[0007] This application provides an ADAS perception optimization method based on multi-task learning for improving the detection effect and accuracy of advanced driver assistance.

[0008] In a first aspect, this application provides an ADAS perception optimization method based on multi-task learning, and the method includes:

[0009] Step S1: Construct a multi-task network model. The multi-task network model includes a shared Backbone part, a Neck part, and a Head part. The Backbone part adopts a CSP-Darknet53 structure with the initial C3 module replaced by a C2F module, and uses 3×3 convolution to replace the initial 6×6 convolution;

[0010] Step S2: Extract features through the improved Backbone part, and input the features into the Neck part of each task for fusion;

[0011] Step S3: Decode the features output by the Neck part of each task through the Head part to generate prediction results for object detection, lane line detection, distance estimation, and scene classification;

[0012] Step S4: Optimize the model parameters using a multi-task loss function, and the loss function includes the weighting of each task loss term;

[0013] In the first implementation manner of the present application, the ADAS perception optimization method includes: the Neck part includes an SPPF module. After extracting features through the improved Backbone part, the receptive field is increased through the SPPF module.

[0014] In the second implementation manner of the present application, the ADAS perception optimization method includes: the Neck part adopts a bidirectional feature pyramid including top-down and bottom-up directions, which is used to fuse Low-Level detailed features and High-Level semantic features.

[0015] In the third implementation manner of the present application, the ADAS perception optimization method includes: the multi-task loss function includes an object detection task loss function, a lane line task loss function, a ranging task loss function, and a scene classification task loss function. The formula of the multi-task loss function is:

[0016] L = L det + L ll + L d + L cc

[0017] Where:

[0018] L det represents the object detection task loss function; L ll represents the lane line task loss function;

[0019] L d represents the ranging task loss function; L cc represents the scene classification task loss function.

[0020] In the fourth implementation manner of the present application, the ADAS perception optimization method includes: the target detection task loss function includes classification loss, object loss, and bounding box loss. The classification loss uses the binary cross-entropy loss function, the object loss uses the DFL loss function, and the bounding box loss uses the CIoU loss function.

[0021] In the fifth implementation manner of the present application, the ADAS perception optimization method includes: the lane line task loss function uses the Focal Loss loss function, the ranging task loss function uses the smooth L1 loss function, and the scene classification task loss function uses the L1_Loss loss function.

[0022] In the sixth implementation manner of the present application, the ADAS perception optimization method includes: the Head part includes a target Head part, a lane line Head part, a distance estimation Head, and a scene classification Head. In the validation and prediction modes, the Head part uses three different resolutions as inputs and outputs a tensor including class predictions and bounding boxes.

[0023] In the seventh implementation manner of the present application, the ADAS perception optimization method includes: in the training mode, the Head part periodically applies convolutional layers to each input and outputs three tensors, and each of the tensors includes class predictions and bounding boxes.

[0024] In the eighth implementation manner of the present application, the ADAS perception optimization method includes: the improved Backbone part adopts the structure of the YOLOv8 model

[0025] In a second aspect, the present application further provides an ADAS perception optimization system based on multi-task learning for implementing the above method, including the following modules:

[0026] A model construction module that constructs a multi-task network model. The multi-task network model includes a shared Backbone part, a Neck part, and a Head part. The Backbone part adopts a CSP-Darknet53 structure in which the initial C3 module is replaced by a C2F module, and uses 3×3 convolutions to replace the initial 6×6 convolutions;

[0027] A feature extraction and fusion module that extracts features through the improved Backbone part and inputs the features into the Neck parts of each task for fusion;

[0028] A multi-task processing module that decodes the features output by the Neck parts of each task through the Head part to generate prediction results for target detection, lane line detection, distance estimation, and scene classification;

[0029] A loss optimization module that uses a multi-task loss function to optimize model parameters, and the loss function includes the weighting of each task loss term.

[0030] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0031] The present invention constructs a multi-task network model based on the improved CSP-Darknet53. After the general features extracted by the shared Backbone part are fused by the two-way feature pyramid in the Neck part, they are decoded and output by the dedicated Head part of each task, and are jointly optimized by combining the multi-task loss function, which significantly improves the detection accuracy and operation efficiency of the ADAS perception system. By replacing the C3 module with the C2F module and optimizing the convolutional kernel design, the feature extraction ability is enhanced while the computational amount is reduced; by fusing the low-level details and high-level semantic features through the two-way feature pyramid, the accuracy of multi-scale object detection and lane line positioning is improved; the loss terms of object detection (CIoU, DFL), lane line detection (Focal Loss), ranging (smooth L1), and scene classification (L1 Loss) are integrated using the multi-task loss function, effectively balancing the learning weights between tasks, reducing the number of model parameters and the consumption of edge-side resources, solving the problems of low efficiency and information loss caused by the traditional serial multi-model, and realizing the efficient collaborative perception optimization in complex scenarios. Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0033] Figure 1 It is a flowchart of the ADAS perception optimization method based on multi-task learning in the embodiments of the present application;

[0034] Figure 2 It is a schematic diagram of the multi-task network structure in the embodiments of the present application;

[0035] Figure 3 It is a schematic diagram of the head structure of the multi-task network algorithm in the embodiments of the present application;

[0036] Figure 4 It is a schematic diagram of the ADAS perception optimization system based on multi-task learning of the present application. Detailed Embodiments

[0037] The embodiments of the present application provide a method, a system, and a device for detecting copper foil surface defects based on machine learning. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "including" or "having" and any of its variations are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0038] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 , in the embodiments of the present application, the ADAS perception optimization method based on multi-task learning includes the following steps:

[0039] Step S1: Construct a multi-task network model. The multi-task network model includes a shared Backbone part, a Neck part, and a Head part. The Backbone part adopts a CSP-Darknet53 structure in which the initial C3 module is replaced by a C2F module, and a 3×3 convolution is used to replace the initial 6×6 convolution.

[0040] Specifically, the Backbone part is the basis of the multi-task network model and the basic component for feature extraction. It is specifically used to extract features from the input data and provide conditions for the processing of subsequent tasks. The Neck part is used for object detection tasks, lane line tasks, ranging tasks, and scene classification tasks. The four Neck parts for different tasks respectively correspond to different specific tasks, and specifically process and fuse the features output by the Backbone part to adapt to the task requirements of each Neck part itself. For example, the Neck part for the object detection task will focus on enhancing the features related to the object contour and position, and the Neck part for the lane line task will highlight the information related to the lane line features, etc. Finally, the Head part is responsible for processing the features processed by the Neck part and outputting specific task prediction results, such as identifying the target object and its position and other information.

[0041] Furthermore, the C2F module has better feature extraction ability and network structure optimization compared to the C3 module. It can capture important information in images more effectively, improve the model's perception ability of complex scenes, and is more streamlined and efficient in structure. While reducing the computational cost, it enhances the feature extraction ability, thereby accelerating the model's running speed and improving performance. Replacing the 6×6 convolution with 3×3 convolution can significantly reduce the computational cost, and stacking multiple 3×3 convolutions can achieve a receptive field effect similar to that of a large convolution kernel, and can also increase the network's non-linear expression ability. Deleting the 10th and 14th convolutional layers can reduce the number of model parameters and computational cost, reduce the risk of overfitting, and accelerate the training and inference speed.

[0042] Step S2: Extract features through the improved Backbone part and input the features into the Neck part of each task for fusion.

[0043] Specifically, the Neck part is responsible for fusing the features extracted from the Backbone part. The features output by the Backbone part contain rich image information, and these features are transmitted to different Neck parts. The Neck part integrates the features from different levels or different branches through feature fusion technology to make full use of multi-scale and multi-angle feature information and enhance the model's perception ability of target objects and lane lines, etc. For example, fusing high-level semantic features with low-level detail features enables the model to accurately identify the category when detecting the target and precisely determine the position and shape.

[0044] Moreover, different tasks have different requirements for features. For example, in the object detection task, it may be necessary to fuse features of different scales to take into account the detection of large and small targets. The lane line task pays more attention to the line features in the image. The Neck part fuses and adjusts the features output by the Backbone to provide the most suitable feature representation for each task.

[0045] Step S3: Decode the features output by the Neck part of each task through the Head part to generate prediction results for object detection, lane line detection, distance estimation, and scene classification.

[0046] Specifically, the decoder of the Head part decodes and transforms the features output by the Neck part according to different task requirements for prediction for each task. For example, in the object detection task, the target category and bounding box position are predicted based on the features. In the lane line task, the position and shape of the lane line are predicted, etc. It is the key link to convert feature information into specific task results. Through the processing of the decoder, complex feature information can be converted into specific task outputs, providing a basis for subsequent ADAS decisions.

[0047] Furthermore, the Head part uses a convolutional layer to convert high-dimensional features into class predictions and bounding boxes. The convolutional layer can effectively extract key information from the features. The Head part further processes the high-dimensional features output by the decoder through the convolutional layer and converts them into specific task prediction results such as class predictions and bounding box predictions in object detection tasks, which is an important step before the final output of the entire model.

[0048] Secondly, the convolutional layer can effectively extract and transform features, enabling the model to more accurately identify object classes and determine object positions. At the same time, the Head part also performs certain post-processing on the prediction results, such as non-maximum suppression, to remove duplicate detection boxes and improve the accuracy and stability of detection.

[0049] The Head part adopts a separation method, following the YOLOv8 Detect Head, excluding the objectness branch. The head of other algorithm tasks, such as Figure 3 shown in the detailed design of the multi-task network algorithm head.

[0050] Step S4: Use a multi-task loss function to optimize the model parameters. The loss function includes the weighting of each task loss term.

[0051] Specifically, use the results output by the Head part to calculate the loss function, adopt an end-to-end method, and train the multi-task network model through the multi-task loss function.

[0052] End-to-end training allows the entire model to directly perform joint training from the input data to the output result, that is, from the input raw image data to the final detection result. The entire process is optimized and learned as a whole, avoiding the error accumulation problem caused by traditional step-by-step training. The multi-task loss function combines the losses of different tasks. By simultaneously optimizing the object detection task loss function, lane line task loss function, ranging task loss function, and scene classification task loss function, the model can achieve good performance on multiple tasks simultaneously. During the training process, according to the feedback of the loss function, continuously adjust the model parameters to minimize the loss function value and improve the accuracy and generalization ability of the model.

[0053] The ADAS perception optimization method based on multi-task learning provided by this invention patent constructs a multi-task network model including a shared Backbone part, a Neck part, and a Head part, improves the Backbone part, and performs end-to-end training using a specific multi-task loss function, achieving efficient fusion and optimization of multiple perception tasks, effectively solving problems such as low operating efficiency, high computational cost, and information loss caused by serial detection of multiple network models in traditional infrared ADAS algorithms, significantly improving the perception accuracy and operating efficiency of the model, reducing the time consumption and computational resource consumption of multiple network models on the board, and simplifying the post-processing process at the same time.

[0054] In a specific embodiment, the Neck part includes an SPPF module. After extracting features using the improved Backbone part, the receptive field is increased through the SPPF module.

[0055] Specifically, the SPPF module is a fast spatial pyramid pooling module. Through pooling operations at different scales, the receptive field of the network is increased. A larger receptive field means that the model can obtain more global information, enabling the model to capture image information in a larger range and having better perception ability for some large-sized target objects or complex scene layouts. The SPPF module fuses features from different receptive fields through pooling operations at different scales, improving the detection performance of the model at different scales. At the same time, the SPPF module can also maintain the resolution of the feature map, avoiding information loss caused by excessive pooling and helping to improve the localization accuracy of the model.

[0056] In a specific embodiment, the Neck part adopts a two-way feature pyramid including top-down and bottom-up directions for fusing Low-Level detailed features and High-Level semantic features.

[0057] Specifically, for the Neck part of the object detection task, by detecting various target objects in the image, such as vehicles, pedestrians, traffic signs, etc., to determine their categories and positions, and being responsible for processing and fusing object detection-related features to improve detection accuracy and recall rate. The Neck part of the lane line task focuses on extracting and fusing lane line-related features, including information such as the position, shape, and type of the lane line, enabling the model to accurately detect the position and shape of the lane line and providing a road boundary reference for the vehicle's driving. The Neck part of the distance estimation task processes features for distance measurement. Combining image information and model learning, through feature analysis of target objects in the image, it estimates the distance between the target object and the vehicle to achieve accurate measurement of the distance to the target object. The Neck part of the scene classification task analyzes and fuses the overall features of the image to determine the current scene category, such as basement, tunnel, and regular scene, etc., so as to adopt corresponding driving strategies according to different scenes.

[0058] Furthermore, through a top-down and bottom-up bidirectional feature pyramid, effective fusion of features at different levels is achieved. Low-Level detailed features contain detailed information such as the edges and textures of the target object, which helps to accurately determine the position and shape of the target. High-Level semantic features contain information such as the category and semantics of the target object, which helps to accurately identify the category of the target and plays a key role in identifying the target category and detecting large targets.

[0059] In a specific embodiment, the multi-task loss function includes an object detection task loss function, a lane line task loss function, a ranging task loss function, and a scene classification task loss function. The formula for the multi-task loss function is:

[0060] L = L det + L ll + L d + L cc

[0061] Where:

[0062] L det represents the object detection task loss function; L ll represents the lane line task loss function;

[0063] L d represents the ranging task loss function; L cc represents the scene classification task loss function.

[0064] Specifically, the multi-task loss function adds different task loss functions. When training the model, the performance of each task is considered simultaneously. By adjusting the weights of different task loss functions, the learning focus of the model on each task can be balanced, and multi-task collaborative optimization can be achieved.

[0065] In a specific embodiment, the object detection task loss function includes a classification loss, an object loss, and a bounding box loss. The classification loss uses a binary cross-entropy loss function, the object loss uses a DFL loss function, and the bounding box loss uses a CIoU loss function.

[0066] Specifically, the classification loss uses the binary cross-entropy loss function to measure the difference between the predicted target class and the true class of the model. In object detection, each object may belong to one of multiple classes. This function can effectively calculate the gap between the predicted probability and the true label, guiding the model to learn correct classification. The object loss uses the DFL loss function, mainly to handle the problem of ambiguous samples in object detection. In the actual scenario, some sample annotations are ambiguous. The DFL loss function can more accurately measure the prediction error of the model for these samples, improving the model's ability to process ambiguous samples. The bounding box loss uses the CIoU loss function. When calculating the loss, it not only considers the overlapping area (IoU) between the predicted bounding box and the true bounding box, but also considers factors such as the distance between their center points and the aspect ratio. Compared with the traditional IoU loss function, it can more accurately measure the difference between the predicted bounding box and the true bounding box, enabling the model to converge to the true bounding box faster and more accurately during training and improving the localization accuracy of the model.

[0067] In a specific embodiment, the lane line task loss function uses the Focal Loss loss function, the ranging task loss function uses the smooth L1 loss function, and the scene classification task loss function uses the L1_Loss loss function.

[0068] Specifically, the lane line task loss function uses the Focal Loss loss function because there is an imbalance problem between positive and negative samples in lane line detection, that is, the background pixels (negative samples) are much more than the lane line pixels (positive samples). The Focal Loss loss function reduces the weight for easy-to-classify samples and increases the weight for difficult-to-classify samples, effectively solving the sample imbalance problem, that is, the ratio difference between lane line pixels and non-lane line pixels, making the model more focused on learning lane line features and improving the accuracy of lane line detection.

[0069] The ranging task loss function uses the smooth L1 loss function. When calculating the error between the predicted distance and the true distance, it uses the L2 loss for small errors and the L1 loss for large errors. This characteristic makes it more robust to outliers and can smoothly handle the error between the predicted distance and the true distance, thereby being able to more accurately measure the prediction error and improving the ranging accuracy.

[0070] The scene classification task loss function uses the L1_Loss loss function to calculate the absolute error between the predicted value and the true value, simply and directly measuring the performance of the model in the scene classification task. By calculating the absolute error between the predicted scene class and the true scene class of the model, it effectively guides the model to learn different scene features and accurately achieve scene classification.

[0071] By merging four networks, namely object detection, lane line detection, ranging, and scene classification, into one network model, the semantic information and spatial location information of the context are utilized to effectively reduce the object detection positioning error, filter the object detection information in the area outside the lane lines, reduce the probability of false detection in object detection, and further improve the detection speed and accuracy.

[0072] In a specific embodiment, the Head part includes an object Head part, a lane line Head part, a distance estimation Head, and a scene classification Head. In the validation and prediction modes, the Head part uses three different resolutions as inputs and outputs a tensor including class predictions and bounding boxes.

[0073] Specifically, the Head part is composed of an object Head part, a lane line Head part, a distance estimation Head, and a scene classification Head, each undertaking different tasks. The object Head part is responsible for object detection, identifying various objects in the image and their positions; the lane line Head part focuses on lane line detection, determining the position, shape, and direction of the lane lines; the distance estimation Head estimates the distance between the detected objects and the vehicle; and the scene classification Head classifies the current driving scene.

[0074] In the validation and prediction modes, in order to capture information at different scales and improve the comprehensiveness of detection, the Head part uses three different resolutions as inputs. The high-resolution input can provide rich details, which is beneficial for detecting small objects or accurately identifying them; the low-resolution input has a fast calculation speed and can capture the overall features of the objects, facilitating the detection of large objects or quickly making scene judgments. Finally, the Head part outputs a tensor containing class predictions and bounding boxes. The class predictions present the class to which each detected object belongs in the form of probabilities, and the bounding boxes are used to determine the position and size of the objects in the image, providing reliable perception data for the ADAS system.

[0075] In a specific embodiment, in the training mode, the Head part periodically applies convolutional layers to each input and outputs three tensors, each tensor including class predictions and bounding boxes.

[0076] Specifically, in the training mode, the Head part of this multi-task network model plays a key role. Its purpose is to decode the features output by the Neck part to obtain the prediction results of each task (object detection, lane line detection, distance estimation, and scene classification). To enable the model to better learn features, the Head part will periodically apply convolutional layers to each input. As the core component of a convolutional neural network, the convolutional layer can automatically extract features from the input data. By sliding the convolutional kernel over the input data for convolution operations, it can capture features at different scales and levels. Periodic application means that the convolutional layer processes the input in stages and multiple times, enabling the model to gradually learn features at different levels and scales, from early shallow edge and texture features to later more complex and abstract semantic features, thereby improving the feature expression ability and generalization ability.

[0077] At the same time, the Head part will output three tensors, each of which contains class prediction and bounding box information. Outputting multiple tensors is to enable the model to learn different feature representations at different stages. By comparing the differences between the output results of different tensors and the ground truth labels, the model can be more finely optimized. Class prediction is the model's judgment of the class to which the target belongs, such as determining whether the target is a pedestrian, vehicle, etc. in object detection, or determining whether the scene is an urban road, highway, etc. in scene classification; the bounding box is used to determine the position and size of the target in the image, usually represented by a rectangular box in object detection, containing the coordinate information of the upper left and lower right corners of the target. This processing method helps the model achieve better performance in multiple perception tasks.

[0078] In a specific embodiment, the improved Backbone part adopts the structure of the YOLOv8 model.

[0079] Specifically, the convolutional layer is a basic component of the Backbone part, which extracts the features of the image through convolution operations. Following the design idea of the YOLOv8 architecture means borrowing some concepts of YOLOv8 in terms of model structure, feature extraction method, etc., such as adopting similar multi-scale feature fusion and efficient network structure design to improve the detection speed and accuracy of the model.

[0080] The above described an ADAS perception optimization method based on multi-task learning in the embodiments of the present application. Next, the ADAS perception optimization system based on multi-task learning in the embodiments of the present application will be described. Please refer to Figure 4 , an embodiment of an ADAS perception optimization system based on multi-task learning in the embodiments of the present application includes:

[0081] A model construction module that constructs a multi-task network model, and the multi-task network model includes a shared Backbone part, a Neck part, and a Head part;

[0082] The Backbone part adopts the CSP-Darknet53 structure that replaces the initial C3 module with a C2F module, and uses a 3×3 convolution to replace the initial 6×6 convolution;

[0083] The feature extraction and fusion module extracts features through the improved Backbone part and inputs the features into the Neck part of each task for fusion;

[0084] The multi-task processing module decodes the features output by the Neck part of each task through the Head part to generate prediction results for object detection, lane line detection, distance estimation, and scene classification;

[0085] The loss optimization module uses a multi-task loss function to optimize the model parameters, and the loss function includes the weighting of each task loss term.

[0086] The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An ADAS perception optimization method based on multi-task learning, characterized in that The method includes: Step S1, constructing a multi-task network model, the multi-task network model includes a shared Backbone part, a Neck part and a Head part, the Backbone part adopts a CSP-Darknet53 structure that replaces the initial C3 module with a C2F module, and uses a 3×3 convolution to replace the initial 6×6 convolution; Step S2, extracting features through the improved Backbone part, and inputting the features into the Neck part of each task for fusion; Step S3, decoding the features output by the Neck part of each task through the Head part to generate prediction results for object detection, lane line detection, distance estimation and scene classification; Step S4, using a multi-task loss function to optimize the model parameters, and the loss function includes the weighting of each task loss term.

2. The method according to claim 1, wherein The Neck part includes an SPPF module, and after extracting features through the improved Backbone part, the receptive field is increased through the SPPF module.

3. The method according to claim 1, wherein The Neck part adopts a two-way feature pyramid including top-down and bottom-up to fuse Low-Level detailed features and High-Level semantic features.

4. The method according to claim 1, wherein The multi-task loss function includes an object detection task loss function, a lane line task loss function, a distance measurement task loss function and a scene classification task loss function, and the formula of the multi-task loss function is: L = L det + L ll + L d + L cc Where: L det represents the loss function of the object detection task; L ll represents the loss function of the lane line task; L d represents the loss function of the ranging task; L cc represents the loss function of the scene classification task.

5. The method according to claim 4, characterized in that, The object detection task loss function includes a classification loss, an object loss and a bounding box loss, the classification loss uses a binary cross-entropy loss function, the object loss uses a DFL loss function, and the bounding box loss uses a CIoU loss function.

6. The method according to claim 4, characterized in that, The lane line task loss function uses a Focal Loss loss function, the distance measurement task loss function uses a smooth L1 loss function, and the scene classification task loss function uses an L1_Loss loss function.

7. The method according to claim 1, characterized in that The Head part includes an object Head part, a lane line Head part, a distance estimation Head and a scene classification Head. In the validation and prediction modes, the Head part uses three different resolutions as inputs and outputs a tensor including class predictions and bounding boxes.

8. The method according to claim 7, wherein In the training mode, the Head part periodically applies convolutional layers to each input and outputs three tensors, and each tensor includes class predictions and bounding boxes.

9. The method according to claim 1, characterized in that The improved Backbone part adopts the structure of the YOLOv8 model.

10. An ADAS perception optimization system based on multi-task learning for implementing the method according to any one of claims 1-9, characterized in that, It includes the following modules: A model construction module that constructs a multi-task network model, the multi-task network model includes a shared Backbone part, a Neck part and a Head part, the Backbone part adopts a CSP-Darknet53 structure that replaces the initial C3 module with a C2F module, and uses a 3×3 convolution to replace the initial 6×6 convolution; A feature extraction and fusion module that extracts features through the improved Backbone part and inputs the features into the Neck part of each task for fusion; The multi-task processing module decodes the features output by the Neck part of each task through the Head part to generate prediction results for object detection, lane line detection, distance estimation, and scene classification; The loss optimization module uses a multi-task loss function to optimize the model parameters, and the loss function includes the weighting of each task loss term.