Multi-task detection method for comprehensive driving perception
By building a multi-task detection model, combining CSPDarknet53 and feature pyramid network, the SimAttention mechanism is introduced, which solves the problem of multi-task processing delay in the existing technology, and realizes efficient driving perception, which is suitable for autonomous driving systems.
Patent Information
- Application Number
- CN202510692155.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-27
AI Technical Summary
In existing driving perception systems, vision-based tasks such as object detection, image segmentation and lane line detection usually require multiple independent network processing, resulting in computational delay and inefficiency, making it difficult to achieve real-time and effective environmental perception.
A multi-task detection model is built, including backbone network, neck network, detection head, segmentation head and deep information extraction head. CSPDarknet53 is used as the backbone network basis, combined with the C2fA module and feature pyramid network, a SimAttention attention mechanism is introduced, feature extraction and fusion is optimized, and end-to-end multi-task processing is realized.
It improves the speed and accuracy of driving perception, and can handle traffic object detection, segmentation and depth generation tasks at the same time, significantly improves the performance and recall of the model, and is suitable for real-time panoramic driving perception.
Smart Images

Figure CN120526397A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of driving perception technology, and in particular to a multi-task detection method for comprehensive driving perception. Background Art
[0002] The rapid development of pure computer vision algorithms and end-to-end deep learning technology has made autonomous driving a research hotspot in the field of computer vision. Despite this, in the process of realizing low-cost autonomous driving systems, vision-based tasks such as object detection and image segmentation still face many challenges. Given the cost-effectiveness of camera solutions, they have dominated the actual deployment of object detection and segmentation tasks. Therefore, how to ensure that the vehicle can achieve effective environmental perception and provide more complete information to the downstream with only a single camera input in front has become a key issue that needs to be solved urgently. Traditionally, there are three key tasks in environmental perception that are considered to be the basis for guiding intelligent vehicles: traffic object detection, drivable area segmentation, and lane line detection.
[0003] In recent years, with the rapid development of deep learning, many well-known object detection algorithms have emerged. Currently, mainstream object detection algorithms can be divided into two-stage and single-stage methods. Two-stage methods complete the detection task in two steps: region proposal and object classification. The most typical example is the R-CNN series of algorithms. While they offer high detection accuracy and stability, they are generally slower than single-stage methods. Therefore, single-stage methods are more popular in real-time detection applications.
[0004] Existing object detection methods include single-shot multi-box detectors and the YOLO family, while semantic segmentation is handled by networks such as U-Net, SegNet, and ERNet. For lane detection, models such as LaneNet and spatial convolutional neural networks (SCNN) are used. However, using three independent networks to process the same image data stream can lead to unnecessary latency.
[0005] To address this issue, many existing techniques integrate these features into a single encoder-decoder architecture, where the backbone and neck act as encoders to collaboratively generate context for three different tasks, such as in MultiNet, DLT-Net, YOLOP, and hybridnet. Multi-task learning networks provide a computationally efficient solution by using an encoder-decoder framework, where the encoder is effectively shared between different tasks, where existing models have shortcomings in speed and performance. Summary of the Invention
[0006] In order to solve the above technical problem or at least partially solve the above technical problem, the present invention provides a multi-task detection method for comprehensive driving perception.
[0007] The present invention provides a multi-task detection method for comprehensive driving perception, comprising: Construct an end-to-end multi-task detection model that collaboratively handles traffic object detection, segmentation, and depth generation tasks; the multi-task detection model includes a backbone network, a neck network, a detection head, a segmentation head, and a depth information extraction head; the backbone network and the neck network cooperate to form an encoder, wherein the backbone network is used to extract image features, and the neck network is used to fuse features generated by the backbone network; The detection head identifies multiple targets, including detection frames and confidence levels for vehicles and pedestrians, and then generates secondary labels representing the status of other vehicles. Based on these secondary labels, the driver determines whether to yield or follow the vehicle. The segmentation head segments multiple categories of targets, including drivable areas, vehicles, pedestrians, and background areas, and based on the classification information, provides label information indicating whether the drivable area ahead is clean; The depth information extraction head generates depth information; The obtained detection information, segmentation information, depth information, and state information are integrated with the target detection information at time t-1 of the previous frame to obtain the detection timing information. Finally, the model integrates all the information to obtain the final result, which is more conducive to perception information and pre-decision information for subsequent tasks.
[0008] Furthermore, the backbone network implements multi-scale feature extraction based on CSPDarknet53, and in order to enhance the information extraction capability, the C2F module connected to the output stage in each stage of CSPDarknet53 is optimized to a C2fA module.
[0009] Furthermore, the neck network uses a simple spatial pyramid fast pooling module to process the deepest feature map from the backbone network. The simple spatial pyramid fast pooling module includes: an initial CBR module, which performs a convolution operation on the input feature map to extract features, and accelerates the training process and improves the stability of the model through batch normalization. The ReLU activation function introduces nonlinearity, enhances the expressiveness of the model, and reduces the number of channels to half of the original number of channels; then, through multiple cascaded maximum pooling operations, the results of the maximum pooling operations of each layer are spliced, and finally, a CBR module is used to convert the fused feature map into a specified number of output channels.
[0010] Furthermore, in the feature fusion process, the neck network adopts a feature pyramid network structure, and the C2F module alternately arranged with the CBS module is optimized into a C2fA module to effectively combine features at different semantic levels.
[0011] Furthermore, any feature map input to the C2fA module is decomposed into two parts by the CBS module, where the CBS module contains a convolutional layer, a batch normalization layer, and a SiLU activation function; one part is processed by a CBS module and input into a convolutional layer with multiple convolution kernel sizes, and multiple convolutional layers process features in parallel; the results of multiple convolution parallel processing are fused by the CBS module, processed by the CBR module, and then spliced and combined with the other part; the CBR module contains a convolutional layer, a batch normalization layer, and a ReLU activation function layer.
[0012] Furthermore, the detection head adopts the anchor-based multi-scale detection head used in YOLOP and integrates a path aggregation network and a feature pyramid network. The SimAttention mechanism is introduced before each detection head of the multi-scale detection head. The detection head uses three anchors with different aspect ratios to predict the position offset and scale change of the target for each grid unit on the multi-scale fusion feature map. The process generates the probability, confidence score and secondary label of each detection category. Among them, the three detection head information in the multi-scale detection head are obtained, with dimensions of [20, 20, 10], [40, 40, 10], and [80, 80, 10] respectively. The first two dimensions represent size information, and the second dimension represents the probability, confidence score and secondary label of each detection category: the secondary label is the binary classification information of dead or alive vehicles. The status information of whether to give way or follow the vehicle is determined based on the detection category of the secondary label.
[0013] Furthermore, the segmentation head and the depth information extraction head have the same structure, including the following combination in sequence: CBS module, upsampling using nearest neighbor interpolation, C2fA module, CBS module, upsampling, CBS module, C2fA module, upsampling, CBS module.
[0014] Furthermore, the depth information extraction head uses the output of a simple spatial pyramid fast pooling module to extract depth information. The depth head defines D discrete depth planes parallel to the image plane along the optical axis. Each depth plane and the image feature plane together form the viewing cone space. Assuming the lowest depth range is a meters, the maximum depth range is b meters, and the interval is c meters, a total of D = (b a) / c + 1 discrete depth planes are generated. Each pixel (h, w) corresponds to D possible 3D points in the viewing frustum, and the depth of each point is a discrete value d; The input feature map is replicated and expanded along the depth dimension into a 4-dimensional tensor as a depth feature, which represents the characteristic response of each pixel at different depth planes; the depth feature is combined with the depth encoding; For each pixel point (h, w) and depth plane d, calculate its 3D coordinates in the camera coordinate system; The model predicts the probability distribution of pixels in the depth plane; The final depth information includes: 3D coordinates, depth features combined with depth coding, and probability distribution.
[0015] Furthermore, the total loss function of the multi-task detection model is as follows: ; in, For traffic object detection loss, is the segmentation loss, is the depth information loss, Both are balanced 、 and The balance factor.
[0016] Furthermore, traffic object detection loss It is defined as the combination of the following components: Penalized classification loss , confidence loss and bounding box regression loss , expressed as: ; in, Balance 、 and The balance factor, where the penalty classification loss and confidence loss Using zoom loss; bounding box regression loss Use Distance-IoU loss; Contains the classification error cross entropy loss between the pixels output by the network and the target and the IoU loss between the segmentation mask output by the network and the true mask; Using cross entropy loss, the depth distance predicted by the model is calculated with the true value to guide the learning process of the model.
[0017] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art: To improve the speed and performance of our model, we adopted CSPDarknet as the basis of the backbone network and optimized the backbone network using the C2fA module to effectively combine features at different semantic levels. The backbone network strikes a balance between accuracy and computational cost. In the decoder, we retained YOLOP's anchor-based multi-scale detection scheme and combined it with a path aggregation network and a feature pyramid network. The integration of the feature pyramid network and the path aggregation network enhances feature fusion by promoting the top-to-bottom propagation of semantic features and the bottom-to-top propagation of positioning features. In addition, we introduced the SimAttention mechanism before the detection head to enhance the overall performance. At an input resolution of 640x640, our network takes slightly longer to converge than YOLOP, but significantly improves accuracy and recall.
[0018] This application simultaneously addresses three driving perception tasks: traffic object detection, traffic object segmentation, and depth generation, and supports end-to-end training. Our model demonstrates outstanding performance on the SDExpressway and BDD100K datasets for traffic object detection and segmentation, surpassing existing technologies in terms of accuracy and efficiency.
[0019] This application integrates all the information to obtain the final result, obtain perception information that is more conducive to subsequent tasks, and pre-decision information, providing effective support for intelligent driving decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0022] Figure 1 A schematic diagram of the architecture of a multi-task detection model provided by an embodiment of the present invention; Figure 2 A schematic diagram of the architecture of a multi-task detection model represented by a feature graph provided in an embodiment of the present invention; Figure 3 A schematic diagram of focus in a backbone network provided by an embodiment of the present invention; Figure 4 Schematic diagram of the CBS and CBR modules provided in an embodiment of the present invention; Figure 5A schematic diagram of a C2fA module provided in an embodiment of the present invention; Figure 6 This is a rendering of the multi-task detection model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0024] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0025] The panoramic driving perception system is crucial for autonomous driving as it provides key traffic-related information. This application constructs a multi-task detection model for comprehensive driving perception, which is an efficient and effective multi-task learning framework designed for real-time panoramic driving perception. The multi-task detection model contains an encoder for feature extraction and three decoders that simultaneously perform traffic object detection, traffic object segmentation, and depth information generation. For feature extraction, a C2fA module is proposed to enhance the extraction capability of the model. To enhance our dataset, we expanded the SDExpressway dataset to include nighttime and severe weather conditions.
[0026] Example 1 The multi-task detection method for comprehensive driving perception provided in this application includes: Build an end-to-end multi-task detection model that can collaboratively handle traffic object detection, segmentation, and depth generation tasks. Figure 1 and Figure 2As shown, the multi-task detection model includes a backbone network, a neck network, a detection head, a segmentation head, and a depth information extraction head. The backbone network and the neck network work together to form an encoder. The backbone network is used to extract image features, and the neck network is used to fuse the features generated by the backbone network. The detection head, segmentation head, and depth information extraction head are used for traffic object detection, traffic object segmentation, and depth information generation, respectively.
[0027] Next, we introduce in detail the design of the multi-task detection model architecture and the optimization strategies implemented to improve multi-task performance.
[0028] The backbone network is designed to achieve multi-scale feature extraction. Feature extraction is a key step in multi-task learning and directly affects the performance of the multi-task detection model on various tasks. This application uses CSPDarknet53 as the basis of the backbone network of the multi-task detection model, and in order to enhance the information extraction capability, the C2F module connected to the output stage in each stage of CSPDarknet53 is optimized to a C2fA module. The C2fA module is an improvement on the C2f module, such as Figure 5 As shown. Any feature map input to the C2fA module is decomposed into two parts by the CBS module. In this application, Figure 4 As shown in the figure, the CBS module contains convolutional layers, batch normalization layers, and SiLU activation functions. After a portion of the data is processed by a CBS module, it is input into multiple convolutional layers with different convolution kernel sizes. Multiple convolutional layers process features in parallel, allowing for more extensive integration of feature information, providing significant advantages. The results of multiple convolutional parallel processing are fused by the CBS module, processed by the CBR module, and then spliced and combined with the other portion. Figure 4 As shown in the figure, the CBR module contains a convolutional layer, a batch normalization layer, and a ReLU activation function layer. The C2fA module also contains an additional convolutional layer compared to C2f, which helps to extract more complex and deeper features. This design not only reduces the gradient duplication in the optimization process, but also improves feature propagation and feature reuse. Experimental results show that the use of multiple parallel convolutional layers and multi-layer perception layers significantly improves the model's ability to extract depth information from images, thereby improving the receptive field and feature representation capabilities. The CSPDarknet53 input is first processed by focus, as shown in the following example. Figure 3 As shown in Figure 2, the focus layer is a spatial-channel reorganization operation used to reduce computational effort while preserving feature information. It first appeared in YOLOv5 as an alternative to traditional, computationally intensive downsampling methods. Focus splits the input image into four subimages along a 2x2 grid, concatenates these four subimages along the channel dimension, and reorganizes the input image slices. This reduces spatial resolution while increasing the number of channels, thereby reducing computational effort and preserving more information.
[0029] The neck network aims to fuse multi-scale features generated by the backbone network.
[0030] The neck network uses a simple spatial pyramid fast pooling module to process the deepest feature maps from the backbone network. Furthermore, an attention layer is added after the simple spatial pyramid fast pooling module to further enhance feature extraction. The simple spatial pyramid fast pooling module includes an initial CBR module for preliminary processing of the input feature map, reducing the number of channels to half the original number. It performs convolution operations on the input feature map to extract features, and batch normalization is used to accelerate training and improve model stability. The ReLU activation function introduces nonlinearity to enhance the model's expressiveness. Multiple cascaded max pooling and concatenation operations are performed: Multiple cascaded max pooling operations and concatenation of the results of each layer's max pooling operations are used to fuse features of different scales. Finally, a CBR module converts the fused feature map to the specified number of output channels. Through this design, the simple spatial pyramid fast pooling module can effectively extract multi-scale features and fuse these features to enhance the model's ability to recognize objects of different sizes. Furthermore, its simplified design improves computational efficiency, making it suitable for computer vision tasks with high real-time requirements. Ablation experiments verify that the simple spatial pyramid fast pooling module balances speed and accuracy.
[0031] In the feature fusion process, the neck network adopts a feature pyramid network structure, and the C2F module, which is arranged alternately with the CBS module, is optimized as a C2fA module to effectively combine features at different semantic levels. This design ultimately produces feature information containing multiple scales and semantic levels, which will be passed to the three decoders.
[0032] The detection head is designed to recognize various object categories, including vehicles, traffic signs, and road markings. In this application, the detection head adopts the anchor-based multi-scale detection head used in YOLOP and integrates a path aggregation network and a feature pyramid network. The feature pyramid network promotes the top-down flow of semantic information, while the path aggregation network enables the effective bottom-up propagation of localization details. This synergistic combination enhances feature fusion capabilities. To improve the detection performance of small objects, a SimAttention mechanism is introduced before each detection head in the multi-scale detection head. Unlike existing channel or spatial attention modules, SimAttention focuses on the spatial dimension, assigning different weights to different spatial locations within the feature map, thereby improving the overall performance of the task. Notably, SimAttention also shows certain advantages in terms of computational cost. The detection head utilizes three anchors with different aspect ratios to predict the position offset and scale change of the object for each grid cell in the multi-scale fused feature map. The process generates the probability, confidence score and secondary label of each detection category; among them, the three detection head information in the multi-scale detection head are obtained, and the dimensions are [20, 20, 10], [40, 40, 10], [80, 80, 10] respectively. The first two dimensions represent the size information, and the second dimension represents the probability, confidence score and secondary label of each detection category: the secondary label is the binary classification information of dead or alive vehicles; the status information of whether to give way or follow the vehicle is determined based on the detection category of the secondary label. Figure 6 As shown, the secondary label is Figure 6 The D and L in the code represent the information of dead or alive vehicles, respectively. The decision of whether to follow the vehicle or take a detour is based on the information of dead or alive vehicles. The segmentation head and the depth information extraction head have the same structure, including the following combination in sequence: CBS module, upsampling using nearest neighbor interpolation, C2fA module, CBS module, upsampling, CBS module, C2fA module, upsampling, CBS module.
[0033] The segmentation head uses the feature maps of the shallow layer of the feature pyramid.
[0034] The depth information extraction head uses the output of a simple spatial pyramid fast pooling module to extract depth information. The depth head defines D discrete depth planes parallel to the image plane along the optical axis. Each depth plane and the image feature plane together form the viewing cone space. Assuming the lowest depth range is a meters, the maximum depth range is b meters, and the interval is c meters, a total of D = (b a) / c + 1 discrete depth planes are generated. Each pixel (h, w) corresponds to D possible 3D points in the viewing frustum, and the depth of each point is a discrete value d; The input feature map is replicated and expanded along the depth dimension into a 4-dimensional tensor as a depth feature, which represents the characteristic response of each pixel at different depth planes; the depth feature is combined with the depth encoding; For each pixel point (h, w) and depth plane d, calculate its 3D coordinates in the camera coordinate system; The model predicts the probability distribution of pixels in the depth plane; The final depth information includes: 3D coordinates, depth features combined with depth coding, and probability distribution.
[0035] In the design of multi-task loss function, traffic object detection loss It is defined as the combination of the following components: Penalized classification loss , confidence loss and bounding box regression loss , expressed as: ; in, Balance 、 and The balance factor, an exemplary The values are 0.5, 1.0, and 0.05. and confidence loss Adopting zoom loss. Traditional focal loss uniformly reduces the weight of negative samples to balance the contribution of foreground and background classes, while zoom loss specifically reduces the weight of negative samples while maintaining the emphasis on positive samples. Zoom loss is particularly beneficial in reducing the loss contribution of well-classified examples, thereby encouraging the network to pay more attention to negative samples that are difficult to classify. Bounding box regression loss The Distance-IoU loss is adopted, which not only considers the overlap between the predicted box and the real box, but also combines the distance, scale and aspect ratio similarity between the predicted box and the real box.
[0036] is the segmentation loss, which includes the classification error cross entropy loss between the pixels output by the network and the target and the IoU loss between the segmentation mask output by the network and the true mask; given that the IoU loss is particularly effective for predicting sparse categories (such as lane lines), in addition to the cross entropy loss with logits, the IoU loss is included middle.
[0037] For depth information loss, cross entropy loss is used, and the depth distance predicted by the model is calculated with the true value.
[0038] The total loss function is as follows: ; in, are all balance factors, an example The values are 0.3, 0.2 and 0.2 respectively.
[0039] use and Performing careful calibration ensures that each component of the total loss function contributes harmoniously.
[0040] The obtained detection information, segmentation information, depth information, and state information are integrated with the target detection information at time t-1 of the previous frame to obtain the detection timing information. Finally, the model integrates all the information to obtain the final result, which is more conducive to perception information and pre-decision information for subsequent tasks.
[0041] In this application, we evaluate the performance of our multi-task detection model on the BDD100K and SDExpressway datasets to demonstrate its effectiveness. To fully assess the model's generalization capabilities and mitigate the relative scarcity of the SDExpressway dataset for highway multi-task detection, we enriched the dataset with 2,000 additional images collected and annotated, including images captured at night and in adverse weather conditions such as rain and fog. Extensive experiments on the challenging BDD100K and enhanced SDExpressway datasets demonstrate that our multi-task detection model achieves state-of-the-art performance, significantly outperforming baseline models.
[0042] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, structure or unit, which can be electrical, mechanical or other forms.
[0043] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0044] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0045] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A multi-task detection method for comprehensive driving perception, characterized in that: include: Construct an end-to-end multi-task detection model that collaboratively handles traffic object detection, segmentation, and depth generation tasks; the multi-task detection model includes a backbone network, a neck network, a detection head, a segmentation head, and a depth information extraction head; the backbone network and the neck network cooperate to form an encoder, wherein the backbone network is used to extract image features, and the neck network is used to fuse features generated by the backbone network; The detection head identifies multiple targets, including detection frames and confidence levels for vehicles and pedestrians, and then generates secondary labels representing the status of other vehicles. Based on these secondary labels, the driver determines whether to yield or follow the vehicle. The segmentation head segments multiple categories of targets, including drivable areas, vehicles, pedestrians, and background areas, and based on the classification information, provides label information indicating whether the drivable area ahead is clean; The depth information extraction head generates depth information; The obtained detection information, segmentation information, depth information, and state information are integrated with the target detection information at time t-1 of the previous frame to obtain the detection timing information. Finally, the model integrates all the information to obtain the final result, which is more conducive to perception information and pre-decision information for subsequent tasks.
2. The multi-task detection method for comprehensive driving perception according to claim 1, characterized in that: The backbone network is based on CSPDarknet53 to implement multi-scale feature extraction. In order to enhance the information extraction capability, the C2F module connected to the output stage in each stage of CSPDarknet53 is optimized to a C2fA module.
3. The multi-task detection method for comprehensive driving perception according to claim 1, characterized in that: The neck network uses a simple spatial pyramid fast pooling module to process the deepest feature map from the backbone network. The simple spatial pyramid fast pooling module includes: an initial CBR module, which performs a convolution operation on the input feature map to extract features, and accelerates the training process and improves the stability of the model through batch normalization. The ReLU activation function introduces nonlinearity, enhances the expressiveness of the model, and reduces the number of channels to half of the original number of channels; then, through multiple cascaded maximum pooling operations, the results of the maximum pooling operations in each layer are spliced, and finally, a CBR module is used to convert the fused feature map into a specified number of output channels.
4. The multi-task detection method for comprehensive driving perception according to claim 3, characterized in that: During the feature fusion process, the neck network adopts a feature pyramid network structure, and the C2F module alternately arranged with the CBS module is optimized into a C2fA module to effectively combine features at different semantic levels.
5. The multi-task detection method for comprehensive driving perception according to claim 2 or 4, characterized in that: Any feature map input to the C2fA module is decomposed into two parts by the CBS module, where the CBS module contains a convolutional layer, a batch normalization layer, and a SiLU activation function; one part is processed by a CBS module and input into a convolutional layer with multiple convolution kernel sizes, and multiple convolutional layers process features in parallel; the results of multiple convolution parallel processing are fused by the CBS module, processed by the CBR module, and then spliced and combined with the other part; the CBR module contains a convolutional layer, a batch normalization layer, and a ReLU activation function layer.
6. The multi-task detection method for comprehensive driving perception according to claim 1, characterized in that: The detection head adopts the anchor-based multi-scale detection head used in YOLOP and integrates a path aggregation network and a feature pyramid network. The SimAttention mechanism is introduced before each detection head of the multi-scale detection head. The detection head uses three anchors with different aspect ratios to predict the position offset and scale change of the target for each grid cell on the multi-scale fusion feature map. The process generates the probability, confidence score and secondary label of each detection category. Among them, the three detection head information in the multi-scale detection head are obtained, with dimensions of [20, 20, 10], [40, 40, 10], and [80, 80, 10] respectively. The first two dimensions represent size information, and the second dimension represents the probability, confidence score and secondary label of each detection category: the secondary label is the binary classification information of dead or alive vehicles. The status information of whether to give way or follow the vehicle is determined based on the detection category of the secondary label.
7. The multi-task detection method for comprehensive driving perception according to claim 1, characterized in that: The segmentation head and the depth information extraction head have the same structure, including the following combination in sequence: CBS module, upsampling using nearest neighbor interpolation, C2fA module, CBS module, upsampling, CBS module, C2fA module, upsampling, CBS module.
8. The multi-task detection method for comprehensive driving perception according to claim 7, characterized in that: The depth information extraction head uses the output of a simple spatial pyramid fast pooling module to extract depth information. The depth head defines D discrete depth planes parallel to the image plane along the optical axis. Each depth plane and the image feature plane together form the viewing cone space. Assuming the lowest depth range is a meters, the maximum depth range is b meters, and the interval is c meters, a total of D = (b a) / c + 1 discrete depth planes are generated. Each pixel (h, w) corresponds to D possible 3D points in the viewing frustum, and the depth of each point is a discrete value d; The input feature map is replicated and expanded along the depth dimension into a 4-dimensional tensor as a depth feature, which represents the characteristic response of each pixel at different depth planes; the depth feature is combined with the depth encoding; For each pixel point (h, w) and depth plane d, calculate its 3D coordinates in the camera coordinate system; The model predicts the probability distribution of pixels in the depth plane; The final depth information includes: 3D coordinates, depth features combined with depth coding, and probability distribution.
9. The multi-task detection method for comprehensive driving perception according to claim 1, characterized in that: The total loss function of the multi-task detection model is as follows: ; in, For traffic object detection loss, is the segmentation loss, is the depth information loss, Both are balanced 、 and The balance factor.
10. The multi-task detection method for comprehensive driving perception according to claim 9, characterized in that: Traffic object detection loss It is defined as the combination of the following components: Penalized classification loss , confidence loss and bounding box regression loss , expressed as: ; in, Balance 、 and The balance factor, where the penalty classification loss and confidence loss Using zoom loss; bounding box regression loss Use Distance-IoU loss; Contains the classification error cross entropy loss between the pixels output by the network and the target and the IoU loss between the segmentation mask output by the network and the true mask; Using cross entropy loss, the depth distance predicted by the model is calculated with the true value to guide the learning process of the model.
Citation Information
Patent Citations
Multi-task network road target detection method for vehicle automatic driving
CN116665176A
Panoramic driving perception method based on deep learning
CN117058641A
Construction method of lightweight remote sensing target instance segmentation model based on state space model knowledge distillation
CN119274051A
Multi-task joint perception network model and detection method for traffic road surface information
US20240420487A1