Coordinated multi-scale feature enhancement network-based driving scene multi-task perception method
By coordinating multi-scale features to enhance the network CMFANet architecture, the problem of large computing overhead and insufficient generalization capabilities of multi-task networks in autonomous driving is solved, and efficient and accurate traffic scene perception is achieved to adapt to complex environments.
Patent Information
- Application Number
- CN202510690027.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing multi-task network is difficult to effectively adapt to the complex requirements of different tasks in autonomous driving, especially traffic object detection, driving area segmentation and lane line detection, which has problems such as large computing overhead and insufficient generalization capabilities.
The CMFANet architecture based on coordinated multi-scale feature enhancement network is adopted, including a shared backbone network layer, an independent task neck layer and a task head. Multi-level features are extracted through hybrid enhancement strategies, and the dynamic deformation module and spatial context fusion module are used to optimize feature representations to achieve multi-task perception.
Efficiently perform multi-task perception in resource-constrained environments, taking into account efficiency and accuracy, adapt to complex geometric deformation scenarios, meet real-time application requirements, and optimize the performance of each task.
Smart Images

Figure CN120564151A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a multi-task perception method for driving scenes based on a coordinated multi-scale feature enhancement network. Background Art
[0002] With the rapid development of deep learning, computer vision technology is increasingly being used in autonomous driving. As a crucial component of autonomous driving systems, surround-view driving perception systems are fundamental to ensuring safe and reliable vehicle operation. These systems utilize onboard sensors (such as lidar or cameras) to extract crucial environmental information, supporting vehicle control and decision-making.
[0003] Camera-based perception systems have attracted significant attention due to their low cost and ease of deployment. Camera-based perception systems process images captured by onboard cameras to provide a comprehensive understanding of the surrounding driving environment. This, in turn, provides real-time perception and decision support for autonomous driving systems. In autonomous driving, surround-view driving perception systems primarily perform three core tasks: traffic object detection, drivable area segmentation, and lane detection. Traffic object detection identifies road objects, enabling obstacle avoidance and path planning. Drivable area segmentation uses semantic segmentation to define the current driving area and alternative lanes, optimizing path selection. Lane detection locates and tracks lane markings, ensuring proper vehicle alignment. These interrelated tasks work together to establish comprehensive environmental awareness, ensuring the safety and efficiency of autonomous driving systems.
[0004] Numerous methods have been proposed for each task. For traffic object detection, methods such as the two-stage Faster R-CNN and the single-stage YOLO family have been developed. For drivable area segmentation, PSPNet and DeeplabV3+ are commonly used, while lane detection methods include SCNN, AdNet, ENet-SAD, PointLaneNet, and MFIALane. These task-specific methods have achieved promising results in their respective fields.
[0005] Due to the limited computing resources of onboard equipment, designing a separate network for each task is clearly unfeasible. Therefore, multi-task networks offer a more practical solution. These networks share a common backbone network to extract semantic information, while employing a neck network for feature fusion. The fused features are then passed to the task-specific head to complete the task. This design reduces computational overhead while significantly accelerating inference speed, meeting the requirements of real-world applications. Existing technologies, YOLOP, employ an encoder-decoder architecture, utilizing a single encoder for feature extraction and three independent decoders to handle different tasks. Building on the YOLOP design, HybridNet employs the BiFPN feature fusion mechanism to further improve model performance.
[0006] While the aforementioned methods achieve satisfactory performance, they still suffer from several limitations. For example, both YOLOP and HybridNets rely on anchor-based approaches to detect traffic objects, which not only increases inference time but also affects the generalization capabilities of the models. Furthermore, these methods typically utilize existing standardized backbone networks that are primarily optimized for single-task applications. Consequently, they struggle to effectively adapt to the diverse requirements of different tasks in multi-task systems. Consequently, developing lightweight, efficient, highly accurate, and more versatile multi-task networks has become an important research area. Summary of the Invention
[0007] To solve the above technical problems, the present invention provides a multi-task perception method for driving scenes based on a coordinated multi-scale feature enhancement network, which can simultaneously solve problems such as traffic target detection, drivable area segmentation and lane line detection, thereby achieving comprehensive perception of traffic scenes.
[0008] The technical solution adopted by the present invention is: a multi-task perception method for driving scenes based on a coordinated multi-scale feature enhancement network, the specific steps of which are as follows:
[0009] S1. Construct a multi-task network model based on the coordinated multi-scale feature enhancement network CMFANet;
[0010] The model adopts a Coordinated Multi-Scale Feature Enhancement Network (CMFANet) architecture, i.e., an encoder-decoder architecture. The model includes a shared backbone network layer, three independent task neck layers, and three independent task heads.
[0011] The encoder consists of a shared backbone network layer and three independent task neck layers, while the decoder consists of three different independent task heads. The three independent tasks are traffic object detection, drivable area segmentation, and lane detection.
[0012] S2, based on the model constructed in step S1, the shared backbone network layer adopts a hybrid enhancement strategy to extract multi-level feature representations from the input image;
[0013] S3: Based on step S2, the three independent task neck layers use feature fusion strategies to process and improve the features extracted in step S2, integrate multi-scale feature representations, and obtain fused features;
[0014] S4. Based on step S3, the three independent task heads receive the fused features from step S3 and generate the final output according to the requirements of each task to achieve comprehensive perception of the traffic scene.
[0015] Furthermore, the step S2 is specifically as follows:
[0016] The shared backbone network layer includes: 1 Conv layer, 2 DCN layers, 4 hybrid aggregation networks Manet, 2 SCDown modules, 1 SPPF module and 1 PSA module.
[0017] Among them, the DCN layer is used to process low-level information; the hybrid aggregation network Manet is the context communication module in the backbone network layer;
[0018] The DCN layer includes: 1 deformable convolution layer, 1 batch normalization layer and 1 SiLU activation function layer. The calculation process of the DCN layer is expressed as follows:
[0019] X out =Si(BN(DConv(X in ))
[0020] Among them, X in represents the input features, X out Represents the output features, DConv represents deformable convolution, BN represents batch normalization, and Si represents SiLU activation function.
[0021] The hybrid aggregation network Manet includes three branches, as follows:
[0022] 1) Conv branch, using 1×1 standard convolution to capture the correlation between different channels;
[0023] 2) Depthwise separable convolution branch to extract spatial and channel information;
[0024] First, depthwise convolution is used to calculate each input channel separately using K×K convolution, and then pointwise convolution is used to mix all channels using 1×1 convolution and adjust the number of output channels.
[0025] 3) C2f branch, using C2f module to enhance the expressiveness of features;
[0026] The main branch extracts deep features through multiple Bottleneck structures, while the diversion branch retains some input features and passes them directly to the output end; the output features of different Bottleneck stages are channel-joined with the initial diversion features through layered fusion, and the number of channels is controlled by 1×1 convolution to achieve a lightweight design.
[0027] Finally, the outputs of all three branches are merged and the channel dimension is adjusted through 1×1 standard convolution to obtain high-dimensional fusion features.
[0028] Among them, the calculation process expression of Manet is as follows:
[0029]
[0030] Among them, x1 represents the intermediate features of the features input into Manet after a conventional convolution layer, X in-Manet represents the Manet input feature, x conv Represents the intermediate features after the Conv branch, x DS represents the intermediate features after depth-wise separable convolution, x C2f represents the intermediate features after the C2f branch, X out-Manet Represents Manet output features, DS represents depth-wise separable convolution branch, DWConv represents deep-dimensional convolution, PWConv represents point-dimensional convolution, and C2f represents C2f module branch.
[0031] The calculation process expression of the C2f branch in Manet is as follows:
[0032]
[0033] Among them, x2, x3 represent the features of x1 after being divided by the number of channels, and Bottle represents the Bottle module. Represents the intermediate features after the n-th layer Bottleneck module; Split and Cat operations are performed on the channel dimension.
[0034] When processing high-level semantic information, the SCDown module is used for fast downsampling, the SPPF module integrates multi-scale information, and the PSA module enhances features. Finally, the shared backbone network outputs multi-level features. The generation process of the backbone network multi-layer feature map is expressed as follows:
[0035]
[0036] Among them, F in Represents the input image, P1 represents the output feature map after the Conv layer; P kRepresents the k-th layer output feature map, Ma represents Manet, Op k-1 Indicates the k-1th corresponding related operation, DCN indicates the DCN layer; SCD indicates the SCDown module; SPPF indicates the SPPF module; PSA indicates the PSA module, and P5 indicates the feature map output by the SPPF module and the PSA module.
[0037] Furthermore, the step S3 is specifically as follows:
[0038] The three independent task neck layers include: 1 neck detection layer and 2 neck segmentation layers.
[0039] The neck detection layer adopts a dynamic deformation enhancement module DDEM, which includes a dynamic pyramid module DPM and a deformable pyramid module DePM; DPM and DePM adopt a top-down and bottom-up architecture design respectively.
[0040] DPM implements dynamic feature upsampling through the dynamic sampling module Dysample. It first generates sampling points, then dynamically adjusts the sampling positions based on the offset learned by the network, accurately captures the detailed features of the input features, and then splices the upsampled feature map with the features of the same level. The specific calculation process of DPM is expressed as follows:
[0041]
[0042] in, represents the k-th layer output feature map of the DPM module, Dy represents the dynamic sampling module, P k Represents the feature map of the corresponding level.
[0043] DePM dynamically adjusts the position of the convolution kernel through deformable convolution, and deformable convolution adaptively adjusts the sampling position according to the geometric shape of the input features. The calculation process of the DePM module is expressed as follows:
[0044]
[0045] Among them, P′3, P′4, and P′5 represent the features input to the detection task head, and C2fCIB represents the C2fCIB module.
[0046] The segmentation neck layer adopts the dynamic spatial context fusion module DSCFM, which includes: a dynamic pyramid module DPM and a spatial context perception module SCAM;
[0047] For lane detection, one segmentation neck layer incorporates a complete dynamic pyramid module (DPM) for precise feature extraction. For drivable area segmentation, another segmentation neck layer uses nearest neighbor interpolation to upsample high-level features and introduces SCAM to optimize the representation of spatial context.
[0048] For the lane detection task, the calculation expression of DPM in the segmented neck layer is as follows:
[0049]
[0050] Where Dy represents the dynamic sampling module.
[0051] Similarly, for the drivable area segmentation task, the calculation expression of DPM in the segmented neck layer is as follows:
[0052]
[0053] Here, Up represents upsampling using the neighborhood value interpolation method. That is, when k = 3 or 4, the drivable area segmentation task uses the neighborhood value interpolation method for upsampling.
[0054] The calculation process of SCAM is expressed as follows:
[0055]
[0056] Among them, X in-SCAM Represents the SCAM input feature map, Max represents Maxpool, Avg represents Avgpool, Soft represents Softmax, x 1-SCAM 、x 2-SCAM 、x 3-SCAM 、 Represents the intermediate feature, Matrix represents matrix multiplication, Hadamard represents Hadamard product, x out-SCAM Represents the SCAM output feature map, i.e., P′1.
[0057] Furthermore, the step S4 is specifically as follows:
[0058] The three independent task heads include: 1 detection head and 2 segmentation heads, forming a multi-task head group, corresponding to traffic object detection, drivable area segmentation and lane line recognition tasks respectively.
[0059] The detection head adopts an anchor-free decoupled design and consists of three branches: a localization branch that predicts the object position and a classification branch that determines the object category and confidence score; it receives multi-scale feature maps from the detection neck layer, namely P′3, P′4, P′5, and outputs a tensor including the category prediction probability and its corresponding bounding box coordinates and confidence score.
[0060] For lane detection and drivable area segmentation, the two segmentation heads use the same structural design and independently process the corresponding high-level semantic information in the network. The segmentation head adopts a lightweight structure, receives the multi-scale feature map from the segmentation neck layer, namely P′1, and outputs the image segmentation result.
[0061] Furthermore, in step S1, the multi-task network model loss function is designed as follows:
[0062] The loss function includes: traffic object detection loss, lane detection task loss and drivable area segmentation task loss, and the expression is as follows:
[0063]
[0064] in, represents the loss function for traffic object detection, represents the loss function of the lane detection task, Represents the loss function for the drivable area segmentation task.
[0065] The loss functions used in traffic object detection tasks include: binary cross entropy loss Distributed focal loss Complete focus loss Responsible for object classification, focuses on distribution differences in bounding box regression, while The difference between the predicted bounding box and the ground truth bounding box is measured. The expression is as follows:
[0066]
[0067] Among them, λ1, λ2, and λ3 represent the corresponding coefficients.
[0068] The specific expression is as follows:
[0069]
[0070] Among them, x n Represents the predicted category of the detected object, y n Indicates the true category of the detected object.
[0071] The specific expression is as follows:
[0072]
[0073]
[0074] The variable y represents the ground truth value of the detected bounding box coordinates.i+1 and y i The values of are the upper and lower limits of y respectively.
[0075] The specific expression is as follows:
[0076]
[0077]
[0078]
[0079]
[0080] Among them, CIoU represents complete intersection over union, IoU represents intersection over union, b and b gt Denote the center point of the predicted box and the center point of the ground truth box respectively. ρ denotes the Euclidean distance between the predicted point and the ground truth point, c denotes the diagonal length of the minimum outer rectangle of the two boxes. α denotes the control factor, and v denotes the aspect ratio consistency penalty. h and w denote the height and width of the predicted box respectively. h gt and w gt Represents the height and width of the box ground truth.
[0081] In the segmentation task, the same general loss function design is used for lane detection and drivable area segmentation tasks, and the loss functions of these two tasks are and collectively referred to as Then the segmentation loss function includes focal loss and Tversky losses The specific expression is as follows:
[0082]
[0083] Among them, α1 and α2 represent the corresponding coefficients.
[0084] The specific expression is as follows:
[0085]
[0086] Among them, p t Represents the probability that the relevant model predicts the positive class. t Represents a weighting coefficient used to balance the relative importance of positive and negative training examples. The focusing parameter γ is used to adjust the weight of each sample's contribution to the loss function.
[0087] The specific expression is as follows:
[0088]
[0089] Among them, TP represents true positive samples, FP represents false positive samples, FN represents false negative samples, and α TL It represents the penalty intensity for controlling missed detection FN, and β represents the penalty intensity for controlling false detection FP.
[0090] Beneficial effects of the present invention: The method of the present invention constructs a multi-task network model based on the Coordinated Multi-Scale Feature Enhancement Network (CMFANet). First, the shared backbone network layer uses a hybrid enhancement strategy to extract multi-level feature representations from the input image. Then, the task neck layer uses a feature fusion strategy to process and refine the extracted features, integrating the multi-scale feature representations to obtain fused features. Finally, the task head receives the fused features and generates the final output according to the requirements of each task, achieving a comprehensive perception of the traffic scene. The CMFANet model constructed by the method of the present invention can simultaneously perform traffic object detection, drivable area segmentation, and lane segmentation in resource-constrained environments. The model balances efficiency and accuracy while meeting the requirements of real-time applications. The constructed efficient backbone network combined with the hybrid enhancement strategy can effectively adapt to scenes with complex geometric deformations and meet the different requirements of each task within the multi-task framework. In addition, the method of the present invention develops powerful and efficient neck layers for each task: a deformation enhancement module for the detection task and a dynamic spatial context fusion module for the segmentation task. These modules further optimize the overall performance of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 This is a flowchart of a multi-task perception method for driving scenes based on a coordinated multi-scale feature enhancement network of the present invention.
[0092] Figure 2 Schematic diagram of the multi-task network model structure based on the coordinated multi-scale feature enhancement network CMFANet in an embodiment of the present invention.
[0093] Figure 3 Schematic diagram of the shared backbone network layer structure in an embodiment of the present invention.
[0094] Figure 4 Schematic diagram of the neck layer structure detected in an embodiment of the present invention.
[0095] Figure 5 Schematic diagram of the neck layer structure segmentation in an embodiment of the present invention.
[0096] Figure 6 Schematic diagram of the effects under different weather conditions in an embodiment of the present invention.
[0097] Figure 7Schematic diagram of the comparison results between the embodiment of the present invention and the existing advanced algorithm under different weather conditions during the day.
[0098] Figure 8 Schematic diagram of the comparison results between the embodiment of the present invention and the existing advanced algorithm under different weather conditions at night. DETAILED DESCRIPTION
[0099] The method of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0100] like Figure 1 As shown in FIG, a flowchart of a multi-task perception method for driving scenes based on a coordinated multi-scale feature enhancement network of the present invention is shown, and the specific steps are as follows:
[0101] S1. Construct a multi-task network model based on the coordinated multi-scale feature enhancement network CMFANet;
[0102] In order to build a high-precision, low-parameter multi-task perception model that can accurately understand traffic driving scenes, this embodiment proposes a multi-task network model based on the coordinated multi-scale feature enhancement network CMFANet, such as Figure 2 As shown in FIG, the model adopts a Coordinated Multi-Scale Feature Enhancement Network (CMFANet) architecture, i.e., an encoder-decoder architecture. The model includes a shared backbone network layer, three independent task neck layers, and three independent task heads.
[0103] The encoder consists of a shared backbone network layer and three independent task-head layers, while the decoder includes three independent task heads. These three independent tasks are traffic object detection, drivable area segmentation, and lane detection. This design enables the model to handle all three tasks simultaneously, effectively providing a comprehensive understanding of traffic driving scenes.
[0104] The backbone network extracts multi-layered feature representations from the input image, capturing both low-level and high-level semantic information. The neck layer further processes and refines these features using a carefully designed feature fusion strategy, integrating multi-scale feature representations to enhance their expressiveness and provide richer information for downstream tasks. Finally, the task head receives the fused features and generates the final output based on the requirements of each task.
[0105] S2, based on the model constructed in step S1, the shared backbone network layer adopts a hybrid enhancement strategy to extract multi-level feature representations from the input image;
[0106] S3: Based on step S2, the three independent task neck layers use feature fusion strategies to process and improve the features extracted in step S2, integrate multi-scale feature representations, and obtain fused features;
[0107] S4. Based on step S3, the three independent task heads receive the fused features from step S3 and generate the final output according to the requirements of each task to achieve comprehensive perception of the traffic scene.
[0108] In this embodiment, step S2 is specifically as follows:
[0109] The backbone network acts as a feature extractor and can capture multi-level features from the input image. This embodiment draws inspiration from YOLOv10 and modularizes the entire backbone network into a series of different modules. These modules include Conv layer, DCN layer, Manet, SCDown, SPPF and PSA modules, and share the overall structure of the backbone network layer as shown below. Figure 3 The shared backbone network layer includes: 1 Conv layer, 2 DCN layers, 4 hybrid aggregation network Manets, 2 SCDown modules, 1 SPPF module and 1 PSA module.
[0110] The DCN layer processes low-level information, while the Manet hybrid aggregation network serves as the contextual communication module within the backbone network layer. The backbone network utilizes a modular design and consists of multiple components. When processing underlying semantic information, the network utilizes the DCN layer to centrally process target features while simultaneously utilizing the Manet layer to establish connections for contextual information, thereby enhancing the representation of feature information.
[0111] In the three tasks of traffic object detection, drivable area segmentation, and lane detection, the objects involved in each task have different geometric characteristics, which brings unique challenges to the detection work. In the traffic object detection task, the objects are often represented by slightly deformed rectangles. Although the shape is relatively regular, geometric deformation may occur due to changes in scale, orientation, and other factors. On the other hand, the drivable area segmentation task deals with irregular polygons whose boundaries are often very irregular and vary greatly, increasing the difficulty of accurate segmentation. Lane detection mainly focuses on narrow and long line segments. These objects have relatively simple shapes, but are very difficult to detect due to factors such as curvature, occlusion, and perspective changes.
[0112] Objects in each task have distinct geometric characteristics, resulting in significantly different detection difficulties and required feature representation methods. Traffic object detection requires precise boundary localization and accurate category identification. Drivable area segmentation requires effective capture of complex, irregular shapes and accurate boundary delineation. Lane detection requires the precise identification of narrow and curved line segments.
[0113] To address these challenges, the shared backbone network described in this embodiment adopts a hybrid enhancement strategy, combining the DCN layer and Manet, so that the network can adapt to complex geometric deformation environments while balancing the feature extraction requirements of different tasks.
[0114] The DCN layer includes: 1 deformable convolution layer, 1 batch normalization layer and 1 SiLU activation function layer. Unlike standard convolution, deformable convolutions dynamically adjust the sampling position by introducing an offset, thereby effectively focusing on the area of interest. In addition, deformable convolutions also allow the network to learn the weight of each sampling point, thereby distinguishing between valid and invalid positions, reducing interference, and significantly enhancing feature extraction capabilities. By combining the DCN layer, the shared backbone network described in this embodiment can effectively adapt to environments with complex geometric deformations and focus on task-related areas during the sampling process.
[0115] The calculation process of the DCN layer is expressed as follows:
[0116] X out =Si(BN(DConv(X in ))
[0117] Among them, X in represents the input features, X out Represents the output features, DConv represents deformable convolution, BN represents batch normalization, and Si represents SiLU activation function.
[0118] After adapting to environments with complex geometric deformations, the next step is to balance and enhance the extracted features to meet the needs of different tasks. Specifically, the spatial feature information in the feature map needs to be corrected to ensure that the feature map accurately reflects the geometric structure of the object. At the same time, the channel feature information must be effectively extracted to capture the details required for each task. In order to further improve the feature representation, it is necessary to strengthen the existing spatial features to accurately locate the boundaries and details of the object. In addition, since the contextual information between different tasks has significant mutual correlation, the consistency of the contextual information must be ensured during the feature extraction process to strengthen the semantic association between tasks.
[0119] This embodiment uses a hybrid aggregation network (Manet) as the context communication module in the backbone network. The hybrid aggregation network Manet includes three branches, as follows:
[0120] 1) Conv branch, using 1×1 standard convolution to capture the correlation between different channels;
[0121] 2) Depthwise separable convolutional branches that can effectively extract spatial and channel information while reducing computational overhead;
[0122] First, depthwise convolution is used to calculate each input channel separately using K×K convolution (channels are not mixed, the number of output channels = the number of input channels, and only spatial features are extracted). Then, pointwise convolution is used to mix all channels using 1×1 convolution and adjust the number of output channels.
[0123] 3) C2f branch, using C2f module to enhance the expressiveness of features;
[0124] The main branch extracts deep features through multiple Bottleneck structures (a combination of 1×1 and 3×3 convolutions). At the same time, the diversion branch retains some input features and passes them directly to the output end (similar to the dense connection of DenseNet) to ensure that shallow information is not lost; through layered fusion, the output features of different Bottleneck stages are concatenated with the initial diversion features, and 1×1 convolution is used to control the number of channels to achieve lightweight design and avoid dimensionality explosion.
[0125] Finally, the outputs of all three branches are merged and the channel dimension is adjusted through 1×1 standard convolution to obtain high-dimensional fusion features.
[0126] Among them, the calculation process expression of Manet is as follows:
[0127]
[0128] Among them, x1 represents the intermediate features of the features input into Manet after a conventional convolution layer, X in-Manet represents the Manet input feature, x conv Represents the intermediate features after the Conv branch, x DS represents the intermediate features after depth-wise separable convolution, x C2f represents the intermediate features after the C2f branch, X out-Manet Represents Manet output features, DS represents depth-wise separable convolution branch, DWConv represents deep-dimensional convolution, PWConv represents point-dimensional convolution, and C2f represents C2f module branch.
[0129] The calculation process expression of the C2f branch in Manet is as follows:
[0130]
[0131] Among them, x2, x3 represent the features of x1 after being divided by the number of channels, and Bottle represents the Bottle module. Represents the intermediate features after the n-th layer Bottleneck module; Split and Cat operations are performed on the channel dimension.
[0132] When processing high-level semantic information, the SCDown module is used for fast downsampling, the SPPF module integrates multi-scale information, and the PSA module enhances features. Finally, the shared backbone network outputs multi-level features. The generation process of the backbone network multi-layer feature map is expressed as follows:
[0133]
[0134] Among them, F in Represents the input image, P1 represents the output feature map after the Conv layer; P k Represents the k-th layer output feature map, Ma represents Manet, Op k-1 Indicates the k-1th corresponding related operation, DCN indicates the DCN layer; SCD indicates the SCDown module; SPPF indicates the SPPF module; PSA indicates the PSA module, and P5 indicates the feature map output by the SPPF module and the PSA module.
[0135] In this embodiment, step S3 is specifically as follows:
[0136] The three independent task neck layers include: 1 neck detection layer and 2 neck segmentation layers.
[0137] like Figure 4 As shown, the neck detection layer adopts a dynamic deformation enhancement module DDEM, including a dynamic pyramid module DPM and a deformable pyramid module DePM; DPM and DePM adopt a top-down and bottom-up architecture design respectively.
[0138] DPM implements dynamic feature upsampling through the dynamic sampling module Dysample. It first generates sampling points, then dynamically adjusts the sampling positions based on the offset learned by the network, accurately captures the detailed features of the input features, and then splices the upsampled feature map with the features of the same level. The specific calculation process of DPM is expressed as follows:
[0139]
[0140] in, represents the k-th layer output feature map of the DPM module, Dy represents the dynamic sampling module, P k Represents the feature map of the corresponding level.
[0141] To address the challenges posed by camera perspectives and vehicle deformation in traffic scenarios, this embodiment uses a deformable feature pyramid module (DePM). DePM dynamically adjusts the convolution kernel position through deformable convolution, and deformable convolution adaptively adjusts the sampling position based on the geometric shape of the input features. The calculation process of the DePM module is expressed as follows:
[0142]
[0143] Among them, P′3, P′4, and P′5 represent the features input to the detection task head, and C2fCIB represents the C2fCIB module.
[0144] The integration of the DPM and DePM modules has been shown to effectively facilitate the fusion of multi-scale feature maps. Furthermore, research has shown that this design enhances the expressiveness of target features, significantly improving the model's performance in traffic object detection. This network architecture enables it to more effectively adapt to complex traffic scenarios and achieve more efficient detection in environments with significant geometric deformation. Ultimately, both the accuracy and efficiency of traffic object detection are simultaneously improved.
[0145] like Figure 5 As shown, the segmented neck layer adopts a dynamic spatial context fusion module DSCFM, including: a dynamic pyramid module DPM and a spatial context perception module SCAM;
[0146] For lane detection, one segmentation neck layer incorporates a complete dynamic pyramid module (DPM) for precise feature extraction. For drivable area segmentation, another segmentation neck layer uses nearest neighbor interpolation to upsample high-level features and introduces SCAM to optimize the representation of spatial context.
[0147] For the lane detection task, the calculation expression of DPM in the segmented neck layer is as follows:
[0148]
[0149] Where Dy represents the dynamic sampling module.
[0150] Similarly, for the drivable area segmentation task, the calculation expression of DPM in the segmented neck layer is as follows:
[0151]
[0152] Here, Up represents upsampling using the neighborhood value interpolation method. That is, when k = 3 or 4, the drivable area segmentation task uses the neighborhood value interpolation method for upsampling.
[0153] The calculation process of SCAM is expressed as follows:
[0154]
[0155] Among them, X in-SCAM Represents the SCAM input feature map, Max represents Maxpool, Avg represents Avgpool, Soft represents Softmax, x 1-SCAM 、x 2-SCAM 、x3-SCAM 、 Represents the intermediate feature, Matrix represents matrix multiplication, Hadamard represents Hadamard product, x out-SCAM Represents the SCAM output feature map, i.e., P′1.
[0156] The SCAM module aims to optimize the expression of contextual information and enhance the spatial correlation of features. Based on the features that have already significantly represented the target features after DPM processing, SCAM further optimizes the spatial contextual information of these features, thereby effectively improving the performance of the segmentation task.
[0157] In this embodiment, step S4 is specifically as follows:
[0158] The three independent task heads include: 1 detection head and 2 segmentation heads, forming a multi-task head group, corresponding to traffic object detection, drivable area segmentation and lane line recognition tasks respectively.
[0159] To enhance versatility, the detection head adopts an anchor-free decoupled design, consisting of three branches: a localization branch that predicts the object position, and a classification branch that determines the object category and confidence score; it receives multi-scale feature maps from the detection neck layer, namely P′3, P′4, P′5, and outputs a tensor including the category prediction probability and its corresponding bounding box coordinates and confidence score.
[0160] The detection head adopts an innovative "decoupled head" design, which is decomposed into three independent branches after unified dimensionality reduction through 1×1 convolution: a classification branch (3×3 convolution + 1×1 convolution to predict category probability), a regression branch (3×3 convolution + 1×1 convolution to predict bbox coordinates) and an optional confidence branch. It receives multi-scale feature maps from the detection neck layer and simultaneously outputs category probability, bounding box coordinates and confidence, significantly improving detection accuracy while maintaining real-time performance.
[0161] For lane detection and drivable area segmentation, the two segmentation heads use the same structural design and independently process the corresponding high-level semantic information in the network. The segmentation head adopts a lightweight structure, receives the multi-scale feature map from the segmentation neck layer, namely P′1, and outputs the image segmentation result.
[0162] The segmentation head uses a lightweight codec structure to achieve real-time instance segmentation. Its core design is as follows: receiving multi-scale feature maps from the segmentation neck layer and using only the highest resolution feature map for mask prediction, first compressing the input channels to 32 dimensions through 3×3 convolution, then using transposed convolution (ConvTranspose2d) to scale the feature map by 2 times (instead of interpolation upsampling), and then further compressing the channels to 16 dimensions through 3×3 convolution, and finally outputting the final mask, and normalizing the mask probability with a Sigmoid function. This design achieves accurate pixel-level segmentation capabilities while maintaining efficient computation. The deconvolution operation is used to restore the output to the original input image size.
[0163] In this embodiment, the multi-task network model loss function in step S1 is designed as follows:
[0164] The loss function includes: traffic object detection loss, lane detection task loss and drivable area segmentation task loss, and the expression is as follows:
[0165]
[0166] in, represents the loss function for traffic object detection, represents the loss function of the lane detection task, Represents the loss function for the drivable area segmentation task.
[0167] The loss functions used in traffic object detection tasks include: binary cross entropy loss Distributed focal loss Complete focus loss Responsible for object classification, focuses on distribution differences in bounding box regression, while The difference between the predicted bounding box and the ground truth bounding box is measured. The expression is as follows:
[0168]
[0169] Among them, λ1, λ2, and λ3 represent the corresponding coefficients.
[0170] The specific expression is as follows:
[0171]
[0172] Among them, x n Represents the predicted category of the detected object, y n Indicates the true category of the detected object.
[0173] The specific expression is as follows:
[0174]
[0175]
[0176] The variable y represents the ground truth value of the detected bounding box coordinates. i+1 and y i The values of are the upper and lower limits of y respectively.
[0177] The specific expression is as follows:
[0178]
[0179]
[0180]
[0181]
[0182] Among them, CIoU represents complete intersection over union, IoU represents intersection over union, b and b gt Denote the center point of the predicted box and the center point of the ground truth box respectively. ρ denotes the Euclidean distance between the predicted point and the ground truth point, c denotes the diagonal length of the minimum outer rectangle of the two boxes. α denotes the control factor, and v denotes the aspect ratio consistency penalty. h and w denote the height and width of the predicted box respectively, and the symbol h gt and w gt Represents the height and width of the box ground truth.
[0183] In the segmentation task, the same general loss function design is used for lane detection and drivable area segmentation tasks, and the loss functions of these two tasks are and collectively referred to as Then the segmentation loss function includes focal loss and Tversky losses The specific expression is as follows:
[0184]
[0185] Among them, α1 and α2 represent the corresponding coefficients.
[0186] The specific expression is as follows:
[0187]
[0188] Among them, p t Represents the probability that the relevant model predicts the positive class. tRepresents a weighting coefficient used to balance the relative importance of positive and negative training examples. The focusing parameter γ is used to adjust the weight of each sample's contribution to the loss function.
[0189] The specific expression is as follows:
[0190]
[0191] Among them, TP represents true positive samples, FP represents false positive samples, FN represents false negative samples, and α TL It represents the penalty intensity for controlling missed detection (FN), and β represents the penalty intensity for controlling false detection (FP).
[0192] This embodiment is further verified by experiments, as follows:
[0193] like Figure 6 As shown in FIG, the CMFANet model of the present invention performs perception tasks in daytime, rainy day and nighttime scenes, where Figure 6 (a) is a sunny day scene. Figure 6 (b) is a rainy day scene. Figure 6 (c) is a night scene, with three perception tasks in the figure: traffic target detection (red bounding box), driving area segmentation task (green segmentation area) and lane line detection (blue line). Figure 6 It can be seen that the CMFANet model of the proposed method has achieved excellent results in daytime, rainy and nighttime scenes, and can effectively perceive complex and changing traffic environments.
[0194] By using the scaling factor to adjust the number of network layers, this embodiment constructs two versions of CMFANet: nano (n) and small (s). The scaling factor settings of different versions are shown in Table 1.
[0195] Table 1
[0196]
[0197] Among them, depth is the scaling factor of the number of repetitions of the repeating module; width is the scaling factor of the number of network channels in each layer of the model; the maximum number of channels is the maximum value of the number of network channels in each layer of the model, and the number of channels in the network layer that exceeds this value is this value.
[0198] As shown in Table 1, the lightweight version has only 3.8 million parameters, significantly fewer than existing models such as YOLOP. This example evaluated the model on the BDD100K public dataset, achieving excellent results: a mAP50 of 82.5% for traffic object detection, a mIoU of 91.5% for drivable area segmentation, and an IoU of 29.8% for lane detection. Experiments confirm that the proposed model achieves a balance of efficiency, lightweight design, and accuracy in real-world driving scenarios.
[0199] The BDD100K (Berkeley DeepDrive 100K) dataset is a large-scale autonomous driving dataset released by the Berkeley Deep Drive (BDD) team. In this example, the dataset is divided into three parts: a training set consisting of 70,000 images, a validation set consisting of 10,000 images, and a test set consisting of 20,000 images. Because the annotations of the test set are not yet publicly available, this example evaluates the CMFANet model proposed by the present invention on the validation set. Consistent with previous research, in the traffic object detection task, only vehicle targets are focused, specifically cars, buses, trucks, and trains.
[0200] During training, this embodiment uses an end-to-end training approach, integrating all task modules into a unified network framework for collective optimization. In this approach, all network parameters are trained simultaneously, without freezing specific layers or alternating optimization strategies. This allows the model to automatically adjust the weight relationships between different tasks during training, ensuring that learning from different tasks complements rather than conflicts with each other. End-to-end optimization facilitates more efficient parameter sharing and feature extraction across tasks, thereby improving resource allocation in multi-task learning. The collective optimization strategy across the entire network strengthens task coordination, reduces error propagation in intermediate stages, and further enhances the model's robustness and generalization capabilities for complex tasks. In the experiments, this embodiment selected the SGD optimizer with an initial learning rate (lr) of 0.01, a weight decay of 0.0005, and a momentum of 0.937. Furthermore, a linear learning rate annealing strategy was employed throughout the training process. This embodiment was conducted on an Ubuntu 20.04 system equipped with a 24GB NVIDIA GeForce RTX 4090 GPU.
[0201] The comparison results of this embodiment under daytime driving conditions are as follows Figure 7 As shown in the figure, in daytime scenes, the model of our method consistently outperforms YOLOP, achieving more accurate traffic object detection results with no false positives. Furthermore, lane detection and drivable area segmentation results are smoother and more accurate than those of YOLOP. False positives and missed negatives detected by other comparison methods are indicated by yellow circles or boxes.
[0202] This example conducts a comprehensive performance evaluation of the CMFANet model proposed by the method of the present invention on the BDD100k dataset and compares it with the two most advanced models YOLOP and A-YOLOM. The evaluation results are shown in Figure 2. Figure 7 To fully demonstrate the model's performance in different scenarios, we selected three representative daytime conditions for testing: rainy, overcast, and snowy. These scenarios encompass a wide range of weather conditions, changing lighting, and complex traffic environments, collectively reflecting the model's performance in real-world applications.
[0203] Depend on Figure 7 It can be seen that both YOLOP and A-YOLOM showed shortcomings in the traffic object detection task, especially in detecting smaller objects. In contrast, the model of the method of the present invention showed remarkable ability in accurately identifying all target objects (including those that are small and occluded). This finding shows that the model of the method of the present invention shows excellent robustness in complex environments and can effectively reduce false detections and missed detections. In the lane line detection task, YOLOP faces challenges in identifying certain lane lines that are difficult to detect due to occlusion or insufficient light. In contrast, the model of the method of the present invention successfully detected these lane lines, further verifying its accuracy in complex traffic scenes. In the drivable area segmentation task, the model of the method of the present invention performed well, with no omissions or false detections, and the segmentation results were smoother and more natural than those of YOLOP and A-YOLOM. This result shows that the model of the method of the present invention has higher accuracy and consistency in fine-grained area segmentation.
[0204] Figure 8 This figure shows the comparison results of this example under nighttime driving conditions. In this nighttime scenario, the model of our method still achieves significantly better results than YOLOP. YOLOP struggles to detect lane markings, while our model successfully identifies them. Furthermore, in the drivable area segmentation task, our model produces smoother segmentation results with no missed or false positives. False positives and missed positives from the other comparison methods are indicated by yellow circles and boxes.
[0205] Depend on Figure 8It can be seen that this embodiment selected three challenging nighttime environments for evaluation: low-light environment, glare environment, and ground reflection environment. In these low-light and complex environments, the model of the method of the present invention consistently outperforms YOLOP and A-YOLOM. Under low-light conditions, the model of the method of the present invention performs well in detecting small vehicles in the distance, which is a challenging task for YOLOP and A-YOLOM. In addition, in the lane line detection task, YOLOP and A-YOLOM have difficulty identifying clear lane lines under such conditions, while the model of the method of the present invention can successfully identify these details, further demonstrating its robustness in nighttime environments. In the drivable area segmentation task, the model demonstrated high performance, avoiding false detections and missed detections even in low-light conditions and complex backgrounds. The resulting segmentation results are significantly more refined and accurate. This performance demonstrates that the model proposed by the method of the present invention can reliably perform tasks and provide reliable decision support.
[0206] This example also quantitatively analyzes the proposed model on the BDD100k dataset. By comparing multiple performance metrics, the model's performance in tasks such as traffic object detection, lane detection, and drivable area segmentation is evaluated, as follows:
[0207] (1) Parameters and inference speed: The comparison of the model of the present invention method with the existing model in terms of parameter size and inference speed is shown in Table 2. Table 2 shows the comparison of the model of the present invention method with the existing model in terms of parameter size and inference speed. The nano version model of the present invention method has only 3.8M parameters, which is significantly smaller than the existing model. Its processing speed (FPS) is 131.8, which is much higher than other models. This gives the model of the present invention method a significant advantage in efficiency and speed, especially in resource-limited environments, where it can provide faster processing performance.
[0208] Table 2
[0209]
[0210] (2) Traffic object detection: The experimental results of traffic object detection are shown in Table 3. Similar to previous studies, this embodiment uses Recall and mAP50 as evaluation indicators, with the confidence threshold set to 0.001 and the NMS threshold set to 0.6. It can be seen that both versions of the model perform well in this task, especially in terms of detection accuracy. In particular, the mAP50 of the small version reaches 82.5%, which is significantly higher than the scores of existing comparison models. This shows that the model has excellent detection performance. Although the nano version has been further optimized in terms of parameter size, it still achieved an excellent mAP50 of 80.4%, surpassing other comparison models. This result shows that although the nano version adopts a lightweight design, its detection performance is not significantly affected. On the contrary, it optimizes the real-time processing capability to better meet the real-time requirements of practical applications. These results not only demonstrate the super detection capability of the model of the method of the present invention in the traffic object detection task, but also highlight its strong adaptability and flexibility.
[0211] Table 3
[0212]
[0213] (3) Drivable area segmentation: The experimental results of drivable area segmentation are shown in Table 4. In order to evaluate the performance of the model, this embodiment uses mIoU as the evaluation indicator. Although the loss function specially designed for drivable area segmentation is not used in this task, the small version model of this embodiment still achieves the same mIoU score as the SOTA model YOLOP, that is, 91.5%. This result shows that even without a targeted loss function, the model of the method of the present invention can maintain excellent performance in this task. Compared with YOLOP, the model of the method of the present invention not only performs well in accuracy, but also has significant advantages in versatility and inference speed. These advantages make the model of the method of the present invention more competitive in practical applications.
[0214] Table 4
[0215]
[0216] (4) Lane line detection: The lane line detection results are shown in Table 5. In this task, this embodiment selected pixel accuracy and IoU as evaluation indicators. The experimental results show that the model of the method of the present invention performed well in lane line detection and obtained the highest score, with a pixel accuracy of 86.6% and an IoU of 29.8%. Both indicators are significantly better than all the comparison models. These results confirm the excellent performance of the model of the method of the present invention in this task. Although this embodiment performs lightweight optimization aimed at improving inference speed and computational efficiency, the accuracy of the model of the method of the present invention in lane line detection remains at a high level, which shows that the model of the method of the present invention has strong detection capabilities in complex and challenging environments.
[0217] Table 5
[0218]
[0219] (5) Full-time multi-task perception: The full-time multi-task perception results of CMFANet(n) are shown in Table 6. The BDD100k dataset provides label information for the time period in which each image was taken. The dataset is mainly divided into four categories: daytime, nighttime, dawn / dusk, and undefined. In order to comprehensively evaluate the performance of the model, this embodiment evaluates all available time period related data. The results show that the model performs best in the daytime category. In the nighttime category, due to the presence of low light conditions, the indicators related to traffic target detection have declined. However, the performance indicators related to lane segmentation and lane detection remain at a high level, thus meeting the driving safety requirements.
[0220] Table 6
[0221]
[0222] (6) All-weather multi-task perception: The all-weather multi-task perception results of CMFANet(n) are shown in Table 7. The meteorological conditions of the data in the BDD100k dataset can be divided into seven main categories: sunny, cloudy, snowy, cloudy, rainy, foggy and undefined weather. The performance of the proposed CMFANet(n) was comprehensively evaluated under different weather conditions. The model can produce accurate results in all weather conditions. It can meet the needs of safe driving in adverse weather conditions such as snowy, cloudy and rainy days. However, it is worth noting that the traffic target detection indicators of the model are not ideal under clear weather conditions, and are only slightly improved compared to foggy weather conditions. This is because the clear weather dataset contains a large amount of night driving scene data, and under the influence of weak light, the indicators under this weather condition are lower than expected.
[0223] Table 7
[0224]
[0225] To verify the effectiveness of the modules designed by the present method, this example conducted an ablation experiment based on the proposed CMFANet-nano version. The experimental results are shown in Table 8. By systematically removing and replacing individual modules, the contribution of each module to the overall performance can be clearly observed.
[0226] Table 8
[0227]
[0228] Table 8 shows that when using only the backbone network designed using the proposed method, performance in traffic object detection and drivable area segmentation tasks is significantly improved. However, the IoU for lane detection shows a significant decrease. This phenomenon can be attributed to the introduction of deformable convolutions in the backbone network to better process low-level feature information. Deformable convolutions concentrate the convolution kernel sampling points more on salient features, such as traffic objects and large drivable areas. However, for slender lanes with large spans, this operation fails to fully capture relevant features, negatively impacting lane detection performance. The introduction of the DDEM module significantly improves traffic object detection performance. However, because this module's design focuses more on optimizing the detection task, lane detection performance slightly decreases. This imbalance is due to the module's focus on improving detection accuracy, which inadvertently affects feature extraction and allocation for the segmentation task, resulting in a performance tradeoff. Similarly, when using only the DSCFM, performance on the segmentation task significantly improves, while performance on the detection task decreases. This performance shift can be attributed to the shift in focus of the model after the addition of the proposed module. Finally, after the complete design was implemented, all three tasks achieved excellent performance. Notably, the performance of the traffic object detection task improved compared to using only the DDEM module. This demonstrates that the end-to-end collaborative training approach of the present invention effectively promotes the synchronization of various tasks within the network, resulting in more balanced performance across tasks. These results demonstrate that the multi-task training strategy of this embodiment promotes efficient resource sharing and performance complementarity between tasks, significantly improving the overall performance of the model.
[0229] In summary, this paper proposes a lightweight, high-precision multi-task network model, CMFANet, specifically designed for traffic scene perception. This model can simultaneously perform three tasks: traffic object detection, drivable area segmentation, and lane recognition, thereby achieving comprehensive perception of the driving environment. Our method utilizes an end-to-end training approach, rather than relying on freezing specific layers for single-task training. Instead, by employing a multi-task parallel training strategy, tasks are coordinated, thereby improving overall performance. To this end, our method designs a modular backbone network with powerful feature extraction capabilities to effectively support the diverse requirements of each task. Furthermore, task-specific neck layers are developed for each task. Through careful task allocation and feature fusion, the performance and accuracy of each task are effectively improved. Finally, our model is qualitatively and quantitatively analyzed on the BDD100k dataset and compared with state-of-the-art models, verifying its superiority. Furthermore, this example also conducts ablation experiments to further demonstrate the effectiveness of the proposed module.
[0230] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. A multi-task driving scene perception method based on a coordinated multi-scale feature enhancement network, with the following specific steps: S1. Construct a multi-task network model based on the coordinated multi-scale feature enhancement network CMFANet; The model adopts a coordinated multi-scale feature enhancement network-CMFANet architecture, that is, the model adopts an encoder-decoder architecture; the model includes: A shared backbone network layer, three independent task neck layers, and three independent task heads; The encoder consists of a shared backbone network layer and three independent task neck layers, and the decoder consists of three different independent task heads; the three independent tasks are traffic object detection, drivable area segmentation, and lane line detection tasks; S2, based on the model constructed in step S1, the shared backbone network layer adopts a hybrid enhancement strategy to extract multi-level feature representations from the input image; S3: Based on step S2, the three independent task neck layers use feature fusion strategies to process and improve the features extracted in step S2, integrate multi-scale feature representations, and obtain fused features; S4. Based on step S3, the three independent task heads receive the fused features from step S3 and generate the final output according to the requirements of each task to achieve comprehensive perception of the traffic scene.
2. The driving scene multi-task perception method based on a coordinated multi-scale feature enhancement network according to claim 1 is characterized in that: The step S2 is specifically as follows: The shared backbone network layer includes: 1 Conv layer, 2 DCN layers, 4 hybrid aggregation network Manets, 2 SCDown modules, 1 SPPF module and 1 PSA module; Among them, the DCN layer is used to process low-level information; the hybrid aggregation network Manet is the context communication module in the backbone network layer; The DCN layer includes: 1 deformable convolution layer, 1 batch normalization layer and 1 SiLU activation function layer; the calculation process of the DCN layer is expressed as follows: X out =Si(BN(DConv(X in )) Among them, X in represents the input features, X out Represents output features, DConv represents deformable convolution, BN represents batch normalization, and Si represents SiLU activation function; The hybrid aggregation network Manet includes three branches, as follows: 1) Conv branch, using 1×1 standard convolution to capture the correlation between different channels; 2) Depthwise separable convolution branch to extract spatial and channel information; First, use depthwise convolution to calculate each input channel separately with K×K convolution, and then use pointwise convolution of 1×1 convolution to mix all channels and adjust the number of output channels; 3) C2f branch, using C2f module to enhance the expressiveness of features; The main branch extracts deep features through multiple Bottleneck structures, while the branch retains some input features and directly transmits them to the output. Through layered fusion, the output features of different Bottleneck stages are combined with the initial branch features for channel splicing, and 1×1 convolution is used to control the number of channels to achieve a lightweight design. Finally, the outputs of all three branches are merged and the channel dimension is adjusted through 1×1 standard convolution to obtain high-dimensional fusion features; Among them, the calculation process expression of Manet is as follows: Among them, x1 represents the intermediate features of the features input into Manet after a conventional convolution layer, X in-Manet represents the Manet input feature, x conv Represents the intermediate features after the Conv branch, x DS represents the intermediate features after depth-wise separable convolution, x C2f represents the intermediate features after the C2f branch, X out-Manet Represents Manet output features, DS represents depth-wise separable convolution branch, DWConv represents deep-dimensional convolution, PWConv represents point-dimensional convolution, and C2f represents C2f module branch; The calculation process expression of the C2f branch in Manet is as follows: Among them, x2, x3 represent the features of x1 after being divided by the number of channels, and Bottle represents the Bottle module. Represents the intermediate features after the n-th layer Bottleneck module; Split and Cat operations are performed on the channel dimension; When processing high-level semantic information, the SCDown module is used to achieve fast downsampling, the SPPF module integrates multi-scale information, and the PSA module enhances features. Finally, the shared backbone network outputs multi-level features. The generation process of the backbone network multi-layer feature map is expressed as follows: Among them, F in Represents the input image, P1 represents the output feature map after the Conv layer; P k Represents the k-th layer output feature map, Ma represents Manet, Op k-1 Indicates the k-1th corresponding related operation, DCN indicates the DCN layer; SCD indicates the SCDown module; SPPF indicates the SPPF module; PSA indicates the PSA module, and P5 indicates the feature map output by the SPPF module and the PSA module.
3. The driving scene multi-task perception method based on a coordinated multi-scale feature enhancement network according to claim 1, characterized in that: The step S3 is specifically as follows: The three independent task neck layers include: 1 detection neck layer, 2 segmentation neck layers; The neck detection layer adopts the dynamic deformation enhancement module DDEM, which includes: dynamic pyramid module DPM and deformable pyramid module DePM; DPM and DePM adopt top-down and bottom-up architecture design respectively; DPM implements dynamic feature upsampling through the dynamic sampling module Dysample. It first generates sampling points, then dynamically adjusts the sampling positions based on the offset learned by the network, accurately captures the detailed features of the input features, and then splices the upsampled feature map with the features of the same level. The specific calculation process of DPM is expressed as follows: in, represents the k-th layer output feature map of the DPM module, Dy represents the dynamic sampling module, P k Represents the feature map of the corresponding level; DePM dynamically adjusts the position of the convolution kernel through deformable convolution, and deformable convolution adaptively adjusts the sampling position according to the geometric shape of the input features. The calculation process of the DePM module is expressed as follows: Among them, P3′, P4′, and P5′ represent the features input to the detection task head, and C2fCIB represents the C2fCIB module; The segmentation neck layer adopts the dynamic spatial context fusion module DSCFM, which includes: a dynamic pyramid module DPM and a spatial context perception module SCAM; For lane detection, one segmentation neck layer incorporates a complete dynamic pyramid module (DPM) for precise feature extraction. For drivable area segmentation, another segmentation neck layer uses nearest neighbor interpolation to upsample high-level features and introduces SCAM to optimize the representation of spatial context information. For the lane detection task, the calculation expression of DPM in the segmented neck layer is as follows: Where Dy represents the dynamic sampling module; Similarly, for the drivable area segmentation task, the calculation expression of DPM in the segmented neck layer is as follows: Where Up represents upsampling by the neighborhood value interpolation method, that is, when k = 3 or 4, the drivable area segmentation task adopts the neighborhood value interpolation method for upsampling; The calculation process of SCAM is expressed as follows: Among them, X in-SCAM Represents the SCAM input feature map, Max represents Maxpool, Avg represents Avgpool, Soft represents Softmax, x 1-SCAM 、x 2-SCAM 、x 3-SCAM 、 Represents the intermediate feature, Matrix represents matrix multiplication, Hadamard represents Hadamard product, x out-SCAM Represents the SCAM output feature map, namely P1′.
4. The driving scene multi-task perception method based on a coordinated multi-scale feature enhancement network according to claim 1, characterized in that: The step S4 is specifically as follows: The three independent task heads include: one detection head and two segmentation heads, forming a multi-task head group, corresponding to traffic object detection, drivable area segmentation and lane line recognition tasks respectively; The detection head adopts an anchor-free decoupled design and consists of three branches: a localization branch that predicts the object position and a classification branch that determines the object category and confidence score. It receives multi-scale feature maps from the detection neck layer, namely P3′, P4′, and P5′, and outputs a tensor including the category prediction probability and its corresponding bounding box coordinates and confidence score. For lane detection and drivable area segmentation, the two segmentation heads use the same structural design and independently process the corresponding high-level semantic information in the network. The segmentation head uses a lightweight structural design, receives the multi-scale feature map from the segmentation neck layer, namely P1′, and outputs the image segmentation result.
5. The driving scene multi-task perception method based on a coordinated multi-scale feature enhancement network according to claim 1, characterized in that: In step S1, the multi-task network model loss function is designed as follows: The loss function includes: traffic object detection loss, lane detection task loss and drivable area segmentation task loss, and the expression is as follows: in, represents the loss function for traffic object detection, represents the loss function of the lane detection task, Represents the loss function for the drivable area segmentation task; The loss functions used in traffic object detection tasks include: binary cross entropy loss Distributed focal loss Complete focus loss Responsible for object classification, focuses on distribution differences in bounding box regression, while The difference between the predicted bounding box and the ground truth bounding box is measured; the expression is as follows: Among them, λ1, λ2, λ3 represent the corresponding coefficients; The specific expression is as follows: Among them, x n Represents the predicted category of the detected object, y n Indicates the true category of the detected object; The specific expression is as follows: Among them, the variable y represents the ground truth value of the detected bounding box coordinates; i+1 and y i The values of are the upper and lower limits of y respectively; The specific expression is as follows: Among them, CIoU represents complete intersection over union, IoU represents intersection over union, b and b gt Denote the center point of the predicted box and the center point of the ground truth box respectively; ρ denotes the Euclidean distance between the predicted point and the ground truth point, c denotes the diagonal length of the minimum outer rectangle of the two boxes; α denotes the control factor, v denotes the aspect ratio consistency penalty; h and w denote the height and width of the predicted box respectively, h gt and w gt represents the height and width of the box ground truth; In the segmentation task, the same general loss function design is used for lane detection and drivable area segmentation tasks, and the loss functions of these two tasks are and collectively referred to as Then the segmentation loss function includes focal loss and Tversky losses The specific expression is as follows: Among them, α1, α2 represent the corresponding coefficients; The specific expression is as follows: Among them, p t Indicates the probability of the relevant model predicting the positive class; α t represents a weighting coefficient used to balance the relative importance of positive and negative training examples; the focusing parameter γ is used to adjust the weight of each sample's contribution to the loss function; The specific expression is as follows: Among them, TP represents true positive samples, FP represents false positive samples, FN represents false negative samples, and α TL It represents the penalty intensity for controlling missed detection FN, and β represents the penalty intensity for controlling false detection FP.
Citation Information
Patent Citations
Multi-task panoramic driving perception method and system based on improved YOLOv5
CN115223130A
Lightweight automatic driving target detection method based on YOLOV5
CN118918557A
Remote sensing target detection method, device, equipment and medium
CN119295965A
Target detection method for shielded vehicles in urban road scene
CN119810418A
Multi-task joint perception network model and detection method for traffic road surface information
US20240420487A1
Cited By
Photovoltaic module EL image intelligent defect identification method and system based on deep learning, storage medium and electronic terminal
CN121147231A
A Deep Learning-Based Intelligent Defect Recognition Method and System for Photovoltaic Module EL Images, Storage Medium, and Electronic Terminal
CN121147231B