Complex scene pavement marker line detection method based on multi-module adaptive fusion
By improving the YOLOv10 algorithm and adopting multi-module adaptive fusion technology, the accuracy and robustness issues of road marking line detection in complex environments are solved, and high-precision and high-robustness road marking line detection is achieved.
Patent Information
- Application Number
- CN202511009452.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-28
AI Technical Summary
Existing road marking detection technologies suffer from reduced performance in complex environments, especially in adverse weather conditions such as low light, rain, and fog, where accuracy decreases, robustness is poor, and it is difficult to maintain stable detection performance.
By improving the YOLOv10 algorithm, the ADown adaptive downsampling module, the SPPF_LSKA spatial pyramid pooling enhancement module, the MAFPN multi-scale adaptive feature pyramid network, the ASF adaptive spatial fusion module and the ScalSeq scale serialization module are adopted to achieve multi-module adaptive fusion and enhance the feature preservation, long-distance dependency modeling, cross-scale feature fusion and spatial dimension feature fusion capabilities.
It significantly improves the detection accuracy and robustness in complex scenes, improves the accuracy and recall rate of sign line detection, and maintains stable detection performance, especially in harsh environments such as low light, rain and fog.
Smart Images

Figure CN120853128A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of road traffic marking detection technology, specifically to a method for detecting road markings in complex scenarios based on multi-module adaptive fusion. Background Technology
[0002] Road marking detection is a key technology in autonomous driving and driver assistance systems, widely used in lane keeping, lane changing, traffic monitoring, and mobile robot navigation. In recent years, with the rapid development of deep learning technology, object detection methods based on convolutional neural networks have made significant progress, especially real-time object detection algorithms such as YOLO, which perform exceptionally well in road marking detection.
[0003] Existing technologies, in pursuit of higher accuracy in lane marking detection, typically improve performance by increasing network depth and employing more complex feature fusion networks. However, these methods still have significant limitations in complex environments. Under adverse weather conditions such as low light, rain, and fog, traditional detection algorithms are prone to decreased accuracy and poor robustness. Furthermore, existing algorithms struggle to maintain stable detection performance in complex scenarios involving blurred, discontinuous, or worn lane markings, lane occlusion, multiple lanes, and curves. For example, while the existing YOLOv10 algorithm performs well under standard conditions, its detection accuracy and recall need improvement in complex scenarios, particularly in multi-scale feature fusion and spatial awareness capabilities.
[0004] In summary, existing road marking detection technologies still have considerable room for improvement in terms of adaptability to complex environments, multi-scale feature fusion capabilities, and spatial perception accuracy, making it difficult to meet the demands for high-precision and robust detection in practical applications. Summary of the Invention
[0005] To address the technical problem of decreased detection performance in harsh environments in existing technologies, this technical solution provides a road marking detection method for complex scenes based on multi-module adaptive fusion. By improving five core modules of the YOLOv10 algorithm, a significant improvement in detection performance is achieved, enabling high-precision and robust multi-module adaptive fusion road marking detection in complex scenes; effectively solving the aforementioned problems.
[0006] This invention is achieved through the following technical solution:
[0007] A method for detecting road markings in complex scenes based on multi-module adaptive fusion includes the following steps:
[0008] Step 1: Obtain the road marking dataset;
[0009] Step 2: Construct a road marking detection model based on YOLOv10, including:
[0010] Step 2.1: In the YOLOv10 backbone network, replace the SCDown downsampling module with the ADown downsampling module, and add the SPPF_LSKA spatial pyramid pooling enhancement module after the PSA module; the ADown module has better feature preservation ability and reduces information loss during downsampling; the SPPF_LSKA module enhances the ability to model long-distance dependencies of the marker lines through the large kernel selection attention mechanism.
[0011] Step 2.2: In the neck network of YOLOv10, it is reconstructed into a MAFPN multi-scale adaptive feature pyramid network, which integrates the ASF adaptive spatial fusion module and the ScalSeq scale serialization module; MAFPN realizes the adaptive fusion of multi-scale features, the ASF module realizes the adaptive splicing of multi-scale features through Zoom_cat, and the ScalSeq module realizes the multi-scale feature serialization processing.
[0012] Step 2.3: In the YOLOv10 detection head section, maintain the original v10Detect detection head structure;
[0013] Step 3: Train the road marking detection model using the dataset;
[0014] Step 4: Use the trained road marking detection model to detect road markings in complex scenes.
[0015] Furthermore, the dataset mentioned in step one uses the CRLDD complex road marking detection dataset.
[0016] Furthermore, the backbone network described in step two includes a first Conv module, a second Conv module, a first C2f module, a first ADown module, a second C2f module, a second ADown module, a third C2f module, a third ADown module, a fourth C2f module, an SPPF-LSKA module, and a PSA self-attention module;
[0017] The neck network includes an upsampling path, a downsampling path, a MAFPN feature fusion path, an ASF adaptive spatial fusion module, a ScalSeq scale serialization module, and an Add residual connection module.
[0018] The detection head portion includes the v10Detect three-scale detection head.
[0019] Furthermore, the data flow of the road marking detection model is as follows:
[0020] Backbone network data flow:
[0021] The original image is used as the input to the first Conv module. The output feature map of the first Conv module is used as the input to the second Conv module. The output feature map of the second Conv module is used as the input to the first C2f module. The output feature map of the first C2f module is used as the input to the first ADown module. The output feature map of the first ADown module is used as the input to the second C2f module. The output feature map of the second C2f module is used as the input to the second ADown module. The output feature map of the second ADown module is used as the input to the third C2f module. The output feature map of the third C2f module is used as the input to the third ADown module. The output feature map of the third ADown module is used as the input to the fourth C2f module. The output feature map of the fourth C2f module is processed sequentially by the SPPF-LSKA module and the PSA self-attention module.
[0022] Neck network MAFPN feature fusion:
[0023] (1) First stage fusion: The output of the third C2f module is processed by Conv and then concatenated with the output of the PSA module. Then, it is processed by the C2fCIB module to obtain the P5 feature.
[0024] (2) Second stage fusion: P5 features are upsampled by Upsample, and the output of the second C2f module is processed by Conv. The two are concatenated with the output of the third C2f module in three ways by Concat, and then processed by the C2f module to obtain P4 features;
[0025] (3) Third stage fusion: P4 features are upsampled by Upsample, the output of the first C2f module is processed by Conv, the two are concatenated with the output of the second C2f module in three ways by Concat, and then processed by the C2f module to obtain the preliminary P3 features;
[0026] (4) ASF adaptive spatial fusion: The output of the second C2f module is processed by Conv to obtain feature_p3, and the output of the first C2f module is processed by Conv to obtain feature_backbone. The three features of feature_p3, the original output of the second C2f module, and feature_backbone are adaptively concatenated by the Zoom_cat module, and then processed by the C2f module to obtain the final P3 feature.
[0027] (5) Downsampling return path: P3 to P4 return: P3 features are downsampled by Conv, concatenated with P4 features from the second stage, and processed by the C2f module to obtain the final P4 features;
[0028] (6) P4 to P5 return: The P4 feature is downsampled by ADown, concatenated with the P5 feature in the first stage, and processed by the C2fCIB module to obtain the final P5 feature;
[0029] (7) ScalSeq scaling and residual connection: Multi-scale serialization: The features of the second C2f module, the third C2f module and the fourth C2f module are scaled and serialized through the ScalSeq module to output 256-dimensional features;
[0030] (8) Residual connection: The output of the ScalSeq module and the P3 features after ASF processing are residually connected through the Add module to obtain enhanced P3 features;
[0031] (9) Final detection output: Three-scale detection: The enhanced P3 feature, the final P4 feature, and the final P5 feature are respectively input into the v10Detect detection head to detect small, medium, and large targets.
[0032] Furthermore, the ADown downsampling module uses an adaptive weight allocation mechanism to dynamically adjust the downsampling strategy based on the spatial distribution of the input features, which can better preserve the detailed features of the marker lines compared to the traditional SCDown module.
[0033] Furthermore, the SPPF_LSKA spatial pyramid pooling enhancement module combines spatial pyramid pooling with a large kernel selective attention mechanism. It enhances the modeling ability of long-distance dependencies of marker lines through large kernel convolution and adaptively focuses on important spatial feature regions through selective attention mechanism.
[0034] Furthermore, the MAFPN multi-scale adaptive feature pyramid network achieves effective fusion of features at different scales through adaptive weight allocation, thereby enhancing cross-scale feature interaction capabilities.
[0035] Furthermore, the ASF adaptive spatial fusion module achieves adaptive stitching and scaling of multi-scale features through the Zoom_cat module, and enhances the perception of spatial location information of the marker line by combining spatial attention mechanism.
[0036] Furthermore, the ScalSeq scale serialization module performs serialization processing on features at multiple scales and achieves feature enhancement through the Add operation.
[0037] Beneficial effects
[0038] The present invention proposes a method for detecting road markings in complex scenes based on multi-module adaptive fusion, which has the following advantages compared with existing technologies:
[0039] (1) This invention replaces the SCDown downsampling module in YOLOv10 with the ADown adaptive downsampling module, achieving a fundamental improvement in feature preservation capability. The ADown adaptive downsampling module adopts a dual-branch architecture. The first branch performs standard downsampling to ensure the stability of basic functions, while the second branch analyzes the spatial distribution characteristics of input features through global average pooling depth, dynamically generates an adaptive weight matrix, and uses a spatial attention fusion layer to intelligently calculate the optimal fusion strategy between the two branches. This adaptive weight allocation mechanism fundamentally solves the technical defect that traditional fixed downsampling strategies cannot differentiate for different feature distributions, enabling precise protection of the edge contour and geometric shape information of the marking lines, greatly improving the ability to preserve detailed features in complex road scenarios; and effectively solving the problem of feature information loss during downsampling.
[0040] (2) This invention addresses the problem of insufficient long-distance dependency modeling caused by the limited receptive field of existing attention mechanisms by designing the SPPF-LSKA large-kernel spatial attention enhancement module, achieving a significant breakthrough in marker line continuity detection. This module employs a dual-branch parallel processing architecture: the spatial pyramid pooling branch systematically captures the complete feature hierarchy from local details to global shape using 5×5, 9×9, and 13×13 multi-scale pooling kernels; the large-kernel spatial attention branch introduces a 7×7 large-kernel convolution to significantly expand the receptive field coverage, and combines global average pooling and global max pooling dual pooling strategies to achieve a comprehensive evaluation of spatial location importance. The two branches are deeply fused through a residual connection mechanism, injecting enhanced long-distance dependency information while maintaining the integrity of the original features, thus solving the problem of continuity detection for complex targets.
[0041] (3) This invention addresses the problem that the fixed fusion strategy of traditional feature pyramid networks cannot adapt to the needs of changing scenarios. It achieves an intelligent upgrade of cross-scale feature fusion by reconstructing and designing the MAFPN multi-scale adaptive feature pyramid network. This network innovatively introduces a dynamic feature importance evaluation mechanism. It extracts global statistical information of features at each scale through global average pooling, learns nonlinear importance mapping relationships using fully connected layers, and dynamically generates normalized weight coefficients to ensure optimal balance in the fusion process. This adaptive weight allocation strategy completely changes the traditional FPN's coarse-grained processing mode that treats all scale features equally. It achieves refined fusion by intelligently adjusting the contribution of features at each scale according to the characteristics of the actual scene, ensuring that both distant small target markers and close-range large target markers obtain the most suitable feature representations, significantly improving the balance and accuracy of multi-scale target detection.
[0042] (4) This invention addresses the problem of false positives and false negatives caused by the lack of specificity in spatial feature fusion. It achieves precise optimization of multi-scale feature spatial alignment through an innovative ASF adaptive spatial fusion module design. This module first performs spatial attention enhancement processing on features at different scales using the Zoom_cat operation, intelligently identifying and highlighting key spatial regions in each scale feature. Then, it achieves perfect uniformity of spatial dimensions through precise upsampling operations, and finally performs adaptive stitching fusion in the channel dimension. This design fully leverages the detail sensitivity of small-scale features and the global perception of large-scale features. Through the precise guidance of the spatial attention mechanism, the complementary spatial information of features at different scales is maximized, fundamentally solving the technical limitations of traditional simple stitching operations that cannot fully exploit feature potential.
[0043] (5) This invention addresses the critical issue of insufficient robustness in detection under complex environments by introducing the ScalSeq scale serialization module, achieving a revolutionary transformation in multi-scale feature processing. This module breaks through the traditional parallel processing mode by converting it into a serialization modeling mode. Through three core steps—scale importance assessment, serialization fusion, and residual connection—it systematically integrates small-scale local details, medium-scale line segment connections, and large-scale overall shape information according to semantic progression. This serialization modeling strategy deeply explores the intrinsic correlation between multi-scale features and achieves a significant improvement in feature representation capabilities through a progressive feature enhancement mechanism, enabling the model to maintain stable detection performance even in harsh environments such as low light and rainy / foggy weather.
[0044] (6) This invention generates a "multi-level adaptive optimization" effect through deep collaboration of the ADown, SPPF-LSKA, MAFPN, ASF, and ScalSeq modules. This not only achieves a dual breakthrough in feature extraction accuracy and environmental adaptability, but more importantly, it forms a complex scene line detection system with self-regulation and intelligent perception capabilities. Through the organic cooperation and information complementarity between modules, this system exhibits collaborative detection performance in complex traffic environments that far exceeds the independent contributions of each module, providing a highly reliable technical guarantee for autonomous driving and intelligent transportation systems. Attached Figure Description
[0045] Figure 1 This is a schematic diagram of the overall process of the present invention.
[0046] Figure 2 This is an overall architecture diagram of the road marking detection model in this embodiment of the invention;
[0047] Figure 3 This is a schematic diagram of the ADown adaptive downsampling module in an embodiment of the present invention;
[0048] Figure 4This is a schematic diagram of the SPPF_LSKA module in an embodiment of the present invention;
[0049] Figure 5 This is a structural diagram of the MAFPN multi-scale adaptive feature pyramid network in an embodiment of the present invention;
[0050] Figure 6 This is a structural diagram of the ASF adaptive spatial fusion module in an embodiment of the present invention;
[0051] Figure 7 This is a schematic diagram illustrating the specific implementation principle of the Zoom_cat module in this embodiment of the invention.
[0052] Figure 8 This is a structural diagram of the ScalSeq scale serialization module in an embodiment of the present invention;
[0053] Figure 9 This is a graph showing the training process of the road marking detection model in an embodiment of the present invention.
[0054] Figure 10 This is a diagram showing the detection results on the dataset in an embodiment of the present invention. Detailed Implementation
[0055] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. The described embodiments are merely some embodiments of the present invention, and not all embodiments. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the design concept of the present invention should fall within the protection scope of the present invention.
[0056] Example 1:
[0057] A method for detecting road markings in complex scenes based on multi-module adaptive fusion includes the following steps:
[0058] Step 1: Obtain the road marking dataset;
[0059] The CRLDD complex road marking detection dataset was used, including marking categories such as BL (double yellow line), CL (center line), DM (dashed line), JB (intersection marking), LA (lane line), PC (pedestrian crossing), RA (right turn arrow), SA (straight arrow), SL (stop line), SLA (straight left arrow), and SRA (straight right arrow). The data covers different lighting conditions, weather conditions, and road scenarios, and was used for model training and validation.
[0060] Step 2: Construct a road marking detection model;
[0061] An improved road marking detection model was built based on YOLOv10, such as... Figure 2 As shown, in the YOLOv10 backbone network, the SCDown module is replaced with the ADown adaptive downsampling module, and an SPPF-LSKA large-kernel spatial attention enhancement module is added after the PSA module; in the YOLOv10 neck network, it is reconstructed into a MAFPN multi-scale adaptive feature pyramid network, which integrates the ASF adaptive spatial fusion module and the ScalSeq scale sequencer module; the detection head part of YOLOv10 retains the original v10Detect detection head structure.
[0062] The specific structure of the road marking detection model is described below:
[0063] like Figure 2 As shown, the backbone network includes a first Conv module, a second Conv module, a first C2f module, a first ADown module, a second C2f module, a second ADown module, a third C2f module, a third ADown module, a fourth C2f module, an SPPF-LSKA module, and a PSA self-attention module.
[0064] The neck network includes an upsampling path, a downsampling path, a Concat splicing operation, a C2f fusion module, a C2fCIB module, an ASF adaptive spatial fusion module, a ScalSeq scale serialization module, and an Add residual connection module.
[0065] The detection head section includes the v10Detect three-scale detection head.
[0066] The specific data flow is as follows:
[0067] Backbone network data flow:
[0068] The original image is used as the input to the first Conv module. The output feature map of the first Conv module is used as the input to the second Conv module. The output feature map of the second Conv module is used as the input to the first C2f module. The output feature map of the first C2f module is used as the input to the first ADown module. The output feature map of the first ADown module is used as the input to the second C2f module. The output feature map of the second C2f module is used as the input to the second ADown module. The output feature map of the second ADown module is used as the input to the third C2f module. The output feature map of the third ADown module is used as the input to the fourth C2f module. The output feature map of the fourth C2f module is processed sequentially by the SPPF-LSKA module and the PSA self-attention module.
[0069] Neck network MAFPN feature fusion:
[0070] (1) First stage fusion: The output of the third C2f module is processed by Conv and then concatenated with the output of the PSA module. Then, the P5 feature is obtained by processing through the C2fCIB module.
[0071] (2) Second stage fusion: P5 features are upsampled by Upsample, and the output of the second C2f module is processed by Conv. The two are concatenated with the output of the third C2f module in three ways by Concat, and then processed by the C2f module to obtain P4 features.
[0072] (3) Third stage fusion: P4 features are upsampled by Upsample, and the output of the first C2f module is processed by Conv. The two are concatenated with the output of the second C2f module in three ways by Concat, and then processed by the C2f module to obtain the preliminary P3 features.
[0073] (4) ASF adaptive spatial fusion processing: The output of the second C2f module is processed by Conv to obtain feature_p3, and the output of the first C2f module is processed by Conv to obtain feature_backbone. The three features of feature_p3, the original output of the second C2f module, and feature_backbone are adaptively concatenated by the Zoom_cat module, and then processed by the C2f module to obtain the final P3 feature.
[0074] (5) Downsampling return path: P3 to P4 return: P3 features are downsampled by Conv, concatenated with P4 features in the second stage, and processed by C2f module to obtain the final P4 features.
[0075] (6) P4 to P5 return: The P4 feature is downsampled by ADown, concatenated with the P5 feature of the first stage, and processed by the C2fCIB module to obtain the final P5 feature.
[0076] (7) ScalSeq scale serialization and residual connection: Multi-scale serialization: The features of the second, third, and fourth C2f modules are scaled and serialized through the ScalSeq module to output 256-dimensional features.
[0077] (8) Residual connection: The output of the ScalSeq module and the P3 features after ASF processing are residually connected through the Add module to obtain enhanced P3 features.
[0078] (9) Final detection output: Three-scale detection: The enhanced P3 feature, the final P4 feature, and the final P5 feature are respectively input into the v10Detect detection head to detect small, medium, and large targets.
[0079] I. ADown Adaptive Downsampling Module
[0080] As slender linear targets, road markings are prone to losing edge details during downsampling. Traditional SCDown modules employ fixed downsampling strategies, failing to adaptively adjust based on the specific distribution characteristics of the input features. This results in the weakening or loss of crucial marking features during downsampling. To address this issue, this invention proposes the ADown adaptive downsampling module. By analyzing the spatial distribution characteristics of the input features, it dynamically generates downsampling weights, achieving a balance between feature preservation and computational efficiency.
[0081] like Figure 3 As shown, the ADown adaptive downsampling module mainly consists of a feature analysis layer, an adaptive weight generation layer, a dual-branch downsampling processing layer, a spatial attention fusion layer, a feature preservation layer, and an information fusion layer. The core idea of this module is to simultaneously perform standard downsampling and adaptive downsampling through a dual-branch structure, and then dynamically fuse the results of the two branches according to the importance of features.
[0082] First, the feature analysis layer processes the input feature map F∈R C*H*W A global spatial analysis is performed, where C, H, and W represent the number, length, and width of channels, respectively. Global average pooling is used to extract global information for each channel, generating a spatial distribution descriptor. This step captures the statistical characteristics of each channel across the entire spatial dimension, providing a basis for subsequent adaptive weight generation. The spatial distribution descriptor S∈R C*1*1 The calculation formula is:
[0083]
[0084] Where F c (i,j) represents the feature value of the c-th channel at position (i,j), and GAP represents the global average pooling operation.
[0085] Next, the adaptive weight generation layer generates importance weights for each channel based on the spatial distribution descriptor. This layer learns the correlation between channels through a fully connected layer and uses a sigmoid activation function to map the weights to the [0,1] interval, ensuring the weights' rationality. The downsampled weights W∈R (C*1*1) The calculation formula is:
[0086] Wc=σ(FC(Sc)) (2)
[0087] Where σ is the Sigmoid activation function, and FC is a fully connected layer that learns the nonlinear mapping relationship of channel importance. Then, the process enters a two-branch downsampling stage. The first branch performs traditional standard downsampling to ensure basic downsampling functionality; the second branch performs adaptive downsampling based on importance weights, prioritizing the preservation of important features. The output of the standard downsampling branch is:
[0088] F std =Conv std (F) (3)
[0089] The adaptive downsampling branch first multiplies the input features element-wise with the weights to highlight the features of important channels, and then performs convolution processing:
[0090]
[0091] in Conv represents the element-wise multiplication operation. std and Conv adp These are standard convolution and adaptive convolution, respectively.
[0092] Subsequently, the spatial attention fusion layer calculates the fusion weights of the two branch results. This layer generates spatial attention weights α∈R by concatenating the features of the two branches and learning the fusion strategy using 1×1 convolutions. (C*H′*W′) :
[0093] α=σ(Conv 1*1 (Concat(F std ,F adp ))) (5)
[0094] Finally, the information fusion layer adaptively fuses the results from the two branches according to attention weights, which preserves the stability of standard downsampling while incorporating the feature protection capabilities of adaptive downsampling.
[0095]
[0096] Compared to the original SCDown module, the ADown module uses an adaptive weight allocation mechanism to dynamically adjust the downsampling strategy based on the spatial distribution of input features, effectively reducing the loss of detailed features of the marker lines during downsampling. Especially for slender marker line features, this module can identify and preserve these important edge information, significantly improving the accuracy of subsequent detection.
[0097] II. SPPF-LSKA Large Core Spatial Attention Enhancement Module
[0098] Road markings typically have long, continuous linear features, requiring a large receptive field to capture complete line information and long-range spatial dependencies. Traditional attention mechanisms often focus on local features, making it difficult to effectively model the global continuity and long-range dependencies of road markings, especially when dealing with long straight lane lines and dashed markings, which can easily result in broken or discontinuous detection results. To address this issue, this invention proposes the SPPF-LSKA large-kernel spatial attention enhancement module, which enhances the model's ability to model long-range dependencies of road markings by combining spatial pyramid pooling and a large-kernel spatial attention mechanism.
[0099] like Figure 4 As shown, the SPPF-LSKA module combines Spatial Pyramid Pooling (SPPF) and Large Kernel Spatial Attention (LSKA) mechanisms, mainly including the Spatial Pyramid Pooling branch and the Large Kernel Spatial Attention branch. The design concept of this module is to capture the features of marker lines at different scales and model long-distance spatial dependencies by combining multi-scale feature extraction and a large receptive field attention mechanism.
[0100] In the spatial pyramid pooling branch, the input feature F is first subjected to channel dimensionality reduction to reduce computational complexity while maintaining feature expressiveness.
[0101] F reduced =Conv 1*1 (F) (7)
[0102] Then, multi-scale pooling is performed, using pooling kernels of different sizes to capture feature information from different receptive fields. Multi-scale pooling can simultaneously focus on both the local details and global shape features of the marker line.
[0103] F pool1 =MaxPool 5*5 (F reduced (8)
[0104] F pool2 =MaxPool 9*9 (F reduced (9)
[0105] F pool3 =MaxPool 13*13 (F reduced (10)
[0106] By employing pooling operations of varying sizes, this branch can capture multi-level feature information from details to the global picture, making it particularly suitable for processing marker line targets of different widths and lengths. Multi-scale feature fusion is achieved through channel concatenation.
[0107] F sppf =Concat(Freduced ,F pool1 ,F pool2 ,F pool3 (11)
[0108] In the large kernel spatial attention branch, a large kernel convolution is used to process the input features, providing a larger receptive field to capture the long-range spatial dependencies of the marker lines. Compared to traditional small kernel convolution, large kernel convolution can cover a larger spatial region at once, which has a significant advantage for modeling slender marker line features.
[0109] F large =Conv 7*7 (F) (12)
[0110] Next, spatial attention weights are generated. By combining the results of global average pooling and global max pooling, both the average response of the features and the most salient activation regions are considered. This dual pooling strategy can more comprehensively evaluate the importance of spatial location.
[0111] A spatial =σ(Conv 1*1 (GAP(F large )+GMP(F large (13)
[0112] GMP stands for Global Max Pooling. Then, selective attention enhancement is performed, applying spatial attention weights to the result of the large-kernel convolution to highlight important spatial regions.
[0113]
[0114] Finally, the dual-branch feature fusion layer fuses the results of the spatial pyramid pooling branch and the large-kernel spatial attention branch, and preserves the original feature information through residual connections:
[0115] F output =Conv 1*1 (F sppf +F lska The SPPF-LSKA module (15) combines multi-scale pooling and large kernel attention to effectively capture feature information under different receptive fields, enhancing its ability to model long-distance dependencies of marker lines. This module is particularly suitable for handling continuous marker line detection tasks and can significantly improve the recognition accuracy of complex marker line patterns such as long straight lines and dashed line sequences.
[0116] III. MAFPN Multi-Scale Adaptive Feature Pyramid Network
[0117] Traditional feature pyramid networks often employ fixed fusion strategies when processing road marking targets at different scales, making it difficult to adaptively adjust based on the importance of features at different scales. Road markings in real-world scenarios exhibit diverse scale characteristics: distant markings are smaller in scale but require precise localization, while nearby markings are larger in scale but prone to detail loss. To better address this multi-scale feature fusion problem, this invention designs the MAFPN multi-scale adaptive feature pyramid network, which achieves multi-path feature fusion through an adaptive weight allocation mechanism, enhancing information interaction between different scales.
[0118] like Figure 5 As shown, the MAFPN network achieves adaptive fusion of multi-scale features, including upsampling paths, downsampling paths, and an adaptive weight allocation mechanism. The core innovation of this network lies in the introduction of a feature importance evaluation mechanism, which dynamically adjusts the contribution weights of features at different scales during the fusion process, thereby achieving more effective multi-scale feature representation.
[0119] In the adaptive weight allocation mechanism, for features F_i∈{P3,P4,P5} at different scales, global information of features at each scale is first extracted through global average pooling, and then a fully connected layer is used to learn the feature importance mapping and calculate the feature importance weights:
[0120] w i =σ(FC(GAP(F) i ))) (16)
[0121] To ensure the rationality of weight allocation, the weights of all scales are normalized so that the sum of the weights is 1. This normalized weight mechanism ensures the balance of features at different scales during the fusion process, preventing one scale feature from dominating and ignoring information from other scales.
[0122] In the upsampling feature fusion path, high-level features are fused with low-level features through upsampling to convey rich semantic information. In the P5 to P4 fusion process, P5 features are first upsampled to match the spatial size of P4, then feature fusion is performed through channel concatenation and convolution, and finally adaptive weights are applied.
[0123]
[0124] Similarly, the transition from P4 to P3 is also the same.
[0125] In the downsampled feature backhaul path, low-level features are fused with high-level features through downsampling to enhance the transmission of detailed information. The backhaul process from P3 to P4 uses standard downsampling operations:
[0126]
[0127] The P4 to P5 backhaul process uses the ADown adaptive downsampling module to better preserve feature information.
[0128]
[0129] MAFPN achieves effective fusion of features at different scales through an adaptive weight allocation mechanism, enhancing cross-scale feature interaction capabilities. Compared to traditional FPN networks, MAFPN can dynamically adjust the fusion strategy according to the actual distribution of input features, making it particularly suitable for handling multi-scale marker line target detection tasks in complex scenes.
[0130] IV. ASF Adaptive Spatial Fusion Module
[0131] In multi-scale feature fusion, features at different spatial scales often possess different spatial distribution characteristics and importance. Simple feature concatenation or addition operations are insufficient to fully utilize the complementary information of features at different scales. This is particularly true for marker line detection tasks, where features at different scales may focus on different aspects of the marker line: small-scale features focus on detailed edge information, while large-scale features focus on overall shape and continuity. To better fuse these multi-scale features across spatial dimensions, this invention designs an ASF adaptive spatial fusion module, which achieves adaptive fusion of features at different spatial scales through the Zoom_cat operation and spatial attention mechanism.
[0132] like Figure 6 As shown, the ASF module achieves adaptive spatial fusion through the Zoom_cat operation and spatial attention mechanism. The core idea of this module is to evaluate the importance of features at different locations through the spatial attention mechanism, and then use the Zoom_cat operation to achieve adaptive feature stitching in spatial dimensions, thereby making full use of the complementary spatial information of multi-scale features.
[0133] During the channel adjustment and attention generation stages, the number of channels for the three different scale features F_1, F_2, and F_3 are first unified to ensure compatibility with subsequent fusion operations.
[0134] A i =σ(Conv(GAP(F') i ))) (20)
[0135] Then, spatial attention weights are generated for each scale feature, global information is extracted through global average pooling, and attention maps are generated through convolution and sigmoid activation.
[0136] A i =σ(Conv(GAP(F) i '))) (twenty one)
[0137] These attention weights can highlight important spatial regions in features at various scales and suppress irrelevant background information.
[0138] like Figure 7 As shown, in the Zoom_cat adaptive stitching stage, attention enhancement processing is first applied to the features to highlight important spatial regions. Features at different scales are upsampled to the same spatial size to ensure the effectiveness of the stitching operation. Then, an adaptive stitching operation is performed, stitching the enhanced features from the three scales together along the channel dimension.
[0139] F concat =Concat(F 1norm ,F 2norm ,F 3norm ) (twenty two)
[0140] Finally, feature alignment and enhancement are performed, using batch normalization and ReLU activation functions for feature normalization and nonlinear transformation.
[0141] F asf =BatchNorm(ReLU(Conv(F concat ))) (twenty three) The ASF module, through the combination of spatial attention and the Zoom_cat operation, effectively fuses feature information from different spatial scales, enhancing its ability to perceive the spatial location information of marker lines. This module is particularly suitable for handling marker line detection tasks in complex backgrounds, fully utilizing the spatial complementarity of multi-scale features to improve detection accuracy and robustness.
[0142] V. ScalSeq Scale Serialization Module
[0143] In multi-scale feature processing, features at different scales are typically processed in parallel, lacking sequential modeling and progressive feature enhancement across scales. For marker line detection tasks, features from small to large scales often exhibit a progressive semantic relationship: small-scale features capture local details, medium-scale features focus on line segment connections, and large-scale features model the overall shape. To better utilize this progressive relationship between scales, this invention designs the ScalSeq scale sequentialization module to sequentially process features at multiple scales and optimize multi-scale feature representations through progressive feature enhancement.
[0144] like Figure 8 As shown, the ScalSeq module performs serialization processing on features at multiple scales. The core idea of this module is to transform multi-scale feature processing from a parallel mode to a serialization mode, achieving more effective multi-scale feature representation by gradually accumulating and enhancing feature information at different scales.
[0145] During the channel mapping and size normalization stages, the number of channels for input features of different scales F_4, F_6, and F_8 is unified to ensure consistency in subsequent processing.
[0146] F i ′=Conv 1*1 (F i ) (twenty four)
[0147] Then, all features are unified to the same spatial size, and size alignment is achieved through upsampling. In the scale importance evaluation stage, the importance weights of features at each scale are calculated, and the relative importance between scales is learned through global average pooling and fully connected layers.
[0148]
[0149] To ensure the reasonableness of the weight allocation, all scale weights are normalized:
[0150]
[0151] The same operation applies to α6' and α8'. During the serialization fusion stage, features at different scales are weighted and fused according to their importance to achieve feature integration in the serialization process.
[0152]
[0153] Finally, the enhanced features are fused with the input features using the Add residual connection operation, preserving the original feature information while introducing serialization enhancement information:
[0154] F final =F enhanced +F input (28)
[0155] Where F input The input features are for residual connections.
[0156] The ScalSeq module, through scale-based serialization and progressive feature enhancement, better utilizes the progressive relationships between multi-scale features, optimizes multi-scale feature representation, and improves overall detection performance. This module is particularly suitable for handling marker line detection tasks with obvious scale hierarchies, enhancing the expressive power and discriminative power of features through serialization modeling.
[0157] Step 3: Train the road marking detection model using the dataset;
[0158] Input the training set image data from the dataset into the network model for training, and save the model parameters with the highest accuracy on the validation set during the training process, naming the file best.pt.
[0159] The experimental environment was configured with Ubuntu 20.04 operating system, NVIDIA GeForce RTX 2060 graphics card, and PyTorch deep learning framework. The specific configuration is shown in Table 1 below:
[0160] Table 1 Experimental Environment Configuration Table
[0161] parameter Configuration GPU NVIDIA GeForce RTX 2060 System Environment Ubuntu 20.04 CUDA version CUDA 12.5 Programming language version Python 3.9 Deep learning framework PyTorch 2.5.1
[0162] The training image data from the road markings dataset was input into the network model. GPU training was used, with improved network model parameters set. The training run consisted of 300 epochs, employing the AdamW optimizer with a batch size of 49. The initial learning rate was 0.01, and a cosine annealing scheduling strategy was used to dynamically adjust the learning rate to a minimum of 0.0001. The weight decay coefficient was 0.00005. During model training, the model parameters with the highest accuracy were saved and named best.pt.
[0163] Figure 9 The figure shows the curves of various metrics during 300 rounds of model training. After the model training was completed, it was tested on the validation set. The validation set showed a precision of 94.8% for various markers, a recall of 94.0%, a mAP50 of 97.1%, and a mAP50-95 of 80.5%. Compared with the original YOLOv10n, the algorithm of this invention improved recall by 4.6% and mAP50-95 by 2.6%. Recall represents the model's recall rate; a higher value indicates fewer missed detections. mAP50-95 represents the model's average precision at different IoU thresholds; a higher value indicates better network performance. Table 2 shows a comparison of the algorithm metrics, demonstrating the feasibility of this invention.
[0164] Table 2 Comparison of experimental results for different models
[0165]
[0166] Step 4: Use the trained road marking detection model to detect road markings in complex scenes.
[0167] The video signal of the road scene is acquired by the camera, the video signal is extracted into an image and preprocessed. The size of the input image is scaled and adjusted to 640×640 pixels. The trained road marking detection model is used to detect double yellow lines, center lines, dashed lines, intersection markings, lane lines, pedestrian crossings, and various arrow markings in the image, and the detection results are displayed on the monitor in real time. Figure 10 This is a diagram showing the detection effect of the present invention.
[0168] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered by the present invention.
Claims
1. A method for detecting road markings in complex scenes based on multi-module adaptive fusion, characterized in that: Including the following steps: Step 1: Obtain the road marking dataset; Step 2: Construct a road marking detection model based on YOLOv10, including: Step 2.1: In the YOLOv10 backbone network, replace the SCDown downsampling module with the ADown downsampling module, and add the SPPF_LSKA spatial pyramid pooling enhancement module after the PSA module; Step 2.2: In the neck network of YOLOv10, it is reconstructed into a MAFPN multi-scale adaptive feature pyramid network, which integrates the ASF adaptive spatial fusion module and the ScalSeq scale sequencer module. Step 2.3: In the YOLOv10 detection head section, maintain the original v10Detect detection head structure; Step 3: Train the road marking detection model using the dataset; Step 4: Use the trained road marking detection model to detect road markings in complex scenes.
2. The method for detecting road markings in complex scenes based on multi-module adaptive fusion according to claim 1, characterized in that: The dataset mentioned in step one uses the CRLDD complex road marking detection dataset.
3. The method for detecting road markings in complex scenes based on multi-module adaptive fusion according to claim 1, characterized in that: The backbone network described in step two includes a first Conv module, a second Conv module, a first C2f module, a first ADown module, a second C2f module, a second ADown module, a third C2f module, a third ADown module, a fourth C2f module, an SPPF-LSKA module, and a PSA self-attention module; The neck network includes an upsampling path, a downsampling path, a MAFPN feature fusion path, an ASF adaptive spatial fusion module, a ScalSeq scale serialization module, and an Add residual connection module. The detection head portion includes the v10Detect three-scale detection head.
4. The method for detecting road markings in complex scenes based on multi-module adaptive fusion according to claim 3, characterized in that: The data flow of the road marking detection model is as follows: Backbone network data flow: The original image is used as the input to the first Conv module. The output feature map of the first Conv module is used as the input to the second Conv module. The output feature map of the second Conv module is used as the input to the first C2f module. The output feature map of the first C2f module is used as the input to the first ADown module. The output feature map of the first ADown module is used as the input to the second C2f module. The output feature map of the second C2f module is used as the input to the second ADown module. The output feature map of the second ADown module is used as the input to the third C2f module. The output feature map of the third C2f module is used as the input to the third ADown module. The output feature map of the third ADown module is used as the input to the fourth C2f module. The output feature map of the fourth C2f module is processed sequentially by the SPPF-LSKA module and the PSA self-attention module. Neck network MAFPN feature fusion: (1) First stage fusion: The output of the third C2f module is processed by Conv and then concatenated with the output of the PSA module. Then, it is processed by the C2fCIB module to obtain the P5 feature; (2) Second stage fusion: P5 features are upsampled by Upsample, and the output of the second C2f module is processed by Conv. The two are concatenated with the output of the third C2f module in three ways by Concat, and then processed by the C2f module to obtain P4 features; (3) Third stage fusion: P4 features are upsampled by Upsample, the output of the first C2f module is processed by Conv, the two are concatenated with the output of the second C2f module in three ways by Concat, and then processed by the C2f module to obtain the preliminary P3 features; (4) ASF adaptive spatial fusion: The output of the second C2f module is processed by Conv to obtain feature_p3, and the output of the first C2f module is processed by Conv to obtain feature_backbone. The three features of feature_p3, the original output of the second C2f module, and feature_backbone are adaptively concatenated by the Zoom_cat module, and then processed by the C2f module to obtain the final P3 feature. (5) Downsampling return path: P3 to P4 return: P3 features are downsampled by Conv, concatenated with P4 features from the second stage, and processed by the C2f module to obtain the final P4 features; (6) P4 to P5 return: The P4 feature is downsampled by ADown, concatenated with the P5 feature in the first stage, and processed by the C2fCIB module to obtain the final P5 feature; (7) ScalSeq scale serialization and residual connection: Multi-scale serialization: The features of the second C2f module, the third C2f module, and the fourth C2f module are processed by the ScalSeq module to scale serialization, and output 256-dimensional features; (8) Residual connection: The output of the ScalSeq module and the P3 features after ASF processing are residually connected through the Add module to obtain enhanced P3 features; (9) Final detection output: Three-scale detection: The enhanced P3 feature, the final P4 feature, and the final P5 feature are respectively input into the v10Detect detection head to detect small, medium, and large targets.
5. The method for detecting road markings in complex scenes based on multi-module adaptive fusion according to claim 1, characterized in that: The ADown downsampling module described above dynamically adjusts the downsampling strategy based on the spatial distribution of the input features through an adaptive weight allocation mechanism.
6. The method for detecting road markings in complex scenes based on multi-module adaptive fusion according to claim 1, characterized in that: The SPPF_LSKA spatial pyramid pooling enhancement module combines spatial pyramid pooling with a large kernel selection attention mechanism.
7. The method for detecting road markings in complex scenes based on multi-module adaptive fusion according to claim 1, characterized in that: The MAFPN multi-scale adaptive feature pyramid network achieves effective fusion of features at different scales through adaptive weight allocation.
8. The method for detecting road markings in complex scenes based on multi-module adaptive fusion according to claim 1, characterized in that: The ASF adaptive spatial fusion module achieves adaptive stitching and scaling of multi-scale features through the Zoom_cat module.
9. The method for detecting road markings in complex scenes based on multi-module adaptive fusion according to claim 1, characterized in that: The ScalSeq scale serialization module serializes features at multiple scales and enhances features through the Add operation.
Citation Information
Cited By
Ceramic tile surface flaw detection method and system based on improved YOLOv8n
CN120525814A
Tile surface defect detection method and system based on improved YOLOv8n
CN120525814B