Road surface damage detection method based on Mangbar simplified model

By introducing a simplified Mamba model with a self-calibrating selection network and spatial channel coordinate attention blocks, pavement damage detection is optimized, solving the problem of fine-grained features and background interference in existing methods, and achieving pavement damage detection with higher accuracy and lower complexity.

CN120876994APending Publication Date: 2025-10-31ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511084508.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing methods for detecting road damage struggle to effectively capture fine-grained features and complex background interference in damaged areas, resulting in insufficient accuracy in detecting minute cracks and blurred edge regions. Furthermore, the reliance on large-scale labeled datasets is costly and limits the generalization ability of the models.

Method used

A road damage detection method based on the Mamba simplified model is adopted. By introducing a self-calibrating selection network and spatial channel coordinate attention blocks, feature representation is enhanced and background interference is reduced. Combined with depthwise separable convolution and multi-scale feature fusion, the feature extraction and classification process is optimized.

Benefits of technology

It significantly improves the ability to identify minute cracks and blurred edges, enhances the accuracy and robustness of road damage classification, reduces computational complexity, and is suitable for real-time detection on vehicle-mounted systems and mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876994A_ABST
    Figure CN120876994A_ABST
Patent Text Reader

Abstract

The invention discloses a road surface damage detection method based on a Mangbar simplified model, and relates to the technical field of road surface damage detection. The method includes: acquiring a to-be-recognized pavement image; performing feature screening according to the importance of the spatial information of each channel to obtain an initial processing image; dividing the preliminarily processed image into a main channel image and a secondary channel image, and performing adaptive weighting to obtain a feature map; performing feature extraction on the feature map by adopting a depth separable convolution operation to obtain a multi-scale feature; performing feature fusion on the multi-scale features by adopting spatial weighting operation to obtain initial features; performing multiple spatial information fusion on the initial features, and taking output features of the last spatial information fusion as fusion features; and performing pavement damage classification according to the fusion features to obtain a pavement damage detection result. According to the method, the distinction degree between similar pavement damage types can be improved, and high-precision detection of pavement tiny damage and fuzzy edges is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pavement damage detection technology, and in particular to a pavement damage detection method based on the Mamba simplified model. Background Technology

[0002] With continued economic growth and increasing traffic volume, road infrastructure is facing increasing pressure. Consequently, various pavement defects such as cracks and potholes are frequently observed. These defects not only affect driving safety but also reduce traffic efficiency. Pavement damage detection and classification play a crucial role in road maintenance and management. Their accuracy and efficiency directly impact road lifespan, driving safety, and overall traffic efficiency.

[0003] In recent years, the Transformer model has demonstrated outstanding performance in computer vision, providing a new research direction for road damage detection. Some studies utilize Transformer-based architectures to improve classification and segmentation accuracy in this field. For example, researchers have proposed a weakly supervised visual transformer, PicT, specifically for road damage classification. Based on the Swin-Transformer framework, this novel image classification Transformer introduces a patch label teacher model that dynamically generates pseudo-labels for patches. This method enables the model to learn key discriminative features in a weakly supervised environment. Similarly, researchers have developed a Transformer-based detection framework, TransCrack, specifically for crack recognition. This framework employs a five-level pure transformer encoder and a contrastive learning attention (CLA) mechanism to enhance global feature extraction of crack regions. Its core idea is to divide the input image into fixed grid cells, embed positional information, and process it through a transformer to capture long-range dependencies. Furthermore, the decoder integrates the CLA mechanism, fusing global and local attention at multiple scales to achieve high-resolution, pixel-level crack segmentation. Additionally, researchers have introduced a Transformer model that combines dense multi-scale feature learning. The model employs a cross-attention mechanism to expand the receptive field of feature extraction, while utilizing a dense multi-scale feature learning module to integrate local information at different scales.

[0004] Most existing methods simply apply traditional object detection or image classification frameworks to road damage detection. However, road damage areas usually exhibit minute geometric features (such as sub-millimeter cracks), blurred edge transitions (such as the gradual boundary between shallow spalling and normal road surface), and high similarity to background noise (such as asphalt texture, repair marks, shadows, etc.). Existing methods fail to fully consider the fine-grained features and complex background interference unique to road damage, making it difficult to effectively capture key damage features and accurately detect minute cracks and blurred edge areas in road damage. Summary of the Invention

[0005] Therefore, it is necessary to provide a road surface damage detection method based on the simplified Mamba model to address the aforementioned technical problems.

[0006] This invention provides a pavement damage detection method based on the Mamba simplified model, comprising: Acquire the image of the road surface to be identified; Feature filtering is performed based on the importance of spatial information of each channel in the road surface image to be identified, in order to remove spatially redundant features and obtain an initial processed image. The initial processed image is divided into a main channel image and a secondary channel image, and adaptive weighting is applied to the main channel image and the secondary channel image to remove temporally redundant features and obtain a feature map. Depthwise separable convolution operation is used to extract features from the feature map to obtain multi-scale features. Spatial weighting operation is used to fuse the multi-scale features to obtain the initial features. The initial features are fused multiple times with spatial information. The output features of the previous spatial information fusion are downsampled and the output features after downsampling are used as the input features of the next spatial information fusion. The output features of the last spatial information fusion are used as the fused features. Each spatial information fusion step includes: grouping the input features according to their channels and determining the global average features of all spatial locations within each channel group to obtain an attention factor; reweighting the input features using the attention factor to obtain intermediate features; performing convolution operations on the intermediate features to update the local features within the intermediate features to obtain intermediate processed features; performing one-dimensional convolution operations on the intermediate processed features in both the horizontal and vertical directions to obtain feature vectors in the two directions, and then fusing the feature vectors in the two directions to obtain the output features. Damage is classified based on fusion characteristics to obtain pavement damage detection results.

[0007] Optionally, pavement damage detection results are obtained by improving the Mamba simplified model. The improved Mamba simplified model includes: an original dry layer with a self-calibrating selection network, multiple feature processing modules and a classification layer connected in sequence. Each feature processing module includes: a gated convolutional neural network block with spatial channel coordinate attention blocks and a downsampling layer with a self-calibrating selection network connected in sequence. The original dry layer of the self-calibrating selection network consists of sequentially connected spatial and channel reconstruction convolutional modules and a large selection kernel; The spatial and channel reconstruction convolutional module includes: spatial reconstruction units and channel reconstruction units connected in sequence. The channel reconstruction unit includes a channel segmentation stage, a channel transformation stage, and a channel fusion stage. The large selection kernel includes: a multi-scale large convolutional kernel and a spatial selection mechanism connected in sequence. The gated convolutional neural network block that introduces spatial channel coordinate attention includes: a spatial grouping enhancement module, a partial convolution module, and a coordinate attention module connected in sequence.

[0008] Optionally, train the improved Mamba simplified model, specifically including: Obtain historical road surface images and corresponding historical road surface damage types; Input historical road surface images into the improved Mamba simplified model to obtain the predicted road surface damage type; With the goal of minimizing the deviation between the predicted pavement damage type and the historical pavement damage type, the improved Mamba simplified model is trained to obtain the trained improved Mamba simplified model.

[0009] Optionally, constructing an original dry layer with a self-calibration selection network includes: using the self-calibration selection network as a parallel branch of the original dry layer, and connecting the output of the self-calibration selection network and the output of the original dry layer through residual connections to obtain the original dry layer with the self-calibration selection network. Constructing a downsampling layer that incorporates a self-calibration selection network specifically includes: using the self-calibration selection network as a parallel branch of the downsampling layer, and connecting the output of the self-calibration selection network and the output of the downsampling layer through residual connections to obtain the downsampling layer that incorporates the self-calibration selection network. Constructing a gated convolutional neural network block that incorporates spatial channel coordinate attention blocks specifically involves: using the spatial channel coordinate attention block as a parallel branch of the gated convolutional neural network block, and connecting the output of the spatial channel coordinate attention block and the output of the gated convolutional neural network block through residual connections to obtain the gated convolutional neural network block that incorporates spatial channel coordinate attention blocks.

[0010] Optionally, feature filtering is performed based on the importance of spatial information of each channel in the road surface image to be identified, in order to remove spatially redundant features and obtain an initial processed image, specifically including: The importance of spatial information for each channel is determined based on the following formula, and group normalization is used to determine the weight of each channel: ; ; in, For the spatial information of the channel, For the first i The initial processed image of the stage, This indicates that a gated convolutional neural network block, incorporating spatial channel coordinate attention, is used for processing. This indicates that the self-calibration network is selected for processing. The spatial channel coordinates require special processing. i For stage indexing; Based on the channel importance weight, the contribution of different channels to spatial information is quantified to obtain the relative importance of each channel in encoding spatial details; Channels with relative importance higher than the spatial weight threshold are retained, while features with relative importance lower than the spatial weight threshold are discarded to obtain the initial processed image.

[0011] Optionally, the pre-processed image is divided into a primary channel image and a secondary channel image, and adaptive weighting is applied to the primary channel image and the secondary channel image to remove temporal redundancy features, resulting in a feature map, specifically including: The initial processed image is divided into a main channel image and a secondary channel image according to a set ratio; Convolution operations at different scales are performed on the two channels to obtain a feature-enhanced main channel image and a feature-enhanced secondary channel image; The main channel image and the secondary channel image of feature enhancement are adaptively weighted according to the attention coefficients based on the following formula to remove temporally redundant features in the road surface image to be identified, thus obtaining the feature map: ; ; ; ; in, For the initial image processing, Main channel image, For secondary channel images, For feature-enhanced main channel images, For feature-enhanced secondary channel images, For feature maps, This indicates a splitting operation. This represents grouped weighted convolution. This represents pyramid-weighted convolution. This represents the Softmax activation function.

[0012] Optionally, feature extraction can be performed on the feature map using a depthwise separable convolution operation based on the following formula to obtain multi-scale features: ; Based on the following formula, spatial weighting operations are used to fuse multi-scale features. The spatial weighting operations include parallel average pooling and max pooling operations to obtain the initial features: ; (Sigmoid( )); in, For multi-scale features, This is the output of multi-scale features after average pooling. This is the output of multi-scale features after max pooling. As initial features, This represents a depthwise separable convolution operation, and Sigmoid represents the Sigmoid activation function. This indicates the average pooling operation. This indicates a max pooling operation.

[0013] Optionally, intermediate features can be obtained by reweighting the input features using an attention factor based on the following formula: Sigmoid(Norm(PDP(GAP( )))); The intermediate features are convolved using the following formula to update local features within the intermediate features, resulting in intermediate processed features: PC )+ ; Based on the following formula, two independent one-dimensional convolution operations are performed on the intermediate processing features in the horizontal and vertical directions respectively to obtain the output features: Sigmoid(Conv2d(Conv2d( )+ )); in, For input features, As an intermediate feature, For intermediate processing features, The feature vector in the horizontal direction, The feature vectors in the vertical direction, For the output features, GAP represents global average pooling, PDP represents probability density prediction, Norm represents normalization, Sigmoid represents the Sigmoid activation function, PC represents principal component features, and Conv2d represents two-dimensional convolution.

[0014] The pavement damage detection method based on the simplified Mamba model provided in this invention has the following advantages compared with the prior art: This invention preserves edge gradient features through multiple spatial information fusions, thereby improving the ability to identify blurred boundaries; and it performs downsampling operations on the output features of each spatial information fusion to remove redundant data and reduce feature confusion between similar damage types, thereby improving the distinguishability between similar road surface damage types.

[0015] Furthermore, feature fusion of multi-scale features enhances the correlation between spatial features at different scales, thereby improving the feature representation of damaged areas of the road surface. This allows for more accurate aggregation of spatial distribution information of small structures such as cracks, enabling high-precision detection of minor road surface damage and blurred edges, and significantly improving the accuracy and robustness of road surface damage classification. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the overall structure of a pavement damage detection method based on the Mamba simplified model provided in one embodiment; Figure 2 This is a structural diagram of a self-calibration selection network for a pavement damage detection method based on a simplified Mamba model, provided in one embodiment. Figure 3 This is a structural diagram of the spatial channel coordinate attention module of a pavement damage detection method based on the Mamba simplified model provided in one embodiment; Figure 4 The image shows a CQU-BPDD dataset example of a pavement damage detection method based on the Mamba simplified model provided in one embodiment. Figure 5 This is a visualization heatmap of a pavement damage detection method based on the Mamba simplified model provided in one embodiment. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0018] Traditional detection methods have several limitations: they are time-consuming, require significant manpower, and have limited accuracy. These limitations prevent them from meeting the demands of modern transportation systems for large-scale, high-precision, and real-time monitoring. Early road damage detection relied primarily on manual inspection. Road maintenance personnel assessed road conditions by walking or using vehicle-mounted inspection equipment. This process involved visual, tactile, and auditory assessments. However, this method was extremely labor-intensive and time-consuming. Furthermore, it was highly subjective and inefficient. Environmental factors such as weather and lighting conditions further affected accuracy. Manual inspection also struggled to accurately measure the geometric features of damaged areas, such as size and depth. This limitation reduced the applicability of manual inspection in large-scale road networks.

[0019] With advancements in imaging technology, some countries have adopted vehicle-mounted cameras to capture road surface images. These images are then manually analyzed to determine road conditions. While this method reduces the need for on-site manual inspection, it still relies heavily on manual image analysis. Therefore, detection efficiency remains low, and subjectivity issues persist. Given these challenges, early road damage detection methods are no longer sufficient to meet the demands of modern transportation systems for large-scale, high-precision, and real-time monitoring.

[0020] In recent years, advancements in image processing technology have brought new opportunities for road damage detection, leading to the development of road inspection systems. These systems then utilize image processing techniques to detect and classify road conditions. By constructing large-scale road image databases and integrating feature extraction algorithms, these methods have achieved automated analysis. However, early image processing techniques remain relatively underdeveloped. The detection process remains complex, with limited robustness and suboptimal accuracy. While these methods have reduced human intervention and improved automation to some extent, their detection capabilities still face significant limitations.

[0021] With the rapid development of artificial intelligence and deep learning technologies, researchers are increasingly exploring deep learning-based methods for road damage detection to improve accuracy and efficiency. Among these, object detection technology has made significant progress in automatic damage identification. For example, the YOLO (You Only Look Once) series of neural networks has become a widely adopted solution due to its high computational efficiency and excellent detection accuracy. Researchers have achieved top rankings in various countries using YOLO and Faster R-CNN as baseline models on the CRDDC dataset. Similarly, researchers have proposed an innovative method for pothole detection using low-resolution cameras or low-quality images and video streams, integrating ESRGAN for image enhancement with YOLOv7 for detection, significantly improving the model's performance in pothole identification. Furthermore, researchers have developed a lightweight YOLOv8-PD algorithm to address the challenge of detecting large-span cracks, introducing a BOT module to enhance global information extraction. Additionally, a large-kernel separable attention mechanism has been employed to improve detection accuracy, and a C2fGhost block has been constructed in the neck network to reduce computational costs while enhancing feature extraction for complex road surface distress. To further optimize detection performance, this method employs a lightweight shared convolutional detection head, which improves feature representation capabilities while reducing parameter size.

[0022] In recent years, the Transformer model has demonstrated outstanding performance in computer vision, providing a new research direction for road damage detection. Some studies utilize Transformer-based architectures to improve classification and segmentation accuracy in this field. For example, researchers have proposed a weakly supervised visual transformer, PicT, specifically for road damage classification. Based on the Swin-Transformer framework, this novel image classification Transformer introduces a patch label teacher model that dynamically generates pseudo-labels for patches. This method enables the model to learn key discriminative features in a weakly supervised environment. Similarly, researchers have developed a Transformer-based detection framework, TransCrack, specifically for crack recognition. This framework employs a five-level pure transformer encoder and a contrastive learning attention mechanism to enhance global feature extraction of crack regions. Its core idea is to divide the input image into fixed grid cells, embed positional information, and process it through a transformer to capture long-range dependencies. Furthermore, the decoder integrates a contrastive learning attention mechanism to fuse global and local attention at multiple scales, thereby achieving high-resolution, pixel-level crack segmentation. Additionally, researchers have introduced a Transformer model that incorporates dense multi-scale feature learning. The model employs a cross-attention mechanism to expand the receptive field of feature extraction, while utilizing a dense multi-scale feature learning module to integrate local information at different scales.

[0023] Existing research largely treats pavement damage detection as a traditional object detection or image classification problem, often neglecting the fine-grained features of damaged areas. Due to the complexity of pavement damage patterns and strong background noise, current methods still struggle to detect minute cracks and blurred edge regions. Furthermore, deep learning models heavily rely on large-scale labeled datasets, but obtaining high-quality labeled pavement damage data is costly. This constraint limits the generalization ability of existing models, posing a challenge to widespread practical deployment.

[0024] Distinguishing between normal and damaged road surfaces is a significant challenge due to the intricate details in road damage images. The visual similarity between normal and damaged road surface images is high, with differences often concentrated in small damaged areas, making accurate detection even more complex. In particular, the extremely low contrast between damaged and undamaged areas, coupled with the fact that damage typically constitutes a small proportion of the overall image, exacerbates this difficulty. Furthermore, different types of cracks (such as alligator cracks, transverse cracks, and longitudinal cracks) share similar morphological features, further complicating accurate classification. To effectively address these challenges, this study proposes an improved approach based on the MambaOut model. A self-calibrating selection network and a spatial channel coordinate attention block are introduced and integrated into the MambaOut framework to develop a more robust damage detection model, termed the Improved Mamba Reduction Model (SCCA-MambaOut). This model adaptively enhances the feature representation of damaged areas while reducing background interference, thereby improving the discriminability between visually similar damage types. Therefore, it improves detection accuracy and classification reliability.

[0025] Extensive experiments demonstrate that the proposed SCCA-MambaOut model significantly outperforms the baseline MambaOut model. Experimental results on the CQU-BPDD pavement damage dataset show a 2% improvement in success rate, a 0.5% improvement in accuracy, and a 1.2% improvement in F1 score. These results highlight the effectiveness of the improved SCCA-MambaOut model in addressing the challenges of pavement damage classification, providing a promising solution for more accurate and reliable pavement condition assessment.

[0026] This invention provides a pavement damage detection method based on the Mamba simplified model, the method comprising: Obtain the image of the road surface to be identified.

[0027] Feature filtering is performed based on the importance of spatial information of each channel in the road surface image to be identified, in order to remove spatially redundant features and obtain an initial processed image. The initial processed image is divided into a primary channel image and a secondary channel image, and adaptive weighting is applied to the primary and secondary channel images to remove temporally redundant features, resulting in a feature map. Depthwise separable convolution operations are used to extract features from the feature map, obtaining multi-scale features. Spatial weighting operations are then used to fuse the multi-scale features, resulting in the initial features.

[0028] The initial features are fused multiple times using spatial information. The output features of the previous spatial information fusion are downsampled and used as the input features of the next spatial information fusion. The output features of the last spatial information fusion are used as the fused features.

[0029] Each spatial information fusion step includes: grouping the input features according to their channels and determining the global average features of all spatial locations within each channel group to obtain an attention factor; reweighting the input features using the attention factor to obtain intermediate features; performing convolution operations on the intermediate features to update local features, resulting in intermediate processed features; performing one-dimensional convolution operations on the intermediate processed features in both the horizontal and vertical directions to obtain feature vectors in both directions; and fusing these two feature vectors to obtain the output features.

[0030] Damage is classified based on fusion characteristics to obtain pavement damage detection results.

[0031] Among them, the road damage detection results are obtained by improving the Mamba simplified model. The improved Mamba simplified model includes: an original dry layer with a self-calibration selection network, multiple sets of feature processing modules and a classification layer connected in sequence. Each set of feature processing modules includes: a gated convolutional neural network block with spatial channel coordinate attention block and a downsampling layer with a self-calibration selection network connected in sequence.

[0032] The original dry layer of the self-calibrating selection network consists of sequentially connected spatial and channel reconstruction convolutional modules and a large selection kernel. The spatial and channel reconstruction convolutional modules consist of sequentially connected spatial reconstruction units and channel reconstruction units; the channel reconstruction units include channel segmentation, channel transformation, and channel fusion stages. The large selection kernel consists of sequentially connected multi-scale large convolutional kernels and a spatial selection mechanism. The gated convolutional neural network block that introduces spatial channel coordinate attention consists of sequentially connected spatial grouping enhancement modules, partial convolutional modules, and coordinate attention modules.

[0033] The spatial reconstruction unit filters features based on the importance of spatial information in each channel of the road surface image to be identified, removing spatially redundant features to obtain an initial processed image. The channel reconstruction unit divides the initial processed image into a primary channel image and secondary channel images, and adaptively weights these images to remove temporally redundant features, resulting in a feature map. Multi-scale features are extracted from the feature map using depthwise separable convolution operations with multi-scale large convolutional kernels. Finally, a spatial selection mechanism uses spatial weighting operations to fuse the multi-scale features, yielding the initial features.

[0034] By introducing a gated convolutional neural network block with spatial channel coordinate attention, multiple spatial information fusions are performed on the initial features. A downsampling layer of a self-calibrating selection network is then introduced to downsample the output features of the previous spatial information fusion, using the downsampled output features as the input features for the next spatial information fusion. The output features of the final spatial information fusion are then used as the fused features. A classification layer performs damage classification based on the fused features to obtain the pavement damage detection results.

[0035] The process of constructing the original dry layer with the introduction of the self-calibration selection network includes: using the self-calibration selection network as a parallel branch of the original dry layer, and connecting the output of the self-calibration selection network and the output of the original dry layer through residual connections to obtain the original dry layer with the introduction of the self-calibration selection network.

[0036] Constructing a downsampling layer that incorporates a self-calibration selection network includes: using the self-calibration selection network as a parallel branch of the downsampling layer, and connecting the output of the self-calibration selection network and the output of the downsampling layer through residual connections to obtain the downsampling layer that incorporates the self-calibration selection network.

[0037] Constructing a gated convolutional neural network block that incorporates spatial channel coordinate attention blocks includes: using the spatial channel coordinate attention block as a parallel branch of the gated convolutional neural network block, and connecting the output of the spatial channel coordinate attention block and the output of the gated convolutional neural network block through residual connections to obtain the gated convolutional neural network block that incorporates spatial channel coordinate attention blocks.

[0038] In addition, training and improving the simplified Mamba model specifically includes: Obtain historical pavement images and their corresponding historical pavement damage types. Input the historical pavement images into the improved Mamba simplified model to obtain the predicted pavement damage type. Train the improved Mamba simplified model with the objective of minimizing the deviation between the predicted pavement damage type and the historical pavement damage type to obtain the trained improved Mamba simplified model.

[0039] The specific implementation is as follows: 1. Damage detection model (SCCA-MambaOut).

[0040] Image classification tasks differ fundamentally from tasks involving long sequence modeling, such as natural language processing, object detection, and semantic segmentation. Therefore, the MambaOut model (without a State Space Model, SSM) is more suitable for classification scenarios. The proposed SCCA-MambaOut model aims to address the uniqueness problem of road damage images by integrating a Self-Calibrated Selection Network (SCSN) and a Spatial Channel Coordinate Attention (SCCA) module. These components are specifically designed to address key challenges in road damage classification. The Self-Calibrated Selection Network (SCSN) is introduced to mitigate the challenge of low proportions of damaged areas and high similarity between crack types in images, which often leads to classification difficulties. This module adaptively enhances the feature representation of damaged areas, thereby increasing the model's attention to subtle cracks and small-scale defects. As a result, classification accuracy is significantly improved. The Spatial Channel Coordinate Attention (SCCA) module addresses visual confusion caused by uneven illumination, overlapping cracks, and overlapping shadows. By leveraging spatial attention mechanisms, this module enhances the model's ability to extract key features, ensuring effective differentiation between real damage and illumination-induced artifacts. Therefore, the model is more adaptable to environmental changes. Both SCSN and SCCA are designed as lightweight modules, further improving classification accuracy while maintaining computational efficiency. Figure 1 The overall architecture of SCCA-MambaOut is shown.

[0041] The SCCA-MambaOut model structure consists of three main modules. Components (a) and (c) are optimized versions of the original Stem and downsampling layers. By integrating a Self-calibrating Selection Network (SCSN), these layers enhance the model's ability to perceive subtle cracks and highly similar damage patterns. Compared to the original design, the improved Stem and downsampling layers can extract and represent pavement damage features more accurately, thereby improving the accuracy of subsequent classification tasks. Meanwhile, module (b) integrates a Spatial Channel Coordinate Attention (SCCA) module on top of the original Gated Convolutional Neural Network Block (GCNN). This improvement enhances the model's robustness in handling challenging conditions such as uneven lighting and overlapping shadow damage. By effectively reducing background noise interference, the model can better isolate damage features, further improving classification accuracy.

[0042] The SCCA-MambaOut model operates as follows. After an image is input into the model, it first undergoes an initial feature extraction stage, including normalization, regularization, and preliminary preprocessing. In this stage, Spatial Channel Coordinate Attention Block (SCCA) is applied to further refine the extracted features. This refinement process focuses on spatial channel enhancement, spatial information fusion, and background noise reduction, collectively referred to as Stage 1. Subsequently, the model performs three consecutive downsampling operations, employing spatial information fusion and background noise suppression strategies at each stage. These processes are defined as Stages 2 to 4, ultimately leading to the classification layer and completing the final image classification task.

[0043] Feature filtering is performed based on the importance of spatial information of each channel in the road surface image to be identified, in order to remove spatially redundant features and obtain an initial processed image. Specifically, this includes: The importance of spatial information for each channel is determined, and the weight of each channel is determined using group normalization, which can be expressed by formulas (1) and (2). For the spatial information of the channel, For the first i The initial processed image of the stage, This indicates that a gated convolutional neural network block, incorporating spatial channel coordinate attention, is used for processing. This indicates that the self-calibration network is selected for processing. The spatial channel coordinates require processing blocks. Stage Index i Corresponding to the subsequent stages 2-4, 1< i ≤4.

[0044] (1) (2) Based on channel importance weights, the contribution of different channels to spatial information is quantified, yielding the relative importance of each channel in encoding spatial details. Channels with relative importance higher than the spatial weight threshold are retained, while features with relative importance lower than the spatial weight threshold are discarded, resulting in the initial processed image.

[0045] The pre-processed image is divided into a primary channel image and a secondary channel image. Adaptive weighting is then applied to the primary and secondary channel images to remove temporal redundancy features, resulting in a feature map. Specifically, this includes: The channel reconstruction unit includes the channel segmentation stage, the channel conversion stage, and the channel fusion stage.

[0046] During the channel segmentation stage, the initial processed image is divided into a main channel image and a secondary channel image according to a set ratio.

[0047] During the channel conversion stage, convolution operations of different scales are performed on the two channels to obtain a feature-enhanced main channel image and a feature-enhanced secondary channel image.

[0048] In the channel fusion stage, the feature-enhanced main channel image and the feature-enhanced secondary channel image are adaptively weighted according to attention coefficients to remove temporally redundant features in the road surface image to be identified, resulting in a feature map: ; ; ; ; in, For the initial image processing, Main channel image, For secondary channel images, For feature-enhanced main channel images, For feature-enhanced secondary channel images, For feature maps, This indicates a splitting operation. This represents grouped weighted convolution. This represents pyramid-weighted convolution. This represents the Softmax activation function.

[0049] 2. Self-calibration selection network (SCSN).

[0050] Compared to other image classification tasks, road damage detection and classification face unique challenges, emphasizing computational efficiency and lightweight model design. This is crucial for real-time detection of road damage when deployed on in-vehicle systems and mobile devices. By promptly detecting and addressing road defects, such systems can help reduce safety risks and improve road maintenance efficiency. To balance efficiency and accuracy, this study introduces a Self-calibrating Selection Network (SCSN) into the framework. Specifically, SCSN optimizes the original Stem layer and downsampling layer, enhancing the model's ability to extract fine-grained damage features while maintaining lightweight and high performance. The original Stem layer in MambaOut consists of two Conv2D layers, two normalized layers, and a GELU activation function. Based on this structure, the original Stem layer is used as the main channel, while the Self-calibrating Selection Network (SCSN) is introduced as a parallel branch. The outputs of the two paths are then fused together through residual connections to form the newly designed SCSN-Stem layer. This design retains the computational efficiency of MambaOut while leveraging SCSN to refine the feature extraction process, ultimately improving the accuracy and robustness of road damage detection. Similarly, when optimizing the downsampling layer, the original depthwise separable convolutional layer and normalization layer are retained as the main channels, while SCSN is used as a parallel branch. The outputs are fused using residual connections to form the SCSN-downsampling layer. This improvement maintains the computational efficiency of MambaOut while significantly enhancing the model's ability to capture fine-grained local details. Therefore, the model exhibits stronger robustness in handling challenging lighting conditions and visually similar crack patterns, thereby improving its ability to identify complex road damage conditions.

[0051] The detailed structure of the proposed Self-calibration Selection Network (SCSN) is as follows: Figure 2As shown, its core components include the Spatial and Channel Reconstruction Convolution (SCConv) and the Large Selective Kernel (LSK). Upon receiving the input image of road damage, SCSN first processes it using SCConv, an efficient convolutional module designed to reduce redundant information, enhance feature representation, and reduce computational complexity, thus making the model more lightweight. Traditional CNN architectures often suffer from spatial and channel redundancy, where certain feature information is repeatedly encoded, leading to increased computational complexity and reduced generalization ability. SCConv addresses this problem through a two-stage reconstruction mechanism consisting of a Spatial Reconstruction Unit (SRU) and a Channel Reconstruction Unit (CRU). First, spatial information is optimized, then channel information is reconstructed and integrated, achieving efficient feature extraction. Spatial redundancy is particularly pronounced in CNN-based feature extraction, especially in large-scale feature maps where some regions may contain irrelevant or highly similar information. This problem is particularly critical in road damage classification because the damaged area occupies a very small proportion of the entire image, with the majority being undamaged road surface. Damage features, such as cracks and potholes, typically exhibit elongated structures. If a standard CNN architecture is directly applied for feature extraction, background information may interfere with meaningful damage features, ultimately affecting detection accuracy. Introducing SRU and CRU to reconstruct spatial information reduces background redundancy, while channel reconstruction enhances feature representation, enabling the model to focus more precisely on damaged areas. This strengthens the feature hierarchy and representational power, effectively mitigating the adverse effects of background noise and feature redundancy on classification results.

[0052] After processing by SCConv, the feature map is fed into LSK, a module originally designed for remote sensing object detection to address the challenge of varying receptive field requirements for different targets. The core of LSK is to dynamically adjust the convolutional receptive field through a spatial selection mechanism, thereby improving detection performance. In road damage detection and classification, the low damage ratio necessitates dynamic adjustment of the receptive field. Furthermore, when cracks appear near zebra crossings or lane boundaries, the edges of the damaged area become less distinct and blend into the background, posing another key challenge for accurate classification. In such cases, MambaOut may struggle to accurately distinguish between damaged and intact road surfaces. LSK addresses this issue by combining large convolutional kernels with an adaptive selection mechanism, enabling the network to dynamically adjust its receptive field for different targets, thus improving detection and classification accuracy. First, multi-scale large convolutional kernels extract features using depthwise separable convolutions, gradually increasing the kernel size to simulate a wider receptive field, rather than directly using fixed large kernels. This approach prevents excessive computational complexity while ensuring the network effectively captures the spatial background. By progressively expanding the receptive field, the LSK module improves classification accuracy while maintaining computational efficiency, making the model more adaptable to damage features at different scales in road surface detection. This improvement enhances the model's robustness and generalization ability.

[0053] Depthwise separable convolution operations are used to extract features from the feature map, resulting in multi-scale features: (3) In the second step of the LSK module, the model's adaptability to different types of pavement damage is further enhanced, resulting in more accurate detection and classification. The core objective of this mechanism is to dynamically adjust the receptive field using spatial attention, enabling the model to adapt to damage areas of varying sizes and shapes. First, global spatial information is calculated by performing average pooling and max pooling on the features extracted from all convolutional kernels. This operation ensures effective capture of both global and local information. Next, a spatially weighted calculation method is used to fuse the extracted multi-scale features, allowing the model to adaptively adjust its focus based on the characteristics of the damage pattern. Finally, this process produces the optimal feature representation, thereby enhancing the model's ability to distinguish fine-grained damage features from background noise.

[0054] Spatial weighting operations are used to fuse multi-scale features. These operations include parallel average pooling and max pooling operations to obtain the initial features. (4) (Sigmoid( )) (5) in, For multi-scale features, This is the output of multi-scale features after average pooling. This is the output of multi-scale features after max pooling. As initial features, This represents a depthwise separable convolution operation, and Sigmoid represents the Sigmoid activation function. This indicates the average pooling operation. This indicates a max pooling operation.

[0055] By introducing SCSN, redundant features are effectively reduced, input data is improved, computational overhead is lowered, and the ability to process pavement damage images is enhanced. During feature selection, SCSN can dynamically select the optimal receptive field, specifically addressing key challenges such as low damage-to-background ratio and high damage similarity. These improvements make the model more suitable for complex pavement damage detection tasks. Building on this optimization, the improved SCSN-Stem layer and SCSN-downsampling layer retain the high computational efficiency and accuracy of MambaOut while providing targeted enhancements for pavement damage images. These improvements enable the model to achieve outstanding performance in pavement damage detection and classification tasks. This optimization also significantly enhances the model's generalization ability, enabling it to better classify pavement damage in complex environments.

[0056] 3. Spatial Channel Coordinates (SCCA) Note the area.

[0057] MambaOut, an improved version of the Mamba model, removes the State Space Model (SSM) from its original architecture to better adapt to image classification tasks. The SSM structure is not essential for general image classification tasks. To address the specific characteristics of road damage images, including uneven illumination, overlapping shadows and damaged areas, low damage-to-background ratios, and high similarity between damage types, MambaOut underwent targeted optimization. In this improvement, the Spatial Channel Coordinate Attention Block (SCCA) is integrated into the gated CNN block architecture of MambaOut. SCCA is added as a separate branch and then fused with the original gated CNN block using residual connections. This design preserves the computational efficiency of the original model while enhancing its adaptability to road damage images. Experimental results show that this improvement significantly improves the classification accuracy of road damage detection and enhances the model's generalization ability. The optimized MambaOut model can more accurately distinguish various types of road damage, ensuring reliable performance in practical applications. Furthermore, this optimization retains the model's lightweight nature, further enhancing robustness while maintaining low computational complexity. Therefore, the improved SCCA-MambaOut model is well-suited for real-time deployment on in-vehicle and mobile devices, providing more accurate and efficient technical support for the maintenance of intelligent transportation infrastructure.

[0058] The specific structure of the Spatial Channel Coordinates Attention Module (SCCA) is as follows: Figure 3 As shown, this module consists of three key parts: Spatial Group-wise Enhancement (SGE), Partial Convolution, and CoordAttention (CoordAtt). Each sub-module optimizes feature representations at different levels, thereby enhancing the model's adaptability and robustness in road damage detection. The SGE module primarily reduces background noise interference and improves the stability of feature learning. It enhances the model's ability to distinguish key features, especially when dealing with uneven lighting, overlapping shadows, and damaged areas. The Partial Convolution module optimizes computational efficiency by reducing redundant computations, ensuring that the model maintains a lightweight design without incurring excessive computational costs. The CoordAtt module employs a coordinate attention mechanism, capturing channel dependencies while preserving spatial location information. When the input feature map enters the SCCA module, it first passes through the SGE module, where channel grouping is applied.

[0059] When the input feature map enters the SCCA module, it first passes through the SGE module, where channels are grouped so that different parts of the network can focus on specific types of features, enhancing their performance in spatially relevant regions. This improves the model's ability to perceive different crack types. Next, global statistical feature calculation is performed. Within each channel group, the model calculates the global average feature across all spatial locations. This statistical information serves as a baseline for subsequent spatial attention calculations. Then, the model performs spatial enhancement processing, generating attention factors through normalization and adaptive adjustment. These attention factors are applied to the original feature map, reweighting key regions to enhance key features while suppressing background noise. The mathematical expression for this process is shown in Equation (6).

[0060] Sigmoid(Norm(PDP(GAP( )))) (6) After feature enhancement using the SGE module, the model's feature representation capability is significantly improved, while background noise is effectively suppressed, laying a solid foundation for subsequent feature processing. However, since SGE does not involve computation, additional mechanisms are needed to further optimize feature extraction without introducing excessive computational overhead. By utilizing feature redundancy, this method improves computational efficiency without increasing complexity.

[0061] Furthermore, some convolutions employ local feature updates, enabling the network to focus more effectively on important feature regions while further reducing redundant computations. This strategy significantly improves the lightweight performance of the model, ensuring that it maintains high classification accuracy while still possessing efficient computational power. The mathematical expression for this process is shown in Equation (7).

[0062] PC )+ (7) To address the issue of low damage-to-background ratios in road damage images—where the damaged area occupies only a small portion of the entire image—a coordinate attention module for feature enhancement is introduced. CoordAtt combines inter-channel feature relationships while preserving spatial location information, enabling the model to more accurately identify damaged areas. Compared to traditional attention mechanisms, CoordAtt is more lightweight and does not significantly increase computational costs. Furthermore, it leverages the superior performance of attention mechanisms in image classification tasks, ensuring the model maintains high computational efficiency while improving feature extraction capabilities. In traditional attention computation, global pooling operations often lead to the loss of spatial information. To mitigate this issue, CoordAtt performs two independent 1D pooling operations along the horizontal (X-axis) and vertical (Y-axis) directions, ensuring sufficient preservation of coordinate information. This allows the model to more effectively simulate spatial relationships and focus more on damaged areas. After pooling along each direction, the extracted one-dimensional feature map undergoes further one-dimensional pooling before being fused through a 1×1 convolutional layer. Next, two independent 1×1 convolutional layers compute attention weights along the X and Y directions. Finally, the attention computation mechanism generates coordinate attention features to further optimize the representation of the damaged region. The mathematical expression for this process is shown in Equation (8).

[0063] Sigmoid(Conv2d(Conv2d( )+ )) (8) in, For input features, As an intermediate feature, For intermediate processing features, The feature vector in the horizontal direction, The feature vectors in the vertical direction, For the output features, GAP represents global average pooling, PDP represents probability density prediction, Norm represents normalization, Sigmoid represents the Sigmoid activation function, PC represents principal component features, and Conv2d represents two-dimensional convolution.

[0064] After processing with Spatial Channel Coordinate Attention Blocks (SCCA), the extracted features maintain computational efficiency while significantly enhancing spatial relationship modeling capabilities. This improvement effectively reduces background noise interference and enhances feature recognition capabilities. Furthermore, by incorporating a gating mechanism, CoordAtt further optimizes channel feature selection, ensuring key features stand out while suppressing irrelevant information. This improvement enhances the model's overall discriminative ability, making it more suitable for road damage detection and classification tasks.

[0065] 4. Experiment.

[0066] All experiments in this study were conducted on a Windows 11 (64-bit) operating system. The experimental setup included an Intel Core i5-12600KF CPU and an NVIDIA RTX 4070 Ti SUPER (16GB) GPU. The model was implemented in Python, developed using PyTorch, and trained using CUDA 11.8 to improve computational efficiency.

[0067] 4.1 Dataset.

[0068] This study uses the publicly available CQU-BPDD dataset for model training and testing. This dataset consists of 60,056 asphalt pavement images, divided into a training set of 10,137 damaged pavement images and a test set of 49,919 normal pavement images. The training set includes 5,137 damaged pavement images (covering all types of defects) and 5,000 normal pavement images. The test set contains 11,589 damaged pavement images and 38,330 normal pavement images. The CQU-BPDD dataset covers seven common pavement damage types and one normal pavement category: transverse cracks, block cracks, alligator cracks, sealing cracks, longitudinal cracks, wrinkles, and repaired pavement. Example images for different categories are shown below. Figure 4 As shown, this dataset presents significant challenges due to its low damage-to-background ratio, overlap of cracks and shadows, damage occurring near lane markings, and uneven lighting conditions. These complexities make it an ideal benchmark for evaluating the robustness of models in real-world pavement damage detection tasks.

[0069] To further evaluate the robustness and generality of the SCCA-MambaOut model, an additional benchmark dataset, Crack500, was added to the ablation study. Crack500 is a classic benchmark dataset widely used in pavement crack detection research. It contains 3363 pavement crack images, divided into training and testing parts in a 7:3 ratio: 2354 images for training and 1005 images for testing. This dataset partitioning ensures a comprehensive evaluation of the model's stability across different image samples, thereby improving the reliability of the experimental results. In all experiments, the training parameters were set as follows: batch size: 50; initial learning rate: 0.00005. To comprehensively evaluate the model's performance on the classification task, the following main evaluation metrics were used: accuracy, F1 score, and precision. These metrics ensure a comprehensive assessment of the model's recognition ability and robustness in pavement damage classification tasks.

[0070] 4.2 Experimental results.

[0071] To comprehensively evaluate the performance of the SCCA-MambaOut model, comparative experiments were conducted on the CQU-BPDD dataset against several mainstream image classification models. The models compared included AlexNet, VGG16, MobileNetV2, MobileNetV3, DDACDN, FFVT, TransFG, VisionTransformer, and SwinTransformer. Detailed comparison results are shown in Table 1. Compared to the baseline model MambaOut, the SCCA-MambaOut model achieved a 2% improvement in accuracy, a 0.5% improvement in precision, and a 1.2% improvement in F1 score. These results demonstrate that the introduction of the Self-calibrating Selection Network (SCSN) and the Spatial Channel Coordinate Attention Block (SCCA) significantly improves classification accuracy while maintaining high computational efficiency within the same architectural framework.

[0072] Table 1: Comparison of experimental results on the CQU-BPDD dataset AlexNet, a groundbreaking deep neural network architecture, marked a breakthrough in deep learning for computer vision. It consists of eight convolutional layers and three fully connected layers, employing ReLU activation and Dropout regularization to mitigate overfitting. Furthermore, AlexNet utilizes large convolutional kernels and pooling layers for feature extraction, while data augmentation enhances its generalization ability. However, AlexNet's structure is relatively complex, computationally expensive, and requires large training datasets and computational resources. Moreover, its deep convolutional architecture is prone to overfitting, especially with insufficient training data. The large number of convolutional layers also increases its dependence on large datasets, limiting computational efficiency and inference speed. In contrast, VGG16 employs a simpler, more unified design, featuring small convolutional kernels and max-pooling layers. While maintaining strong feature extraction capabilities, it progressively deepens the network to improve classification performance. Due to its deep structure and small convolutional kernels, VGG16 is particularly effective at extracting high-level visual features, making it well-suited for high-precision visual tasks. However, its main drawback is its high computational cost. The deep architecture and large number of parameters increase the demand for training resources and storage capacity. Furthermore, its fully connected layers contain millions of parameters, resulting in a massive model size and significantly increasing computational overhead when processing large-scale images, posing challenges in terms of efficiency and resource consumption. Compared to AlexNet and VGG16, SCCA-MambaOut employs a lightweight modular structure, effectively reducing computational complexity, while integrating a self-calibrating selection network (SCSN) and a spatial channel coordinate attention block (SCCA), maintaining a balance between computational efficiency and classification accuracy. Similarly, compared to VGG16, SCCA-MambaOut achieves a 15.37% improvement in accuracy, a 5.23% improvement in precision, and an 11.72% improvement in F1 score. These improvements demonstrate that SCCA-MambaOut, through optimized model design, improves classification performance while maintaining high efficiency, making it more suitable for resource-constrained environments.

[0073] Both FFVT and Swin-T are Transformer-based models that employ various architectural optimizations to improve computational efficiency, accuracy, or specific performance aspects. ViT uses a standard transformer architecture, including self-attention mechanisms and feedforward neural networks. Its main advantage lies in its ability to capture global information, focusing on any region in the image while effectively modeling long-range dependencies. However, compared to CNNs, ViT performs poorly on small datasets because Transformers require larger datasets for proper training to ensure generalization. FFVT focuses on optimizing computational efficiency for visual tasks. It incorporates various architectural improvements, such as reducing computational complexity and minimizing memory consumption, thus ensuring faster inference performance. However, FFVT still struggles with small datasets because Transformer-based models typically require large amounts of data, while CNN architectures are generally more data-efficient when trained on small datasets. Swin Transformer introduces a window attention mechanism, dividing the input image into multiple local windows and performing self-attention computation within each window. Compared to ViT's global attention mechanism, Swin Transformer significantly reduces computational complexity while improving training and inference efficiency. Despite these advantages, the training process for Swin-T remains complex, especially on large-scale datasets, requiring longer training times and more computational resources. Unlike Transformer-based architectures that rely on computationally expensive self-attention mechanisms, the SCCA-MambaOut model employs a different approach. Instead of utilizing the highly complex Transformer attention mechanism, SCCA-MambaOut introduces a lightweight coordinate attention (CoordAtt) module into a CNN-based framework. This design maintains lower computational complexity while ensuring higher classification accuracy. Compared to Swin-T and FFVT, SCCA-MambaOut avoids an explosive increase in computational complexity and parameter number while still achieving significant improvements in classification performance. This makes SCCA-MambaOut a more practical and efficient solution for road damage image classification tasks, achieving an optimal balance between efficiency and accuracy.

[0074] Comparative experiments were also conducted with CNN-based architectures, particularly MobileNet and DDACDN. MobileNet is an efficient and lightweight CNN model designed specifically for mobile devices and embedded systems. Its core optimization lies in the use of depthwise separable convolutions, which significantly reduces computational complexity, allowing MobileNet to maintain high performance even in environments with limited computational resources. In contrast, DDACDN is a CNN-based image denoising model that integrates a dual-domain attention mechanism. By leveraging multi-scale feature learning and attention mechanisms, DDACDN demonstrates excellent denoising capabilities, especially in complex noisy environments. Although SCCA-MambaOut is also a CNN-based model, it is specifically optimized for road damage classification tasks. To address the challenges of low damage-to-background ratios, high similarity between crack types, and inherent classification difficulty, a self-calibrating selection network (SCSN) was introduced to enhance the model's ability to focus on subtle damage features. Furthermore, a Spatial Channel Coordinate Attention Block (SCCA) was added to optimize the model's response to uneven lighting conditions and overlapping shadows in damaged areas, which pose a significant challenge to classification accuracy. These improvements ensure that SCCA-MambaOut can more accurately identify road damage types while maintaining high computational efficiency. Experimental results demonstrate that SCCA-MambaOut significantly improves success rate, accuracy, and F1 score in pavement damage classification tasks. Compared to traditional CNNs and Transformer-based models, the optimized method maintains high recognition accuracy and more stable classification performance even under limited computational resources. These findings further validate the superiority of SCCA-MambaOut in real-world applications, making it an efficient and practical solution for pavement damage classification tasks.

[0075] Visual Mamba employs a bidirectional state-space model (SSM) to encode image sequences, simultaneously capturing forward and backward contextual information. This mechanism enhances the model's ability to understand the global structure of an image, making it particularly effective in long-sequence tasks such as object detection and semantic segmentation, where long-distance dependency modeling is crucial for improving overall visual task performance. However, research on MambaOut reveals that image classification tasks are fundamentally different from NLP, object detection, and semantic segmentation. These tasks inherently involve long-sequence features and require building long-term dependency models, properties that image classification lacks. Therefore, by removing the SSM component, MambaOut is better suited for image classification tasks, including road damage classification. Building upon MambaOut, the proposed SCCA-MambaOut introduces further optimizations to address the unique challenges of road damage classification. Its architecture integrates a self-calibrating selected convolutional network (SCSN) specifically designed to address the low damage-to-background ratio and high similarity issues between different crack types, ensuring accurate classification even when the damaged area occupies only a small portion of the image and exhibits similar morphological features. Furthermore, a Spatial Channel Coordinate Attention Block (SCCA) was introduced to effectively address issues related to uneven lighting conditions and overlapping shadows, which severely impact damage detection accuracy. Through these optimizations, SCCA-MambaOut outperforms both Visual Mamba and MambaOut on the CQU-BPDD pavement damage dataset, showing improvements in success rate, accuracy, and F1 score. Experimental results demonstrate that compared to Visual Mamba, SCCA-MambaOut improves success rate by 2.27%, accuracy by 0.96%, and F1 score by 1.63%. Similarly, compared to the baseline Mamba model, SCCA-MambaOut improves success rate by 2%, accuracy by 0.5%, and F1 score by 1.2%. These findings further confirm the effectiveness of SCCA-MambaOut in pavement damage classification tasks, demonstrating its robustness and accuracy in handling real-world pavement assessment challenges. To further illustrate the superior performance of SCCA-MambaOut, a visualization heatmap was generated using Grad-CAM, as shown below. Figure 5As shown, the decision-making process of SCCA-MambaOut and [other models] is intuitively explained. The heatmap clearly shows the specific image regions that each model focuses on during classification, thus revealing the basic decision-making mechanism and feature extraction strategy of the model. When applied to low-damage pavement images, the improved SCCA-MambaOut exhibits higher accuracy in localization and feature extraction. Compared to [other models], SCCA-MambaOut can more accurately identify damaged areas and perform more refined and detailed feature extraction. This enhanced capability allows SCCA-MambaOut to capture smaller and more subtle damaged areas, thus more accurately reflecting pavement conditions. The improved ability to locate and analyze damaged areas enhances the stability and reliability of the classification task. Experimental results further confirm that SCCA-MambaOut provides an efficient and accurate solution for pavement damage classification, offering a novel method that can effectively address the challenges in practical pavement assessment.

[0076] 4.3 Ablation studies.

[0077] The main objective of this ablation study is to verify the different impacts of the introduced Self-calibrating Selective Convolutional Network (SCSN) and Spatial Channel Coordinate Attention Block (SCCA) on the pavement damage classification task, and their contribution to the overall model performance. To this end, two variant models were constructed: SC-MambaOut, a MambaOut version containing only the Self-calibrating Selective Convolutional Network (SCSN), and CA-MambaOut, which contains only the Spatial Channel Coordinate Attention Block (SCCA). This experimental setup aims to analyze the individual effects of these two modules on model performance and determine their role in improving classification accuracy, robustness, and computational efficiency.

[0078] Experiments were conducted on the large-scale CQU-BPDD dataset and the small-scale Crack500 dataset to evaluate the impact of each module on the final classification performance, and to verify the robustness and generality of the model on datasets of different sizes. Detailed experimental results are shown in Table 2. The results show that SC-MambaOut and CA-MambaOut achieve improvements in accuracy, precision, and F1 score compared to the Mamba baseline model. This indicates that the Self-calibrated Selective Convolutional Network (SCSN) and the Spatial Channel Coordinate Attention Block (SCCA) play complementary roles in improving classification accuracy and feature extraction capabilities. Furthermore, experimental results on the small Crack500 dataset further validate the robustness of SCCA-MambaOut in scenarios with limited data. Even with a relatively small number of samples, the model maintains high classification accuracy and stability. This finding demonstrates that the proposed improvements are effective not only for large-scale datasets but also for datasets with limited data, thereby further enhancing the model's generalization ability. Ablation studies have confirmed that the combined effect of SCSN and SCCA significantly improves the performance of SCCA-MambaOut in pavement damage classification tasks, while ensuring its adaptability to different dataset sizes.

[0079] Table 2: Ablation experimental results based on the CQU-BPDD and Crack500 datasets 5. Conclusion.

[0080] A lightweight deep learning model, SCCA-MambaOut, is proposed and optimized for pavement damage classification. Compared to traditional CNNs and transformer-based architectures, this model, based on the Mamba, introduces a self-calibrating selection convolutional network (SCSN) and a spatial channel coordinate attention block (SCCA) to address key challenges in pavement damage detection. Specifically, SCSN improves classification performance under low damage rates, while SCCA alleviates problems caused by uneven illumination and overlap between damaged areas and background interference. Experimental results show that SCCA-MambaOut significantly improves success rate, accuracy, and F1 score, outperforming the benchmark Mamba by 2%, 0.5%, and 1.2%, respectively. Furthermore, compared to transformer-based models such as Swin-T and FFVT, and traditional CNN models such as AlexNet, VGG16, and MobileNetV3, SCCA-MambaOut achieves an excellent balance between computational efficiency and classification accuracy, demonstrating strong adaptability and practical value. Through ablation studies, extensive testing was conducted on the large-scale CQU-BPDD dataset and the small-scale Crack500 dataset, validating the independent contributions of SCSN and SCCA. These experiments further confirm that SCCA-MambaOut exhibits strong robustness and generalization ability on datasets of different sizes. Furthermore, Grad-CAM visualization analysis shows that SCCA-MambaOut effectively reduces background noise interference while capturing damaged areas with higher accuracy, thereby enhancing its interpretability and practical applicability in pavement damage detection tasks.

[0081] In summary, SCCA-MambaOut provides an efficient and accurate pavement damage detection solution, offering valuable technical support for intelligent transportation systems and road maintenance applications. Future research will further explore lightweight model architectures to enhance the model's applicability to edge devices, while integrating multimodal information sources such as LiDAR and thermal imaging data to improve the overall performance of damage detection and contribute to the development of intelligent transportation systems.

[0082] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.

Claims

1. A pavement damage detection method based on the Mamba simplified model, characterized in that, include: Acquire the image of the road surface to be identified; Feature filtering is performed based on the importance of spatial information of each channel in the road surface image to be identified, in order to remove spatially redundant features and obtain the initial processed image; The pre-processed image is divided into a main channel image and a secondary channel image, and adaptive weighting is applied to the main channel image and the secondary channel image to remove temporal redundancy features and obtain a feature map. Depthwise separable convolution operations are used to extract features from the feature map to obtain multi-scale features; Spatial weighting is used to fuse multi-scale features to obtain initial features; The initial features are fused multiple times with spatial information. The output features of the previous spatial information fusion are downsampled and the output features after downsampling are used as the input features of the next spatial information fusion. The output features of the last spatial information fusion are used as the fused features. Each spatial information fusion step includes: grouping the input features according to their channels and determining the global average features of all spatial locations within each channel group to obtain an attention factor; reweighting the input features using the attention factor to obtain intermediate features; performing convolution operations on the intermediate features to update the local features within the intermediate features to obtain intermediate processed features; performing one-dimensional convolution operations on the intermediate processed features in both the horizontal and vertical directions to obtain feature vectors in the two directions, and then fusing the feature vectors in the two directions to obtain the output features. Damage is classified based on fusion characteristics to obtain pavement damage detection results.

2. The pavement damage detection method based on the Mamba simplified model as described in claim 1, characterized in that, The road surface damage detection results are obtained by improving the Mamba simplified model. The improved Mamba simplified model includes: an original dry layer with a self-calibration selection network, multiple feature processing modules and a classification layer connected in sequence. Each feature processing module includes: a gated convolutional neural network block with spatial channel coordinate attention block and a downsampling layer with a self-calibration selection network connected in sequence. The original dry layer of the self-calibrating selection network includes: a spatial and channel reconstruction convolutional module and a large selection kernel connected in sequence; The spatial and channel reconstruction convolutional module includes: a spatial reconstruction unit and a channel reconstruction unit connected in sequence, wherein the channel reconstruction unit includes a channel segmentation stage, a channel transformation stage, and a channel fusion stage; the large selection kernel includes: a multi-scale large convolutional kernel and a spatial selection mechanism connected in sequence. The gated convolutional neural network block that introduces spatial channel coordinate attention includes: a spatial grouping enhancement module, a partial convolution module, and a coordinate attention module connected in sequence.

3. The pavement damage detection method based on the Mamba simplified model as described in claim 2, characterized in that, Training the improved Mamba simplified model specifically includes: Obtain historical road surface images and corresponding historical road surface damage types; By inputting historical road surface images into the improved Mamba simplified model, the predicted road surface damage type can be obtained. With the goal of minimizing the deviation between the predicted pavement damage type and the historical pavement damage type, the improved Mamba simplified model is trained to obtain the trained improved Mamba simplified model.

4. The pavement damage detection method based on the Mamba simplified model as described in claim 2, characterized in that, Constructing the original dry layer with the self-calibration selection network introduced specifically includes: using the self-calibration selection network as a parallel branch of the original dry layer, and connecting the output of the self-calibration selection network and the output of the original dry layer through residual connection to obtain the original dry layer with the self-calibration selection network introduced. Constructing the downsampling layer with the self-calibration selection network specifically includes: using the self-calibration selection network as a parallel branch of the downsampling layer, and connecting the output of the self-calibration selection network and the output of the downsampling layer through residual connections to obtain the downsampling layer with the self-calibration selection network. Constructing the gated convolutional neural network block with the spatial channel coordinate attention block specifically includes: using the spatial channel coordinate attention block as a parallel branch of the gated convolutional neural network block, and connecting the output of the spatial channel coordinate attention block and the output of the gated convolutional neural network block through residual connections to obtain the gated convolutional neural network block with the spatial channel coordinate attention block.

5. The pavement damage detection method based on the Mamba simplified model as described in claim 4, characterized in that, The step of feature filtering based on the importance of spatial information of each channel in the road surface image to be identified, in order to remove spatially redundant features and obtain an initial processed image, specifically includes: The importance of spatial information for each channel is determined based on the following formula, and group normalization is used to determine the weight of each channel: ; ; in, For the spatial information of the channel, For the first i The initial processed image of the stage, This indicates that a gated convolutional neural network block, incorporating spatial channel coordinate attention, is used for processing. This indicates that the self-calibration network is selected for processing. The spatial channel coordinates require special processing. i For stage indexing; Based on the channel importance weight, the contribution of different channels to spatial information is quantified to obtain the relative importance of each channel in encoding spatial details; Channels with relative importance higher than the spatial weight threshold are retained, while features with relative importance lower than the spatial weight threshold are discarded to obtain the initial processed image.

6. The pavement damage detection method based on the Mamba simplified model as described in claim 1, characterized in that, The process of dividing the pre-processed image into a primary channel image and a secondary channel image, and then adaptively weighting the primary and secondary channel images to remove temporal redundancy features and obtain a feature map, specifically includes: The initial processed image is divided into a main channel image and a secondary channel image according to a set ratio; Convolution operations at different scales are performed on the two channels to obtain a feature-enhanced main channel image and a feature-enhanced secondary channel image; The main channel image and the secondary channel image of feature enhancement are adaptively weighted according to the attention coefficients based on the following formula to remove temporally redundant features in the road surface image to be identified, thus obtaining the feature map: ; ; ; ; in, For the initial image processing, Main channel image, For secondary channel images, For feature-enhanced main channel images, For feature-enhanced secondary channel images, For feature maps, This indicates a splitting operation. This represents grouped weighted convolution. This represents pyramid-weighted convolution. This represents the Softmax activation function.

7. The pavement damage detection method based on the Mamba simplified model as described in claim 1, characterized in that, Based on the following formula, depthwise separable convolution operations are used to extract features from the feature map, resulting in multi-scale features: ; Based on the following formula, a spatial weighting operation is used to fuse multi-scale features. The spatial weighting operation includes parallel average pooling and max pooling operations to obtain the initial features: ; (Sigmoid( )); in, For multi-scale features, This is the output of multi-scale features after average pooling. This is the output of multi-scale features after max pooling. As initial features, This represents a depthwise separable convolution operation, and Sigmoid represents the Sigmoid activation function. This indicates the average pooling operation. This indicates a max pooling operation.

8. The pavement damage detection method based on the Mamba simplified model as described in claim 1, characterized in that, The intermediate features are obtained by reweighting the input features using an attention factor based on the following formula: Sigmoid(Norm(PDP(GAP( )))); The intermediate features are convolved using the following formula to update local features within the intermediate features, resulting in intermediate processed features: PC( )+ ; Based on the following formula, two independent one-dimensional convolution operations are performed on the intermediate processing features in the horizontal and vertical directions respectively to obtain the output features: Sigmoid(Conv2d(Conv2d( )+ )); in, For input features, As an intermediate feature, For intermediate processing features, The feature vector in the horizontal direction, The feature vectors in the vertical direction, For the output features, GAP represents global average pooling, PDP represents probability density prediction, Norm represents normalization, Sigmoid represents the Sigmoid activation function, PC represents principal component features, and Conv2d represents two-dimensional convolution.