Pavement disease detection method and system
By improving the YOLOv11 network model, the C3k2_Strip and C3k2_Kat modules were built, and combined with the P2 detection head, the problems of low accuracy and poor efficiency of road surface disease detection are solved, and more efficient road surface disease recognition is achieved.
Patent Information
- Application Number
- CN202510422298.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the road surface disease detection method has problems such as low detection accuracy and poor detection efficiency.
Using the improved YOLOv11 network model, by building the C3k2_Strip module and the C3k2_Kat module, the feature extraction capability and feature representation are enhanced, and combined with the P2 detection head, the network structure is optimized to improve detection accuracy and efficiency.
It significantly improves the feature representation and processing capabilities of road disease detection, improves the recognition speed and efficiency, and is suitable for complex road disease recognition scenarios.
Smart Images

Figure CN120279380A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer vision technology, and particularly relates to a road disease detection method and system. Background Art
[0002] While the scale of traffic infrastructure is continuously expanding, road maintenance work faces new challenges. Restricted by human and financial resources, road maintenance work often struggles to keep up with the pace of infrastructure expansion, resulting in relatively common pavement damage; therefore, it is of great practical significance to deeply study the impact of pavement damage on road traffic safety.
[0003] Currently, existing machine learning-based road disease detection methods mainly include two categories. One is the single-stage detection framework represented by SSD (Single Shot MultiBox Detector) and RetinaNet, and the other is the two-stage detection framework centered around MaskR-CNN and Faster R-CNN; however, the above detection methods are prone to false recognition or missed recognition, resulting in low detection accuracy and poor detection efficiency. Summary of the Invention The purpose of the embodiments of this application is to provide a road disease detection method and system, which can solve the technical problems of low detection accuracy and poor detection efficiency in road disease detection in the prior art.
[0004] To solve the above technical problems, this application is implemented as follows: In a first aspect, the embodiments of this application provide a road disease detection method, and the method includes: Obtain a road surface image dataset, preprocess the road surface image dataset to obtain a target dataset; Construct a YOLOv11 network model, and add a P2 detection head in the detection layer of the YOLOv11 network model; Construct a C3k2_Strip module, and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2_Strip module; Construct a C3k2_Kat module, and replace the C3k2 module in the neck network of the YOLOv11 network model with the C3k2_Kat module to obtain an improved YOLOv11 network model; Train the improved YOLOv11 network model according to a part of the data in the target dataset, and test the improved YOLOv11 network model according to another part of the data; Input the road surface image to be detected into the improved YOLOv11 network model after testing, and output the detection result.
[0005] As an alternative implementation of the first aspect of the present application, the process of the C3k2_Strip module for data processing is as follows: The initial feature vector is input into the C3k2_Strip module, and the chained feature extraction sub-module in the C3k2_Strip module performs chained processing on the initial feature vector to obtain a chained feature map; According to the convolutional sub-module in the C3k2_Strip module, convolutional processing is performed on the initial feature vector to obtain a convolutional feature map; According to the residual connection sub-module in the C3k2_Strip module, residual connection is performed on the chained feature map and the convolutional feature map to obtain a target feature map.
[0006] As an alternative implementation of the first aspect of the present application, the chained feature extraction sub-module performs chained processing on the initial feature vector to obtain a chained feature map; specifically: The initial feature vector is input into the chained attention layer in the chained feature extraction sub-module. The chained attention layer performs feature extraction on the initial feature vector to obtain an initial feature map, and performs a first learnable scaling parameter multiplication process on the initial feature map to obtain a first learnable feature map; According to the first random dropout layer in the chained feature extraction sub-module, random dropout processing is performed on the first learnable feature map, and according to the first residual connection layer in the chained feature extraction sub-module, residual connection processing is performed on the learnable feature map after random dropout processing and the initial feature vector to obtain a residual feature map; According to the chained multi-layer perceptron layer in the chained feature extraction sub-module, processing is performed on the residual feature map to obtain a multi-perceptual feature map, and a second learnable scaling parameter multiplication process is performed on the multi-layer perceptual feature map to obtain a second learnable feature map; According to the second random dropout layer in the chained feature extraction sub-module, random dropout processing is performed on the second learnable feature map, and according to the second residual connection layer in the chained feature extraction sub-module, residual connection processing is performed on the second learnable feature map after random dropout processing and the residual feature connection map to obtain the chained feature map.
[0007] As an alternative implementation of the first aspect of the present application, the chained attention layer performs feature extraction on the initial feature vector to obtain an initial feature map; specifically: Perform 1×1 convolutional processing on the initial feature vector to obtain a first convolutional feature map, and perform feature projection processing on the first convolutional feature map; Perform activation processing on the first convolutional feature map after feature projection processing to obtain a first activation feature map; Perform 1×1 convolution processing on the first activation feature map to obtain a second convolution feature map, and perform feature projection processing on the second convolution feature map; Perform residual connection on the second convolution feature map after the feature projection processing and the initial feature vector to obtain the initial feature map; The chained multi-layer perceptron layer processes the residual feature map to obtain a multi-perception feature map; specifically: Perform 1×1 convolution processing on the residual feature map, and then perform depthwise separable convolution processing on the residual feature map after the 1×1 convolution processing to obtain a third convolution feature map; Perform activation processing on the third convolution feature map to obtain a second activation feature map, and perform random dropout processing on the second activation feature map to obtain a dropped feature map; Perform feature reconstruction processing on the dropped feature map, and perform random dropout processing on the dropped feature map after the feature reconstruction processing to obtain the multi-perception feature map.
[0008] As an optional implementation manner of the first aspect of the present application, the process of the C3k2_Kat module for processing data is as follows: Input the input feature map into the C3k2_Kat module, and the preprocessing sub-module in the C3k2_Kat module performs convolution processing on the input feature map to obtain a fourth convolution feature map; According to the Kat feature extraction sub-module in the C3k2_Kat module, perform deep feature step-by-step extraction processing on the fourth convolution feature map to obtain a deep extraction feature map; According to the residual connection sub-module in the C3k2_Kat module, perform residual connection on the deep extraction feature and the input feature map to obtain an output feature map.
[0009] As an optional implementation manner of the first aspect of the present application, the Kat feature extraction sub-module performs deep feature step-by-step extraction processing on the fourth convolution feature map to obtain a deep extraction feature map; specifically: According to the tensor reshaping layer in the Kat feature extraction sub-module, perform dimension reshaping processing and transpose processing on the fourth convolution feature map to obtain a serialized feature map; According to the Kat attention layer in the Kat feature extraction sub-module, perform attention weighting processing on the serialized feature map to obtain an attention feature map; The first residual connection layer in the Kat feature extraction sub-module performs residual connection processing on the attention feature map and the serialized feature map to obtain a residual attention feature map; The multi-layer perceptron layer in the Kat feature extraction sub-module performs non-linear transformation processing on the residual attention feature map to obtain a local attention feature map; The second residual connection layer in the Kat feature extraction sub-module performs residual connection on the local attention feature map and the residual attention feature map to obtain the depth extraction feature map.
[0010] As an optional implementation manner of the first aspect of the present application, the Kat attention layer performs attention weighting processing on the serialized feature map to obtain an attention feature map; specifically: The Kat attention layer performs linear transformation processing on the serialized feature map to obtain a query feature map, a key feature map, and a value feature map corresponding to the serialized feature map; Perform attention head decomposition and calculation on the query feature map, the key feature map, and the value feature map respectively to obtain multiple corresponding query head features, key head features, and value head features; Perform normalization processing on multiple query head features and key head features, and calculate the feature weights of each query head feature and each key head feature; Perform weighted summation on the value head features according to the feature weights of each query head feature and each key head feature to obtain the attention feature map.
[0011] In a second aspect, an embodiment of the present application provides a road surface disease detection system, and the system includes: An acquisition module: acquire a road surface image data set, and preprocess the road surface image data set to obtain a target data set; A first improvement module: construct a YOLOv11 network model, and add a P2 detection head to the detection layer of the YOLOv11 network model; A second improvement module: construct a C3k2_Strip module, and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2_Strip module; A third improvement module: construct a C3k2_Kat module, and replace the C3k2 module in the neck network of the YOLOv11 network model with the C3k2_Kat module to obtain an improved YOLOv11 network model; A training module: train the improved YOLOv11 network model according to a part of the data in the target data set, and test the improved YOLOv11 network model according to another part of the data; A prediction module: input the road surface image to be detected into the improved YOLOv11 network model after testing, and output a detection result.
[0012] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0013] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0014] In the embodiments of the present application, compared with the prior art, the following technical effects are achieved: (1) The C3k2_Strip module enhances the spatial feature extraction ability of the model by introducing a chain attention mechanism, and solves the problem of poor globality in convolutional features of the traditional C3k2. (2) The C3k2_Kat module can calculate multiple attention heads in parallel through the multi-head self-attention mechanism, solving the problem of sequential calculation bottlenecks in traditional convolutions when processing long sequences and the problem that a single attention head cannot fully and flexibly capture complex features.
[0015] (3) Compared with the unimproved YOLOv11 network model, the improved YOLOv11 network model of the present application has significantly improved in feature representation, processing ability, recognition speed and efficiency, and can be applied to more complex road surface disease recognition scenarios. Description of the Drawings
[0016] Figure 1 is a flowchart of a road surface disease detection method provided by some embodiments of the present application; Figure 2 is a structural diagram of an improved YOLOv11 network model of a road surface disease detection method provided by some embodiments of the present application; Figure 3 is a structural diagram of a C3k2_Strip module of a road surface disease detection method provided by some embodiments of the present application; Figure 4 is a structural diagram of a chain feature extraction sub-module of a road surface disease detection method provided by some embodiments of the present application; Figure 5 is a structural diagram of a chain attention layer of a road surface disease detection method provided by some embodiments of the present application; Figure 6 is a structural diagram of a chain multi-layer perceptron layer of a road surface disease detection method provided by some embodiments of the present application; Figure 7 is a structural diagram of a C3k2_Kat module of a road surface disease detection method provided by some embodiments of the present application; Figure 8 It is a structural diagram of the Kat feature extraction sub-module of a road surface disease detection method provided by some embodiments of the present application; Figure 9 It is a structural diagram of the Kat attention layer of a road surface disease detection method provided by some embodiments of the present application Detailed implementation manners Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0017] The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, the "and / or" in the description and claims indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0018] Next, in conjunction with the accompanying drawings, a road surface disease detection method and system provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.
[0019] Embodiment A road surface disease detection method includes the following steps: S100: Obtain a road surface image dataset, and preprocess the road surface image dataset to obtain a target dataset; It should be noted that S100 specifically is: S110: Obtain a large number of road surface images with diseases to construct a road surface image dataset according to each road surface image; S120: Perform disease annotation on each road surface image in the road surface image dataset, and mark the disease position of each road surface image; S130: Perform data augmentation and enhancement processing on the road surface image dataset after disease annotation to obtain a target dataset.
[0020] Furthermore, the road surface image dataset used in this embodiment is a self - collected dataset and a publicly available network dataset. A mirrorless camera of model Nikon Z9 is used to take 300 road surface images with disease defects on - site. These images are then merged with the publicly available RDD2022 road damage image dataset to obtain the road surface image dataset. Then, the 300 road surface images taken are imported into the X - AnyLabeling annotation tool, and the disease defect areas in these 300 images are marked in the yolo format. The annotation file contains information such as the class number of each disease defect target. Data augmentation methods such as randomly enhancing contrast, adding noise, flipping, and scaling are used to expand the image data and labels to 3000 images, simulating the pictures recognized by the camera in various extreme situations, so as to improve the generalization ability of the training model. Subsequently, the expanded 3000 images are merged with the RDD2022 road damage image dataset to obtain the target dataset.
[0021] S200: Construct a YOLOv11 network model, and add a P2 detection head in the detection layer of the YOLOv11 network model; Furthermore, adding a p2 detection head enables 4 different feature maps to be extracted from the road surface disease images in the network model. The sizes of the 4 feature maps are 120×120, 64×64, 32×32, and 16×16 respectively, which are used to detect road surface disease defect targets of 4 different sizes: tiny, small, medium, and large.
[0022] S300: Construct a C3k2_Strip module and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2_Strip module; It should be noted that the process of the C3k2_Strip module for data processing in S300 is as follows: S310: The initial feature vector is input into the C3k2_Strip module, and the chained feature extraction sub - module in the C3k2_Strip module performs chained processing on the initial feature vector to obtain a chained feature map; S320: According to the convolutional sub - module in the C3k2_Strip module, convolutional processing is performed on the initial feature vector to obtain a convolutional feature map; S330: According to the residual connection sub - module in the C3k2_Strip module, residual connection is performed on the chained feature map and the convolutional feature map to obtain the target feature map.
[0023] Furthermore, the feature extraction module C3k2_Strip module (C3k2_StripBlock) of the feed-forward dual-stream collaborative structure enhances the spatial feature extraction ability of the model by introducing a chained feature extraction sub-module, and solves the problem of poor globality in convolutional features of the traditional C3k2. First, the initial feature vector is split into two paths through the convolution module splitting logic: the main branch undergoes chained processing through the chained feature extraction sub-module to extract features and generate a chained feature map, and the second branch performs basic convolution operations on the initial feature vector to generate a convolutional feature map. Finally, the convolutional feature map and the chained feature map generated by the two branches are subjected to residual connection through the residual connection sub-module, and the channel number is adjusted and then the target feature map is output. Through the collaborative innovation of the direction-sensitive convolution design and the dynamic weight fusion mechanism, this module breaks through the bottleneck of the feature expression ability of the traditional module while maintaining light weight.
[0024] It should be noted that in S310, the chained feature extraction sub-module performs chained processing on the initial feature vector to obtain a chained feature map; specifically: S311: The initial feature vector is input into the chained attention layer in the chained feature extraction sub-module. The chained attention layer extracts features from the initial feature vector to obtain an initial feature map, and multiplies the initial feature map by the first learnable scaling parameter to obtain the first learnable feature map; S312: Randomly discard the first learnable feature map according to the first random dropout layer in the chained feature extraction sub-module, and perform residual connection processing on the randomly discarded learnable feature map and the initial feature vector according to the first residual connection layer in the chained feature extraction sub-module to obtain a residual feature map; S313: Process the residual feature map according to the chained multi-layer perceptron layer in the chained feature extraction sub-module to obtain a multi-percept feature map, and multiply the multi-layer percept feature map by the second learnable scaling parameter to obtain the second learnable feature map; S314: Randomly discard the second learnable feature map according to the second random dropout layer in the chained feature extraction sub-module, and perform residual connection processing on the randomly discarded second learnable feature map and the residual feature connection map according to the second residual connection layer in the chained feature extraction sub-module to obtain a chained feature map.
[0025] Furthermore, the initial feature vector is first input into the chained attention layer. This layer extracts features from the features, extracts important feature information, and generates an initial feature map. Then, the initial feature map is multiplied by the first learnable scaling parameter to obtain the first learnable feature map. This process enables the model to adaptively adjust the importance of the features. The first learnable feature map is subjected to dropout processing to reduce the risk of overfitting. The dropout layer randomly selects a part of the neurons to be "switched off", thereby enhancing the generalization ability of the model. The feature map after dropout processing is subjected to residual connection with the initial feature vector to obtain a residual feature map. Residual connection helps the flow of information and avoids the problem of vanishing gradients in deep networks. The residual feature map is then input into the Strip MLP layer. The Strip MLP layer processes the features, extracts more complex feature patterns, and obtains a multi-perception feature map. The multi-perception feature map is multiplied by the second learnable scaling parameter to obtain the second learnable feature map. The second learnable feature map is subjected to dropout processing to further enhance the robustness of the model. The second learnable feature map after dropout processing is subjected to residual connection with the residual feature map, and finally a chained feature map is obtained. This chained feature map fuses multi-level feature information and has a richer expressive ability.
[0026] It should be noted that in S311, the chained attention layer extracts features from the initial feature vector to obtain an initial feature map; specifically: S3111: Perform 1×1 convolution processing on the initial feature vector to obtain a first convolutional feature map, and perform feature projection processing on the first convolutional feature map; S3112: Perform activation processing on the first convolutional feature map after feature projection processing to obtain a first activation feature map; S3113: Perform 1×1 convolution processing on the first activation feature map to obtain a second convolutional feature map, and perform feature projection processing on the second convolutional feature map; S3114: Perform residual connection on the second convolutional feature map after feature projection processing and the initial feature vector to obtain an initial feature map; Furthermore, first, the initial feature vector is subjected to 1×1 convolution processing through a 1x1 convolutional layer, enabling the initial feature vector to undergo a linear transformation to generate the first convolutional feature map. 1×1 convolution can mix information from different channels without changing the spatial resolution, enhancing the feature expression ability with fewer convolutional parameters and higher computational efficiency. The first convolutional feature map is subjected to feature projection processing, and then the GELU activation function is applied to the first convolutional feature map after feature projection processing for activation processing to generate the first activation feature map, enabling the model to learn more complex features. Then, a second 1×1 convolution processing is performed on the first activation feature map through a 1×1 convolutional layer, enabling the first activation feature map to undergo a linear transformation to obtain the second convolutional feature map. Feature projection processing is performed on the second convolutional feature map to further adjust the feature distribution, enhance useful information, suppress noise, and flexibly adjust the number of output channels to meet the requirements of subsequent modules. Finally, the initial feature vector and the second convolutional feature map after feature projection processing are connected through a residual connection to obtain and output the initial feature map. The original input information is retained through the residual connection, alleviating the gradient vanishing problem and improving the training stability; the design of the chain attention layer enhances the performance and stability of the overall feature extraction.
[0027] It should be noted that in S313, the chained multi-layer perceptron layer processes the residual feature map to obtain the multi-perceptual feature map; specifically: S3131: Perform 1×1 convolution processing on the residual feature map, and then perform depthwise separable convolution processing on the residual feature map after 1×1 convolution processing to obtain the third convolutional feature map; S3132: Perform activation processing on the third convolutional feature map to obtain the second activation feature map, and perform random dropout processing on the second activation feature map to obtain the dropout feature map; S3133: Perform feature reconstruction processing on the dropout feature map, and perform random dropout processing on the dropout feature map after feature reconstruction processing to obtain the multi-perceptual feature map.
[0028] Further, the residual feature map first passes through the first fully connected layer, and a 1×1 convolution operation is used for linear transformation in the channel dimension to expand or compress the number of feature channels, thereby enhancing the feature expression ability. Then, depthwise separable convolution is performed on it to obtain the third convolutional feature map, and then the third convolutional feature map is activated to obtain the second activation feature map; while introducing non-linear features, spatial features are extracted to enhance the local expression ability of the features. A random dropout operation is performed on the second activation feature map, randomly setting a part of the elements to 0 to obtain the dropout feature map; a random dropout operation is performed on the second activation feature map to prevent overfitting. Then, a fully connected layer is used to reconstruct the features of the dropout feature map, and finally, a random dropout process is performed on the dropout feature map after feature reconstruction to obtain the multi-sensory feature map, reducing the co-adaptability between neurons through random dropout.
[0029] Specifically, the C3k2_Strip module is represented by the following formula: , where, represents the initial feature vector, represents the random dropout process, represents the feature processing by the chain attention layer, represents the feature processing by the chain multi-layer perceptron layer, and both represent the normalization process, represents the first learnable scaling parameter, represents the second learnable scaling parameter, represents the element-wise multiplication of channel dimension broadcasting, represents the residual feature map, represents the chain feature map; Specifically, the chain attention layer is represented by the following formula: , where, represents the initial feature vector, represents the convolution operation, represents the convolution kernel weight, represents the activation process, represents the first convolutional feature map, represents the first activation feature map, represents the second convolutional feature map, represents the residual connection process, represents the initial feature map.
[0030] S400: Construct the C3k2_Kat module, replace the C3k2 module in the neck network of the YOLOv11 network model with the C3k2_Kat module to obtain the improved YOLOv11 network model; It should be noted that the process of the C3k2_Kat module C3k2_Kat (C3k2_Kolmogorov - Arnold Transformer) for data processing in S400 is as follows: S410: Input the input feature map into the C3k2_Kat module, and the preprocessing sub - module in the C3k2_Kat module performs convolution processing on the input feature map to obtain the fourth convolutional feature map; S420: According to the Kat feature extraction sub - module in the C3k2_Kat module, perform deep - level feature extraction processing on the fourth convolutional feature map step by step to obtain the depth - extracted feature map; S430: According to the residual connection sub - module in the C3k2_Kat module, perform residual connection on the depth - extracted feature and the input feature map to obtain the output feature map.
[0031] Furthermore, first, perform convolution processing on the input feature map according to the preprocessing sub - module to obtain the fourth convolutional feature map; adjust the number of channels of the input feature map to the number of channels required by the subsequent module through convolution operations to provide a suitable data format for subsequent processing. Then, perform deep - level feature extraction processing on the fourth convolutional feature map step by step through the Kat feature extraction sub - module to gradually extract higher - level features to obtain the depth - extracted feature map, and perform residual connection on the depth - extracted feature map and the input feature map through the residual connection sub - module to obtain the output feature map. Through residual connection, it alleviates gradient disappearance and enhances the feature reuse ability.
[0032] It should be noted that in S420, the Kat feature extraction sub - module performs deep - level feature extraction processing on the fourth convolutional feature map step by step to obtain the depth - extracted feature map; specifically: S421: According to the tensor reshaping layer in the Kat feature extraction sub - module, perform dimension reshaping processing and transpose processing on the fourth convolutional feature map to obtain the serialized feature map; S422: According to the Kat attention layer in the Kat feature extraction sub - module, perform attention weighting processing on the serialized feature map to obtain the attention feature map; S423: The first residual connection layer in the Kat feature extraction sub - module performs residual connection processing on the attention feature map and the serialized feature map to obtain the residual attention feature map; S424: According to the multi - layer perceptron layer in the Kat feature extraction sub - module, perform non - linear transformation processing on the residual attention feature map to obtain the local attention feature map; S425: The second residual connection layer in the Kat feature extraction sub-module performs a residual connection on the local attention feature map and the residual attention feature map to obtain a depth extraction feature map.
[0033] Further, the fourth convolutional feature map is subjected to dimension reshaping and transpose processing according to the tensor reshaping layer to obtain a serialized feature map. The fourth convolutional feature map is flattened from the original (N, C, H, W) shape into a two-dimensional form (N, C, H×W), and is converted into a serialized feature map (N, H×W, C) through a transpose operation to adapt to the subsequent feature processing flow. This step converts the information in the spatial dimension into a sequence form, providing a unified data structure for the subsequent Kat attention mechanism and multi-layer perceptron layer. Subsequently, the serialized feature map is weighted processed through the Kat attention mechanism in the Kat attention layer to obtain an attention feature map; the Kat attention mechanism captures the long-range dependence information between features by calculating the interaction relationship between the query, key, and value, using the self-attention mechanism. Specifically, this mechanism first performs a linear transformation on the serialized feature map to generate query, key, and value matrices, then calculates the attention weights through a dot product, and normalizes the weights using the Softmax function. Finally, the attention feature map is generated through weighted summation; this process can effectively capture the global interaction information between features, significantly improving the module's representation ability and the ability to extract defect features of road surface diseases; then, according to the first residual connection layer in the Kat feature extraction sub-module, the attention feature map and the serialized feature map are subjected to residual connection processing to obtain a residual attention feature map; afterwards, the module further introduces a multi-layer perceptron (MLP) to perform a non-linear transformation on the residual attention feature map to obtain a local attention feature map; the MLP consists of multiple fully connected layers and activation functions, and can perform complex non-linear mappings on the residual attention feature map to extract a higher-level local attention feature map. Through the processing of the MLP, the module can further enhance the generalization ability of the model, enabling it to better adapt to diverse task requirements. The combination of the Ka attention mechanism and the MLP enables the Kat feature extraction sub-module to take into account both global context information and local detail features. Specifically, the Kat attention mechanism is responsible for capturing global dependencies, while the MLP focuses on extracting local feature patterns, and the two work together to achieve the efficient fusion of multi-level features. In addition, a Layer Scale and a Dropout are introduced in the module, which are used to stabilize the training process and enhance the robustness of the model respectively. Finally, according to the second residual connection layer, the local attention feature map after layer scaling and dropout processing and the residual attention feature map after layer scaling and dropout processing are subjected to residual connection to obtain a depth extraction feature map. The depth extraction feature map not only retains the spatial structure information of the fourth convolutional feature map, but also incorporates rich global and local feature representations, providing strong feature support for downstream tasks.
[0034] It should be noted that in S422, the Kat attention layer performs attention weighting processing on the serialized feature map to obtain the attention feature map; specifically: S4221: The Kat attention layer performs linear transformation processing on the serialized feature map to obtain a query feature map, a key feature map, and a value feature map corresponding to the serialized feature map; S4222: Respectively perform attention head decomposition and calculation on the query feature map, the key feature map, and the value feature map to obtain multiple corresponding query head features, key head features, and value head features; S4223: Perform normalization processing on multiple query head features and key head features, and calculate the feature weights of each query head feature and each key head feature; S4224: Perform weighted summation on the value head features according to the feature weights of each query head feature and each key head feature to obtain the attention feature map.
[0035] Furthermore, the Kat attention mechanism can enhance the model's global context modeling ability and generalization ability while maintaining efficient computation by using multiple attention heads. First, the serialized feature map is mapped to the query feature map (Query, Q), key feature map (Key, K), and value feature map (Value, V) through linear transformation; the purpose of this step is to provide a basic representation for the subsequent attention mechanism calculation. By decomposing the serialized feature map into Q, K, and V, the model can capture the correlations between different positions in the sequence (represented by Q and K) and the feature information of each position (represented by V) respectively. This decomposition not only provides the necessary input for attention calculation but also reduces the model complexity and improves the computational efficiency by sharing linear transformation parameters. Next, Q, K, and V are split into multiple attention heads (Multi-Head), resulting in multiple corresponding query head features, key head features, and value head features; each head calculates attention independently; the design of this multi-head mechanism allows the model to learn diverse feature representations from different subspaces, thus enhancing the model's expressive ability. To further improve the computational stability and efficiency, normalization is performed on each query head feature and key head feature. The normalization operation can effectively alleviate the problems of gradient explosion or disappearance, ensure the numerical stability of the model during training, and accelerate the convergence speed. In addition, normalization can also make the feature distribution smoother, thereby improving the model's generalization ability. In the multi-head attention calculation stage, the model significantly improves the computational efficiency by calculating the attention scores of each head in parallel. Specifically, for each attention head, the model first calculates the dot product of the query head feature and the key head feature and scales the result by a scaling factor (ScaleFactor) to prevent the dot product value from being too large and causing gradient instability. Subsequently, the Softmax function is applied to the scaled dot product result to obtain the normalized feature weights. These feature weights are used to perform weighted summation on the value head features to obtain the attention feature map, thereby fusing the information of different positions in the sequence. This process not only captures the global dependencies within the sequence but also can dynamically adjust the importance of each position, enhancing the model's ability to model context information. Finally, the outputs of all attention heads are concatenated and mapped back to the original dimension to maintain the consistency of input and output. This step is achieved through linear transformation, ensuring that the output of multi-head attention can be seamlessly connected to the subsequent network layers. To further improve the robustness and generalization ability of the model, the Dropout operation is applied in the output stage. Dropout prevents the model from overfitting by randomly discarding some neurons, enhancing its robustness to road disease noise data. Finally, the module outputs the result, whose shape is the same as the input, facilitating direct use in downstream tasks or further processing.
[0036] Specifically, the Kat attention mechanism is represented by the following formula: , Among them, represents the convolutional kernel weight, represents the bias term, represents the dimension of the attention head, represents the query feature map, represents the key feature map, represents the value feature map, represents the transposed map of the key feature map, represents the Softmax function, represents the feature weight, represents the attention feature map S500: Train the improved YOLOv11 network model according to a part of the data in the target dataset, and test the improved YOLOv11 network model according to another part of the data; It should be noted that S500 is specifically S510: Divide the target dataset into a training set, a validation set and a test set according to a preset ratio; S520: Construct a loss function, and train the improved YOLOv11 network model according to the training set and the loss function; S530: Validate the trained improved YOLOv11 network model according to the validation set, and optimize the parameters of the improved YOLOv11 network model according to the difference between the validation output of the improved YOLOv11 network model and the true label.
[0037] S540: Test the improved YOLOv11 network model with optimized parameters according to the test set to obtain the test results.
[0038] Further, first, divide the target dataset into three parts according to a preset ratio: a training set, a validation set, and a test set. The training set is used for model training, the validation set is used for model tuning and selection, and the test set is used for final evaluation of the model's performance. Dividing the target dataset into a training set, a validation set, and a test set can effectively evaluate the generalization ability of the model. By dividing the dataset according to the preset ratio, the performance of the model on different datasets is ensured, and the overfitting phenomenon is avoided. The training set is used for model learning, the validation set is used for tuning hyperparameters, and the test set is used for final evaluation. For example, in this instance, the target dataset can be divided into a training set: validation set: test set = 8:1:1; The loss function of YOLOv11 (such as the combination of localization loss and classification loss) can effectively guide the model to learn the object detection task. A reasonable loss function can accelerate convergence and improve the accuracy of the model for object positions and categories. By training the improved YOLOv11 network model with the training set and the loss function, the model can gradually learn the features and patterns of the objects. As the training progresses, the loss value of the model should gradually decrease, indicating that the model is learning and optimizing. Using the validation set to validate the trained model can timely detect problems in the training process of the model, such as overfitting or underfitting. By monitoring performance metrics on the validation set (such as mAP, precision, recall, etc.), necessary adjustments and optimizations can be made to the model. Optimizing the parameters according to the difference between the output of the validation set and the true labels can significantly improve the performance of the model. The performance of the optimized model on the validation set should be better than that of the unoptimized model, indicating that the generalization ability of the model has been improved. Using the test set to finally test the model with optimized parameters can obtain the true performance of the model on unseen data. The test results will provide an evaluation of the effectiveness and reliability of the model in practical applications, including metrics such as accuracy, recall, and F1-score.
[0039] S600: Input the pavement image to be detected into the improved YOLOv11 network model after testing, and output the detection result.
[0040] A pavement disease detection method according to this embodiment can achieve real-time detection, has strong adaptability, and is suitable for large-scale pavement monitoring by using an improved YOLOv11 network model for pavement disease detection. The improved YOLOv11 network model can better extract and process features in pavement images by introducing new detection heads and feature extraction modules, thereby improving the accuracy and reliability of detection. Preprocessing the pavement image dataset can remove noise and enhance image quality, thereby improving the effect of subsequent model training. The preprocessed target dataset provides a clearer and more accurate input for the model, promoting the learning and generalization ability of the model. The C3k2_Strip module can effectively extract multi-level feature information through chained feature extraction and convolutional processing, enhancing the model's ability to identify complex pavement diseases. The introduction of residual connections enables better transmission of feature information, reduces information loss, and thus improves the learning efficiency of the model. The C3k2_Kat module can capture more detailed pavement disease features by gradually extracting deep features, enhancing the model's ability to detect different types of diseases. Through residual connections and the combination of input feature maps and deeply extracted feature maps, the model can comprehensively utilize feature information at different levels, improving the accuracy of detection. The chained attention layer can make the model pay more attention to important features through weighted processing of feature maps, enhancing the ability to identify key disease features. The combination of random dropout and residual connections enhances the robustness of the model, reduces the risk of overfitting, and improves the generalization ability of the model.
[0041] It should be noted that for a pavement disease detection method provided in an embodiment of this application, the execution subject can be a pavement disease detection system, or a control module in the pavement disease detection system for executing and loading a pavement disease detection method. In an embodiment of this application, taking a pavement disease detection system as an example for executing and loading a pavement disease detection method, a pavement disease detection method provided in an embodiment of this application is described.
[0042] A pavement disease detection system includes: An acquisition module: acquires a pavement image dataset, preprocesses the pavement image dataset, and obtains a target dataset; A first improvement module: constructs a YOLOv11 network model, and adds a P2 detection head to the detection layer of the YOLOv11 network model; A second improvement module: constructs a C3k2_Strip module, and replaces the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2_Strip module; Third improvement module: Construct the C3k2_Kat module, replace the C3k2 module in the neck network of the YOLOv11 network model with the C3k2_Kat module, and obtain the improved YOLOv11 network model; Training module: Train the improved YOLOv11 network model according to a part of the data in the target dataset, and test the improved YOLOv11 network model according to another part of the data; Prediction module: Input the pavement image to be detected into the improved YOLOv11 network model after testing, and output the detection result. A pavement disease detection system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a laptop computer, a palmtop computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc., which are not specifically limited in the embodiments of the present application.
[0043] A pavement disease detection system in an embodiment of the present application may be a device with an operating system. The operating system may be a windows operating system or other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0044] A pavement disease detection system provided by an embodiment of the present application can implement Figures 1 to 9 each process implemented by a pavement disease detection method in the method embodiment. To avoid repetition, it will not be elaborated here.
[0045] A pavement disease detection system according to this embodiment adds a P2 detection head to the detection layer of the YOLOv11 network model, which can enhance the model's detection ability for small targets and complex scenarios and improve the overall detection performance. Replacing the C3k2 module in the YOLOv11 backbone network with the C3k2_Strip module can extract features in pavement images more effectively, especially when dealing with complex pavement diseases, improving the accuracy and robustness of the model. By replacing the C3k2 module in the neck network with the C3k2_Kat module, different levels of feature information can be better fused, enhancing the model's ability to identify pavement diseases, especially when dealing with diverse diseases. By introducing the P2 detection head and the improved C3k2 module, the overall model shows higher accuracy and robustness in the pavement disease detection task. This system can adapt to the detection requirements of different types of pavement diseases, has a wide range of application prospects, and is suitable for monitoring various scenarios such as urban roads and rural roads. The improved YOLOv11 network model can achieve real-time detection, is suitable for different types of complex pavements, and improves work efficiency.
[0046] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of the pavement disease detection method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0047] An embodiment of the present application also provides a readable storage medium with a program or instruction stored thereon. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of the pavement disease detection method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0048] Among them, the processor is the processor in the electronic device in the above embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0049] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising such element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0050] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0051] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. A road surface disease detection method, characterized in that, The method includes: Obtain a road surface image dataset, preprocess the road surface image dataset to obtain a target dataset; Construct a YOLOv11 network model, and add a P2 detection head to the detection layer of the YOLOv11 network model; Construct a C3k2_Strip module, and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2_Strip module; Construct a C3k2_Kat module, and replace the C3k2 module in the neck network of the YOLOv11 network model with the C3k2_Kat module to obtain an improved YOLOv11 network model; Train the improved YOLOv11 network model according to a part of the data in the target dataset, and test the improved YOLOv11 network model according to another part of the data; Input the road surface image to be detected into the improved YOLOv11 network model after testing, and output the detection result.
2. The pavement disease detection method according to claim 1, characterized in that, The process of the C3k2_Strip module for data processing is as follows: The initial feature vector is input into the C3k2_Strip module, and the chained feature extraction sub-module in the C3k2_Strip module performs chained processing on the initial feature vector to obtain a chained feature map; Perform convolutional processing on the initial feature vector according to the convolutional sub-module in the C3k2_Strip module to obtain a convolutional feature map; Perform residual connection on the chained feature map and the convolutional feature map according to the residual connection sub-module in the C3k2_Strip module to obtain a target feature map.
3. The pavement disease detection method according to claim 2, characterized in that, The chained feature extraction sub-module performs chained processing on the initial feature vector to obtain a chained feature map; specifically: The initial feature vector is input into the chained attention layer in the chained feature extraction sub-module, and the chained attention layer performs feature extraction on the initial feature vector to obtain an initial feature map, and performs a first learnable scaling parameter multiplication process on the initial feature map to obtain a first learnable feature map; Perform random dropout processing on the first learnable feature map according to the first random dropout layer in the chained feature extraction sub-module, and perform residual connection processing on the learnable feature map after random dropout processing and the initial feature vector according to the first residual connection layer in the chained feature extraction sub-module to obtain a residual feature map; Process the residual feature map according to the chained multi-layer perceptron layer in the chained feature extraction sub-module to obtain a multi-perceptual feature map, and perform a second learnable scaling parameter multiplication process on the multi-layer perceptual feature map to obtain a second learnable feature map; Perform random dropout processing on the second learnable feature map according to the second random dropout layer in the chained feature extraction sub-module, and perform residual connection processing on the second learnable feature map after random dropout processing and the residual feature connection map according to the second residual connection layer in the chained feature extraction sub-module to obtain the chained feature map.
4. A pavement disease detection method according to claim 3, characterized in that, The chained attention layer performs feature extraction on the initial feature vector to obtain an initial feature map; specifically: Perform 1×1 convolution processing on the initial feature vector to obtain a first convolutional feature map, and perform feature projection processing on the first convolutional feature map; Perform activation processing on the first convolutional feature map after feature projection processing to obtain a first activation feature map; Perform 1×1 convolution processing on the first activation feature map to obtain a second convolutional feature map, and perform feature projection processing on the second convolutional feature map; Perform residual connection on the second convolutional feature map after feature projection processing and the initial feature vector to obtain the initial feature map; The chained multi-layer perceptron layer processes the residual feature map to obtain a multi-perceptual feature map; specifically: Perform 1×1 convolution processing on the residual feature map, and then perform depthwise separable convolution processing on the residual feature map after 1×1 convolution processing to obtain a third convolutional feature map; Perform activation processing on the third convolutional feature map to obtain a second activation feature map, and perform dropout processing on the second activation feature map to obtain a dropped feature map; Perform feature reconstruction processing on the dropped feature map, and perform dropout processing on the dropped feature map after feature reconstruction processing to obtain the multi-perceptual feature map.
5. A pavement disease detection method according to claim 1, characterized in that, The process of the C3k2_Kat module for data processing is as follows: Input the input feature map into the C3k2_Kat module, and the preprocessing sub-module in the C3k2_Kat module performs convolution processing on the input feature map to obtain a fourth convolutional feature map; According to the Kat feature extraction sub-module in the C3k2_Kat module, perform deep feature gradual extraction processing on the fourth convolutional feature map to obtain a deep extraction feature map; According to the residual connection sub-module in the C3k2_Kat module, perform residual connection on the deep extraction feature and the input feature map to obtain an output feature map.
6. The pavement disease detection method according to claim 5, wherein The Kat feature extraction sub-module performs deep feature gradual extraction processing on the fourth convolutional feature map to obtain a deep extraction feature map; specifically: According to the tensor reshaping layer in the Kat feature extraction sub-module, perform dimension reshaping processing and transpose processing on the fourth convolutional feature map to obtain a serialized feature map; According to the Kat attention layer in the Kat feature extraction sub-module, perform attention weighting processing on the serialized feature map to obtain an attention feature map; The first residual connection layer in the Kat feature extraction sub-module performs residual connection processing on the attention feature map and the serialized feature map to obtain a residual attention feature map; According to the multi-layer perceptron layer in the Kat feature extraction sub-module, perform non-linear transformation processing on the residual attention feature map to obtain a local attention feature map; The second residual connection layer in the Kat feature extraction sub-module performs residual connection on the local attention feature map and the residual attention feature map to obtain the deep extraction feature map.
7. A pavement disease detection method according to claim 6, characterized in that, The Kat attention layer performs attention weighting processing on the serialized feature map to obtain an attention feature map; specifically: The Kat attention layer performs a linear transformation on the serialized feature map to obtain a query feature map, a key feature map, and a value feature map corresponding to the serialized feature map; The query feature map, the key feature map, and the value feature map are respectively subjected to attention head decomposition and calculation to obtain a plurality of corresponding query head features, key head features, and value head features; The plurality of query head features and key head features are normalized, and the feature weights of each query head feature and each key head feature are calculated; The value head features are weighted and summed according to the feature weights of each query head feature and each key head feature to obtain the attention feature map.
8. A road surface disease detection system implementing the road surface disease detection method described in any one of claims 1-7, characterized in that, The system includes: An acquisition module: acquiring a road surface image data set, preprocessing the road surface image data set to obtain a target data set; A first improvement module: constructing a YOLOv11 network model, and adding a P2 detection head to the detection layer of the YOLOv11 network model; A second improvement module: constructing a C3k2_Strip module, and replacing the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2_Strip module; A third improvement module: constructing a C3k2_Kat module, and replacing the C3k2 module in the neck network of the YOLOv11 network model with the C3k2_Kat module to obtain an improved YOLOv11 network model; A training module: training the improved YOLOv11 network model according to a part of the data in the target data set, and testing the improved YOLOv11 network model according to another part of the data; A prediction module: inputting the road surface image to be detected into the improved YOLOv11 network model after testing, and outputting a detection result.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a road surface disease detection method as described in any one of claims 1-7 are implemented.
10. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of a road surface disease detection method as described in any one of claims 1-7 are implemented.
Citation Information
Cited By
Target detection model and method based on deep learning
CN121095726A
Method and system for rapidly detecting abnormal articles on expressway
CN121259766A