Asphalt pavement typical disease multi-target identification method based on deep learning

By constructing a shuttle structure network based on repeated codec wheels and embedding the asphalt pavement multi-objective recognition model with convolutional block attention and adaptive global feature fusion module, the problem of insufficient recognition accuracy in the prior art is solved, efficient and accurate disease recognition is achieved, and effective support for road maintenance is provided.

CN120451719APending Publication Date: 2025-08-08SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510531996.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing multi-objective recognition method based on deep learning is insufficient in complex scenarios, difficult to meet practical application requirements, and has limited ability to extract image details features.

Method used

A spindle-shaped structure network with repeated codec wheels is used as the backbone network, and a convolutional block attention module and an improved adaptive global feature fusion module are embedded in the backbone network to build a multi-objective recognition model for typical diseases of asphalt pavement, and the recognition results are obtained through training and testing.

Benefits of technology

It realizes efficient and accurate multi-objective identification of typical asphalt pavement diseases, providing rapid detection and maintenance support for road diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451719A_ABST
    Figure CN120451719A_ABST
Patent Text Reader

Abstract

The invention discloses an asphalt pavement typical disease multi-target identification method based on deep learning. The method comprises the following steps: acquiring two-dimensional image data and three-dimensional image data of an asphalt pavement; making an asphalt pavement typical disease sample image set; based on deep learning, taking a fusiform structure network of repeated coding and decoding wheels as a trunk network, and embedding a convolution block attention module and an improved self-adaptive global feature fusion module in the trunk network to construct a multi-target recognition model for the typical diseases of the asphalt pavement; and training, verifying and testing the asphalt pavement typical disease multi-target identification model, and scanning the asphalt pavement by using the tested asphalt pavement typical disease multi-target identification model to obtain an asphalt pavement typical disease identification result. According to the method, multi-target identification of typical diseases of the asphalt pavement can be efficiently and accurately realized, and effective technical support is provided for rapid detection and maintenance of road diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road detection and disease identification, and in particular to a multi-target identification method for typical asphalt pavement diseases based on deep learning. Background Art

[0002] Accurately identifying and predicting asphalt pavement defects is a crucial component of road maintenance. Traditional methods rely on manual inspection, which suffers from low efficiency, high subjectivity, and high costs. In recent years, deep learning-based image recognition technology has provided a new solution for automated defect detection. However, existing methods lack accuracy in multi-target recognition in complex scenarios and have limited ability to extract detailed image features, making them difficult to meet practical application requirements. Summary of the Invention

[0003] In response to the above-mentioned deficiencies in the prior art, the present invention provides a multi-target identification method for typical asphalt pavement diseases based on deep learning.

[0004] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0005] A multi-target recognition method for typical asphalt pavement defects based on deep learning includes the following steps:

[0006] S1, collecting two-dimensional image data and three-dimensional image data of the asphalt road surface;

[0007] S2. Producing a sample image set of typical asphalt pavement disease defects based on the two-dimensional image data and three-dimensional image data of the asphalt pavement;

[0008] S3. Based on deep learning, a spindle-structured network with repeated encoding and decoding cycles is used as the backbone network. A convolutional block attention module and an improved adaptive global feature fusion module are embedded in the backbone network to build a multi-target recognition model for typical asphalt pavement defects.

[0009] S4. Use the sample image set of typical asphalt pavement diseases to train, verify and test the multi-target recognition model of typical asphalt pavement diseases, and use the tested multi-target recognition model of typical asphalt pavement diseases to scan the asphalt pavement to obtain the recognition results of typical asphalt pavement diseases.

[0010] Furthermore, step S2 includes the following steps:

[0011] S21. Perform pixel-level typical disease annotation on the two-dimensional image data and the three-dimensional image data of the asphalt pavement, unify the RGB values of different typical disease categories, generate label data, and perform noise reduction processing on the label data to obtain noise-reduced label data;

[0012] S22, performing a pooling operation on the two-dimensional image data and the three-dimensional image data of the asphalt pavement, adjusting the image sizes to a uniform specification, and converting them into a standard image format to obtain two-dimensional standard image data and three-dimensional standard image data of the asphalt pavement;

[0013] S23: Integrate the denoised label data, the two-dimensional standard image data, and the three-dimensional standard image data of the asphalt pavement to produce a sample image set of typical asphalt pavement disease.

[0014] Furthermore, in step S3, a convolutional block attention module is embedded in the backbone network, specifically: the convolutional block attention module is embedded after the convolution operation of each level of the encoding path in the backbone network; the convolutional block attention module includes a channel attention mechanism and a spatial attention mechanism connected in sequence; the convolutional block attention module is used to perform channel weighting and spatial weighting on the input feature map of the convolutional block attention module to obtain a channel and spatial weighted feature map.

[0015] Furthermore, the data processing process of the channel attention mechanism is as follows: global average pooling and global maximum pooling are performed on the input feature map of the convolution block attention module to obtain the first path feature map and the second path feature map respectively, and then the number of channels of the first path feature map and the second path feature map are compressed through the first fully connected layer of the shared multilayer perceptron, and then the number of channels of the first path feature map and the second path feature map are restored through the second fully connected layer of the shared multilayer perceptron. Finally, the first path feature map and the second path feature map with restored channel numbers are added and the channel attention weight is generated through the Sigmoid activation function. The channel attention weight is used to perform channel weighting on the feature map of the input convolution block attention module to obtain the channel-weighted feature map.

[0016] Furthermore, the data processing process of the spatial attention mechanism is as follows: perform global average pooling and global maximum pooling on the channel-weighted feature map to obtain the first compressed feature map and the second compressed feature map in the spatial dimension respectively, then concatenate the first compressed feature map and the second compressed feature map in the spatial dimension, pass through a convolution layer, and finally generate the spatial attention weight through the Sigmoid activation function, and use the spatial attention weight to spatially weight the channel-weighted feature map to obtain the channel- and spatially weighted feature maps.

[0017] Furthermore, in step S3, an improved adaptive global feature fusion module is embedded in the backbone network, specifically: the improved adaptive global feature fusion module is embedded into each output level of the decoding path in the backbone network; the improved adaptive global feature fusion module includes a convolution submodule, an adaptive pooling layer and a self-attention mechanism; the improved adaptive global feature fusion module is used to perform convolution operations, adaptive pooling operations, flattening operations and weighted operations on the input feature map of the improved adaptive global feature fusion module to output a fused feature map.

[0018] Furthermore, the data processing process of the convolution submodule is as follows: three independent 1x1 convolution layers are used to perform convolution operations on the input feature map of the improved adaptive global feature fusion module to generate query feature map, key feature map and value feature map.

[0019] Furthermore, the data processing process of the adaptive pooling layer is: performing an adaptive pooling operation on the key feature map and the value feature map to obtain the key feature map and the value feature map after the adaptive pooling operation.

[0020] Furthermore, the data processing process of the self-attention mechanism is as follows: flatten the query feature map, the key feature map after the adaptive pooling operation, and the value feature map, calculate the similarity of the query feature map and the key feature map after the flattening operation through matrix multiplication to generate attention weights, and normalize the attention weights through the Softmax function, multiply the normalized attention weights with the value feature map after the flattening operation to obtain a weighted feature map, and add the weighted feature map to the input feature map of the improved adaptive global feature fusion module to output a fused feature map.

[0021] The present invention has the following beneficial effects:

[0022] The present invention uses a shuttle structure network with repeated encoding and decoding rounds as the backbone network, and embeds a convolutional block attention module and an improved adaptive global feature fusion module in the backbone network to construct a multi-target recognition model for typical asphalt pavement diseases. The tested multi-target recognition model for typical asphalt pavement diseases is used to scan the asphalt pavement to obtain the recognition results of typical asphalt pavement diseases. The entire process can efficiently and accurately realize multi-target recognition and prediction of typical asphalt pavement diseases, providing effective technical support for the rapid detection and maintenance of road diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a flowchart of a multi-target identification method for typical asphalt pavement defects based on deep learning;

[0024] Figure 2 This is a schematic diagram of a sample image of typical asphalt pavement diseases;

[0025] Figure 3 Schematic diagram of the convolutional block attention module;

[0026] Figure 4 Schematic diagram of the improved adaptive global feature fusion module;

[0027] Figure 5 This is a schematic diagram of a single encoding and decoding wheel structure;

[0028] Figure 6 This is a schematic diagram of the multi-target identification model structure for typical asphalt pavement diseases;

[0029] Figure 7 The loss trend diagram of the validation process of the multi-objective identification and prediction model for typical asphalt pavement diseases;

[0030] Figure 8 This is the average IOU trend chart of the verification process of the multi-target recognition and prediction model for typical asphalt pavement diseases. DETAILED DESCRIPTION

[0031] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0032] like Figure 1 As shown in FIG, a multi-target recognition method for typical asphalt pavement defects based on deep learning includes the following steps:

[0033] S1. Collect two-dimensional image data and three-dimensional image data of the asphalt road surface.

[0034] In an optional embodiment of the present invention, the present invention utilizes a road inspection vehicle equipped with a three-dimensional laser imaging system to collect two-dimensional image data and three-dimensional image data of an asphalt road surface.

[0035] S2. Generate a sample image set of typical asphalt pavement disease defects based on the two-dimensional image data and three-dimensional image data of the asphalt pavement.

[0036] In an optional embodiment of the present invention, step S2 includes the following steps:

[0037] S21. Perform pixel-level typical disease annotation on the two-dimensional image data and the three-dimensional image data of the asphalt pavement, unify the RGB values of different typical disease categories, generate label data, and perform noise reduction processing on the label data to obtain noise-reduced label data.

[0038] S22 , performing pooling operations on the two-dimensional image data and the three-dimensional image data of the asphalt pavement, adjusting the image sizes to a uniform specification, and converting them into a standard image format to obtain two-dimensional standard image data and three-dimensional standard image data of the asphalt pavement.

[0039] S23: Integrate the denoised label data, the two-dimensional standard image data, and the three-dimensional standard image data of the asphalt pavement to produce a sample image set of typical asphalt pavement disease.

[0040] like Figure 2 As shown, the present invention provides a schematic diagram of sample images of typical asphalt pavement diseases. Figure 2 The left part is a two-dimensional standard image of asphalt road surface. Figure 2 The middle part is a 3D standard image of an asphalt road surface. Figure 2 The right side of the diagram shows the label data after noise reduction, which includes two types of defects: marking and grouting.

[0041] S3. Based on deep learning, a shuttle structure network with repeated encoding and decoding rounds is used as the backbone network, and a convolutional block attention module and an improved adaptive global feature fusion module are embedded in the backbone network to construct a multi-target recognition model for typical asphalt pavement diseases.

[0042] In an optional embodiment of the present invention, the present invention embeds a convolutional block attention module in the backbone network, specifically: embedding the convolutional block attention module after the convolution operation of each level of the encoding path in the backbone network; Figure 3 As shown, the convolutional block attention module includes a channel attention mechanism and a spatial attention mechanism connected in sequence; the convolutional block attention module is used to perform channel weighting and spatial weighting on the input feature map of the convolutional block attention module to obtain a channel and spatial weighted feature map.

[0043] like Figure 3 As shown in the figure, the data processing process of the channel attention mechanism is as follows: perform global average pooling and global maximum pooling on the input feature map of the convolution block attention module to obtain the first path feature map and the second path feature map respectively, and then compress the number of channels of the first path feature map and the second path feature map through the first fully connected layer of the shared multilayer perceptron, and then restore the number of channels of the first path feature map and the second path feature map through the second fully connected layer of the shared multilayer perceptron. Finally, add the first path feature map and the second path feature map with restored channel numbers and generate channel attention weights through the Sigmoid activation function. Use the channel attention weights to perform channel weighting on the feature map of the input convolution block attention module to obtain the channel-weighted feature map.

[0044] The expression of the channel attention mechanism is:

[0045]

[0046] Where: M c (F) is the channel weighted feature map, F is the input feature map of the convolutional block attention module, AvgPool is the global average pooling, MLP is the shared multi-layer perceptron, σ is the sigmoid activation function, MaxPool is the global maximum pooling, W1 and W0 are both learnable channel attention weights, is the first path feature map, is the second path feature graph.

[0047] like Figure 3 As shown in the figure, the data processing process of the spatial attention mechanism is as follows: perform global average pooling and global maximum pooling on the channel-weighted feature map to obtain the first compressed feature map and the second compressed feature map in the spatial dimension respectively, then concatenate the first compressed feature map and the second compressed feature map in the spatial dimension, pass through a convolution layer, and finally generate the spatial attention weight through the Sigmoid activation function, and use the spatial attention weight to perform spatial weighting on the channel-weighted feature map to obtain the channel- and spatial-weighted feature maps.

[0048] The present invention multiplies the feature map of the input convolutional block attention module by the channel attention weight to enhance the importance of the channel dimension. The result is then multiplied by the spatial attention weight to enhance the importance of the spatial dimension. The final output is a channel- and spatial-weighted feature map, which retains the importance of both the channel and spatial dimensions. The present invention finally uses this channel- and spatial-weighted feature map as the output of the encoder.

[0049] The present invention embeds an improved adaptive global feature fusion module in the backbone network, specifically: embedding the improved adaptive global feature fusion module into each output level of the decoding path in the backbone network; Figure 4 As shown, the improved adaptive global feature fusion module includes a convolution submodule, an adaptive pooling layer and a self-attention mechanism; the improved adaptive global feature fusion module is used to perform convolution operations, adaptive pooling operations, flattening operations and weighting operations on the input feature map of the improved adaptive global feature fusion module to output a fused feature map.

[0050] like Figure 4 As shown in Figure 1, the data processing process of the convolution submodule is as follows: three independent 1x1 convolution layers are used to perform convolution operations on the input feature map of the improved adaptive global feature fusion module to generate query feature maps, key feature maps, and value feature maps.

[0051] like Figure 4 As shown in the figure, the data processing process of the adaptive pooling layer is: performing adaptive pooling operations on the key feature map and the value feature map to obtain the key feature map and the value feature map after the adaptive pooling operation.

[0052] like Figure 4 As shown in the figure, the data processing process of the self-attention mechanism is: flatten the query feature map, the key feature map after the adaptive pooling operation, and the value feature map, calculate the similarity of the query feature map and the key feature map after the flattening operation through matrix multiplication to generate the attention weight, and normalize the attention weight through the Softmax function, multiply the normalized attention weight with the value feature map after the flattening operation to obtain the weighted feature map, and add the weighted feature map to the input feature map of the improved adaptive global feature fusion module to output the fused feature map.

[0053] like Figure 5 As shown in FIG, the present invention provides a schematic diagram of a single encoding and decoding wheel structure. Figure 6 As shown in the figure, the multi-target recognition model for typical asphalt pavement defects uses a shuttle-shaped network with repeated encoding and decoding cycles as the backbone network, where the repeated encoding and decoding cycles are multiple identical single encoding and decoding cycles. The specific working process of the multi-target recognition model for typical asphalt pavement defects includes the following steps:

[0054] A1. Input asphalt pavement image data with 2 channels and a resolution of 256x512. To extract multi-scale feature information, average pooling is used to downsample the input 256x512 asphalt pavement image data by a factor of 2, 4, 8, 16, and 32, resulting in multi-scale input feature maps with resolutions of 128×256, 64×128, 32×64, 16×32, and 8×16, respectively. These multi-scale input feature maps serve as input to a single encoder-decoder round to extract feature information at different levels.

[0055] A2. Encoding operation. The asphalt pavement image data with 2 channels and a resolution of 256x512 from step A1 and the multi-scale input feature maps of 128×256, 64×128, 32×64, 16×32, and 8×16 generated by the downsampling operation are used as the encoder input. The first level (256x512) is directly convolved and the convolved feature map is downsampled using max pooling. The downsampled image is concatenated with the feature map of the second level (128x256). The convolution and downsampling operations are repeated on the concatenated feature map until the sixth level, obtaining outputs of sizes (minimum number of channels, 256, 512), (2×minimum number of channels, 128, 256), (4×minimum number of channels, 64, 128), (8×minimum number of channels, 32, 64), (16×minimum number of channels, 16, 32), and (32×minimum number of channels, 8, 16). The above 6 outputs are then used as the input feature maps of the convolutional block attention module (CBAM), which performs channel weighting and spatial weighting on the input feature maps of the convolutional block attention module to obtain channel and spatial weighted feature maps.

[0056] A3. Decoding: The channel- and spatially weighted feature maps are used as decoder input. Each layer is upsampled through deconvolution, and the feature maps from different layers are concatenated. Feature fusion is performed using the improved adaptive global feature fusion module. The improved adaptive global feature fusion module performs convolution, adaptive pooling, flattening, and weighting on the input feature maps of the improved adaptive global feature fusion module to output a fused feature map.

[0057] A4. Memory gate operation. The fused feature maps of the previous and current rounds are used as input. Using learnable weight parameters, a Sigmoid activation function is used to generate the weights of the fused feature map of the previous round. The weights of the fused feature map of the previous round are multiplied by the fused feature map of the previous round to obtain a weighted fused feature map. This weighted feature map is then added to the fused feature map of the current round to obtain a fused feature map that enhances global connectivity.

[0058] A5. Repeat the encoder, decoder, and memory gate operations for a total of six rounds. Map the output of the final decoder through a 3x3 convolutional layer to the channels corresponding to the target types being detected. Apply the Softmax activation function to generate the final classification probabilities, thereby identifying typical asphalt pavement defects.

[0059] S4. Use the sample image set of typical asphalt pavement diseases to train, verify and test the multi-target recognition model of typical asphalt pavement diseases, and use the tested multi-target recognition model of typical asphalt pavement diseases to scan the asphalt pavement to obtain the recognition results of typical asphalt pavement diseases.

[0060] In an optional embodiment of the present invention, the present invention divides the sample image set of typical asphalt pavement diseases into training set data, validation set data and test set data. The training process of the present invention uses a shuffling function to randomly perturb the training set data to enhance the generalization ability of the model, and at the same time introduces a pre-training mechanism to accelerate convergence and improve the performance of the multi-target recognition model of typical asphalt pavement diseases. In addition, the present invention introduces a learning rate scheduler to enable the training process to have a dynamic learning rate adjustment mechanism. When the performance of the multi-target recognition model of typical asphalt pavement diseases stagnates or the loss function converges slowly, the learning rate is automatically reduced to prompt the multi-target recognition model of typical asphalt pavement diseases to jump out of the local optimal solution, improve the global optimization ability, and further enhance the training effect of the multi-target recognition model of typical asphalt pavement diseases.

[0061] During the verification process, the model enters the evaluation mode and traverses each batch of data in the verification set. For each batch, the verification loss, accuracy and IOU index are calculated and recorded to evaluate the model performance. After each verification round, the training and verification loss and IOU index are written to a CSV file, and the learning rate is updated according to the IOU index of the verification set. If the verification IOU of the current round is better than the historical best value, the current model weight is saved as the best model; if Figure 7 As shown in Figure 2, the Loss value of the validation set of the multi-target disease recognition and prediction model for asphalt pavement shows a decreasing trend with the iterative process of training; Figure 8 As shown in the figure, the average IOU value increases and stabilizes, with the highest average IOU reaching 0.793, proving that this scheme has considerable recognition and prediction effects.

[0062] The present invention then integrates the multi-target recognition model of typical asphalt pavement diseases that has passed the test into a road inspection vehicle, and uses the road inspection vehicle with the integrated model to scan the asphalt pavement to obtain the recognition results of the typical asphalt pavement diseases.

[0063] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0064] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0066] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

[0067] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A multi-target recognition method for typical asphalt pavement defects based on deep learning, characterized by: The following steps are involved: S1, collecting two-dimensional image data and three-dimensional image data of the asphalt road surface; S2. Producing a sample image set of typical asphalt pavement disease defects based on the two-dimensional image data and three-dimensional image data of the asphalt pavement; S3. Based on deep learning, a spindle-structured network with repeated encoding and decoding cycles is used as the backbone network. A convolutional block attention module and an improved adaptive global feature fusion module are embedded in the backbone network to build a multi-target recognition model for typical asphalt pavement defects. S4. Use the sample image set of typical asphalt pavement diseases to train, verify and test the multi-target recognition model of typical asphalt pavement diseases, and use the tested multi-target recognition model of typical asphalt pavement diseases to scan the asphalt pavement to obtain the recognition results of typical asphalt pavement diseases.

2. The multi-target identification method for typical asphalt pavement defects based on deep learning according to claim 1 is characterized in that: Step S2 includes the following steps: S21. Perform pixel-level typical disease annotation on the two-dimensional image data and the three-dimensional image data of the asphalt pavement, unify the RGB values of different typical disease categories, generate label data, and perform noise reduction processing on the label data to obtain noise-reduced label data; S22, performing a pooling operation on the two-dimensional image data and the three-dimensional image data of the asphalt pavement, adjusting the image sizes to a uniform specification, and converting them into a standard image format to obtain two-dimensional standard image data and three-dimensional standard image data of the asphalt pavement; S23: Integrate the denoised label data, the two-dimensional standard image data, and the three-dimensional standard image data of the asphalt pavement to produce a sample image set of typical asphalt pavement disease.

3. The multi-target identification method for typical asphalt pavement defects based on deep learning according to claim 1 is characterized in that: In step S3, a convolutional block attention module is embedded in the backbone network, specifically: the convolutional block attention module is embedded after the convolution operation of each level of the encoding path in the backbone network; the convolutional block attention module includes a channel attention mechanism and a spatial attention mechanism connected in sequence; the convolutional block attention module is used to perform channel weighting and spatial weighting on the input feature map of the convolutional block attention module to obtain a channel and spatial weighted feature map.

4. The multi-target identification method for typical asphalt pavement defects based on deep learning according to claim 3 is characterized in that: The data processing process of the channel attention mechanism is as follows: perform global average pooling and global maximum pooling on the input feature map of the convolution block attention module to obtain the first path feature map and the second path feature map respectively, and then compress the number of channels of the first path feature map and the second path feature map through the first fully connected layer of the shared multilayer perceptron, and then restore the number of channels of the first path feature map and the second path feature map through the second fully connected layer of the shared multilayer perceptron. Finally, add the first path feature map and the second path feature map with restored channel numbers and generate channel attention weights through the Sigmoid activation function. Use the channel attention weights to perform channel weighting on the feature map of the input convolution block attention module to obtain the channel-weighted feature map.

5. The multi-target identification method for typical asphalt pavement defects based on deep learning according to claim 3 is characterized in that: The data processing process of the spatial attention mechanism is as follows: perform global average pooling and global maximum pooling on the channel-weighted feature map to obtain the first compressed feature map and the second compressed feature map in the spatial dimension respectively, then concatenate the first compressed feature map and the second compressed feature map in the spatial dimension, pass through a convolution layer, and finally generate the spatial attention weight through the Sigmoid activation function, and use the spatial attention weight to perform spatial weighting on the channel-weighted feature map to obtain the channel- and spatial-weighted feature maps.

6. The multi-target identification method for typical asphalt pavement defects based on deep learning according to claim 1 is characterized in that: In step S3, an improved adaptive global feature fusion module is embedded in the backbone network, specifically: the improved adaptive global feature fusion module is embedded in each output layer of the decoding path in the backbone network; the improved adaptive global feature fusion module includes a convolution submodule, an adaptive pooling layer and a self-attention mechanism; the improved adaptive global feature fusion module is used to perform convolution operations, adaptive pooling operations, flattening operations and weighting operations on the input feature map of the improved adaptive global feature fusion module to output a fused feature map.

7. The multi-target identification method for typical asphalt pavement defects based on deep learning according to claim 6 is characterized in that: The data processing process of the convolution submodule is as follows: three independent 1x1 convolution layers are used to perform convolution operations on the input feature map of the improved adaptive global feature fusion module to generate query feature map, key feature map and value feature map.

8. The multi-target identification method for typical asphalt pavement defects based on deep learning according to claim 6 is characterized in that: The data processing process of the adaptive pooling layer is: perform adaptive pooling operations on the key feature map and the value feature map to obtain the key feature map and the value feature map after the adaptive pooling operation.

9. The multi-target identification method for typical asphalt pavement defects based on deep learning according to claim 6 is characterized in that: The data processing process of the self-attention mechanism is as follows: flatten the query feature map, the key feature map after the adaptive pooling operation, and the value feature map, calculate the similarity of the query feature map and the key feature map after the flattening operation through matrix multiplication to generate the attention weight, and normalize the attention weight through the Softmax function, multiply the normalized attention weight with the value feature map after the flattening operation to obtain the weighted feature map, and add the weighted feature map to the input feature map of the improved adaptive global feature fusion module to output the fused feature map.

Citation Information

Patent Citations

  • Grouting control method and system for pavement patching based on image semantic segmentation

    CN117974617A

  • Road and bridge deck disease identification method based on deep learning

    CN119339236A