Rapid detection and visualization method and system for road surface apparent damage based on infrared thermal imager

By improving the backbone network of the YOLO model to Swin Transformer and introducing a smooth gradient class activation mapping algorithm, the problem of poor interpretability and low detection efficiency of the apparent road damage detection in extreme scenarios is solved, and high-precision and robustness detection under extreme low light conditions is achieved.

CN119579572BActive Publication Date: 2025-07-08HARBIN INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411770346.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-07-08
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Existing deep learning models have problems of poor interpretability, missed detection, missed detection and low efficiency in road apparent damage detection in extreme scenarios, especially in extreme low-light conditions, it is difficult to effectively detect micro-disease targets.

Method used

The improved YOLO model is adopted to replace the C2f module in the backbone network as the Swin Transformer module, and combined with the smooth gradient class activation mapping algorithm for visualization. By introducing Gaussian noise smooth gradient information, the interpretability and small object detection capabilities of the model are enhanced.

Benefits of technology

It improves the accuracy and robustness of apparent damage detection on roads, can effectively identify small targets and weak information under extreme low light conditions, and enhances the accuracy and real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119579572B_ABST
    Figure CN119579572B_ABST
Patent Text Reader

Abstract

A rapid detection and visualization method and system for road surface damage based on an infrared thermal imager, belonging to the field of non-destructive road testing technology. To solve the problem of poor interpretability of existing deep learning models for road surface damage, the present invention first obtains an image of road surface damage, then replaces at least one C2f module in the backbone network of the YOLO model with a Swin Transformer module, and inputs the image into the improved YOLO model for damage detection; during the processing of the improved YOLO model, a smooth gradient class activation mapping algorithm considering loss is used for visualization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of non-destructive road detection, and in particular relates to a method and system for rapid detection and visualization of road surface apparent damage. Background Art

[0002] An infrared thermal imager measures the radiation on the surface of an object through an infrared detector and an optical imaging objective lens, establishes a mutual mapping relationship between the radiation and the surface temperature, and reflects it onto the photosensitive element of the infrared detector to obtain an infrared thermal image. This thermal image corresponds to the thermal distribution field on the surface of the object and can directly reflect the heat distribution status of the measured target and its surroundings. When there are defects inside an object, a temperature difference will be formed between the defective area and the non-defective area of the object. This temperature difference depends not only on the thermophysical properties of the object material but also on the size of the defect, the distance from the surface, and its thermophysical properties. The infrared thermal imager can detect these temperature changes and then judge the situation of the defects, which also provides the possibility for the application of the infrared thermal imager in non-destructive road detection.

[0003] Most of the existing traditional road damage detections adopt digital imaging technology, which is also the mainstream method for studying the characteristics of road diseases. However, in cases of bad weather, uneven illumination, road shadows, and partial occlusion, etc., the traditional road image acquisition system cannot well obtain the key information of the image, resulting in unsatisfactory detection effects in extreme scenarios. The infrared imaging detection technology overcomes some drawbacks of the conventional detection. Based on the internal relationship among the infrared radiation - surface temperature - material properties of the road, through the analysis of thermal image features, the surface temperature distribution of the road can be intuitively understood, and then the diseases can be located and detected by the apparent damage of the road and the surrounding temperature differences.

[0004] At present, after a large number of graphic data sets are collected in actual projects, the location and type of damage are still judged and located manually. This method is prone to misjudgment and missed judgment, and completely relies on manual work during the disease interpretation process, which will prolong the detection cycle and is difficult to meet the growing needs of road engineering detection. With the improvement of computer performance, machine learning technologies such as deep learning have been widely used in image recognition and target detection due to their powerful feature automatic extraction ability. However, there is little research on road damage detection in extremely weak light scenarios, and most of them have problems such as inconsistent sample image features, lack of image data, low accuracy, and poor stability, and cannot well meet the needs of modern road detection and maintenance.

[0005] Although deep learning technology has made great progress in image tracking, target detection, etc. at the present stage, there is still room for improvement in the target detection of road disease images in extreme scenarios. Especially for the tiny disease targets in infrared images and the real-time problem during detection, it is easy to have problems such as missed detection, misdetection, and low efficiency. Summary of the Invention

[0006] The purpose of the present invention is to solve the problem of poor interpretability in the existing deep learning models for detecting road surface damage.

[0007] A rapid detection and visualization method for road surface damage based on an infrared thermal imager. First, an image of the road surface damage is obtained, and then the image is input into an improved YOLO model with an attention mechanism for damage detection to obtain the detection result;

[0008] The improved YOLO model with an attention mechanism is obtained by replacing at least one C2f module in the backbone network of the YOLO model with a Swin Transformer module;

[0009] During the processing of the improved YOLO model with an attention mechanism, a Smooth Grad-CAM algorithm considering loss is used for visualization. The Smooth Grad-CAM algorithm considering loss includes the following steps:

[0010] Gaussian noise is introduced in the gradient calculation to smooth the gradient. Assume that there are n' convolutional kernels in the last convolutional layer of the network of the YOLO model, and each convolutional kernel generates a feature map. For the k-th feature map, the gradient information y c is calculated as follows:

[0011]

[0012] where m represents the number of iterations, n represents the number of times of adding Gaussian noise, x is the input feature information, is the introduced noise, and y c represents the prediction score of the model for class c;

[0013] Then, global average pooling is performed on the gradient information in the spatial dimension M×N to obtain the weight δ k of the feature map;

[0014] Calculate the loss during the propagation process:

[0015] L k = L O-classify + λL cam (k)

[0016]

[0017] where L O-classify represents the classification loss, L cam (k) represents the interpretability metric loss, camk represents the k-th feature map, and mask k represents the corresponding target region mask, Denote the L2 norm, and λ represents the proportion of the interpretability metric loss;

[0018] Comprehensively consider the weight δ k and the gradient loss to obtain the discriminative localization map F of the object of interest:

[0019]

[0020] Among them, A k represents the k-th feature map, and ReLU is the activation function.

[0021] Furthermore, the weight δ of the feature map k is as follows:

[0022]

[0023] In the formula, M×N represents the spatial dimension of the feature map A k and A ij is the element in the i-th row and j-th column of the feature map.

[0024] Furthermore, improving the YOLO model with the added attention mechanism is to replace the last C2f module in the backbone network of the YOLO model with a Swin Transformer module.

[0025] Furthermore, the processing process of the backbone network of the improved YOLO model with the added attention mechanism is as follows:

[0026] The input is first sent to the first CBS convolution module for processing, and then has been processed by the second CBS convolution module, the first C2f module, the third CBS convolution module, the second C2f module, the fourth CBS convolution module, the third C2f module, the fifth CBS convolution module, the first Swin Transformer module, and the SPPF module;

[0027] The CBS convolution module includes a convolutional layer, a BN layer, and a SiLU activation function.

[0028] Furthermore, the obtained image of the road apparent damage is an infrared image.

[0029] A rapid detection and visualization system for road apparent damage based on an infrared thermal imager, comprising:

[0030] A road apparent damage image acquisition unit: used to first acquire an image of road apparent damage;

[0031] A road apparent damage detection unit: input the image into the improved YOLO model with the added attention mechanism for damage detection to obtain the detection result;

[0032] The improved YOLO model with an attention mechanism is to replace at least one C2f module in the backbone network of the YOLO model with a Swin Transformer module;

[0033] Visualization unit: During the processing of the improved YOLO model with an attention mechanism, the smooth gradient class activation mapping algorithm considering loss is used for visualization. The smooth gradient class activation mapping algorithm considering loss includes the following steps:

[0034] Introduce Gaussian noise in gradient calculation to smooth the gradient. Assume that there are n' convolutional kernels in the last convolutional layer of the network, and each convolutional kernel generates a feature map. For the k-th feature map, calculate the gradient information y c as follows:

[0035]

[0036] In the formula, m represents the number of iterations, n represents the number of times Gaussian noise is added, x is the input feature information, is the introduced noise, and y c represents the predicted score of the model for class c;

[0037] Then perform global average pooling on this gradient information in the spatial dimension M×N to obtain the weight δ of the feature map k ;

[0038] Calculate the loss during the propagation process:

[0039] L k = L O-classify + λL cam (k)

[0040]

[0041] In the formula, L O-classify represents the classification loss, L cam (k) represents the interpretability metric loss, cam k represents the k-th feature map, mask k represents the corresponding target region mask, represents the L2 norm, and λ represents the proportion of the interpretability metric loss;

[0042] Comprehensively consider the weight δ k and the gradient loss to obtain the discriminant localization map F of the attention object:

[0043]

[0044] where A k represents the k-th feature map, and ReLU is the activation function.

[0045] Further, the weight δ of the feature map k is as follows:

[0046]

[0047] where M×N represents the spatial dimension of the feature map A k and \(A_{ij}\) is the element in the \(i\)-th row and \(j\)-th column of the feature map. ij

[0048] Further, the improved YOLO model with an added attention mechanism replaces the last C2f module in the backbone network of the YOLO model with a Swin Transformer module.

[0049] Further, the processing procedure of the backbone network of the improved YOLO model with an added attention mechanism is as follows:

[0050] The input is first sent to the first CBS convolution module for processing, and then has been processed by the second CBS convolution module, the first C2f module, the third CBS convolution module, the second C2f module, the fourth CBS convolution module, the third C2f module, the fifth CBS convolution module, the first Swin Transformer module, and the SPPF module;

[0051] The CBS convolution module includes a convolutional layer, a BN layer, and a SiLU activation function.

[0052] Further, the obtained image of the road apparent damage is an infrared image.

[0053] Beneficial effects:

[0054] Compared with the prior art, the model of the present application improves YOLOv8, adds a Swin Transformer attention mechanism for multi-feature perception fusion, so that the model obtains more accurate and useful information about the attention features, and can effectively improve the recognition effect. The present invention introduces a smooth gradient class activation mapping algorithm considering loss to enhance the interpretability of the model and improve the accuracy of small target detection. At the same time, it locates and focuses on the feature region to enhance the localization of the model decision, enriches the semantic information, and thus improves the detection accuracy. In addition, the present invention uses infrared images for processing, which can well solve the problem of difficult road detection in extremely weak light scenes. Even in extremely weak light scenes, the present invention also has good road detection effects. The present invention improves the accuracy and robustness of the road apparent damage recognition model in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Flowchart of the method for rapid detection and visualization of road apparent damage based on an infrared thermal imager of the present invention; ​

[0056] Figure 2 Schematic diagram of adding the target detection network structure at different positions of the improved attention mechanism of the present invention;

[0057] Figure 3 To add the Swin Transformer structure;

[0058] Figure 4 Visualization diagram comparison of the target detection network effect of the improved YOLOv8 road apparent damage infrared image of the present invention and the infrared three-dimensional comparison analysis image. Detailed implementation manners

[0059] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be described below through specific embodiments shown in the drawings. However, it should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0060] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.

[0061] Detailed implementation manner one: Combining Figure 1 To illustrate this implementation manner,

[0062] The method for rapid detection and visualization of road apparent damage based on an infrared thermal imager according to this implementation manner includes the following steps:

[0063] Step 1: Use an infrared thermal imager to collect infrared images of road apparent damage on site.

[0064] In this step, the vehicle-mounted and hand-held methods are used to collect the road damage data set. Considering that the temperature change throughout the day has a significant impact on the image temperature difference effect, the data at different time periods of a day are collected.

[0065] Step 2: Manually annotate the disease areas of the infrared images obtained in Step 1, and use Labelimg to annotate the damages in the images to make a data set.

[0066] Screen the images collected in Step 1, and determine the apparent damaged areas of the road through on-site inspections, manual image analysis, and expert confirmation methods; when manually searching for damaged targets in the infrared thermal image, due to the complex information contained in the original image, the damaged features may not be obvious. Therefore, it is necessary to carefully screen and annotate the original image, and use Labelimg software to annotate the image to form a corresponding txt file. This file contains the type, size, and location information of the target of interest to ensure the accuracy and reliability of the dataset.

[0067] Step 3: Use the dataset obtained in Step 2 for data augmentation to expand the image dataset, and then divide it into a training set, a validation set, and a test set according to a certain ratio (such as 8:1:1);

[0068] During the process of using the dataset obtained in Step 2 for data augmentation, the augmentation methods include data augmentation methods such as rotating, scaling, adding random noise, horizontal mirroring, and color conversion to the dataset.

[0069] Data augmentation is an important means to improve the generalization ability of deep learning models. By methods such as rotation, scaling, adding random noise, horizontal mirroring, and color inversion, the number of samples can be effectively increased, the background information of the images can be enriched, and the network can learn more robust and deeper target features, thereby avoiding problems such as non-convergence and overfitting during model training, and ultimately improving the detection accuracy and precision of the model.

[0070] Step 4: Use the training set divided in Step 3 to train and optimize the improved YOLO model with an attention mechanism to obtain a trained network model;

[0071] In this embodiment, the improved YOLO model with an attention mechanism is an improvement based on the YOLOv8 model. Use the training set obtained in Step 3 to train and optimize the improved YOLOv8 model, and the network structure is as Figure 2 shown. First, replace C2f at different positions in the Backbone with Swin Transformer (there are three ways to replace the positions, Figure 2ST1, ST2, and ST3 in it), model training is carried out respectively, and then deep feature extraction is realized through multi-layer convolution, C2f, Swin Transformer, and SPPF modules in sequence to achieve efficient fusion of local features and global features; the convolutional layer of the Backbone is CBS, and CBS includes a convolutional layer, a BN layer, and a SiLU activation function. It should be noted that: Swin Transformer has achieved good results when verified on other datasets, but for cracks, the foreground information in the image is limited. Therefore, the present invention attempts to change the position of Swin Transformer to find an optimal solution so that the network can effectively extract crack features.

[0072] Next, the feature map extracted by the Backbone is passed to the Neck network. Through the multi-scale feature pyramid and realizing the beneficial transfer of cross-layer information, the network can better perceive the features of targets at different scales. The model uses anchor boxes to predict the bounding boxes on the feature maps at multiple scales, predicting the position offset, confidence, and class probability of each anchor box. Finally, it enters the detection network. In this area, the optimized feature maps are extracted and superimposed to obtain the feature map with region proposals, and the fully connected operation is used for target localization, so as to perform the bounding box regression and classification regression of image diseases. The prediction results are post-processed, including threshold filtering and non-maximum suppression, to remove low-confidence predictions and merge overlapping boxes, and finally output the class, bounding box coordinates, and confidence of the target, and finally obtain the accurate information of the detected underground target space body. To enhance the interpretability of the model and the trust in the detection results, the present invention proposes a new smooth gradient class activation mapping algorithm considering loss. During the forward propagation process, the model outputs the class scores; subsequently, during the backward propagation, the gradient of the feature map of the last convolutional layer with respect to the class scores is calculated, and a visual heat map is generated through global average pooling and weighted processing, showing the regions that the model focuses on when making decisions, especially the visualization of small targets and weak information regions in the apparent damage of the road surface.

[0073] The network architecture of Swin Transformer is as Figure 3As shown. Swin Transformer is an advanced vision Transformer architecture. Traditional vision models may not extract complex structural and detailed features in images sufficiently when processing images, while Swin Transformer largely solves this problem. When the input image passes through the previous convolutional layer to output a feature map, it first passes through the window multi-head self-attention (W-MSA) module, which divides the feature map into multiple windows and independently performs multi-head self-attention calculations within each window to capture local features. Then, through the shifted window multi-head self-attention (SW-MSA) module, a shifted window mechanism is introduced to achieve information interaction between windows, capture global context information, and enhance the model's perception ability. At the same time, the multi-layer perceptron (MLP) further extracts and transforms the image features, the LayerNorm layer normalizes the features, and the residual connection helps to alleviate the gradient problem and improve the model performance.

[0074] W-MSA divides the input feature map into multiple windows of a fixed size and independently performs multi-head self-attention calculations within each window. This operation reduces the computational complexity while ensuring the local dependence of features. For the self-attention calculation within each window, it can be expressed as:

[0075]

[0076] where Q, K, and V are the query, key, and value matrices respectively, which are obtained from the input feature map through linear transformation. d k is the dimension of the key, which is used to scale the dot product result to prevent gradient vanishing or explosion.

[0077] SW-MSA is an improvement over W-MSA. It introduces a shifted window mechanism to make adjacent windows overlap, thus achieving information interaction between windows. Through multiple shifts and calculations, SW-MSA can capture global context information and improve the model's perception ability. Before each MSA module and MLP, Swin Transformer applies a LayerNorm layer (LN) to normalize the input features and accelerate the model training process. At the same time, a residual connection is applied after each module, adding the input and output of the module, which helps to alleviate the gradient vanishing problem and further improve the model performance. The calculation formula for this part can be expressed as:

[0078]

[0079] In the formula, μ is the sample mean, σ 2is the sample variance, ε is a very small number, and γ and β are learnable parameters for rescaling and offsetting.

[0080] During the model training process, the present invention uses Precision, Recall, and mean Average Precision (mAP) as model evaluation metrics. Precision represents the proportion of samples predicted as damaged that are actually damaged; Recall represents the proportion of samples actually damaged that are correctly predicted as damaged; mAP is a comprehensive metric obtained by averaging the detection precisions of multiple classes. By continuously optimizing the model structure and parameters, these metrics are made optimal to ensure the detection performance and accuracy of the model. To fully utilize the advantages of the improved YOLOv8 architecture, the present invention continuously iterates the model structure and parameter configuration to achieve the best balance among these three metrics. The calculation formulas for average precision AP, precision P, recall R, harmonic mean F1, and mean average precision mAP are as follows:

[0081] Precision P:

[0082]

[0083] Recall R:

[0084]

[0085] Average precision and mean average precision value mAP:

[0086]

[0087] Among them, TP - the number of positive examples of this target and predicted as positive examples in the object detection task; FP - the number of negative examples of this target and predicted as positive examples in the object detection task; P(r) - the curve corresponding to both precision and recall; N - the number of classes in the multi-class object detection task.

[0088] After the model is trained, the obtained weight file is stored.

[0089] Step 5: Using the network model obtained in Step 4, extract the feature vector, and through ReLu activation and normalization operations, obtain the class activation mapping diagram, and further obtain the visualization effect diagram;

[0090] Step 5 utilizes the improved YOLOv8 model trained in Step 4 and introduces class activation mapping to enhance the visual interpretation of the model. Conventional methods such as CAM, Grad-CAM, Grad-CAM++ mostly use the feature maps of the last convolutional layer for backpropagation to obtain gradients, and use the gradient information as weights to weight the feature maps, thereby obtaining a heatmap for a specific class, mainly for interpreting and analyzing the prediction results of models based on convolutional networks. However, they do not fully address problems such as gradient disappearance and too rapid descent during the information propagation process, and the losses during the information propagation process are not located and quantified. Based on this, the present invention proposes a new visualization scheme, namely the smooth gradient class activation mapping algorithm considering losses.

[0091] In the gradient calculation, Gaussian noise is introduced to smooth the gradient. Assume that there are n' convolutional kernels in the last convolutional layer of the network, and each convolutional kernel generates a feature map. For the k-th feature map, the category c is obtained through analysis, and noise is introduced The subsequent gradient information y c (x):

[0092]

[0093] In the formula, m represents the number of iterations, n represents the number of times Gaussian noise is added, x is the input feature information, and y c represents the prediction score of the model for category c.

[0094] Then, global average pooling is performed on this gradient information in the spatial dimension M×N to obtain the weight δ of the feature map k , that is:

[0095]

[0096] In the formula, M×N represents the spatial dimension of the feature map A k The element at the i-th row and j-th column of the feature map A ij

[0097] Considering the losses during the gradient information propagation process, including the original classification loss and the interpretability metric loss, the present invention uses the following formula to consider the gradient loss problem, that is:

[0098] L k = L O-classify + λL cam (k)

[0099]

[0100] In the formula, L O-classify represents the classification loss, and L cam ​(k) represents the interpretability metric loss, camk represents the k-th feature map, and mask k represents the corresponding target region mask, represents the L2 norm, and λ represents the proportion of the interpretability metric loss.

[0101] Comprehensively consider the weight δ k and the gradient loss to further measure the importance of feature map A k for object detection, classification, and prediction. Use this weight to weight all the feature maps of the last convolutional layer, and after ReLU activation, obtain the discriminative localization map F of the object of interest, that is:

[0102]

[0103] Step 6: Use a script developed based on Python to process the infrared data and obtain a comparative analysis of the three-dimensional disease image and the model detection effect.

[0104] For the collected data, use a script developed based on Python to process it for three-dimensional temperature difference visualization, and conduct a comparative analysis with the results obtained in Step 5 to further verify the feasibility of using an infrared thermal imager for road damage detection.

[0105] Embodiment

[0106] Use an infrared thermal imager to detect about 10 km of urban roads and campus roads. Preprocess the collected infrared data according to the above steps, then conduct model design. At the same time, data augmentation and image annotation are required to make a dataset, and finally, model experiments and comparative analysis are carried out.

[0107] The results of the model experiment are shown in Table 1 below:

[0108] Table 1 Comparison of iterative evaluation indicators of each model

[0109]

[0110] It can be seen that when improving the model, it is not that the more Swin Transformer modules are added, the better the effect, and the improvement position also has a significant impact on the final detection effect. Therefore, select Yolov8+ST-3 (that is, replace the last C2f module in the Backbone with the Swin Transformer module) for subsequent model detection, experimental verification, and visualization interpretation.

[0111] Furthermore, the recognition effects of each model on the test set are compared as Figure 4 shown.

[0112] Based on rich practical experience and professional knowledge in such products, the present invention designs a rapid detection and visualization method for road apparent damage based on an infrared thermal imager, which can, to a certain extent, solve problems such as poor road damage detection effect, difficult data interpretation, large workload of target recognition and classification, and poor real-time performance under extremely low light conditions. Compared with the prior art, the model of this application uses infrared modal image data, and then improves the traditional model by introducing the Swin Transformer attention mechanism and the Class Activation Mapping (Grad-CAM) module. While better mapping more target features to the feature map, it can also well extract the feature information of some small targets in the diseases, enrich the semantic information, and thus be more applicable to the detection of small targets and weak edge information of road apparent damage. Specific Embodiment 2:

[0114] A rapid detection and visualization system for road apparent damage based on an infrared thermal imager described in this embodiment includes:

[0115] Road apparent damage image acquisition unit: used to first acquire images of road apparent damage; in this embodiment, the acquired images of road apparent damage are infrared images.

[0116] Road apparent damage detection unit: input the image into the improved YOLO model with an added attention mechanism for damage detection to obtain the detection result;

[0117] The improved YOLO model with an added attention mechanism is to replace at least one C2f module in the backbone network of the YOLO model with a Swin Transformer module. Preferably, the improved YOLO model with an added attention mechanism is to replace the last C2f module in the backbone network of the YOLO model with a Swin Transformer module. The processing process of the backbone network of the improved YOLO model with an added attention mechanism is as follows:

[0118] The input is first sent to the first CBS convolution module for processing, and then has been processed by the second CBS convolution module, the first C2f module, the third CBS convolution module, the second C2f module, the fourth CBS convolution module, the third C2f module, the fifth CBS convolution module, the first Swin Transformer module, and the SPPF module;

[0119] The CBS convolution module includes a convolutional layer, a BN layer, and a SiLU activation function.

[0120] Visualization unit: In the processing process of the improved YOLO model with an added attention mechanism, the Smooth Grad-CAM algorithm considering loss is used for visualization, and the Smooth Grad-CAM algorithm considering loss includes the following steps:

[0121] Introduce Gaussian noise in the gradient calculation to smooth the gradient. Assume that there are n' convolutional kernels in the last convolutional layer of the network, and each convolutional kernel generates a feature map. For the k-th feature map, calculate the gradient information y c as follows:

[0122]

[0123] where m represents the number of iterations, n represents the number of times Gaussian noise is added, x is the input feature information, is the introduced noise, and y c represents the prediction score of the model for class c;

[0124] Then perform global average pooling on this gradient information in the spatial dimension M×N to obtain the weight δ of the feature map k ; the weight δ of the feature map k is as follows:

[0125]

[0126] where M×N represents the spatial dimension of the feature map A k , and A ij is the element in the i-th row and j-th column of the feature map.

[0127] Calculate the loss during the propagation process:

[0128] L k = L O-classify + λL cam (k)

[0129]

[0130] where L O-classify represents the classification loss, L cam (k) represents the interpretability metric loss, cam k represents the k-th feature map, mask k represents the corresponding target region mask, represents the L2 norm, and λ represents the proportion of the interpretability metric loss;

[0131] Comprehensively consider the weight δ k and the gradient loss to obtain the discriminative localization map F of the object of interest:

[0132]

[0133] where A k represents the k-th feature map, and ReLU is the activation function.

[0134] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A rapid detection and visualization method for road surface apparent damage based on an infrared thermal imager, characterized in that First, obtain the image of the apparent damage of the road surface, and then input the image into the improved YOLO model with an attention mechanism for damage detection to obtain the detection result; The improved YOLO model with an attention mechanism is to replace at least one C2f module in the backbone network of the YOLO model with a Swin Transformer module; During the processing of the improved YOLO model with an attention mechanism, the Smooth Grad-CAM algorithm considering loss is used for visualization. The Smooth Grad-CAM algorithm considering loss includes the following steps: Introduce Gaussian noise in gradient calculation to smooth the gradient. Assume that there are n' convolutional kernels in the last convolutional layer of the YOLO model network, and each convolutional kernel generates a feature map. For the k-th feature map, calculate the gradient information y c as follows: Wherein, m represents the number of iterations, n represents the number of times of adding Gaussian noise, x is the input feature information, is the introduced noise, and y c represents the prediction score of the model for class c; Then, global average pooling is performed on the gradient information in the spatial dimension M×N to obtain the weight δ of the feature map k ; Calculate the loss during the propagation process: L k = L O-classify + λL cam (k) Where, L O-classify represents the classification loss, L cam (k) represents the interpretability metric loss, cam k represents the k-th feature map, mask k represents the corresponding target region mask, represents the L2 norm, and λ represents the proportion of the interpretability metric loss; Comprehensively consider the weight δ k and the gradient loss to obtain the discriminant localization map F of the object of interest: Among them, A k represents the k-th feature map, and ReLU is the activation function.

2. The rapid detection and visualization method for road surface apparent damage based on an infrared thermal imager according to claim 1, wherein The weight δ of the feature map k is as follows: Wherein, M×N represents the spatial dimension of the feature map A k , and A ij is the element at the i-th row and j-th column in the feature map.

3. A rapid detection and visualization method for road surface apparent damage based on an infrared thermal imager according to claim 2, characterized in that The improved YOLO model with an attention mechanism is to replace the last C2f module in the backbone network of the YOLO model with a Swin Transformer module.

4. A rapid detection and visualization method for road surface apparent damage based on an infrared thermal imager according to claim 3, characterized in that, The processing process of the backbone network of the improved YOLO model with an attention mechanism is as follows: The input is first sent to the first CBS convolution module for processing, and then has been processed by the second CBS convolution module, the first C2f module, the third CBS convolution module, the second C2f module, the fourth CBS convolution module, the third C2f module, the fifth CBS convolution module, the first Swin Transformer module, and the SPPF module; The CBS convolution module includes a convolutional layer, a BN layer, and a SiLU activation function.

5. A rapid detection and visualization method for road surface apparent damage based on an infrared thermal imager according to any one of claims 1 to 4, characterized in that, The obtained image of the apparent damage of the road surface is an infrared image.

6. A rapid detection and visualization system for road surface damage based on an infrared thermal imager, characterized in that, Including: Road surface apparent damage image acquisition unit: used to first obtain the image of the apparent damage of the road surface; Road surface apparent damage detection unit: input the image into the improved YOLO model with an attention mechanism for damage detection to obtain the detection result; The improved YOLO model with an attention mechanism is to replace at least one C2f module in the backbone network of the YOLO model with a Swin Transformer module; Visualization unit: During the processing of the improved YOLO model with an attention mechanism, the Smooth Grad-CAM algorithm considering loss is used for visualization. The Smooth Grad-CAM algorithm considering loss includes the following steps: Introduce Gaussian noise in gradient calculation to smooth the gradient. Assume that there are n' convolutional kernels in the last convolutional layer of the network, and each convolutional kernel generates a feature map. For the k-th feature map, calculate the gradient information y c as follows: Wherein, m represents the number of iterations, n represents the number of times of adding Gaussian noise, x is the input feature information, is the introduced noise, and y c represents the predicted score of the model for class c; Then, global average pooling is performed on the gradient information in the spatial dimension M×N to obtain the weight δ of the feature map k ; Calculate the loss during the propagation process: L k = L O-classify + λL cam (k) where L O-classify represents the classification loss, and L cam (k) represents the interpretability metric loss, cam k represents the k-th feature map, and mask k represents the corresponding target region mask, represents the L2 norm, and λ represents the proportion of the interpretability metric loss; Comprehensively consider the weight δ k and the gradient loss to obtain the discriminant localization map F of the object of interest: Among them, A k represents the k-th feature map, and ReLU is the activation function.

7. A rapid detection and visualization system for road surface apparent damage based on an infrared thermal imager according to claim 6, characterized in that, The weight δ of the feature map k is as follows: Wherein, M×N represents the spatial dimension of the feature map A k and A ij is the element at the i-th row and j-th column in the feature map.

8. The rapid detection and visualization system for road surface apparent damage based on an infrared thermal imager according to claim 7, characterized in that The improved YOLO model with an attention mechanism is to replace the last C2f module in the backbone network of the YOLO model with a Swin Transformer module.

9. The rapid detection and visualization system for road surface apparent damage based on an infrared thermal imager according to claim 8, characterized in that, The processing process of the backbone network of the improved YOLO model with an attention mechanism is as follows: The input is first sent to the first CBS convolution module for processing, and then has been processed by the second CBS convolution module, the first C2f module, the third CBS convolution module, the second C2f module, the fourth CBS convolution module, the third C2f module, the fifth CBS convolution module, the first Swin Transformer module, and the SPPF module; The CBS convolution module includes a convolutional layer, a BN layer, and a SiLU activation function.

10. A rapid detection and visualization method for road surface apparent damage based on an infrared thermal imager according to any one of claims 6 to 9, characterized in that, The obtained image of the apparent damage of the road surface is an infrared image.