Visible Light-Infrared Image Structure Adaptive Fusion-Based Crack Detection Method and System

Through the crack detection method of adaptive fusion of visible-infrared image structures, a dual-branch encoder and edge structure guide decoder are used, combined with a multi-mode gating attention module and a Token-KAN module, the problem of insufficient crack detection accuracy in low light and complex environments is solved, and high-precision crack detection is achieved.

CN120047432BActive Publication Date: 2025-07-01HUNAN ARCHITECTURAL DESIGN INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510484009.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-01
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing crack detection methods have low detection accuracy in low light, complex environments or harsh environments, and the modal fusion effect is poor, resulting in insufficient crack detection accuracy.

Method used

The crack detection method of adaptive fusion of visible-infrared image structure is adopted. Through the visible-infrared dual-branch encoder and edge structure guide decoder, combined with the multi-mode gating attention module and the Token-KAN module, multi-level feature extraction and fusion are realized, and the consistency estimation decision weight of edge prediction results is used for weighted fusion, and crack detection is optimized.

Benefits of technology

High-precision crack detection is achieved in low light and complex environments, which improves the accuracy and robustness of crack detection, can effectively capture different characteristics of cracks and refine the edges of cracks, significantly improving detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047432B_ABST
    Figure CN120047432B_ABST
Patent Text Reader

Abstract

The present application provides a crack detection method and system for visible light-infrared image structure adaptive fusion. The method includes: inputting the visible light image and the infrared image of the object to be detected into a trained multi-level adaptive fusion detection model; the model includes a visible light-infrared dual-branch encoder and an edge structure-guided decoder; each encoder includes n feature encoding modules for extracting encoded features of multiple levels of the visible light image and the infrared image; the visible light and infrared encoded features are fused through n multi-modal gated attention modules to output n features of different sizes; the edge structure-guided decoder correspondingly includes n feature decoding modules, which respectively decode the features of different sizes and input them into the Token-KAN module; n levels of crack prediction results are obtained through n Token-KAN modules; the n levels of crack prediction results are weighted and fused to obtain the final crack prediction result. The present application can improve the accuracy of crack detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of crack detection, and specifically relates to a crack detection method and system for visible light-infrared image structure adaptive fusion. Background Art

[0002] Segmentation and detection of crack images are of great significance in infrastructure maintenance, especially in civil engineering fields such as buildings, bridges, tunnels, and roads. Cracks are common damage forms in these structures. Besides affecting the service life of facilities, they may also pose potential public safety hazards. Therefore, it is particularly important to use advanced technologies to improve detection efficiency and reduce labor costs. The formation of cracks is usually the result of infrastructure aging, overloading, or environmental factors. As cracks expand, they may cause further damage to the structure and even lead to catastrophic consequences. Detecting and analyzing cracks in a timely manner can help engineers evaluate the structural health status and take repair or reinforcement measures. Automated crack detection technology can identify potential crack hazards at an early stage and prevent catastrophic accidents. Taking bridges as an example, automated crack detection can quickly locate weak links in the bridge structure and perform maintenance or reinforcement in a timely manner, thereby extending its service life.

[0003] Traditional crack detection usually relies on manual inspections. This method not only requires a large amount of time and manpower but is also easily affected by human factors, resulting in missed detections or false detections. Due to the complex and diverse shapes, sizes, and distributions of cracks and their frequent occurrence in different locations, it is difficult to guarantee the accuracy of manual detection.

[0004] In the crack detection task, some automated segmentation methods have been studied. These methods usually adopt an encoder-decoder structure, where the encoder is responsible for extracting multi-scale features, and the decoder is responsible for gradually upsampling to restore the spatial resolution and generate the final semantic segmentation map. Currently, the mainstream deep learning methods include U-Net (U-Net: ConyolutionalNetworks for Biomedical Image Segmentation. Springer International Publishing(2015): 234-241) and its variants. U-Net is a network with an encoder-decoder structure and is widely used in medical imaging, remote sensing imaging, and crack detection tasks. Its basic principle is to gradually extract multi-scale features through convolutional layers and pooling layers, and gradually restore the spatial resolution of the image through the decoder, and generate the final semantic segmentation map. In addition, CrackFormer-II (CrackFormer Network for Pavementrack Segmentation. lEEE Transactions onIntelligent Transportation Systems (2023): 9240-9252) enhances the global feature modeling ability through the Transformer structure and combines U-Net-style skip connections to restore high-resolution spatial information, thereby generating the final semantic segmentation map.

[0005] In the multi-modal crack detection task, the well-known technology usually adopts the modal fusion method to make full use of the complementarity between visible light and infrared information. For example, the dual-encoder architecture based on ResNet has also been used in the crack detection task, such as GMNet (Graded-feature multilabel-learning network for RGB-thermal urban scenesemantic segmentation, lEEE Transactions on Image Processing (2021):7790-7802) and SGFNet (SGFNet: Semantic-Guided Fusion Network for RGB-Thermal SemanticSegmentation. IEEE Transactions onCircuits and Systems for Video Technology(2023):7737-7748). These methods design independent encoders for the visible light and infrared branches, share features in the low-level encoders, extract features independently in the high-level encoders, and finally generate the semantic segmentation result map through the decoder.

[0006] Regarding the existing crack detection methods, although models such as U-Net, CrackFormer-II, GMNet, and SGFNet have made certain progress in the semantic segmentation task of visible light and infrared crack images, there are still the following technical problems: The U-Net and CrackFormer-II networks are designed for the segmentation of visible light images, resulting in a decline in detection performance in low-light, complex, or harsh environments, and possible misclassification problems; The GMNet and SGFNet networks perform feature fusion through simple fusion modules during the modal fusion process, and the fusion effect is not good. The above problems will affect the accuracy of crack detection. Therefore, it is necessary to provide a crack detection method with better performance. Summary of the Invention

[0007] The purpose of this application is to provide a crack detection method and system for adaptive fusion of visible light-infrared image structures, which can improve the accuracy of crack detection.

[0008] In the first aspect, this application provides a crack detection method for adaptive fusion of visible light-infrared image structures, including:

[0009] Collect the visible light image and infrared image of the object to be detected;

[0010] Input the visible light image and the infrared image into the trained multi-level adaptive fusion detection model for visible light-infrared image cracks; the multi-level adaptive fusion detection model for visible light-infrared image cracks outputs the crack prediction result of the object to be detected;

[0011] Among them, the multi-level adaptive fusion detection model for visible light-infrared image cracks includes a visible light-infrared double-branch encoder and an edge structure-guided decoder;

[0012] The visible light-infrared double-branch encoder includes a visible light branch encoder and an infrared branch encoder;

[0013] Among them, both the visible light branch encoder and the infrared branch encoder include n feature encoding modules connected in sequence, which are used to extract the encoded features of multiple levels of the visible light image and the infrared image;

[0014] The visible light-infrared double-branch encoder also includes n multi-modal gated attention modules;

[0015] The output of the i-th feature encoding module of the visible light branch encoder and the output of the i-th feature encoding module of the infrared branch encoder serve as the input of the i-th multi-modal gated attention module; among them, ;

[0016] The n multi-modal gated attention modules respectively output n features of different sizes;

[0017] The edge structure-guided decoder correspondingly includes n feature decoding modules connected in sequence;

[0018] The output of the i-th multi-modal gated attention module serves as the input of the i-th feature decoding module of the edge structure-guided decoder;

[0019] The edge structure-guided decoder also includes n Token-KAN modules;

[0020] The output of the i-th feature decoding module of the edge structure-guided decoder serves as the input of the i-th Token-KAN module;

[0021] The crack prediction results of n levels are obtained through the n Token-KAN modules; the crack prediction results of n levels are weighted and fused to obtain the final crack prediction result.

[0022] In a possible implementation, each feature encoding module of the visible light branch encoder and the infrared branch encoder includes a convolutional block and a downsampling module connected in sequence; each feature decoding module of the edge structure-guided decoder includes a convolutional block and an upsampling module connected in sequence.

[0023] In a possible implementation, each multi-modal gated attention module includes five convolutional layers, three gated fusion modules, and one squeeze-and-excitation attention module;

[0024] Among them, the five convolutional layers are the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer respectively;

[0025] The three gated fusion modules are the first gated fusion module, the second gated fusion module, and the third gated fusion module respectively;

[0026] For the i-th multi-modal gated attention module, the is input into the first convolutional layer for feature extraction to obtain visible light features ; the is input into the second convolutional layer for feature extraction to obtain infrared features ; the and are jointly input into the third convolutional layer for feature extraction to obtain concatenated features , where , means concatenating two groups of features with the same dimension along the channel dimension;

[0027] The squeeze-and-excitation attention module includes a global pooling layer, a 1×1 convolutional layer, a RELU activation function, a 1×1 convolutional layer, and a Sigmoid activation function connected in sequence; the Sigmoid activation function outputs channel attention weights; based on the channel attention weights, the mixed features are weighted channel by channel to obtain enhanced features ;

[0028] The first gated fusion module weights and fuses the visible light features and the features to obtain the first fusion feature ;

[0029] The second gated fusion module weights and fuses the infrared features and the features to obtain the second fusion feature ;

[0030] and are respectively subjected to convolution extraction through the fourth convolutional layer and the fifth convolutional layer to obtain and ; the third gated fusion module fuses and to obtain the third fusion feature .

[0031] In a possible implementation, the first gated fusion module weights and fuses the visible light features and the features to obtain the first fusion feature , and the calculation formula is:

[0032] ;

[0033] The second gated fusion module weights and fuses the infrared features and the features to obtain the second fusion feature , and the calculation formula is:

[0034] ;

[0035] Among them, is the gated weight calculated by the first gated fusion module, ; is the gated weight calculated by the second gated fusion module, ; among them, represents the Sigmoid activation function;

[0036] and are respectively obtained by convolution extraction through the fourth convolutional layer and the fifth convolutional layer and ; The third gated fusion module fuses and to obtain the third fusion feature , and the calculation formula is:

[0037] ;

[0038] Among them, is the gated weight calculated by the third gated fusion module, ;

[0039] The obtained is the output of the i-th multi-modal gated attention module.

[0040] In a possible implementation, the Token-KAN module includes a patch embedding layer, a tokenization layer, a KAN network layer, a depthwise separable convolutional layer, and layer normalization connected in sequence; after performing a residual connection on the output of the depthwise separable convolutional layer and the output of the patch embedding layer, the features with a normalized distribution are obtained through layer normalization.

[0041] In a possible implementation, the crack prediction result includes a crack edge prediction result and a crack area prediction result;

[0042] The visible light-infrared multi-modal crack dataset is used to train the constructed multi-level adaptive fusion detection model for visible light-infrared image cracks, and the trained multi-level adaptive fusion detection model for visible light-infrared image cracks is obtained; during the training process, the loss function adopted is the joint loss function for edge and crack detection, including the mean square error loss function and the focal loss function;

[0043] The calculation formula of the mean square error loss function is as follows:

[0044]

[0045] where, represents the mean square error loss of the crack edge prediction result of the i-th level, N is the number of training samples, is the crack edge prediction result of the i-th level corresponding to the k-th training sample, is the true crack edge label corresponding to the k-th training sample;

[0046] The calculation formula of the focal loss function is as follows:

[0047]

[0048] where, N is the number of training samples, is the weighting factor, is the focusing parameter, is the final crack prediction result of the k-th training sample, is the true image label of the k-th training sample;

[0049] The joint loss function for edge and crack detection, that is, the total loss function The calculation formula is as follows:

[0050] ;

[0051] where, and are the weights of the mean square error loss and the focal loss respectively.

[0052] In a possible implementation manner, the weighted fusion of the crack prediction results of n levels to obtain the final crack prediction result includes:

[0053] First, estimate the weights of the crack prediction results of n levels based on the consistency of the edge prediction results;

[0054] Based on the mean square error loss of the crack edge prediction result of the i-th level, calculate the inverse loss value of the i-th level:

[0055] ;

[0056] Sum the inverse loss values of each level to obtain the total inverse loss : :

[0057] ;

[0058] Calculate the weight of the crack prediction result of the i-th level :

[0059] ;

[0060] Then, for the visible light image and the infrared image of the object to be detected input, based on the weight Obtain the final crack prediction result of the object to be detected :

[0061] ;

[0062] Wherein is the crack region prediction result of the i-th level corresponding to the visible light image and the infrared image of the object to be detected.

[0063] In a second aspect, the present application provides a system, including: a memory and a processor;

[0064] The memory is used to store a computer program;

[0065] The processor is used to call the computer program to execute the method as described above.

[0066] In a third aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on an electronic device, the electronic device is enabled to implement the method as described above.

[0067] In a fourth aspect, the present application provides a computer program product, including a computer program. When the computer program runs on an electronic device, the electronic device is enabled to implement the method as described above.

[0068] For the specific implementation manners of the second to fourth aspects of the present application, reference may be made to the implementation manner of the first aspect above, and details are not described herein again.

[0069] Beneficial effects:

[0070] This application extracts multi-level features of visible light and infrared images through a dual-branch convolutional neural network encoder, effectively utilizing the complementary information of the two modalities, thereby enabling high-precision crack detection in low-light, complex, or harsh environments. The encoder adopts a visible light-infrared dual-branch structure to extract features from visible light images and infrared images respectively, fully mining the crack information in both modalities, and effectively enhancing the model's crack detection ability under different lighting and environmental conditions. The crack features are captured through layer-by-layer convolution and downsampling modules, effectively enhancing the model's feature extraction ability for complex crack morphologies.

[0071] This application realizes the adaptive fusion of multi-level features of different modality images through a multi-modal gating attention module, achieving the interactive utilization and collaborative learning of visible light-infrared image information, and improving the recognition accuracy of crack detection. By using the attention mechanism, it can adaptively focus on key feature regions while suppressing irrelevant background noise. By using the gating fusion module, the model can dynamically adjust the fusion ratio of visible light features and infrared features to achieve efficient integration of information between modalities. By adopting the method of multi-level feature fusion, the model can effectively capture different features of cracks, thereby improving the segmentation detection accuracy. The visible light-infrared dual-branch encoder realizes deep feature extraction and fuses crack information from different sources through the effective fusion of multi-level visible light-infrared features, improving the feature expression ability, and providing rich spatial structure information for the subsequent decoder.

[0072] This application designs a decoder guided by edge structure information for simultaneously predicting crack regions and crack edges. The edge structure-guided decoder helps to accurately segment crack edges by upsampling layer by layer and fusing the features of the corresponding layers of the encoder (skip connection). The decoder uses multi-level upsampling and edge-guided strategies to effectively refine crack edges and achieve high-precision crack edge segmentation. The decoder can estimate decision weights based on the consistency of edge prediction results to guide the weighted fusion of multi-level crack prediction results. It guides the network to optimize the learning process through the joint loss function of edge and crack detection, guiding the model to focus on the edge details of cracks. Combining the self-learning ability of deep networks and the prior knowledge of edge structures, through structure-adaptive perception of multi-modal feature fusion and decision fusion, it optimizes the crack edge detail prediction results and further improves the boundary detection accuracy, thus significantly improving the accuracy of crack detection. Description of the Drawings

[0073] Figure 1 It is a schematic diagram of the basic process of the method in the embodiment of this application.

[0074] Figure 2 It is a framework diagram of the visible light-infrared image crack multi-level adaptive fusion detection model provided by the embodiment of this application.

[0075] Figure 3 Schematic diagram of the convolutional block structure in the embodiment of the present application.

[0076] Figure 4 Schematic diagram of the multi-modal gated attention module structure in the embodiment of the present application.

[0077] Figure 5 Schematic diagram of the Token-KAN module structure in the embodiment of the present application.

[0078] Figure 6 Visualization comparison chart of the experimental results of the embodiment of the present application and other models on the same dataset. Detailed implementation manners

[0079] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution of the present application will be further described in detail below in conjunction with the embodiments of the present application and the accompanying drawings.

[0080] The present application proposes a crack detection method and system for visible light-infrared image structure adaptive fusion. The present application performs crack detection based on a visible light-infrared image crack multi-level adaptive fusion detection model. The visible light-infrared image crack multi-level adaptive fusion detection model includes a visible light-infrared double-branch encoder and an edge structure-guided decoder. The encoder includes multiple levels of feature encoding modules, and each feature encoding module includes a convolutional block and a downsampling module; the encoder also includes a multi-modal gated attention module, which introduces an edge-guided adaptive gating fusion strategy to learn the crack structure features with variable shapes and different scales, further improving the refinement degree of crack structure extraction and the detection accuracy and robustness of road cracks. The edge structure-guided decoder correspondingly includes multiple levels of feature decoding modules, and each feature decoding module includes a convolutional block and an upsampling module. The edge structure-guided decoder also includes multiple Token-Kolmogorov-Arnold (Token-KAN) modules. Using this module can enable the model proposed by the present application to more accurately capture slender crack features and maintain the coherence of the crack structure. This module can also reduce the computational complexity and reduce the interference of background noise. Using the edge and crack detection joint loss function, the decision weights are estimated based on the consistency of the edge prediction results, and the weighted fusion of the multi-level crack prediction results is guided to obtain the crack prediction result. The technical solution of the present application will be specifically described below in conjunction with specific embodiments.

[0081] Embodiment 1

[0082] As Figure 1 shown, the embodiment of the present application provides a crack detection method for visible light-infrared image structure adaptive fusion, including:

[0083] Collect the visible light image and infrared image of the object to be detected;

[0084] Input the visible light image and the infrared image into the trained multi-level adaptive fusion detection model for visible light-infrared image cracks; the multi-level adaptive fusion detection model for visible light-infrared image cracks outputs the crack prediction result of the object to be detected;

[0085] Among them, the multi-level adaptive fusion detection model for visible light-infrared image cracks includes a visible light-infrared double-branch encoder and an edge structure-guided decoder;

[0086] The visible light-infrared double-branch encoder includes a visible light branch encoder and an infrared branch encoder;

[0087] Among them, both the visible light branch encoder and the infrared branch encoder include n feature encoding modules connected in sequence, which are used to extract the encoded features of multiple levels of the visible light image and the infrared image;

[0088] The visible light-infrared double-branch encoder also includes n multi-modal gated attention modules;

[0089] The output of the i-th feature encoding module of the visible light branch encoder and the output of the i-th feature encoding module of the infrared branch encoder serve as the input of the i-th multi-modal gated attention module; among them, ;

[0090] The n multi-modal gated attention modules respectively output n features of different sizes;

[0091] The edge structure-guided decoder correspondingly includes n feature decoding modules connected in sequence;

[0092] The output of the i-th multi-modal gated attention module serves as the input of the i-th feature decoding module of the edge structure-guided decoder;

[0093] The edge structure-guided decoder also includes n Token-KAN modules;

[0094] The output of the i-th feature decoding module of the edge structure-guided decoder serves as the input of the i-th Token-KAN module;

[0095] The crack prediction results of n levels (stages) are obtained through the n Token-KAN modules; the crack prediction results of the n levels are weighted and fused to obtain the final crack prediction result.

[0096] In some embodiments, the crack prediction results of n levels (stages) are obtained through the n Token-KAN modules, including that the outputs of the n Token-KAN modules pass through an upsampling module and a convolutional block in sequence to obtain the crack prediction results of n levels (stages). Among them, the upsampling module upsamples the feature size to the same size and obtains its binary classification probability map through a 1×1 convolutional layer, which is convenient for subsequent calculation of the loss function.

[0097] It should be understood that the sequential connection means that the input of the previous module / layer serves as the input of the subsequent module / layer.

[0098] In some embodiments, each feature encoding module of the visible light branch encoder and the infrared branch encoder may include a convolutional block and a downsampling module connected in sequence; correspondingly, each feature decoding module of the edge structure guiding decoder may include a convolutional block and an upsampling module connected in sequence.

[0099] In some embodiments, n = 4. The framework diagram of the visible light-infrared image crack multi-level adaptive fusion detection model corresponding to this embodiment is as Figure 2 shown.

[0100] In some embodiments, the specific structure of the convolutional block included in the encoder is as Figure 3 shown, and each convolutional block includes two 3×3 convolutional layers, two batch normalization layers, and two RELU activation functions.

[0101] Specifically, as Figure 2 shown, the core module of the visible light-infrared dual-branch encoder is the multi-modal gated attention module, and the specific structure of the multi-modal gated attention module is as Figure 4 shown. Each multi-modal gated attention module includes five convolutional layers, three gated fusion modules, and one squeeze-and-excitation attention module;

[0102] Among them, the five convolutional layers are the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer respectively;

[0103] The three gated fusion modules are the first gated fusion module, the second gated fusion module, and the third gated fusion module respectively;

[0104] For the i-th multi-modal gated attention module, the input passes through the first convolutional layer for feature extraction to obtain visible light features ; the input passes through the second convolutional layer for feature extraction to obtain infrared features ; the and The third convolutional layer is jointly input for feature extraction to obtain concatenated features , where , denotes concatenating two sets of features with the same dimension along the channel dimension.

[0105] The multi-modal gated attention module utilizes the squeeze-and-excitation attention mechanism to extract the weights between channels through global pooling (used to squeeze spatial information and provide global context information) and convolutional operations, applies channel attention, and strengthens key information. Specifically, the squeeze-and-excitation attention module includes a global pooling layer, a 1×1 convolutional layer, a ReLU activation function, a 1×1 convolutional layer, and a Sigmoid activation function connected in sequence; the squeeze-and-excitation attention module squeezes each channel through the global pooling layer, then uses a 1×1 convolution for dimensionality reduction and feature extraction, increases non-linearity through the ReLU activation function, and then uses a 1×1 convolution for dimensionality increase to restore the original number of channels. Finally, channel attention weights are generated through the Sigmoid activation function. Based on the generated channel attention weights, the mixed features are weighted channel by channel to obtain enhanced features , and the enhanced features can highlight the key features of the concatenated multi-channel features;

[0106] The multi-modal gated attention module uses the gated mechanism to dynamically balance the weights of the two sets of features, generates a gated weight between 0 and 1 through the Sigmoid activation function, and the gated weight determines the fusion ratio of the two sets of features.

[0107] The first gated fusion module weights and fuses the visible light features and the features to obtain the first fused feature :

[0108] ;

[0109] The second gated fusion module weights and fuses the infrared features and the features to obtain the second fused feature :

[0110] ;

[0111] where is the gated weight calculated by the first gated fusion module, ; is the gated weight calculated by the second gated fusion module, ; where denotes the Sigmoid activation function;

[0112] and are respectively subjected to convolution extraction through the fourth convolutional layer and the fifth convolutional layer to obtain and ; The third gated fusion module fuses and to obtain the third fusion feature :

[0113] ;

[0114] wherein, is the gating weight calculated by the third gated fusion module, ;

[0115] The obtained is the output of the i-th multi-modal gated attention module.

[0116] As Figure 5 shown, the core module of the edge structure-guided decoder is the Token-KAN module. The Token-KAN module includes a patch embedding layer, a tokenization layer, a KAN network layer, a depthwise separable convolutional layer, and a layer normalization connected in sequence. The input features are divided into multiple small blocks through patch embedding and mapped into a high-dimensional space to form preliminary embedded features. This process is equivalent to transforming local regions into higher-level representations, which helps capture the local structural features of cracks. The input is transformed into structured semantic units through the tokenization layer, enabling efficient attention calculation and feature fusion within the framework of the Kolmogorov-Arnold (KAN) module. The KAN network layer deeply models the embedded features through mathematical mapping (such as Kolmogorov-Arnold representation), strengthens the characterization of crack morphology, and combines depth convolution and non-linear mapping to efficiently extract the edge information and morphological details of cracks. In feature processing, a lightweight depthwise separable convolutional layer is adopted to further reduce the number of parameters and computational complexity while maintaining the expressive power of the model. Finally, after performing a residual connection between the output of the depthwise separable convolutional layer and the output of the patch embedding layer, the features with a normalized distribution are obtained through layer normalization to avoid gradient vanishing, accelerate the convergence process, and improve the model stability.

[0117] In some embodiments, the crack prediction results include crack edge prediction results and crack region prediction results.

[0118] In some embodiments, the constructed visible-infrared image crack multi-level adaptive fusion detection model is trained using a visible-infrared multi-modal crack dataset to obtain a trained visible-infrared image crack multi-level adaptive fusion detection model. During the training process, the loss function used is the edge and crack detection joint loss function, including the mean square error loss function and the focal loss function (Focal Loss function). The mean square error loss function is mainly used to optimize the pixel prediction results in the edge detection task, making the prediction of the crack edge more accurate. Here, the mean square error loss function is used to calculate the error between the crack edge prediction result and the true crack edge label. Its calculation formula is as follows:

[0119]

[0120] where, represents the mean square error loss of the crack edge prediction result at the i-th level, N is the number of training samples, is the crack edge prediction result at the i-th level corresponding to the k-th training sample, is the true crack edge label corresponding to the k-th training sample;

[0121] The focal loss function makes the model pay more attention to the prediction of the crack area because these areas are usually small and have a small proportion in the data. By increasing the loss of difficult-to-classify areas (such as crack areas), the focal loss function helps to improve the detection ability of the model, especially when the contrast between the background and the crack area is very unbalanced. It avoids the background class (i.e., non-crack area) occupying too much training attention. Its calculation formula is as follows:

[0122]

[0123] where, N is the number of training samples, is the weighting factor, is the focusing parameter, is the final crack prediction result of the k-th training sample, is the true image label of the k-th training sample; the true image label can take values of 0 or 255, where 0 corresponds to the non-crack area and 255 corresponds to the crack area.

[0124] In some embodiments, since the proportion of positive samples (crack areas) is small, a larger weight is given, and the proportion of negative samples (non-crack areas) is large. The weighting factor = 0.25.

[0125] In some embodiments, = 2. Taking = 2 in the binary semantic segmentation task can enhance the learning ability of the model.

[0126] The combined loss function for edge and crack detection, i.e., the total loss function has the following calculation formula:

[0127] ;

[0128] Among them, and are the weights of the mean squared error loss and the focal loss respectively.

[0129] In some embodiments, the weights of the crack prediction results of n levels are estimated based on the consistency of the edge prediction results, so as to guide the weighted fusion of the crack prediction results of n levels to obtain the final crack prediction result, which specifically includes:

[0130] First, estimate the weights of the crack prediction results of n levels based on the consistency of the edge prediction results;

[0131] reflects the error between the crack edge prediction result of the i-th level and the true crack edge label. To avoid numerical instability, take the reciprocal of to obtain the inverse loss value of the i-th level, and its calculation formula is as follows:

[0132] ;

[0133] Sum the inverse loss values of levels to obtain the total inverse loss , and its calculation formula is as follows:

[0134] ;

[0135] The weight of the crack prediction result of the i-th level has the following calculation formula:

[0136] ;

[0137] Since the weight is proportional to the inverse loss value, the better the prediction effect (the smaller the loss), the greater the weight of the level.

[0138] The final crack prediction result of the k-th training sample is generated by weighted summation, and its calculation formula is as follows:

[0139] ;

[0140] Among them, is the crack region prediction result of the i-th level corresponding to the k-th training sample;

[0141] For the visible light image and infrared image of the object to be detected in the input model, the final crack prediction result of the object to be detected output by the model is as follows:

[0142] ;

[0143] wherein, is the crack region prediction result of the i-th layer corresponding to the visible light image and infrared image of the object to be detected.

[0144] The crack detection method provided by the above embodiment has the following characteristics:

[0145] (1) A visible light-infrared multimodal image structure adaptive fusion crack detection method based on a convolutional neural network is proposed. This method adopts an encoder-decoder architecture, extracts the features of visible light and infrared images through a dual-branch encoder, and then decodes the features through a decoder. The prediction result is optimized through an adaptive weighted fusion strategy, thereby effectively improving the accuracy of crack detection.

[0146] (2) A multimodal gated attention fusion module is proposed. This module adaptively fuses visible light and infrared features, and uses a gated mechanism to dynamically adjust the fusion ratio of the two modal features, thereby realizing efficient information integration. At the same time, by introducing a multi-level feature fusion strategy, the model can fully capture multi-scale crack features and further improve the detection effect.

[0147] (3) An edge structure-guided decoder is proposed, which simultaneously predicts the crack region and the crack edge. The decision weight is further guided by the consistency estimation based on the edge prediction result to weight the fusion of multi-layer crack prediction results. This method effectively refines the crack edge, thereby realizing high-precision crack edge segmentation.

[0148] As Figure 6 shown, compared with the prior art, it mainly has the following advantages:

[0149] The classic U-Net network is designed for visible light images, and its network architecture only contains a visible light branch. The network architecture of this application adopts a dual-branch convolutional neural network encoder, which extracts multi-level features of visible light and infrared images through two branches respectively, effectively utilizing the complementary information of the two modalities, so as to achieve high-precision crack detection in low-light, complex or harsh environments. The U-Net network extracts features through the encoder. On this basis, this application introduces a multi-modal gated attention fusion module to achieve adaptive fusion of multi-level features of different modal images. The U-Net network directly generates a crack segmentation map through the decoder's upsampling. This application designs a decoder guided by edge structure information, which can simultaneously predict the crack area and the crack edge, and estimate the decision weights based on the consistency of the edge prediction results, guiding the weighted fusion of multi-level crack prediction results. This application can guide the network to optimize the learning process through the joint loss function of edge and crack detection, thus significantly improving the accuracy of crack detection.

[0150] The CrackFormer-II network is also designed for the segmentation of visible light images, resulting in a decline in detection performance in low-light, complex or harsh environments, and possible misclassification problems.

[0151] In the process of modal fusion, the GMNet and SGFNet networks perform feature fusion through a simple fusion module, and the fusion effect is not good. This application introduces a multi-modal gated attention module to achieve adaptive fusion of multi-level features of different modal images, realizes the interactive utilization and collaborative learning of visible light-infrared image information, and improves the recognition accuracy of crack detection.

[0152] When dealing with small and discontinuous targets such as cracks, the existing technologies may cause problems such as blurred or broken edges. This application designs a decoder guided by edge structure information to simultaneously predict the crack area and the crack edge, and estimate the decision weights based on the consistency of the edge prediction results, guiding the weighted fusion of multi-level crack prediction results. A joint loss function for edge and crack detection is designed to guide the network to optimize the learning process, thus significantly improving the accuracy of crack detection. This application combines the self-learning ability of deep networks and the prior knowledge of edge structures, and optimizes the crack edge detail prediction results through structure-adaptive perception of multi-modal feature fusion and decision fusion.

[0153] The automated crack segmentation and detection technology based on deep learning in this application can process a large amount of image data in a short time, significantly improving the detection efficiency. This automated system can work continuously for 24 hours, quickly identifying the specific location, length, width, and depth of cracks, thus saving a large amount of time and effort for manual inspection. Through the application of convolutional neural networks (CNNs) in image segmentation, the deep learning model can efficiently process complex crack images. With the support of a large amount of labeled data, the model can automatically learn crack features and reduce human errors. Compared with traditional methods, the deep learning model is more accurate and robust in identifying cracks under complex backgrounds, greatly improving the accuracy of detection.

[0154] It should be noted that the terms "first", "second", etc. in the specification, claims, and above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of this application described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0155] Experiments were conducted using the Asphalt Pavement Crack Detection Based on Convolutional Neural Network and Infared Thermography dataset. The training set contains 358 groups of crack images, and the test set contains 90 groups of crack images. Each group of images includes a visible light image, an infrared image, a fused image, and a label image. In this experiment, the Adam optimizer was used, with an initial learning rate of 0.0001 and a batch size of 4.

[0156] As Figure 6 shown, it is the visualization result of crack segmentation. From left to right are the visible light image, infrared image, fused image, label image, U-Net, Crackformer-II, GMNet, SGFNet, and the crack segmentation results of the method of the present invention. It can be seen from Figure 6 this that for different types of cracks, the method of the present invention has achieved more accurate segmentation and has the highest accuracy in detecting the edge shape, width, etc. of cracks.

[0157] Embodiment 2

[0158] This embodiment provides an electronic device, including: a memory and a processor;

[0159] The memory is used for storing a computer program;

[0160] The processor is used for calling the computer program to execute the method described in Embodiment 1.

[0161] Embodiment 3

[0162] This embodiment provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program runs on an electronic device, the electronic device implements the method described in Embodiment 1.

[0163] Embodiment 4

[0164] This embodiment provides a computer program product, including a computer program. When the computer program runs on an electronic device, the electronic device implements the method described in Embodiment 1.

[0165] The specific implementation manners of a system, an electronic device, a computer-readable storage medium, and a computer program product provided in the embodiments of this application may refer to the specific embodiments of the above method, and will not be elaborated herein.

[0166] Obviously, those skilled in the art should understand that the above units or steps of this application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, this application is not limited to any specific combination of hardware and software.

[0167] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.

Claims

1. A crack detection method based on adaptive fusion of visible light and infrared image structure, characterized in that: include: Collecting visible light images and infrared images of the object to be detected; Inputting the visible light image and the infrared image into a trained visible light-infrared image crack multi-level adaptive fusion detection model; the visible light-infrared image crack multi-level adaptive fusion detection model outputs the crack prediction result of the object to be detected; Wherein, the visible light-infrared image crack multi-level adaptive fusion detection model includes a visible light-infrared dual-branch encoder and an edge structure guided decoder; The visible light-infrared dual-branch encoder includes a visible light branch encoder and an infrared branch encoder; The visible light branch encoder and the infrared branch encoder each include n feature encoding modules connected in sequence, which are used to extract encoding features of multiple levels of the visible light image and the infrared image; The visible light-infrared dual-branch encoder also includes n multi-mode gated attention modules; The output of the i-th feature encoding module of the visible light branch encoder and the output of the ith feature encoding module of the infrared branch encoder As the input of the i-th multi-modal gated attention module; where, ; The n multi-mode gated attention modules output n features of different sizes respectively; The edge structure guided decoder correspondingly comprises n feature decoding modules connected in sequence; The output of the i-th multi-mode gated attention module is used as the input of the i-th feature decoding module of the edge structure guided decoder; The edge structure guided decoder also includes n Token-KAN modules; The output of the i-th feature decoding module of the edge structure guided decoder is used as the input of the i-th Token-KAN module; The n Token-KAN modules are used to obtain n-level crack prediction results; the n-level crack prediction results are weighted and fused to obtain the final crack prediction result.

2. The method according to claim 1, characterized in that Each feature encoding module of the visible light branch encoder and the infrared branch encoder includes a convolution block and a downsampling module connected in sequence; each feature decoding module of the edge structure guided decoder includes a convolution block and an upsampling module connected in sequence.

3. The method according to claim 1, characterized in that Each multi-mode gated attention module consists of five convolutional layers, three gated fusion modules, and a compression and excitation attention module; Among them, the five convolutional layers are the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer and the fifth convolutional layer; The three gated fusion modules are respectively a first gated fusion module, a second gated fusion module and a third gated fusion module; For the i-th multi-mode gated attention module, the Input the first convolutional layer for feature extraction to obtain visible light features ; Input the second convolutional layer for feature extraction to obtain infrared features ; and Input the third convolution layer together for feature extraction to obtain the splicing features ,in , It means concatenating two sets of features with the same dimension along the channel dimension; The compression and excitation attention module includes a global pooling layer, a 1×1 convolution layer, a RELU activation function, a 1×1 convolution layer and a Sigmoid activation function connected in sequence; the Sigmoid activation function outputs a channel attention weight; based on the channel attention weight, the mixed feature The enhanced features are obtained by channel-by-channel weighting ; The first gated fusion module weightedly fuses the visible light features and Features Get the first fusion feature ; The second gated fusion module weightedly fuses infrared features and Features Get the second fusion feature ; and The features are obtained by convolution extraction through the fourth and fifth convolution layers respectively. and ; The third gated fusion module fusion and Get the third fusion feature .

4. The method according to claim 3, characterized in that The first gated fusion module weightedly fuses the visible light features and Features Get the first fusion feature , the calculation formula is: ; The second gated fusion module weightedly fuses infrared features and Features Get the second fusion feature , the calculation formula is: ; in, is the gating weight calculated by the first gating fusion module, ; is the gating weight calculated by the second gating fusion module, ;in, Represents the Sigmoid activation function; and Convolution extraction is performed through the fourth and fifth convolution layers respectively. and ; The third gated fusion module fusion and Get the third fusion feature , the calculation formula is: ; in, is the gating weight calculated by the third gating fusion module, ; The resulting That is the output of the i-th multi-mode gated attention module.

5. The method according to claim 1, characterized in that The Token-KAN module includes a patch embedding layer, a tokenization layer, a KAN network layer, a depthwise separable convolution layer and layer normalization, which are connected in sequence. After residual connection is performed on the output of the depthwise separable convolution layer and the output of the patch embedding layer, the features of the canonical distribution are obtained through layer normalization.

6. The method according to any one of claims 1 to 5, characterized in that The crack prediction results include crack edge prediction results and crack area prediction results; The constructed visible light-infrared multimodal crack dataset is used to train the constructed visible light-infrared image crack multi-level adaptive fusion detection model to obtain a trained visible light-infrared image crack multi-level adaptive fusion detection model; during the training process, the loss function used is the joint loss function of edge and crack detection, including the mean square error loss function and the focus loss function; The mean square error loss function calculation formula is as follows: ; in, represents the mean square error loss of the crack edge prediction result at the i-th level, N is the number of training samples, is the crack edge prediction result of the i-th level corresponding to the k-th training sample, is the true crack edge label corresponding to the kth training sample; The focal loss function calculation formula is as follows: ; Where N is the number of training samples, is the weighting factor, is the focusing parameter, is the final crack prediction result of the kth training sample, is the real image label of the kth training sample; The joint loss function of edge and crack detection, that is, the total loss function The calculation formula is as follows: ; in, and are the weights of mean square error loss and focal loss, respectively.

7. The method according to claim 6, characterized in that The weighted fusion of the crack prediction results at the n levels to obtain the final crack prediction result includes: Firstly, the weights of the crack prediction results at n levels are estimated based on the consistency of the edge prediction results; Mean square error loss of crack edge prediction results based on the i-th level , calculate the inverse loss value of the i-th level : ; Will The total inverse loss is obtained by summing the inverse loss values ​​of each level. : ; Calculate the weight of the crack prediction results at the i-th level : ; Then, for the input visible light image and infrared image of the object to be detected, based on the weights Get the final crack prediction result of the object to be tested : ; in, It is the crack area prediction result of the i-th level corresponding to the visible light image and infrared image of the object to be detected.

8. A crack detection system with adaptive fusion of visible light and infrared image structure, characterized in that: include: Memory and processor; The memory is used to store computer programs; The processor is configured to call the computer program to execute the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed on an electronic device, the electronic device implements the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed on an electronic device, the electronic device implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Unmanned aerial vehicle multi-modal data fusion curtain wall supporting block state detection method and system

    CN118658087A

  • Construction method of road crack segmentation system fusing infrared and visible light images

    CN118762264A