Construction Method of Road Crack Segmentation System Integrating Infrared and Visible Light Images
By improving the RTFormer model, introducing infrared image branches and feature fusion modules, combined with knowledge distillation technology, the accuracy and computational burden problems of crack detection in complex environments are solved, and efficient crack identification and detection are achieved.
Patent Information
- Application Number
- CN202410904750.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-07-08
AI Technical Summary
The prior art lacks the accuracy and robustness of the identification of fine cracks under different lighting and weather conditions, and the traditional feature fusion method has information redundancy and calculation burden.
Improved RTFormer model, introduce infrared image branches, use visible light and infrared fusion module IFF and improved DAPPM module, and combine knowledge distillation technology for training to optimize feature fusion and model design.
Improve the accuracy and reliability of crack detection under different lighting and weather conditions, reduce model calculation requirements, and is suitable for on-site applications of resource-constrained equipment.
Smart Images

Figure CN118762264B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a construction method of a crack segmentation system that fuses infrared and visible light images, and belongs to the technical field of crack detection. Background Art
[0002] With the rapid development of transportation infrastructure, the apparent disease detection of roads and bridges has become particularly important. The early detection of cracks plays a crucial role in extending the service life of roads and ensuring driving safety. In recent years, image segmentation technology based on deep learning has achieved remarkable results in the field of crack detection. Image segmentation technology in the field of computer vision is a fundamental and key technology. Especially in the inspection of roads and bridges, high-precision image segmentation tasks such as crack detection are crucial for early identification and prevention of potential structural problems. Although the application of deep learning technology, especially convolutional neural networks (CNNs) and Transformer models, has significantly improved the accuracy of image segmentation, the detection accuracy and robustness under complex lighting and diverse environmental conditions still face challenges. These models still have deficiencies in processing images in complex environments, such as the identification of small cracks under different lighting and weather conditions.
[0003] Existing research has paid less attention to how to effectively fuse infrared and visible light images to improve crack detection performance under variable lighting and different weather conditions. Relying solely on visible light images is restricted by the environment such as shadow occlusion and texture similarity, resulting in a reduced recognition rate. In addition, traditional feature fusion methods, such as direct splicing, are prone to information redundancy in the feature space, increasing the computational burden of subsequent processing.
[0004] Previous scholar Ma fused infrared and visible light images through gradient transfer and total variation minimization methods, improving the visibility and information richness of images under various lighting conditions; DenseFuse (a new deep learning framework for infrared and visible light image fusion) uses a deep learning framework for feature extraction and fusion, enhancing the feature integration ability of infrared and visible light images; although both DenseFuse and RTFormer are advanced technologies in the field of image fusion, they adopt different technical paths. DenseFuse focuses on optimizing the feature extraction process through the DenseBlock structure, while RTFormer uses the self-attention mechanism of Transformer to process image data. Both have their own advantages in the field of image fusion and are suitable for different application scenarios.
[0005] Therefore, there is an urgent need to propose a construction method of a road crack segmentation system that fuses infrared and visible light images to solve the above technical problems. Summary of the Invention
[0006] The object of the present invention is to solve the problem of insufficient recognition of fine cracks under different lighting and weather conditions, and to provide a construction method for a road crack segmentation system that fuses infrared and visible light images. A brief overview of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention.
[0007] The technical solution of the present invention:
[0008] A construction method for a road crack segmentation system that fuses infrared and visible light images, comprising the following steps:
[0009] Step 1: Improve the RTFormer model;
[0010] Step 2: Propose a visible light and infrared fusion module IFF;
[0011] Step 3: DDAPPM: Improve the DAPPM module;
[0012] Step 4: Use knowledge distillation for training.
[0013] Preferably: Step 1 includes the following steps:
[0014] Step 11: Introduction of the infrared branch:
[0015] Add a branch dedicated to processing infrared images to the RTFormer model; infrared images can provide information different from visible light images, such as temperature distribution, and are particularly suitable for identifying cracks under low light or other complex lighting conditions;
[0016] Step 12: Optimization of feature fusion:
[0017] Select the Layer-3 layer in the RTFormer model to fuse the features of infrared images and visible light images; the Layer-3 layer is at a relatively deep position in the network and can process higher-level feature information. Experiments show that fusing at this layer can maximize the complementarity of the two image features and improve the accuracy of crack detection;
[0018] Step 13: Deepening of the network structure: Add additional layers compression3e and Layer-3-1e to the infrared image branch and the visible light branch respectively to enhance the depth and feature extraction ability of the network; these newly added layers enable the network to better capture and fuse complex features from different sources while maintaining a relatively low complexity.
[0019] Preferably: Step 2 includes the following steps:
[0020] Step 21. Feature Enhancement:
[0021] IFF uses the attention mechanism to preprocess and enhance the features from visible and infrared images; the attention mechanism can automatically identify and emphasize the most important information in both images. For example, in crack detection, the attention focuses on the crack edges or areas with obvious contrast to the environment.
[0022] Specifically, preprocess the visible light image to improve its contrast. At the same time, perform denoising and enhancement operations on the infrared image; during the feature fusion process, apply channel attention operations to amplify important features and suppress irrelevant features.
[0023] Step 22. Fusion Strategy:
[0024] The enhanced features are not simply concatenated but fused through an addition operation. The fusion method is element-wise addition, which helps to integrate multi-scale features and improve the detection performance; this method effectively combines local and global information, keeps the number of channels unchanged, reduces the model complexity and computational burden, and at the same time avoids the feature redundancy problem introduced by traditional fusion methods.
[0025] Step 23. Module Integration:
[0026] IFF is integrated into the backbone of the RTFormer network to ensure that the features of visible light and infrared data can be effectively fused throughout the network, thereby improving the model's ability to identify cracks.
[0027] Specifically, the above-mentioned processed feature fusion result is used as part of the IFF module to further remove interference features and enhance the fused features; the output of the IFF module is integrated into the backbone of the RTFormer network to achieve more efficient feature fusion and target detection.
[0028] Preferably, Step 3 includes the following steps:
[0029] Step 31. Improvement of DAPPM Structure:
[0030] Based on the original DAPPM, more dense connections and pooling operations are added to better aggregate features from different network depths; such a design helps the model capture richer context information, especially when dealing with images of large scales and complex backgrounds, and thus DDAPPM is obtained.
[0031] Step 32. Feature Reuse:
[0032] Through an improved dense connection method, each layer is connected not only to the previous layer but also to all previous layers, that is, all layers are interconnected. For example, the third layer that receives the features of the first layer is connected to the first layer, which means retransmitting the already extracted information to the first layer, namely feature reuse. This can maximize the information flow and feature reuse, and this method helps to improve the feature representation ability, especially in differentiating between crack regions and non-crack regions in images. Here, the layer refers to the internal layer of RTFormer;
[0033] Step 33, Multi-scale information fusion:
[0034] The DDAPPM module adopts pooling layers of different scales and can process multi-scale information; this multi-scale fusion strategy enables the model to better understand the overall scene structure while retaining key details; the pooling layer reduces the spatial size of the feature map through the pooling operation, and the two are inseparable in function and structure. The pooling operation is the core calculation process of the pooling layer, and the pooling layer is the specific layer type for implementing the pooling operation.
[0035] Preferably: Step 4 includes the following steps:
[0036] Step 41, Teacher model training:
[0037] First, train a teacher model, which has a large number of parameters and high complexity and can achieve high performance on the training data; the process of training the teacher model can adopt the process of training an ordinary deep learning model. Here, it means the best model improved so far, which is defined as the teacher model;
[0038] Step 42, Student model design and training: Network structure design. Design a student model, which is simplified in structure but learns from the teacher model through the distillation process. During the training process, the student model not only learns the standard training objective but also learns to imitate the output of the teacher model, which is achieved by adding a distillation loss function. The student model is smaller than the teacher model; in knowledge distillation, the KL divergence loss is usually used to measure the difference between the probability distribution of the output of the student model and the probability distribution of the output of the teacher model. The KL divergence loss is used for the soft labels of the output of the teacher model, and the cross-entropy loss is used for the true labels.
[0039] The present invention has the following beneficial effects:
[0040] The infrared image branch introduced by the present invention effectively overcomes the environmental limitations faced by relying solely on visible light images, such as visual recognition problems of shadow occlusion and texture similarity, and can still maintain a high recognition rate under insufficient light or reflective conditions;
[0041] The addition of the infrared imaging technology in the present invention can utilize spectral information of different wavelengths, enhance the visualization of cracks, and thus improve the accuracy and reliability of detection;
[0042] The present invention applies the knowledge distillation technology in the model design, effectively reducing the number of model parameters and computational requirements, making the model more suitable for running on resource-constrained devices; this is particularly important for on-site applications, especially in mobile monitoring devices or remote monitoring systems, enabling real-time and efficient crack detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a structural diagram of a road crack segmentation system that fuses infrared and visible light images.
[0044] Figure 2 is a diagram of the RTFormer-F4 network architecture.
[0045] Figure 3a is a schematic diagram of the structure of the interactive fusion module.
[0046] Figure 3b is a schematic diagram of the structure of the feature interactive fusion module.
[0047] Figure 4a is a schematic diagram of the DAPPM structure.
[0048] Figure 4b is a schematic diagram of the DDAPPM structure.
[0049] Figure 5 is a diagram of the distillation learning process. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described below through specific embodiments shown in the drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0051] DETAILED DESCRIPTION OF THE EMBODIMENT 1: In combination with Figures 1-5 To illustrate this embodiment, a method for constructing a road crack segmentation system that fuses infrared and visible light images in this embodiment is characterized by including the following steps:
[0052] Step 1: RTFormer14 is a deep learning model for image segmentation tasks; RTFormer is based on the Transformer architecture, a model for processing sequential data, especially in the field of natural language processing; in image segmentation, Transformer can effectively capture long-range dependencies in images, thereby improving the accuracy of segmentation; however, this invention improves the RTFormer model inspired by DenseFuse and applies it to road crack segmentation that fuses infrared and visible light images based on the RTFormer-slim model.
[0053] Step 1 duplicates the original single visible light input channel into an identical channel to process infrared images, including the following steps:
[0054] Step 11: Introduction of the infrared branch:
[0055] Add a branch in the RTFormer model specifically for processing infrared images; infrared images can provide information different from visible light images, such as temperature distribution, and are especially suitable for identifying cracks under low-light or other complex lighting conditions.
[0056] Step 12: Optimization of feature fusion:
[0057] Select the Layer-3 layer in the RTFormer model for fusing the features of infrared images and visible light images; the Layer-3 layer is in a relatively deep position in the network and can process higher-level feature information. Experiments show that fusing at this layer can maximize the complementarity of the two types of image features and improve the accuracy of crack detection.
[0058] Step 13: Deepening the network structure: Add additional layers (compression3e and Layer-3-1e) to the infrared image branch and the visible light branch respectively. Compression3e is a component in the RTFormer model used to quickly extract local information of images in the first three Stages. Layer-3-1e is another important component in the RTFormer model used to efficiently obtain the global context information required for semantic segmentation tasks in the latter two stages. Layer-3-1e is the same module in the network structure, and this name is used to distinguish it from the original module of the same name. These additional layers enable the network to better capture and fuse complex features from different sources while maintaining a relatively low complexity.
[0059] By fusing infrared images, it can effectively overcome the environmental limitations such as shadow occlusion and texture similarity that occur when solely relying on visible light images. Infrared images can provide temperature distribution information and maintain high performance even under low-light or other complex lighting conditions, resulting in an improved recognition rate.
[0060] Step 2: Propose the visible light and infrared fusion module IFF (Interaction Feature Fusion, internal fusion module);
[0061] To effectively extract features, a feature fusion module Interaction Feature Fusion (IFF) based on the attention mechanism is proposed; this module is used to fuse the features of the visible light and infrared branches in the backbone network; the structure of IFF is similar to that of the Interaction Fusion Module (IFM), the difference being that IFM enhances the visible light and infrared features respectively using the attention features and then concatenates them; while IFF enhances the fused features of visible light and infrared using the attention features, and uses the Add operation to replace the Concat operation, keeping the number of channels unchanged, which is beneficial to reducing the subsequent computational complexity; IFF uses the following strategies, which specifically include the following steps:
[0062] Step 21: Feature enhancement:
[0063] IFF preprocesses and enhances the features from visible light and infrared images using the attention mechanism; the attention mechanism can automatically identify and emphasize the most important information in the two images. For example, in crack detection, the attention is concentrated on the crack edges or areas with obvious contrast to the environment;
[0064] Specifically, preprocess the visible light image to improve its contrast. At the same time, perform denoising and enhancement operations on the infrared image; during the feature fusion process, apply channel attention operations (such as CBAM) to amplify important features and suppress irrelevant features;
[0065] Step 22: Fusion strategy:
[0066] The enhanced features are fused not by simple concatenation but by an addition operation (Add). The fusion method is element-wise addition (Add), which helps to integrate multi-scale features and improve the detection performance; this method effectively combines local and global information, keeps the number of channels unchanged, reduces the model complexity and computational burden, and at the same time avoids the feature redundancy problem introduced by traditional fusion methods; using the addition operation to replace the concatenation method can achieve effective integration of features without increasing the computational complexity, improve the recognition ability and segmentation accuracy of the model;
[0067] Step 23: Module integration:
[0068] IFF is integrated into the backbone of the RTFormer network to ensure that the features of visible light and infrared data can be effectively fused throughout the network, thereby improving the model's ability to identify cracks;
[0069] Specifically, the above-mentioned processed feature fusion result is used as part of the IFF module to further remove interference features and enhance the fused features; the output of the IFF module is integrated into the backbone (main network) of the RTFormer network to achieve more efficient feature fusion and object detection;
[0070] Step 3: DDAPPM: Improved DAPPM module;
[0071] The DDAPPM module is used for feature fusion of the low-resolution infrared image branch, extracts context information, inputs feature maps of different resolutions into the next convolutional layer, and finally connects the outputs of all layers together; in order to strengthen feature transmission and increase the feature reuse rate, the DDAPPM module uses a densely connected form. Step 3 specifically includes the following steps:
[0072] Step 31. Improvement of the DAPPM structure:
[0073] Based on the original DAPPM ( Figure 4a , Deep Fusion Pyramid Pooling Module), more dense connections and pooling operations are added to better aggregate features from different network depths (connections between different sub-blocks in the module represent information integration at different depths); such a design helps the model capture richer context information, especially when processing images with large scales and complex backgrounds, and thus the DDAPPM ( Figure 4b , DDAPPM is an improvement of DAPPM, and the connections of the DDAPPM module are denser than those of DAPPAM, increasing the mutual complementarity of information in the network circulation) is obtained;
[0074] Step 32. Feature reuse:
[0075] Through the improved dense connection method, each layer is connected not only to the previous layer but also to all previous layers, that is, all layers are interconnected. For example, the third layer that receives the features of the first layer is connected to the first layer, which means retransmitting the already extracted information to the first layer, that is, feature reuse. This can maximize the information flow and feature reuse, and this method helps to improve the feature representation ability, especially in distinguishing between crack regions and non-crack regions in the image. Here, the layer refers to the internal layer of RTFormer;
[0076] Step 33. Multi-scale information fusion:
[0077] The DDAPPM module adopts pooling layers of different scales and can process multi-scale information. This multi-scale fusion strategy enables the model to better understand the overall scene structure while retaining key details. The pooling layer reduces the spatial size of the feature map through pooling operations. The two are inseparable in terms of function and structure. The pooling operation is the core calculation process of the pooling layer, and the pooling layer is the specific layer type that implements the pooling operation.
[0078] Step 4: Use knowledge distillation for training.
[0079] Knowledge distillation ( Figure 5 ) is a model training technique that improves the performance of a small network by transferring knowledge from a large, complex (teacher model) to a small, simplified (student model) network. This method is particularly suitable for applications that need to be deployed in resource-constrained environments, such as mobile devices or edge computing devices. In the image segmentation task of first fusing infrared and visible light and then detecting road cracks, using knowledge distillation can reduce the model size and computational requirements while maintaining high accuracy. Liu's research proposed a structured knowledge distillation method for semantic segmentation. Combining with other steps of the present invention, through the knowledge transfer from the teacher model to the student model, the performance of the student model is effectively improved. Step 4 specifically includes the following steps:
[0080] Step 41: Teacher model training:
[0081] First, train a teacher model (a large RTFormer model. The RTFormer-slim model is the model in the paper "Rtformer: Efficient design for real-time semantic segmentation with transformer", translated as "Efficient design for real-time semantic segmentation with transformer"). This model has a large number of parameters and high complexity and can achieve high performance on the training data. The process of training the teacher model can adopt the process of training an ordinary deep learning model. Here, it means the best model improved so far, which is defined as the teacher model.
[0082] Step 42. Design and training of the student model: Design the network structure, and design a student model (a smaller and more efficient RTFormer model). This model is structurally simplified but learns from the teacher model through the distillation process. During the training process, the student model not only learns the standard training objectives (such as cross-entropy loss), but also learns to imitate the output of the teacher model, which is achieved by adding a distillation loss function (KLLoss). The student model is smaller than the teacher model; in knowledge distillation, the KL divergence loss (KLDivLoss) is usually used to measure the difference between the probability distribution of the output of the student model and the probability distribution of the output of the teacher model. The KL divergence loss is used for the soft labels of the output of the teacher model, and the cross-entropy loss is used for the true labels;
[0083] The present invention has been verified through extensive experiments. The improved RTFormerF10 model performs better than the basic model on the self-built dataset; especially in terms of the mIoU evaluation index, this model has increased by 1% compared to the basic model and reached 0.8541; this improvement in performance is attributed to the model's ability to more accurately identify and segment the crack area, and it can maintain a high recognition rate even under insufficient light or reflective conditions; the infrared image branch introduced in the present invention effectively overcomes the environmental limitations faced by relying solely on visible light images, such as shadow occlusion and visual recognition problems with similar textures; the addition of infrared imaging technology enables the system to utilize spectral information of different wavelengths, enhances the visualization of cracks, and thus improves the accuracy and reliability of detection; the application of knowledge distillation technology in the model design effectively reduces the number of model parameters and computational requirements, making the model more suitable for running on resource-constrained devices; this is particularly important for on-site applications, especially in mobile monitoring devices or remote monitoring systems, where real-time and efficient crack detection can be achieved.
[0084] It should be noted that in the above embodiments, as long as the technical solutions are not contradictory, they can be arranged and combined. Those skilled in the art can exhaust all possibilities based on the mathematical knowledge of permutations and combinations. Therefore, the present invention will no longer describe the technical solutions after permutation and combination one by one, but it should be understood that the technical solutions after permutation and combination have been disclosed by the present invention.
[0085] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Construction method of road crack segmentation system integrating infrared and visible light images, characterized in that: It includes the following steps: Step 1: Improve the RTFormer backbone network; Step 1 includes the following steps: Step 11: Introduction of the infrared branch: Add an infrared image branch to the RTFormer model; Step 12: Optimization of feature fusion: Select to perform the fusion of infrared image features and visible light image features at the Layer-3 layer of the model; Step 13: Deepening of the network structure: Add additional layers, compression3e and Layer-3-1e, respectively, after the Layer-3-1 layer in the infrared image branch and the visible light branch; Step 2: Propose the visible light image and infrared fusion module IFF; Step 2 includes the following steps: Step 21: Feature enhancement: First, perform an addition operation on the feature maps of the infrared image and the visible light image, and then perform a multiplication operation on the result of the addition with the attention weights; Step 22: Fusion strategy: Perform an addition operation fusion on the result of the multiplication operation and the result of the addition of the feature maps of the infrared image and the visible light image; Step 23: Module integration: The first IFF performs the fusion of the features of the infrared image and the visible light image at the Layer-3 layer of the model. The output of the first IFF is used as the input of the Layer-4 layer. The second IFF performs the fusion of the features of the infrared image and the visible light image at the Layer-3-1e layer. The output of the second IFF is used as another input of the Layer-4 layer; Step 3: Improve the DAPPM module; Step 3 includes the following steps: Step 31: DAPPM improvement: Based on DAPPM, add dense connections and pooling operations to obtain DDAPPM; The positions to add connections are: Add connections between Conv and the second, third, and fourth Conv upsamples; Add connections between the first Conv upsample and the third and fourth Conv upsamples; Add connections between the second Conv upsample and the fourth Conv upsample; Step 32: Feature reuse: All layers are interconnected; Through the improved dense connection method, each layer is connected not only to the previous layer but also to all previous layers, that is, all layers are interconnected; Step 33: Multi-scale information fusion: DDAPPM adopts pooling layers of different scales; Step 4: Use knowledge distillation to train the improved RTFormer.
2. The method for constructing a road crack segmentation system that fuses infrared and visible light images according to claim 1, wherein: Step 4 includes the following steps: Step 41: Teacher model training: Train a teacher model, and the process of training the teacher model adopts the process of training a deep learning model; Step 42: Student model design and training: Design a student model, which learns from the teacher model through the distillation process. During the training process, the student model learns the standard training objective, which is the cross-entropy loss, and learns to imitate the output of the teacher model, which is achieved by adding a distillation loss function.