A method for detecting road cracks by thermal-optical matrix fusion

The dual-modal fusion of visible light and temperature matrices with an improved YOLOv8 network and AMPN enhances road crack detection precision and efficiency, addressing the limitations of single-modal methods under complex environmental conditions.

CN119919416BActive Publication Date: 2025-07-15SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510409571.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-15
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing road crack detection methods are insufficient in complex environments, especially under conditions such as insufficient light or surface reflection, which can easily lead to false inspection and missed inspection, which increases the cost of manual review and the burden of later repair.

Method used

Using a detection method based on thermal-light matrix fusion, combining the dual-modal information of visible light images and temperature matrix, the feature weight adaptive weighted fusion and regional block-level dynamic fusion are performed through the improved YOLOv8-AMPN network model to achieve multi-scale feature enhancement and dynamic resolution adjustment, and improve detection accuracy and efficiency.

Benefits of technology

It significantly improves the accuracy and efficiency of crack detection, reduces manual participation, reduces maintenance costs, and extends the service life of the road.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919416B_ABST
    Figure CN119919416B_ABST
Patent Text Reader

Abstract

The present invention designs a road crack detection method based on thermo-optical matrix fusion. The method includes: obtaining a temperature matrix and a visible light image through a visible light-temperature matrix acquisition device, extracting crack edge features through temperature gradient analysis, and performing binarization processing to obtain a binarized temperature matrix map; proposing a thermo-optical matrix image fusion module to effectively fuse the temperature matrix and the visible light image; using the thermo-optical matrix fusion crack image obtained based on the fusion module as the input of an improved YOLOv8-AMPN network model, and further enhancing the semantic enhancement effect of the temperature matrix on the visible light image and highlighting crack features through the feature weight adaptive weighted fusion and regional block-level dynamic fusion methods in the network. The advantages of the present invention are as follows: introducing a temperature matrix branch effectively overcomes the limitations of visible light images under environmental changes, solves the problem of redundant and complex multi-modal image information, and improves the sensitivity and detection response speed of the model to crack features through temperature information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a road crack detection technology, especially an automated detection method based on thermal-optical matrix fusion. The main innovation lies in the dual-modal fusion of visible light images and temperature matrices, which solves the problem of insufficient robustness of single-modal detection in complex road environments. The temperature matrix has stability under changing lighting conditions and can effectively distinguish the temperature difference information between cracks and normal road surfaces, while visible light images have higher resolution for crack details. By introducing an improved YOLOv8 network architecture, combined with the Adaptive Multimodal Pyramid Network (AMPN) and the feature weight adaptive weighted fusion and regional block-level dynamic fusion methods in the network, this system realizes multi-scale feature enhancement and dynamic resolution adjustment, further enhancing the semantic enhancement of the temperature matrix on visible light images, highlighting crack features, adapting to various detection scenarios, significantly improving the accuracy and efficiency of crack detection, reducing manual participation, and lowering maintenance costs. Background Art

[0002] With the rapid development of road traffic, the aging of infrastructure and the increase in vehicle loads, road surface crack detection and maintenance have become crucial. Cracks are manifestations of early road damage. If not repaired in time, they will lead to further structural damage and safety hazards, significantly increasing repair costs and accident risks. Traditional manual inspection methods are inefficient, subjective, and prone to missed and false detections, resulting in small cracks deteriorating into large-scale damages and increasing the resource consumption of later maintenance.

[0003] Existing crack detection methods mainly detect based on single-modal images (such as visible light images) or the combination of visible light and infrared images. Although they can detect cracks to a certain extent, due to the influence of complex road environments and lighting changes, there are limitations in the robustness and accuracy of crack detection. Especially in complex environments such as insufficient lighting or surface reflection, it is easy to cause a decrease in detection accuracy. When dealing with complex scenarios, the current detection methods often increase the cost of manual recheck and the burden of later repair due to false and missed detections. Summary of the Invention

[0004] To overcome these problems, the present invention proposes a road crack detection method based on thermal-optical matrix fusion. By combining the dual-modal information of visible light images and temperature matrices, and utilizing the stability of the temperature matrix under changing lighting conditions and the high-detail resolution advantage of visible light images, accurate detection of cracks is achieved.

[0005] Specifically, the present invention provides a road crack detection method based on thermal-optical matrix fusion, which includes:

[0006] Based on the visible light-temperature matrix acquisition device, the temperature matrix and visible light image are obtained. The crack edge features are extracted through temperature gradient analysis and binarized to obtain the binarized temperature matrix image;

[0007] The temperature matrix and the visible light image are fused through the thermal-optical matrix image fusion module;

[0008] The thermal-optical matrix fusion crack image obtained based on the thermal-optical matrix image fusion module is used as the input of the improved YOLOv8-AMPN network model. Through the feature weight adaptive weighted fusion and regional block-level dynamic fusion in the network model, an enhanced thermal-optical matrix fusion crack image is obtained;

[0009] Crack identification is performed based on the enhanced thermal-optical matrix fusion crack image.

[0010] Further, the visible light-temperature matrix acquisition device includes: a visible light camera, a coaxial alignment adjusting rod, a hand push handle, a synchronous trigger, an infrared thermal imaging camera, a level, a laser indicator, and a movable triangular stable bracket.

[0011] Further, through the coaxial alignment adjusting rod and the level, the optical axes of the visible light camera and the infrared thermal imaging camera are adjusted to be parallel; through the laser indicator, the visible light camera and the infrared thermal imaging camera are adjusted to maintain the same field of view at different focal lengths.

[0012] Further, the obtaining of the temperature matrix and the visible light image based on the visible light-temperature matrix acquisition device, the extraction of the crack edge features through temperature gradient analysis, and the binarization processing to obtain the binarized temperature matrix image include:

[0013] The visible light image is obtained, adjusted to a preset resolution size, and the picture size is unified;

[0014] Based on the visible light-temperature matrix acquisition device, the temperature matrix data is obtained. The temperature matrix is subjected to grayscale image generation and temperature matrix gradient calculation to segment the crack edge, and further a binarized temperature matrix image is generated and corresponding to the visible light picture;

[0015] Data enhancement is performed on the binarized temperature matrix image and the visible light image.

[0016] Further, the fusion of the temperature matrix and the visible light image through the thermal-optical matrix image fusion module includes:

[0017] (1) Feature preliminary alignment

[0018] First, the temperature matrix features and the visible light image features are respectively input into the thermal-optical matrix image fusion module, and a shared convolutional alignment layer is used for feature exchange and feature alignment;

[0019] (2) Feature complementarity enhancement

[0020] Calculate the similarity score through the similarity calculation layer, and adjust the enhancement between the thermal matrix features and the visible light image features through a learnable parameter:

[0021] (3) Multi-layer refinement fusion

[0022] Fuse the enhanced features of the visible light image and the thermal matrix layer by layer, and adopt channel-level and spatial-level attention mechanisms;

[0023] (4) Thermal-optical enhancement module

[0024] Through a preset convolutional layer, perform feature refinement extraction on the fused features;

[0025] Use the complementary weight layer to calculate the contribution of each modality to the final feature, and obtain the fused feature according to the contribution;

[0026] Input the fused feature into the saliency prediction layer to obtain the crack saliency map.

[0027] Furthermore, taking the thermal-optical matrix fusion crack image obtained based on the thermal-optical matrix image fusion module as the input of the improved YOLOv8-AMPN network model, and obtaining the enhanced thermal-optical matrix fusion crack image through the feature weight adaptive weighted fusion and region block-level dynamic fusion in the network model, including:

[0028] S41: Multi-scale convolutional kernel fusion;

[0029] S42: Dynamic resolution adjustment;

[0030] S43: Global feature enhancement.

[0031] Furthermore, the multi-scale convolutional kernel fusion includes:

[0032] Introduce convolutional kernels of different sizes in the convolutional module of the YOLOv8-AMPN network model;

[0033] In the YOLOv8-AMPN network model, splice the feature maps extracted by different convolutional kernels in different modalities to achieve multi-scale fusion of visible light features and thermal matrix features;

[0034] In the spliced multi-scale features, according to the differences between the visible light features and the thermal matrix features, adjust the feature contributions of the two modalities through feature weight adaptive weighted fusion to achieve multi-modal feature expression;

[0035] Among them, the feature weight adaptive weighted fusion includes:

[0036] Feature partitioning, splitting the visible light image and the temperature matrix features into regional blocks;

[0037] By calculating the average brightness and temperature difference of the regional blocks, an adaptive weight matrix is generated:

[0038] For each regional block, the visible light image and the temperature matrix feature blocks are weighted and fused through the adaptive weight matrix;

[0039] Through the fused feature blocks of all regional blocks, the overall fused feature map is obtained.

[0040] Furthermore, the dynamic resolution adjustment includes:

[0041] For each regional block, an edge detection is performed using a directional filter to generate a direction matrix for determining the main direction of the cracks within the regional block;

[0042] According to the direction matrix, an adaptive convolution kernel is designed for each regional block and applied to the visible light image and the temperature matrix feature extraction respectively; wherein, in the design of the adaptive convolution kernel, the shape and size of the convolution kernel are dynamically adjusted according to the direction and width features of the cracks;

[0043] Within each regional block, an adaptive fusion weight is calculated according to the temperature difference and the edge sharpness;

[0044] Based on the adaptive fusion weight, the visible light and the temperature matrix features within each regional block are fused to obtain a fused feature block:

[0045] According to all regional blocks, the overall feature map of dynamic fusion is obtained.

[0046] Furthermore, the global feature enhancement includes:

[0047] The overall feature map of the adaptive weighted fusion of the regional blocks and the overall feature map of the dynamic fusion at the regional block level are stitched together to form a multi-scale feature map;

[0048] A global pooling operation is performed on the multi-scale feature map to generate a globally enhanced multi-scale feature map.

[0049] The advantages of the present invention are as follows: The temperature matrix can effectively distinguish the temperature difference information between the crack area and the normal road surface, compensating for the deficiencies of the visible light image under complex lighting conditions; while the visible light image has a higher resolution for the fine structure of the cracks, and the combination of the two can effectively improve the robustness and accuracy of the detection.

[0050] Furthermore, by introducing an improved YOLOv8 network architecture and combining it with the Adaptive Multimodal Pyramid Network (AMPN) and the multimodal feature fusion module, the present invention can adaptively fuse the visible light and temperature matrix features, achieve multi-scale feature enhancement and dynamic resolution adjustment, and meet the crack detection requirements under different environments and conditions. This method significantly improves the efficiency and accuracy of crack detection, reduces manual participation, lowers maintenance costs, and extends the service life of roads. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0052] Figure 1 is a flowchart of the method for detecting road cracks by thermal-optical matrix fusion of the present invention.

[0053] Figure 2 is a schematic diagram of the visible light-temperature matrix acquisition device of the present invention.

[0054] Figure 3 is a schematic diagram of the related preprocessed image of the present invention.

[0055] Figure 4 is a binary feature fusion map of the visible light-temperature matrix of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0057] The present invention designs a method for detecting road cracks by thermal-optical matrix fusion to improve the accuracy and efficiency of road crack detection. As Figure 1 shown, it is a flowchart of the method for detecting road cracks by thermal-optical matrix fusion of the present invention. The specific description is as follows:

[0058] S1: Setup of the data acquisition device

[0059] The visible light (VIS) image and infrared (IR) thermal imaging temperature matrix of road cracks are obtained by a dual-modal imaging device. The relevant information collected by each modality is stored in a database for subsequent multi-modal feature processing.

[0060] In the data collection section, an innovative device setup method and shooting process are adopted to ensure accurate data of visible light (VIS) and infrared (IR) thermal imaging temperature matrix are obtained, so as to achieve high-quality multi-modal information acquisition for the same crack.

[0061] S11: Install a visible light camera (i.e., visible light camera) and an infrared thermal imaging device (i.e., infrared camera) on the same device to form a dual-modal device for synchronous imaging. These cameras have adjustable focal lengths and angles to ensure that the crack images in the same field of view are captured simultaneously.

[0062] S12: To ensure the stability of shooting, the device includes designing a triangular stable bracket with shock absorption function that carries two shooting devices at the same time, ensuring that the device is not affected by vibration or external interference during shooting. The bracket has the function of adjusting height and angle to adapt to different road surfaces and crack positions. As Figure 2 shown, the visible light-temperature matrix acquisition device of the present invention includes a visible light camera 1, a coaxial alignment adjusting rod 2, a hand-pushing handle 3, a synchronous trigger 4, an infrared thermal imaging camera 5, a level 6, a laser indicator 7, and a movable triangular stable bracket 8. The two cameras can also be called cameras in the following description.

[0063] S13: To ensure that the visible light and infrared thermal imaging cameras can capture the overlapping area of the same crack, the device adopts a coaxial alignment system to adjust the optical axes of the two cameras to be parallel. Through a mechanical adjustment device, the two devices are made to have the same field of view at different focal lengths.

[0064] To ensure the field of view consistency of the visible light and infrared thermal imaging cameras at different focal lengths, the following calculation formula can be introduced for calibration based on the relationship between the field of view (Field of View, FOV) and the focal length (Focal Length, f). The calculation formula for field of view consistency is as follows:

[0065]

[0066] Calculation formula for field of view width:

[0067] where, is the field of view angle of the camera in the horizontal direction, ω is the width of the sensor, and f is the focal length. This formula can be used to determine whether the field of view widths of the visible light and infrared thermal imaging cameras are the same.

[0068]

[0069] Field of view height calculation formula:

[0070] Wherein, is the field of view angle of the camera in the vertical direction, and h is the height of the sensor.

[0071] Field of view overlap calibration formula: To ensure the field of view overlap area of visible light and infrared thermal imaging images at different focal lengths, the difference in field of view width and height can be used to calculate the overlap ratio (Overlap Ratio, OR):

[0072]

[0073] Wherein, and are the horizontal field of view angles of the visible light and infrared thermal imaging cameras respectively, and are the vertical field of view angles of the visible light and infrared thermal imaging cameras respectively. When OR is close to 1, the field of view overlap is optimal.

[0074] In the present invention, the visual field consistency mechanical adjustment device includes:

[0075] (1) Laser indicator

[0076] A laser indicator is provided at the top of the triangular bracket device, pointing to the crack target area, for accurately calibrating the shooting position to ensure that the device is correctly aligned with the crack. The role of the laser indicator is to enable the visible light and infrared thermal imaging cameras to be accurately aligned before shooting, reduce subsequent alignment errors, and improve the spatial consistency and detection accuracy of multi-modal data.

[0077] (2) Mechanical structure adjustment part

[0078] Adjustment device: A fine adjustment knob and a slide rail are added to the camera bracket to enable the camera to accurately adjust its position and angle in the horizontal and vertical directions. The fine adjustment knob can finely adjust the optical axis of the camera to make the fields of view of the two cameras overlap on the calibration mirror mark. After the preliminary adjustment, the collimation mark on the calibration mirror can be observed through the human eye or an auxiliary camera to confirm whether the fields of view of the two cameras completely overlap. This step can be achieved by comparing the position of the mark in the camera's field of view with the naked-eye observation through a semi-transparent mirror.

[0079] (3) Formula consistency check

[0080] Field of view consistency formula check: Through calibration, the field of view consistency visible to the human eye is achieved, and the field of view consistency maintenance formula can be used for accuracy verification. For example:

[0081] Calculation of width field of view overlap degree:

[0082]

[0083] Calculation of the high - altitude field - of - view overlap degree:

[0084] The consistency of the two fields of view is determined by the calculated overlap degree approaching 1.

[0085] (5)Long - term consistency test

[0086] Regular image inspection: During the use of the device, by observing the collimation - mark images taken, regularly check for any deviation to ensure long - term field - of - view consistency. If the collimation mark shows deviation, the fine - adjustment screw of the bracket can be readjusted to restore the field - of - view consistency.

[0087] Verification of the mark alignment formula: After long - term use, if it is necessary to confirm the accuracy, the field - of - view overlap degree can be recalculated through the formula to ensure the calibration accuracy at each focal length and in different environments.

[0088] S14: To capture visible - light and infrared thermal - imaging images at the same time, the device is integrated with a synchronous trigger. Each time a shot is taken, the trigger starts the two cameras simultaneously, ensuring the time synchronization of the two acquired images, thus guaranteeing the consistency of crack information at the same moment.

[0089] S15: Before crack shooting, install the laser indicator at the center of the base of the dual - mode device to calibrate the center position of the crack and ensure that the device is directly facing the crack area.

[0090] S16: During the acquisition process, using the real - time monitoring function built into the shooting device, it is possible to check at any time whether the quality of the acquired images meets the detection requirements. If there are situations such as blurring, deviation, or insufficient exposure, the operator can make adjustments.

[0091] S2: Pre - processing of visible - light images and temperature - matrix data

[0092] S21: Adjust the visible - light images to a suitable resolution size and unify the picture sizes.

[0093] S22: Based on the temperature - matrix data obtained by the visible - light - temperature - matrix acquisition device, generate a grayscale image of the temperature matrix, calculate the gradient of the temperature matrix to segment the crack edge, and further generate a binary image of the temperature matrix, corresponding to the visible - light picture. As Figure 3 shown, it includes a grayscale image of the temperature matrix, a segmentation map of the temperature - matrix gradient calculation, a binary image of the temperature matrix, and a visible - light image.

[0094] S23: Data enhancement

[0095] Data augmentation performs data geometric augmentation through methods such as random cropping, flipping, rotation, and scaling in geometric transformation. It enhances the diversity of data to enable the model to adapt to cracks from different perspectives.

[0096] After the above data preprocessing and augmentation, it is ensured that the data distributions of the two modalities are consistent and corresponding. The final output is paired visible light images and binary images of temperature matrices, which are input into the thermal-optical matrix fusion module.

[0097] S3: Thermal-optical matrix fusion module

[0098] This module mainly focuses on the fusion of thermal imaging matrices and visible light images. It aims to enable the thermal matrix and the optical matrix to each play their advantages in the fusion through feature alignment, feature compensation, and depth fusion. The design of this module focuses on enhancing the feature representation in the crack area and reflecting the complementarity of temperature and optical information in the final output.

[0099] S31: Initial feature alignment

[0100] First, the temperature matrix feature F _t and the visible light image feature F_ v are respectively input into this module. To ensure that the two modality features maintain the same scale and feature resolution before fusion, a shared convolutional alignment layer (SCAL) is used for feature exchange.

[0101]

[0102] Among them, SCAL is a multi-level convolutional layer containing 3×3 and 1×1 convolutions, which is used to effectively refine and align the input visible light image and temperature matrix features.

[0103] S32: Feature complementary enhancement

[0104] To understand the feature similarity between the visible light image and the temperature matrix, a similarity computation layer (SCL) is designed:

[0105]

[0106] where ϵ is a value to prevent the denominator from being zero.

[0107] According to the similarity score, the enhancement between different modality features is adjusted through a learnable parameter :

[0108]

[0109] S33: Multi - layer Refinement and Fusion

[0110] The enhanced features of the visible - light image and the temperature matrix are fused layer by layer, and channel - level and spatial - level attention mechanisms are adopted.

[0111]

[0112] Among them, represents the channel - level feature splicing and fusion.

[0113] S34: Thermal - Optical Enhancement Module

[0114] By combining the uniqueness of the visible - light image and the temperature matrix features, a dedicated convolutional layer is designed to first perform a feature refinement extraction on the fused features:

[0115]

[0116] A complementary weight layer is used to automatically calculate the contribution of each modality to the final feature:

[0117]

[0118] The final fused feature is:

[0119]

[0120] S35: Final Fused Feature The crack saliency map is obtained through the saliency prediction layer:

[0121]

[0122] where σ represents the Sigmoid activation function, which is used to map the output to [0, 1] to represent the crack saliency probability.

[0123] During the thermal - optical matrix fusion process, the temperature matrix can effectively compensate for the deficiencies of the visible - light image under ambient light changes. The temperature matrix provides temperature difference information about the crack area, and this information remains valid under changing lighting conditions, shadows, or blurry conditions, thus enhancing the robustness of the model to crack features. This compensation characteristic enables the temperature matrix to play a key role in complex road environments, ensuring more stable and accurate crack detection.

[0124] Compared with the visible - light image, the temperature matrix is more sensitive to cracks under temperature difference changes. Therefore, during the feature alignment and fusion process, through the Shared Convolutional Alignment Layer (SCAL) and the Feature Complementary Enhancement Module, the unique thermal contrast features in the temperature features are effectively extracted, thereby enhancing the saliency of the crack area. As Figure 4As shown, it is the visible light - temperature matrix binary feature fusion diagram of the present invention.

[0125] S4: Dual - modal multi - scale feature enhanced YOLOv8 - Adaptive Multimodal Pyramid Network (YOLOv8 - AMPN)

[0126] According to the characteristics of the actual road scene, the YOLOv8 network structure is optimized by introducing Adaptive Multimodal Pyramid Network (AMPN) to enhance the multi - scale crack feature extraction ability, enabling the model to effectively process large - scale cracks while maintaining sensitivity to small cracks. AMPN adaptively fuses visible light image and temperature matrix features at the multi - modal and multi - scale levels, thus improving the detection effect of the model on complex crack features.

[0127] S41: Multi - scale convolution kernel fusion

[0128] 1. To adapt to multi - scale features, we introduce convolution kernels of different sizes in the convolution module, such as 1x1, 3x3, and 5x5, to capture small and large - scale crack features. 1x1 convolution kernel: used for integrating information between channels, compressing high - dimensional features to lower dimensions to avoid loss of small - feature information; 3x3 convolution kernel: used for general feature extraction, being sensitive to medium - and small - scale features of cracks (such as edge textures); 5x5 convolution kernel: suitable for extracting large - scale crack features, capable of capturing features in a wider crack area.

[0129] 2. Feature concatenation: In AMPN, the feature maps extracted by different convolution kernels are concatenated in different modalities to achieve multi - scale fusion of visible light features and temperature matrix features. The formula is as follows:

[0130]

[0131] In the concatenated multi - scale features, AMPN adjusts the feature contributions of the two modalities through adaptive weight according to the differences between visible light features and temperature matrix features to achieve a richer multi - modal feature expression.

[0132] Among them, an innovative feature weight adaptive weighted fusion method is proposed, and the specific steps are as follows:

[0133] Feature partitioning: Assume that and are partitioned into regional blocks of size , and the size of each feature map is , then the following set of regional blocks can be obtained:

[0134]

[0135] Among them, and represent the regional blocks at position . H is the height of the feature map, and W is the width of the feature map, representing the number of pixels in the vertical and horizontal directions respectively.

[0136] Weight generation: By calculating the average brightness and temperature difference of the regional blocks, an adaptive weight matrix is generated:

[0137]

[0138] Among them, and represent the eigenvalues of the visible light image and the temperature matrix at position respectively.

[0139] Adaptive weighted fusion of regional blocks: For each regional block, we use the corresponding weights and to perform weighted fusion on the feature blocks of the visible light image and the temperature matrix:

[0140]

[0141] Finally, the overall fused feature map can be obtained by aggregating all regional blocks :

[0142]

[0143] S42: Dynamic resolution adjustment strategy

[0144] To adapt to the feature requirements of different scales in crack detection, AMPN designs a dynamic resolution adjustment strategy to flexibly process the resolution of the feature maps at different levels of the feature pyramid, so as to enhance the model's detection ability for large-scale and fine cracks.

[0145] 1. Downsampling of large-scale feature maps: For the extraction of large-scale crack features, a 2x2 downsampling convolution (stride 2) is used to reduce the size of the feature map to increase the receptive field.

[0146] 2. Upsampling of small-scale feature maps: For fine crack features, an upsampling layer is used to restore the resolution and retain the texture information of the fine cracks.

[0147] The refinement steps in combination with the regional block-level dynamic fusion method are as follows:

[0148] Directional detection calculation: For each regional block, a directional filter is used for edge detection to generate a direction matrix to determine the main direction of the cracks within the regional block.

[0149]

[0150] Among them, DF represents Direction Filter, which is a direction filter. The direction matrix provides directional information for determining the shape of the strip convolution kernel suitable for the crack morphology, so that the filter can capture the continuous features along the crack direction.

[0151] Adaptive convolution kernel: According to the direction matrix , design an adaptive convolution kernel for each region block, and apply it to the visible light image and temperature matrix feature extraction respectively:

[0152]

[0153] Among them, AC represents Adaptive Conv, which represents the convolution operation that dynamically adjusts the size and direction of the convolution kernel according to the dynamic adjustment of the convolution kernel size and direction.

[0154] In the design of the adaptive convolution kernel, the shape and size of the convolution kernel will be dynamically adjusted according to the direction and width characteristics of the crack. For example:

[0155] For the detected main crack direction being a transverse crack, the selected convolution kernel shape is 1×5 or 1×7 to ensure capturing the continuous crack features along the transverse direction.

[0156] For longitudinal cracks or blocks with a wider crack area, the selected convolution kernel shape is 5×1 or 7×1 to enhance the capture of longitudinal features.

[0157] In the crack area with a smaller width, a smaller convolution kernel (such as 3×3) is selected to enhance the extraction of detail features.

[0158] This selection strategy of the adaptive convolution kernel ensures that the features of cracks with different morphologies in complex scenes can be fully extracted, further enhancing the robustness of the model in crack detection.

[0159] Adaptive fusion weight calculation based on region blocks: Within each region block, calculate the adaptive fusion weight according to the temperature difference and edge sharpness .

[0160]

[0161] Among them, ES represents Edge Sharpness, which is used to calculate the edge sharpness of the visible light image features within the region block; TD represents Temperature Difference, which is used to calculate the temperature difference of the temperature matrix features within the region block.

[0162] Region Block Dynamic Fusion: Based on Adaptive Fusion Weights , fuse the visible light and temperature matrix features within each region block to obtain a fused feature block :

[0163]

[0164] Finally, the overall fused feature map can be obtained by relying on all region blocks :

[0165]

[0166] The dynamic resolution adjustment strategy of AMPN enables each layer of the feature pyramid to flexibly choose upsampling or downsampling according to the actual scale requirements of the cracks. In this way, the model can increase the receptive field when processing large-scale cracks and retain the detail resolution when detecting small cracks.

[0167] S43: Global Feature Enhancement

[0168] Concatenate the region block adaptive weighted fusion with the region block-level dynamic fusion at the global feature map level to form the final fused feature map .

[0169]

[0170] Within each region block, simultaneously retain the multi-modal feature information obtained by adaptive weighted fusion and dynamic fusion, and form the final fused feature map through concatenation.

[0171] Perform global pooling operations on the multi-scale feature maps to generate globally enhanced multi-scale feature maps , enabling the model to effectively aggregate global information in different modalities and scales, thereby improving the crack detection accuracy and robustness to complex scenes. Specifically:

[0172] During the crack detection process, used to supplement the environmental information lacking in local features, thereby reducing false detections or missed detections caused by changes in perspective or road conditions. After concatenating the global feature map with the local region block features, the model's crack recognition ability at each scale is improved, and the model's robustness in complex scenes (such as environments with a lot of background noise) is enhanced. Through global feature enhancement, the model can better understand the contrast between cracks and the surrounding background, ensuring higher accuracy and consistency when detecting crack edges.

[0173] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims described above.

Claims

1. A method for detecting road cracks by thermo-optical matrix fusion, characterized in that, The method includes: Based on a visible light - temperature matrix acquisition device, obtaining a temperature matrix and a visible light image, extracting crack edge features through temperature gradient analysis, and performing binarization processing to obtain a binarized temperature matrix image; Through a thermal - optical matrix image fusion module, fusing the binarized temperature matrix image with the visible light image to obtain a thermal - optical matrix fusion crack image, including: (1) Preliminary feature alignment Inputting the temperature matrix features and the visible light image features into the thermal - optical matrix image fusion module respectively, and using a shared convolutional alignment layer for feature exchange and feature alignment; (2) Feature complementary enhancement Calculating a similarity score through a similarity calculation layer, and adjusting the enhancement between the temperature matrix features and the visible light image features through learnable parameters; (3) Multi - layer refinement fusion Adopting channel - level and spatial - level attention mechanisms to layer - by - layer fuse the enhanced features of the visible light image and the temperature matrix; (4) Thermal - optical enhancement module Through a preset convolutional layer, performing feature refinement extraction on the fused features; Calculating the contribution of each modality to the final feature using a complementary weight layer, and obtaining a fused feature according to the contribution; Inputting the fused feature into a saliency prediction layer to obtain a crack saliency map; Using the thermal - optical matrix fusion crack image as the input of the YOLOv8 - AMPN network model, performing feature weight adaptive weighted fusion and region - block - level dynamic fusion through the network model to obtain an enhanced thermal - optical matrix fusion crack image; Performing crack recognition based on the enhanced thermal - optical matrix fusion crack image.

2. A road crack detection method based on thermal - optical matrix fusion according to claim 1, wherein The visible light - temperature matrix acquisition device includes: a visible light camera, a coaxial alignment adjusting rod, a hand - push handle, a synchronous trigger, an infrared thermal imaging camera, a level, a laser indicator, and a movable triangular stabilizing bracket.

3. A road crack detection method based on thermal - optical matrix fusion according to claim 2, wherein Through the coaxial alignment adjusting rod and the level, adjusting the optical axes of the visible light camera and the infrared thermal imaging camera to be parallel; through the laser indicator, adjusting the visible light camera and the infrared thermal imaging camera to maintain the same field of view at different focal lengths.

4. A road crack detection method based on thermal - optical matrix fusion according to any one of claims 1 - 3, wherein The obtaining of the temperature matrix and the visible light image based on the visible light - temperature matrix acquisition device, extracting crack edge features through temperature gradient analysis, and performing binarization processing to obtain a binarized temperature matrix image includes: Obtaining a visible light image, adjusting the visible light image to a preset resolution size, and unifying the picture size; Based on the visible light - temperature matrix acquisition device, obtaining temperature matrix data, generating a grayscale image from the temperature matrix data and calculating the temperature matrix gradient to segment the crack edge, further generating a binarized temperature matrix image corresponding to the visible light image; Performing data enhancement on the binarized temperature matrix image and the visible light image.

5. A road crack detection method based on thermal - optical matrix fusion according to claim 1, wherein Using the fused thermal-optical matrix crack image as the input of the YOLOv8-AMPN network model, and performing feature weight adaptive weighted fusion and regional block-level dynamic fusion through the network model to obtain an enhanced fused thermal-optical matrix crack image, including: S41: Multi-scale convolutional kernel fusion; S42: Dynamic resolution adjustment; S43: Global feature enhancement.

6. The method for detecting road cracks by thermal-optical matrix fusion according to claim 5, wherein the multi-scale convolutional kernel fusion includes: introducing convolutional kernels of different sizes into the convolutional module of the YOLOv8-AMPN network model; in the YOLOv8-AMPN network model, splicing the feature maps extracted by different convolutional kernels in different modalities to achieve multi-scale fusion of visible light features and temperature matrix features; in the spliced multi-scale features, according to the differences between the visible light features and the temperature matrix features, adjusting the feature contributions of the two modalities through feature weight adaptive weighted fusion to achieve multi-modal feature expression; wherein, the feature weight adaptive weighted fusion includes: feature partitioning, dividing the visible light image and the temperature matrix features into regional blocks; generating an adaptive weight matrix by calculating the average brightness and temperature difference of the regional blocks: for each regional block, performing weighted fusion on the visible light image and the temperature matrix features through the adaptive weight matrix; obtaining the fused overall feature map through the fused feature blocks of all regional blocks.

7. The method for detecting road cracks by thermal-optical matrix fusion according to claim 6, wherein the dynamic resolution adjustment includes: for each regional block, using a directional filter for edge detection to generate a direction matrix for determining the main direction of cracks within the regional block; designing an adaptive convolutional kernel for each regional block according to the direction matrix and applying it to the extraction of visible light image and temperature matrix features respectively; wherein, in the design of the adaptive convolutional kernel, the shape and size of the convolutional kernel are dynamically adjusted according to the direction and width features of the cracks; calculating an adaptive fusion weight within each regional block according to the temperature difference and edge sharpness; fusing the visible light and temperature matrix features within each regional block based on the adaptive fusion weight to obtain a fused feature block: obtaining the overall feature map of dynamic fusion according to all regional blocks.

8. The method for detecting road cracks by thermal-optical matrix fusion according to claim 7, wherein the global feature enhancement includes: splicing the overall feature map of regional block adaptive weighted fusion and the overall feature map of regional block-level dynamic fusion to form a multi-scale feature map; performing a global pooling operation on the multi-scale feature map to generate a globally enhanced multi-scale feature map.

Citation Information

Patent Citations

  • Infrared image and visible light image fusion method

    CN117391983A

  • Salient target detection method based on infrared and visible light image fusion

    CN117935006A