Concrete crack depth detection method and system based on multi-modal data fusion
By using multimodal data fusion technology, combining visible light and thermal imaging images, pseudo-infrared images are generated and temperature field transfer is performed, which solves the efficiency and accuracy problems of concrete crack depth detection and realizes pixel-level depth prediction and three-dimensional map visualization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OCEAN UNIVERSITY
- Filing Date
- 2025-10-28
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies for detecting the depth of concrete cracks suffer from problems such as low detection efficiency, discrete results, difficulty in achieving large-area batch detection and rapid data processing, and poor versatility of image-based detection methods, which cannot achieve pixel-level depth prediction.
By simultaneously acquiring visible light and thermal imaging images, and utilizing multimodal data fusion technology, including active infrared thermal imaging, multimodal data registration, and heat conduction inversion models, pseudo-infrared images are generated and temperature field migration is performed. The crack depth is calculated by combining the temperature decay characteristic curve, and a three-dimensional crack map is generated.
It improves the versatility and accuracy of concrete crack depth detection, realizes pixel-level depth prediction and visual archiving, and solves the problems of insufficient model versatility and intuitive data presentation in traditional technologies.
Smart Images

Figure CN121033018B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of concrete crack detection technology, and in particular to a method and system for detecting the depth of concrete cracks based on multimodal data fusion. Background Technology
[0002] In the construction industry, concrete structures are widely used in various infrastructures and buildings, and their safety is directly related to the safety of people's lives and property. Concrete cracks, as one of the most common hidden dangers in concrete structures, not only affect the appearance of the structure, but more seriously, they weaken its strength and durability, leading to a decrease in its load-bearing capacity and even causing safety accidents. Therefore, accurately detecting the depth of concrete cracks is crucial for assessing the safety and remaining lifespan of a structure.
[0003] Among existing technologies for detecting the depth of concrete cracks, acoustic testing is the most widely used mainstream method, and its high-precision detection capability for deep cracks has made it a common choice for engineering inspections. However, this technology has obvious limitations. The detection process requires data collection point by point, and it can only process a single crack or a small number of cracks at a time, making it impossible to quickly process data from large-area batch detections. At the same time, the detection results are mostly presented in the form of discrete data, lacking intuitive visualization records, and the data archiving process is cumbersome, making it inconvenient for subsequent traceability, comparative analysis, and engineering file management.
[0004] To address the efficiency and archiving issues of acoustic wave detection, image-based crack depth detection technology has been gradually proposed and explored for application. However, current image-based depth detection is easily limited to surface feature detection, while the complex topological structure inside cracks necessitates the construction of different models for targeted depth estimation based on different crack morphologies. This results in poor versatility of existing algorithms, making it difficult to adapt to the complex and diverse crack scenarios in engineering. Although introducing multimodal data fusion technology can alleviate the problem of poor model versatility to some extent, most current data fusion techniques simply stitch together or weight data from different modalities, only obtaining a rough depth prediction result, failing to achieve pixel-level depth prediction, and struggling to pinpoint the depth of each crack branch within the concrete crack. Summary of the Invention
[0005] This invention provides a method and system for detecting the depth of concrete cracks based on multimodal data fusion, which can improve the versatility and accuracy of concrete crack depth detection.
[0006] In a first aspect, embodiments of the present invention provide a method for detecting the depth of concrete cracks based on multimodal data fusion, comprising:
[0007] Within a preset time interval, visible light images of concrete cracks in the target area are continuously acquired to obtain a visible light image sequence. Simultaneously, thermal imaging images under the same scene are acquired using a preset active infrared thermal imaging technology to obtain a thermal imaging image sequence. The active infrared thermal imaging technology obtains the thermal imaging image sequence under the corresponding temperature change by determining the optimal thermal excitation control mode.
[0008] By using a preset multimodal data registration algorithm and a visible light image, a registration displacement vector field is predicted to generate a pseudo-infrared image. The temperature field of a thermal imaging image at the same time and in the same frame is then transferred to the pseudo-infrared image to generate a fused modal image, resulting in a registered fused image sequence. The pseudo-infrared image includes thermal imaging texture features generated through simulation and geometric structural features consistent with the visible light image.
[0009] Based on the fused image sequence, a temperature decay characteristic curve within the time interval is generated, and crack depth data is obtained through a preset heat conduction inversion model and the temperature decay characteristic curve; wherein, the crack depth data can be used to generate a three-dimensional crack map.
[0010] This invention ensures temporal and spatial consistency of data from two modalities through synchronous acquisition, avoiding errors caused by data misalignment. Active infrared thermal imaging, through optimal thermal excitation control, acquires complete data on crack temperature changes, overcoming the limitation of visible light only capturing surface geometric features and providing fundamental thermal conduction data for depth detection. By combining pseudo-infrared images with visible light geometry and simulated thermal imaging textures, the geometric mismatch problem of cross-modal data is solved. Combined with temperature field migration, pixel-level fusion of the two modalities is achieved, overcoming the limitations of traditional simple splicing or weighted fusion, ensuring that the temperature data of each pixel can be accurately correlated to the specific spatial location of the crack. The thermal diffusion law of the crack region is captured through temperature decay characteristic curves, and the thermal conduction inversion model is based on this law to achieve depth quantification calculation. This process utilizes the objective physical law that different crack shapes have different thermal conduction effects (essentially, concrete and air have different conduction effects) to calculate crack depth, thus solving the problem of insufficient model versatility caused by depth estimation based on topology in traditional technologies. Furthermore, the three-dimensional crack map visualizes the depth data, solving the problem of no intuitive archiving of discrete data from acoustic wave detection, and providing more comprehensive data support for structural safety assessment. Compared with existing technologies, the present invention can improve the versatility and accuracy of concrete crack depth detection.
[0011] Furthermore, the continuous acquisition of visible light images of concrete cracks within the target area yields a visible light image sequence. Simultaneously, using a pre-set active infrared thermal imaging technology, thermal images of the same scene are acquired to obtain a thermal image sequence. Specifically:
[0012] A global image of the target area is acquired, and features are extracted from the global image to obtain initial crack features; wherein, the initial crack features include crack surface width, crack direction, and texture complexity;
[0013] Based on the initial characteristics of the crack, the optimal mode is selected from a variety of preset thermal excitation modes, and an adjustable thermal excitation control command is generated according to the optimal mode.
[0014] The adjustable thermal excitation control command controls the infrared thermal excitation device to apply corresponding thermal excitation to the concrete cracks, and simultaneously acquires visible light images and thermal imaging images of the concrete cracks to obtain visible light image sequences and thermal imaging image sequences, respectively.
[0015] This invention analyzes initial features such as crack surface width and orientation to accurately identify crack morphological differences, avoiding the blind selection of thermal excitation modes. It matches the optimal thermal excitation mode to different crack characteristics, solving the problem of poor adaptability of traditional fixed thermal excitation to complex cracks. Adjustable commands ensure precise application of thermal excitation, guaranteeing that temperature change data accurately reflects crack depth information. Based on the data acquired using the optimized thermal excitation mode, it improves the clarity of crack details in visible light images and the temperature change recognition of thermal imaging images, laying the foundation for subsequent registration and fusion.
[0016] Furthermore, the step of predicting the registration displacement vector field using a preset multimodal data registration algorithm and a visible light image to generate a pseudo-infrared image specifically involves:
[0017] The visible light image sequence is enhanced by a preset first image enhancement technique to obtain an enhanced visible light image sequence;
[0018] For each enhanced visible light image in the enhanced visible light image sequence, a preset generative adversarial network is used to extract geometric structural features from the enhanced visible light image, and the thermal imaging texture features of the enhanced visible light image are predicted based on the pre-learned geometric-thermal imaging texture mapping relationship, so as to generate an initial pseudo-infrared image based on the geometric structural features and the thermal imaging texture features.
[0019] The enhanced visible light image and the corresponding initial pseudo-infrared image are input into a preset deformation field prediction network. Through a preset multi-scale attention mechanism, the geometric structural feature similarity between the two images is calculated, and a registration displacement vector field is generated based on the geometric structural feature similarity.
[0020] Using a preset spatial transformer network, the geometric structure features of the initial pseudo-infrared image are adjusted according to the registration displacement vector field to generate a pseudo-infrared image.
[0021] This invention addresses the issues of noise interference and feature blurring in the original image by enhancing the grayscale levels, edge details, and texture clarity of the visible light image, thus providing a high-quality image for geometric structure extraction. By accurately extracting the geometric features of the visible light image and combining them with pre-learned mapping relationships to generate a thermal imaging texture, it achieves preliminary fusion of features from two modalities and generates an initial simulated thermal imaging texture. This provides an initial migration benchmark for subsequent temperature field transfer, resolving the problem of large differences in cross-modal features and ensuring the accuracy of subsequent registration. Furthermore, by calculating feature similarity through a multi-scale attention mechanism, it generates a precise displacement vector field and performs geometric adjustments on the initial pseudo-infrared image to ensure a complete match with the geometric structure of the visible light image, guaranteeing pixel-level precise fusion for subsequent temperature field transfer.
[0022] Furthermore, the visible light image sequence is enhanced using a preset first image enhancement technique to obtain an enhanced visible light image sequence, specifically as follows:
[0023] For each visible light image in the visible light image sequence, grayscale distribution analysis is performed on the visible light image to obtain grayscale distribution results. Then, grayscale adjustment is performed on the visible light image using a preset grayscale correction algorithm and the grayscale distribution results to obtain a first intermediate visible light image.
[0024] By using a preset multi-scale edge detection algorithm, edge features are extracted from the first intermediate visible light image, and edge enhancement is performed on the first intermediate visible light image based on the edge features to obtain a second intermediate visible light image.
[0025] By using a preset local grayscale difference amplification algorithm, the texture of the second intermediate visible light image is enhanced to obtain an enhanced visible light image.
[0026] This invention addresses the issue of feature overload caused by uneven grayscale distribution in the original image by adjusting the image grayscale range, thereby improving the overall image contrast. It also solves the problem of geometric structure extraction error caused by edge blurring by accurately capturing the edge contour of the crack, thus enhancing the distinction between the crack and the background. Furthermore, it solves the problem of inaccurate thermal imaging texture mapping caused by unclear texture by magnifying the texture detail differences on the crack surface, ensuring the texture authenticity of the subsequently generated initial pseudo-infrared image.
[0027] Furthermore, the step of transferring the temperature field of the thermal imaging image at the same moment and in the same frame to the pseudo-infrared image to generate a fused modal image, resulting in a registered fused image sequence, specifically involves:
[0028] The thermal imaging image sequence is enhanced by a preset second image enhancement technique to obtain an enhanced thermal imaging image sequence.
[0029] The pseudo-infrared image and the enhanced thermal imaging image at the same time and in the same frame are input into a preset migration network, so that a migration temperature field is generated through the migration network according to the thermodynamic laws.
[0030] The migration temperature field and the geometric structural features of the pseudo-infrared image are fused through a preset multi-scale feature fusion network to generate a fused modal image and obtain an initial fused modal image sequence.
[0031] Inter-frame temperature verification is performed on the initial fused modal image sequence to smooth temperature fluctuations between adjacent fused modal images, thus obtaining a fused image sequence.
[0032] This invention addresses the issues of weak temperature signals and high noise in original thermal imaging images by improving the temperature gradient recognition of thermal imaging images and smoothing inter-frame temperature fluctuations. It precisely transfers the temperature field of thermal imaging images to pseudo-infrared images based on thermodynamic principles, avoiding the misalignment of temperature information with geometric structures in traditional fusion processes, thus achieving pixel-level temperature information matching. Furthermore, it generates a fused modal image by deeply fusing the temperature field with geometric features. Inter-frame verification smooths temperature fluctuations, ensuring the temperature stability of the fused image sequence and providing support for the accurate generation of subsequent temperature decay characteristic curves.
[0033] Furthermore, the thermal imaging image sequence is enhanced using a preset second image enhancement technique to obtain an enhanced thermal imaging image sequence, specifically as follows:
[0034] For each thermal imaging image in the thermal imaging image sequence, temperature analysis is performed on the thermal imaging image to obtain temperature analysis results, and based on the temperature analysis results, the thermal imaging image is amplified by temperature gradient to obtain an intermediate thermal imaging image sequence.
[0035] Inter-frame temperature verification is performed on the intermediate thermal imaging image sequence to smooth temperature fluctuations between adjacent thermal imaging images, resulting in an enhanced thermal imaging image sequence.
[0036] This invention addresses the problems of gentle temperature gradients and indistinct temperature features in crack regions in original thermal imaging images by amplifying temperature differences in the images, thereby improving the accuracy of temperature field migration. It also solves the problem of temperature instability in thermal imaging image sequences caused by environmental interference by smoothing temperature fluctuations between adjacent frames, ensuring that the temperature changes in the enhanced image sequence are continuous and reliable.
[0037] Furthermore, the step of generating a temperature decay characteristic curve within the time interval based on the fused image sequence, and obtaining crack depth data through a preset heat conduction inversion model and the temperature decay characteristic curve, specifically involves:
[0038] Using a preset image segmentation algorithm, semantic segmentation is performed on any fused image in the fused image sequence to obtain several consecutive ROI templates;
[0039] For each ROI template, the average temperature of all pixels in the corresponding region is calculated in each frame of the fused image sequence to generate the temperature decay characteristic curve of the ROI template, thus obtaining a set of temperature decay characteristic curves.
[0040] Using a preset heat conduction inversion model, crack depth data is output based on the set of temperature decay characteristic curves.
[0041] This invention provides a method for semantically segmenting fused images to obtain continuous Regions of Interest (ROI) templates, accurately locating crack-related regions of interest, eliminating interference from irrelevant regions, and focusing on effective detection areas. By calculating the average temperature of each ROI across different frames, a set of temperature decay characteristic curves is generated, establishing a correlation between regional temperature changes and time, providing a data foundation for depth calculation. By utilizing the physical laws of heat conduction to establish a mathematical correlation between temperature changes and crack depth, crack depth data is output, achieving quantitative depth detection.
[0042] Furthermore, the generation of the three-dimensional crack map specifically involves:
[0043] Extract the fused image at the feature time from the fused image sequence to obtain the key fused image;
[0044] Using a preset multi-view 3D reconstruction algorithm, a 3D geometric mesh skeleton of the concrete surface is constructed based on the visible light image sequence;
[0045] The key fused image is extracted using a pre-defined multi-channel feature extraction network, and the multi-channel features are mapped onto the three-dimensional geometric mesh skeleton to obtain a three-dimensional geometric mesh model; wherein, the multi-channel features include visible light texture features, temperature field features, and temperature decay rate features;
[0046] Based on the crack depth data, the crack region of the three-dimensional geometric mesh model is geometrically shifted to generate a three-dimensional crack map of the concrete cracks in the target region.
[0047] This invention constructs an initial three-dimensional geometric mesh model of a concrete surface based on a visible light image sequence, providing a basic geometric framework for the three-dimensional map and reflecting the overall morphology of the concrete surface. By extracting multi-channel features such as visible light texture, temperature field, and temperature decay rate from key fused images and mapping them onto the initial mesh model, the information dimensions of the three-dimensional model are enriched, allowing the model to simultaneously include geometric and thermal features. By adjusting the geometric displacement of the crack region according to crack depth data, a three-dimensional crack map is generated, accurately presenting the depth distribution of cracks in space, realizing the visualization of detection results, and providing more comprehensive data support for structural safety assessment.
[0048] Furthermore, the fused image at the feature time point is extracted from the fused image sequence to obtain the key fused image, specifically as follows:
[0049] The temperature decay characteristic curve is subjected to multi-scale curvature analysis using a preset Gaussian difference pyramid algorithm to obtain a set of key feature points. The set of key feature points is then clustered using a preset clustering algorithm to obtain a set of candidate feature times.
[0050] The candidate set of feature moments is solved using a preset simulated annealing algorithm to obtain the global optimal solution, and the global optimal solution is determined as the feature moment; wherein, the feature moment is the moment when the fused image sequence best matches the crack depth data;
[0051] Based on the characteristic time points, key fused images are extracted from the fused image sequence.
[0052] This invention employs multi-scale curvature analysis on temperature decay curves to extract key feature point sets and capture important temperature change nodes related to crack depth in the curves. By clustering key feature points, a candidate set is obtained, and representative potential feature moments are selected to narrow down the search range for the optimal moment. By solving the candidate set, the global optimal solution is obtained to determine the feature moment that best matches the crack depth data, ensuring that the extracted key fused image accurately reflects the crack depth characteristics.
[0053] Secondly, embodiments of the present invention provide a concrete crack depth detection device based on multimodal data fusion, comprising an image sequence acquisition module, a cross-modal registration module, and a depth data acquisition module, wherein...
[0054] The image sequence acquisition module is used to continuously acquire visible light images of concrete cracks in the target area within a preset time interval to obtain a visible light image sequence, and simultaneously acquire thermal imaging images of the same scene using a preset active infrared thermal imaging technology to obtain a thermal imaging image sequence; wherein, the active infrared thermal imaging technology acquires the thermal imaging image sequence under the corresponding temperature change by determining the optimal thermal excitation control mode.
[0055] The cross-modal registration module is used to predict the registration displacement vector field using a preset multimodal data registration algorithm and a visible light image to generate a pseudo-infrared image, and to transfer the temperature field of a thermal imaging image at the same moment and in the same frame to the pseudo-infrared image to generate a fused modal image, thus obtaining a registered fused image sequence; wherein, the pseudo-infrared image includes thermal imaging texture features generated by simulation and geometric structural features consistent with the visible light image;
[0056] The depth data acquisition module is used to generate a temperature decay characteristic curve within the time interval based on the fused image sequence, and to acquire crack depth data through a preset heat conduction inversion model and the temperature decay characteristic curve; wherein, the crack depth data can be used to generate a three-dimensional crack map.
[0057] This invention employs an image sequence acquisition module to synchronously acquire data from two modalities, ensuring temporal and spatial consistency and avoiding errors caused by data misalignment. Active infrared thermal imaging, through an optimal thermal excitation control mode, acquires complete data on crack temperature changes, overcoming the limitation of visible light only capturing surface geometric features and providing fundamental thermal conduction data for depth detection. A cross-modal registration module combines pseudo-infrared images with visible light geometry and simulated thermal imaging textures to resolve the geometric mismatch problem of cross-modal data. Combined with temperature field migration, pixel-level fusion of the two modalities is achieved, overcoming the limitations of traditional simple stitching or weighted fusion and ensuring that the temperature data of each pixel is consistent. It can accurately correlate the specific spatial location of cracks; through the depth data acquisition module, it uses the temperature decay characteristic curve to capture the heat diffusion law of the crack area, and the heat conduction inversion model realizes the depth quantification calculation based on this law. In this process, it uses the objective physical law that cracks of different shapes have different heat conduction effects (essentially, concrete and air have different conduction effects) to calculate the crack depth, thereby solving the problem of insufficient model universality caused by depth estimation based on topology in traditional technology. The three-dimensional crack map visualizes the depth data, solves the problem of no intuitive archiving of discrete data from acoustic wave detection, and provides more comprehensive data support for structural safety assessment.
[0058] The above description is merely an overview of the technical solutions of the embodiments of the present invention. In order to better understand the technical means of the embodiments of the present invention and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0059] Figure 1 A schematic diagram of a concrete crack depth detection method based on multimodal data fusion provided in an embodiment of the present invention;
[0060] Figure 2 This is a structural diagram of a concrete crack depth detection system based on multimodal data fusion, provided for an embodiment of the present invention. Detailed Implementation
[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] Example 1:
[0063] like Figure 1 As shown in the figure, a method for detecting the depth of concrete cracks based on multimodal data fusion is provided by an embodiment of the present invention, including the following steps:
[0064] S101, within a preset time interval, continuously acquire visible light images of concrete cracks in the target area to obtain a visible light image sequence, and simultaneously acquire thermal imaging images of the same scene using a preset active infrared thermal imaging technology to obtain a thermal imaging image sequence; wherein, the active infrared thermal imaging technology obtains the thermal imaging image sequence under the corresponding temperature change by determining the optimal thermal excitation control mode.
[0065] In this embodiment, the continuous acquisition of visible light images of concrete cracks within the target area to obtain a visible light image sequence, and the simultaneous acquisition of thermal imaging images of the same scene using a preset active infrared thermal imaging technology to obtain a thermal imaging image sequence, specifically involves: acquiring a global image of the target area and extracting features from the global image to obtain initial crack features; wherein, the initial crack features include crack surface width, crack direction, and texture complexity; based on the initial crack features, selecting the optimal mode from a variety of preset thermal excitation modes, and generating an adjustable thermal excitation control command based on the optimal mode; using the adjustable thermal excitation control command to control the infrared thermal excitation device to apply corresponding thermal excitation to the concrete cracks, and simultaneously acquiring visible light images and thermal imaging images of the concrete cracks within a preset time interval to obtain visible light image sequences and thermal imaging image sequences respectively.
[0066] In one specific embodiment, a high-resolution visible light camera with a recognition accuracy of 0.05mm is selected as the image acquisition device, and it is fixed in front of the target area with a tripod. The center of the camera lens is kept horizontally aligned with the center of the target area, and the distance is adjusted according to the size of the target area to ensure that a single image can completely cover the target area (if the target area exceeds 0.5㎡, 2-3 images can be taken by panning the camera, and then the global image can be synthesized by image stitching technology).
[0067] Furthermore, under natural or auxiliary lighting, a visible light camera acquires a global image of the target area (covering the entire crack area, typically ranging from 100mm × 100mm to 500mm × 500mm). Simultaneously, an infrared thermal imaging camera acquires an initial thermal image under the same field of view for subsequent reference.
[0068] It should be noted that the infrared thermal imaging camera used is FLIR A655sc (resolution 640×512 pixels, temperature measurement range -20-150℃, accuracy ±2%), and it is synchronized with the visible light camera through hardware triggering (the trigger signal is output by STM32F407, triggered on the rising edge, and the synchronization error is <10μs).
[0069] Furthermore, digital image processing is performed on the visible light global image to extract the initial features of the crack, specifically including:
[0070] (1) Crack surface width: The crack pixel width is calculated by edge detection algorithm (such as Canny operator) and converted into physical width by combining camera calibration parameters (e.g., each pixel corresponds to an actual size of 0.05mm), with an accuracy of ±0.01mm.
[0071] (2) Crack direction: The main direction of the crack is extracted by using Histogram of Oriented Gradients (HOG) or skeletonization algorithm (angle range 0°-180°, accuracy ±1).
[0072] (3) Texture complexity: Calculated based on gray-level co-occurrence matrix (GLCM), with a window size of 5×5 pixels and a step size of 1 pixel. The contrast (range 0-1000) and entropy (range 0-8) are calculated. The result is mapped to the 0-1 range (0 for smooth and 1 for high complexity) by normalization formula (texture complexity = (contrast / 1000 + entropy / 8) / 2).
[0073] Furthermore, multiple thermal excitation modes are preset, including continuous wave mode, pulse mode, and harmonic mode. Each mode corresponds to different thermal excitation parameters (such as power, duration, and frequency). For example:
[0074] (1) Continuous wave mode: suitable for wide cracks (width > 0.5 mm), power range of 100W-500W, duration of 5-10 seconds.
[0075] (2) Pulse mode: suitable for narrow cracks (width ≤ 0.5 mm), with a peak power of up to 1000 W and a pulse width of 1-5 milliseconds.
[0076] (3) Harmonic mode: suitable for complex texture cracks, with a power modulation frequency of 0.1Hz-10Hz.
[0077] Furthermore, based on the initial characteristics of the crack, the optimal mode is selected through a decision tree algorithm: if the crack width is large (>0.5mm) and the direction is straight, the continuous wave mode is selected to ensure uniform thermal excitation; if the crack width is small (≤0.5mm) and the texture complexity is high, the pulse mode is selected to enhance thermal contrast; if the crack direction is curved and the texture is complex, the harmonic mode is selected to balance penetration depth and resolution.
[0078] Furthermore, based on the selected mode, a digital control signal (such as a PWM signal) is generated, and the power and timing of the infrared thermal excitation device (such as a halogen lamp or laser) are adjusted by a microcontroller (such as an ARM Cortex-M series). The power adjustment accuracy is ±5W, and the timing control accuracy is ±0.1 seconds.
[0079] Furthermore, the infrared thermal excitation device applies thermal excitation to the concrete crack area according to control commands. The duration of thermal excitation is consistent with the preset time interval (usually 10-60 seconds, adjusted according to the characteristics of the crack) to ensure that the temperature change range is within 10°C-50°C, avoiding overheating damage to the concrete surface.
[0080] It should be noted that during thermal excitation, the visible light camera and the infrared thermal imaging camera synchronously acquire images at a fixed frame rate (e.g., 10-30 fps for visible light image acquisition and 5-15 fps for thermal imaging image acquisition). The acquisition time interval is controlled by a timer to ensure consistent sequence length (typically 100-500 frames).
[0081] Preferably, the visible light image sequence is saved in RGB format with a resolution of not less than 1920×1080, and is used to record the changes in crack geometry over time.
[0082] Preferably, the thermal imaging image sequence is saved as a temperature matrix format (each pixel corresponds to a temperature value, with an accuracy of ±0.1°C) to record the temperature field distribution in the crack region.
[0083] Furthermore, the acquired image sequences are transmitted in real time to a processing unit (such as an industrial computer) and undergo preliminary verification (such as frame alignment check and brightness equalization) to ensure data quality.
[0084] S102, using a preset multimodal data registration algorithm and a visible light image, a registration displacement vector field is predicted to generate a pseudo-infrared image, and the temperature field of a thermal imaging image at the same moment and in the same frame is transferred to the pseudo-infrared image to generate a fused modal image, resulting in a registered fused image sequence; wherein, the pseudo-infrared image includes thermal imaging texture features generated by simulation and geometric structural features consistent with the visible light image;
[0085] In this embodiment, the step of predicting the registration displacement vector field to generate a pseudo-infrared image using a preset multimodal data registration algorithm and a visible light image specifically involves: enhancing the visible light image sequence using a preset first image enhancement technique to obtain an enhanced visible light image sequence; for each enhanced visible light image in the enhanced visible light image sequence, extracting geometric structural features using a preset generative adversarial network, and predicting the thermal imaging texture features of the enhanced visible light image based on a pre-learned geometry-thermal imaging texture mapping relationship, thereby generating an initial pseudo-infrared image based on the geometric structural features and thermal imaging texture features; inputting the enhanced visible light image and the corresponding initial pseudo-infrared image into a preset deformation field prediction network, calculating the similarity of the geometric structural features of the two images using a preset multi-scale attention mechanism, and generating a registration displacement vector field based on the similarity of the geometric structural features; and adjusting the geometric structural features of the initial pseudo-infrared image using a preset spatial transformer network based on the registration displacement vector field to generate the pseudo-infrared image.
[0086] In this embodiment, the step of enhancing the visible light image sequence using a preset first image enhancement technique to obtain an enhanced visible light image sequence specifically involves: for each visible light image in the visible light image sequence, performing grayscale distribution analysis on the visible light image to obtain a grayscale distribution result, and adjusting the grayscale of the visible light image using a preset grayscale correction algorithm and the grayscale distribution result to obtain a first intermediate visible light image; extracting edge features from the first intermediate visible light image using a preset multi-scale edge detection algorithm, and enhancing the edge of the first intermediate visible light image based on the edge features to obtain a second intermediate visible light image; and enhancing the texture of the second intermediate visible light image using a preset local grayscale difference amplification algorithm to obtain an enhanced visible light image.
[0087] In one specific embodiment, the acquired visible light image sequence (resolution not less than 1920×1080, frame rate 10-30 fps) is enhanced, including:
[0088] (1) Gray-scale distribution analysis: Perform gray-scale histogram statistics on each frame of visible light image to analyze its dynamic range (usually 0-255 gray levels). If the image is too dark (average gray value < 50) or too bright (average gray value > 200), then perform gray-scale adjustment.
[0089] (2) Gray-scale correction: Adaptive histogram equalization (CLAHE) is used, the block size is set to 8×8 pixels, and the cliplimit is 2.0. The gray-scale of the visible light image is adjusted to control the average gray-scale value in the range of 120-180, and the first intermediate visible light image is obtained.
[0090] (3) Edge feature extraction: A multi-scale edge detection algorithm is applied, using Gaussian kernels of different sizes (such as 3×3, 5×5, and 7×7 pixels) for convolution to extract edge features from coarse to fine. The crack edges are enhanced in a targeted manner to increase the edge width by 1-2 pixels and improve the contrast by 20%-30%.
[0091] (4) Texture enhancement: Calculate the gray difference (maximum value - minimum value in the window) within a 5×5 local window. For windows with a gray difference > 10, enlarge them using the formula (enhanced gray value = original gray value + (gray difference - 10) × 1.8). For windows with a gray difference ≤ 10, keep them unchanged to obtain an enhanced visible light image sequence.
[0092] Furthermore, a Conditional Generative Adversarial Network (cGAN) was used, where the generator was a U-Net structure, the input was an enhanced visible light image (256×256×3), the encoder contained 5 convolutional layers (3×3 kernels, stride 2, LeakyReLU activation function, slope 0.2), the decoder contained 5 deconvolutional layers (3×3 kernels, stride 2, ReLU activation function), and skip connections passed shallow geometric features; the discriminator adopted the PatchGAN architecture, containing 4 convolutional layers (4×4 kernels, stride 2, LeakyReLU activation function, slope 0.2), and the output was a 30×30 patch matching probability map; the dataset consisted of 5000 pairs of visible light-thermal imaging images (covering different crack types), the optimizer was Adam (learning rate 0.0002, β1=0.5), the loss function was adversarial loss (BCEWithLogitsLoss) + L1 loss (weight 0.5), and the number of iterations was 100 epochs.
[0093] During data processing, the enhanced visible light image is input into the generator. The trained GAN generator extracts the geometric structural features (edge coordinates, curvature, etc.) of the enhanced visible light image. Based on the pre-learned mapping table (e.g., straight edges correspond to continuous temperature gradients in thermal imaging, and corners correspond to temperature change points), the thermal imaging texture features are predicted and fused to generate the initial pseudo-infrared image (resolution 256×256×1, temperature range 20-60℃).
[0094] Furthermore, the enhanced visible light image and the corresponding initial pseudo-infrared image are input into two parallel encoder branches of a deformation field prediction network structure based on the VoxelMorph architecture. The geometric structural feature similarity between the two images is calculated in a multi-scale feature space (original scale, 1 / 2 downsampled, 1 / 4 downsampled): cosine similarity is used to measure the correlation between feature vectors; key feature regions (such as crack edges and corners) are highlighted by an attention weight matrix; the similarity threshold is set to 0.7, and regions with a value lower than this are marked as regions that need to be registered in detail; based on the feature similarity calculation results, the displacement vector of each pixel (including displacement components in the x and y directions) is predicted by a regression network to generate a dense registration displacement vector field.
[0095] It should be noted that the displacement vector field resolution is consistent with the input image, and the displacement range is controlled within ±15 pixels to ensure that the deformation is within a reasonable range.
[0096] It should be noted that the deformation field prediction network adopts an encoder-decoder architecture, which includes 3 downsampling layers (3×3 convolutional kernels, stride 2) and 3 upsampling layers (3×3 transposed convolutional kernels, stride 2), and embeds a multi-scale attention mechanism (attention weights are calculated via Softmax).
[0097] Furthermore, based on the registration displacement vector field, a pixel coordinate mapping relationship is established from the input image to the output image; nonlinear spatial transformations, including translation, rotation, and local deformation, are performed on the initial pseudo-infrared image to ensure that its geometric structure is precisely aligned with the visible light image; a bilinear interpolation algorithm is used to ensure that the transformed image remains smooth and continuous, avoiding holes and jagged edges; and the final pseudo-infrared image is generated.
[0098] It should be noted that the spatial transformer network adopts a differentiable architecture and supports end-to-end training.
[0099] In this embodiment, the step of transferring the temperature field of the thermal imaging image at the same moment and in the same frame to the pseudo-infrared image to generate a fused modal image and obtain a registered fused image sequence specifically involves: enhancing the thermal imaging image sequence using a preset second image enhancement technique to obtain an enhanced thermal imaging image sequence; inputting the pseudo-infrared image and the enhanced thermal imaging image at the same moment and in the same frame to a preset transfer network to generate a transferred temperature field according to thermodynamic laws; fusing the transferred temperature field with the geometric structural features of the pseudo-infrared image using a preset multi-scale feature fusion network to generate a fused modal image and obtain an initial fused modal image sequence; and performing inter-frame temperature verification on the initial fused modal image sequence to smooth temperature fluctuations between adjacent fused modal images to obtain the fused image sequence.
[0100] In this embodiment, the step of enhancing the thermal imaging image sequence using a preset second image enhancement technique to obtain an enhanced thermal imaging image sequence specifically involves: for each thermal imaging image in the thermal imaging image sequence, performing temperature analysis on the thermal imaging image to obtain a temperature analysis result, and based on the temperature analysis result, performing temperature gradient amplification on the thermal imaging image to obtain an intermediate thermal imaging image sequence; and performing inter-frame temperature verification on the intermediate thermal imaging image sequence to smooth temperature fluctuations between adjacent thermal imaging images to obtain the enhanced thermal imaging image sequence.
[0101] In one specific embodiment, the acquired thermal imaging image sequence (temperature resolution ±0.1°C, spatial resolution not less than 320×240 pixels) is input. Temperature distribution statistics are performed on each frame of the thermal imaging image to calculate the global temperature range (typically 15°C-80°C) and temperature gradient distribution. The distribution characteristics of hot spots (the top 5% of pixels with the highest temperature) and cold spots (the bottom 5% of pixels with the lowest temperature) are analyzed. An adaptive temperature enhancement algorithm is used to non-linearly amplify areas with significant temperature gradients (gradient value > 2°C / pixel) with a magnification factor of 1.5-3.0 times, enhancing the temperature contrast between crack edges and the background. A temperature consistency verification mechanism is established between adjacent frames (time interval 0.1-0.5 seconds), and a sliding window (window size 5-10 frames) is used to smooth temperature fluctuations, eliminating abnormal temperature jumps (areas with single-frame temperature changes > 5°C are smoothed and corrected). The resulting enhanced thermal imaging image sequence has a 30%-50% improvement in temperature contrast and an inter-frame temperature fluctuation standard deviation < 0.5°C.
[0102] Furthermore, the temperature field migration network adopts an encoder-decoder structure. The encoder part includes a thermodynamic feature extraction module and sets the following thermodynamic constraints: (1) Heat conduction constraint: The heat conduction equation is introduced into the loss function to ensure that the temperature field after migration conforms to the physical laws of heat conduction; (2) Boundary conditions: The temperature discontinuity at the crack boundary is maintained, and the boundary temperature gradient is maintained at 3-8°C / mm; (3) Energy conservation: Ensure that the total thermal energy change in the region before and after migration is <5%.
[0103] Specifically, a pseudo-infrared image and a simultaneously enhanced thermal image are input into a transfer network. A feature-level attention mechanism aligns the spatial features of the two modalities at multiple scales. Based on the aligned features, the decoder generates a transfer temperature field that precisely matches the spatial location of the pseudo-infrared image. The transfer temperature field maintains the temperature range of the real thermal image (20°C-60°C) while perfectly aligning with the geometry of the pseudo-infrared image.
[0104] Furthermore, the multi-scale feature fusion network adopts a pyramid structure, which includes three fusion scales (original scale, half scale, and quarter scale).
[0105] Specifically, for low-level feature fusion, at the 1 / 4 scale, global temperature distribution features and basic geometric structures are mainly fused; for mid-level feature fusion, at the 1 / 2 scale, local temperature gradient features and edge orientation features are fused; and for high-level feature fusion, at the original scale, detailed temperature change features and texture features are fused.
[0106] Specifically, for the crack region, the geometric structure weight is 0.6 and the temperature field weight is 0.4; for the background region, the geometric structure weight is 0.3 and the temperature field weight is 0.7; for the transition region, an adaptive weight is adopted, which is dynamically adjusted based on the temperature gradient and edge intensity.
[0107] The feature reconstruction module upsamples the fused multi-scale features to the original resolution to generate a fused modality image.
[0108] It should be noted that the inter-frame temperature verification mechanism includes: (1) temporal consistency verification: performing time series analysis on 5-10 consecutive fused images to detect abnormal temperature fluctuation points; (2) spatial continuity verification: checking the temperature continuity between adjacent pixels and smoothing out temperature change areas (gradient > 5°C / pixel).
[0109] It should be noted that the image sequence optimization includes: (1) Temperature smoothing: using temporal Gaussian filtering (time window 3-5 frames, σ=1.0) to smooth temperature fluctuations; (2) Edge preservation: using an edge-preserving filtering algorithm in the crack edge area to ensure that the geometric structure is not blurred; (3) Anomaly removal: correcting or removing pixels whose temperature values exceed the reasonable range (<15°C or >80°C).
[0110] Furthermore, the registered fused image sequence is obtained, with each frame containing: accurate geometric structure information (derived from pseudo-infrared images); real temperature field data (derived from thermal imaging images); and temporally smooth and continuous temperature change characteristics.
[0111] S103, Based on the fused image sequence, a temperature decay characteristic curve within the time interval is generated, and crack depth data is obtained through a preset heat conduction inversion model and the temperature decay characteristic curve; wherein, the crack depth data can be used to generate a three-dimensional crack map.
[0112] In this embodiment, the step of generating a temperature decay characteristic curve within the time interval based on the fused image sequence, and obtaining crack depth data through a preset heat conduction inversion model and the temperature decay characteristic curve, specifically involves: performing semantic segmentation on any fused image in the fused image sequence using a preset image segmentation algorithm to obtain several consecutive ROI templates; for each ROI template, calculating the average temperature of all pixels in the corresponding region on each frame of the fused image sequence to generate the temperature decay characteristic curve of the ROI template, thus obtaining a set of temperature decay characteristic curves; and outputting crack depth data based on the set of temperature decay characteristic curves using a preset heat conduction inversion model.
[0113] In one specific embodiment, the step of generating a temperature decay characteristic curve within the time interval based on the fused image sequence, and obtaining crack depth data through a preset heat conduction inversion model and the temperature decay characteristic curve, specifically involves:
[0114] First, perform data preprocessing and quality verification time alignment checks, including the following:
[0115] (1) Check the continuity of timestamps in the fused image sequence: correct inter-frame time jitter to ensure uniform time intervals (deviation < 0.1 seconds); establish a unified time axis with the starting point being the start of thermal excitation.
[0116] (2) Temperature data calibration: Temperature calibration is performed based on known reference temperature points; the influence of ambient temperature fluctuations is corrected (reference temperature 20°C ± 5°C); abnormal temperature values (data points that exceed the reasonable range of 15°C-80°C) are removed.
[0117] (3) Spatial consistency verification: check the stability of image registration between adjacent frames; correct pixel offset caused by device micro-movement (maximum correction range ±3 pixels).
[0118] Furthermore, a pre-trained DeepLabv3+ segmentation model (trained based on 10,000+ concrete crack samples) is loaded; a representative fused image of a single frame is input (usually the frame with the highest temperature contrast is selected); pixel-level semantic segmentation is performed, and a binary mask of the crack region is output.
[0119] Furthermore, ROI template optimization is performed: independent crack segments are identified, and isolated areas with an area of less than 10 square millimeters are removed; the correct connection relationship between the main crack and its branches is ensured; the segmentation boundaries are smoothed to eliminate jagged edges; and ROIs are systematically numbered from near to far according to their spatial location.
[0120] Furthermore, for each ROI template, its temperature change is tracked over the entire time series, with the acquisition frequency synchronized with the image frame rate (5-15Hz), while the arithmetic mean of the temperature of all pixels within the ROI region is calculated.
[0121] Furthermore, each ROI generates an independent time-temperature curve, and a temporal Gaussian filter (σ=1.5 frames) is applied to smooth random fluctuations; feature point identification is performed, including the starting temperature point (at the start of thermal excitation), the peak temperature point (the moment when the temperature is highest), and the half-life point (the moment when the temperature drops to the median of the peak and the end point); the temperature data is normalized to the [0,1] interval to eliminate the influence of absolute temperature values.
[0122] Furthermore, the material parameter ranges for the heat conduction inversion model are set as follows: thermal conductivity of concrete: 1.5-3.0 W / (m·K), specific heat capacity: 900-1000 J / (kg·K), density: 2300-2500 kg / m³ 3 Setting boundary conditions: Surface thermal convection coefficient 8-12 W / (m²) 2 ·K).
[0123] In the iterative inversion process, an initial depth value (within the range of 0-50 mm) is given based on the temperature decay rate; the theoretical temperature decay curve is calculated based on the current depth estimate; the root mean square error of the theoretical curve and the measured curve is compared; the depth estimate is adjusted using the adaptive gradient descent method; iteration stops when the error is <0.5°C or the number of iterations exceeds 300; finally, a depth value is output for each ROI, with the accuracy retained to 0.1 mm, and a confidence index (based on the goodness-of-fit R-squared) is calculated. 2 The confidence index is used to output a depth distribution map, which shows the depth changes at different locations of the crack.
[0124] It should be noted that the measurement range of the crack depth data obtained by this invention is 0-50 mm, with an accuracy of ±0.5 mm.
[0125] In this embodiment, generating a three-dimensional crack map specifically involves: extracting the fused image at a feature time from the fused image sequence to obtain a key fused image; constructing a three-dimensional geometric mesh skeleton of the concrete surface based on the visible light image sequence using a preset multi-view three-dimensional reconstruction algorithm; extracting multi-channel features of the key fused image using a preset multi-channel feature extraction network, and mapping the multi-channel features onto the three-dimensional geometric mesh skeleton to obtain a three-dimensional geometric mesh model; wherein the multi-channel features include visible light texture features, temperature field features, and temperature decay rate features; and adjusting the geometric displacement of the crack region of the three-dimensional geometric mesh model based on the crack depth data to generate a three-dimensional crack map of the concrete cracks in the target area.
[0126] In one specific embodiment, 20 frames are randomly selected from the visible light image sequence, and the SIFT (Scale Invariant Feature Transform) algorithm is used to extract feature points from each frame (no less than 500 feature points are extracted per frame to ensure coverage of the concrete surface and crack areas). The FLANN matcher is used to match the feature points of adjacent frames, setting a matching distance threshold of 50, discarding incorrect matching pairs (such as matching points whose distance exceeds the threshold), and retaining correctly matched feature point pairs.
[0127] Based on the matched feature point pairs, the Structure for Motion Restoration (SfM) algorithm is used to calculate the camera's intrinsic and extrinsic parameters (intrinsic parameters include focal length and principal point coordinates, extrinsic parameters include camera position and attitude). The 3D coordinates of the feature points are then calculated through triangulation to generate a sparse point cloud on the concrete surface. The density of the sparse point cloud must meet the requirement of "no less than 10 points per square centimeter" to ensure coverage of the fine details of cracks (e.g., at least 3 points are needed to locate the edge of a 0.05mm wide crack).
[0128] The Poisson reconstruction algorithm was used to densify the sparse point cloud, with the voxel size of the dense point cloud set to 0.1 mm (matching the 0.05 mm recognition accuracy of a visible light camera), generating a dense point cloud covering the concrete surface. Subsequently, a greedy projection triangulation algorithm was used to convert the dense point cloud into a three-dimensional geometric mesh skeleton—setting the triangular facet side length threshold to 0.2 mm to ensure that the mesh accurately reflects the unevenness of the concrete surface, and that the mesh density in crack areas is higher than that in non-crack areas (facet side length ≤ 0.1 mm in crack areas).
[0129] Furthermore, the key fused image is input into a pre-defined multi-channel feature extraction network (based on a CNN architecture, containing 5 convolutional layers and 2 pooling layers) to extract three types of core features:
[0130] (1) Visible light texture features: The surface texture of the crack in the image (such as the roughness of the crack edge) is extracted by the first and second convolutional layers. The output feature map size is 1 / 4 of the original image, and the number of feature channels is 64.
[0131] (2) Temperature field features: The temperature difference between the cracked area and the non-cracked area in the image (e.g., the temperature at the crack is 2-5℃ lower than the surrounding area) is extracted by the 3rd-4th convolutional layer. The output feature map size is 1 / 8 of the original image, and the number of feature channels is 128.
[0132] (3) Temperature decay rate feature: The temperature decay rate of each pixel in the corresponding image (e.g., the decay rate at the crack is 0.01℃ / min higher than the surrounding area) is calculated by combining the temperature decay curve through the 5th convolutional layer. The output feature map size and temperature field feature are then used. Figure 1 The number of channels is 64.
[0133] Three types of multi-channel features are matched to their corresponding positions in the 3D geometric mesh skeleton using a coordinate mapping algorithm—using the pixel coordinates of the visible light image and the 3D coordinates of the mesh skeleton as the mapping reference to ensure a one-to-one correspondence between features and each triangular facet on the mesh. For example, the texture features of the crack edge are mapped to the crack edge facet of the mesh skeleton, and the temperature field features of the low-temperature region are mapped to the crack interior facet of the mesh skeleton, ultimately generating a 3D geometric mesh model with multi-channel feature information.
[0134] Furthermore, the crack depth data (e.g., a crack segment with a depth of 15mm ± 0.5mm) is associated with the three-dimensional geometric mesh model—by matching the mesh patches with the ROI template, the corresponding crack area in the model is located and marked as the area to be adjusted.
[0135] Based on the crack depth data, the displacement along the normal vector direction of the concrete surface is calculated for each grid patch in the area to be adjusted—the displacement direction is inside the concrete (i.e., away from the surface), and the displacement magnitude is equal to the crack depth value of that area (e.g., for an area with a depth of 15mm, the displacement is set to 15mm). If the crack depth in a certain area fluctuates (e.g., 14.5-15.5mm), the average depth is taken as the displacement to ensure a continuous transition of displacement within the area.
[0136] A mesh deformation algorithm (such as Laplacian deformation) is used to adjust the marked areas to be adjusted, keeping the mesh shape of non-crack areas unchanged during the adjustment process to avoid distortion of the overall model. After adjustment, the 3D geometric mesh model is textured (e.g., blue represents crack areas, with the color depth increasing as the crack depth increases; gray represents non-crack areas), finally generating a 3D crack map that includes surface morphology (width, orientation), depth distribution, and temperature characteristics. The map can be rotated and zoomed for viewing (e.g., when zoomed in 10 times, crack details as small as 0.05mm can be clearly observed).
[0137] In this embodiment, the key fused image is obtained by extracting the fused image at a characteristic moment from the fused image sequence. Specifically, the temperature decay characteristic curve is subjected to multi-scale curvature analysis using a preset Gaussian difference pyramid algorithm to obtain a set of key feature points. The set of key feature points is then clustered using a preset clustering algorithm to obtain a candidate set of characteristic moments. The candidate set of characteristic moments is then solved using a preset simulated annealing algorithm to obtain a global optimal solution, and the global optimal solution is determined as the characteristic moment. The characteristic moment is the moment when the fused image sequence best matches the crack depth data. Based on the characteristic moment, the key fused image is extracted from the fused image sequence.
[0138] In one specific embodiment, the multi-scale curvature analysis of the temperature decay characteristic curves is carried out by: importing the set of temperature decay characteristic curves into data processing software (such as MATLAB), smoothing each curve by using the moving average method, setting the window size to 5 data points, eliminating random noise in the curves, and ensuring that the curve trend is continuous and stable.
[0139] For each smoothed temperature decay curve, a Gaussian difference pyramid algorithm was used for multi-scale analysis. The pyramid was set to have 4 layers, with Gaussian filter standard deviations of 1.0, 2.0, 4.0, and 8.0 for each layer. Through filtering at different scales, the characteristic changes of the curve in different stages, such as rapid temperature decrease in the short term, gradual temperature decay in the medium term, and stable temperature in the long term, were captured.
[0140] The curvature values of the processed curves for each layer of the pyramid are calculated. The points corresponding to the peak curvature are the key feature points (representing the moment when the temperature decay rate changes significantly, such as the inflection point where the temperature changes from rapid to slow after thermal excitation stops). A curvature threshold of 0.02 is set (calibrated according to the crack depth detection range of 0-50mm). Points with curvature values exceeding the threshold are selected to form a subset of key feature points for each curve. Finally, all subsets are summarized to form the key feature point set.
[0141] Furthermore, the process of obtaining the clustering of the candidate set of feature moments is as follows: extract the time coordinates corresponding to each point in the key feature point set (i.e., the frame number of the fused image sequence, such as the 10th frame corresponding to the acquisition time of 15 minutes), and use the time coordinates as the clustering input data.
[0142] The K-means clustering algorithm was used to cluster the time coordinates. The number of clusters K was determined based on the total number of frames in the fused image sequence. If the sequence contained 60 frames (acquisition time 30 min, 1 frame every 30 s), then K was set to 5-8 classes to ensure that the time points contained in each class had similar temperature decay characteristics. During the clustering process, the standard deviation of the time points within a class was less than 2 frames (i.e. 1 min) as the termination condition to avoid excessive dispersion of data within the class.
[0143] Calculate the average value for each type of time point, and use the time corresponding to the average value as the "candidate feature time" (e.g., if a certain type of time point is the 12th, 13th, and 14th frames, the average value is 13 frames, and the corresponding candidate time is 19.5 min). Summarize all candidate feature times to form a candidate feature time set (e.g., {5.5 min, 12 min, 19.5 min, 25 min}).
[0144] Furthermore, the process of determining the globally optimal feature time is as follows: an objective function is constructed based on the matching degree between the fused image and the crack depth data corresponding to the candidate time. The matching degree is calculated in two aspects: first, the deviation between the temperature value of the ROI region in the image at that time and the theoretical temperature value output by the heat conduction inversion model (the smaller the deviation, the higher the matching degree); second, the degree of overlap between the geometric shape (width, direction) of the crack in the image at that time and the average shape of the crack in the visible light image sequence (overlap ≥ 90% is excellent).
[0145] The simulated annealing algorithm was used to optimize the objective function. The initial temperature was set to 100°C, the cooling rate to 0.95, and the number of iterations to 100. During the algorithm iteration process, the candidate time with the highest matching degree was gradually selected. In the initial stage, time with a lower matching degree was allowed to be selected. As the temperature decreased, the algorithm gradually focused on the optimal solution. Finally, when the iteration converged, the output global optimal solution was the feature time (e.g., 19.5 min, corresponding to the 13th frame of the fused image sequence).
[0146] Based on the determined feature time, the corresponding frame image is extracted from the fused image sequence and used as the key fused image.
[0147] This invention ensures temporal and spatial consistency of data from two modalities through synchronous acquisition, avoiding errors caused by data misalignment. Active infrared thermal imaging, through optimal thermal excitation control, acquires complete data on crack temperature changes, overcoming the limitation of visible light only capturing surface geometric features and providing fundamental thermal conduction data for depth detection. By combining pseudo-infrared images with visible light geometry and simulated thermal imaging textures, the geometric mismatch problem of cross-modal data is solved. Combined with temperature field migration, pixel-level fusion of the two modalities is achieved, overcoming the limitations of traditional simple splicing or weighted fusion, ensuring that the temperature data of each pixel can be accurately correlated to the specific spatial location of the crack. The thermal diffusion law of the crack region is captured through temperature decay characteristic curves, and the thermal conduction inversion model is based on this law to achieve depth quantification calculation. This process utilizes the objective physical law that different crack shapes have different thermal conduction effects (essentially, concrete and air have different conduction effects) to calculate crack depth, thus solving the problem of insufficient model versatility caused by depth estimation based on topology in traditional technologies. Furthermore, the three-dimensional crack map visualizes the depth data, solving the problem of no intuitive archiving of discrete data from acoustic wave detection, and providing more comprehensive data support for structural safety assessment. Compared with existing technologies, the present invention can improve the versatility and accuracy of concrete crack depth detection.
[0148] Example 2:
[0149] like Figure 2As shown, this embodiment provides a concrete crack depth detection system based on multimodal data fusion, including an image sequence acquisition module 201, a cross-modal registration module 202, and a depth data acquisition module 203, wherein...
[0150] The image sequence acquisition module 201 is used to continuously acquire visible light images of concrete cracks in the target area within a preset time interval to obtain a visible light image sequence, and simultaneously acquire thermal imaging images of the same scene using a preset active infrared thermal imaging technology to obtain a thermal imaging image sequence; wherein, the active infrared thermal imaging technology acquires the thermal imaging image sequence under the corresponding temperature change by determining the optimal thermal excitation control mode.
[0151] In this embodiment, the image sequence acquisition module 201 continuously acquires visible light images of concrete cracks within the target area to obtain a visible light image sequence. Simultaneously, using a preset active infrared thermal imaging technology, it acquires thermal imaging images of the same area to obtain a thermal imaging image sequence. Specifically, the image sequence acquisition module 201 acquires a global image of the target area and extracts features from the global image to obtain initial crack features. These initial crack features include crack surface width, crack direction, and texture complexity. Based on the initial crack features, an optimal mode is selected from a set of preset thermal excitation modes to generate an adjustable thermal excitation control command. Using the adjustable thermal excitation control command, the infrared thermal excitation device is controlled to apply corresponding thermal excitation to the concrete cracks. Within a preset time interval, visible light images and thermal imaging images of the concrete cracks are simultaneously acquired to obtain visible light image sequences and thermal imaging image sequences, respectively.
[0152] The cross-modal registration module 202 is used to predict the registration displacement vector field using a preset multimodal data registration algorithm and a visible light image to generate a pseudo-infrared image, and to transfer the temperature field of a thermal imaging image at the same moment and in the same frame to the pseudo-infrared image to generate a fused modal image, thereby obtaining a registered fused image sequence; wherein, the pseudo-infrared image includes thermal imaging texture features generated by simulation and geometric structural features consistent with the visible light image;
[0153] In this embodiment, the cross-modal registration module 202 predicts the registration displacement vector field using a preset multimodal data registration algorithm and a visible light image to generate a pseudo-infrared image. Specifically, the cross-modal registration module 202 enhances the visible light image sequence using a preset first image enhancement technique to obtain an enhanced visible light image sequence. For each enhanced visible light image in the enhanced visible light image sequence, a preset generative adversarial network extracts geometric structural features from the enhanced visible light image and predicts the thermal imaging texture features of the enhanced visible light image based on a pre-learned geometry-thermal imaging texture mapping relationship. Based on the geometric structural features and thermal imaging texture features, an initial pseudo-infrared image is generated. The enhanced visible light image and the corresponding initial pseudo-infrared image are input into a preset deformation field prediction network to calculate the similarity of the geometric structural features of the two images using a preset multi-scale attention mechanism, and a registration displacement vector field is generated based on the similarity of the geometric structural features. Based on the registration displacement vector field, the geometric structural features of the initial pseudo-infrared image are adjusted using a preset spatial transformer network to generate a pseudo-infrared image.
[0154] In this embodiment, the cross-modal registration module 202 transfers the temperature field of the thermal imaging image at the same moment and in the same frame to the pseudo-infrared image to generate a fused modal image, resulting in a registered fused image sequence. Specifically, the cross-modal registration module 202 enhances the thermal imaging image sequence using a preset second image enhancement technique to obtain an enhanced thermal imaging image sequence; the pseudo-infrared image and the enhanced thermal imaging image at the same moment and in the same frame are input into a preset transfer network to generate a transfer temperature field according to thermodynamic laws; the transfer temperature field is fused with the geometric structural features of the pseudo-infrared image using a preset multi-scale feature fusion network to generate a fused modal image, resulting in an initial fused modal image sequence; inter-frame temperature verification is performed on the initial fused modal image sequence to smooth temperature fluctuations between adjacent fused modal images, resulting in a fused image sequence.
[0155] The depth data acquisition module 203 is used to generate a temperature decay characteristic curve within the time interval based on the fused image sequence, and to acquire crack depth data through a preset heat conduction inversion model and the temperature decay characteristic curve; wherein, the crack depth data can be used to generate a three-dimensional crack map.
[0156] In this embodiment, the depth data acquisition module 203 generates a temperature decay characteristic curve within the time interval based on the fused image sequence, and acquires crack depth data through a preset heat conduction inversion model and the temperature decay characteristic curve. Specifically, the depth data acquisition module 203 performs semantic segmentation on any fused image in the fused image sequence using a preset image segmentation algorithm to obtain several consecutive ROI templates; for each ROI template, the average temperature of all pixels in the corresponding region is calculated in each frame of the fused image sequence to generate the temperature decay characteristic curve of the ROI template, thus obtaining a set of temperature decay characteristic curves; and the crack depth data is output based on the set of temperature decay characteristic curves using a preset heat conduction inversion model.
[0157] For a more detailed explanation of the working principle and procedures of this embodiment, please refer to the relevant description in Embodiment 1.
[0158] In this embodiment of the invention, the image sequence acquisition module 201 synchronously acquires data to ensure temporal and spatial consistency between the two modalities, avoiding errors caused by data misalignment. Active infrared thermal imaging, through an optimal thermal excitation control mode, acquires complete data on crack temperature changes, overcoming the limitation of visible light only capturing surface geometric features and providing fundamental thermal conduction data for depth detection. The cross-modal registration module 202 combines pseudo-infrared images with visible light geometry and simulated thermal imaging textures to solve the problem of geometric mismatch in cross-modal data. Combined with temperature field migration, pixel-level fusion of the two modalities is achieved, overcoming the limitations of traditional simple stitching or weighted fusion, ensuring the temperature data of each pixel. All of these methods can accurately correlate to the specific spatial location of the crack. Through the depth data acquisition module 203, the thermal diffusion law of the crack area is captured by the temperature decay characteristic curve. The thermal conduction inversion model realizes the depth quantification calculation based on this law. In this process, the objective physical law that cracks of different shapes have different thermal conduction effects (essentially, concrete and air have different conduction effects) is used to calculate the crack depth, thereby solving the problem of insufficient model universality caused by depth estimation based on topology in traditional technology. The three-dimensional crack map visualizes the depth data, solves the problem of no intuitive archiving of discrete data from acoustic wave detection, and provides more comprehensive data support for structural safety assessment.
[0159] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0160] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for detecting the depth of concrete cracks based on multimodal data fusion, characterized in that, include: Within a preset time interval, visible light images of concrete cracks in the target area are continuously acquired to obtain a visible light image sequence. Simultaneously, thermal imaging images under the same scene are acquired using a preset active infrared thermal imaging technology to obtain a thermal imaging image sequence. The active infrared thermal imaging technology obtains the thermal imaging image sequence under the corresponding temperature change by determining the optimal thermal excitation control mode. By using a preset multimodal data registration algorithm and a visible light image, a registration displacement vector field is predicted to generate a pseudo-infrared image. The temperature field of a thermal imaging image at the same time and in the same frame is then transferred to the pseudo-infrared image to generate a fused modal image, resulting in a registered fused image sequence. The pseudo-infrared image includes thermal imaging texture features generated through simulation and geometric structural features consistent with the visible light image. Specifically, the step of predicting a registration displacement vector field to generate a pseudo-infrared image using a preset multimodal data registration algorithm and a visible light image involves: enhancing the visible light image sequence using a preset first image enhancement technique to obtain an enhanced visible light image sequence; for each enhanced visible light image in the enhanced visible light image sequence, extracting geometric structural features using a preset generative adversarial network, and predicting the thermal imaging texture features of the enhanced visible light image based on a pre-learned geometry-thermal imaging texture mapping relationship, thereby generating an initial pseudo-infrared image based on the geometric structural features and thermal imaging texture features; inputting the enhanced visible light image and the corresponding initial pseudo-infrared image into a preset deformation field prediction network, calculating the similarity of the geometric structural features of the two images using a preset multi-scale attention mechanism, and generating a registration displacement vector field based on the similarity of the geometric structural features; and adjusting the geometric structural features of the initial pseudo-infrared image using a preset spatial transformer network based on the registration displacement vector field to generate the pseudo-infrared image. Based on the fused image sequence, a temperature decay characteristic curve within the time interval is generated, and crack depth data is obtained through a preset heat conduction inversion model and the temperature decay characteristic curve; wherein, the crack depth data is used to generate a three-dimensional crack map.
2. The method for detecting concrete crack depth based on multimodal data fusion as described in claim 1, characterized in that, The process involves continuously acquiring visible light images of concrete cracks within the target area to obtain a visible light image sequence. Simultaneously, using a pre-set active infrared thermal imaging technology, thermal images of the same area are acquired to obtain a thermal image sequence. Specifically: A global image of the target area is acquired, and features are extracted from the global image to obtain initial crack features; wherein, the initial crack features include crack surface width, crack direction, and texture complexity; Based on the initial characteristics of the crack, the optimal mode is selected from a variety of preset thermal excitation modes, and an adjustable thermal excitation control command is generated according to the optimal mode. The adjustable thermal excitation control command controls the infrared thermal excitation device to apply corresponding thermal excitation to the concrete cracks, and simultaneously acquires visible light images and thermal imaging images of the concrete cracks within a preset time interval to obtain visible light image sequences and thermal imaging image sequences, respectively.
3. The method for detecting concrete crack depth based on multimodal data fusion as described in claim 1, characterized in that, The process of enhancing the visible light image sequence using a preset first image enhancement technique to obtain an enhanced visible light image sequence specifically involves: For each visible light image in the visible light image sequence, grayscale distribution analysis is performed on the visible light image to obtain grayscale distribution results. Then, grayscale adjustment is performed on the visible light image using a preset grayscale correction algorithm and the grayscale distribution results to obtain a first intermediate visible light image. By using a preset multi-scale edge detection algorithm, edge features are extracted from the first intermediate visible light image, and edge enhancement is performed on the first intermediate visible light image based on the edge features to obtain a second intermediate visible light image. The second intermediate visible light image is texture-enhanced using a preset local grayscale difference amplification algorithm to obtain an enhanced visible light image.
4. The method for detecting concrete crack depth based on multimodal data fusion as described in claim 1, characterized in that, The process of transferring the temperature field of the thermal imaging image at the same moment and in the same frame to the pseudo-infrared image to generate a fused modal image and obtain a registered fused image sequence is as follows: The thermal imaging image sequence is enhanced by a preset second image enhancement technique to obtain an enhanced thermal imaging image sequence; The pseudo-infrared image and the enhanced thermal imaging image at the same time and in the same frame are input into a preset migration network, so that a migration temperature field is generated through the migration network according to the thermodynamic laws. The migration temperature field and the geometric structural features of the pseudo-infrared image are fused through a preset multi-scale feature fusion network to generate a fused modal image and obtain an initial fused modal image sequence. Inter-frame temperature verification is performed on the initial fused modal image sequence to smooth temperature fluctuations between adjacent fused modal images, thus obtaining a fused image sequence.
5. The method for detecting concrete crack depth based on multimodal data fusion as described in claim 4, characterized in that, The process involves enhancing the thermal imaging image sequence using a preset second image enhancement technique to obtain an enhanced thermal imaging image sequence, specifically as follows: For each thermal imaging image in the thermal imaging image sequence, temperature analysis is performed on the thermal imaging image to obtain temperature analysis results, and based on the temperature analysis results, the thermal imaging image is amplified by temperature gradient to obtain an intermediate thermal imaging image sequence. Inter-frame temperature verification is performed on the intermediate thermal imaging image sequence to smooth temperature fluctuations between adjacent thermal imaging images, resulting in an enhanced thermal imaging image sequence.
6. The method for detecting concrete crack depth based on multimodal data fusion as described in claim 1, characterized in that, The step involves generating a temperature decay characteristic curve within the time interval based on the fused image sequence, and obtaining crack depth data using a preset heat conduction inversion model and the temperature decay characteristic curve. Specifically: Using a preset image segmentation algorithm, semantic segmentation is performed on any fused image in the fused image sequence to obtain several consecutive ROI templates; For each ROI template, the average temperature of all pixels in the corresponding region is calculated in each frame of the fused image sequence to generate the temperature decay characteristic curve of the ROI template, thus obtaining a set of temperature decay characteristic curves. Using a preset heat conduction inversion model, crack depth data is output based on the set of temperature decay characteristic curves.
7. The method for detecting concrete crack depth based on multimodal data fusion as described in claim 1, characterized in that, The generation of the three-dimensional crack map is specifically as follows: Extract the fused image at the feature time from the fused image sequence to obtain the key fused image; Specifically, extracting the fused image at a characteristic moment from the fused image sequence to obtain the key fused image involves: performing multi-scale curvature analysis on the temperature decay characteristic curve using a preset Gaussian difference pyramid algorithm to obtain a set of key feature points; clustering the key feature point set using a preset clustering algorithm to obtain a candidate set of characteristic moments; solving the candidate set of characteristic moments using a preset simulated annealing algorithm to obtain the global optimal solution, and determining the global optimal solution as the characteristic moment; wherein, the characteristic moment is the moment when the fused image sequence best matches the crack depth data; and extracting the key fused image from the fused image sequence based on the characteristic moment. Using a preset multi-view 3D reconstruction algorithm, a 3D geometric mesh skeleton of the concrete surface is constructed based on the visible light image sequence; The key fused image is extracted using a pre-defined multi-channel feature extraction network, and the multi-channel features are mapped onto the three-dimensional geometric mesh skeleton to obtain a three-dimensional geometric mesh model; wherein, the multi-channel features include visible light texture features, temperature field features, and temperature decay rate features; Based on the crack depth data, the crack region of the three-dimensional geometric mesh model is geometrically shifted to generate a three-dimensional crack map of the concrete cracks in the target region.
8. A concrete crack depth detection system based on multimodal data fusion, characterized in that, It includes an image sequence acquisition module, a cross-modal registration module, and a depth data acquisition module, among which, The image sequence acquisition module is used to continuously acquire visible light images of concrete cracks in the target area within a preset time interval to obtain a visible light image sequence, and simultaneously acquire thermal imaging images of the same scene using a preset active infrared thermal imaging technology to obtain a thermal imaging image sequence; wherein, the active infrared thermal imaging technology acquires the thermal imaging image sequence under the corresponding temperature change by determining the optimal thermal excitation control mode. The cross-modal registration module is used to predict the registration displacement vector field using a preset multimodal data registration algorithm and a visible light image to generate a pseudo-infrared image, and to transfer the temperature field of a thermal imaging image at the same moment and in the same frame to the pseudo-infrared image to generate a fused modal image, thus obtaining a registered fused image sequence; wherein, the pseudo-infrared image includes thermal imaging texture features generated by simulation and geometric structural features consistent with the visible light image; Specifically, the step of predicting a registration displacement vector field to generate a pseudo-infrared image using a preset multimodal data registration algorithm and a visible light image involves: enhancing the visible light image sequence using a preset first image enhancement technique to obtain an enhanced visible light image sequence; for each enhanced visible light image in the enhanced visible light image sequence, extracting geometric structural features using a preset generative adversarial network, and predicting the thermal imaging texture features of the enhanced visible light image based on a pre-learned geometry-thermal imaging texture mapping relationship, thereby generating an initial pseudo-infrared image based on the geometric structural features and thermal imaging texture features; inputting the enhanced visible light image and the corresponding initial pseudo-infrared image into a preset deformation field prediction network, calculating the similarity of the geometric structural features of the two images using a preset multi-scale attention mechanism, and generating a registration displacement vector field based on the similarity of the geometric structural features; and adjusting the geometric structural features of the initial pseudo-infrared image using a preset spatial transformer network based on the registration displacement vector field to generate the pseudo-infrared image. The depth data acquisition module is used to generate a temperature decay characteristic curve within the time interval based on the fused image sequence, and to acquire crack depth data through a preset heat conduction inversion model and the temperature decay characteristic curve; wherein, the crack depth data is used to generate a three-dimensional crack map.
Citation Information
Patent Citations
Railway sleeper crack detection method and system based on three-dimensional feature construction
CN120147315A
Power line identification method and device based on image fusion, equipment and medium
CN120689709A