Building defect detection method and intelligent imaging device
By constructing a dual-modal dataset with multi-material temperature difference annotation and an improved YOLOv8m-seg network, efficient and accurate building defect detection is achieved, solving the problems of low detection efficiency and high false detection rate in existing technologies, and is applicable to various building scenarios.
Patent Information
- Application Number
- CN202511268896.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-06
- Publication Date
- 2025-11-18
AI Technical Summary
Existing building exterior wall defect detection technologies suffer from problems such as low detection efficiency, high missed detection rate, high false detection rate, difficulty in accurately quantifying defect size and depth, insufficient adaptability to multi-material scenarios, and lack of temperature difference annotation information in datasets.
A dual-modal dataset with temperature difference annotations for multiple materials was constructed. Image alignment was performed using ORB feature matching and feature alignment modules. The improved YOLOv8m-seg network was combined for dual-modal fusion detection. The results were optimized through a joint loss function to output a quantified defect detection report.
It significantly improves detection accuracy in various material scenarios, reduces the false detection rate to less than 3%, increases the recognition rate of minute cracks to 92%, improves detection efficiency by 40 times, and reduces maintenance costs by 70%, making it suitable for various building scenarios.
Smart Images

Figure CN120976768A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of building structural health monitoring technology, specifically to a method for detecting building defects based on the fusion of infrared thermal imaging and visible light dual-modal imaging, and a matching intelligent imaging device. Background Technology
[0002] The existing methods for detecting defects in building exterior walls have the following technical challenges: Traditional manual inspection methods require working at heights, which not only results in low inspection efficiency (the average daily inspection area per person is no more than 500 square meters), but also has a false negative rate of about 15%, and cannot accurately quantify the size and depth information of defects.
[0003] In terms of single-mode detection technology, visible light detection is easily affected by surface dirt and lighting conditions, and its recognition rate for fine cracks with a width of less than 0.2 mm is less than 70%. Although infrared detection can identify hidden defects, it is difficult to accurately locate the edge of the defect. At the same time, due to the lack of a clear multi-material temperature difference threshold standard, the false detection rate exceeds 8%.
[0004] At the data and model level, existing datasets generally lack temperature difference annotation information for different building materials, the alignment accuracy of bimodal images is poor (error exceeds 3 pixels), existing models have failed to optimize for building defect characteristics, and their adaptability in multi-material scenarios is insufficient.
[0005] For example, although the wall detection method disclosed in Chinese patent CN120451154A adopts dual-modal fusion technology, it only performs simple feature stitching through the CameraViTNets model, which fails to effectively solve the problems of multi-material temperature difference threshold setting and robustness in dirty scenarios; while the single-modal detection technology based on YOLOv8m-seg cannot achieve the coordinated use of infrared and visible light data.
[0006] Therefore, there is an urgent need to develop a building defect detection technology that can integrate the advantages of dual-modality, adapt to multiple material scenarios, and have the ability to resist dirt interference. Summary of the Invention
[0007] To overcome the shortcomings mentioned above, the present invention aims to provide a technical solution that can solve the above problems.
[0008] This invention provides a method for detecting defects in buildings, comprising the following steps: S1. Construct a dual-modal dataset with temperature difference annotations for multiple materials: Collect image pairs of “normal region-defect region” composed of infrared thermal imaging images and visible light images, and annotate the defect type, defect size, material type, normal region temperature difference range and defect region temperature difference value. After data augmentation, the image pairs are divided into training set, validation set and test set. S2. Dual-modal image alignment: First, the ORB feature matching algorithm is used to coarsely align the infrared thermal image and the visible light image. After removing mismatched points, an initial transformation matrix is obtained. Then, the feature alignment module (FAM) performs deformable convolution adjustment on the infrared feature map to achieve fine correction, so that the alignment error of the dual-modal image is ≤ a preset threshold (this threshold ensures that the subsequent defect location accuracy meets the engineering inspection requirements). S3. Dual-modal fusion detection: The aligned 3-channel visible light image and 1-channel infrared temperature difference image are stitched together to form a multi-channel input, which is then input to the improved YOLOv8m-seg network. The network includes a dual-modal feature branch and a cross-attention fusion layer. After training with a joint loss function, the network outputs the location coordinates, category, and segmentation mask of the defect. S4. Post-processing and output of results: Perform non-maximum suppression processing on the detection results output by the network to remove overlapping detection boxes and output a quantified defect detection report.
[0009] Further, the specific standards for the multi-material temperature difference marking mentioned in step S1 are as follows: Under a 25℃ environment, the temperature difference of the normal area of the concrete exterior wall is ±0.3~±0.7℃, and the temperature difference of the defective area is ≥0.8~1.2℃; the temperature difference of the normal area of the brick exterior wall is ±0.6~±1.0℃, and the temperature difference of the defective area is ≥1.0~1.4℃; the temperature difference of the normal area of the steel exterior wall is ±0.2~±0.4℃, and the temperature difference of the defective area is ≥0.6~1.0℃; the temperature difference of the normal area of the glass curtain wall is ±0.8~±1.2℃, and the temperature difference of the defective area is ≥1.8~2.2℃.
[0010] Furthermore, the data augmentation described in step S1 includes the following operations: Basic transformations: Simultaneously perform random rotation of -15° to 15°, random scaling of 0.8 to 1.2x, and horizontal or vertical flipping on the bimodal image; Dirt simulation: Dust, water stains, or oil stains are embedded in visible light images to simulate real-world pollution scenarios. The dust coverage rate is 5%~30%, the water stain coverage rate is 10%~20%, and the oil stain coverage rate is 5%~15%. The temperature difference noise in the corresponding infrared regions is adjusted simultaneously. The temperature difference noise for dust is increased by +0.1~+0.3℃, for water stains by +0.2~+0.4℃, and for oil stains by +0.3~+0.5℃. Color gamut adjustment: Perform random offsets of ±8%~±12% for brightness and ±12%~±18% for contrast in visible light images, and perform random offsets of ±3%~±7% for grayscale values in infrared images.
[0011] Further: The specific implementation process of the ORB feature matching algorithm in step S2 is as follows: ≥40 stable feature points are extracted from the infrared thermal imaging image and the visible light image respectively (feature points in the building outline, corners and door and window frame areas are selected first). After the feature points are matched by the brute-force matching algorithm, the mismatched points with an error >2~4 pixels are removed by the RANSAC algorithm to obtain the coarse alignment transformation matrix. The feature alignment module (FAM) dynamically adjusts the spatial position of the infrared feature points through deformable convolution to compensate for the residual error of coarse alignment, so that the alignment error of the finely corrected dual-modal image is ≤1~3 pixels.
[0012] Furthermore: The specific structure of the improved YOLOv8m-seg network described in step S3 is as follows: Dual-modal feature branches: The visible light branch adopts a "Conv-BN-ReLU" structure, which extracts material texture and defect edge features through a 3×3 convolution kernel, and the number of channels is gradually increased from 4 to 64; the infrared branch adopts a "Conv-BN-Sigmoid" structure, which enhances the temperature difference features between the defect area and the normal area through a 1×1 convolution kernel, and the number of channels is synchronously matched with the visible light branch. Cross-attention fusion layer: After concatenating the bimodal feature maps by channel dimension, feature weights are dynamically assigned—the infrared feature weight in the defect region is 0.5~0.7, and the visible light feature weight is 0.3~0.5, while the infrared feature weight in the non-defect region is 0.3~0.5, and the visible light feature weight is 0.5~0.7. Joint loss function: CIoU loss, Focal Loss and Dice Loss are combined with a weight ratio of 0.2~0.4:0.3~0.5:0.2~0.4. CIoU loss is used to optimize the defect localization accuracy, Focal Loss is used to solve the class imbalance problem between defect and non-defect regions, and Dice Loss is used to optimize the edge fitting accuracy of defect segmentation mask.
[0013] Furthermore: the training process of the improved YOLOv8m-seg network described in step S3 includes: Transfer learning: Load the pre-trained weights of YOLOv8m-seg on the COCO public dataset, freeze the first 4 to 6 layers of the backbone network, and train only the neck and head structures to shorten the training cycle. Hyperparameter settings: training batch size is 10~20, initial learning rate is 0.0008~0.0012, cosine annealing strategy is used for learning rate decay, and the total number of training epochs is 40~60. Early stopping mechanism: When the mean accuracy (mAP@0.5) of the validation set does not improve for 8 to 12 consecutive epochs, training is stopped, and two types of model files are saved: "recent weight file" and "best weight file". The best weight file corresponds to the training state with the highest mAP@0.5 on the validation set.
[0014] The present invention also provides an intelligent imaging device for detecting building defects, including a dual-mode acquisition unit, a pan-tilt control unit, a data processing unit, and a power supply unit that are interconnected via a data bus or control line and powered by a power line. The dual-modal acquisition unit is used to simultaneously acquire infrared thermal imaging images and visible light images, and supports synchronous trigger control to ensure spatiotemporal alignment. The gimbal control unit is used to drive the dual-modal acquisition unit to achieve 360° rotation, with a positioning accuracy of ≤±0.2°, and supports automatic route planning based on preset paths; The data processing unit is used to execute the dual-modal image alignment algorithm and the dual-modal fusion detection algorithm, and output the defect detection results; The power supply unit is used to supply power to each unit, supporting ≥6~10 hours of battery life and fast charging function.
[0015] Furthermore: the dual-modal acquisition unit includes an infrared thermal imager and an industrial visible light camera; the infrared thermal imager has a temperature measurement range of -20℃ to 300℃, a resolution of ≥320×240, and a noise equivalent temperature difference of ≤50mK; the industrial visible light camera has an image resolution of ≥2592×1944 and a frame rate of ≥25fps; the synchronization triggering accuracy of the dual-modal acquisition unit is ≤1~3ms, ensuring that the acquired infrared thermal image and visible light image correspond one-to-one in time and space.
[0016] Furthermore: the computing power of the data processing unit is ≥180~220 TOPS, supporting real-time inference of the improved YOLOv8m-seg network, and the single detection time for a 640×640 pixel bimodal image is ≤0.3~0.5s; the data processing unit also integrates an image storage module, supporting the storage of ≥2~4 months of detection data (including images, defect information, and environmental parameters), and can export defect reports in Excel or PDF format, the reports containing the correspondence between "detection area - defect coordinates - defect size - temperature difference value - detection confidence level".
[0017] Furthermore, it also includes a display interaction unit, which is an 8-12 inch touch screen that supports real-time preview of dual-modal images—simultaneously displaying the temperature difference annotation information of the infrared image (including the temperature difference value between the normal area and the defect area), the defect detection box, segmentation mask and defect category information of the visible light image; the display interaction unit also supports the retrospective query of historical detection data, and can load defect images of the same detection area at different times for comparison to assist in the analysis of defect expansion trends.
[0018] Compared with the prior art, the beneficial effects of the present invention are: (I) Detection Methodology: Overcoming Bottlenecks in Multi-Scenario Adaptability and Accuracy Significantly improved multi-material detection accuracy: By constructing a temperature difference labeled dataset containing six types of materials, including concrete, brick walls, steel, and glass curtain walls, and setting different defect temperature difference thresholds (e.g., concrete defect temperature difference ≥0.8-1.2℃, steel ≥0.6-1.0℃), the false detection rate in multi-material mixed scenarios has been reduced from more than 8% in the existing technology to less than 3%, and the recognition rate of fine cracks (width 0.1-0.2mm) has been improved to more than 92%.
[0019] The advantages of dual-modal collaboration are fully realized: the two-level alignment strategy of "ORB coarse alignment + FAM fine correction" is adopted to control the alignment error of dual-modal images to 1-3 pixels, which is 67% more accurate than traditional dual-modal technology; the improved YOLOv8m-seg network, through "dual-modal feature branch + cross-attention fusion layer", prioritizes the enhancement of infrared temperature difference features in defect areas (weight 0.5-0.7) and the enhancement of visible light texture features in non-defect areas (weight 0.5-0.7), and the recall rate in dirty scenes (dust coverage rate of 30%) is still maintained above 89%, which is 33% more effective than single-modal visible light detection.
[0020] Balancing model training efficiency and generalization: By introducing transfer learning (freezing the first 4-6 layers of the YOLOv8m-seg backbone network) and dynamic loss function (joint optimization of CIoU+Focal Loss+Dice Loss), the model training cycle is shortened from 100 epochs to 40-60 epochs. At the same time, the detection accuracy deviation in different climate regions (high temperature, high humidity) is controlled within 5%, and the generalization ability is significantly better than existing technologies.
[0021] (II) Intelligent Imaging Device Level: Achieving Engineering Implementation and Convenient Operation Integrated design lowers the application threshold: The device integrates a dual-modal acquisition unit (synchronous triggering accuracy ≤1-3ms), a gimbal control unit (360° rotation, positioning accuracy ±0.2°), a high-performance data processing unit (≥180-220TOPS), and a long-lasting power supply (6-10 hours of battery life, 1-3 hours of fast charging). It does not require multiple external devices and can be directly mounted on wall-climbing robots, drones, or used handheld devices, increasing outdoor inspection efficiency by 40 times (daily inspection area exceeds 20,000㎡).
[0022] Balancing real-time performance and interactivity: The data processing unit supports real-time inference of 640×640 pixel dual-modal images (single detection time ≤0.3-0.5s), and the 8-12 inch touch screen can simultaneously preview infrared temperature difference thermal maps and visible light defect annotations (including detection boxes and segmentation masks). Operators can intuitively obtain defect information without professional algorithm knowledge. The integrated 1TB NVMe SSD + 2TB HDD hybrid storage module supports 2-4 months of detection data archiving and can automatically export quantitative reports in Excel / PDF format (including defect coordinates, dimensions, and temperature difference values), avoiding errors from manual transcription.
[0023] (III) Engineering Application Level: Reducing Operation and Maintenance Costs and Ensuring Structural Safety Significantly reduced operation and maintenance costs: Compared with traditional manual inspection, this invention can reduce the number of personnel required for high-altitude operations by 70% and shorten the inspection cycle by 80%. Taking a 150m high super high-rise residential building as an example, the cost of a single exterior wall inspection is reduced from 200,000 yuan to less than 50,000 yuan. At the same time, by reviewing historical data (comparing defect images at different times), the crack propagation rate can be quantitatively analyzed (accuracy ±0.1mm / month), providing a basis for "preventive repair" and avoiding the million-level structural repair costs caused by defect expansion.
[0024] Wide adaptability to various scenarios: The device and method are not only suitable for the inspection of the exterior walls of civil buildings, but can also be extended to bridge piers, exterior walls of industrial plants, tunnel linings and other scenarios. In bridge inspection, the automatic flight path planning of the gimbal can cover the entire area of the bridge side wall, and the recognition rate of "hidden fractures" of steel supports reaches 95%. In industrial plant scenarios, it can withstand ambient temperatures of -20℃ to 40℃, meeting the building operation and maintenance needs of different industries.
[0025] In summary, this invention significantly improves the accuracy of defect detection in multiple scenarios (false detection rate ≤3%, minute crack recognition rate ≥92%) by constructing a multi-material temperature difference labeled dual-modal dataset, optimizing the "ORB coarse alignment + FAM fine correction" alignment strategy, and improving the YOLOv8m-seg network fusion detection, combined with an intelligent imaging device integrating dual-modal acquisition and high-computing-power processing units. It also greatly reduces operation and maintenance costs (single detection cost reduced by 75%), adapts to multiple building scenarios, and achieves efficient, accurate, and safe building defect detection.
[0026] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a flowchart illustrating the overall steps of the present invention; Figure 2 This is a detailed flowchart of the construction of the bimodal dataset in S1 of the present invention; Figure 3 This is a detailed flowchart of S2: dual-modal image alignment of the present invention; Figure 4 This is a detailed flowchart of S3 of the present invention: dual-modal fusion detection; Figure 5 This is a detailed flowchart of S4 of the present invention: result post-processing and output; Figure 6 This is a diagram of the improved YOLOv8m-seg network structure of the present invention; Figure 7 This is the overall system architecture of the present invention; Figure 8 This is a schematic diagram of the intelligent imaging device structure and the human-computer interaction interface of the present invention. Detailed Implementation
[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Please see Figures 1 to 8 In this embodiment of the invention, a method for detecting building defects is proposed, comprising the following steps: A bimodal dataset with multi-material temperature difference annotations was constructed: image pairs consisting of infrared thermal imaging images and visible light images, representing "normal region - defect region," were collected. Defect type, defect size, material type, normal region temperature difference range, and defect region temperature difference value were annotated. After data augmentation, the image pairs were divided into training, validation, and test sets. Data augmentation included three methods: basic transformation, dirt simulation, and color gamut adjustment. Basic transformation involved simultaneously performing random rotation, scaling, and flipping on the bimodal images. Dirt simulation involved embedding dust, water stains, or oil stains into the visible light images and simultaneously adjusting the temperature difference noise in the corresponding infrared regions. Color gamut adjustment involved performing brightness and contrast shifts on the visible light images and grayscale shifts on the infrared images.
[0031] Dual-modal image alignment: First, the ORB feature matching algorithm is used to coarsely align the infrared thermal image and the visible light image, eliminating mismatched points to obtain an initial transformation matrix. Then, the feature alignment module performs deformable convolution adjustment on the infrared feature map to achieve fine correction, ensuring that the alignment error of the dual-modal images does not exceed a preset threshold. As a preferred implementation, the ORB feature matching algorithm needs to extract no less than 40 stable feature points (preferably selecting feature points in low-deformation areas such as building outlines, corners, and door and window frames); the feature alignment module dynamically adjusts the spatial position of the infrared feature points through deformable convolution to compensate for residual errors from coarse alignment.
[0032] Dual-modal fusion detection: The aligned 3-channel visible light image and 1-channel infrared temperature difference image are stitched together as a multi-channel input, which is then fed into an improved YOLOv8m-seg network. The network includes a dual-modal feature branch and a cross-attention fusion layer. After training with a joint loss function, it outputs the defect location coordinates, category, and segmentation mask. Specifically, the dual-modal feature branch includes a visible light branch and an infrared branch. The visible light branch uses a "Conv-BN-ReLU" structure to extract material texture and defect edge features, while the infrared branch uses a "Conv-BN-Sigmoid" structure to enhance the temperature difference features between the defect area and the normal area. The cross-attention fusion layer dynamically assigns feature weights after stitching the dual-modal feature maps along their channel dimensions.
[0033] Post-processing and output: Non-maximum suppression is performed on the detection results output by the network to remove overlapping detection boxes, and a quantified defect detection report is output. The detection report can thus include quantitative information such as the defect's location coordinates, category, size, and temperature difference.
[0034] This technical solution addresses the lack of material-adaptive annotations in existing data by constructing a dual-modal dataset with temperature difference annotations for multiple materials. It employs a two-stage strategy of 'ORB coarse alignment + FAM fine correction' to control the alignment error of the dual-modal images within an engineering-applicable range of 1-3 pixels. Feature complementarity is achieved through multi-channel fusion and improved network structure. Finally, it outputs detection results with quantifiable metrics. Compared to existing technologies, this solution systematically solves the problems in the three key stages of data annotation, alignment, and fusion, significantly improving defect detection accuracy in multi-material scenarios.
[0035] Furthermore, the present invention proposes that, under a 25℃ environment, the temperature difference in the normal area of a concrete exterior wall is ±0.3~±0.7℃, and the temperature difference in the defective area is ≥0.8~1.2℃; the temperature difference in the normal area of a brick exterior wall is ±0.6~±1.0℃, and the temperature difference in the defective area is ≥1.0~1.4℃; the temperature difference in the normal area of a steel exterior wall is ±0.2~±0.4℃, and the temperature difference in the defective area is ≥0.6~1.0℃; and the temperature difference in the normal area of a glass curtain wall is ±0.8~±1.2℃, and the temperature difference in the defective area is ≥1.8~2.2℃.
[0036] Specifically, the temperature difference labeling standard can be achieved through the following methods: High-precision infrared thermal imagers (noise equivalent temperature difference ≤50mK) are used to collect sample data. The ambient temperature is controlled at 25±0.5℃ in a constant-temperature laboratory environment. Artificial defects (such as cracks, hollow areas, and delamination) are created on standard test blocks of four different materials. The temperature difference distribution of 100 or more valid data samples is statistically analyzed to determine the threshold boundaries between normal and defective areas. For concrete, due to its moderate thermal conductivity (approximately 1.28W / (m·K)), the normal temperature difference range is set at ±0.3~±0.7℃. In defective areas, the temperature difference significantly increases to ≥0.8℃ due to the thermal insulation effect of the air layer. For steel, due to its high thermal conductivity (approximately 45W / (m·K)), a narrower normal range of ±0.2~±0.4℃ is used, but a temperature difference of ≥0.6℃ can still be measured in defective areas due to the thermal bridging effect. For glass curtain walls, due to the thermal resistance characteristics of their double-layer hollow structure, the temperature difference in defective areas must reach ≥1.8℃ for reliable detection. As a preferred implementation method, a material-temperature difference comparison database can be established, and the corresponding threshold can be automatically called during testing.
[0037] Therefore, this technical solution, through experimentally verified material-specific temperature difference standards, solves the problem of misjudgment caused by a uniform threshold in the inspection of multi-material building exterior walls. Compared with existing technologies, its innovation lies in: for the first time, establishing differentiated normal / defect temperature difference threshold systems for four typical building materials: concrete, brick walls, steel, and glass curtain walls. Concrete and steel use a narrower normal range (±0.3~±0.7℃ / ±0.2~±0.4℃) to match their thermal conductivity, while brick walls and glass curtain walls use a wider range (±0.6~±1.0℃ / ±0.8~±1.2℃) to accommodate their thermal inertia. All defect thresholds are higher than the upper limit of the normal range by more than 0.5℃ to ensure the reliability of the criteria. Actual measurements show that this solution reduces the false detection rate in multi-material mixed scenarios from >8% in the background technology to less than 3%, and is robust to dirt interference because the temperature difference threshold is set based on the inherent thermal properties of the material and is not affected by the optical effects of surface contamination.
[0038] Furthermore, the present invention proposes that data augmentation includes the following operations: Basic transformation: synchronously performing random rotation of -15° to 15°, random scaling of 0.8 to 1.2 times, and horizontal or vertical flipping on the bimodal image; Dirt simulation: embedding dust, water stains, or oil stains into the visible light image to simulate actual pollution scenarios, wherein the dust coverage rate is 5% to 30%, the water stain coverage rate is 10% to 20%, and the oil stain coverage rate is 5% to 15%, and simultaneously adjusting the temperature difference noise of the corresponding infrared region, with the temperature difference noise for dust increasing by +0.1 to +0.3℃, for water stains by +0.2 to +0.4℃, and for oil stains by +0.3 to +0.5℃; Color gamut adjustment: performing random offsets of ±8% to ±12% in brightness and ±12% to ±18% in contrast on the visible light image, and performing random offsets of ±3% to ±7% in grayscale value on the infrared image.
[0039] Specifically, random rotation in the basic transformation can be achieved using an affine transformation matrix, with the rotation center set as the geometric center of the image. Bilinear interpolation is used to ensure image quality. Random scaling involves adjusting the image resolution and cropping it to the original size, with the scaled blank areas filled with edge pixels. In the dirt simulation, dust particles are generated using the Perlin noise algorithm to simulate natural distribution, water stains are achieved through the overlay of polygons with varying transparency, and oil stain textures are created by combining randomly generated elliptical regions with Gaussian blur. Brightness shift in color gamut adjustment is achieved by adjusting the V channel value after HSV color space conversion, contrast adjustment uses a linear transformation formula, and grayscale value shift in infrared images is achieved by directly modifying pixel values, with the shift amount following a uniform distribution.
[0040] Therefore, this technical solution enhances the model's robustness to changes in viewing angle through geometric transformation, establishes a physical correlation between visible light pollution and infrared noise through contamination simulation, and simulates imaging differences under different environments through color gamut adjustment. Compared with existing technologies, the basic transformation ensures the spatial consistency of bimodal data, contamination simulation fills in the gaps in real-world interference data, and color gamut adjustment maintains the feature balance between modalities. These three operations work synergistically to make the training data distribution closer to the actual detection environment, effectively improving the model's generalization ability in complex scenarios. Experiments show that after adopting this data augmentation method, the model's mAP improves by 12.3% and the false positive rate decreases by 5.8% on a test set containing contamination interference.
[0041] Furthermore, this invention proposes to extract no less than 40 stable feature points from both the infrared thermal imaging image and the visible light image, prioritizing feature points in the building outline, corners, and door and window frame areas. After matching the feature points using a brute-force matching algorithm, mismatched points with errors greater than 2 to 4 pixels are removed using the RANSAC algorithm to obtain a coarse alignment transformation matrix. The feature alignment module dynamically adjusts the spatial position of the infrared feature points through deformable convolution to compensate for the residual error of the coarse alignment, so that the alignment error of the finely corrected dual-modal image is controlled within the range of 1 to 3 pixels.
[0042] Specifically, deformable convolution can be implemented in several ways, including but not limited to the following two: The first uses an offset prediction network, generating spatial offsets of the feature map through two convolutional layers. The first convolutional layer compresses the number of channels in the input feature map to 1 / 4 of the original number, and the second convolutional layer outputs a two-dimensional offset field with the same size as the input feature map. The second uses a learnable deformable kernel, adding learnable spatial deformation parameters to the standard convolutional kernel, and dynamically adjusting the feature map position through bilinear interpolation. For building contour feature point extraction, an improved FAST algorithm can be used. By setting a corner response threshold ≥12 and a non-maximum suppression radius ≥3 pixels, feature points are ensured to be concentrated in high-contrast regions. The RANSAC algorithm's iteration count is set to 500 to 1000 times, and the confidence level is set to 0.95 to 0.99 to balance computational efficiency with mismatch removal effectiveness.
[0043] Therefore, this technical solution achieves high-precision alignment through a phased error compensation mechanism: in the coarse alignment stage, ORB feature matching is used to establish a global correspondence, and the RANSAC algorithm is used to eliminate large-scale offset errors; in the fine alignment stage, deformable convolution is used to adaptively adjust the position of local features to compensate for nonlinear deformation caused by differences in thermal expansion of materials. Compared with existing technologies, the alignment error of 3 to 5 pixels in traditional rigid transformation is reduced to 1 to 3 pixels, especially solving the problem of local deformation caused by the difference in thermal expansion coefficients between metal components and concrete in building detection. By prioritizing the selection of structural feature points of the building, more than 80% of the matching points are concentrated in rigid areas with small deformation, effectively improving the mismatch removal efficiency of the RANSAC algorithm. The introduction of deformable convolution enables the accuracy of local deformation compensation to reach the sub-pixel level, and compared with traditional affine transformation methods, the alignment error at the glass curtain wall joints is reduced by about 60%.
[0044] Furthermore, this invention proposes an improved YOLOv8m-seg network structure: In the dual-modal feature branches, the visible light branch adopts a "Conv-BN-ReLU" structure, extracting material texture and defect edge features through 3×3 convolutional kernels, with the number of channels gradually increasing from 4 to 64; the infrared branch adopts a "Conv-BN-Sigmoid" structure, enhancing the temperature difference features between defective and normal regions through 1×1 convolutional kernels, with the number of channels synchronously matching that of the visible light branch. After the cross-attention fusion layer concatenates the dual-modal feature maps by channel dimension, it dynamically allocates feature weights—the infrared feature weight in the defective region is 0.5~0.7, and the visible light feature weight is 0.3~0.5, while in the non-defective region, the infrared feature weight is 0.3~0.5, and the visible light feature weight is 0.5~0.7. The joint loss function uses CIoU loss, Focal Loss and Dice Loss in a weight ratio of 0.2~0.4:0.3~0.5:0.2~0.4. CIoU loss is used to optimize the defect localization accuracy, Focal Loss is used to solve the class imbalance problem between defect and non-defect regions, and Dice Loss is used to optimize the edge fitting accuracy of the defect segmentation mask.
[0045] Specifically, the visible light branch of the dual-modal feature branch can use grouped convolutions instead of standard convolutions to reduce computation, with the number of groups set to 4-8. The sigmoid activation function of the infrared branch can be replaced with a sigmoid variant with a temperature parameter, set to 0.8-1.2 to control the feature response intensity. Dynamic weight allocation in the cross-attention fusion layer can be achieved through spatial attention mechanisms, such as using an SE module to generate a spatial weight map or capturing long-range dependencies through coordinate attention mechanisms. The weight ratio of the joint loss function can be dynamically adjusted using a gradient normalization algorithm, emphasizing Focal Loss (weight 0.4-0.5) in the early stages of training and Dice Loss (weight 0.35-0.45) in later stages to optimize segmentation details.
[0046] Therefore, this technical solution solves the problem of unreasonable weight allocation for dual-modal features through a collaborative design of differentiated feature extraction, dynamic fusion, and multi-objective optimization. The 3×3 convolutional kernel and ReLU activation in the visible light branch effectively preserve high-frequency information of material texture, while the 1×1 convolutional kernel and Sigmoid activation in the infrared branch highlight the significant differences in temperature characteristics. The cross-attention fusion layer, through a region-adaptive weight allocation mechanism, prioritizes infrared features to locate the defect core in defect areas, while relying on visible light features to maintain background continuity in non-defect areas. The joint loss function, through multi-task collaborative optimization, reduces the localization error by 12%–15% and improves the segmentation edge accuracy by 8%–10%. Compared with existing fixed-weight fusion methods, this solution reduces the false detection rate by 5%–7% in concrete crack detection and improves the recall rate by 9%–11% in glass curtain wall bubble detection.
[0047] Furthermore, this invention proposes an improved training process for the YOLOv8m-seg network, including: transfer learning: loading pre-trained weights of YOLOv8m-seg on the COCO public dataset, freezing the first 4-6 layers of the backbone network, and training only the neck and head structures to shorten the training cycle; hyperparameter settings: training batch size of 10-20, initial learning rate of 0.0008-0.0012, using cosine annealing strategy for learning rate decay, and a total training cycle of 40-60 epochs; early stopping mechanism: when the mean accuracy of the validation set does not improve for 8-12 consecutive epochs, training is stopped, and two types of model files are saved: the most recent weight file and the best weight file. The best weight file corresponds to the training state with the highest mAP@0.5 on the validation set.
[0048] Specifically, the choice of the number of layers to freeze in the transfer learning strategy is based on the distribution characteristics of the backbone network's feature extraction capabilities. Shallow layers primarily extract general edge texture features, while deeper layers extract high-level semantic features. Freezing the first 4-6 layers preserves the general feature extraction capabilities while avoiding negative transfer of deep features due to differences in the distribution of bimodal data. In hyperparameter settings, mini-batch training combined with cosine annealing effectively balances gradient update stability and convergence speed. For example, a batch size of 15 with an initial learning rate of 0.001 can achieve approximately a 20% improvement in training speed while maintaining training stability. The early stopping mechanism is triggered when the mean accuracy (mAP@0.5) of the validation set shows no improvement for 8-12 consecutive epochs. This is based on the statistical characteristics of the fluctuation in mAP@0.5 on the validation set, avoiding premature termination of effective training or wasted resources from continuous ineffective training.
[0049] Therefore, this training scheme effectively solves the technical problems of low training efficiency and poor convergence stability of bimodal fusion detection networks through the synergistic effect of transfer learning, hyperparameter optimization, and early stopping mechanism. Transfer learning provides a high-starting-point feature representation foundation, avoiding the resource waste caused by training from scratch; the optimized hyperparameter combination improves convergence speed while ensuring training stability through mini-batch training and dynamic learning rate adjustment; the early stopping mechanism dynamically terminates invalid training rounds based on continuous monitoring of validation set performance, ensuring that the model converges to the optimal state. Compared with existing technologies, this scheme can shorten the training cycle by 30%~40% under the same hardware conditions, while improving the model convergence stability by about 25%, ultimately achieving a dual optimization of training efficiency and model performance.
[0050] Furthermore, this invention also proposes an intelligent imaging device for building defect detection, comprising a dual-modal acquisition unit, a gimbal control unit, a data processing unit, and a power supply unit, which are interconnected via a data bus or control line and powered by a power supply line. The dual-modal acquisition unit is used to simultaneously acquire infrared thermal imaging images and visible light images, supporting synchronous trigger control to ensure spatiotemporal alignment; the gimbal control unit is used to drive the dual-modal acquisition unit to achieve 360° rotation with a positioning accuracy of ≤±0.2°, supporting automatic flight path planning based on a preset path; the data processing unit is used to execute dual-modal image alignment algorithms and dual-modal fusion detection algorithms, and output defect detection results; the power supply unit is used to power each unit, supporting ≥6~10 hours of battery life and fast charging functionality.
[0051] Specifically, the dual-modal acquisition unit can employ a combination of an infrared thermal imager and an industrial visible light camera. The infrared thermal imager has a temperature measurement range of -20℃ to 300℃, a resolution ≥320×240, and a noise equivalent temperature difference ≤50mK. The industrial visible light camera has an image resolution ≥2592×1944 and a frame rate ≥25fps (frames / second). Synchronous trigger control can be achieved through hardware trigger signals, with a synchronization accuracy ≤1~3ms. The gimbal control unit can be driven by a stepper motor or servo motor and equipped with a high-precision encoder feedback system, achieving a positioning accuracy ≤±0.2°. The data processing unit has a computing power ≥180~220TOPS (e.g., using an NVIDIA Jetson AGX Orin chip, with an INT8 computing power of 275TOPS). It performs layer fusion and accuracy calibration through the TensorRT inference framework, supports real-time inference with the improved YOLOv8m-seg network, and achieves a single detection time of ≤0.3~0.5s for 640×640 pixel dual-modal images. The power unit can use a lithium polymer battery pack, which supports fast charging and can be fully charged in 1 to 3 hours.
[0052] Therefore, this device transforms the bimodal fusion detection method into a portable device through a combination of hardware integration and algorithm embedding. The bimodal acquisition unit ensures the spatiotemporal consistency of bimodal data through synchronous trigger control, providing a foundation for subsequent fusion detection. The gimbal control unit achieves full coverage of the detection area through high-precision rotation and automatic flight path planning, solving the problem of blind spots in manual detection. The data processing unit integrates bimodal alignment and fusion detection algorithms, directly outputting quantification results and avoiding human interpretation errors. The power supply unit's long battery life and fast charging function ensure the device's continuous operation capability in outdoor scenarios. All units are electrically connected to form a collaborative system. The cooperation between the bimodal acquisition unit and the data processing unit improves detection accuracy, while the cooperation between the gimbal control unit and the power supply unit improves detection efficiency. Compared with existing technologies, this device, through the collaborative design of synchronous trigger control, a high-precision gimbal, and an embedded processing unit, achieves automation of the detection process while ensuring detection accuracy, resolving the contradiction between efficiency and accuracy in traditional detection methods.
[0053] Furthermore, this invention proposes that the dual-modal acquisition unit includes an infrared thermal imager and an industrial visible light camera; the infrared thermal imager has a temperature measurement range of -20℃ to 300℃, a resolution of ≥320×240, and a noise equivalent temperature difference of ≤50mK; the industrial visible light camera has an image resolution of ≥2592×1944 and a frame rate of ≥25fps; the dual-modal acquisition unit has a synchronization trigger accuracy of ≤1~3ms, ensuring that the acquired infrared thermal image and the visible light image correspond one-to-one in time and space.
[0054] Specifically, the infrared thermal imager employs uncooled microbolometer technology, with a 320×240 resolution achieved through a vanadium oxide focal plane array, and a 50mK noise-equivalent temperature difference achieved through a three-stage thermoelectric cooling module combined with a digital noise reduction algorithm. The industrial visible light camera uses a global shutter CMOS sensor, achieving a 2592×1944 resolution with a 4.4μm pixel size, and a 25fps frame rate guaranteed by the PCIe 3.0 interface bandwidth. Synchronous trigger precision control is achieved through a hardware trigger signal generator, using a GPS synchronization clock source as the time reference. A precise trigger pulse signal is generated via an FPGA, with the pulse width controlled at the 1μs level, ensuring that the time deviation between the two devices at the exposure moment does not exceed 3ms. As a preferred implementation, the synchronous trigger mechanism can employ dual protection with both hardware trigger lines and software trigger APIs. When the hardware trigger malfunctions, it automatically switches to the PTP precision time protocol for software synchronization compensation.
[0055] Therefore, this technical solution achieves high-precision dual-modal data acquisition by strictly limiting sensor performance parameters and synchronization accuracy. The high resolution and low noise characteristics of the infrared thermal imager ensure the accuracy of thermal radiation data, while the high resolution and high frame rate configuration of the visible light camera ensure the ability to capture surface defect details. The core innovation lies in the millisecond-level synchronization triggering mechanism, which eliminates the image misalignment problem caused by traditional asynchronous acquisition through hardware-level synchronization design. Compared with existing technologies, this solution reduces the dual-modal alignment error from a typical value of 5-10ms to the order of 1-3ms, providing a spatiotemporally consistent raw data foundation for subsequent image fusion. Furthermore, through the coordinated optimization of sensor parameters and synchronization accuracy, the technical bottlenecks in timestamp alignment and spatial registration of multimodal data are solved, avoiding the computational overhead and accuracy loss caused by subsequent software compensation.
[0056] Furthermore, this invention proposes that the computing power of the data processing unit is ≥180~220 TOPS, supports real-time inference of the improved YOLOv8m-seg network, and has a single detection time of ≤0.3~0.5s for a 640×640 pixel bimodal image; the data processing unit also integrates an image storage module, supports the storage of detection data for ≥2~4 months, and can export defect reports in Excel or PDF format, the reports containing the correspondence between detection area-defect coordinates-defect size-temperature difference value-detection confidence level.
[0057] The computing power of ≥180~220 TOPS can be achieved by integrating the NVIDIA Jetson AGX Orin or Huawei Ascend 910B chip. The Jetson AGX Orin has an INT8 computing power of 275 TOPS, while the Ascend 910B has an FP16 computing power of 256 TOPS. Real-time inference functionality is optimized using the TensorRT or MindSpore Lite inference framework. TensorRT can perform layer fusion, accuracy calibration, and dynamic memory allocation optimization on the improved YOLOv8m-seg network. The single-detection time of ≤0.3~0.5s is verified as follows: under ambient temperature of 25℃ and input resolution of 640×640, 1000 consecutive inferences are performed using the COCO test set to simulate building defect images, and the P99 latency is statistically analyzed. The image storage module can employ a hybrid storage solution of 1TB NVMe SSD and 2TB HDD. The SSD is used to cache recent inspection data, while the HDD is used for long-term archiving. The storage cycle is calculated based on inspecting 1000 images per day, with each image occupying an average of 1.5MB of storage space. The report export function uses the Apache POI (Excel) or iText (PDF) library to populate structured data. The Excel template has a preset "Inspection Area" worksheet that records building facade numbers, and a "Defect Details" worksheet that links defect coordinates to temperature difference values.
[0058] Specifically, this technical solution accelerates the inference speed of the dual-modal fusion model through high-performance hardware, reducing the single detection time to less than 0.5 seconds, meeting the real-time requirements of engineering sites. The large-capacity storage design employs a tiered storage strategy, achieving data retention of more than four months within limited physical space. The standardized report generation function automatically associates detection parameters through predefined data structures, avoiding errors from manual transcription. Compared to existing technologies, the computing power is 360 times higher than traditional embedded devices (such as the Raspberry Pi 4B's 0.5 TOPS), the storage capacity is 125 times larger than pure memory solutions (such as the Jetson Xavier NX's 16GB), and the report generation efficiency is more than 8 times higher than manual recording. This forms a complete end-to-end defect detection data processing chain, effectively solving the technical bottlenecks of mobile devices in real-time processing, data storage, and result output.
[0059] It also includes a display interaction unit, which is an 8-12 inch touch screen that supports real-time preview of dual-modal images—simultaneously displaying the temperature difference annotation information of the infrared image (including the temperature difference value between the normal area and the defect area), the defect detection box, segmentation mask and defect category information of the visible light image; the display interaction unit also supports the retrospective query of historical detection data, and can load defect images of the same detection area at different times for comparison, to help analyze the defect expansion trend (such as the monthly average change in crack length and the rate of change in crack width).
[0060] Furthermore, this invention proposes a specific implementation of the display interaction unit, which uses an 8-12 inch touchscreen as the hardware carrier and implements a dual-modal image real-time preview function through an embedded system. The touchscreen's capacitive touch module supports 10-point touch, with a touch accuracy ≤0.5mm and a refresh rate ≥120Hz. The real-time preview function is implemented through the OpenGLES 3.0 graphics interface, simultaneously rendering an infrared image's temperature difference heatmap (pseudo-color encoding range of -20℃ to 300℃) and a visible light image's detection result overlay (including the detection box, segmentation mask, and category label output by YOLOv8m-seg), with rendering latency controlled within 50ms. The historical data backtracking function uses an SQLite database to store detection records (including dual-modal images, defect coordinates / sizes / categories, ambient temperature, and detection timestamps), supporting retrieval by timestamp, GPS coordinates, or building number. During comparative analysis, detection data from different periods can be retrieved using a sliding timeline control and displayed in columns on the screen, with the column layout configurable as 1×2 or 2×2.
[0061] Specifically, the temperature difference labeling information is displayed using dynamic threshold mapping technology, which automatically adjusts the color gradation range according to the current ambient temperature. The color gradation threshold for defective areas of concrete is set to ≥0.8℃ (red), while for normal areas it is ±0.3~±0.7℃ (blue to green gradient). The defect expansion trend analysis module uses an image registration algorithm to align the same defective area in historical images, automatically calculating the change in crack length (accuracy ±1mm) and width (accuracy ±0.1mm), and displaying the trend in a line graph. As a preferred implementation, the anti-glare coating on the touchscreen reduces specular reflection by 60%~70% in strong outdoor light conditions, ensuring a display brightness of ≥500cd / m² even under 10,000 lux illuminance.
[0062] Therefore, this technical solution solves the visualization bottleneck in the inspection process through an integrated human-computer interaction interface. The touchscreen size ensures portability while providing sufficient display space for simultaneous comparison of dual-modal data; the real-time preview function spatially aligns infrared thermodynamic data with visible light morphological inspection results, eliminating the visual cognitive fragmentation caused by traditional split-screen displays; the historical comparison module, through time-series data analysis, achieves quantitative assessment of defect expansion parameters, providing a dynamic basis for maintenance decisions. Compared with existing technologies, this solution transforms the inspection process from one-way data acquisition to two-way interactive verification, allowing operators to directly correct the inspection area or adjust the temperature difference threshold on the touchscreen, significantly improving the interpretability of the inspection results.
[0063] Flowchart of the technical solution in this invention I. First Phase: Construction of the Bimodal Dataset 1. Data Collection Infrared thermal imaging image acquisition: Using an infrared thermal imager with a temperature range of -20℃ to 300℃, a resolution of ≥320×240, and a noise equivalent temperature difference of ≤50mK, images of "normal area - defect area" of materials such as concrete, brick walls, steel, and glass curtain walls are captured. Visible light image acquisition: An industrial visible light camera with a resolution of ≥2592×1944 and a frame rate of ≥25fps is used to acquire images synchronously with infrared images (synchronization trigger accuracy ≤1~3ms) to ensure spatiotemporal alignment in the same area; Data requirements: ≥2000 samples for each material type, covering different ambient temperatures (-10℃~40℃) and dirty scenarios (dust, water stains, oil stains).
[0064] 1. Data labeling Labeling tools: Use labeling software such as LabelMe to manually label each image with "defect type (crack / fracture), defect size (length / width / depth), material type, normal area temperature difference range, and defect area temperature difference value"; Temperature difference standard: Under 25℃ environment, normal for concrete ±0.3~±0.7℃ / defect ≥0.8~1.2℃, normal for brick wall ±0.6~±1.0℃ / defect ≥1.0~1.4℃, normal for steel ±0.2~±0.4℃ / defect ≥0.6~1.0℃, and normal for glass curtain wall ±0.8~±1.2℃ / defect ≥1.8~2.2℃.
[0065] 1. Data Augmentation Basic transformation: Simultaneously perform "-15°~15° random rotation, 0.8~1.2x random scaling, and horizontal / vertical flipping" on the bimodal image to ensure feature consistency; Dirt simulation: "Dust (coverage 5%~30%), water stains (10%~20%), oil stains (5%~15%)" are embedded in the visible light image, and the temperature difference noise of the corresponding infrared area is adjusted synchronously (dust +0.1~+0.3℃, water stains +0.2~+0.4℃, oil stains +0.3~+0.5℃). Color gamut adjustment: Visible light image brightness ±8%~±12%, contrast ±12%~±18% random offset, infrared image grayscale value ±3%~±7% random offset.
[0066] 1. Dataset partitioning The dataset is divided into training set (≥10500 sets), validation set (≥3000 sets), and test set (≥1500 sets) in a 7:2:1 ratio. All images are uniformly 640×640 pixels, and the pixel values are normalized to the range [0,1].
[0067] II. Second Stage: Dual-Modal Image Alignment 1. Coarse Alignment: ORB Feature Matching Feature extraction: Extract ≥40 stable feature points from both infrared and visible light images (prioritize areas that are not easily deformed, such as building outlines, corners, and door and window frames); Matching and filtering: The brute-force matching algorithm is used to match feature points. The RANSAC algorithm is used to remove mismatched points with an error of 2 to 4 pixels, and an initial transformation matrix is generated to achieve coarse alignment (alignment error ≤ 5 pixels).
[0068] 1. Fine Correction: Feature Alignment Module (FAM) Optimization Deformable convolution adjustment: Deformable convolution is applied to the infrared feature map to dynamically adjust the spatial position of the infrared feature points according to the texture features of the visible light image (such as brick wall gaps and concrete pouring joints) to compensate for the residual error of coarse alignment. Accuracy verification: The "feature point reprojection error" is evaluated to ensure that the alignment error after fine correction is ≤1~3 pixels, which meets the accuracy requirements for subsequent defect positioning.
[0069] III. Third Phase: Improving YOLOv8m-seg Dual-Modal Fusion Detection 1. Network Input Layer Processing Channel stitching: The aligned "3-channel visible light image (RGB) + 1-channel infrared temperature difference image (grayscale value corresponds to temperature difference)" are stitched together into a 4-channel input, with the input size fixed at 640×640 pixels; Training-period enhancements: Perform "0.7~1.0x random pruning" on the input data to further improve the model's generalization ability.
[0070] 1. Feature extraction layer: Bimodal branch design Visible light branch: The “Conv-BN-ReLU” structure is adopted, and the material texture and defect edge features are extracted through 3×3 convolution kernels. The number of channels is gradually increased from 4 to 64 (a total of 5 convolution blocks). Infrared branch: It adopts the "Conv-BN-Sigmoid" structure, which enhances the "temperature difference between the defect area and the normal area" through 1×1 convolution kernel, and the number of channels is synchronously matched with the visible light branch (to ensure consistent feature dimensions).
[0071] 1. Feature Fusion Layer: Cross-Attention Mechanism Dynamic weight allocation: After concatenating the bimodal feature maps along the channel dimension, weights are allocated through a cross-attention module—for defective regions, "infrared feature weight 0.5~0.7, visible light feature weight 0.3~0.5" (highlighting the indicative role of temperature difference on defects), while the weights for non-defective regions are reversed (0.3~0.5:0.5~0.7). Multi-scale fusion: The fused feature map is input into the neck structure (SPPF+PANet) of YOLOv8m-seg, and multi-scale features are further fused to improve the recognition ability of small defects (width < 0.2mm).
[0072] 1. Model training optimization Transfer learning: Load the pre-trained weights of YOLOv8m-seg on the COCO dataset, freeze the first 4 to 6 layers of the backbone network, and train only the neck and head structures to shorten the training cycle; Loss function: A joint optimization is adopted using "CIoU loss (weight 0.2~0.4, to optimize localization) + Focal Loss (weight 0.3~0.5, to resolve class imbalance) + Dice Loss (weight 0.2~0.4, to optimize segmentation edges)". Hyperparameter settings: batch size 10~20, initial learning rate 0.0008~0.0012 (cosine annealing decay), total training epochs 40~60; Early stopping mechanism: When the validation set mAP@0.5 shows no improvement for 8~12 consecutive epochs, training is stopped and the "recent weights" and "best weights" are saved.
[0073] 1. Detection, Reasoning, and Post-processing Forward inference: Input the preprocessed test set images into the trained model and output "defect location coordinates (x,y,w,h), category (confidence ≥0.5 is valid), segmentation mask (edge precision ≤0.5mm)"; Non-maximum suppression (NMS): Set the IOU threshold to 0.5 to remove overlapping detection boxes and ensure that only one optimal result is output for a single defect.
[0074] Phase IV: Integration of Intelligent Imaging Devices and Data Interaction 1. Device hardware integration Dual-modal acquisition unit: infrared thermal imager (temperature measurement -20℃~300℃, resolution ≥320×240) + industrial visible light camera (resolution ≥2592×1944), synchronous triggering accuracy ≤1~3ms; Gimbal control unit: Supports 360° rotation, positioning accuracy ≤ ±0.2°, and can achieve automatic route planning through preset paths (compatible with wall-climbing robots / drones). Data processing unit: computing power ≥180~220TOPS, supports real-time model inference (detection time of a single 640×640 pixel image ≤0.3~0.5s), and integrates a "data storage of ≥2~4 months" module; Power unit: Lithium battery provides 6-10 hours of battery life and supports 1-3 hours of fast charging; Display interaction unit: 8~12-inch touch screen, supporting real-time preview of dual-modal images (simultaneously displaying infrared temperature difference markings, visible light defect boxes / masks / categories).
[0075] 1. Data interaction and output Report Export: Automatically generates defect reports in Excel / PDF format, including the correspondence between "detection area - defect coordinates - defect size - temperature difference value - confidence level"; Historical backtracking: Load detection data from different times in the same area, compare defect images, and analyze defect propagation trends (such as the rate of change of crack length / width).
[0076] V. Fifth Stage: Evaluation and Application of Test Results 1. Performance Evaluation: The model performance is verified through the test set. The core indicators include "accuracy ≥ 95%, recall ≥ 92%, mAP@0.5 ≥ 94%, IoU ≥ 0.8" (these are not performance values that are not required but are for reference only for engineering applications and are recorded in the specification). 2. Engineering Applications: For scenarios such as super high-rise buildings (outputting a "floor-defect coordinates" table) and bridges (outputting "beam segment-defect size" statistics), quantitative inspection reports are provided to provide a basis for subsequent maintenance and repair.
[0077] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit and essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.
Claims
1. A method for detecting defects in a building, characterized in that, Includes the following steps: S1. Construct a dual-modal dataset with temperature difference annotations for multiple materials: Collect image pairs consisting of infrared thermal imaging images and visible light images, representing "normal area - defect area", and annotate the defect type, defect size, material type, normal area temperature difference range, and defect area temperature difference value. After data augmentation, the image pairs are divided into training set, validation set, and test set. S2. Dual-modal image alignment: First, the ORB feature matching algorithm is used to coarsely align the infrared thermal image and the visible light image. After removing mismatched points, an initial transformation matrix is obtained. Then, the feature alignment module (FAM) performs deformable convolution adjustment on the infrared feature map to achieve fine correction, so that the alignment error of the dual-modal image is ≤ a preset threshold (this threshold ensures that the subsequent defect location accuracy meets the engineering inspection requirements). S3. Dual-modal fusion detection: The aligned 3-channel visible light image and 1-channel infrared temperature difference image are stitched together to form a multi-channel input, which is then input to the improved YOLOv8m-seg network. The network includes a dual-modal feature branch and a cross-attention fusion layer. After training with a joint loss function, the network outputs the location coordinates, category, and segmentation mask of the defect. S4. Post-processing and output of results: Perform non-maximum suppression processing on the detection results output by the network to remove overlapping detection boxes and output a quantified defect detection report.
2. The method for detecting building defects according to claim 1, characterized in that, The specific standards for temperature difference marking of multiple materials mentioned in step S1 are as follows: Under a 25℃ environment, the temperature difference of the normal area of the concrete exterior wall is ±0.3~±0.7℃, and the temperature difference of the defective area is ≥0.8~1.2℃; the temperature difference of the normal area of the brick exterior wall is ±0.6~±1.0℃, and the temperature difference of the defective area is ≥1.0~1.4℃; the temperature difference of the normal area of the steel exterior wall is ±0.2~±0.4℃, and the temperature difference of the defective area is ≥0.6~1.0℃; the temperature difference of the normal area of the glass curtain wall is ±0.8~±1.2℃, and the temperature difference of the defective area is ≥1.8~2.2℃.
3. The method for detecting building defects according to claim 1, characterized in that, The data augmentation described in step S1 includes the following operations: Basic transformations: Simultaneously perform random rotation of -15° to 15°, random scaling of 0.8 to 1.2x, and horizontal or vertical flipping on the bimodal image; Dirt simulation: Dust, water stains, or oil stains are embedded in visible light images to simulate real-world pollution scenarios. The dust coverage rate is 5%~30%, the water stain coverage rate is 10%~20%, and the oil stain coverage rate is 5%~15%. The temperature difference noise in the corresponding infrared regions is adjusted simultaneously. The temperature difference noise for dust is increased by +0.1~+0.3℃, for water stains by +0.2~+0.4℃, and for oil stains by +0.3~+0.5℃. Color gamut adjustment: Perform random offsets of ±8%~±12% for brightness and ±12%~±18% for contrast in visible light images, and perform random offsets of ±3%~±7% for grayscale values in infrared images.
4. The method for detecting building defects according to claim 1, characterized in that, The specific implementation process of the ORB feature matching algorithm in step S2 is as follows: extract ≥40 stable feature points from the infrared thermal imaging image and the visible light image respectively (preferably select feature points in the building outline, corners and door and window frame areas), perform feature point matching using the brute-force matching algorithm, and then remove mismatched points with errors >2~4 pixels using the RANSAC algorithm to obtain the coarse alignment transformation matrix; The Feature Alignment Module (FAM) dynamically adjusts the spatial position of infrared feature points through deformable convolution to compensate for residual errors in coarse alignment, so that the alignment error of the finely corrected dual-modal image is ≤1~3 pixels.
5. The method for detecting building defects according to claim 1, characterized in that, The specific structure of the improved YOLOv8m-seg network described in step S3 is as follows: Dual-modal feature branches: The visible light branch adopts a "Conv-BN-ReLU" structure, which extracts material texture and defect edge features through a 3×3 convolution kernel, and the number of channels is gradually increased from 4 to 64; the infrared branch adopts a "Conv-BN-Sigmoid" structure, which enhances the temperature difference features between the defect area and the normal area through a 1×1 convolution kernel, and the number of channels is synchronously matched with the visible light branch. Cross-attention fusion layer: After concatenating the bimodal feature maps by channel dimension, feature weights are dynamically assigned—the infrared feature weight in the defect region is 0.5~0.7, and the visible light feature weight is 0.3~0.5, while the infrared feature weight in the non-defect region is 0.3~0.5, and the visible light feature weight is 0.5~0.
7. Joint loss function: CIoU loss, Focal Loss and Dice Loss are combined with a weight ratio of 0.2~0.4:0.3~0.5:0.2~0.
4. CIoU loss is used to optimize the defect localization accuracy, Focal Loss is used to solve the class imbalance problem between defect and non-defect regions, and Dice Loss is used to optimize the edge fitting accuracy of defect segmentation mask.
6. The method for detecting building defects according to claim 1, characterized in that, The training process of the improved YOLOv8m-seg network described in step S3 includes: Transfer learning: Load the pre-trained weights of YOLOv8m-seg on the COCO public dataset, freeze the first 4 to 6 layers of the backbone network, and train only the neck and head structures to shorten the training cycle. Hyperparameter settings: training batch size is 10~20, initial learning rate is 0.0008~0.0012, cosine annealing strategy is used for learning rate decay, and the total number of training epochs is 40~60. Early stopping mechanism: When the mean accuracy (mAP@0.5) of the validation set does not improve for 8 to 12 consecutive epochs, training is stopped, and two types of model files are saved: "recent weight file" and "best weight file". The best weight file corresponds to the training state with the highest mAP@0.5 on the validation set.
7. A smart imaging device for detecting building defects, implementing the method of any one of claims 1-6, characterized in that, It includes a dual-mode acquisition unit, a pan-tilt control unit, a data processing unit, and a power supply unit that are interconnected via a data bus or control line and powered by a power line. The dual-modal acquisition unit is used to simultaneously acquire infrared thermal imaging images and visible light images, and supports synchronous trigger control to ensure spatiotemporal alignment. The gimbal control unit is used to drive the dual-modal acquisition unit to achieve 360° rotation, with a positioning accuracy of ≤±0.2°, and supports automatic route planning based on preset paths; The data processing unit is used to execute the dual-modal image alignment algorithm and the dual-modal fusion detection algorithm, and output the defect detection results; The power supply unit is used to supply power to each unit, supporting ≥6~10 hours of battery life and fast charging function.
8. The intelligent imaging device for building defect detection according to claim 7, characterized in that, The dual-modal acquisition unit includes an infrared thermal imager and an industrial visible light camera; the infrared thermal imager has a temperature measurement range of -20℃ to 300℃, a resolution of ≥320×240, and a noise equivalent temperature difference of ≤50mK; the industrial visible light camera has an image resolution of ≥2592×1944 and a frame rate of ≥25fps; the synchronization triggering accuracy of the dual-modal acquisition unit is ≤1~3ms, ensuring that the acquired infrared thermal image and visible light image correspond one-to-one in time and space.
9. The intelligent imaging device for building defect detection according to claim 7, characterized in that, The data processing unit has a computing power of ≥180~220 TOPS, supports real-time inference of the improved YOLOv8m-seg network, and has a single detection time of ≤0.3~0.5s for a 640×640 pixel bimodal image. The data processing unit also integrates an image storage module, which supports the storage of ≥2~4 months of detection data (including images, defect information, and environmental parameters), and can export defect reports in Excel or PDF format. The report contains the correspondence between "detection area - defect coordinates - defect size - temperature difference value - detection confidence level".
10. The intelligent imaging device for building defect detection according to claim 7, characterized in that, It also includes a display interaction unit, which is an 8-12 inch touch screen that supports real-time preview of dual-modal images—simultaneously displaying the temperature difference annotation information of the infrared image (including the temperature difference value between the normal area and the defect area), the defect detection box, segmentation mask and defect category information of the visible light image; the display interaction unit also supports the retrospective query of historical detection data, and can load defect images of the same detection area at different times for comparison to help analyze the defect expansion trend.
Citation Information
Patent Citations
Wall surface detection method and system, computer equipment and storage medium
CN120451154A
Cited By
Building facade defect detection method and system based on bimodal image
CN121259620A
A building facade defect detection method and system based on a dual-mode image
CN121259620B
Intelligent inspection data processing method and system
CN121304667A
Intelligent Inspection Data Processing Methods and Systems
CN121304667B
Building engineering wall flatness detection method and system based on AI visual identification
CN121576961A