Glass curtain wall intelligent detection system based on multi-modal data fusion

Through multimodal data fusion and intelligent detection systems, efficient and accurate detection of glass curtain walls has been achieved, solving the problems of low efficiency, high cost and low automation of traditional detection methods, and providing three-dimensional visualized detection reports.

CN121640283APending Publication Date: 2026-03-10ZHEJIANG COLLEGE OF SECURITY TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-04
Publication Date
2026-03-10

Smart Images

  • Figure CN121640283A_ABST
    Figure CN121640283A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of building glass curtain wall detection, in particular to a glass curtain wall intelligent detection system based on multi-modal data fusion, which comprises a visible light image acquisition module, an infrared image acquisition module, a three-dimensional model data module, a multi-modal data fusion processing module, an artificial intelligence recognition module and an output module, and the multi-modal data fusion processing module comprises a data registration unit and a feature fusion analysis unit. According to the method, accurate registration of the visible light image, the infrared image and the three-dimensional model data is realized through the multi-modal data fusion processing module, the problem of difficulty in data alignment in a traditional method is effectively solved, and the accuracy and comprehensiveness of defect detection are improved. The artificial intelligence recognition module is used for automatic defect recognition and classification, the automation degree is remarkably improved, manual intervention is reduced, and the safety risk is reduced. The system can detect surface defects and internal defects of the glass curtain wall at the same time, and the detection reliability and efficiency are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of building glass curtain wall inspection technology, and more particularly to an intelligent inspection system for glass curtain walls based on multimodal data fusion. Background Technology

[0002] With urban development, glass curtain wall buildings are increasingly common, and the safety hazards they pose after long-term use (such as glass breakage, sealant aging, and internal hollowing) have become a major public safety issue. Traditional inspection methods mainly rely on manual inspection from close range using suspended platforms or by tapping, which is inefficient, costly, dangerous, and heavily influenced by subjective factors. In recent years, technologies using drones for building exterior inspection have emerged, but most only use a single visible light camera for taking pictures, still requiring manual interpretation of massive amounts of images, resulting in low automation and the inability to detect internal defects (such as hollowing). Some research has attempted to introduce infrared thermal imagers, but how to accurately register visible light images with infrared images and use artificial intelligence algorithms for collaborative analysis and decision-making of multimodal data remains a technical challenge. Existing technologies lack a complete solution integrating automatic acquisition, multi-source data fusion, intelligent analysis, and precise positioning. Therefore, there is an urgent need in this field for a new glass curtain wall inspection technology that can achieve high efficiency, accuracy, automation, and the ability to detect both internal and external defects. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this application provides an intelligent inspection system for glass curtain walls based on multimodal data fusion.

[0004] To achieve the above objectives, the technical solution of this application is as follows: An intelligent inspection system for glass curtain walls based on multimodal data fusion, comprising: Visible light image acquisition module, used to acquire visible light images of the curtain wall surface; Infrared image acquisition module, used to acquire infrared images of the curtain wall surface; A 3D model data acquisition module, which pre-stores 3D model data based on buildings and / or is used to acquire 3D model data based on buildings in real time; A multimodal data fusion processing module is used to accurately register the acquired multi-source data, which includes visible light images acquired by the visible light image acquisition module, infrared images acquired by the infrared image acquisition module, and 3D model data stored and / or acquired in real time by the 3D model data module. It includes a data registration unit and a feature fusion analysis unit. The data registration unit performs pixel-level spatial registration of the visible light images, infrared images, and 3D model data stored and / or acquired in the 3D model data module at the same time. The feature fusion analysis unit dynamically assigns different weights to different types of defects based on a pre-built STDFusionNet neural network model. The artificial intelligence recognition module is used for defect identification and classification. The output module is used to output the detection results.

[0005] Furthermore, it includes a drone inspection module, which includes a drone equipped with a visible light image acquisition module, an infrared image acquisition module, and a three-dimensional model data acquisition module, for inspecting the exterior of buildings. The visible light image acquisition module, infrared image acquisition module, and three-dimensional model data acquisition module simultaneously acquire visible light images, infrared images, and three-dimensional model data of the building's glass curtain wall.

[0006] Furthermore, the drone inspection module includes an RTK positioning unit for achieving centimeter-level precise positioning of the drone.

[0007] Furthermore, it includes an intelligent route planning module, which is used to automatically plan a full-coverage inspection route and real-time obstacle avoidance based on the 3D model data of the stored buildings or the 3D model data acquired by the 3D model data acquisition module.

[0008] Furthermore, the data registration unit processes the acquired data using the Shearlet-SURF algorithm.

[0009] Furthermore, the Shearlet-SURF algorithm is used to process the collected data, including the following steps: S211: Perform Shearlet transform on visible light and infrared images respectively, and filter high-frequency subbands; S212: Using the high-frequency sub-bands selected in S21 as input, execute the SURF feature point detection process to extract SURF feature points; S213: The bidirectional nearest neighbor distance ratio algorithm is used for initial feature point matching; then, a mismatch elimination algorithm based on the consistency of the spatial distribution of feature points is applied. Using the idea of ​​the RANSAC algorithm, the geometric relationship of feature points in their Shearlet domain is used as a consistency constraint to further purify the matching pairs. S214: Based on the purified matching pairs, calculate the affine transformation matrix between the visible light image and the infrared image to achieve pixel-level fine registration; then map the registered infrared image data and the acquired 3D model data to the same visual coordinate system through coordinate system one.

[0010] Furthermore, the STDFusionNet neural network model includes: Two independent feature extraction paths are used to process the registered visible light image and infrared image respectively. Each feature extraction path is built based on a lightweight GhostNet. The dynamic weight allocation module takes as input feature maps of the same level on two independent feature extraction paths, and is used to learn and generate spatial weight maps and channel weight vectors in real time. The feature fusion and recognition module fuses the dual-path features after they have been weighted by the DWAM module by adding them element by element.

[0011] Furthermore, the artificial intelligence recognition module performs defect recognition and classification based on the Ghost convolution and bottleneck structure optimized YOLOv5 model. It replaces the standard Bottleneck (C3) module in the YOLOv5 model with the Ghost Bottleneck module, and in the detection head part, some standard convolutions are replaced with depthwise separable convolutions.

[0012] Furthermore, the loss function of the artificial intelligence recognition module introduces Focal Loss as a classification loss function based on the original CIoU Loss.

[0013] Furthermore, the output module associates the detected defects with the stored 3D model data of the building, performs visual annotation on the 3D model, and automatically generates a standardized inspection report containing defect information, location map, and maintenance suggestions.

[0014] The beneficial effects of this application are as follows: By using a multimodal data fusion processing module to achieve precise registration of visible light images, infrared images, and 3D model data, the system effectively solves the problem of data alignment difficulties in traditional methods, improving the accuracy and comprehensiveness of defect detection. The use of an artificial intelligence recognition module for automatic defect identification and classification significantly enhances the degree of automation, reduces manual intervention, and lowers safety risks. The system can simultaneously detect surface defects (such as breakage) and internal defects (such as hollow areas) in glass curtain walls, enhancing the reliability and efficiency of detection. The 3D visualization function facilitates defect location and maintenance planning, optimizing the detection process. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, obtaining other drawings based on these drawings without creative effort still falls within the scope of the present invention.

[0016] Figure 1 This is a schematic diagram of the intelligent inspection system for glass curtain walls provided in an embodiment of this application; Figure 2 A flowchart of the intelligent detection method for glass curtain walls provided in the embodiments of this application; Figure 3 This is a flowchart of the multimodal data fusion processing module provided in an embodiment of this application; Figure 4 A flowchart of the data registration unit processing provided in the embodiments of this application; Figure 5 A flowchart of the feature fusion analysis unit provided in the embodiments of this application; Figure 6 This is a visualization of the inspection results on a digital twin model (BIM). Detailed Implementation

[0017] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this application.

[0018] With urban development, glass curtain wall buildings are increasingly common, and the safety hazards they pose after long-term use (such as glass breakage, sealant aging, and internal hollowing) have become a major public safety issue. Traditional inspection methods mainly rely on manual inspection from close range using suspended platforms or by tapping, which is inefficient, costly, dangerous, and heavily influenced by subjective factors. In recent years, technologies using drones for building exterior inspection have emerged, but most only use a single visible light camera for taking pictures, still requiring manual interpretation of massive amounts of images, resulting in low automation and the inability to detect internal defects (such as hollowing). Some research has attempted to introduce infrared thermal imagers, but how to accurately register visible light images with infrared images and use artificial intelligence algorithms for collaborative analysis and decision-making of multimodal data remains a technical challenge. Existing technologies lack a complete solution integrating automatic acquisition, multi-source data fusion, intelligent analysis, and precise positioning. Therefore, there is an urgent need in this field for a new glass curtain wall inspection technology that can achieve high efficiency, accuracy, automation, and the ability to detect both internal and external defects.

[0019] Based on this, to improve upon the problems in related technologies, embodiments of this application provide an intelligent inspection system and method for glass curtain walls that features high detection efficiency, high accuracy, strong security, and the ability to simultaneously detect surface and internal defects. Specifically, as... Figure 1 As shown, the system includes, Visible light image acquisition module, used to acquire visible light images of the curtain wall surface; Infrared image acquisition module, used to acquire infrared images of the curtain wall surface; A 3D model data module, which pre-stores building-based 3D model data and / or is used for real-time acquisition of building-based 3D model data; A multimodal data fusion processing module is used to accurately register the acquired multi-source data, which includes visible light images acquired by the visible light image acquisition module, infrared images acquired by the infrared image acquisition module, and 3D model data stored and / or acquired in real time by the 3D model data module. It includes a data registration unit and a feature fusion analysis unit. The data registration unit performs pixel-level spatial registration of the visible light images, infrared images, and 3D model data stored and / or acquired in the same time with the 3D model data module. The feature fusion analysis unit, based on a pre-built STDFusionNet neural network model, dynamically assigns different weights to different types of defects. The artificial intelligence recognition module is used for defect identification and classification. The output module is used to output the detection results.

[0020] The method is based on the above system implementation and specifically includes the following steps: S1: The visible light image acquisition module and the infrared image acquisition module simultaneously acquire visible light images and infrared images of the glass curtain wall, and the three-dimensional model data module pre-stores three-dimensional model data based on the building and / or acquires three-dimensional model data based on the building at the same time. S2: The multi-modal data fusion processing module performs precise registration of the collected multi-source data, including: S21: Through the data registration unit, pixel-level spatial registration is performed on the visible light image, infrared image and 3D model data pre-stored in the 3D model data module and / or the 3D model data acquired at the same time. S22: Through the feature fusion analysis unit, based on the pre-built STDFusionNet neural network model, different data weights are dynamically assigned to different types of defects; S3: Input the registered data into the artificial intelligence recognition module to perform defect identification and classification; S4: Output module outputs test report.

[0021] Specifically, the visible light image acquisition module typically employs a high-resolution digital camera to capture visual details of the curtain wall surface, such as cracks or stains; the infrared image acquisition module utilizes an infrared thermal imager to detect temperature distribution, thereby identifying defects such as internal voids; the 3D model data module is either a storage module that pre-stores 3D model data based on the building, or a 3D model data acquisition module that acquires 3D model data based on the building in real time, or a combination of a storage module that pre-stores 3D model data based on the building and a 3D model data acquisition module. The 3D model data acquisition module can acquire point cloud data of the building using LiDAR or photogrammetry technology to construct a 3D geometric model. By fusing visible light, infrared, and 3D point cloud data, a comprehensive, integrated detection of glass curtain walls—from surface defects (cracks) to internal hazards (voids) to structural deformation—is achieved, overcoming the limitations of single-sensor detection.

[0022] The multimodal data fusion processing module operates as follows: First, spatial alignment is performed by the data registration unit to ensure that data from different sources are in the same coordinate system, resolving data misalignment issues caused by different sensor positions and viewing angles. Second, the feature fusion analysis unit dynamically assigns weights to different types of sensor data for different types of defects by dynamically weighting and fusing features using the STDFusionNet model. For example, visible light data is given higher weight for surface cracks, while infrared thermal imaging data is given higher weight for internal voids.

[0023] This application's embodiments, through precise registration and intelligent fusion, enable the system to automatically and efficiently detect internal and external defects, reducing reliance on manual labor, improving detection accuracy, and avoiding the subjectivity and risks of traditional methods. Specifically, pixel-level registration ensures the consistency of multi-source data; the dynamic weight allocation of feature fusion units adapts to different defect types, improving identification accuracy; and the entire process achieves end-to-end automation, significantly shortening the detection cycle from data acquisition to result output.

[0024] Optionally, in some embodiments, a drone inspection module is included. This module comprises a drone equipped with a visible light image acquisition module, an infrared image acquisition module, and a 3D model data acquisition module. It is used for inspecting the building's exterior, simultaneously acquiring visible light images, infrared images, and 3D model data of the building's glass curtain wall through these modules. Specifically, the drone flies along a preset path, simultaneously triggering each sensor to acquire data, achieving time synchronization and spatial consistency. The drone enables fully automated data acquisition, reducing the risks and costs of manual climbing. Simultaneous acquisition avoids data misalignment issues, improving the reliability of subsequent analysis. Furthermore, the drone platform can cover high-altitude areas, reducing blind spots and enhancing the overall system's applicability.

[0025] Optionally, in some embodiments, the UAV inspection module includes an RTK positioning unit for achieving centimeter-level precise positioning of the UAV. The RTK positioning unit typically consists of a base station and a rover station. The base station provides correction signals, and the rover station receives and processes GPS data to achieve high-precision positioning. During UAV flight, the RTK unit receives satellite signals and correction data in real time to calculate the precise position. As an alternative embodiment, those skilled in the art can also use other high-precision positioning technologies such as PPK (post-processing dynamic positioning) or visual SLAM systems, but it is necessary to ensure that the positioning accuracy meets data alignment requirements. Specifically, in some embodiments, the RTK unit is integrated into the UAV control system, including an antenna, receiver, and processing chip, eliminating positioning errors through real-time differential calculation; in other embodiments, visual-assisted positioning based on ground control points is used, i.e., marker points are set around buildings, and the UAV identifies these points through a camera and fuses them with RTK data to achieve redundant calibration.

[0026] Optionally, in some embodiments, an intelligent flight path planning module is included to automatically plan a full-coverage inspection route and perform real-time obstacle avoidance based on stored 3D model data of the building or a 3D model generated from 3D model data acquired by a 3D model data acquisition module. The intelligent flight path planning module analyzes obstacles and surface features in the 3D model to generate the optimal flight path. After inputting the 3D model, the module calculates the minimum path covering all curtain wall areas and, during flight, detects obstacles in real time through the 3D model data acquisition module and adjusts the flight path accordingly.

[0027] Optionally, in some embodiments, the data registration unit processes the acquired data using the Shearlet-SURF algorithm. Specifically, the Shearlet-SURF algorithm is a feature point descriptor fusion registration method based on Shearlet transform and improved SURF. The Shearlet-SURF algorithm combines multi-scale analysis of Shearlet transform with SURF feature detection for image registration; it transforms visible light and infrared images, extracts feature points, and then matches them. The Shearlet-SURF algorithm improves the accuracy and efficiency of image registration, is particularly suitable for multimodal data, and reduces mismatches.

[0028] The Shearlet-SURF algorithm is used to process the collected data, which includes the following steps: S211: Perform Shearlet transform on both the visible light and infrared images, filtering for high-frequency subbands across multiple scales and directions. Specifically, the Shearlet transform decomposition process generates one low-frequency subband and several sets of high-frequency subbands. The low-frequency subband corresponds to the overall grayscale contour of the image and is easily affected by the difference in illumination / thermal radiation between the visible light and infrared images; while the high-frequency subbands correspond to singular features such as edges, textures, and abrupt changes in the image, precisely the concentrated areas of key targets such as cracks and object boundaries. After decomposing the visible light and infrared images separately, all effective high-frequency subbands need to be selectively filtered (using energy thresholding to remove weak-energy high-frequency subbands dominated by noise) to avoid invalid information interfering with subsequent feature extraction. The multi-directional characteristics of the Shearlet transform provide optimal representation capabilities for singular features such as edges and textures, effectively capturing linear features such as glass cracks and boundaries. This step provides "purified" input for subsequent feature point extraction, allowing feature points to focus on consistent target details in both types of images, rather than redundant information caused by modal differences.

[0029] S212: Using the high-frequency sub-band selected in S21 as input, the SURF (Speeded-Up RobustFeatures) feature point detection process is executed to extract SURF feature points. Specifically, the pixel response value is quickly calculated using the Hessian matrix of SURF, local extrema are selected as candidate feature points, and then non-maximum suppression and sub-pixel localization are used to remove blurred or unstable points, thereby obtaining stable feature point coordinates. In this process, the high-frequency sub-band has filtered out most of the modal difference noise, and the repeatability of feature points is significantly improved.

[0030] First, the neighborhood is rotated along its principal direction around each feature point (ensuring rotation invariance). Next, the gradient magnitude and direction of pixels within the neighborhood are calculated, and the neighborhood is divided into cells and blocks. The gradient direction histogram (HOG) information of each block is calculated. Finally, a weighted fusion strategy is adopted (since HOG is more robust to illumination, it can be assigned a weight of 0.6-0.7) to combine the HOG histogram with the Haar wavelet response features of traditional SURF to generate a hybrid descriptor.

[0031] S213: The Bi-directional Nearest Neighbor Distance Ratio (BNNDR) algorithm is used for initial feature point matching; then, a mismatch elimination algorithm based on the consistency of the spatial distribution of feature points is applied. Using the idea of ​​RANSAC (Random Sampling Consensus) algorithm, the geometric relationship of feature points in their Shearlet domain is used as a consistency constraint to further purify the matching pairs.

[0032] The specific process of using the bidirectional nearest neighbor distance ratio algorithm for initial feature point matching is as follows: For feature point A in the visible light image, find the nearest neighbor B and the second nearest neighbor C in the infrared image; simultaneously, for point B in the infrared image, find the nearest neighbor A' and the second nearest neighbor C' in the visible light image. Only when A and B are mutually nearest neighbors, and the distance ratios (A-B distance / A-C distance, B-A distance / B-C' distance) are both less than a set threshold, is the matching pair retained, significantly reducing one-way false matching.

[0033] The specific process for further refining the matching is as follows: During screening, a small number of matching pairs are randomly selected to construct an initial affine model. Then, the scale consistency and orientation correspondence of other matching pairs in the Shearlet domain are verified to ensure they conform to the model. Those that conform are considered interior points, and those that do not are considered exterior points. After multiple iterations, the matching pair corresponding to the model with the most interior points is selected as the refined result.

[0034] S214: Based on the purified matching pairs, calculate the affine transformation matrix between the visible light image and the infrared image to achieve pixel-level fine registration; then map the registered infrared image data and the acquired 3D model data to the same visual coordinate system through coordinate system one.

[0035] The process of solving the affine transformation matrix is ​​as follows: The purified matching pair contains the pixel coordinates of corresponding points in the visible light and infrared images. Since the deformations of the two images are mostly combinations of translation, rotation, and scaling, the affine transformation matrix (containing 6 parameters describing linear transformations and translations) can be solved using the least squares method. After solving, all pixels of the infrared image are mapped to the pixel coordinate system of the visible light image through this matrix, completing the pixel-level fine registration between the two.

[0036] The coordinate system process is as follows: Infrared and visible light data are two-dimensional image data, while LiDAR data is three-dimensional point cloud data. The coordinate systems of these three types differ (infrared / visible light uses an image coordinate system, while LiDAR uses a three-dimensional world coordinate system). Transformation parameters between these coordinate systems need to be obtained through sensor calibration: First, the registered infrared-visible light image is converted from the image coordinate system to the camera coordinate system. Then, combining the extrinsic parameters (rotation matrix, translation vector) of the LiDAR and camera, the LiDAR point cloud is also converted to the camera coordinate system, ultimately achieving the mapping and fusion of the three types of data under the same visual coordinate system.

[0037] The improved Shearlet-SURF registration algorithm, compared with the traditional SIFT / SURF, improves the feature point matching accuracy by about 35% in the visible light-infrared image registration task, and controls the registration error within 2 pixels, laying a precise spatial foundation for subsequent fusion.

[0038] Optionally, in some embodiments, during model training, the Dynamic Weight Assignment Module (DWAM) of the STDFusionNet neural network model learns the importance of different modal data through backpropagation. For example, when the training samples are image pairs containing "internal glass voids," the loss function guides the network to assign higher channel weights to infrared image features because infrared thermal imaging is sensitive to temperature differences and can more effectively characterize void defects; while for "surface scratch" samples, the network assigns higher weights to visible light image features. Traditional multimodal fusion is mostly early (data-level) or late (decision-level) fusion, failing to fully utilize the complementarity between modalities. This invention constructs a Spatio-Temporal Dynamic Fusion Net (STDFusionNet) based on an attention mechanism to achieve optimal fusion at the feature level. Specifically, the STDFusionNet neural network model includes: Two independent feature extraction paths (Backbone) are used to process the registered visible light image and infrared image respectively. Each feature extraction path is built based on a lightweight GhostNet. The Dynamic Weight Allocation Module (DWAM) takes as input feature maps of the same level on two independent feature extraction paths and learns in real time to generate spatial weight maps and channel weight vectors. The feature fusion and recognition module fuses the dual-path features after they have been weighted by the DWAM module by adding them element-wise.

[0039] GhostNet, as a lightweight network, boasts a core advantage in generating redundant feature maps (Ghost features) through "ordinary convolution + inexpensive linear transformation," significantly reducing the number of parameters and computational cost while maintaining feature expressive power. Two independent paths process the registered visible light and infrared images separately, outputting multi-scale feature maps (such as Conv2, Conv3, Conv4, etc., corresponding to texture, contour, and semantic features from shallow to deep), providing a foundation for subsequent feature pyramid fusion. Simultaneously, the independent path design avoids mutual interference between the two modalities during the extraction stage, preserving the unique characteristics of each modality.

[0040] DWAM takes feature maps of the same level along two paths (such as visible light Conv3 feature map Fv and infrared Conv3 feature map Fi) as input and is itself a lightweight quantum network (which can be composed of 1-2 layers of convolution, batch normalization and activation functions to avoid increasing too much computational cost).

[0041] Spatial weight maps are used to emphasize key regions in the image related to defects (such as the edges of cracks) and suppress irrelevant background. Specifically, for Fv and Fi, DWAM compresses the channel dimension through convolution operations, and then outputs a spatial weight map Ws with the same size as the feature map after passing through a sigmoid activation function. Each pixel value in Ws is between 0 and 1. Key defect regions (such as crack edges) have strong feature responses and corresponding weight values ​​close to 1, and are thus emphasized; background regions have weak feature responses and corresponding weight values ​​close to 0, and are effectively suppressed. This design allows the network to automatically focus on effective regions and reduce the interference of background redundancy on recognition.

[0042] The channel weight vector is used to evaluate the importance of different modalities (visible light vs. infrared) at the current feature level. For example, when identifying surface cracks, the network automatically assigns higher weights to visible light features; while when identifying internal cavities, it assigns higher weights to infrared features. Specifically, DWAM compresses the channel features of Fv and Fi into a one-dimensional vector through global average pooling, and then learns the importance between modalities through fully connected layers, outputting a channel weight vector Wc (with the same dimension as the number of channels in the current feature map). For example, when identifying surface cracks, the texture details of visible light features are more critical, and Wc will assign higher weights to each channel of the visible light features; when identifying internal cavities, the thermal conductivity differences of infrared features are more crucial, and Wc will be biased towards infrared features.

[0043] After DWAM weighting, the visible light feature map Fv' = Fv × Ws × Wc_v (Wc_v is the visible light channel weight), and the infrared feature map Fi' = Fi × Ws × Wc_i (Wc_i is the infrared channel weight). The feature fusion and recognition module adopts element-wise addition fusion. Compared with stitching, this method can reduce the feature dimension and allow complementary features of the two modalities to enhance each other. For example, the crack texture of visible light and the crack thermal trace features of infrared light are superimposed to form a more recognizable fused feature. The fused multi-scale features are integrated by the Feature Pyramid (FPN) to achieve full coverage of defects of different scales (such as large-sized hollows and small-sized microcracks).

[0044] Specifically, the processing procedure of the STDFusionNet neural network model is as follows: S221: Input the registered visible light image and infrared image, and send them to two feature extraction paths based on GhostNet respectively; S222: Two paths output multi-scale feature maps to construct a dual-modal feature pyramid; S223: For each level of the feature pyramid, input the DWAM module to generate the spatial weight map and channel weight vector for the corresponding level; S224: Weight the bimodal feature map and then perform feature fusion at this level by adding elements one by one.

[0045] The AI ​​recognition module analyzes and fuses features based on the optimized algorithm, outputting defect classification results. Optionally, in some embodiments, to enable real-time inference on the drone, this invention performs a deep lightweight modification to the neck and head of YOLOv5s. Using an improved YOLOv5 deep learning model as a foundation, a Ghost convolutional module is introduced to replace some standard convolutional layers, significantly reducing the number of model parameters and computational load, enabling deployment on edge computing devices such as NVIDIA Jetson to achieve real-time analysis with a detection frame rate ≥15fps.

[0046] Specifically, the standard Bottleneck (C3) module in the YOLOv5 model is replaced with the GhostBottleneck module to achieve neck optimization. The standard C3 module contains multiple standard Bottlenecks, resulting in high computational overhead from 3×3 convolutions, high feature map redundancy, and low computational resource utilization. This embodiment modifies the C3 module definition in models / yolo.py, replacing Bottleneck with GhostBottleneck. It generates "phantom feature maps" using "a few standard convolutions + a large number of inexpensive linear transformations," which are then concatenated with the feature maps generated by a small number of convolutions. This approach preserves rich feature information with extremely low computational cost.

[0047] Specifically, in the detection head portion, this embodiment replaces some standard convolutions with depthwise separable convolutions. Depthwise separable convolutions split the standard convolution into depthwise convolutions and pointwise convolutions. The depthwise convolution operates independently on each channel, capturing spatial features, while the pointwise convolution handles channel fusion. This further reduces the number of parameters and computational cost.

[0048] The YOLOv5 model, after being deeply lightweighted, reduces the number of parameters by 43% while keeping the average accuracy (mAP) down by no more than 1%, and improves the inference speed on edge computing devices to 18.6 FPS, truly meeting the real-time requirements of engineering sites.

[0049] Optionally, in some embodiments, to address the problem of extreme imbalance between positive and negative samples in glass curtain wall defects (the defect area is much smaller than the background), Focal Loss is introduced as a classification loss function based on the original CIoU Loss (bounding box regression), which effectively improves the model's attention to and detection ability of difficult samples (small defects such as cracks and bubbles have a small pixel ratio, and standard cross-entropy is difficult to effectively capture).

[0050] Focal Loss reduces the weight of easily classified samples (high-confidence samples) by modulating the factor (1-p_t)^γ, forcing the model to focus on difficult samples, and introduces a balancing factor α. t To alleviate the imbalance between the number of positive and negative samples.

[0051] The classification loss function is FL(p) t )=−α t (1−p t ) γ log(p t ); where p t The predicted probability for a sample is close to 1 when the sample is correctly classified and close to 0 when it is incorrect; α t γ controls the balance between positive and negative samples; γ controls the degree of focus on difficult samples.

[0052] Optionally, in some embodiments, the output module associates the detected defects with stored 3D model data of the building, performs visual annotation on the 3D model, and automatically generates a standardized inspection report containing defect information, location map, and maintenance suggestions. In some embodiments of this application, the 3D model data is a BIM model, which integrates multi-disciplinary data to construct a digital twin spanning the entire lifecycle. Figure 6 The image shown is a visualization of the detection results obtained after detection in an embodiment of this application on a digital twin model (BIM).

[0053] This application's embodiments closely address the industry pain points in glass curtain wall inspection, namely "small defect size, diverse shapes, complex backgrounds, and the need for simultaneous detection of internal and external defects." It systematically integrates the hardware combination of "drone + RTK + multiple sensors (visible light, infrared, lidar)" with the software algorithm of "multimodal data registration and fusion + lightweight AI recognition + digital twin visualization," providing a complete algorithm solution from data alignment to intelligent decision-making, demonstrating significant technological advancement.

[0054] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0055] This application also provides an electronic device, which includes a processor and a memory, and the various devices can be connected through one or more buses.

[0056] The electronic device also includes a computer program stored in a memory. When the computer program is executed by a processor, it causes the electronic device to perform the above-described method. All relevant details regarding the steps in the above method embodiments can be found in the functional descriptions of the corresponding physical devices, and will not be repeated here.

[0057] This application also provides a computer-readable storage medium including a computer program that, when run on a computer, causes the computer to perform the methods provided in the above-described method embodiments.

[0058] This application also provides a computer program product, including: computer program code, which, when run on an electronic device, causes the electronic device to execute the method provided in the above-described method embodiments.

[0059] It should be understood that in the embodiments of this application, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0060] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0061] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0062] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0063] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0064] It should be understood that the term "embodiment" used throughout this specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, various embodiments throughout this specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0065] Those skilled in the art will understand that the various numerical designations such as "first," "second," etc., involved in this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application, nor do they indicate the order of sequence.

[0066] In this application, the use of singular pronouns to denote "one or more" rather than "one and only one," unless otherwise specified. In this application, unless otherwise specified, "at least one" is intended to mean "one or more," and "more than" is intended to mean "two or more."

[0067] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0068] The following is a specific embodiment of this application.

[0069] Take the inspection of the glass curtain wall of a 5,000-square-meter office building as an example.

[0070] Preliminary preparation: Obtain the BIM model of the building and import it into the system. The system automatically generates a full-coverage inspection route with centimeter-level accuracy.

[0071] Data Acquisition: Control the DJI Matrice 350 RTK drone equipped with RTK, visible light camera, infrared thermal imager and lidar to automatically fly along the flight path and collect data simultaneously.

[0072] Data processing and recognition: Registration: The system uses the Shearlet-SURF algorithm to process the acquired data. For example, it precisely aligns thermal anomaly regions in the infrared image with the specific locations of glass plates in the visible light image.

[0073] Identification: The registered data blocks are fed into a pre-trained lightweight YOLOv5 model (STDFusionNet). Based on the fused features, the model identifies 35 defects, including 25 surface cracks, 7 sealant failures, and 3 internal voids.

[0074] Results Display and Report Generation: All identified defects are automatically marked on the building's BIM model. Clicking on a defect point allows you to view details (e.g., the fifth pane of glass on the east side of the 3rd floor has a crack approximately 10cm long). The system automatically generates a PDF report within 10 minutes, including a defect statistics table, a distribution heatmap, and repair recommendations.

[0075] Results Verification: On-site verification showed that the system's detection accuracy reached 98.2%, with a false negative rate of less than 2.9%, which is significantly better than traditional manual detection.

[0076] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the claims. In conclusion, the above description is merely a preferred embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A glass curtain wall intelligent detection system based on multi-modal data fusion, characterized in that, The system comprises: a visible light image acquisition module for acquiring visible light images of the curtain wall surface; an infrared image acquisition module for acquiring infrared images of the curtain wall surface; a three-dimensional model data module pre-stored with three-dimensional model data based on the building and / or for real-time acquisition of three-dimensional model data based on the building; a multi-modal data fusion processing module for accurate registration of multi-source data, including the visible light images acquired by the visible light image acquisition module, the infrared images acquired by the infrared image acquisition module, and the three-dimensional model data stored and / or real-time acquired by the three-dimensional model data module; the module comprises a data registration unit and a feature fusion analysis unit; the data registration unit is configured to perform pixel-level spatial registration on the visible light images, the infrared images, and the three-dimensional model data stored by the three-dimensional model data module and / or the three-dimensional model data acquired at the same time; the feature fusion analysis unit is configured to dynamically assign different data weights to different types of defects based on a pre-constructed STDFusionNet neural network model; an artificial intelligence recognition module for defect recognition and classification; an output module for outputting the detection results.

2. The intelligent glass curtain wall detection system based on multi-modal data fusion according to claim 1, characterized in that the system comprises a UAV inspection module, which comprises a UAV carrying the visible light image acquisition module, the infrared image acquisition module, and the three-dimensional model data acquisition module, and is configured to perform building appearance inspection and synchronously acquire visible light images, infrared images, and three-dimensional model data of the building glass curtain wall through the visible light image acquisition module, the infrared image acquisition module, and the three-dimensional model data acquisition module. 3.The multi-modal data fusion based intelligent detection system for glass curtain wall according to claim 2, characterized in that, The UAV inspection module comprises an RTK positioning unit for cm-level accurate positioning of the UAV.

4. The multi-modal data fusion based intelligent detection system for glass curtain wall according to claim 2, characterized in that, The system comprises an intelligent flight path planning module configured to automatically plan a full-coverage inspection flight path and perform real-time obstacle avoidance based on the stored three-dimensional model data of the building or the three-dimensional model generated based on the three-dimensional model data acquired by the three-dimensional model data acquisition module. 5.The multi-modal data fusion based intelligent detection system for glass curtain wall according to claim 1, characterized in that, The data registration unit is configured to process the acquired data using the Shearlet-SURF algorithm. 6.The multi-modal data fusion based intelligent detection system for glass curtain wall according to claim 5, characterized in that, Processing the acquired data using the Shearlet-SURF algorithm comprises the following steps: S211: performing Shearlet transformation on the visible light images and the infrared images respectively and screening high-frequency subbands; S212: taking the high-frequency subbands screened in S21 as input, performing a feature point detection process of SURF, and extracting SURF feature points; S213: performing preliminary matching of the feature points using a bidirectional nearest neighbor distance ratio algorithm, and further purifying the matching pairs by applying a false matching elimination algorithm based on the spatial distribution consistency of the feature points and using the geometric relationship of the feature points in their Shearlet domains as a consistency constraint. S214: According to the matching pair after purification, the affine transformation matrix between the visible light image and the infrared image is calculated to realize pixel-level precise registration; then the registered infrared image data and the collected three-dimensional model data are mapped to the same visual coordinate system through coordinate system unification. 7.The multi-modal data fusion based intelligent detection system for glass curtain wall according to claim 1, wherein, The STDFusionNet neural network model comprises: Two independent feature extraction paths are used to process the registered visible light image and infrared image respectively, and each feature extraction path is based on a lightweight GhostNet; A dynamic weight allocation module, which inputs are the feature maps of the same level in the two independent feature extraction paths, is used to learn and generate a spatial weight map and a channel weight vector in real time; A feature fusion and recognition module is used to fuse the double-path features weighted by the DWAM module through element-by-element addition. 8.The multi-modal data fusion based intelligent detection system for glass curtain wall according to claim 1, wherein, The artificial intelligence recognition module is based on a Ghost convolution and bottleneck structure optimized YOLOv5 model for defect recognition and classification, which replaces the standard Bottleneck (C3) module in the YOLOv5 model with a GhostBottleneck module, and replaces part of the standard convolution with a depth separable convolution in the detection head part. 9.The multi-modal data fusion based intelligent detection system for glass curtain wall according to claim 8, characterized in that, The loss function of the artificial intelligence recognition module introduces Focal Loss as a classification loss function on the basis of the original CIoU Loss. 10.The multi-modal data fusion based intelligent detection system for glass curtain wall according to claim 8, characterized in that, The output module associates the detected defects with the stored three-dimensional model data of the building, visualizes the labeling on the three-dimensional model, and automatically generates a standardized detection report containing defect information, location map and repair suggestions.

Citation Information

Patent Citations

  • Building defect detection method and intelligent imaging device

    CN120976768A