Visual diagnosis system for building surface layer defects
By combining multimodal data acquisition and dynamic learning optimization modules, high-precision detection of building surface defects and generation of three-dimensional visualization reports are achieved, solving the problems of low efficiency and inaccurate positioning in traditional detection methods, and improving detection accuracy and maintenance guidance value.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-24
AI Technical Summary
Traditional building surface defect detection relies on manual visual inspection, which is inefficient and highly subjective. The two-dimensional image diagnosis results are difficult to accurately map to three-dimensional physical space, resulting in inaccurate defect location and insufficient maintenance guidance value.
A multimodal data acquisition module, including an imaging unit, a thermodynamic detection unit, and a 3D detection unit, is used. Combined with a heterogeneous data processing module, it performs spatiotemporal registration and feature pyramid network fusion. A dynamic learning optimization module performs defect diagnosis and generates a 3D visualization diagnostic report.
It achieves sub-millimeter level defect detection, improves defect identification accuracy, reduces false alarm rate, and enhances maintenance decision-making efficiency and reduces maintenance costs through 3D mapping and intelligent report generation.
Smart Images

Figure CN121724920A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of building surface quality inspection technology, and in particular to a visual diagnostic system for building surface defects. Background Technology
[0002] As an important protective layer and aesthetic carrier of building structure, the quality defects of architectural decorative surfaces directly affect the safety and functionality of buildings. Traditional inspection methods rely on manual visual inspection and contact measurement, which have problems such as low efficiency, strong subjectivity, and insufficient identification rate of hidden defects.
[0003] With the development of artificial intelligence and multimodal perception technology, vision-based automated defect diagnosis has become possible. This type of technology typically uses deep learning models to intelligently analyze acquired two-dimensional images, effectively identifying surface texture anomalies such as cracks and stains. However, the diagnostic results relying on two-dimensional images are difficult to accurately map to three-dimensional physical space, leading to inaccurate defect localization and insufficient repair guidance value. Therefore, this application proposes a new technical solution. Summary of the Invention
[0004] To improve the accuracy of defect location, this application provides a visual diagnostic system for building surface defects.
[0005] This application provides a visual diagnostic system for defects in building surface layers, employing the following technical solution:
[0006] A visual diagnostic system for defects in building surface layers includes a multimodal data acquisition module and a diagnostic analysis module. The multimodal data acquisition module includes an imaging unit and a three-dimensional detection unit, or includes a thermodynamic detection unit, an imaging unit, and a three-dimensional detection unit.
[0007] The diagnostic analysis module includes:
[0008] The heterogeneous data processing module is used to perform spatiotemporal registration on the data acquired by the multimodal data acquisition module using a parallel computing architecture, and to extract multi-scale fusion features using a feature pyramid network.
[0009] A dynamic learning optimization module is used to process fused features through a defect diagnosis model with an incremental learning mechanism, implementing the association between local features and global structure; and,
[0010] The diagnostic decision output module is used to generate diagnostic reports;
[0011] The heterogeneous data processing module includes a spatiotemporal registration unit for executing the spatiotemporal registration process. The spatiotemporal registration unit is used to spatially align the laser point cloud with the image data using the ICP algorithm and establish a two-dimensional-three-dimensional coordinate mapping relationship through projection transformation.
[0012] The diagnostic report of the diagnostic decision output module shall at least: combine the three-dimensional point cloud coordinates and the spatial topology relationship of the BIM model to generate a three-dimensional visualization diagnostic report containing at least two of the following: defect type, location coordinates, severity, and maintenance suggestions.
[0013] Optionally, the heterogeneous data processing module further includes:
[0014] An adaptive data augmentation unit is used to dynamically adjust the image enhancement strategy based on the ambient illumination and surface material parameters of the detected object.
[0015] Cross-modal correlation unit, which is used to establish the mapping relationship between visible light defects and thermodynamic anomalies using CycleGAN network;
[0016] A 3D reconstruction unit, used to generate point cloud models with specified accuracy based on the Poisson surface reconstruction algorithm.
[0017] Optionally, the dynamic learning optimization module includes:
[0018] The edge-end online learning module is used to filter new abnormal samples in real time using the locality-sensitive hashing algorithm and to fine-tune the model parameters.
[0019] The cloud-based periodic training module is used to reconstruct the feature space and optimize the network topology by periodically aggregating data from each edge node.
[0020] Optionally, the defect diagnosis model includes a cascaded attention convolutional neural network and a graph neural network, and it includes:
[0021] Spatial attention subnet, which is used to calculate cracks and voids using deformable convolution kernels;
[0022] The channel attention subnetwork is used to enhance chromatic aberration and contamination features using a squeeze-excitation network, and to generate a feature channel weight matrix.
[0023] Graph reasoning subnets are used to construct a correlation graph between the physical properties of the surface material and environmental parameters.
[0024] Optionally, the three-dimensional visualization diagnostic report of the diagnostic decision output module includes: multi-view three-dimensional visualization rendering implemented through the WebGL engine.
[0025] Optionally, the multimodal data acquisition module further includes a housing, a mobile vehicle for mounting the housing, and a stabilizing mechanism for the acquisition units. The mobile vehicle includes a drone or a smart car. The stabilizing mechanism includes a gimbal structure and is mounted on the mobile vehicle. The housing is mounted on the gimbal structure. At least two of the thermodynamic detection unit, imaging unit, and three-dimensional detection unit are mounted on the housing.
[0026] Optionally, if the mobile vehicle is a drone, the container is mounted below the drone, and the stabilization mechanism further includes:
[0027] Electric winding devices, of which there are at least three and are evenly distributed around the central axis of the housing;
[0028] A link-type composite tube, comprising multiple single tubes that are hinged together in sequence, with one single tube at the very end fixed to an electric winder.
[0029] The positioning inner rod is longer than a single tube and is adapted to the inner diameter of the single tube;
[0030] The base is located directly below the drone and is used to temporarily fix the end of the link-type composite tube away from the drone.
[0031] A rod feeder, mounted on a base, is used to feed the positioning inner rod into a chain-link assembly tube that is dragged below the housing.
[0032] Optionally, the base has a hollow structure and a closed cylindrical body at the top. The top of the cylindrical body has a vertically penetrating positioning groove. The positioning groove is elongated and has a circular hole at one end with a diameter greater than the width of the main body.
[0033] The outer wall of the bottommost single tube is fixed with an anti-detachment ring. The diameter of the anti-detachment ring is smaller than the diameter of the circular hole at one end of the positioning groove and larger than the width of the main body of the positioning groove. The positioning inner rod is a ferromagnetic structure.
[0034] The rod feeder includes:
[0035] An electromagnet, which is fixed to the base and located on the side of the single tube that is snapped into the positioning groove;
[0036] The nozzle is fixed to the base and is located directly below the bottommost single tube;
[0037] A water pump, which is connected to the nozzle via a pipe;
[0038] The conveyor belt is positioned laterally above the nozzle with one end close to the upward extension line of the nozzle.
[0039] End plate, which is located in front of the conveyor belt and allows the upward extension line of the nozzle to pass between it and the conveyor belt;
[0040] Side baffles, located on the conveyor belt, at least two in number, with a positioning inner rod placed between them.
[0041] Optionally, the rod feeder further includes a controller electrically connected to the water pump and electromagnet, and electrically connected to a laser beam sensor. The laser beam sensor is installed inside the base, and the beam path is traversed by the single tube leaving the conveyor belt as it moves. The controller can be configured as follows:
[0042] The delivery rod length is calculated based on the feedback from the laser beam sensor and the preset length of the positioning inner rod.
[0043] The power of the water pump and electromagnet is controlled according to the length of the feed rod.
[0044] In summary, this application includes at least the following beneficial technical effects:
[0045] First, by acquiring and spatiotemporally registering imaging, thermodynamic detection, and 3D detection data, combined with multi-scale feature fusion of the feature pyramid network, sub-millimeter-level precision detection of defects such as cracks and hollow areas is achieved. The dynamic learning optimization module utilizes a two-stage update mechanism to enable the diagnostic model to respond to new defects in real time at the edge, and continuously optimize the generalization ability in the cloud to improve the defect recognition accuracy and reduce the false alarm rate.
[0046] Secondly, the 3D point cloud coordinates are linked to the BIM model topology, accurately mapping 2D defect features to 3D solid space, reducing defect location deviation. At the same time, it automatically outputs a 3D visualization report containing expansion trend prediction and maintenance process matching, improving maintenance decision-making efficiency, reducing maintenance costs, and increasing maintenance guidance value. Attached Figure Description
[0047] Figure 1 This is a system block diagram of an embodiment of this application;
[0048] Figure 2 This is a flowchart of an embodiment of this application;
[0049] Figure 3 This is a schematic diagram of the overall structure when the mobile vehicle is an unmanned aerial vehicle (UAV).
[0050] Figure 4 This is a side view of the enclosure and base structure.
[0051] Figure 5 This is a structural diagram of the rod feeder area.
[0052] Explanation of reference numerals in the attached drawings: 1. Box body; 2. Mobile carrier; 3. Gimbal structure; 4. Electric rewinder; 5. Chain link combined pipe; 51. Single pipe; 511. Anti-derailment ring; 6. Positioning inner rod; 7. Base; 71. Cylinder; 8. Rod feeder; 81. Electromagnet; 82. Nozzle; 83. Conveyor belt; 84. End plate; 85. Side baffle; 86. Laser beam sensor. Detailed Implementation
[0053] The following is in conjunction with the appendix Figures 1-2 This application will be described in further detail.
[0054] This application discloses a visual diagnostic system for defects in building surface layers.
[0055] Reference Figure 1 A visual diagnostic system for building surface defects includes a multimodal data acquisition module and a diagnostic analysis module.
[0056] The multimodal data acquisition module includes an imaging unit and a three-dimensional detection unit, or includes a thermodynamic detection unit, an imaging unit, and a three-dimensional detection unit. Specifically, it acquires two-dimensional image data (i.e., surface texture images, such as cracks and color differences) of the decorative surface layer through visible light imaging equipment (such as a camera); it acquires thermodynamic data of the decorative surface layer, i.e., temperature distribution, through infrared thermal imaging equipment, and identifies defects invisible to the naked eye, such as hollowness and water seepage, through temperature difference; and it acquires three-dimensional deformation data (such as unevenness and warping) of the decorative surface layer through three-dimensional laser scanning equipment.
[0057] Based on the above settings, the multimodal data acquisition module can collect multi-dimensional raw data of the building decoration surface as a basis, and considering the purpose, it should at least have two-dimensional images and three-dimensional scan data.
[0058] The diagnostic analysis module uses devices with data analysis and processing capabilities, such as host computers and servers, as hardware carriers, and loads corresponding computer programs. These programs must include at least the following modules: The diagnostic analysis module includes:
[0059] 1) Heterogeneous data processing module, which uses a parallel computing architecture to perform spatiotemporal registration on the data collected by the multimodal data acquisition module and uses a feature pyramid network to extract multi-scale fusion features.
[0060] The heterogeneous data processing module includes a spatiotemporal registration unit, which uses the ICP algorithm to spatially align laser point clouds with image data and establishes a two-dimensional to three-dimensional coordinate mapping relationship through projection transformation; Example workflow:
[0061] An improved ICP algorithm (50 iterations, matching error threshold of 0.8 mm) is used to spatially align point cloud and image data, that is, to accurately align 2D image and 3D point cloud, and to establish a two-dimensional to three-dimensional coordinate mapping relationship through projection transformation.
[0062] The heterogeneous data processing module also includes:
[0063] An adaptive data augmentation unit is used to dynamically adjust the image enhancement strategy based on the ambient illumination and surface material parameters of the detected object. Example: Using adaptive gamma correction, the Retinex algorithm is used to enhance low-light images and increase the brightness of low-light areas.
[0064] Cross-modal correlation unit, which is used to establish the mapping relationship between visible light defects and thermodynamic anomalies using CycleGAN network; Example: Construct an improved CycleGAN network (add gradient consistency loss term) to establish a pixel-level correspondence between crack regions (visible light) and temperature difference anomaly regions (infrared);
[0065] A 3D reconstruction unit is used to generate a point cloud model with specified (e.g., sub-millimeter) accuracy based on the Poisson surface reconstruction algorithm; Example: using an improved Poisson reconstruction algorithm (depth 10, sample spacing 0.4mm), a normal constraint term is introduced to eliminate scan shadows.
[0066] Regarding the use of feature pyramid networks to extract multi-scale fusion features, it includes:
[0067] Deformable convolutional layers are embedded in the feature pyramid network to extract multi-scale fusion features including crack width, temperature gradient of hollow area, and surface curvature change, generating a feature tensor with dimensions of 256×256×128.
[0068] Based on the above settings, this system eliminates spatiotemporal synchronization errors by performing spatiotemporal registration and alignment, data augmentation, 3D reconstruction, and feature fusion on the collected multi-source data, providing a more accurate data foundation for subsequent defect feature correlation analysis.
[0069] 2) Dynamic learning optimization module, which processes the fused features using a defect diagnosis model with an incremental learning mechanism, and associates local features with the global structure. The dynamic learning optimization module includes:
[0070] The edge-end online learning module is used to filter new abnormal samples in real time using the Locality Sensitive Hash (LSH) algorithm and fine-tune model parameters to adapt to new situations (such as new defect patterns). For example, when a suspected defect with a confidence level below 0.85 is detected, the LSH index is automatically triggered to fine-tune model parameters at the edge, with the update step size set to 0.001-0.01 for adaptive adjustment.
[0071] The cloud-based periodic training module is used to periodically (e.g., monthly) aggregate data from each edge node, reconstruct the feature space, and optimize the network topology. This enables regular aggregation of all data for global model optimization and version iteration, improving overall performance and generalization ability. It addresses the problem that statically trained models cannot adapt to the dynamic evolution of defect patterns caused by material aging and environmental stress changes. Example:
[0072] By aggregating 30 days of edge node data on a cloud platform, using knowledge distillation technology to compress the model size, reconstructing the Euclidean distance metric matrix of the feature space, and optimizing the edge weight calculation function of the graph neural network, the system can reduce computing resource consumption by 70% while maintaining high accuracy, thus adapting to the deployment needs of large-scale building clusters.
[0073] As can be seen from the above, this system adopts an edge-cloud collaborative architecture. For example, the edge computing node is equipped with a lightweight diagnostic model to achieve real-time defect detection within 200ms, and the cloud platform integrates a digital twin engine to establish a virtual mirror system that includes material aging models and environmental stress models.
[0074] As can be seen from the above, the dynamic learning optimization module can dynamically adjust the learning rate and optimize the algorithm parameters based on the historical defect data of the building decoration surface and the current environmental parameters, so as to improve the model's ability to identify new defect types and reduce the false alarm rate.
[0075] Regarding the defect diagnosis model mentioned in the dynamic learning optimization module, it includes cascaded attention convolutional neural networks and graph neural networks, specifically including:
[0076] Spatial attention subnet, which is used to calculate cracks and voids (regional heat maps) using deformable convolution kernels.
[0077] The channel attention subnetwork is used to enhance chromatic aberration and contamination features using a squeeze-excitation network, and to generate a feature channel weight matrix.
[0078] Graph reasoning subnets are used to construct a correlation map between the physical properties (such as the material's elastic modulus) of the surface material and environmental parameters (such as temperature and humidity parameters).
[0079] 3) Diagnostic decision output module, which is used to generate diagnostic reports; Regarding the diagnostic report: Combining the 3D point cloud coordinates and the spatial topology relationship of the BIM model, it generates a 3D visualization diagnostic report containing at least two of the following: defect type (such as cracks, hollows, etc.), location coordinates (such as coordinates on the 3D model / BIM model), severity (such as crack width, hollow area), and maintenance suggestions (such as recommended processes, material usage, trend prediction).
[0080] Example: By calling a preset material failure knowledge graph and combining the 3D point cloud coordinates with the spatial topology of the BIM model, an interactive diagnostic report is generated, which includes crack propagation trend prediction, hollow area ratio statistics, and repair process matching degree. Multi-view 3D visualization rendering is achieved through the WebGL engine.
[0081] Understandably, the UI display effect of the diagnostic decision output module is the intelligent report generation module (section), which is used to automatically generate diagnostic report templates that meet industry standards based on the diagnostic results, and supports custom editing functions to meet the needs of different users.
[0082] In summary, this system:
[0083] First, by simultaneously acquiring and spatiotemporally registering visible light, infrared thermal imaging, and 3D laser scanning, and combining multi-scale feature fusion with feature pyramid network, sub-millimeter-level precision detection of defects such as cracks and hollow areas is achieved; the dynamic learning optimization module utilizes a two-stage update mechanism to enable the diagnostic model to respond to new defects in real time at the edge, and continuously optimize the generalization ability in the cloud to improve the defect recognition accuracy and reduce the false alarm rate.
[0084] Secondly, based on the improved Poisson surface reconstruction algorithm and BIM model topology association, two-dimensional defect features are accurately mapped to three-dimensional solid space, with positioning errors controlled within ±1.5mm. The intelligent report generation engine, combined with a material failure knowledge graph, automatically outputs a three-dimensional visualization report containing extended trend predictions and maintenance process matching, improving maintenance decision-making efficiency and reducing maintenance costs.
[0085] Furthermore, this invention features a lightweight model at the edge to achieve real-time detection within 200ms, while a cloud-based digital twin engine integrates material aging and environmental stress models, supporting long-term defect evolution analysis. By compressing the model size and optimizing graph neural network parameters through knowledge distillation technology, the system can reduce computational resource consumption by 70% while maintaining high accuracy, adapting to the deployment needs of large-scale building complexes.
[0086] In another embodiment of the application, based on the above settings, the implementation of the present invention is specifically described using the inspection of the exterior wall decorative surface of a high-rise commercial complex as an example application scenario; the exterior wall of the example building adopts a composite system of dry-hanging stone and glass curtain wall, which has typical defects such as cracks, hollowness, and aging of the sealant.
[0087] (a) Multimodal data acquisition:
[0088] It is equipped with a FLIR A700 infrared thermal imager (640×480 resolution, thermal sensitivity 0.05℃), a Nikon D850 visible light camera (45.75 million pixels) and a FARO Focus S350 3D laser scanner (range measurement error ±1mm); during use, a target ball is used for assisted positioning, and the spatial positioning error is controlled within ±1.2mm at a detection distance of 10 meters.
[0089] Acquired data: (1) Visible light image: resolution 8256×5504, stored in RAW format; (2) Infrared thermal image: temperature resolution 0.1℃, generating a 256-level pseudo-color image; (3) Point cloud data: point spacing 0.5mm, single-station scanning time 2.5 minutes.
[0090] (II) Heterogeneous Data Processing:
[0091] (1) Spatiotemporal registration: The improved ICP algorithm (50 iterations, matching error threshold of 0.8 mm) is used to align the point cloud and image data, and a two-dimensional-three-dimensional coordinate mapping relationship is established through projection transformation.
[0092] (2) Adaptive enhancement: In low-light areas (<50 lux), the Retinex enhancement algorithm (scale parameter σ=15 / 80 / 160) is enabled, and the CLAHE algorithm (clip limit=2.0, tile=8×8) is used to improve the texture contrast of marble material.
[0093] (3) Cross-modal association: Construct an improved CycleGAN network (add gradient consistency loss term), train for 200 epochs, establish pixel-level correspondence between crack region (visible light) and temperature difference abnormal region (infrared), and achieve cross-validation accuracy of 92.3%.
[0094] (4) 3D reconstruction: The improved Poisson reconstruction algorithm (depth 10, sample spacing 0.4mm) was adopted, and the normal constraint term was introduced to eliminate the scanning shadow. The average error of the reconstructed surface was 0.28mm.
[0095] (III) Dynamic Learning Optimization:
[0096] Deploying a lightweight model at the edge (MobileNetV3 + GraphSAGE, 4.7M parameters):
[0097] (1) Online incremental learning: When the detection confidence p < 0.85, the radius of the LSH hash bucket is set to 0.45, and samples with feature differences > 15% are selected and stored in the buffer. Fine-tuning is triggered every 50 samples accumulated (learning rate 0.003, Adam optimizer).
[0098] (2) 200,000 sets of edge data are aggregated in the cloud every month. Knowledge distillation (temperature coefficient T=6) is used to compress the ResNet152 teacher model (AP=0.891) to the EfficientNet-B4 student model (AP=0.877), reducing the model size by 68%.
[0099] (iv) Defect diagnosis and analysis: For example, the detection process for a certain hollow defect:
[0100] (1) Spatial attention subnet (deformable convolution deform_groups=4) locates abnormal regions, and the peak intensity of the heat map reaches 0.93.
[0101] (2) The graph reasoning subnet is constructed to form an association graph containing nodes (stone thickness 25mm, elastic modulus 50GPa, ambient humidity 72%RH), and the probability of hollow area expansion is calculated to be 34% / year.
[0102] (3) BIM integration: The defect coordinates (X=35.6m, Y=12.8m, Z=48.2m) were mapped to the Revit model, and the deviation from the structural beam position was detected to be 2.3mm.
[0103] (v) Diagnostic Report Generation: The intelligent engine calls the ASTM E3030 standard template to automatically generate a report including:
[0104] (1) 3D visualization: WebGL rendering of the hollow area (area 1.2m², temperature difference ΔT=4.3℃);
[0105] (2) Repair recommendations: Epoxy resin injection is recommended (87% matching accuracy), with an estimated material consumption of 2.6 kg;
[0106] (3) Trend prediction: Based on the LSTM model, it is predicted that the area of the hollow area may expand to 1.8m² after 6 months.
[0107] (vi) Performance verification
[0108] In a 3000m² testing area, the data compared with traditional methods are shown in the table below, and the performance verification table is also shown.
[0109] Performance Verification Table:
[0110] index This invention Manual inspection Traditional system Crack detection rate 98.2% 76.4% 89.1% Hollow drum positioning accuracy (mm) ±1.3 ±5.0 ±2.8 Single-area detection time (s) 8.7 1800 23.5 New defect adaptation cycle 2 hours N / A 72 hours
[0111] This embodiment verifies the effectiveness of the system in practical engineering, particularly in achieving precise location of 0.9mm-level cracks on complex curved surfaces (glass curtain wall joints), and reducing the false alarm rate caused by seasonal temperature changes from 12.3% to 3.8% through dynamic learning. Compared with traditional manual inspection methods, this system demonstrates significant advantages in crack detection rate, hollow area location accuracy, single-area inspection time, and adaptation cycle for new defects.
[0112] Specifically, the crack detection rate of this invention reaches 98.2%, far exceeding the 76.4% of manual inspection and the 89.1% of traditional AI systems, effectively avoiding the omission of defects. Regarding the accuracy of hollow area positioning, this invention achieves a high-precision positioning of ±1.3mm, superior to the ±5.0mm of manual inspection and the ±2.8mm of traditional AI systems, providing more accurate location information for maintenance work. Furthermore, the system takes only 8.7 seconds to inspect a single area, greatly improving inspection efficiency compared to 1800 seconds for manual inspection and 23.5 seconds for traditional AI systems, demonstrating a significant time advantage. More importantly, the system's adaptation period to new defects is only 2 hours, enabling rapid response to newly emerging defect types, while traditional AI systems require 72 hours, and manual inspection cannot adapt to new defects.
[0113] Reference Figure 2 In another embodiment of this application, a visual diagnosis method for building surface defects based on the above system includes the following steps:
[0114] Step 1: Simultaneously acquire multimodal data of the decorative surface layer through visible light imaging unit, infrared thermal imaging unit and three-dimensional laser scanning unit, and use hardware synchronous triggering mechanism to ensure data timestamp alignment, and control spatial positioning error within ±1.5mm;
[0115] Step 2: Perform spatiotemporal registration on multi-source heterogeneous data, use the improved ICP algorithm to spatially align point cloud and image data, use adaptive gamma correction and Retinex algorithm to enhance low-light images, and establish the mapping relationship between visible light texture and thermodynamic anomaly through cross-modal association units.
[0116] Step 3: Embed deformable convolutional layers in the feature pyramid network to extract multi-scale fusion features including crack width, temperature gradient of hollow area, and surface curvature change, and generate a feature tensor with dimensions of 256×256×128.
[0117] Step 4: Input the fused features into a cascaded attention convolutional neural network and a graph neural network. Calculate the heat map of the defect area through the spatial attention subnetwork, generate the feature channel weight matrix through the channel attention subnetwork, and construct the graph inference subnetwork to build a correlation map of nodes containing material elastic modulus and environmental temperature and humidity parameters.
[0118] Step 5: Based on the online incremental learning mechanism of the dynamic learning optimization module, when a suspected defect with a confidence level below 0.85 is detected, the local sensitive hash index is automatically triggered to fine-tune the model parameters at the edge, and the update step size is set to 0.001-0.01 for adaptive adjustment.
[0119] Step Six: The diagnostic decision output module calls the material failure knowledge graph, combines the 3D point cloud coordinates with the spatial topology of the BIM model, and generates an interactive diagnostic report that includes crack propagation trend prediction, hollow area ratio statistics, and repair process matching degree. It also uses the WebGL engine to achieve multi-view 3D visualization rendering.
[0120] Step 7: Aggregate the accumulated edge node data of 30 days on the cloud platform, use knowledge distillation technology to compress the model size, reconstruct the Euclidean distance metric matrix of the feature space, optimize the edge weight calculation function of the graph neural network, and complete the global model iterative update.
[0121] In summary, this application achieves comprehensive and accurate perception of decorative surface defects through multimodal data acquisition and heterogeneous data processing technologies, providing a solid foundation for subsequent defect diagnosis and analysis. The introduction of a dynamic learning optimization module enables the diagnostic model to continuously learn and optimize, thereby adapting to new defect types, improving recognition accuracy, and reducing false alarm rates. The defect diagnosis and analysis module, through the collaborative work of spatial attention subnets, channel attention subnets, and graph reasoning subnets, achieves accurate judgment of defect type, location coordinates, severity, and maintenance suggestions, providing a scientific basis for maintenance decisions. The addition of an intelligent report generation engine further presents the diagnostic results in the form of a three-dimensional visual report, greatly improving the efficiency of maintenance decisions and reducing maintenance costs.
[0122] In another embodiment of this application, relying solely on the heterogeneous data processing module for spatiotemporal registration using algorithms, without adjusting the operating modes of the thermodynamic detection unit, imaging unit, and 3D detection unit, still results in significant registration difficulty and errors. Therefore, synchronous acquisition is preferred, which requires hardware improvements. Specifically:
[0123] Reference Figure 3 The multimodal data acquisition module also includes a housing 1, a mobile carrier 2 for mounting the housing 1, and a stabilizing mechanism for the acquisition unit. The mobile carrier 2 includes a drone or a smart car. This embodiment uses a drone as an example for illustration.
[0124] The stabilization mechanism includes a gimbal structure 3 and is mounted on the mobile vehicle 2. The gimbal structure 3, such as a camera gimbal, has its base fixed to the bottom bracket of the drone by bolts. The stabilization unit is fixed to the housing 1 by bolts, clips, etc. The housing 1 can be a frame structure. At least two of the thermodynamic detection unit, imaging unit, and 3D detection unit are mounted on the housing 1. Taking all three as an example: the thermodynamic detection unit and imaging unit are distributed to the left and right of the 3D detection unit, and the position is above. The reason for this arrangement is that, for example, the 3D laser scanner of the FARO series, the laser emission part rotates continuously during operation to generate lasers emitted up, down, left and right. The above distribution method can reduce interference.
[0125] Regarding the remote and synchronous start / stop triggering of electronic devices, the power-on of the three devices and the sending of start / stop control commands can be controlled by an additional control board, or by adding a network device for remote access and control via a personal terminal. This is existing technology and will not be elaborated further.
[0126] Reference Figure 4 and Figure 5 In another embodiment of this application, the stabilizing mechanism further includes:
[0127] Electric winders 4, of which there are at least three and are evenly distributed around the central axis of the housing 1; in this embodiment, three are used as an example: they are distributed at 120° intervals.
[0128] The chain-link composite tube 5 includes multiple single tubes 51 that are hinged to each other in sequence, and one single tube 51 located at the very end is fixed to the electric winder 4.
[0129] The positioning inner rod 6 is longer than the single tube 51 and is adapted to the inner diameter of the single tube 51;
[0130] The base 7 is located directly below the drone and is used to temporarily fix the end of the link-type composite tube 5 away from the drone.
[0131] The rod feeder 8 is mounted on the base 7 and is used to feed the positioning inner rod 6 into the chain-link combination tube 5 that is dragged below the housing 1.
[0132] In use, at low positions, such as within 5m, the box 1 is directly supported by a triangular bracket, and then the building surface to be tested is detected. At higher positions, the box 1 is first installed on the gimbal structure 3 below the drone, and then one end of the link-type combined tube 5 is connected to the electric rewinder 4. After that, the drone is launched and controlled to hover at the target height. Then, the electric rewinder 4 is controlled to release the link-type combined tube 5. The base 7 is placed below the drone, and the lowest section of the link-type combined tube 5 is locked with the base 7. Then, the rod feeder 8 sends the positioning inner rods 6 into the positioning inner rods 6. Each positioning inner rod 6 is partly in the upper single tube 51 and partly in the lower adjacent single tube 51. In this way, the link-type combined tube 5 can be transformed from a flexible rod structure into a rigid rod structure, so that the three link-type combined tubes 5 can work together to stabilize the drone and reduce the erroneous detection data and duplicate data caused by the movement of the drone during data collection.
[0133] Regarding the electric winder 4, an electric rope winder can be used.
[0134] The link-type combined tube 5 can be a thin flexible tube rather than a rigid or metal tube, to facilitate winding. The reason it's not a complete tube structure is that winding a single tube would significantly affect the part near the electric winder 4, causing the upper part to be flattened, affecting the insertion of the positioning inner rod 6. Furthermore, after the positioning inner rod 6 is inserted, even slight loosening can cause it to shift vertically, affecting its positioning capability. The hinges between the individual flexible tubes can be either a one-piece narrow strip structure extending from both sides, or they can form an ear structure, with adjacent ear structures connected by a pivot.
[0135] In one embodiment, the base 7 has a hollow structure; the top support fixes the closed upper cylinder 71, and the top of the cylinder 71 is provided with a positioning groove. The positioning groove is elongated when viewed from above and has a round hole at one end. The diameter of the round hole is larger than the main width of the positioning groove.
[0136] The outer wall of the lowest single tube 51 of the chain-link combined tube 5 is fixed with an anti-detachment ring 511. The diameter of the anti-detachment ring 511 is smaller than the diameter of the round hole at one end of the positioning groove, and the diameter is larger than the width of the main body of the positioning groove.
[0137] That is, the single tube 51 can be inserted into the round hole of the positioning groove, and then moved toward the end away from the round hole to initially fix the bottom single tube 51. It can be understood that the number of cylinders 71 and corresponding mechanisms is three, to match the number of chain-link combined tubes 5.
[0138] The rod feeder 8 is located in the base 7 and below the positioning slot, and includes:
[0139] Electromagnet 81 is fixed to base 1 and located to the side of single tube 51 that is inserted into positioning groove, and is in close proximity to it; single tube 51 is non-magnetic, while positioning inner rod 51 is ferromagnetic and can be magnetically attracted.
[0140] The nozzle 82 is fixed to the base 7 and is located directly below the bottommost single pipe 51;
[0141] A water pump, which is connected to the nozzle 82 via a pipe;
[0142] The conveyor belt 83 is arranged laterally above the nozzle 82 with one end close to the upward extension line of the nozzle 82.
[0143] End plate 84, which is located in front of conveyor belt 83 and causes the upward extension line of nozzle 82 to pass between it and conveyor belt 83;
[0144] Side baffles 85, which are located on conveyor belt 83 and there are at least two of them, with a positioning inner rod placed between them.
[0145] When in use, the staff places the positioning inner rod 6 vertically onto the conveyor belt 83 from the feeding hole pre-opened in the base 7. The movement of the conveyor belt 83 causes the positioning inner rod 6 to move toward the end plate 84. After it leaves the conveyor belt 83, it falls onto the nozzle 82. The water pump works, and the nozzle 82 sprays water to push the positioning inner rod 6 upward to insert into the single pipe 51.
[0146] When a positioning inner rod 6 needs to be inserted, the nozzle 82 stops the high-pressure water jet; the electromagnet 81 temporarily holds the positioning inner rod 6 in the single tube 51 to prevent it from sliding down; when the nozzle 82 is turned on, the electromagnet 81 is de-energized and loses its magnetism, and no longer obstructs the positioning inner rod 6.
[0147] Based on the above setup, the link-type combined tube 5 can be easily fed in using the rod feeder 8.
[0148] Furthermore, to prevent the positioning inner rod 6 from tipping over on the conveyor belt 83, the lower end of the positioning inner rod 6 is designed as an inverted hollow cone, which facilitates interlocking and also makes it less prone to tipping over; if necessary, the conveyor belt 83 can even be tilted towards the nozzle 82. The lower end of the bottommost single tube 51 can also be made into an inverted cone shape for easy insertion.
[0149] Furthermore, the nozzle 82 is connected to a three-way pipe. One end of the three-way pipe is connected to a solenoid valve, which in turn is connected to a water pump. The third end of the three-way pipe is connected to a pressure valve. This design prevents the water pump from frequently starting and stopping after it is turned on, and closing the solenoid valve will stop water from flowing from the nozzle 82. Also, if the internal pressure of the pipe is too high when the pump is turned off, it can be released through the pressure valve.
[0150] In another embodiment of this application, the rod feeder 8 also includes a controller, which is electrically connected to the water pump, solenoid valve, conveyor belt and electromagnet 81, thereby enabling automated rod feeding through settings; it should be noted that the controller is also connected to a laser beam sensor 86, which is installed inside the base 7. The beam path is passed by the single tube 51 leaving the conveyor belt as it moves, that is, the controller can complete the counting of the positioning inner rod 6 based on the feedback from the laser beam sensor 86.
[0151] Therefore, the controller can be configured as follows:
[0152] The delivery rod length is calculated based on the feedback from the laser beam sensor 86 and the preset length of the positioning inner rod 6.
[0153] The power of the water pump and electromagnet 81 is controlled according to the length of the feed rod; the correspondence between length and water pump power, and length and electromagnet 81 power is pre-verified and recorded, and can be retrieved when needed.
[0154] Understandably, the longer the rod is, the heavier it is, and the greater the attraction force of the electromagnet 81 and the water supply pressure of the water pump are required; if it is set according to a larger standard from the beginning, it will lead to energy waste.
[0155] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A visual diagnostic system for building surface defects, characterized in that, It includes a multimodal data acquisition module and a diagnostic analysis module. The multimodal data acquisition module includes an imaging unit and a three-dimensional detection unit, or includes a thermodynamic detection unit, an imaging unit, and a three-dimensional detection unit. The diagnostic analysis module includes: The heterogeneous data processing module is used to perform spatiotemporal registration on the data acquired by the multimodal data acquisition module using a parallel computing architecture, and to extract multi-scale fusion features using a feature pyramid network. A dynamic learning optimization module is used to process fused features through a defect diagnosis model with an incremental learning mechanism, implementing the association between local features and global structure; and, The diagnostic decision output module is used to generate diagnostic reports; The heterogeneous data processing module includes a spatiotemporal registration unit for executing the spatiotemporal registration process. The spatiotemporal registration unit is used to spatially align the laser point cloud with the image data using the ICP algorithm and establish a two-dimensional-three-dimensional coordinate mapping relationship through projection transformation. The diagnostic report of the diagnostic decision output module shall at least: combine the three-dimensional point cloud coordinates and the spatial topology relationship of the BIM model to generate a three-dimensional visualization diagnostic report containing at least two of the following: defect type, location coordinates, severity, and maintenance suggestions.
2. The visual diagnostic system for building surface defects according to claim 1, characterized in that: The heterogeneous data processing module further includes: An adaptive data augmentation unit is used to dynamically adjust the image enhancement strategy based on the ambient illumination and surface material parameters of the detected object. Cross-modal correlation unit, which is used to establish the mapping relationship between visible light defects and thermodynamic anomalies using CycleGAN network; A 3D reconstruction unit, used to generate point cloud models with specified accuracy based on the Poisson surface reconstruction algorithm.
3. The visual diagnostic system for building surface defects according to claim 1, characterized in that: The dynamic learning optimization module includes: The edge-end online learning module is used to filter new abnormal samples in real time using the locality-sensitive hashing algorithm and to fine-tune the model parameters. The cloud-based periodic training module is used to reconstruct the feature space and optimize the network topology by periodically aggregating data from each edge node.
4. The visual diagnostic system for building surface defects according to claim 1, characterized in that: The defect diagnosis model comprises a cascaded attention convolutional neural network and a graph neural network, and includes: Spatial attention subnet, which is used to calculate cracks and voids using deformable convolution kernels; The channel attention subnetwork is used to enhance chromatic aberration and contamination features using a squeeze-excitation network, and to generate a feature channel weight matrix. Graph reasoning subnets are used to construct a correlation graph between the physical properties of the surface material and environmental parameters.
5. The visual diagnostic system for building surface defects according to claim 1, characterized in that: The diagnostic decision output module provides a 3D visualization diagnostic report, which includes multi-view 3D visualization rendering achieved through a WebGL engine.
6. The visual diagnostic system for building surface defects according to claim 1, characterized in that: The multimodal data acquisition module also includes a housing (1), a mobile vehicle (2) for mounting the housing (1), and a stabilizing mechanism for the acquisition unit. The mobile vehicle (2) includes a drone or a smart car. The stabilizing mechanism includes a gimbal structure (3) and is installed on the mobile vehicle (2). The housing (1) is installed on the gimbal structure (3). At least two of the thermodynamic detection unit, imaging unit, and three-dimensional detection unit are installed on the housing (1).
7. The visual diagnostic system for building surface defects according to claim 6, characterized in that: If the mobile vehicle (2) is a drone, the housing (1) is mounted below the drone, and the stabilization mechanism further includes: Electric winding devices (4), of which there are at least three and are evenly distributed around the central axis of the housing (1); The link-type composite tube (5) includes multiple single tubes (51) that are hinged to each other in sequence, and one single tube (51) at the very end is fixed to the electric winder (4). The positioning inner rod (6) is longer than the single tube (51) and is adapted to the inner diameter of the single tube (51); The base (7) is located directly below the UAV and is used to temporarily fix the link-type composite tube (5) away from the UAV. A rod feeder (8) is mounted on a base (7) and is used to feed the positioning inner rod (6) into a chain-link assembly tube (5) that is dragged under the housing (1).
8. The visual diagnostic system for building surface defects according to claim 7, characterized in that: The base (7) has a hollow structure and a closed cylinder (71) is provided on the top. The top of the cylinder (71) has a vertically penetrating positioning groove. The positioning groove is long and has a circular hole with a diameter greater than the width of the main body at one end. The outer wall of the bottommost single tube (51) is fixed with an anti-detachment ring (511). The diameter of the anti-detachment ring (511) is smaller than the diameter of the circular hole at one end of the positioning groove and larger than the width of the main body of the positioning groove. The positioning inner rod (6) is a ferromagnetic structure. The rod feeder (8) includes: An electromagnet (81) is fixed to the base (7) and located to the side of the single tube (51) that is inserted into the positioning groove; The nozzle (82) is fixed to the base (7) and located directly below the bottommost single tube (51); A water pump, which is connected to the nozzle (82) via a pipe; The conveyor belt (83) is arranged laterally above the nozzle (82) with one end close to the upward extension line of the nozzle (82); End plate (84), which is located in front of the conveyor belt (83) and allows the upward extension line of the nozzle (82) to pass between it and the conveyor belt (83); Side baffles (85), located on the conveyor belt (83), at least two in number, with a positioning inner rod (6) placed between them.
9. The visual diagnostic system for building surface defects according to claim 8, characterized in that: The rod feeder (8) also includes a controller electrically connected to the water pump, the electromagnet (81), and a laser beam sensor electrically connected to it. The laser beam sensor (86) is installed inside the base (7), and the beam path is passed by the single tube (51) leaving the conveyor belt (83) during its movement. The controller can be configured as follows: The length of the delivery rod is calculated based on the feedback from the laser beam sensor (86) and the preset length of the positioning inner rod (6); The power of the water pump and electromagnet (81) is controlled according to the length of the feed rod.