Urban building disease detection method and device, electronic equipment and storage medium

By using pixel-level registration and hierarchical deep learning recognition of multimodal image data, combined with three-dimensional spatial mapping, the accuracy of UAV identification and autonomous obstacle avoidance in high-density urban areas has been solved, enabling efficient building defect detection and maintenance guidance.

CN121937451APending Publication Date: 2026-04-28SHENZHEN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN UNIV
Filing Date
2026-03-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing drone-based exterior wall inspection technology struggles to accurately identify building defects in high-density urban areas, especially in complex urban canyon environments where the lack of prior 3D information makes autonomous obstacle avoidance and accurate data collection difficult. Furthermore, traditional detection methods cannot penetrate the wall surface to identify internal hazards, resulting in a high false alarm rate and a lack of 3D spatial correlation in the detection results.

Method used

A method combining multimodal data acquisition and hierarchical deep learning recognition is adopted. By pixel-level registration of visible light and thermal infrared image data, combined with a hierarchical deep learning model, building defects are identified and three-dimensional spatial mapping is performed to generate a heat map of defect distribution, thereby realizing the three-dimensional spatial positioning and visualization analysis of defects.

Benefits of technology

It improves the accuracy and flexibility of building defect identification, enabling accurate location of defects in complex backgrounds, generating intuitive three-dimensional spatial distribution maps to guide maintenance work, and improving the accuracy and effectiveness of inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937451A_ABST
    Figure CN121937451A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of building disease detection, in particular to an urban building disease detection method and device, electronic equipment and a storage medium. Multi-modal image data formed by original visible light and thermal infrared image data is obtained, and an original thermal infrared image is subjected to geometric correction; calculating a mapping relation with an original visible light image so as to complete pixel-level registration, obtaining target multi-modal image data, inputting the target multi-modal image data into a hierarchical deep learning recognition model, recognizing building disease information, then performing three-dimensional space mapping, generating a building three-dimensional mesh model containing disease three-dimensional space setting coordinates, and finally performing three-dimensional mesh modeling. And then calculating a relationship between a model surface grid vertex and a disease point cloud density, generating a disease distribution thermodynamic diagram, analyzing disease aggregation characteristics in multiple dimensions according to the thermodynamic diagram, and quantitatively analyzing spatial correlation between the disease and a building construction node in combination with building component information. According to the invention, the urban building disease detection efficiency and precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of building defect detection technology, and in particular to a method, device, electronic equipment and storage medium for detecting urban building defects. Background Technology

[0002] In densely populated urban areas, the aging of existing buildings and the problem of exterior wall peeling have become serious public safety hazards. Urban building exterior walls mainly include decorated tile walls, glass curtain walls, and undecorated cement mortar plastered walls. Regardless of the type, all exterior walls will gradually age over time, and exterior wall cladding tiles may become hollow and collapse due to defects in the adhesive materials and construction. Hollow spots and cracks in building exterior walls are core risk sources threatening public safety; therefore, exterior wall inspection has become a crucial part of daily building maintenance.

[0003] However, existing drone inspection technology for building exteriors faces two major constraints in practical applications. Firstly, traditional close-range flight inspection methods are no longer suitable for high-density residential areas, forcing drones to operate at a greater, legally permissible safe distance. Long-distance inspection reduces the resolution of visible light images, making it difficult to identify minor defects; it also weakens the ability of thermal infrared sensors to capture internal temperature differences; and telephoto shots are susceptible to camera shake, further increasing the difficulty of accurate identification and positioning. Secondly, older buildings lack a digital foundation. Most of the existing buildings requiring inspection were constructed a long time ago and generally lack Building Information Modeling (BIM) or high-precision as-built drawings. Because older residential areas lack BIM models, existing automated flight path planning processes based on BIM models cannot be directly applied. This makes it difficult for drones to achieve safe autonomous obstacle avoidance and accurate close-fitting data collection in complex urban canyon environments due to the lack of prior 3D environmental information. To improve inspection efficiency and accuracy, drone technology has been introduced into the field of intelligent building exterior inspection, significantly improving inspection efficiency and flexibility, and effectively overcoming the limitations of manual inspection in high-altitude and complex environments. However, traditional detection methods mainly rely on manual visual inspection or single visible light photography, which has the problems of limited collection environment and high risk. In high-density urban areas, the buildings are close together and vegetation is severely obstructed. Traditional drone aerial photography strategies are limited and cannot cover low-level obstructed areas. High-rise operations are inefficient and manual high-altitude operations are risky. At the same time, the ability to identify hidden defects is weak. Single visible light sensors cannot penetrate to identify early hidden problems such as hollowness and leakage inside the wall, which can easily lead to missed detections. Moreover, the identification algorithm has poor anti-interference ability. In complex urban backgrounds, existing general algorithms are difficult to accurately segment small building defects, resulting in a high false alarm rate. In addition, the existing detection results are mostly a large number of two-dimensional photo folders, lacking three-dimensional spatial correlation, and cannot clearly identify the specific floor and location of the defect, making it difficult to directly guide maintenance work. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies, this invention provides a method, device, electronic device, and storage medium for detecting urban building defects. It can achieve accurate synchronization and pixel-level registration of multimodal data, construct a hierarchical deep learning recognition architecture to improve the accuracy of defect recognition in complex backgrounds, and achieve accurate three-dimensional spatial positioning and visualization analysis of defects through reverse ray tracing.

[0005] The first aspect of this application provides a method for detecting urban building defects, the method comprising: Acquire raw multimodal image data; the raw multimodal image data includes raw visible light image data and raw thermal infrared image data, the raw visible light image data is acquired by a visible light camera, and the raw thermal infrared image data is acquired by a thermal infrared camera; Geometric correction is performed on the original thermal infrared image data, and the mapping relationship between the original visible light image data and the original thermal infrared image data is calculated to perform pixel-level registration on the original thermal infrared image data to obtain target multimodal image data. The target multimodal image data is input into a preset hierarchical deep learning recognition model to identify the corresponding building defects information through the hierarchical deep learning recognition model. Multimodal image data containing the building defects information is mapped in three-dimensional space to obtain a three-dimensional mesh model of the building containing the three-dimensional spatial coordinates of the building defects information. Calculate the density relationship between the grid vertices on the surface of the building's 3D mesh model and the point cloud of the disease, and remap the calculated disease density values ​​into a color gradient to generate a heat map of disease distribution covering the surface of the building's 3D mesh model. Based on the heat map of disease distribution, the clustering characteristics of diseases are analyzed from multiple dimensions, and the spatial correlation between the disease information and the building structure nodes is quantitatively analyzed in combination with the building component information.

[0006] In one optional implementation, the hierarchical deep learning recognition model includes a first-level recognition model and a second-level recognition model. The first-level recognition model is constructed using the SegFormer algorithm, and the second-level recognition model is constructed using a combination of K-Net and UPerNet algorithms.

[0007] In an optional implementation, the step of inputting the target multimodal image data into a preset hierarchical deep learning recognition model to identify the corresponding building defect information through the hierarchical deep learning recognition model includes: The target multimodal image data is input into the first-level recognition model, and multi-scale feature maps are extracted by the Transformer encoder and combined with the decoder to generate a binary mask for the wall area and window area. The wall region in the target multimodal image data is cropped according to the binarized mask to generate visible light wall shadow and thermal infrared wall image; The visible light wall image and the thermal infrared wall image are input into the second-level recognition model. The linear features of cracks and the blocky features of detachment in the visible light wall image are captured by the dynamic kernel convolution of K-Net, and the first binary disease mask is output. The low-temperature cold spots of leakage and the high-temperature hot spots of hollowing in the thermal infrared wall image are identified by the multi-scale feature fusion of UPerNet, and the second binary disease mask is output. The first binarized defect mask and the second binarized defect mask are fused to generate a comprehensive defect distribution map that includes the building defect information.

[0008] In an optional implementation, the step of performing geometric correction on the original thermal infrared image data and calculating the mapping relationship between the original visible light image data and the original thermal infrared image data to perform pixel-level registration on the original thermal infrared image data to obtain target multimodal image data includes: Obtain the raw intrinsic parameter data of the thermal infrared camera and calculate the resolution scaling ratio between the thermal infrared image and the visible light image; The original intrinsic parameter data is scaled proportionally according to the resolution scaling ratio to obtain the adjusted effective intrinsic parameter data. Based on the effective intrinsic parameter data, the original thermal infrared image data is geometrically corrected, and the target thermal infrared image data after distortion correction is output. In the original visible light image data, a building corner point is selected as the first feature point, and in the target thermal infrared image data, a corresponding building corner point is selected as the second feature point. Calculate the homography matrix between the original visible light image data and the target thermal infrared image data based on the first feature point and the second feature point; The target thermal infrared image data is transformed by the homography matrix to the visible light coordinate system, thereby achieving pixel-level registration between the original visible light image data and the target thermal infrared image data. The original visible light image data and the target thermal infrared image data after pixel-level registration are integrated into the thermal infrared image data after pixel-level registration.

[0009] In an optional implementation, the step of mapping the multimodal image data containing the building defects information into three-dimensional space to obtain a three-dimensional mesh model of the building containing the three-dimensional spatial coordinates of the building defects information includes: Establish the transformation relationship from WGS84 geographic coordinate system to UTM projected coordinate system, and then to ENU engineering coordinate system; Obtain the camera pose of any one of the visible light camera and the thermal infrared camera, and construct a rotation matrix based on the three-dimensional rotation principle; Radial and tangential distortion corrections are applied to the normalized image coordinates to obtain the corrected image coordinates; Normalized coordinates are calculated based on the corrected image coordinates, principal point coordinates, and given camera focal length, and local ray direction vectors are constructed based on the normalized coordinates. The local ray direction vector is transformed according to the rotation matrix to obtain the ray direction in the ENU engineering coordinate system; A spatial ray is generated with the camera's position in the ENU engineering coordinate system as the starting point and the ray direction as the direction. The intersection points of the spatial ray and the three-dimensional mesh model of the building are calculated, and the building defect information is mapped to three-dimensional space to obtain the three-dimensional mesh model of the building.

[0010] In an optional implementation, the step of calculating the density relationship between the grid vertices on the surface of the building 3D mesh model and the disease point cloud, and remapping the calculated disease density values ​​into a color gradient to generate a disease distribution heatmap covering the surface of the building 3D mesh model includes: The three-dimensional building model is parsed into a triangular mesh structure, and the set of mesh vertices is extracted as nodes for calculating the density of defects. For each grid vertex, a disease point cloud dataset within the neighborhood range corresponding to each grid vertex is retrieved according to a preset search radius; the grid vertex set includes multiple grid vertices, and the disease point cloud dataset includes multiple disease points within the neighborhood. Calculate the distance between each grid vertex and each disease point in its neighborhood, and obtain the disease density value of each grid vertex through mass accumulation operation to obtain a set of density values; The density value set is normalized, and the normalized density value set is mapped to a preset color gradient range to generate a color value set; The color value set is assigned to the vertices of the triangular grid to generate the heat map of the disease distribution.

[0011] In an optional implementation, the multi-dimensional aspect includes a vertical dimension, a horizontal dimension, and an orientation dimension. The step of analyzing disease clustering characteristics from multiple dimensions based on the disease distribution heatmap, and quantitatively analyzing the spatial correlation between the disease information and building structural nodes in conjunction with building component information, includes: When the vertical dimension is selected, the three-dimensional mesh model of the building is divided into the top layer, standard layer and bottom layer according to the building floor height information. The set of disease density values ​​of the corresponding mesh vertices of each layer is extracted. The vertical distribution pattern of diseases is analyzed by statistically analyzing the average density of each layer, and the vertical dimension analysis results are obtained. When the horizontal dimension is selected, the number of defect points of the grid vertex corresponding to each construction node is extracted by combining the semantic information of the building components, and the spatial correlation between the number of defect points and each construction node is calculated to obtain the horizontal dimension analysis results. When the orientation dimension is selected, the grid vertex set is divided according to the orientation of the building facade, and the total number of defects corresponding to each building facade orientation is counted. By comparing the total number of defects, the influence of the building facade orientation on the formation of defects is analyzed, and the analysis results of the orientation dimension are obtained. The vertical dimension analysis results, the horizontal dimension analysis results, and the orientation dimension analysis results are weighted and fused to generate a spatial distribution feature vector of building defects.

[0012] A second aspect of this application provides an urban building defects detection device, the device comprising: The data acquisition module is used to acquire raw multimodal image data; the raw multimodal image data includes raw visible light image data and raw thermal infrared image data, the raw visible light image data is acquired by a visible light camera, and the raw thermal infrared image data is acquired by a thermal infrared camera; The data preprocessing module is used to perform geometric correction on the original thermal infrared image data and calculate the mapping relationship between the original visible light image data and the original thermal infrared image data to perform pixel-level registration on the original thermal infrared image data to obtain target multimodal image data. The defect identification module is used to input the target multimodal image data into a preset hierarchical deep learning identification model, so as to identify the corresponding building defect information through the hierarchical deep learning identification model; The 3D mapping module is used to perform 3D spatial mapping on multimodal image data containing the building defect information to obtain a 3D mesh model of the building containing the 3D spatial coordinates of the building defect information. The remapping module is used to calculate the density relationship between the grid vertices on the surface of the building 3D mesh model and the disease point cloud, and remap the calculated disease density values ​​into color gradients to generate a disease distribution heat map covering the surface of the building 3D mesh model. The visualization analysis module is used to analyze the clustering characteristics of diseases from multiple dimensions based on the disease distribution heat map, and to quantitatively analyze the spatial correlation between the disease information and the building structure nodes by combining the building component information.

[0013] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the urban building defects detection method.

[0014] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method for detecting urban building defects.

[0015] In summary, the urban building defects detection method, device, electronic equipment, and storage medium provided in this application have at least one of the following beneficial effects: 1. Simultaneously acquire raw visible light image data and raw thermal infrared image data. The raw visible light image data visually presents the building's exterior appearance, while the raw thermal infrared image data captures the temperature difference information inside the building. Although long-distance detection will reduce the resolution of the visible light image, combining it with the thermal infrared image data can make up for the shortcomings of a single visible light image in identifying internal hazards. 2. Geometric correction is performed on the original thermal infrared image data, and the mapping relationship between the original visible light image data and the original thermal infrared image data is calculated. Pixel-level registration is then performed on the original thermal infrared image data to obtain multimodal image data of the target. This solves the problems of weakened ability of thermal infrared sensors to capture internal temperature differences during long-distance detection and difficulty in accurate identification and positioning due to camera shake during telephoto shooting. Through registration, the visible light image and the thermal infrared image are spatially correlated, enabling more accurate combination of the two image information for disease identification and improving identification accuracy. 3. Instead of relying on BIM models, the target multimodal image data is input into a pre-defined hierarchical deep learning recognition model, which identifies the corresponding building defects. The deep learning model possesses powerful feature extraction and recognition capabilities, enabling it to automatically learn and identify building defects from multimodal image data without requiring prior 3D environmental information. This solves the problem of drones lacking prior information in complex urban canyon environments, hindering safe autonomous obstacle avoidance and precise data acquisition, and improving the flexibility of inspections. 4. Through three-dimensional spatial mapping, the location of defects in the three-dimensional space of the building can be accurately determined, directly guiding maintenance work and solving the problem that traditional detection results are difficult to use directly for maintenance; the density relationship between the grid vertices on the surface of the three-dimensional mesh model of the building and the point cloud of defects is calculated, and the calculated defect density value is remapped into a color gradient to generate a defect distribution heat map covering the surface of the three-dimensional mesh model of the building. The defect distribution heat map can intuitively show the distribution of defects, and the density of defects is reflected by the depth of color, helping staff to understand the clustering characteristics of defects more clearly and improve the ability to identify and analyze defects; 5. Through multi-dimensional analysis and quantitative analysis, we can gain a deeper understanding of the relationship between defects and building structure, providing more comprehensive and accurate information for building safety assessment and maintenance, and further improving the accuracy and effectiveness of inspections. Attached Figure Description

[0016] Figure 1 This is a schematic flowchart illustrating a method for detecting urban building defects, as shown in an embodiment of this application. Figure 2 This is another schematic flowchart illustrating a method for detecting urban building defects, as shown in an embodiment of this application. Figure 3 This is a schematic diagram illustrating a multimodal data registration process according to an embodiment of this application; Figure 4 This is a schematic diagram illustrating a hierarchical deep learning recognition architecture according to an embodiment of this application; Figure 5 This is a schematic flowchart illustrating a reverse ray tracing principle according to an embodiment of this application; Figure 6 This is a functional block diagram of an urban building defects detection device shown in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application. Detailed Implementation

[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0018] The following will clearly and completely describe the concept, specific structure, and technical effects of the present invention in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention. Furthermore, all connections / linkages involved in the patent do not simply refer to direct contact between components, but rather to the ability to form a better connection structure by adding or reducing connecting accessories according to specific implementation conditions. The various technical features in this invention can be combined interactively without contradicting each other.

[0019] To facilitate understanding of the inventive concept of this application, the following embodiments describe a drone-based intelligent inspection of a high-density, old residential area. This area comprises multiple high-rise residential buildings with exterior facades including tiled walls, painted walls, and glass windows, and the buildings are 15-25 years old. The area is densely wooded, and the distance between buildings is 8-15 meters.

[0020] Reference Figure 1The diagram shown is a flowchart illustrating a method for detecting urban building defects according to an embodiment of this application. The method for detecting urban building defects includes the following steps.

[0021] S11, acquire raw multimodal image data.

[0022] The original multimodal image data includes original visible light image data and original thermal infrared image data. The original visible light image data is acquired by a visible light camera, and the original thermal infrared image data is acquired by a thermal infrared camera.

[0023] Refer to together Figure 2 In some embodiments, the electronic device can pre-build a two-stage UAV operation process of rough model surveying and fine data acquisition. In the first stage, the UAV is controlled to use oblique photography to establish a rough three-dimensional model of the survey area as a digital base for environmental surveying and obstacle avoidance analysis. In the second stage, fine flight path planning is carried out based on the rough model. For multi-story buildings, a close-range interlacing flight strategy is adopted, and for high-rise buildings, a medium-range balanced acquisition strategy is adopted. The visible light camera and the thermal infrared camera are simultaneously controlled to be turned on to acquire multimodal image data.

[0024] In practice, the inspection was conducted using a drone system comprising a DJI M3T industrial-grade drone and its integrated wide-angle visible light camera and FLIR thermal infrared camera. The wide-angle visible light camera has a resolution of 4000×3000 pixels; the FLIR thermal infrared camera has a resolution of 640×512 pixels and a thermal sensitivity of less than or equal to 0.05 degrees Celsius. The drone also incorporates a high-precision RTK positioning system with a positioning accuracy of ±2 centimeters. After the drone system was deployed, in the first phase of rough model surveying, the drone conducted oblique photography at a height of 30 meters above the tallest building in the survey area, with a forward overlap of 80% and a lateral overlap of 70%. Aerial triangulation and texture mapping were then performed using DJI Terra software to generate a rough 3D model of the survey area (digital base), with an accuracy of approximately ±10 centimeters. In the second phase of refined data acquisition, a close-proximity photographic flight path was planned within the 3D view based on the digital base. For multi-story buildings or highly obstructed areas, a close-up photography mode is used, with a sampling distance of 5-10 meters and an overlap rate greater than 80%. For high-rise buildings or open areas, a medium-distance balanced sampling mode is used, with a sampling distance of 20-25 meters and an overlap rate of 70%. Visible light and thermal infrared cameras are activated simultaneously. The sampling time for a single high-rise building is approximately 1.2 hours.

[0025] After two-stage UAV operations, multimodal images containing visible light and thermal infrared images can be collected, which are called raw multimodal image data.

[0026] By constructing a two-stage operation process of "rough model survey - fine data collection" in the complex environment of high-density urban areas, and combining close-range interspersed flight and medium-range balanced data collection strategies, we have effectively overcome the blind spots caused by vegetation obstruction and the bottleneck of high-rise operation efficiency, and achieved full coverage and safe collection of building facade data.

[0027] S12, geometric correction is performed on the original thermal infrared image data, and the mapping relationship between the original visible light image data and the original thermal infrared image data is calculated to perform pixel-level registration on the original thermal infrared image data to obtain target multimodal image data.

[0028] Refer to together Figure 2 In some embodiments, the electronic device establishes a heterogeneous image spatial alignment mechanism, uses the Brown-Conrady distortion model to perform geometric correction on the thermal infrared image to eliminate barrel distortion, calculates the mapping relationship between visible light and thermal infrared images based on the homography matrix to achieve pixel-level registration, and constructs a multimodal dataset containing both texture and temperature information.

[0029] In an optional implementation, the step of performing geometric correction on the original thermal infrared image data and calculating the mapping relationship between the original visible light image data and the original thermal infrared image data to perform pixel-level registration on the original thermal infrared image data to obtain target multimodal image data includes: Obtain the raw intrinsic parameter data of the thermal infrared camera and calculate the resolution scaling ratio between the thermal infrared image and the visible light image; The original intrinsic parameter data is scaled proportionally according to the resolution scaling ratio to obtain the adjusted effective intrinsic parameter data. Based on the effective intrinsic parameter data, the original thermal infrared image data is geometrically corrected, and the target thermal infrared image data after distortion correction is output. In the original visible light image data, a building corner point is selected as the first feature point, and in the target thermal infrared image data, a corresponding building corner point is selected as the second feature point. Calculate the homography matrix between the original visible light image data and the target thermal infrared image data based on the first feature point and the second feature point; The target thermal infrared image data is transformed by the homography matrix to the visible light coordinate system, thereby achieving pixel-level registration between the original visible light image data and the target thermal infrared image data. The original visible light image data and the target thermal infrared image data after pixel-level registration are integrated into the thermal infrared image data after pixel-level registration.

[0030] In some embodiments, for the thermal infrared image in the raw thermal infrared image data, the electronic device can acquire raw intrinsic parameter data from the thermal infrared camera, including focal length fx, fy, and principal point coordinates. The system obtains the distortion coefficients k1, k2, p1, and p2, and the resolution information of the thermal infrared and visible light images. Next, it calculates the resolution scaling ratio between the thermal infrared and visible light images, and proportionally scales the original intrinsic parameter data of the thermal infrared camera to obtain the adjusted effective intrinsic parameter data. Then, the electronic device uses the Brown-Conrady distortion model to perform geometric correction on the thermal infrared image based on the adjusted effective intrinsic parameter data, eliminating barrel distortion in the thermal infrared image and outputting the distortion-corrected thermal infrared image (also called the target thermal infrared image). For the visible light image in the original visible light image data, the electronic device can select building corner points (at least 4 pairs) as feature points (referred to as first feature points for easy distinction), and simultaneously select the corresponding building corner points as feature points (referred to as second feature points) in the distortion-corrected thermal infrared image. Based on these feature points, the homography matrix H between the visible light and thermal infrared images is calculated. Next, the thermal infrared image is transformed using the calculated homography matrix to the visible light coordinate system, achieving pixel-level registration between the visible light and thermal infrared images. In other words, the thermal infrared image is transformed into the visible light coordinate system, and the thermal infrared image and the visible light image are in the same coordinate system, achieving a pixel-level registration error of less than 2 pixels.

[0031] Furthermore, the electronic device fuses the original visible light image data (containing texture information) with the pixel-level registered thermal infrared image data (containing temperature information), that is, the visible light image and the registered thermal infrared image are superimposed by channel to generate a multimodal image dataset containing both texture and temperature information, also known as target multimodal image data.

[0032] It should be noted that the principal point coordinates are the position of the origin of the imaging plane coordinate system in the pixel coordinate system. The origin of the imaging plane coordinate system is the intersection of the optical axis and the imaging plane, with the x-axis to the right and the y-axis downwards; the origin of the pixel coordinate system is the upper left corner of the image, with the x-axis to the right and the y-axis downwards.

[0033] Through the above optional implementation methods, by fusing visible light texture and thermal infrared temperature field data, and using Brown-Conrady distortion correction and homography matrix registration to achieve pixel-level alignment, collaborative identification of four typical defects—cracks, peeling, hollowing, and leakage—is achieved, filling the technical blind spot where traditional single sensors cannot detect hidden defects inside the wall coating.

[0034] S13, the target multimodal image data is input into a preset hierarchical deep learning recognition model to identify the corresponding building defects information through the hierarchical deep learning recognition model.

[0035] The information on defects may include, but is not limited to, cracks, peeling, leakage, and hollow areas.

[0036] Refer to together Figure 2 In some embodiments, a hierarchical deep learning recognition architecture of "macroscopic component extraction - microscopic defect segmentation" is constructed. This hierarchical deep learning recognition model includes a first-level recognition model and a second-level recognition model. The first-level recognition model uses the SegFormer algorithm to train a wall extraction model to accurately separate wall and window areas from complex urban backgrounds, removing interference from the sky and vegetation. The second-level recognition model uses a combination of K-Net and UPerNet algorithms within the extracted wall area to identify cracks and peeling in visible light images, and leaks and hollow areas in thermal infrared images. The training and implementation process of this hierarchical deep learning recognition model is described in detail below: (1) Training of the first-level recognition model.

[0037] The electronic device first collects a dataset of complex urban background images containing walls, windows, sky, and vegetation, and labels the wall and window regions in the images to generate corresponding binary mask labels. Using the SegFormer algorithm, the collected image dataset and its corresponding binary mask labels are used as training data to train the SegFormer model, enabling it to accurately output binary masks of walls and windows from complex urban backgrounds, thereby defining regions of interest. During training, conventional deep learning training strategies are employed, such as using a stochastic gradient descent (SGD) optimizer, setting appropriate hyperparameters such as learning rate and batch size, and continuously adjusting model parameters through backpropagation until the model's performance metrics (such as Intersection over Union (IoU) on the validation set reach the preset requirements.

[0038] (2) Training of the second-level recognition model.

[0039] Simultaneously, the electronic device first collects image datasets containing cracks and detachments from visible light images, and image datasets containing leaks and voids from thermal infrared images. Cracked areas, detached areas in the visible light images, and leaked and voided areas in the thermal infrared images are labeled. A combined K-Net and UperNet algorithm architecture is constructed, and ADE20K pre-trained weights are loaded using transfer learning to accelerate model convergence and improve initial performance. For the collected visible light and thermal infrared image datasets, data augmentation strategies such as rotation, cropping, and color dithering are employed to expand the dataset size and improve the model's robustness under complex lighting conditions. The K-Net and UperNet combined model is trained using the augmented visible light and thermal infrared image datasets and their corresponding labels as training data. During training, for the visible light image branch, cracks and detachments are identified by capturing linear features; for the thermal infrared image branch, leaks are determined by identifying low-temperature cold spot areas, and voids are determined by identifying high-temperature hot spot areas. Using conventional deep learning training strategies, such as the Adam optimizer, setting appropriate hyperparameters such as learning rate and batch size, and continuously adjusting model parameters through backpropagation until the model's performance metrics (such as accuracy, recall, F1 score, etc.) on the validation set reach the preset requirements.

[0040] Once the hierarchical deep learning recognition model is built and trained, the electronic device can input the target multimodal image data obtained in step S12 into the hierarchical deep learning recognition model to initiate the automated identification process of building defects. This process strictly follows the hierarchical processing logic and makes full use of the characteristics of multimodal data to achieve efficient and accurate defect detection.

[0041] In an optional implementation, the step of inputting the target multimodal image data into a preset hierarchical deep learning recognition model to identify the corresponding building defect information through the hierarchical deep learning recognition model includes: The target multimodal image data is input into the first-level recognition model, and multi-scale feature maps are extracted by the Transformer encoder and combined with the decoder to generate a binary mask for the wall area and window area. The wall region in the target multimodal image data is cropped according to the binarized mask to generate visible light wall shadow and thermal infrared wall image; The visible light wall image and the thermal infrared wall image are input into the second-level recognition model. The linear features of cracks and the blocky features of detachment in the visible light wall image are captured by the dynamic kernel convolution of K-Net, and the first binary disease mask is output. The low-temperature cold spots of leakage and the high-temperature hot spots of hollowing in the thermal infrared wall image are identified by the multi-scale feature fusion of UPerNet, and the second binary disease mask is output. The first binarized defect mask and the second binarized defect mask are fused to generate a comprehensive defect distribution map that includes the building defect information.

[0042] Refer to together Figure 3 In some embodiments, in the first-level recognition model, where the first-level recognition model adopts the SegFormer-B5 model, when the target multimodal image data is input, a multi-scale feature map is extracted by the Transformer encoder, and a binary mask of the wall and window areas is generated by the decoder. (That is, the building body mask), with a wall extraction accuracy of no less than 95% (verified on the test set), achieving automatic removal of complex urban backgrounds (including sky, vegetation, and vehicles). Next, using... To mask the wall regions in the target multimodal image data, the main building area is located. Next, K-Net uses dynamic convolutional kernels to capture local features of the wall edges, enhancing its adaptability to irregular wall shapes. UpperNet integrates multi-level features (such as shallow textures and deep semantics) to improve the segmentation accuracy of the wall regions, ultimately outputting a binary mask of the wall regions (i.e., the wall mask), further narrowing the analysis scope to the wall surface.

[0043] Furthermore, after accurately locating the wall area, the second-level recognition model conducts targeted building defect detection based on multimodal image data. It fully utilizes the different characteristics of visible light and thermal infrared images to achieve comprehensive identification and precise location of defects such as cracks, peeling, leakage, and hollow areas. The specific detection process is as follows: (1) Visible light channel: detect cracks and detachment.

[0044] Wall masks are applied to visible light images to extract sub-images of wall areas. Linear features of cracks (width ≥ 1 cm) and blocky features of detached tissue (area ≥ 10 cm²) are captured using dynamic kernel convolution of K-Net, outputting a binarized disease mask. and To facilitate differentiation, a binary disease mask is used. and This is called the first binary disease mask.

[0045] (2) Thermal infrared channel: detects leakage and hollowness.

[0046] Wall masks are applied to thermal infrared images to extract sub-images of wall areas. Multi-scale feature fusion using UPerNet is used to identify low-temperature cold spots indicating leakage and high-temperature hot spots indicating hollow areas (temperature difference ≥ 3℃), outputting a binarized defect mask. and To facilitate differentiation, a binary disease mask is used. and This is called the second binarized disease mask.

[0047] Finally, the identification results from the visible light and thermal infrared channels are weighted and fused to generate a comprehensive defect distribution map that includes building defect information (i.e., any one or a combination of defect types such as cracks, peeling, leakage, and hollowing). ; The weighting coefficients α and β are dynamically adjusted based on the modal confidence level (α+β=1).

[0048] In addition to defect types, the comprehensive defect distribution map also includes the coordinates corresponding to those defect types. By spatially aligning the defect types and coordinates of the visible light and thermal infrared branches, positional consistency is ensured through coordinate mapping or homography transformation. Furthermore, non-maximum suppression (NMS) is used to merge overlapping detection boxes, avoiding duplicate marking.

[0049] In addition, the electronic device can generate a final inspection report, which includes the location coordinates of the defect type (crack, detachment) and confidence score.

[0050] Through the above optional implementation methods, in order to achieve high-precision identification in complex backgrounds, a hierarchical deep learning model system of "macro-component extraction - micro-disease segmentation" was established. The SegFormer algorithm was used to first extract the wall area to remove interference from the sky and trees, and then the K-Net and UPerNet combined algorithm was used to capture fine features, which greatly improved the identification accuracy of micro-cracks, peeling, hollowing and leakage in complex urban backgrounds.

[0051] S14, perform three-dimensional spatial mapping on the multimodal image data containing the building defect information to obtain a three-dimensional mesh model of the building containing the three-dimensional spatial coordinates of the building defect information.

[0052] Refer to together Figure 2 In some embodiments, after identifying the disease information and obtaining multimodal image data containing the disease information, the electronic device can achieve accurate three-dimensional spatial mapping of the disease information through a reverse ray tracing algorithm. This involves establishing a transformation from the WGS84 geographic coordinate system to the UTM projection coordinate system and then to the ENU engineering coordinate system. Using the pinhole imaging principle and camera pose parameters, a spatial ray is constructed that starts from the camera optical center and passes through the disease pixels. The intersection of this ray with the three-dimensional mesh model of the building is then calculated, thereby achieving accurate positioning of the disease information in three-dimensional space.

[0053] In an optional implementation, the step of mapping the multimodal image data containing the building defects information into three-dimensional space to obtain a three-dimensional mesh model of the building containing the three-dimensional spatial coordinates of the building defects information includes: Establish the transformation relationship from WGS84 geographic coordinate system to UTM projected coordinate system, and then to ENU engineering coordinate system; Obtain the camera pose of any one of the visible light camera and the thermal infrared camera, and construct a rotation matrix based on the three-dimensional rotation principle; Radial and tangential distortion corrections are applied to the normalized image coordinates to obtain the corrected image coordinates; Normalized coordinates are calculated based on the corrected image coordinates, principal point coordinates, and given camera focal length, and local ray direction vectors are constructed based on the normalized coordinates. The local ray direction vector is transformed according to the rotation matrix to obtain the ray direction in the ENU engineering coordinate system; A spatial ray is generated with the camera's position in the ENU engineering coordinate system as the starting point and the ray direction as the direction. The intersection points of the spatial ray and the three-dimensional mesh model of the building are calculated, and the building defect information is mapped to three-dimensional space to obtain the three-dimensional mesh model of the building.

[0054] Refer to together Figure 5 In some embodiments, the electronic device establishes a coordinate system transformation chain, projecting the WGS84 latitude and longitude coordinates of the UAV POS data to the UTM coordinate system, and establishing a local ENU coordinate system based on the reference points in the SfM reconstruction report, thus establishing a transformation chain from WGS84 to UTM to local ENU. At the end of this chain, the local ENU coordinate system is set as the final computing platform to support the building's 3D mesh model. Next, the camera's attitude Euler angles are read to construct a rotation matrix R, and the camera's main line-of-sight direction vector is calculated. After correcting the radial and tangential distortion of the defect pixel coordinates, normalized coordinates are calculated based on the pinhole imaging model to construct a local ray direction vector. The ray direction in the ENU engineering coordinate system is obtained through rotation matrix transformation. The intersection points of the rays and the building's 3D mesh model (Mesh, containing 2.56 million triangular faces) are calculated. A total of 49 defects were mapped in this inspection.

[0055] Specifically, the electronic device first establishes a transformation relationship from the WGS84 geographic coordinate system to the UTM projected coordinate system and then to the ENU engineering coordinate system. That is, it converts (λ, φ, h) to (E, N, U) using the UTM projection formula (such as the transverse Mercator projection), where E is the eastward coordinate, N is the northward coordinate, and U is the elevation. Next, a point in the UTM coordinate system is selected. As the origin of ENU, the directions of the ENU axes are defined, where the E-axis is due east (consistent with the E-axis of UTM), the N-axis is due north (consistent with the N-axis of UTM), and the U-axis is perpendicular to the EN plane and points upwards (consistent with the U-axis of UTM). Then, the transformation formula is used... , , Achieve the conversion.

[0056] Simultaneously, the camera pose (represented by Euler angles) recorded in the SfM reconstruction report is read, and a rotation matrix R is constructed to describe the relationship between the camera coordinate system and the ENU engineering coordinate system.

[0057] Next, the normalized image coordinates are corrected using radial and tangential distortion correction methods to obtain the corrected image coordinates. The core of normalized image coordinates is projecting 3D spatial points onto the normalized plane (Z=1) of the camera coordinate system, achieving scale uniformity of the coordinates by eliminating the influence of depth (Z value). Specifically, for spatial points in the real world... Its corresponding camera coordinate system coordinates are The corresponding image coordinate system coordinates are The corresponding pixel coordinates are The imaging plane coordinates are determined by the principal point coordinates. Replace with pixel coordinates (u, v). , dx and dy are the physical dimensions of a pixel (unit: millimeters / pixel). For a real-world spatial point... The coordinates are transformed to the camera coordinate system using camera extrinsic parameters (rotation matrix R and translation vector t) to obtain... : Next, the points in the camera coordinate system Projecting onto the normalized plane (Z=1) yields normalized coordinates. ,in , At this point, the third dimension of the normalized coordinates is 1, that is... .

[0058] Furthermore, based on the distortion-corrected image coordinates (also known as pixel coordinates) and principal point coordinates... Calculate the normalized coordinates with the focal length f to obtain the local ray direction vector. ,in These are normalized coordinates.

[0059] Then, the local ray direction vector is obtained through the camera pose matrix R. The ray direction in the ENU engineering coordinate system is obtained by transformation. The position of the camera in the ENU engineering coordinate system. Starting from, with Rays are generated for the direction, and finally the defect information in the two-dimensional image is mapped to the three-dimensional space through the above rays to obtain a three-dimensional mesh model of the building containing the three-dimensional spatial coordinates of the building defect information.

[0060] The above step S14 is completed in the visualization tool, which is developed based on the Rhino / Grasshopper platform. After the user reads the target image and file, the program automatically completes a series of processes such as coordinate transformation, attitude calculation, distortion correction and ray generation, and outputs the position point of the camera in three-dimensional space, the camera geometric model and the ray set corresponding to the labeled area.

[0061] Through the above optional implementation methods, a mapping mechanism based on reverse ray tracing is proposed to achieve accurate three-dimensional spatial tracing and positioning. It only requires a 2D photo and a rough model. By using reverse ray tracing to establish a projection chain, the amount of computation is greatly reduced, the requirements for hardware equipment are lowered, and the computational efficiency is improved. It breaks down the conversion barrier between two-dimensional image pixels and three-dimensional geographic space, and automatically converts discrete detection images into three-dimensional defect point clouds attached to the surface of the building model, directly answering the question of where the defect is located and providing maintenance personnel with precise spatial coordinate guidance.

[0062] S15, calculate the density relationship between the grid vertices on the surface of the building 3D mesh model and the disease point cloud, and remap the calculated disease density value into a color gradient to generate a disease distribution heat map covering the surface of the building 3D mesh model.

[0063] Refer to together Figure 2 In some embodiments, after the identification and location of building defects are completed, in order to more intuitively display the distribution of defects on the building surface, electronic devices can develop automated mapping tools based on a parametric design platform, use point gravity algorithm to calculate the density relationship between the grid vertices on the building model surface and the defect point cloud, and remap the density values ​​into color gradients to generate a defect distribution heatmap covering the model surface.

[0064] The parametric visualization tool, also developed on the Rhino / Grasshopper platform, constructs a 3D mesh model of the building using 3D scanning technology (such as LiDAR scanning) or based on design drawings, ensuring the model accurately reflects the building's geometry and surface features. Next, the building defects identified in step S14 (including the location coordinates of defects such as cracks, peeling, hollow areas, and leaks) are converted into point cloud data. Each defect point contains its coordinate information in 3D space. Then, the Pull Point algorithm is used to treat each defect point as an object with a certain mass, and the mesh vertices as points subject to gravity. The magnitude of the gravitational force exerted by a defect point on surrounding mesh vertices is inversely proportional to the distance between the defect point and the mesh vertex, and directly proportional to the "mass" of the defect point (which can be understood as the severity or weight of the defect). For each mesh vertex, the sum of the gravitational forces exerted on it by all defect points is calculated. The sum of the gravitational forces reflects the density of defects around that mesh vertex, i.e., the density relationship. In the specific calculation process, appropriate gravity formulas and related parameters are set; for example, the gravity formula can be expressed as... Where F is the magnitude of gravity and G is the gravitational constant. and Here, r represents the "mass" of the defect point and the mesh vertex (which can be set according to the actual situation), respectively, and r is the distance between them. The algorithm iterates through all surface mesh vertices of the 3D building mesh model, calculates the total gravitational force on each vertex, and obtains the density value corresponding to each mesh vertex.

[0065] Furthermore, based on the calculated density value range, the color mapping interval is determined. For example, the minimum density value is mapped to blue, the maximum density value to red, and intermediate density values ​​are mapped to a transitional color between blue and red using linear interpolation. Simultaneously, a correspondence table between density values ​​and colors is established based on the color mapping range. In the automated mapping tool of the parametric design platform, density values ​​can be converted into corresponding color values ​​according to set rules through programming. Based on the density value of each mesh vertex, the corresponding color value is looked up in the color mapping table and assigned to the corresponding mesh vertex. Finally, in the parametric design platform, the colored mesh vertex information is used to render the 3D building mesh model. After rendering, the surface of the 3D building mesh model will exhibit different color distributions based on the colors of the mesh vertices, forming a heatmap of the disease distribution. Darker areas in the heatmap indicate higher disease density, while lighter areas indicate lower disease density. By constructing a parametric three-dimensional visualization analysis system, it is possible to generate multi-dimensional heat maps of disease distribution, which intuitively reveal the clustering characteristics of diseases in vertical, horizontal and directional dimensions, providing scientific data support for the preventive maintenance, repair site selection and cause analysis of existing buildings.

[0066] The above-mentioned method of topological correlation analysis between defects and components based on "point gravitational field" not only focuses on "where the building is broken", but also uses the Pull Point algorithm to calculate the topological distance density from the defect point to building components such as "air conditioner unit / window corner", thereby automatically deducing the cause of the defect.

[0067] S16. Based on the heat map of disease distribution, analyze the disease clustering characteristics from multiple dimensions, and combine the information of building components to quantitatively analyze the spatial correlation between the disease information and the building structure nodes.

[0068] The multi-dimensional approach includes vertical, horizontal, and orientation dimensions. Vertically, the density of defects on the top, standard, and bottom floors is statistically analyzed to determine the vertical distribution patterns. Horizontally, combined with component information, the number of defects at structural nodes such as air conditioner locations, window corners, and balconies is statistically analyzed to reveal the correlation mechanism between defects and their location. Orientation-wise, the distribution of defects on the east, south, west, and north facades is statistically analyzed to understand the impact of sunlight and wind / rain on defect formation.

[0069] Refer to together Figure 2 Electronic devices can perform multi-dimensional spatial distribution feature analysis based on the generated three-dimensional heat map, analyze the clustering characteristics of diseases from the vertical, horizontal and orientation dimensions, and combine building component information to quantitatively analyze the spatial correlation between diseases and building structural nodes such as air conditioning unit positions, window corners, and balconies.

[0070] In an optional implementation, the multi-dimensional aspect includes a vertical dimension, a horizontal dimension, and an orientation dimension. The step of analyzing disease clustering characteristics from multiple dimensions based on the disease distribution heatmap, and quantitatively analyzing the spatial correlation between the disease information and building structural nodes in conjunction with building component information, includes: When the vertical dimension is selected, the three-dimensional mesh model of the building is divided into the top layer, standard layer and bottom layer according to the building floor height information. The set of disease density values ​​of the corresponding mesh vertices of each layer is extracted. The vertical distribution pattern of diseases is analyzed by statistically analyzing the average density of each layer, and the vertical dimension analysis results are obtained. When the horizontal dimension is selected, the number of defect points of the grid vertex corresponding to each construction node is extracted by combining the semantic information of the building components, and the spatial correlation between the number of defect points and each construction node is calculated to obtain the horizontal dimension analysis results. When the orientation dimension is selected, the grid vertex set is divided according to the orientation of the building facade, and the total number of defects corresponding to each building facade orientation is counted. By comparing the total number of defects, the influence of the building facade orientation on the formation of defects is analyzed, and the analysis results of the orientation dimension are obtained. The vertical dimension analysis results, the horizontal dimension analysis results, and the orientation dimension analysis results are weighted and fused to generate a spatial distribution feature vector of building defects.

[0071] In some embodiments, to delve deeper into the potential relationships between the analysis results of various dimensions and to comprehensively characterize the spatial distribution features of building defects, the specific steps are as follows: (1) Vertical dimension analysis: Based on the building's floor height information, the model is divided into the top floor, standard floor, and bottom floor, and the set of disease density values ​​for the corresponding grid vertices of each floor is extracted. By statistically analyzing the average density of each layer ( Analyze the vertical distribution pattern of diseases and generate a vertical density gradient curve.

[0072] (2) Horizontal dimension analysis: By combining semantic information of building components (such as air conditioner unit locations, window corners, and balconies), the number of defect points corresponding to the grid vertices of each structural node is extracted. (k is the component type), calculate the spatial correlation between defects and components. : ; in, The surface area of ​​the component. Generate a heat map showing the correlation between defects and components, using the component exposure coefficient (value range [0,1]).

[0073] Component exposure factor Determined by a weighted combination of component orientation and height, satisfying: ; in, Let be the angle between the component's normal and the horizontal plane. For the height of the component, The total building height is denoted by α, and β are weighting coefficients.

[0074] (3) Orientation dimension analysis: The grid vertex set is divided according to the building facade orientation (east, south, west, north). ( ), count the total number of disease points in each orientation (τ is the density threshold, dynamically adjusted to 1.5 times the global density mean), through comparison Analyze the differences in the effects of sunlight and wind and rain on disease formation.

[0075] Furthermore, the analysis results of vertical, horizontal, and oriented dimensions are weighted and fused to generate a spatial distribution feature vector F of building defects (also known as multi-dimensional analysis results). ,in , , .

[0076] Among them, the multi-dimensional analysis results show that the density of defects in the top layer is twice that of the standard layer in the vertical dimension (mainly cracks, peeling, leakage and hollowing caused by thermal stress); and in the horizontal dimension, 70% of the leakage is concentrated below the air conditioning unit.

[0077] Through steps S11 to S16 above, the identification results of this inspection are as follows: 16 cracks; 3 peelings; 22 leaks; and 8 hollow areas. The total time taken was 2.5 days (including 1 day for data collection and 1.5 days for processing and analysis), which is about 8 times more efficient than traditional manual inspection (estimated to take 3 weeks). The generated 3D heat map of the defects can be directly imported into the property management system to provide precise coordinate guidance for repair location.

[0078] Reference Figure 6 The diagram shown is a functional block diagram of an urban building defect detection device according to an embodiment of this application.

[0079] In some embodiments, the urban building defects detection device 60 may include multiple functional modules composed of computer program segments. The computer programs for each program segment of the urban building defects detection device 60 may be stored in the memory of an electronic device and executed by at least one processor to perform (see details). Figure 1 (Description) The function of urban building defect detection. Based on its functions, it can be divided into multiple functional modules. These modules may include: a data acquisition module 601, a data preprocessing module 602, a defect identification module 603, a 3D mapping module 604, a remapping module 605, and a visualization analysis module 606. The module referred to in this application is a series of computer program segments that can be executed by at least one processor and perform a fixed function, stored in memory. In this embodiment, the functions of each module will be detailed in subsequent embodiments.

[0080] The data acquisition module 601 is used to acquire raw multimodal image data; the raw multimodal image data includes raw visible light image data and raw thermal infrared image data, the raw visible light image data is acquired by a visible light camera, and the raw thermal infrared image data is acquired by a thermal infrared camera.

[0081] The data preprocessing module 602 is used to perform geometric correction on the original thermal infrared image data and calculate the mapping relationship between the original visible light image data and the original thermal infrared image data to perform pixel-level registration on the original thermal infrared image data to obtain target multimodal image data.

[0082] The defect identification module 603 is used to input the target multimodal image data into a preset hierarchical deep learning identification model, so as to identify the corresponding building defect information through the hierarchical deep learning identification model.

[0083] The three-dimensional mapping module 604 is used to perform three-dimensional spatial mapping on the multimodal image data containing the building defect information to obtain a three-dimensional mesh model of the building containing the three-dimensional spatial coordinates of the building defect information.

[0084] The remapping module 605 is used to calculate the density relationship between the grid vertices on the surface of the building 3D mesh model and the disease point cloud, and remap the calculated disease density value into a color gradient to generate a disease distribution heat map covering the surface of the building 3D mesh model.

[0085] The visualization analysis module 606 is used to analyze the clustering characteristics of diseases from multiple dimensions based on the disease distribution heat map, and to quantitatively analyze the spatial correlation between the disease information and the building structure nodes by combining the building component information.

[0086] It should be understood that the various variations and specific embodiments of the urban building defects detection method provided in the above embodiments are also applicable to the urban building defects detection device of this embodiment. Through the foregoing detailed description of the urban building defects detection method, those skilled in the art can clearly understand the implementation method of the urban building defects detection device of this embodiment. For the sake of brevity, it will not be described in detail here.

[0087] See Figure 7 The diagram shown is a schematic representation of the structure of an electronic device according to an embodiment of this application. In a preferred embodiment of this application, the electronic device 7 includes a memory 71, at least one processor 72, and at least one communication bus 73.

[0088] Those skilled in the art should understand that Figure 7 The structure of the electronic device shown does not constitute a limitation of the embodiments of this application. It can be a bus structure or a star structure. The electronic device 7 may also include more or fewer other hardware or software than shown, or different component arrangements.

[0089] In some embodiments, the electronic device 7 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), digital processors, and embedded devices. The electronic device 7 may also include user equipment, which includes, but is not limited to, any electronic product capable of human-computer interaction with a user via a keyboard, mouse, remote control, touchpad, or voice control device, such as personal computers, tablets, smartphones, digital cameras, and drones.

[0090] In the embodiments provided in this application, it should be understood that the disclosed methods, apparatuses, computer-readable storage media, and electronic devices can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple components or modules may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices, components, or modules may be electrical, mechanical, or other forms.

[0091] The components described as separate parts may or may not be physically separate. The components shown as components may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the components can be selected to achieve the purpose of this embodiment according to actual needs.

[0092] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each component can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0093] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0094] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0095] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0096] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A method for detecting urban building defects, characterized in that, The method includes: Acquire raw multimodal image data; the raw multimodal image data includes raw visible light image data and raw thermal infrared image data, the raw visible light image data is acquired by a visible light camera, and the raw thermal infrared image data is acquired by a thermal infrared camera; Geometric correction is performed on the original thermal infrared image data, and the mapping relationship between the original visible light image data and the original thermal infrared image data is calculated to perform pixel-level registration on the original thermal infrared image data to obtain target multimodal image data. The target multimodal image data is input into a preset hierarchical deep learning recognition model to identify the corresponding building defects information through the hierarchical deep learning recognition model. Multimodal image data containing the building defects information is mapped in three-dimensional space to obtain a three-dimensional mesh model of the building containing the three-dimensional spatial coordinates of the building defects information. Calculate the density relationship between the grid vertices on the surface of the building's 3D mesh model and the point cloud of the disease, and remap the calculated disease density values ​​into a color gradient to generate a heat map of disease distribution covering the surface of the building's 3D mesh model. Based on the heat map of disease distribution, the clustering characteristics of diseases are analyzed from multiple dimensions, and the spatial correlation between the disease information and the building structure nodes is quantitatively analyzed in combination with the building component information.

2. The method for detecting urban building defects according to claim 1, characterized in that, The hierarchical deep learning recognition model includes a first-level recognition model and a second-level recognition model. The first-level recognition model is constructed using the SegFormer algorithm, and the second-level recognition model is constructed using a combination of K-Net and UPerNet algorithms.

3. The method for detecting urban building defects according to claim 2, characterized in that, The step of inputting the target multimodal image data into a preset hierarchical deep learning recognition model, so as to identify the corresponding building defect information through the hierarchical deep learning recognition model, includes: The target multimodal image data is input into the first-level recognition model, and multi-scale feature maps are extracted by the Transformer encoder and combined with the decoder to generate a binary mask for the wall area and window area. The wall region in the target multimodal image data is cropped according to the binarized mask to generate visible light wall shadow and thermal infrared wall image; The visible light wall image and the thermal infrared wall image are input into the second-level recognition model. The linear features of cracks and the blocky features of detachment in the visible light wall image are captured by the dynamic kernel convolution of K-Net, and the first binary disease mask is output. The low-temperature cold spots of leakage and the high-temperature hot spots of hollowing in the thermal infrared wall image are identified by the multi-scale feature fusion of UPerNet, and the second binary disease mask is output. The first binarized defect mask and the second binarized defect mask are fused to generate a comprehensive defect distribution map that includes the building defect information.

4. The method for detecting urban building defects according to claim 1, characterized in that, The step of performing geometric correction on the original thermal infrared image data and calculating the mapping relationship between the original visible light image data and the original thermal infrared image data to perform pixel-level registration on the original thermal infrared image data to obtain target multimodal image data includes: Obtain the raw intrinsic parameter data of the thermal infrared camera and calculate the resolution scaling ratio between the thermal infrared image and the visible light image; The original intrinsic parameter data is scaled proportionally according to the resolution scaling ratio to obtain the adjusted effective intrinsic parameter data. Based on the effective intrinsic parameter data, the original thermal infrared image data is geometrically corrected, and the target thermal infrared image data after distortion correction is output. In the original visible light image data, a building corner point is selected as the first feature point, and in the target thermal infrared image data, a corresponding building corner point is selected as the second feature point. Calculate the homography matrix between the original visible light image data and the target thermal infrared image data based on the first feature point and the second feature point; The target thermal infrared image data is transformed by the homography matrix to the visible light coordinate system, thereby achieving pixel-level registration between the original visible light image data and the target thermal infrared image data. The original visible light image data and the target thermal infrared image data after pixel-level registration are integrated into the thermal infrared image data after pixel-level registration.

5. The method for detecting urban building defects according to claim 1, characterized in that, The step of mapping the multimodal image data containing the building defects information into three-dimensional space to obtain a three-dimensional mesh model of the building containing the three-dimensional spatial coordinates of the building defects information includes: Establish the transformation relationship from WGS84 geographic coordinate system to UTM projected coordinate system, and then to ENU engineering coordinate system; Obtain the camera pose of any one of the visible light camera and the thermal infrared camera, and construct a rotation matrix based on the three-dimensional rotation principle; Radial and tangential distortion corrections are applied to the normalized image coordinates to obtain the corrected image coordinates; Normalized coordinates are calculated based on the corrected image coordinates, principal point coordinates, and given camera focal length, and local ray direction vectors are constructed based on the normalized coordinates. The local ray direction vector is transformed according to the rotation matrix to obtain the ray direction in the ENU engineering coordinate system; A spatial ray is generated with the camera's position in the ENU engineering coordinate system as the starting point and the ray direction as the direction. The intersection points of the spatial ray and the three-dimensional mesh model of the building are calculated, and the building defect information is mapped to three-dimensional space to obtain the three-dimensional mesh model of the building.

6. The method for detecting urban building defects according to claim 1, characterized in that, The step of calculating the density relationship between the grid vertices on the surface of the building's 3D mesh model and the point cloud of the disease, and remapping the calculated disease density values ​​into a color gradient to generate a heat map of disease distribution covering the surface of the building's 3D mesh model includes: The three-dimensional building model is parsed into a triangular mesh structure, and the set of mesh vertices is extracted as nodes for calculating the density of defects. For each grid vertex, a disease point cloud dataset within the neighborhood range corresponding to each grid vertex is retrieved according to a preset search radius; the grid vertex set includes multiple grid vertices, and the disease point cloud dataset includes multiple disease points within the neighborhood. Calculate the distance between each grid vertex and each disease point in its neighborhood, and obtain the disease density value of each grid vertex through mass accumulation operation to obtain a set of density values; The density value set is normalized, and the normalized density value set is mapped to a preset color gradient range to generate a color value set; The color value set is assigned to the vertices of the triangular grid to generate the heat map of the disease distribution.

7. The method for detecting urban building defects according to claim 1, characterized in that, The multi-dimensional analysis includes vertical, horizontal, and orientation dimensions. The analysis of disease clustering characteristics from multiple dimensions based on the disease distribution heatmap, combined with quantitative analysis of the spatial correlation between disease information and building structural nodes using building component information, includes: When the vertical dimension is selected, the three-dimensional mesh model of the building is divided into the top layer, standard layer and bottom layer according to the building floor height information. The set of disease density values ​​of the corresponding mesh vertices of each layer is extracted. The vertical distribution pattern of diseases is analyzed by statistically analyzing the average density of each layer, and the vertical dimension analysis results are obtained. When the horizontal dimension is selected, the number of defect points of the grid vertex corresponding to each construction node is extracted by combining the semantic information of the building components, and the spatial correlation between the number of defect points and each construction node is calculated to obtain the horizontal dimension analysis results. When the orientation dimension is selected, the grid vertex set is divided according to the orientation of the building facade, and the total number of defects corresponding to each building facade orientation is counted. By comparing the total number of defects, the influence of the building facade orientation on the formation of defects is analyzed, and the analysis results of the orientation dimension are obtained. The vertical dimension analysis results, the horizontal dimension analysis results, and the orientation dimension analysis results are weighted and fused to generate a spatial distribution feature vector of building defects.

8. A device for detecting urban building defects, characterized in that, The device includes: The data acquisition module is used to acquire raw multimodal image data; the raw multimodal image data includes raw visible light image data and raw thermal infrared image data, the raw visible light image data is acquired by a visible light camera, and the raw thermal infrared image data is acquired by a thermal infrared camera; The data preprocessing module is used to perform geometric correction on the original thermal infrared image data and calculate the mapping relationship between the original visible light image data and the original thermal infrared image data to perform pixel-level registration on the original thermal infrared image data to obtain target multimodal image data. The defect identification module is used to input the target multimodal image data into a preset hierarchical deep learning identification model, so as to identify the corresponding building defect information through the hierarchical deep learning identification model; The 3D mapping module is used to perform 3D spatial mapping on multimodal image data containing the building defect information to obtain a 3D mesh model of the building containing the 3D spatial coordinates of the building defect information. The remapping module is used to calculate the density relationship between the grid vertices on the surface of the building 3D mesh model and the disease point cloud, and remap the calculated disease density values ​​into color gradients to generate a disease distribution heat map covering the surface of the building 3D mesh model. The visualization analysis module is used to analyze the clustering characteristics of diseases from multiple dimensions based on the disease distribution heat map, and to quantitatively analyze the spatial correlation between the disease information and the building structure nodes by combining the building component information.

9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the urban building defects detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the urban building defects detection method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method and device for monitoring and demonstrating weathering of immovable cultural relics and storage medium

    CN122156503A

  • Methods, devices and storage media for weathering monitoring and demonstration of immovable cultural relics

    CN122156503B