Holographic reconstruction system and method for building apparent disease based on unmanned aerial vehicle vision
Patent Information
- Application Number
- CN202610735109.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-08-18
AI Technical Summary
[0006]本发明的目的在于提供一种基于无人机视觉的建筑表观病害全息重建系统和方法,能够解决现有技术中缺乏对病害在三维空间中精确重建的能力、无法实现病害的全息可视化与结构化表达的问题
[0112]This invention addresses the limitations of traditional manual inspections, such as incomplete coverage and restricted viewing angles of fixed sensors, by constructing a multi-view collaborative acquisition system using unmanned aerial vehicles (UAVs) and incorporating a high-precision spatiotemporal synchronization mechanism. This achieves fully automated, comprehensive, and efficient acquisition of building facade data. A multispectral visual sensor array and hardware-level time synchronization module ensure strict alignment of multimodal data across the spatiotemporal dimensions, laying a data foundation for subsequent high-precision 3D reconstruction. A three-stage reconstruction process—sparse point cloud construction, dense point cloud generation, and Poisson surface reconstruction—overcomes the shortcomings of existing technologies, such as rough 3D models, surface discontinuities, and topological errors, generating a high-fidelity, closed, and smooth 3D mesh model of the building facade. The introduction of a semantic segmentation network and spatial mapping mechanism accurately maps the lesion identification results from 2D images to 3D space, resolving issues of ambiguous spatial positioning and distorted morphological representation of lesions. By extracting multidimensional geometric parameters of the lesion region using discrete differential geometry methods, the invention achieves precise quantification of key indicators such as lesion size, shape, and curvature, surpassing the limitations of existing technologies that only provide pixel-level length or area estimations. Finally, through texture mapping and structured data output, a holographic visualization model containing spatial location, geometric shape, surface texture, and category attributes was constructed. This model provides high-precision, high-completeness, and high-visualization data support for building structural safety assessment, disease evolution analysis, and repair scheme formulation, significantly improving the technical level and engineering application value of building surface disease detection.
Smart Images

Figure CN122597709A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and intelligent detection technology, and in particular to a system and method for holographic reconstruction of building surface defects based on UAV vision. Background Technology
[0002] With the acceleration of urbanization and the exacerbation of infrastructure aging, the automated detection and visual assessment of structural defects in buildings has become a core requirement for intelligent operation and maintenance in civil engineering. Traditional manual inspections rely on experience-based judgment and static image recording, with their core principles based on visual observation and localized measurements. However, structural defects exhibit a highly non-uniform spatial distribution: crack direction, width variations, and surface peeling areas often span facades, corners, and high-altitude areas, making comprehensive manual coverage difficult and prone to overlooking hidden defects. Static image acquisition suffers from limited viewing angles, lighting interference, and lack of scale, resulting in large errors in extracting geometric parameters of defects and fragmented three-dimensional spatial relationships, failing to support a holistic deduction of structural safety status.
[0003] In addition, existing visual inspection systems based on fixed cameras or handheld devices lack the ability to continuously perceive space and the dynamic modeling mechanism, making it difficult to track the evolution of diseases in a timely manner and to visualize them in multiple dimensions. This seriously restricts the accuracy of structural health assessment and the timeliness of decision response.
[0004] The identification and measurement of building surface defects mainly rely on manual inspections or fixed visual sensors, which suffers from limited coverage, low operational efficiency, discontinuous data acquisition, lack of three-dimensional spatial information, and insufficient accuracy in quantifying defect morphology. Especially in high-rise buildings, complex facades, or hazardous areas, manual inspection is difficult to implement, and the deployment of fixed sensors is costly and has limited viewing angles, resulting in incomplete defect identification results, discrete measurement data, and the inability to construct a complete spatial defect distribution model. Some existing technologies attempt to introduce drones equipped with visual devices for inspection, but their data processing remains at the two-dimensional image level, lacking the ability to accurately reconstruct the true morphology, spatial coordinates, geometric dimensions, and surface texture of defects in three-dimensional space. This prevents the realization of holographic visualization and structured representation of defects, resulting in a lack of accurate data support for subsequent assessment and repair decisions.
[0005] Therefore, there is a need to provide a holographic reconstruction system and method for building surface defects based on UAV vision, which can solve the problems of existing technologies lacking the ability to accurately reconstruct defects in three-dimensional space and failing to achieve holographic visualization and structured expression of defects. Summary of the Invention
[0006] The purpose of this invention is to provide a system and method for holographic reconstruction of building surface defects based on UAV vision, which can solve the problems of existing technologies that lack the ability to accurately reconstruct defects in three-dimensional space and cannot achieve holographic visualization and structured expression of defects.
[0007] This invention is implemented as follows:
[0008] A holographic reconstruction system for building surface defects based on UAV vision includes a UAV vision acquisition module, a spatiotemporal synchronization and pose binding module, an image preprocessing module, a sparse point cloud construction module, a dense point cloud generation module, a surface reconstruction module, a semantic segmentation module, a 3D semantic mapping module, a geometric parameter extraction module, a texture mapping and visualization module, and a structured data output module.
[0009] The aforementioned UAV visual acquisition module includes a UAV platform and a multispectral visual sensor array, an airborne inertial navigation system, and a global positioning system module mounted on the UAV.
[0010] The multispectral vision sensor array includes a visible light imaging unit, a near-infrared imaging unit, and a polarization imaging unit. The visible light imaging unit, near-infrared imaging unit, and polarization imaging unit achieve millisecond-level frame synchronization acquisition through the time synchronization module in the hardware-level spatiotemporal synchronization and pose binding module.
[0011] The spatiotemporal synchronization and pose binding module includes a hardware-level time synchronization module and a spatial pose synchronization module. The spatial pose synchronization module is equipped with a high-precision pose synchronization system, which, combined with the airborne inertial navigation system and global positioning system module, acquires the spatial pose data of the UAV in real time when acquiring each frame of image, including three-dimensional coordinates, pitch angle, roll angle, and yaw angle.
[0012] The image preprocessing module is used to preprocess visual data packets with pose labels to generate standardized visual datasets.
[0013] The sparse point cloud construction module is used to construct a sparse point cloud model of a building facade based on a standardized visual dataset.
[0014] The dense point cloud generation module is used to generate a dense point cloud model of a building facade based on a sparse point cloud model using a multi-view stereo matching algorithm.
[0015] The surface reconstruction module is used to reconstruct the surface of the dense point cloud model and generate a triangular mesh model of the building facade.
[0016] The semantic segmentation module is used to perform pixel-level identification and annotation of building surface defects on the triangular mesh model through a semantic segmentation network.
[0017] The three-dimensional semantic mapping module is used to project the disease category label map onto the surface of the triangular mesh model, establish the mapping relationship between the disease area and the three-dimensional mesh vertices, and generate a three-dimensional mesh model with disease semantic labels.
[0018] The geometric parameter extraction module is used to extract geometric parameters for each type of disease region in a 3D mesh model with disease semantic labels;
[0019] The texture mapping and visualization module is used to perform texture mapping on a 3D mesh model with disease semantic labels to generate a holographic visualization model of building surface diseases.
[0020] The structured data output module is used to output a holographic visualization model of building surface defects and its corresponding structured data file.
[0021] A method for holographic reconstruction of building appearance defects using the aforementioned UAV vision-based building appearance defect holographic reconstruction system includes the following steps:
[0022] Step 1: Collect raw visual data sequences covering the entire facade of the building using a multispectral visual sensor array and a spatiotemporal synchronization module;
[0023] Step 2: Using the airborne inertial navigation system and the global positioning system module, the spatial pose data of the UAV is acquired in real time when each frame of image is captured, forming a visual data package with pose labels;
[0024] Step 3: The image preprocessing module preprocesses the visual data packets with pose labels to generate a standardized visual dataset;
[0025] Step 4: The sparse point cloud construction module constructs a sparse point cloud model of the building facade based on a standardized visual dataset;
[0026] Step 5: The dense point cloud generation module generates a dense point cloud model of the building facade based on the sparse point cloud model using a multi-view stereo matching algorithm.
[0027] Step 6: The surface reconstruction module performs surface reconstruction on the dense point cloud model to generate a triangular mesh model of the building facade;
[0028] Step 7: The semantic segmentation module performs pixel-level identification and annotation of building surface defects on the triangular mesh model using a semantic segmentation network;
[0029] Step 8: The 3D semantic mapping module projects the disease category label map onto the surface of the triangular mesh model, establishes the mapping relationship between the disease area and the 3D mesh vertices, and generates a 3D mesh model with disease semantic labels.
[0030] Step 9: The geometric parameter extraction module extracts geometric parameters for each type of disease region in the 3D mesh model with disease semantic labels;
[0031] Step 10: The texture mapping and visualization module performs texture mapping on the 3D mesh model with disease semantic labels to generate a holographic visualization model of building surface diseases;
[0032] Step 11: The structured data output module outputs a holographic visualization model of building surface defects and its corresponding structured data file.
[0033] Step 1 includes the following sub-steps:
[0034] Step 11: The UAV visual acquisition module uses a multispectral visual sensor array mounted on the UAV platform to continuously acquire images of the target building facade under multiple angles, heights, and lighting conditions along a preset flight path, obtaining a sequence of original visual data covering the entire building facade.
[0035] In step 11, the flight trajectory of the UAV is pre-planned by the ground control station, and the trajectory parameters include flight altitude, horizontal spacing, heading angle, and image overlap rate.
[0036] Step 12: The raw visual data sequence is stored in an uncompressed bitmap format, with each frame image accompanied by a timestamp accurate to the microsecond level;
[0037] Step 2 includes the following sub-steps:
[0038] Step 21: The spatial pose synchronization module of the spatiotemporal synchronization and pose binding module uses a high-precision pose synchronization system, combined with the airborne inertial navigation system and the global positioning system module, to acquire the spatial pose data of the UAV in real time when acquiring each frame of image, including three-dimensional coordinates, pitch angle, roll angle and yaw angle.
[0039] Step 22: Bind the spatial pose data to the original visual data sequence one-to-one using a timestamp matching mechanism to form a visual data package with pose labels;
[0040] In step 22, the bound visual data packet with pose label is stored in the form of structured data blocks. Each data block contains six fields: image data pointer, timestamp, three-dimensional coordinates, pitch angle, roll angle, and yaw angle.
[0041] Step 3 includes the following sub-steps:
[0042] Step 31: Perform distortion correction on the images in the visual data package;
[0043] Specifically, the camera factory calibration parameters of the multispectral vision sensor array are first read, including focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient. For each pixel, its corresponding position in the ideal distortion-free image is calculated using a reverse mapping model based on its position in the image coordinate system, and then the pixel value is obtained through bilinear interpolation. The distortion compensation model uses a Brownian model for distortion correction, which includes three radial distortion coefficients and two tangential distortion coefficients.
[0044] Step 32: Perform illumination normalization on the distortion-corrected image;
[0045] Specifically, the local adaptive histogram equalization algorithm divides the image into several local regions, performs histogram equalization on each region independently, and then smooths the region boundaries through bilinear interpolation.
[0046] Step 33: Perform noise filtering on the image after illumination normalization;
[0047] Specifically, a three-dimensional block matching filter algorithm is used to treat multiple adjacent images as a three-dimensional data cube, and similar image blocks are searched simultaneously in the spatial and temporal domains. Gaussian noise and salt-and-pepper noise are suppressed through collaborative filtering.
[0048] Step 34: Perform resolution unification processing on the noise-filtered image;
[0049] Specifically, a cubic spline interpolation algorithm is used to scale all images to a uniform size while maintaining edge sharpness;
[0050] Step 35: The preprocessed standardized visual dataset is stored in a compressed format.
[0051] Step 4 includes the following sub-steps:
[0052] Step 41: Construct a sparse point cloud model of the building facade;
[0053] Step 42: Feature point extraction;
[0054] Specifically, a Gaussian scale space is first constructed for each image to generate a multi-scale image pyramid. At each scale level, keypoints with scale invariance are detected by calculating the determinant response of the Hessian matrix of each pixel. For each keypoint, its principal orientation is calculated to ensure rotation invariance.
[0055] Step 43: Feature descriptor generation;
[0056] Specifically, a 16×16 pixel neighborhood is defined around the key point in step 42, which is divided into four 4×4 sub-regions. Gradient histograms in eight directions are calculated for each sub-region, ultimately forming a 128-dimensional feature vector.
[0057] Step 44: Feature matching;
[0058] Specifically, the nearest neighbor distance ratio method is used to find the nearest and second nearest matching points for each feature point in the feature descriptor subsets of the two images. If the ratio of the nearest distance to the second nearest distance is less than 0.8, it is determined to be a valid match; otherwise, it is determined to be an invalid match.
[0059] Step 45: Fundamental matrix estimation;
[0060] Specifically, the random sampling consensus algorithm is used to randomly select eight pairs of matching points from the matching point pairs, calculate the fundamental matrix, and then evaluate the quality of the sparse point cloud model by the number of interior points, iterating until convergence.
[0061] Step 46: Triangulation calculation;
[0062] Specifically, by using the matching point pairs and their corresponding camera poses, the three-dimensional coordinates of the spatial points are solved using the linear least squares method.
[0063] Step 47: Pose optimization;
[0064] Specifically, the bundle adjustment algorithm is used, with the coordinates of all three-dimensional points and the camera pose as optimization variables, and the reprojection error as the cost function. The Levenberg-Marquardt algorithm is used for iterative optimization to minimize the sum of squared reprojection errors of all observation points.
[0065] Step 5 includes the following sub-steps:
[0066] Step 51: Construct a cost space for each image, where the dimensions of the cost space are the image width, image height, and number of depth levels;
[0067] Step 52: Path aggregation;
[0068] Specifically, dynamic programming is performed along eight directions, with the cumulative cost calculated independently for each direction. Finally, the cumulative costs of the eight directions are added together to obtain the total cost of each pixel at each depth level.
[0069] Step 53: Parallax Selection;
[0070] Specifically, the depth level with the lowest total cost is selected as the optimal parallax, and then converted into three-dimensional coordinates based on the camera pose.
[0071] Step 54: Obtain color information by sampling from the original visual data sequence through back projection;
[0072] Step 55: Generate a dense point cloud model.
[0073] Step 6 includes the following sub-steps:
[0074] Step 61: Calculate the normal vector for each point in the point cloud;
[0075] Specifically, the normal vector calculation adopts the local plane fitting method. Taking the current point cloud point as the center, it searches for its fifty nearest neighbor point cloud points, fits the best plane, and takes the plane normal as the normal of the point cloud point, thereby calculating the normal vector of each point cloud point.
[0076] Step 62: Unify the direction of the normal vector;
[0077] Step 63: Poisson Reconstruction.
[0078] Specifically, the point cloud and normal vector are treated as samples of a vector field. An indicator function is constructed such that the indicator function is one inside the object and zero outside, and its gradient field is consistent with the sampled normal vector field. By solving the Poisson equation, the isosurface of the indicator function is obtained, which is the reconstructed surface.
[0079] Step 64: Isosurface extraction, using the moving cube algorithm to generate a mesh model composed of triangular facets;
[0080] Step 65: Mesh simplification;
[0081] Specifically, an edge-folding algorithm is used to reduce the number of triangular facets to 20% to 30% of the original number while maintaining geometric features;
[0082] Step 66: Generate a triangular mesh model.
[0083] In step 7, the semantic segmentation network is an encoder-decoder architecture. The encoder uses a residual neural network, with the 50th layer serving as the backbone to extract multi-scale feature maps. The decoder uses a pyramid pooling module and a skip connection structure to output disease category label maps. The pyramid pooling module pools the feature maps output by the encoder at different scales, then upsamples them to their original size, and after concatenation, fuses multi-scale contextual information through convolutional layers. The skip connection structure adds the feature maps from different levels of the encoder to the feature maps from the corresponding levels of the decoder, enhancing the ability to recover details.
[0084] The defects are categorized into five types: cracks, peeling, water seepage, rust, and bulging.
[0085] The semantic segmentation network takes a 2D image as input and renders a 3D mesh model into a multi-view 2D image. The rendering process uses orthogonal projection to generate a frontal view along the building facade normal, with the resolution consistent with the original visual data sequence. The output of the semantic segmentation network is a label map of the same size as the input image, with each pixel assigned a category label. The training dataset for the semantic segmentation network contains manually annotated images of building defects, covering different building types, materials, lighting conditions, and defect morphologies. The training process uses a cross-entropy loss function, with an adaptive moment estimation algorithm as the optimizer. The initial learning rate is 0.001, decreasing by half every ten rounds. During inference, predictions are made for the multi-view rendered images of the same region, and then the results are fused using majority voting to improve segmentation accuracy. Finally, the defect category label map output by the semantic segmentation network is strictly aligned with the rendered images of the original visual data sequence.
[0086] Step 8 includes the following sub-steps:
[0087] Step 81: Use perspective projection transformation algorithm and nearest neighbor interpolation algorithm to project the disease category label map onto the surface of the triangular mesh model to ensure the spatial consistency of semantic labels in three-dimensional space;
[0088] Specifically, the perspective projection transformation algorithm establishes the correspondence between 2D pixel coordinates and 3D mesh vertices based on the camera pose during rendering. For each 2D pixel, its viewing direction in 3D space is calculated using the perspective projection transformation formula, and its intersection with the triangular mesh model is obtained to obtain the corresponding 3D vertex. If the viewing direction intersects with multiple triangular faces, the closest intersection point is selected. For each 3D vertex, the category labels of all pixels projected to that vertex are counted, and the final semantic label of the vertex is determined by majority voting. The nearest neighbor interpolation algorithm is used to process vertices that are not directly projected, assigning them the semantic labels of the spatial nearest neighbor labeled vertices.
[0089] Step 82: Perform spatial smoothing on semantic tags to eliminate tag noise;
[0090] Step 83: Using the three-dimensional median filtering algorithm, with each vertex as the center, search for the ten nearest neighbor vertices in its spatial neighborhood, and take the mode as the final label of the vertex.
[0091] Step 84: Generate a 3D mesh model with semantic labels for defects based on the final labels of each vertex. Each vertex is accompanied by a semantic label, which fully marks the spatial distribution of all defect areas on the building facade.
[0092] Step 9 includes the following sub-steps:
[0093] Step 91: Extract geometric parameters, including the area of the diseased region, maximum length, maximum width, perimeter, depth gradient, and mean surface curvature.
[0094] The area of the diseased region is calculated by summing the areas of all triangular faces within the region, and the area of the triangular faces is calculated using Heron's formula.
[0095] The calculation method for the maximum length and maximum width is as follows: first, extract the boundary vertices of the diseased area, construct the convex hull, calculate the Euclidean distance between all pairs of vertices on the convex hull, take the maximum value as the maximum length, and take the maximum distance perpendicular to the direction of the maximum length as the maximum width.
[0096] The perimeter is calculated by summing the lengths of all sides on the boundary of the diseased area.
[0097] The method for calculating the depth change gradient is as follows: First, define the reference plane of the building facade, fit the best plane of all healthy area vertices through principal component analysis, and use it as the reference; then calculate the signed distance from each diseased vertex to the reference plane, and then calculate the difference in distance between adjacent vertices, and take the average of the absolute values as the depth change gradient.
[0098] The method for calculating the average surface curvature is as follows: using the discrete Gaussian curvature formula, for each vertex, calculate the angular defects of all triangular facets in its neighborhood, divide by the neighborhood area, and obtain the average curvature.
[0099] Step 92: All geometric parameters are statistically analyzed according to disease category to generate a structured parameter table containing eight fields: disease number, category, area, maximum length, maximum width, perimeter, depth variation gradient, and average surface curvature.
[0100] Step 10 includes the following sub-steps:
[0101] Step 101: Select the optimal viewpoint image from the original visual data sequence as the texture source;
[0102] The method for selecting texture sources is as follows: calculate the normal vector of each triangular facet, calculate the angle between it and the optical axis of all cameras, and select the camera with the smallest angle and no occlusion as the texture source of that facet.
[0103] Step 102: Using an automatic texture coordinate unfolding and seamless stitching algorithm, the real surface texture is attached to the 3D mesh surface;
[0104] The method for automatically unfolding texture coordinates is to use existing parametric mapping algorithms to flatten the 3D mesh surface into a 2D parametric domain.
[0105] Specifically, the seamless stitching algorithm processes the texture seams between adjacent patches, employs Poisson image editing technology, solves the gradient domain equation to make the gradient at the seam continuous and eliminate visual discontinuities; the texture image is cropped from the corresponding region of the original visual data, and after distortion correction and color correction, it is pasted into the two-dimensional parameter domain.
[0106] Step 103: Generate a holographic visualization model;
[0107] The final holographic visualization model contains three layers of information: geometric mesh, semantic labels, and real texture.
[0108] In step 11, the structured data file contains the spatial coordinates of the disease, geometric parameters, category labels, texture index, and the code of the building area to which it belongs. It supports direct loading by 3D visualization software and storage in a structured database.
[0109] The holographic visualization model is exported in an open format, consisting of three parts: geometry file, material file, and texture file. The geometry file stores vertex coordinates, face indices, normal vectors, and texture coordinates; the material file defines texture mapping paths and rendering parameters; and the texture file stores compressed texture images.
[0110] The structured data file is stored in tabular form, with each row corresponding to an independent disease instance. Fields include a unique disease identifier, three-dimensional coordinates of the center point, coordinates of the corner points of the smallest bounding rectangle, area, maximum length, maximum width, perimeter, depth gradient, mean surface curvature, disease category, texture file path, floor number, and facade orientation.
[0111] Compared with the prior art, the present invention has the following advantages:
[0112] This invention addresses the limitations of traditional manual inspections, such as incomplete coverage and restricted viewing angles of fixed sensors, by constructing a multi-view collaborative acquisition system using unmanned aerial vehicles (UAVs) and incorporating a high-precision spatiotemporal synchronization mechanism. This achieves fully automated, comprehensive, and efficient acquisition of building facade data. A multispectral visual sensor array and hardware-level time synchronization module ensure strict alignment of multimodal data across the spatiotemporal dimensions, laying a data foundation for subsequent high-precision 3D reconstruction. A three-stage reconstruction process—sparse point cloud construction, dense point cloud generation, and Poisson surface reconstruction—overcomes the shortcomings of existing technologies, such as rough 3D models, surface discontinuities, and topological errors, generating a high-fidelity, closed, and smooth 3D mesh model of the building facade. The introduction of a semantic segmentation network and spatial mapping mechanism accurately maps the lesion identification results from 2D images to 3D space, resolving issues of ambiguous spatial positioning and distorted morphological representation of lesions. By extracting multidimensional geometric parameters of the lesion region using discrete differential geometry methods, the invention achieves precise quantification of key indicators such as lesion size, shape, and curvature, surpassing the limitations of existing technologies that only provide pixel-level length or area estimations. Finally, through texture mapping and structured data output, a holographic visualization model containing spatial location, geometric shape, surface texture, and category attributes was constructed. This model provides high-precision, high-completeness, and high-visualization data support for building structural safety assessment, disease evolution analysis, and repair scheme formulation, significantly improving the technical level and engineering application value of building surface disease detection. Attached Figure Description
[0113] Figure 1 This is an architecture diagram of the holographic reconstruction system for building surface defects based on UAV vision, as described in this invention.
[0114] Figure 2 This is a flowchart of the method for holographic reconstruction of building surface defects based on UAV vision according to the present invention;
[0115] Figure 3 This is a flowchart illustrating the four-stage logical process of the UAV-based holographic reconstruction method for building surface defects, which involves UAV acquisition, 3D reconstruction, semantic mapping, and parameter extraction. Detailed Implementation
[0116] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0117] Please see the appendix Figure 1 A holographic reconstruction system for building surface defects based on UAV vision includes a UAV vision acquisition module, a spatiotemporal synchronization and pose binding module, an image preprocessing module, a sparse point cloud construction module, a dense point cloud generation module, a surface reconstruction module, a semantic segmentation module, a 3D semantic mapping module, a geometric parameter extraction module, a texture mapping and visualization module, and a structured data output module.
[0118] The UAV visual acquisition module includes a UAV platform (including UAV, landing platform, charging equipment, etc.) and a multispectral visual sensor (i.e. camera) array, an airborne inertial navigation system and a global positioning system module mounted on the UAV. The UAV platform, equipped with a multispectral visual sensor array, continuously acquires images of the target building facade under multiple angles, heights and lighting conditions along a preset flight trajectory, and obtains the original visual data sequence covering the entire building facade.
[0119] The multispectral vision sensor array includes a visible light imaging unit, a near-infrared imaging unit, and a polarization imaging unit. The visible light imaging unit, the near-infrared imaging unit, and the polarization imaging unit achieve millisecond-level frame synchronization acquisition through the time synchronization module in the hardware-level spatiotemporal synchronization and pose binding module.
[0120] The visible light imaging unit is responsible for acquiring color and texture information of the building surface. Its sensor resolution is no less than 40 million pixels, and the frame rate is no less than 15 frames per second. The near-infrared imaging unit can penetrate surface dust and shallow coatings to reveal potential structural cracks and moisture penetration areas. Its operating wavelength is 700 to 1000 nanometers, and its resolution is no less than 20 million pixels. The polarization imaging unit can eliminate specular reflection interference from glass curtain walls or metal surfaces and enhance the contrast of damaged areas. Its polarization angle is adjustable, supporting simultaneous acquisition of four channels at 0 degrees, 45 degrees, 90 degrees, and 135 degrees. The hardware-level time synchronization module can use a field-programmable gate array (FPGA) chip as the main control clock source. It triggers the shutters of each sensor through pulse width modulation (PWM) signals, ensuring that all polarization imaging units complete exposure within the same millisecond time window, with a time synchronization error of less than 0.5 milliseconds.
[0121] The spatiotemporal synchronization and pose binding module includes a hardware-level time synchronization module and a spatial pose synchronization module. The spatial pose synchronization module is equipped with a high-precision pose synchronization system, combined with an onboard inertial navigation system and a global positioning system module, to acquire the spatial pose data of the UAV in real time when acquiring each frame of image data, including three-dimensional coordinates, pitch angle, roll angle, and yaw angle. The spatial pose data and the original visual data sequence are bound one-to-one through a timestamp matching mechanism to form a visual data package with pose tags.
[0122] Preferably, the airborne inertial navigation system can consist of a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer, with a sampling frequency of 200 Hz. It fuses data from various sensors in a multispectral visual sensor array using an extended Kalman filter algorithm to output high-frequency attitude angle information. The global positioning system module can employ dual-frequency carrier phase differential positioning technology, achieving a positioning accuracy better than five centimeters, with an update frequency of 10 Hz. The core of the timestamp matching mechanism lies in establishing a precise correspondence between visual data frames and pose data frames.
[0123] The image preprocessing module is used to preprocess visual data packets with pose labels. The preprocessing includes image distortion correction, illumination normalization, noise filtering, and resolution unification to generate a standardized visual dataset.
[0124] Preferably, image distortion correction can be preprocessed using a radial and tangential distortion joint compensation model based on the camera intrinsic parameter matrix, illumination normalization can be preprocessed using a local adaptive histogram equalization algorithm, and noise filtering can be preprocessed using a three-dimensional block matching filter algorithm.
[0125] The sparse point cloud construction module is used to construct a sparse point cloud model of a building facade based on a standardized visual dataset.
[0126] The construction process of the sparse point cloud model includes feature point extraction, feature descriptor generation, feature matching, fundamental matrix estimation, triangulation calculation, and pose optimization. Preferably, feature point extraction can employ an accelerated robust feature algorithm, feature descriptor generation can employ the oriented gradient histogram descriptor, and pose optimization can employ the bundle adjustment algorithm.
[0127] The dense point cloud generation module is used to generate a dense point cloud model of a building facade based on a sparse point cloud model using a multi-view stereo matching algorithm. Preferably, the multi-view stereo matching algorithm can adopt a semi-global matching strategy, with pixel-level cost calculation and path aggregation as the core, and output dense point cloud data containing three-dimensional coordinates and color information.
[0128] The surface reconstruction module is used to reconstruct the surface of the dense point cloud model, generating a triangular mesh model of the building facade. Preferably, the surface reconstruction can employ the Poisson surface reconstruction algorithm, which obtains a closed, smooth, and topologically correct three-dimensional mesh surface by solving the Poisson equation of the indicator function.
[0129] The semantic segmentation module is used to perform pixel-level identification and annotation of building surface defects on the triangular mesh model through a semantic segmentation network. Preferably, the semantic segmentation network is an encoder-decoder architecture, where the encoder part can adopt a residual neural network and the decoder part can adopt a pyramid pooling module and a skip connection structure to output a defect category label map; the defect categories include five types: cracks, peeling, water seepage, corrosion, and bulging.
[0130] The aforementioned 3D semantic mapping module is used to project the disease category label map onto the surface of the triangular mesh model, establish a mapping relationship between the disease area and the vertices of the 3D mesh, and generate a 3D mesh model with disease semantic labels. Preferably, the projection process can employ perspective projection transformation and nearest neighbor interpolation algorithms to ensure the spatial consistency of semantic labels in 3D space.
[0131] The geometric parameter extraction module is used to extract geometric parameters for each type of disease region in a 3D mesh model with disease semantic labels, including the area, maximum length, maximum width, perimeter, depth gradient, and mean surface curvature of the disease region. Preferably, the geometric parameter extraction is based on the calculation of mesh vertex coordinates and triangular facet normal vectors, and can be implemented using discrete differential geometry methods.
[0132] The texture mapping and visualization module is used to perform texture mapping on a 3D mesh model with disease semantic labels to generate a holographic visualization model of building surface diseases. Preferably, the texture mapping can be performed by selecting the optimal viewpoint image from the original visual data sequence as the texture source, and then using an automatic texture coordinate unfolding and seamless stitching algorithm to attach the real surface texture to the 3D mesh surface.
[0133] The structured data output module is used to output a holographic visualization model of building surface defects and its corresponding structured data file; the structured data file includes the spatial coordinates of the defects, geometric parameters, category labels, texture index, and the code of the building area to which it belongs, and supports direct loading by 3D visualization software and storage in a structured database.
[0134] Please see the appendix Figure 2 and attached Figure 3 A method for holographic reconstruction of building surface defects based on UAV vision includes four stages: UAV data acquisition, 3D reconstruction, semantic mapping, and parameter extraction. Specifically, it includes the following steps:
[0135] Step 1: Collect raw visual data sequences covering the entire facade of the building using a multispectral visual sensor array and a spatiotemporal synchronization module.
[0136] Step 1 includes the following sub-steps:
[0137] Step 11: The UAV visual acquisition module uses a multispectral visual sensor array mounted on the UAV platform to continuously acquire images of the target building facade under multiple angles, heights, and lighting conditions along a preset flight path, obtaining a sequence of original visual data covering the entire building facade.
[0138] In step 11, the flight trajectory of the UAV is pre-planned by the ground control station, and the trajectory parameters include flight altitude, horizontal spacing, heading angle, and image overlap rate.
[0139] The flight altitude is set based on the total building height, typically between 0.8 and 1.2 times the building height, to ensure full facade coverage and image resolution that meets subsequent processing requirements. The horizontal spacing is set to the horizontal distance between adjacent shooting points, typically 0.3 to 0.5 times the distance between the drone and the building facade, to ensure sufficient overlap between adjacent images for feature matching. The heading angle is automatically adjusted based on the building facade's orientation, ensuring the multispectral vision sensor array's optical axis is always perpendicular to the currently scanned building facade. The shooting overlap rate can be set to 70%-90% to ensure sufficient redundant information for point cloud densification and surface optimization during 3D reconstruction.
[0140] Step 12: The raw visual data sequence is stored in an uncompressed bitmap format, with each frame image accompanied by a timestamp accurate to the microsecond level, providing a basis for subsequent binding with pose data.
[0141] Step 2: Using the airborne inertial navigation system and the global positioning system module, the spatial pose data of the UAV is acquired in real time when each frame of image is captured, forming a visual data package with pose labels.
[0142] Step 2 includes the following sub-steps:
[0143] Step 21: The spatial pose synchronization module of the spatiotemporal synchronization and pose binding module uses a high-precision pose synchronization system, combined with the airborne inertial navigation system and the global positioning system module, to acquire the spatial pose data of the UAV in real time when acquiring each frame of image, including three-dimensional coordinates, pitch angle, roll angle and yaw angle.
[0144] Because the frame rate of the multispectral vision sensor array is inconsistent with the frame rate of the pose sensor in the spatial pose synchronization module, linear interpolation is required to upsample the pose data. Specifically, for each frame of visual image, linear interpolation is performed between two adjacent pose data points centered on its timestamp to calculate the precise three-dimensional coordinates and pose angle corresponding to that timestamp.
[0145] The three-dimensional coordinates are represented in the geocentric and geofixed coordinate system, and the attitude angles are represented by the Euler angles of the UAV's body coordinate system relative to the local horizontal coordinate system.
[0146] Step 22: Bind the spatial pose data to the original visual data sequence one-to-one using a timestamp matching mechanism to form a visual data package with pose labels.
[0147] In step 22, the bound visual data packets with pose labels are stored in the form of structured data blocks. Each data block contains six fields: image data pointer, timestamp, three-dimensional coordinates, pitch angle, roll angle, and yaw angle, ensuring that subsequent processing modules can directly call the spatial pose information without additional parsing.
[0148] Step 3: The image preprocessing module preprocesses the visual data packets with pose labels to generate a standardized visual dataset.
[0149] Step 3 includes the following sub-steps:
[0150] Step 31: Perform distortion correction on the images in the visual data package.
[0151] Preferably, image distortion correction can be achieved using a joint radial and tangential distortion compensation model based on the camera intrinsic parameter matrix.
[0152] Specifically, the camera's factory calibration parameters for the multispectral vision sensor array are first read, including focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient. For each pixel, its corresponding position in the ideal distortion-free image is calculated using an existing inverse mapping model based on its position in the image coordinate system, and then the pixel value is obtained through bilinear interpolation. The distortion compensation model can use the existing Brownian model for distortion correction, which includes three radial distortion coefficients and two tangential distortion coefficients, effectively eliminating barrel and pincushion distortion.
[0153] Step 32: Perform illumination normalization on the distortion-corrected image.
[0154] Preferably, illumination normalization can be achieved using a local adaptive histogram equalization algorithm. Illumination normalization processing eliminates the influence of differences in illumination intensity and color temperature on subsequent feature extraction for images acquired at different times and under different weather conditions.
[0155] Specifically, the local adaptive histogram equalization algorithm divides the image into several local regions, performs histogram equalization on each region independently, and then smooths the region boundaries through bilinear interpolation to avoid block artifacts.
[0156] Step 33: Perform noise filtering on the image after illumination normalization.
[0157] Preferably, noise filtering can be achieved using a three-dimensional block matched filtering algorithm.
[0158] Specifically, a three-dimensional block matching filter algorithm is used to treat multiple adjacent images as a three-dimensional data cube, and similar image blocks are searched simultaneously in the spatial and temporal domains. Gaussian noise and salt-and-pepper noise are suppressed through collaborative filtering.
[0159] Step 34: Perform resolution unification processing on the noise-filtered image.
[0160] Specifically, all images are scaled to a uniform size, typically 4096×2160 pixels, using a cubic spline interpolation algorithm to maintain edge sharpness during the scaling process.
[0161] Step 35: The preprocessed standardized visual dataset is stored in a compressed format with a compression ratio controlled within 5:1 to ensure that the image quality loss is less than 3%.
[0162] Step 4: The sparse point cloud construction module constructs a sparse point cloud model of the building facade based on a standardized visual dataset.
[0163] Step 4 includes the following sub-steps:
[0164] Step 41: Construct a sparse point cloud model of the building facade.
[0165] Step 42: Feature point extraction.
[0166] Specifically, a Gaussian scale space is first constructed for each image to generate a multi-scale image pyramid. At each scale level, scale-invariant keypoints are detected by calculating the determinant response of the Hessian matrix of each pixel. For each keypoint, its principal orientation is calculated to ensure rotation invariance.
[0167] Step 43: Feature descriptor generation.
[0168] Specifically, a 16×16 pixel neighborhood is defined around the key point in step 42, which is then divided into four 4×4 sub-regions. Gradient histograms in eight directions are calculated for each sub-region, ultimately forming a 128-dimensional feature vector.
[0169] Step 44: Feature matching.
[0170] Specifically, the nearest neighbor distance ratio method is used. In the feature descriptor subsets of the two images, the nearest and second nearest matching points are found for each feature point. If the ratio of the nearest distance to the second nearest distance is less than 0.8, it is determined to be a valid match; otherwise, it is determined to be an invalid match.
[0171] Step 45: Estimation of the fundamental matrix.
[0172] Specifically, the random sampling consensus algorithm is used to randomly select eight pairs of matching points from the matching point pairs, calculate the fundamental matrix, and then evaluate the quality of the sparse point cloud model by the number of interior points, iterating until convergence.
[0173] Step 46: Triangulation calculation.
[0174] Specifically, by using the matching point pairs and their corresponding camera poses, the three-dimensional coordinates of the spatial points are solved using the linear least squares method.
[0175] Step 47: Pose optimization.
[0176] Specifically, the bundle adjustment algorithm is used, with the coordinates of all three-dimensional points and the camera pose as optimization variables, and the reprojection error as the cost function. The algorithm is then iteratively optimized using the existing Levenberg-Marquardt algorithm to minimize the sum of squared reprojection errors of all observation points.
[0177] The sparse point cloud model ultimately contains tens of thousands to hundreds of thousands of 3D points. Each 3D point is accompanied by color information and a visibility list, which is used to record which cameras observed the corresponding 3D point.
[0178] Step 5: The dense point cloud generation module generates a dense point cloud model of the building facade based on the sparse point cloud model using a multi-view stereo matching algorithm.
[0179] The dense point cloud generation module adopts a semi-global matching strategy through a multi-view stereo matching algorithm, with pixel-level cost calculation and path aggregation as the core, and outputs dense point cloud data containing three-dimensional coordinates and color information based on the sparse point cloud model.
[0180] Step 5 includes the following sub-steps:
[0181] Step 51: Construct a cost space for each image, with dimensions of image width, image height, and depth level.
[0182] The number of depth levels is set according to the depth range of the building facade and the required precision, and is usually between five hundred and one thousand levels.
[0183] Step 52: Path aggregation.
[0184] Specifically, dynamic programming is performed along eight directions, with the cumulative cost calculated independently for each direction. Finally, the cumulative costs of the eight directions are added together to obtain the total cost for each pixel at each depth level.
[0185] Pixel-level cost calculation can use the normalized cross-correlation coefficient as a similarity measure to calculate the degree of matching between the current pixel and corresponding pixels from other viewpoints under different depth assumptions.
[0186] Step 53: Parallax selection.
[0187] Specifically, the depth level with the lowest total cost is selected as the optimal parallax, and then converted into three-dimensional coordinates based on the camera pose.
[0188] Step 54: Obtain color information by sampling from the original visual data sequence through back projection.
[0189] Step 55: Generate a dense point cloud model based on the information from steps 51-54.
[0190] The dense point cloud model contains millions to tens of millions of three-dimensional points with a point spacing of less than five millimeters, completely covering all visible surfaces of the building facade.
[0191] Preferably, to improve computational efficiency, the densification process adopts a block processing strategy, that is, the large-size image is divided into several overlapping sub-blocks, and stereo matching is performed on each sub-block. Finally, the point cloud data of the overlapping areas are fused by weighted averaging to form a dense point cloud model.
[0192] Step 6: The surface reconstruction module performs surface reconstruction on the dense point cloud model to generate a triangular mesh model of the building facade.
[0193] Surface reconstruction can be performed using the Poisson surface reconstruction algorithm, which obtains a closed, smooth, and topologically correct 3D mesh surface by solving the Poisson equation of the indicator function.
[0194] Step 6 includes the following sub-steps:
[0195] Step 61: Calculate the normal vector for each point cloud point.
[0196] Specifically, the normal vector can be calculated using a local plane fitting method. Taking the current point cloud point as the center, search for its fifty nearest neighbor point cloud points, fit the best plane, and take the plane normal as the normal vector of the point cloud point, thereby calculating the normal vector of each point cloud point.
[0197] Step 62: Unify the direction of the normal vector.
[0198] Preferably, the minimum spanning tree algorithm of the existing technology is used to unify the direction of the normal vector, that is, starting from the seed point, the direction of the normal vector is gradually propagated to ensure that all normal vectors point to the outside of the object.
[0199] Step 63: Poisson Reconstruction.
[0200] Specifically, the point cloud and normal vectors are treated as samples of a vector field. An indicator function is constructed such that the indicator function is one inside the object and zero outside, and its gradient field is consistent with the sampled normal vector field. By solving the Poisson equation, the isosurface of the indicator function is obtained, which is the reconstructed surface.
[0201] Step 64: Isosurface Extraction. The existing moving cube algorithm can be used to generate a mesh model composed of triangular facets.
[0202] Step 65: Grid simplification.
[0203] Specifically, by using existing edge-folding algorithms, the number of triangular facets can be reduced to 20% to 30% of the original number while maintaining geometric features, thereby reducing the computational burden of subsequent processing.
[0204] Step 66: Generate a triangular mesh model.
[0205] The final generated triangular mesh model contains hundreds of thousands to millions of triangular faces with smooth, continuous surfaces, no holes or self-intersections, and correct topological structure, which can be directly used for subsequent semantic segmentation and texture mapping.
[0206] Step 7: The semantic segmentation module performs pixel-level identification and annotation of building surface defects on the triangular mesh model through a semantic segmentation network.
[0207] In step 7, the semantic segmentation network employs an encoder-decoder architecture. The encoder uses a residual neural network, with its fiftieth layer serving as the backbone to extract multi-scale feature maps. The decoder uses a pyramid pooling module and skip connections to output disease category label maps. The pyramid pooling module pools the feature maps output by the encoder at different scales, then upsamples them to their original size, concatenates them, and fuses multi-scale contextual information through convolutional layers. The skip connections add the feature maps from different levels of the encoder to the corresponding feature maps from the decoder, enhancing detail recovery capabilities.
[0208] The aforementioned disease categories include five types: cracks, peeling, water seepage, rust, and bulging.
[0209] The semantic segmentation network takes a 2D image as input and renders a 3D mesh model into a multi-view 2D image. The rendering process uses orthogonal projection to generate a frontal view along the building facade normal, with the resolution consistent with the original visual data sequence. The output of the semantic segmentation network is a label map of the same size as the input image, with each pixel assigned a category label. The training dataset for the semantic segmentation network contains 5,000 manually annotated images of building defects, covering different building types, materials, lighting conditions, and defect morphologies. The training process uses a cross-entropy loss function, with an adaptive moment estimation algorithm as the optimizer. The initial learning rate is 0.001, decreasing by half every ten rounds. During inference, predictions are made for the multi-view rendered images of the same region, and then the results are fused using majority voting to improve segmentation accuracy. Finally, the defect category label map output by the semantic segmentation network is strictly aligned with the rendered images of the original visual data sequence, providing a foundation for subsequent spatial mapping.
[0210] Step 8: The 3D semantic mapping module projects the disease category label map onto the surface of the triangular mesh model, establishes the mapping relationship between the disease area and the vertices of the 3D mesh, and generates a 3D mesh model with disease semantic labels.
[0211] Step 8 includes the following sub-steps:
[0212] Step 81: Use perspective projection transformation algorithm and nearest neighbor interpolation algorithm to project the disease category label map onto the surface of the triangular mesh model to ensure the spatial consistency of semantic labels in three-dimensional space.
[0213] Specifically, the perspective projection transformation algorithm establishes a correspondence between 2D pixel coordinates and 3D mesh vertices based on the camera pose during rendering. For each 2D pixel, its viewing direction in 3D space is calculated using the perspective projection transformation formula, and its intersection with the triangular mesh model is obtained to obtain the corresponding 3D vertex. If the viewing direction intersects with multiple triangular faces, the closest intersection point is selected. For each 3D vertex, the category labels of all pixels projected to that vertex are counted, and the final semantic label of the vertex is determined by majority voting. The nearest neighbor interpolation algorithm is used to handle vertices not directly projected to, assigning them the semantic labels of their spatial nearest neighbor labeled vertices.
[0214] Step 82: Perform spatial smoothing on semantic tags to eliminate tag noise.
[0215] Step 83: Using a three-dimensional median filtering algorithm, search for the ten nearest neighbor vertices in the spatial neighborhood of each vertex, and take the mode as the final label of the vertex.
[0216] Step 84: Generate a 3D mesh model with semantic labels for defects based on the final labels of each vertex. Each vertex is accompanied by a semantic label, which fully marks the spatial distribution of all defect areas on the building facade.
[0217] Step 9: The geometric parameter extraction module extracts geometric parameters for each type of disease region in the 3D mesh model with disease semantic labels.
[0218] Step 9 includes the following sub-steps:
[0219] Step 91: Extract geometric parameters, including the area of the diseased region, maximum length, maximum width, perimeter, depth gradient, and mean surface curvature.
[0220] Preferably, the geometric parameter extraction is based on the calculation of the grid vertex coordinates and the normal vectors of the triangular facets, and is implemented using the discrete differential geometry method.
[0221] The area of the diseased region can be calculated by summing the areas of all triangular facets within the region. The area of the triangular facets can be calculated using the Heron formula, which is based on existing technology.
[0222] The calculation method for the maximum length and maximum width is as follows: first, extract the boundary vertices of the diseased area, construct the convex hull, calculate the Euclidean distance between all vertex pairs on the convex hull, take the maximum value as the maximum length, and take the maximum distance perpendicular to the direction of the maximum length as the maximum width.
[0223] The perimeter can be calculated by summing the lengths of all sides on the boundary of the diseased area.
[0224] The method for calculating the depth variation gradient is as follows: First, define a reference plane for the building facade. Then, fit the optimal plane of all healthy area vertices using principal component analysis, and use it as the reference. Next, calculate the signed distance from each diseased vertex to the reference plane, and then calculate the difference in distance between adjacent vertices. Take the average of the absolute values as the depth variation gradient.
[0225] The method for calculating the average surface curvature is as follows: the discrete Gaussian curvature formula of the existing technology can be used to calculate the angular defects of all triangular facets in the neighborhood of each vertex, and divide by the neighborhood area to obtain the average curvature.
[0226] Step 92: All geometric parameters are statistically analyzed according to disease category to generate a structured parameter table containing eight fields: disease number, category, area, maximum length, maximum width, perimeter, depth variation gradient, and average surface curvature.
[0227] Step 10: The texture mapping and visualization module performs texture mapping on the 3D mesh model with disease semantic labels to generate a holographic visualization model of building surface diseases.
[0228] Step 10 includes the following sub-steps:
[0229] Step 101: Select the optimal viewpoint image from the original visual data sequence as the texture source.
[0230] The method for selecting the texture source is as follows: calculate the normal vector of each triangular facet, calculate the angle between the normal vector and the optical axis of all cameras, and select the camera with the smallest angle and no obstruction as the texture source of that facet.
[0231] Step 102: Using an automatic texture coordinate unfolding and seamless stitching algorithm, the real surface texture is attached to the 3D mesh surface.
[0232] The method for automatically unfolding texture coordinates is to use existing parametric mapping algorithms to flatten the three-dimensional mesh surface into a two-dimensional parameter domain, ensuring that the texture is free from stretching and tearing.
[0233] Specifically, the seamless stitching algorithm processes the texture seams between adjacent patches by employing existing Poisson image editing techniques to solve the gradient domain equation, ensuring gradient continuity at the seams and eliminating visual discontinuities. The texture image is cropped from the original visual data, and after distortion correction and color correction, it is then pasted into the two-dimensional parameter domain.
[0234] Step 103: Generate a holographic visualization model.
[0235] The resulting holographic visualization model contains three layers of information: geometric mesh, semantic tags, and real texture. It can be freely rotated, scaled, and sectioned in existing 3D visualization software to intuitively display the spatial morphology and distribution of the disease.
[0236] Step 11: The structured data output module outputs a holographic visualization model of building surface defects and its corresponding structured data file.
[0237] The structured data file contains the spatial coordinates of the lesion, geometric parameters, category labels, texture index, and the code of the building area to which it belongs. It supports direct loading by existing 3D visualization software and storage in the structured database.
[0238] The holographic visualization model is exported in an open format and consists of three parts: geometry file, material file, and texture file.
[0239] The geometry file stores vertex coordinates, face indices, normal vectors, and texture coordinates. The material file defines the texture map path and rendering parameters. The texture file stores the compressed texture image.
[0240] The structured data file is stored in tabular form, with each row corresponding to an independent disease instance. The fields include a unique disease identifier, three-dimensional coordinates of the center point, coordinates of the corner point of the smallest bounding rectangle, area, maximum length, maximum width, perimeter, depth gradient, average surface curvature, disease category, texture file path, floor number, and facade orientation.
[0241] The structured data files can be directly imported into existing building information modeling platforms or structural health monitoring databases for tracking disease evolution, risk assessment, and support for repair decisions.
[0242] This invention utilizes multi-view collaborative data acquisition from unmanned aerial vehicles (UAVs), combined with a high-precision spatiotemporal synchronization mechanism and multi-source data fusion algorithm, to achieve accurate reconstruction of building surface defects from two-dimensional images to three-dimensional holographic models. It outputs a structured dataset containing the spatial location, geometric dimensions, surface texture, morphological features, and spatial distribution relationships of the defects, enabling spatial localization, morphological quantification, and high-fidelity visualization of defects. This significantly improves the accuracy and efficiency of building health assessment, providing high-precision, high-completeness, and high-visualization technical support for building structural safety assessment. It addresses the problems of incomplete coverage, missing three-dimensional information, inaccurate morphological quantification, and low visualization in existing technologies for detecting building surface defects. The entire methodology, from data acquisition to result output, achieves fully automated, high-precision, and holographic reconstruction of building surface defects, providing a novel technical means for modern building operation and maintenance.
[0243] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the invention. Therefore, any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A holographic reconstruction system for architectural surface defects based on UAV vision, characterized by: It includes a UAV visual acquisition module, a spatiotemporal synchronization and pose binding module, an image preprocessing module, a sparse point cloud construction module, a dense point cloud generation module, a surface reconstruction module, a semantic segmentation module, a 3D semantic mapping module, a geometric parameter extraction module, a texture mapping and visualization module, and a structured data output module; The aforementioned UAV visual acquisition module includes a UAV platform and a multispectral visual sensor array, an airborne inertial navigation system, and a global positioning system module mounted on the UAV. The multispectral vision sensor array includes a visible light imaging unit, a near-infrared imaging unit, and a polarization imaging unit. The visible light imaging unit, near-infrared imaging unit, and polarization imaging unit achieve millisecond-level frame synchronization acquisition through the time synchronization module in the hardware-level spatiotemporal synchronization and pose binding module. The spatiotemporal synchronization and pose binding module includes a hardware-level time synchronization module and a spatial pose synchronization module. The spatial pose synchronization module is equipped with a high-precision pose synchronization system, which, combined with the airborne inertial navigation system and global positioning system module, acquires the spatial pose data of the UAV in real time when acquiring each frame of image, including three-dimensional coordinates, pitch angle, roll angle, and yaw angle. The image preprocessing module is used to preprocess visual data packets with pose labels to generate standardized visual datasets. The sparse point cloud construction module is used to construct a sparse point cloud model of a building facade based on a standardized visual dataset. The dense point cloud generation module is used to generate a dense point cloud model of a building facade based on a sparse point cloud model using a multi-view stereo matching algorithm. The surface reconstruction module is used to reconstruct the surface of the dense point cloud model and generate a triangular mesh model of the building facade. The semantic segmentation module is used to perform pixel-level identification and annotation of building surface defects on the triangular mesh model through a semantic segmentation network. The three-dimensional semantic mapping module is used to project the disease category label map onto the surface of the triangular mesh model, establish the mapping relationship between the disease area and the three-dimensional mesh vertices, and generate a three-dimensional mesh model with disease semantic labels. The geometric parameter extraction module is used to extract geometric parameters for each type of disease region in a 3D mesh model with disease semantic labels; The texture mapping and visualization module is used to perform texture mapping on a 3D mesh model with disease semantic labels to generate a holographic visualization model of building surface diseases. The structured data output module is used to output a holographic visualization model of building surface defects and its corresponding structured data file.
2. A method for holographic reconstruction of building appearance defects using the UAV vision-based building appearance defect holographic reconstruction system as described in claim 1, characterized in that: Includes the following steps: Step 1: Collect raw visual data sequences covering the entire facade of the building using a multispectral visual sensor array and a spatiotemporal synchronization module; Step 2: Using the airborne inertial navigation system and the global positioning system module, the spatial pose data of the UAV is acquired in real time when each frame of image is captured, forming a visual data package with pose labels; Step 3: The image preprocessing module preprocesses the visual data packets with pose labels to generate a standardized visual dataset; Step 4: The sparse point cloud construction module constructs a sparse point cloud model of the building facade based on a standardized visual dataset; Step 5: The dense point cloud generation module generates a dense point cloud model of the building facade based on the sparse point cloud model using a multi-view stereo matching algorithm. Step 6: The surface reconstruction module performs surface reconstruction on the dense point cloud model to generate a triangular mesh model of the building facade; Step 7: The semantic segmentation module performs pixel-level identification and annotation of building surface defects on the triangular mesh model using a semantic segmentation network; Step 8: The 3D semantic mapping module projects the disease category label map onto the surface of the triangular mesh model, establishes the mapping relationship between the disease area and the 3D mesh vertices, and generates a 3D mesh model with disease semantic labels. Step 9: The geometric parameter extraction module extracts geometric parameters for each type of disease region in the 3D mesh model with disease semantic labels; Step 10: The texture mapping and visualization module performs texture mapping on the 3D mesh model with disease semantic labels to generate a holographic visualization model of building surface diseases; Step 11: The structured data output module outputs a holographic visualization model of building surface defects and its corresponding structured data file.
3. The method for holographic reconstruction of architectural surface defects according to claim 2, characterized in that: Step 1 includes the following sub-steps: Step 11: The UAV visual acquisition module uses a multispectral visual sensor array mounted on the UAV platform to continuously acquire images of the target building facade under multiple angles, heights, and lighting conditions along a preset flight path, obtaining a sequence of original visual data covering the entire building facade. In step 11, the flight trajectory of the UAV is pre-planned by the ground control station, and the trajectory parameters include flight altitude, horizontal spacing, heading angle, and image overlap rate. Step 12: The raw visual data sequence is stored in an uncompressed bitmap format, with each frame image accompanied by a timestamp accurate to the microsecond level; Step 2 includes the following sub-steps: Step 21: The spatial pose synchronization module of the spatiotemporal synchronization and pose binding module uses a high-precision pose synchronization system, combined with the airborne inertial navigation system and the global positioning system module, to acquire the spatial pose data of the UAV in real time when acquiring each frame of image, including three-dimensional coordinates, pitch angle, roll angle and yaw angle. Step 22: Bind the spatial pose data to the original visual data sequence one-to-one using a timestamp matching mechanism to form a visual data package with pose labels; In step 22, the bound visual data packet with pose label is stored in the form of structured data blocks. Each data block contains six fields: image data pointer, timestamp, three-dimensional coordinates, pitch angle, roll angle, and yaw angle.
4. The method for holographic reconstruction of architectural surface defects according to claim 2, characterized in that: Step 3 includes the following sub-steps: Step 31: Perform distortion correction on the images in the visual data package; Specifically, the camera factory calibration parameters of the multispectral vision sensor array are first read, including focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient. For each pixel, its corresponding position in the ideal distortion-free image is calculated using a reverse mapping model based on its position in the image coordinate system, and then the pixel value is obtained through bilinear interpolation. The distortion compensation model uses a Brownian model for distortion correction, which includes three radial distortion coefficients and two tangential distortion coefficients. Step 32: Perform illumination normalization on the distortion-corrected image; Specifically, the local adaptive histogram equalization algorithm divides the image into several local regions, performs histogram equalization on each region independently, and then smooths the region boundaries through bilinear interpolation. Step 33: Perform noise filtering on the image after illumination normalization; Specifically, a three-dimensional block matching filter algorithm is used to treat multiple adjacent images as a three-dimensional data cube, and similar image blocks are searched simultaneously in the spatial and temporal domains. Gaussian noise and salt-and-pepper noise are suppressed through collaborative filtering. Step 34: Perform resolution unification processing on the noise-filtered image; Specifically, a cubic spline interpolation algorithm is used to scale all images to a uniform size while maintaining edge sharpness; Step 35: The preprocessed standardized visual dataset is stored in a compressed format.
5. The method for holographic reconstruction of architectural surface defects according to claim 2, characterized in that: Step 4 includes the following sub-steps: Step 41: Construct a sparse point cloud model of the building facade; Step 42: Feature point extraction; Specifically, firstly, a Gaussian scale space is constructed for each image to generate a multi-scale image pyramid; at each scale level, keypoints with scale invariance are detected by calculating the determinant response of the Hessian matrix of the pixels; for each keypoint, its principal direction is calculated to ensure rotation invariance. Step 43: Feature descriptor generation; Specifically, a 16×16 pixel neighborhood is defined around the key point in step 42, which is divided into four 4×4 sub-regions. Gradient histograms in eight directions are calculated for each sub-region, ultimately forming a 128-dimensional feature vector. Step 44: Feature matching; Specifically, the nearest neighbor distance ratio method is used to find the nearest and second nearest matching points for each feature point in the feature descriptor subsets of the two images. If the ratio of the nearest distance to the second nearest distance is less than 0.8, it is determined to be a valid match; otherwise, it is determined to be an invalid match. Step 45: Fundamental matrix estimation; Specifically, the random sampling consensus algorithm is used to randomly select eight pairs of matching points from the matching point pairs, calculate the fundamental matrix, and then evaluate the quality of the sparse point cloud model by the number of interior points, iterating until convergence. Step 46: Triangulation calculation; Specifically, by using the matching point pairs and their corresponding camera poses, the three-dimensional coordinates of the spatial points are solved using the linear least squares method. Step 47: Pose optimization; Specifically, the bundle adjustment algorithm is used, with the coordinates of all three-dimensional points and the camera pose as optimization variables, and the reprojection error as the cost function. The Levenberg-Marquardt algorithm is used for iterative optimization to minimize the sum of squared reprojection errors of all observation points. Step 5 includes the following sub-steps: Step 51: Construct a cost space for each image, where the dimensions of the cost space are the image width, image height, and number of depth levels; Step 52: Path aggregation; Specifically, dynamic programming is performed along eight directions, with the cumulative cost calculated independently for each direction. Finally, the cumulative costs of the eight directions are added together to obtain the total cost of each pixel at each depth level. Step 53: Parallax Selection; Specifically, the depth level with the lowest total cost is selected as the optimal parallax, and then converted into three-dimensional coordinates based on the camera pose. Step 54: Obtain color information by sampling from the original visual data sequence through back projection; Step 55: Generate a dense point cloud model.
6. The method for holographic reconstruction of architectural surface defects according to claim 2, characterized in that: Step 6 includes the following sub-steps: Step 61: Calculate the normal vector for each point in the point cloud; Specifically, the normal vector calculation adopts the local plane fitting method. Taking the current point cloud point as the center, it searches for its fifty nearest neighbor point cloud points, fits the best plane, and takes the plane normal as the normal of the point cloud point, thereby calculating the normal vector of each point cloud point. Step 62: Unify the direction of the normal vector; Step 63: Poisson Reconstruction; Specifically, the point cloud and normal vector are treated as samples of a vector field. An indicator function is constructed such that the indicator function is one inside the object and zero outside, and its gradient field is consistent with the sampled normal vector field. By solving the Poisson equation, the isosurface of the indicator function is obtained, which is the reconstructed surface. Step 64: Isosurface extraction, using the moving cube algorithm to generate a mesh model composed of triangular facets; Step 65: Mesh simplification; Specifically, an edge-folding algorithm is used to reduce the number of triangular facets to 20% to 30% of the original number while maintaining geometric features; Step 66: Generate a triangular mesh model.
7. The method for holographic reconstruction of architectural surface defects according to claim 2, characterized in that: In step 7, the semantic segmentation network is an encoder-decoder architecture. The encoder uses a residual neural network, with the 50th layer serving as the backbone to extract multi-scale feature maps. The decoder uses a pyramid pooling module and a skip connection structure to output disease category label maps. The pyramid pooling module pools the feature maps output by the encoder at different scales, then upsamples them to their original size, and after concatenation, fuses multi-scale contextual information through convolutional layers. The skip connection structure adds the feature maps from different levels of the encoder to the feature maps from the corresponding levels of the decoder, enhancing the ability to recover details. The defects are categorized into five types: cracks, peeling, water seepage, rust, and bulging. The semantic segmentation network takes a 2D image as input and renders a 3D mesh model into a multi-view 2D image. The rendering process uses orthogonal projection to generate a frontal view along the building facade normal, with the resolution consistent with the original visual data sequence. The output of the semantic segmentation network is a label map of the same size as the input image, with each pixel assigned a category label. The training dataset for the semantic segmentation network contains manually annotated images of building defects, covering different building types, materials, lighting conditions, and defect morphologies. The training process uses a cross-entropy loss function, with an adaptive moment estimation algorithm as the optimizer. The initial learning rate is 0.001, decreasing by half every ten rounds. The inference process predicts the multi-view rendered images of the same region separately, then fuses the multi-view results using majority voting to improve segmentation accuracy. Finally, the defect category label map output by the semantic segmentation network is strictly aligned with the rendered images of the original visual data sequence. Step 8 includes the following sub-steps: Step 81: Use perspective projection transformation algorithm and nearest neighbor interpolation algorithm to project the disease category label map onto the surface of the triangular mesh model to ensure the spatial consistency of semantic labels in three-dimensional space; Specifically, the perspective projection transformation algorithm establishes the correspondence between 2D pixel coordinates and 3D mesh vertices based on the camera pose during rendering. For each 2D pixel, its viewing direction in 3D space is calculated using the perspective projection transformation formula, and its intersection with the triangular mesh model is obtained to obtain the corresponding 3D vertex. If the viewing direction intersects with multiple triangular faces, the closest intersection point is selected. For each 3D vertex, the category labels of all pixels projected to that vertex are counted, and the final semantic label of the vertex is determined by majority voting. The nearest neighbor interpolation algorithm is used to process vertices that are not directly projected, assigning them the semantic labels of the spatial nearest neighbor labeled vertices. Step 82: Perform spatial smoothing on semantic tags to eliminate tag noise; Step 83: Using the three-dimensional median filtering algorithm, with each vertex as the center, search for the ten nearest neighbor vertices in its spatial neighborhood, and take the mode as the final label of the vertex. Step 84: Generate a 3D mesh model with semantic labels for defects based on the final labels of each vertex. Each vertex is accompanied by a semantic label, which fully marks the spatial distribution of all defect areas on the building facade.
8. The method for holographic reconstruction of architectural surface defects according to claim 2, characterized in that: Step 9 includes the following sub-steps: Step 91: Extract geometric parameters, including the area of the diseased region, maximum length, maximum width, perimeter, depth gradient, and mean surface curvature. The area of the diseased region is calculated by summing the areas of all triangular faces within the region, and the area of the triangular faces is calculated using Heron's formula. The calculation method for the maximum length and maximum width is as follows: first, extract the boundary vertices of the diseased area, construct the convex hull, calculate the Euclidean distance between all pairs of vertices on the convex hull, take the maximum value as the maximum length, and take the maximum distance perpendicular to the direction of the maximum length as the maximum width. The perimeter is calculated by summing the lengths of all sides on the boundary of the diseased area. The method for calculating the depth change gradient is as follows: First, define the reference plane of the building facade, fit the best plane of all healthy area vertices through principal component analysis, and use it as the reference; then calculate the signed distance from each diseased vertex to the reference plane, and then calculate the difference in distance between adjacent vertices, and take the average of the absolute values as the depth change gradient. The method for calculating the average surface curvature is as follows: using the discrete Gaussian curvature formula, for each vertex, calculate the angular defects of all triangular facets in its neighborhood, divide by the neighborhood area, and obtain the average curvature. Step 92: All geometric parameters are statistically analyzed according to disease category to generate a structured parameter table containing eight fields: disease number, category, area, maximum length, maximum width, perimeter, depth variation gradient, and average surface curvature.
9. The method for holographic reconstruction of architectural surface defects according to claim 2, characterized in that: Step 10 includes the following sub-steps: Step 101: Select the optimal viewpoint image from the original visual data sequence as the texture source; The method for selecting texture sources is as follows: calculate the normal vector of each triangular facet, calculate the angle between it and the optical axis of all cameras, and select the camera with the smallest angle and no occlusion as the texture source of that facet. Step 102: Using an automatic texture coordinate unfolding and seamless stitching algorithm, the real surface texture is attached to the 3D mesh surface; The method for automatically unfolding texture coordinates is to use existing parametric mapping algorithms to flatten the 3D mesh surface into a 2D parametric domain. Specifically, the seamless stitching algorithm processes the texture seams between adjacent patches, employs Poisson image editing technology, solves the gradient domain equation to make the gradient at the seam continuous and eliminate visual discontinuities; the texture image is cropped from the corresponding region of the original visual data, and after distortion correction and color correction, it is pasted into the two-dimensional parameter domain. Step 103: Generate a holographic visualization model; The final holographic visualization model contains three layers of information: geometric mesh, semantic labels, and real texture.
10. The method for holographic reconstruction of architectural surface defects according to claim 2, characterized in that: In step 11, the structured data file contains the spatial coordinates of the disease, geometric parameters, category labels, texture index, and the code of the building area to which it belongs. It supports direct loading by 3D visualization software and storage in a structured database. The holographic visualization model is exported in an open format, consisting of three parts: geometry file, material file, and texture file. The geometry file stores vertex coordinates, face indices, normal vectors, and texture coordinates; the material file defines texture mapping paths and rendering parameters; and the texture file stores compressed texture images. The structured data file is stored in tabular form, with each row corresponding to an independent disease instance. Fields include a unique disease identifier, three-dimensional coordinates of the center point, coordinates of the corner points of the smallest bounding rectangle, area, maximum length, maximum width, perimeter, depth gradient, mean surface curvature, disease category, texture file path, floor number, and facade orientation.