A building outer wall defect identification and quantitative evaluation system based on multi-modal fusion

The multimodal fusion building exterior wall defect identification and quantitative assessment system solves the problems of low efficiency and inaccurate assessment in existing technologies, realizes comprehensive identification and scientific quantitative assessment of building exterior wall defects, and improves the automation level and safety of detection.

CN122153583APending Publication Date: 2026-06-05AILUO (WUHAN) TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AILUO (WUHAN) TECHNOLOGY CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Current methods for detecting defects in building exterior walls rely on manual inspections, which are inefficient and highly subjective. Single detection methods are insufficient to comprehensively and accurately identify and assess various defects, and the assessment results lack scientific quantification, affecting building safety and maintenance decisions.

Method used

The building exterior wall defect identification and quantitative assessment system adopts multimodal fusion. Through feature extraction, feature fusion, defect identification, spatial correlation and quantitative assessment modules, it combines multiple detection data for comprehensive analysis, identifies defect categories and spatial distribution, and generates quantitative assessment parameters.

Benefits of technology

It improves the accuracy and reliability of defect identification, overcomes the limitations of single detection methods, realizes objective quantitative assessment of defects, reduces the influence of subjective human factors, and provides a scientific basis for maintenance decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122153583A_ABST
    Figure CN122153583A_ABST
Patent Text Reader

Abstract

The application discloses a kind of building outer wall defect identification and quantitative evaluation system based on multi-modal fusion, it is related to building detection and artificial intelligence technical field, including: feature extraction module, feature fusion module, defect identification module, spatial correlation module, quantitative evaluation module and grade determination module.The system is projected to unified feature space by obtaining multi-modal detection data and extracting feature representation, and is fused based on response consistency, to realize the intelligent identification of defect;Space influence factor is generated by calculating defect space proximity and distribution density, combined with feature numerical distribution for quantitative evaluation, based on the distance between evaluation vector and preset grade standard to realize defect grade determination, to provide a comprehensive and reliable technical solution for building outer wall defect evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of building inspection and artificial intelligence technology, specifically to a system for identifying and quantitatively evaluating defects in building exterior walls based on multimodal fusion. Background Technology

[0002] With the acceleration of urbanization and the increase in the service life of buildings, various defects in building exterior walls are becoming increasingly prominent. These defects not only affect the aesthetics of buildings but also endanger the structural safety and service life of buildings. Currently, the detection of defects in building exterior walls mainly relies on manual inspections, which suffers from low efficiency, strong subjectivity, and high safety risks. Although automated detection methods using single detection methods have been developed, due to the complexity and diversity of the types and manifestations of defects in building exterior walls, single detection methods are often insufficient to comprehensively and accurately identify and assess various defects.

[0003] In existing technologies, some studies employ single detection methods such as optics, infrared, and ultrasound for defect identification. However, each method has its limitations: optical detection is susceptible to environmental factors such as lighting and weather; infrared detection is highly dependent on ambient temperature; and ultrasonic detection is sensitive to material properties. These limitations make it difficult to guarantee the reliability and accuracy of the detection results. Existing defect assessment methods are mostly based on empirical judgment or simple image processing algorithms, lacking in-depth analysis and scientific quantification of defect characteristics, making it difficult to objectively assess the degree of defects. Moreover, the assessment results often fail to fully reflect the actual condition of the defects, affecting the scientific nature of subsequent maintenance decisions. Furthermore, there is a lack of systematic consideration of the spatial distribution characteristics of multiple defects and their synergistic effects, making it difficult to comprehensively assess the overall impact of defects on building safety. Summary of the Invention

[0004] The purpose of this invention is to provide a building exterior wall defect identification and quantitative assessment system based on multimodal fusion, which aims to solve at least one of the technical problems existing in the prior art.

[0005] The technical solution of this invention is: a building exterior wall defect identification and quantitative assessment system based on multimodal fusion, comprising the following steps: The feature extraction module is used to acquire multimodal detection data of building exterior walls and extract features from the multimodal detection data to obtain feature representations of each modality. The feature fusion module is used to project the modal feature representations onto the same feature dimension space, calculate the response consistency of each modal feature representation, assign fusion weights according to the response consistency and perform weighted concatenation to obtain a fused feature vector; The defect identification module is used to identify defects based on the fused feature vector and obtain defect category identifiers and spatial location information; The spatial association module is used to calculate the spatial proximity of multiple defects based on the spatial location information, divide the defects whose spatial proximity satisfies the clustering condition into a defect set, calculate the spatial distribution density of the defect set and generate a spatial influence factor. The quantitative evaluation module is used to extract a feature subset corresponding to the defect category identifier from the fused feature vector, calculate the numerical distribution statistics of the feature subset, and weight the numerical distribution statistics with the spatial influence factor and map them to physical size and damage depth to obtain quantitative evaluation parameters. The level determination module is used to combine the quantitative evaluation parameters with the defect category identifier into a judgment vector, calculate the distance between the judgment vector and each level standard vector in the preset level standard set, and select the level corresponding to the smallest distance as the defect level determination result.

[0006] The specific operation process of the feature extraction module is as follows: Acquire multimodal detection data of building exterior walls, including visual detection data of the building exterior wall surface and physical detection data of the building exterior wall interior; Noise filtering is performed on each modal data in the multimodal detection data to obtain filtered modal data. Multi-level encoding operations are performed on the filtered modal data, feature responses are extracted at each encoding level, and the feature responses at each encoding level are concatenated to obtain a multi-level feature representation. The multi-level feature representations are subjected to scale normalization to obtain normalized feature representations; The normalized feature representation is compared with the preset feature template by dimension-wise similarity calculation. Feature dimensions with similarity higher than the preset screening threshold are selected. Feature components corresponding to the feature dimensions are extracted from the normalized feature representation to form the feature representations of each modality.

[0007] The specific operation process of the feature fusion module is as follows: Receive modal feature representations output by the feature extraction module; Each modal feature representation is subjected to dimensional transformation, and each modal feature representation is projected onto the same feature dimension space to obtain the projected modal feature representations. Calculate the feature response differences between the modal feature representations after projection in the corresponding regions of spatial locations, identify the sensitivity of each mode to building exterior wall defects based on the feature response differences, and determine the response consistency of each modal feature representation based on the sensitivity. After projection, feature regions whose feature activation intensity exceeds a preset activation threshold are extracted from each modal feature representation. The feature numerical statistics within the feature region are calculated as signal intensity. The response consistency and the signal intensity are normalized and weighted to obtain the fusion weight corresponding to each modal feature representation. The modal feature representations after projection are weighted element-wise according to the fusion weights, and the weighted modal feature representations are concatenated to obtain the fused feature vector.

[0008] The specific operation process of the defect identification module is as follows: Receive the fused feature vector output by the feature fusion module; The fused feature vector is matched with the standard features of each defect category in the preset defect category feature library. The defect category corresponding to the standard feature with the highest similarity is selected as the defect category identifier. Spatial structure analysis is performed on the fused feature vector to identify spatial regions in which the feature activation values ​​in the fused feature vector are continuously distributed, and the boundary coordinates and geometric center coordinates of the spatial regions are calculated. Convert the boundary coordinates to the boundary position in the physical coordinate system of the building's exterior wall surface, and convert the geometric center coordinates to the center position in the physical coordinate system of the building's exterior wall surface. Output the boundary and center positions as spatial location information.

[0009] The specific operation process of the spatial association module is as follows: Receive spatial location information of multiple defects output by the defect identification module; Calculate the Euclidean distance between the spatial location information of any two defects among multiple defects, use the Euclidean distance as the spatial proximity, filter defect pairs with spatial proximity less than a preset distance threshold, and mark the defects in the defect pairs as defects that satisfy the clustering condition. For the defects that meet the aggregation conditions, a connectivity analysis is performed, and the interconnected defects are grouped into the same defect set. The number of defects in the defect set and the area of ​​the spatial region covered by the defect set are counted, and the ratio of the number of defects to the area of ​​the spatial region is calculated as the spatial distribution density. The difference between the spatial distribution density and the preset benchmark density value is calculated, and the difference is exponentially calculated and normalized to obtain the spatial influence factor.

[0010] The specific operation process of the quantitative evaluation module is as follows: Receive the fused feature vector and defect category identifier output by the defect identification module and the spatial influence factor output by the spatial association module; Semantic segmentation is performed on the fused feature vector based on the defect category identifier, and the feature dimension components that are semantically associated with the defect category identifier are extracted from the fused feature vector. The feature dimension components are then combined into a feature subset. Calculate the mean and variance of each feature value within the feature subset, and combine the mean and variance into a numerical distribution statistic; The numerical distribution statistics and spatial influence factors are subjected to tensor product operation. The result is linearly transformed and the horizontal component is extracted as the size correlation quantity, and the vertical component is extracted as the depth correlation quantity. Based on the defect category identifier, the corresponding size conversion coefficient and depth conversion coefficient are determined. The size correlation quantity is scaled according to the size conversion coefficient to obtain the physical size. The depth correlation quantity is scaled according to the depth conversion coefficient to obtain the damage depth. The physical size and damage depth are combined into a quantitative evaluation parameter.

[0011] The specific operation process of the level determination module is as follows: Receive the quantitative evaluation parameters output by the quantitative evaluation module and the defect category identifier output by the defect identification module; The defect category identifier is converted into a category code value, and the quantitative evaluation parameter is vector-concatenated with the category code value to form an evaluation vector. Extract the standard vectors of each level sequentially from the preset set of level standards, calculate the Euclidean distance between the evaluation vector and each level standard vector, and establish a mapping relationship between each Euclidean distance and the corresponding level standard vector; In the mapping relationship, the Euclidean distance with the smallest value is selected, the level corresponding to the smallest Euclidean distance is extracted, and the level is used as the defect level determination result.

[0012] This invention improves the accuracy and reliability of defect identification through multimodal data fusion, effectively overcoming the limitations of single detection methods. The response consistency-based feature fusion mechanism adaptively adjusts the weights of each modality's features, enhancing the robustness of feature fusion. By introducing spatial correlation analysis and spatial influence factors, it achieves a quantitative representation of the distribution characteristics of defect groups, overcoming the shortcomings of traditional methods that only focus on individual defects. An objective quantitative evaluation mechanism for defect physical parameters is established by combining numerical distribution statistics of feature subsets with spatial influence factors. Through the calculation method of the distance between the evaluation vector and the standard vector, automatic defect level determination is achieved, eliminating the influence of subjective human factors. This invention not only improves the automation level of building exterior wall defect identification and evaluation but also provides a more scientific and reliable basis for building maintenance decisions. Attached Figure Description

[0013] Figure 1This is a schematic diagram of a building exterior wall defect identification and quantitative assessment system based on multimodal fusion, provided in an embodiment of the present invention. Figure 2 This is a comparative analysis diagram of the spatial distribution density of defect sets in an embodiment of the present invention. Detailed Implementation

[0014] like Figure 1 As shown, Figure 1 This is a schematic diagram of a building exterior wall defect identification and quantitative assessment system based on multimodal fusion provided in an embodiment of the present invention. The system includes: The feature extraction module is used to acquire multimodal detection data of building exterior walls and extract features from the multimodal detection data to obtain feature representations of each modality. The feature fusion module is used to project the modal feature representations onto the same feature dimension space, calculate the response consistency of each modal feature representation, assign fusion weights according to the response consistency and perform weighted concatenation to obtain a fused feature vector; The defect identification module is used to identify defects based on the fused feature vector and obtain defect category identifiers and spatial location information; The spatial association module is used to calculate the spatial proximity of multiple defects based on the spatial location information, divide the defects whose spatial proximity satisfies the clustering condition into a defect set, calculate the spatial distribution density of the defect set and generate a spatial influence factor. The quantitative evaluation module is used to extract a feature subset corresponding to the defect category identifier from the fused feature vector, calculate the numerical distribution statistics of the feature subset, and weight the numerical distribution statistics with the spatial influence factor and map them to physical size and damage depth to obtain quantitative evaluation parameters. The level determination module is used to combine the quantitative evaluation parameters with the defect category identifier into a judgment vector, calculate the distance between the judgment vector and each level standard vector in the preset level standard set, and select the level corresponding to the smallest distance as the defect level determination result.

[0015] The specific operation process of the feature extraction module is as follows: Acquire multimodal detection data of building exterior walls, including visual detection data of the building exterior wall surface and physical detection data of the building exterior wall interior; Noise filtering is performed on each modal data in the multimodal detection data to obtain filtered modal data. Multi-level encoding operations are performed on the filtered modal data, feature responses are extracted at each encoding level, and the feature responses at each encoding level are concatenated to obtain a multi-level feature representation. The multi-level feature representations are subjected to scale normalization to obtain normalized feature representations; The normalized feature representation is compared with the preset feature template by dimension-wise similarity calculation. Feature dimensions with similarity higher than the preset screening threshold are selected. Feature components corresponding to the feature dimensions are extracted from the normalized feature representation to form the feature representations of each modality.

[0016] In this embodiment, the feature extraction module acquires multimodal detection data of the building's exterior wall, including visual detection data of the building's exterior wall surface and physical detection data of the building's exterior wall interior. The visual detection data is acquired using a high-definition camera with an image resolution of 4096×3072 pixels, while the physical detection data is acquired using an ultrasonic testing device and an infrared thermal imager. The ultrasonic testing sampling frequency is 500kHz, and the infrared thermal imaging temperature accuracy is ±0.05℃.

[0017] When performing noise filtering on multimodal detection data, a bilateral filtering algorithm was used for visual detection data, with the edge-preserving parameter σd set to 5, the smoothing parameter σr set to 50, and the window size set to 21×21 pixels. Wavelet transform filtering was used for ultrasonic data, employing the db4 wavelet basis function, with a decomposition level of 4 layers. Adaptive median filtering was used for infrared thermal imaging data, with a window size ranging from 3×3 to 11×11 pixels. After filtering, noise in the visual data was smoothed while preserving edge details; high-frequency interference in the ultrasonic data was effectively suppressed; and salt-and-pepper noise in the infrared thermal imaging data was removed, improving the clarity of the thermal image.

[0018] In the multi-level encoding process, a three-layer convolutional encoding structure is used for the filtered visual data. The first layer uses 64 3×3 convolutional kernels with a stride of 2; the second layer uses 128 3×3 convolutional kernels with a stride of 2; and the third layer uses 256 3×3 convolutional kernels with a stride of 1. Each convolutional layer is followed by a ReLU activation function for nonlinear transformation. Frequency domain decomposition is used for the ultrasonic data, extracting feature responses from the 0-50kHz, 50-200kHz, and 200-500kHz frequency bands respectively. A temperature gradient pyramid is constructed for the infrared thermal imaging data, containing three levels: original size, half size, and quarter size. Temperature gradient and directional features are extracted from each level. The coded feature dimensions for each layer of visual data are 64×1024×768, 128×512×384, and 256×512×384, respectively; the feature dimensions for each frequency band of ultrasonic data are 128×1; and the feature dimensions for each layer of infrared thermal imaging are 2×4096×3072, 2×2048×1536, and 2×1024×768. By concatenating the feature responses of each coded layer according to their dimensions, visual data forms a 448×512×384-dimensional feature representation, ultrasonic data forms a 384×1-dimensional feature representation, and infrared thermal imaging data forms a 6×1024×768-dimensional feature representation.

[0019] When performing scale normalization on multi-level feature representations, visual data uses the maximum-minimum normalization method to map feature values ​​to the [0, 1] interval; ultrasonic data is standardized to make the feature mean 0 and the standard deviation 1; infrared thermal imaging data is Z-score standardized to make the temperature gradient feature distribution more concentrated. After normalization, the feature scales of the three modalities are consistent, which is helpful for subsequent similarity calculations. After normalization, visual data retains a 448×512×384-dimensional feature representation, ultrasonic data retains a 384×1-dimensional feature representation, and infrared thermal imaging data retains a 6×1024×768-dimensional feature representation.

[0020] The pre-set feature template library contains standard feature representations of 200 typical building exterior wall defects, each template annotated and verified by experts. For the visual modality, the cosine similarity between the normalized feature representation and the template is calculated; for the ultrasonic modality, Euclidean distance is calculated as the similarity index; for the infrared thermal imaging modality, Pearson correlation coefficient is calculated as the similarity index. Pre-set screening thresholds are set to 0.85, 0.80, and 0.75, respectively. After similarity calculation, there are 125 feature dimensions with similarity higher than 0.85 in the visual modality, 87 feature dimensions with similarity higher than 0.80 in the ultrasonic modality, and 112 feature dimensions with similarity higher than 0.75 in the infrared thermal imaging modality. The feature components corresponding to these high-similarity dimensions are extracted from the normalized feature representations to form the feature representations of each modality: 125×512×384-dimensional feature representation for the visual modality, 87×1-dimensional feature representation for the ultrasonic modality, and 112×1024×768-dimensional feature representation for the infrared thermal imaging modality.

[0021] During the inspection of the exterior wall of a commercial building, visual data captured surface features such as cracks, weathering, and stains; ultrasonic data detected an internal cavity, 15 mm thick, located 30 mm from the surface; and infrared thermal imaging revealed an abnormal temperature area with a maximum temperature difference of 4.8℃. After feature extraction processing, the similarity of crack features reached 0.92, the similarity of cavity features reached 0.88, and the similarity of temperature anomaly features reached 0.86, all exceeding the preset thresholds. The corresponding feature components were successfully extracted to form the modal feature representations.

[0022] This invention achieves comprehensive detection of surface and internal defects in building exterior walls through multimodal detection data fusion and multi-level feature extraction. Different filtering algorithms are used to process the data from each modality, effectively suppressing noise interference; multi-level coding captures defect features at different scales; and scale normalization and similarity filtering are employed to extract the most representative feature components. The entire technical solution is adaptable to different environmental conditions and building material types, significantly improving the accuracy and robustness of building exterior wall defect identification, providing a reliable basis for building safety assessment, reducing the subjectivity and uncertainty of traditional manual inspection, and lowering inspection costs and time consumption.

[0023] The specific operation process of the feature fusion module is as follows: Receive modal feature representations output by the feature extraction module; Each modal feature representation is subjected to dimensional transformation, and each modal feature representation is projected onto the same feature dimension space to obtain the projected modal feature representations. Calculate the feature response differences between the modal feature representations after projection in the corresponding regions of spatial locations, identify the sensitivity of each mode to building exterior wall defects based on the feature response differences, and determine the response consistency of each modal feature representation based on the sensitivity. After projection, feature regions whose feature activation intensity exceeds a preset activation threshold are extracted from each modal feature representation. The feature numerical statistics within the feature region are calculated as signal intensity. The response consistency and the signal intensity are normalized and weighted to obtain the fusion weight corresponding to each modal feature representation. The modal feature representations after projection are weighted element-wise according to the fusion weights, and the weighted modal feature representations are concatenated to obtain the fused feature vector.

[0024] The feature fusion module receives modal feature representations output by the feature extraction module, including visual modal feature representations, ultrasonic modal feature representations, and infrared thermal imaging modal feature representations. The visual modal feature representation has a dimension of 125×512×384, the ultrasonic modal feature representation has a dimension of 87×1, and the infrared thermal imaging modal feature representation has a dimension of 112×1024×768. Since the dimensions of the modal feature representations are inconsistent, a dimension transformation process is required to project each modal feature representation onto the same feature dimension space.

[0025] The dimensionality transformation process employs projection mapping technology to uniformly project the feature representations of each modality onto a 256×256×3 three-dimensional feature space. For visual modal feature representations, a bilinear interpolation method is used to transform the spatial scale, adjusting the 512×384 spatial scale to 256×256. Simultaneously, a fully connected layer compresses the 125 channels into 3 channels, forming a 256×256×3 feature representation. For ultrasonic modal feature representations, the 87×1 one-dimensional feature is upsampled to 87×256, then expanded to 256×256 through transpose convolution, and finally the number of channels is adjusted to 3 through 1×1 convolution, forming a 256×256×3 feature representation. For infrared thermal imaging modal feature representations, a bilinear interpolation method is first used to adjust the spatial scale from 1024×768 to 256×256, and then a 1×1 convolution compresses the 112 channels into 3 channels, forming a 256×256×3 feature representation. During the projection process, key information of the original feature representation is preserved to ensure that different modal features correspond to the same area of ​​the building's exterior wall at the same location.

[0026] The feature response differences of each modal feature representation after projection are calculated using a region comparison method. The projection space is uniformly divided into 16×16 regions, each region being 16×16 pixels in size. Within each region, the feature differences between three modal pairs—visual and ultrasonic, visual and infrared, and ultrasonic and infrared—are calculated. The feature differences are calculated using Euclidean distance. The feature values ​​of the three channels within each region are extracted, the distances between corresponding channels of the modal pairs are calculated, and the average of the three channel distances is used to obtain the feature response difference value for that region. Taking the visual and ultrasonic modal pair as an example, the mean value of the visual modal features within a certain region is [0.75, 0.62, 0.43], and the mean value of the ultrasonic modal features is [0.82, 0.58, 0.39], with a calculated Euclidean distance of 0.09.

[0027] The sensitivity of each modality to building exterior wall defects is identified based on the differences in characteristic responses. For each defect type, a response difference threshold standard is established. For crack-type defects, the visual modality has high sensitivity, with a threshold of 0.1; for void-type defects, the ultrasonic modality has high sensitivity, with a threshold of 0.08; and for moisture-permeable defects, the infrared thermal imaging modality has high sensitivity, with a threshold of 0.12. When the characteristic response difference is below the threshold, it indicates that the corresponding modality has high sensitivity to this type of defect. For example, in a certain area, the difference between the visual and ultrasonic modalities is 0.05, the difference between the visual and infrared modalities is 0.15, and the difference between the ultrasonic and infrared modalities is 0.18; therefore, the visual modality is determined to have the highest sensitivity in this area.

[0028] The response consistency of each modality's feature representation is determined based on sensitivity. Response consistency measures the degree of consistency in detection across different modalities in the same defect area. A consistency calculation window size of 32×32 pixels is set, and the proportion of regions within the window whose feature response differences are below a threshold is used as the response consistency index. For the visual modality and ultrasonic modality, the proportion of regions with feature response differences below the threshold within a certain window is 75%, resulting in a response consistency of 0.75; the response consistency between the visual modality and infrared modality is 0.62; and the response consistency between the ultrasonic modality and infrared modality is 0.58. The overall response consistency of each modality is calculated by weighted averaging: visual modality 0.685, ultrasonic modality 0.665, and infrared thermal imaging modality 0.600.

[0029] The activation threshold was set to the 90th percentile of the feature values: 0.82 for the visual modality, 0.78 for the ultrasonic modality, and 0.85 for the infrared thermal imaging modality. Feature regions exceeding the activation threshold were extracted: the activation region area was 6553 pixels for the visual modality, 4096 pixels for the ultrasonic modality, and 5242 pixels for the infrared thermal imaging modality.

[0030] For each modality's activation region, four statistical measures—mean, standard deviation, kurtosis, and skewness—are calculated, and the signal strength index is obtained by weighted summation. The visual modality's feature region has a mean of 0.88, a standard deviation of 0.05, a kurtosis of 3.2, and a skewness of 0.3, resulting in a calculated signal strength of 0.726; the ultrasonic modality's signal strength is 0.693; and the infrared thermal imaging modality's signal strength is 0.712.

[0031] The response consistency weighting coefficient is set to 0.6, and the signal strength weighting coefficient is set to 0.4. The visual modality fusion weight is 0.685×0.6+0.726×0.4=0.701; the ultrasonic modality fusion weight is 0.665×0.6+0.693×0.4=0.676; and the infrared thermal imaging modality fusion weight is 0.600×0.6+0.712×0.4=0.645. The fusion weights of the three modalities are normalized so that their sum is 1, resulting in a normalized fusion weight of 0.347 for the visual modality, 0.335 for the ultrasonic modality, and 0.319 for the infrared thermal imaging modality.

[0032] The visual modal feature representation is multiplied by 0.347, the ultrasonic modal feature representation by 0.335, and the infrared thermal imaging modal feature representation by 0.319 to obtain the weighted modal feature representations. The three weighted modal feature representations are then concatenated to form the final fused feature vector, with dimensions of 256×256×9. The first three channels represent the weighted visual features, the middle three channels represent the weighted ultrasonic features, and the last three channels represent the weighted infrared thermal imaging features.

[0033] This invention achieves effective fusion of multimodal detection data through unified feature space projection, difference analysis, and adaptive weight calculation. By fully utilizing the complementarity of each modality's data and adjusting the weights of each modality for different types of defects, the comprehensiveness and accuracy of defect detection are improved. The blind zone problem of single-modality detection is solved through feature response difference and sensitivity analysis; and the contribution of each modality in the fusion process is balanced through a weight calculation strategy combining response consistency and signal strength. This technology can adapt to building exterior walls of different materials and structures, effectively identifying surface and internal defects, and improving the accuracy and robustness of defect identification.

[0034] The specific operation process of the defect identification module is as follows: Receive the fused feature vector output by the feature fusion module; The fused feature vector is matched with the standard features of each defect category in the preset defect category feature library. The defect category corresponding to the standard feature with the highest similarity is selected as the defect category identifier. Spatial structure analysis is performed on the fused feature vector to identify spatial regions in which the feature activation values ​​in the fused feature vector are continuously distributed, and the boundary coordinates and geometric center coordinates of the spatial regions are calculated. Convert the boundary coordinates to the boundary position in the physical coordinate system of the building's exterior wall surface, and convert the geometric center coordinates to the center position in the physical coordinate system of the building's exterior wall surface. Output the boundary and center positions as spatial location information.

[0035] The defect identification module receives a fused feature vector output from the feature fusion module. This feature vector has dimensions of 256×256×9 and represents the multimodal fused feature information of the building's exterior wall. The fused feature vector contains feature information from visual, ultrasonic, and infrared thermal imaging modes, with each mode occupying three channels. In practical applications, the fused feature vector is transmitted to the defect identification module via a high-speed data bus at a transmission rate of 10Gbps to ensure lossless data transmission.

[0036] The pre-defined defect category feature library stores 15 common standard defect categories for building exterior walls. Each defect feature has a dimension of 256×256×9, matching the dimensions of the fused feature vector. Defect categories include surface cracks, structural cracks, weathering and spalling, material aging, concrete carbonization, steel reinforcement corrosion, component deformation, hollow areas and detachment, water seepage and leakage, thermal bridging, insulation layer damage, coating peeling, salt precipitation damage, dirt contamination, and joint damage. The pre-defined defect category feature library was trained using 2000 typical defect samples, covering defect features of building exterior walls under different materials and environmental conditions.

[0037] The cosine similarity method is used to perform similarity matching between the fused feature vector and the standard features of each defect category in the preset defect category feature library. The 256×256×9 dimensional fused feature vector is flattened into a 589824 dimensional one-dimensional vector, and the standard features are also flattened into one-dimensional vectors. The cosine similarity between the two vectors is then calculated. The cosine similarity value ranges from [-1, 1], with values ​​closer to 1 indicating greater similarity between the two feature vectors. For the inspection of a specific building's exterior wall, the cosine similarity between the fused feature vector and 15 standard defect features was as follows: surface cracks 0.82, structural cracks 0.93, weathering and spalling 0.75, material aging 0.68, concrete carbonation 0.62, steel reinforcement corrosion 0.56, component deformation 0.48, hollowing and detachment 0.65, water seepage and leakage 0.71, thermal bridging 0.59, insulation layer damage 0.64, coating peeling 0.69, salt precipitation damage 0.54, dirt contamination 0.63, and joint damage 0.67. Structural cracks showed the highest similarity, reaching 0.93; therefore, structural cracks were used as the defect category identifier for this area.

[0038] After similarity matching, spatial structure analysis is performed on the fused feature vector to identify spatial regions where feature activation values ​​are continuously distributed. Feature activation values ​​refer to the numerical values ​​of each element in the fused feature vector, representing the intensity of the defect feature at that location. An activation threshold of 0.75 is set; a feature value greater than 0.75 is considered to indicate the presence of a defect feature activation at that location. For a 256×256×9 fused feature vector, the average of the nine channels is first calculated to obtain a 256×256 feature activation map. A region growing algorithm is then applied to this activation map to identify continuously distributed high-activation regions. The seed point for region growing is selected from the location with the highest activation value in the activation map. The expansion condition is that the activation values ​​of adjacent points are greater than the threshold and the difference in activation values ​​is less than 0.1. Region growing stops when the expansion condition is not met. For the building's exterior wall, a spatial region with continuously distributed feature activation values ​​was identified, with an area of ​​1283 pixels.

[0039] The feature activation region is binarized to generate a binary mask, and a contour tracking algorithm is applied to extract the region boundary point set. For this detection, contour tracking yields 78 boundary point coordinates, forming a closed polygonal contour. The boundary point set is simplified, retaining key inflection points, resulting in 12 simplified boundary point coordinates. The simplified boundary coordinates are represented in feature space as: (124, 87), (138, 92), (145, 103), (148, 117), (143, 128), (135, 135), (122, 138), (109, 133), (102, 123), (100, 111), (104, 98), (114, 91).

[0040] The geometric center coordinates were obtained by averaging the coordinates of the boundary points. For this detection, the average coordinates of the 12 boundary points were (124, 113), which was used as the geometric center coordinates of the defect area. To improve the accuracy of the geometric center, the distribution weight of the activation values ​​was considered, and a weighted geometric center was calculated. Based on the activation value of each point within the region as the weight, the weighted average geometric center coordinates were calculated to be (125, 114), which is close to the result obtained by simple averaging, indicating that the activation value distribution in the defect area is relatively uniform.

[0041] Boundary coordinates and geometric center coordinates are pixel coordinates in feature space, which need to be converted to the physical coordinate system of the building's exterior surface. During the conversion, camera parameters and shooting distance are considered to establish a mapping relationship between pixel coordinates and physical coordinates. The camera focal length is 24mm, the sensor size is 36mm×24mm, the shooting distance is 5m, and the pixel resolution is 4096×3072. Based on the principle of similar triangles, the ratio of pixel to actual physical size is calculated to be 1:4, meaning one pixel in feature space corresponds to 4mm in actual physical space. Simultaneously, camera lens distortion correction is considered, using a polynomial distortion model with distortion coefficients k1=-0.02 and k2=0.004.

[0042] Using the upper left corner of the building's exterior wall as the origin of the physical coordinate system, with the horizontal direction as the x-axis and the vertical direction as the y-axis, the units are millimeters. For this detection, the boundary point coordinates in the feature space are converted to the boundary positions in the physical coordinate system as follows: (496, 348), (552, 368), (580, 412), (592, 468), (572, 512), (540, 540), (488, 552), (436, 532), (408, 492), (400, 444), (416, 392), (456, 364), all in millimeters. These coordinates constitute the physical boundary of the defect area on the building's exterior wall surface.

[0043] The geometric center coordinates (125, 114) in feature space are converted to the center position (500, 456) in the physical coordinate system, in millimeters. This position is 500mm horizontally and 456mm vertically from the upper left corner of the building's exterior wall, representing the center position of the defect area.

[0044] The spatial location information output by the defect identification module, including boundary and center locations, is transmitted to subsequent processing modules in a standardized data format for defect visualization and assessment report generation. The spatial location information is encapsulated in JSON data format, containing the defect type, an array of boundary point coordinates, and center point coordinates; the data size is approximately 2KB.

[0045] This invention achieves accurate identification of building exterior wall defect types by fusing feature vectors with similarity matching of standard defect features, and precisely locates the physical position of defects through spatial structure analysis and coordinate transformation. It overcomes the limitations of traditional single-modal recognition methods, improving the accuracy and reliability of defect identification by utilizing multi-modal fusion features. Through a matching mechanism based on a pre-set defect category feature library, the system can identify various types of building exterior wall defects, adapting to different materials and environmental conditions. The spatial structure analysis algorithm effectively distinguishes defective areas from normal areas, extracting the geometric features and location information of defects. The coordinate transformation mechanism maps feature space coordinates to physical space coordinates, providing precise location references for defect assessment and repair.

[0046] The specific operation process of the spatial association module is as follows: Receive spatial location information of multiple defects output by the defect identification module; Calculate the Euclidean distance between the spatial location information of any two defects among multiple defects, use the Euclidean distance as the spatial proximity, filter defect pairs with spatial proximity less than a preset distance threshold, and mark the defects in the defect pairs as defects that satisfy the clustering condition. For the defects that meet the aggregation conditions, a connectivity analysis is performed, and the interconnected defects are grouped into the same defect set. The number of defects in the defect set and the area of ​​the spatial region covered by the defect set are counted, and the ratio of the number of defects to the area of ​​the spatial region is calculated as the spatial distribution density. The difference between the spatial distribution density and the preset benchmark density value is calculated, and the difference is exponentially calculated and normalized to obtain the spatial influence factor.

[0047] The spatial association module receives spatial location information of multiple defects output by the defect identification module. Taking the inspection of a building's exterior wall as an example, 15 defect points are identified. Each defect point includes a defect type identifier, a set of boundary location coordinates, and center location coordinates. The defect types include 5 surface cracks, 4 structural cracks, 3 weathering and peeling, 2 hollow detachment, and 1 water seepage / leakage. The center coordinates of the defects are as follows: surface cracks (500, 456) mm, (650, 490) mm, (820, 510) mm, (780, 620) mm, (520, 680) mm; structural cracks (580, 520) mm, (620, 580) mm, (680, 630) mm, (750, 680) mm; weathering and spalling (900, 400) mm, (950, 450) mm, (980, 510) mm; hollow flaking (350, 550) mm, (380, 620) mm; water seepage and leakage (850, 750) mm. The boundary coordinate set is the coordinates of the vertices of a polygon, and each defect contains an average of 12 boundary point coordinates.

[0048] Spatial proximity is calculated based on the center coordinates of the defects. For the 15 defects mentioned above, the spatial proximity of 105 defect pairs needs to be calculated. When calculating the Euclidean distance, the difference between the horizontal and vertical distances is considered to obtain the straight-line distance between the two points. For example, the Euclidean distance between surface crack 1 (500, 456) mm and structural crack 1 (580, 520) mm is 103.08 mm; the Euclidean distance between surface crack 1 (500, 456) mm and surface crack 2 (650, 490) mm is 155.24 mm; and the Euclidean distance between structural crack 1 (580, 520) mm and structural crack 2 (620, 580) mm is 75.00 mm.

[0049] Different distance thresholds are set for different types of building exterior wall defects: 150mm for structural cracks, 120mm for surface cracks, 100mm for weathering and peeling, 80mm for hollow areas and detachments, and 200mm for water seepage and leakage. If the Euclidean distance between two defects is less than the corresponding preset distance threshold, they are marked as defect pairs that meet the aggregation condition. In this example, the spatial proximity of defect pairs such as surface crack 1 and structural crack 1, structural crack 1 and structural crack 2, weathering and peeling 1 and weathering and peeling 2, and hollow areas and detachments 1 and hollow areas and detachments 2 is less than their respective preset distance thresholds; therefore, these defects are marked as defects that meet the aggregation condition.

[0050] Connectivity analysis employs the connected component algorithm from graph theory, treating each defect as a node in the graph and establishing edges between defect pairs that satisfy the clustering condition. A depth-first search traverses the entire graph to identify all connected components, each of which represents a defect set. A disjoint-set data structure is used for efficient connectivity analysis. For the aforementioned 15 defect points, connectivity analysis yields 5 defect sets: set 1 contains 7 defects; set 2 contains 2 defects; set 3 contains 3 defects; set 4 contains 2 defects; and set 5 contains 1 defect.

[0051] The area of ​​the spatial region covered by the defect set is obtained by calculating the convex hull of the defect set. The coordinates of the boundary points of all defects within the set are merged, and the Graham scan algorithm is used to calculate the convex hull of these points. The area of ​​the convex hull polygon is the area of ​​the spatial region covered by the defect set. For defect set 1, which contains 7 defects, the convex hull area is 98500 mm². 2 Defect set 2 contains 2 defects with a convex hull area of ​​12300 mm². 2 Defect set 3 contains 3 defects with a convex hull area of ​​14600 mm². 2 Defect set 4 contains 2 defects with a convex hull area of ​​7800 mm². 2 Defect set 5 contains 1 defect with a convex hull area of ​​3200 mm². 2 .

[0052] The ratio of the number of defects to the area of ​​the spatial region is used as the spatial distribution density, with units of defects / m². 2 For ease of calculation and comparison, the area unit is converted to square meters. The spatial distribution density of defect set 1 is 7 / (0.0985) = 71.07 defects / m². 2 The spatial distribution density of defect set 2 is 2 / (0.0123) = 162.60 defects / m. 2 The spatial distribution density of defect set 3 is 3 / (0.0146) = 205.48 defects / m. 2 The spatial distribution density of defect set 4 is 2 / (0.0078) = 256.41 defects / m. 2 The spatial distribution density of defect set 5 is 1 / (0.0032) = 312.50 defects / m. 2 .

[0053] The preset baseline density value is a reference value determined comprehensively based on the building's exterior wall material, age, and environmental conditions. For concrete exterior walls, the preset baseline density value is 100 particles / m². 2 For brick and stone exterior walls, the preset baseline density value is 80 particles / m². 2 For metal exterior walls, the preset baseline density value is 50 particles / m². 2For glass curtain walls, the preset baseline density value is 30 units / m². 2 In this example, the building's exterior walls are made of concrete, with a preset baseline density of 100 particles / m³. 2 .

[0054] Calculate the density difference for each defect set: The density difference for defect set 1 is 71.07 - 100 = -28.93 defects / m². 2 The density difference of defect set 2 is 162.60 - 100 = 62.60 defects / m². 2 The density difference of defect set 3 is 205.48 - 100 = 105.48 defects / m. 2 The density difference of defect set 4 is 256.41 - 100 = 156.41 defects / m. 2 The density difference of defect set 5 is 312.50 - 100 = 212.50 defects / m. 2 .

[0055] The density difference is exponentially calculated using a curve transformation similar to the Sigmoid function, mapping the density difference to the (0, 1) interval. A negative density difference indicates that the defect density is lower than the baseline value, with a small spatial impact, and the spatial impact factor is close to 0. A positive density difference indicates that the defect density is higher than the baseline value, with a larger spatial impact, and the spatial impact factor is close to 1. For the five defect sets in this example, the calculated spatial impact factors are: 0.28 for defect set 1, 0.65 for defect set 2, 0.78 for defect set 3, 0.85 for defect set 4, and 0.92 for defect set 5. Normalization ensures that the sum of all spatial impact factors is 1, resulting in the following normalized spatial impact factors: 0.08 for defect set 1, 0.19 for defect set 2, 0.22 for defect set 3, 0.24 for defect set 4, and 0.27 for defect set 5.

[0056] like Figure 2 The diagram shows a comparative analysis of the spatial distribution density of defect sets in this invention. Although defect set 5 (water seepage / leakage) contains only one defect, its calculated area is based on a compact convex hull (0.0032m²). 2 This results in a density as high as 312.50 particles / m³. 2 Significantly higher than the benchmark value of 100 / m 2 In contrast, while defect set 1 contains the most defects (7), their distribution is more scattered, with an area of ​​0.0985 m². 2 The calculated density is 71.07 cells / m². 2 If the rectangular bounding box calculation is used according to existing technology, the bounding box area is usually larger than the convex hull area (for example, the bounding box area of ​​set 1 may reach 0.12m). 2This would cause the calculated density to be diluted (e.g., reduced to approximately 58 particles / m³). 2 This underestimates the severity of defects in the area. This technical solution uses precise convex hull area calculations to obtain density values ​​(e.g., 256.41 units / m² for set 4). 2 This more accurately reflects the degree of stress concentration or damage density in local areas. The figure clearly shows that the densities of sets 2, 3, 4, and 5 all exceed the preset benchmark density of concrete exterior walls (100 particles / m²). 2 This indicates that these areas require special attention.

[0057] This invention provides an important reference for defect risk assessment by analyzing the spatial distribution characteristics of defects in building exterior walls. Through calculations of spatial proximity, connectivity analysis, and spatial distribution density among defects, it accurately identifies defect clusters and quantifies the potential impact of defect clusters on building safety using spatial impact factors. This technology overcomes the limitations of traditional defect assessment methods that focus only on individual defects while ignoring the spatial distribution characteristics of defect groups, effectively identifying high-risk defect clusters.

[0058] The specific operation process of the quantitative evaluation module is as follows: Receive the fused feature vector and defect category identifier output by the defect identification module and the spatial influence factor output by the spatial association module; Semantic segmentation is performed on the fused feature vector based on the defect category identifier, and the feature dimension components that are semantically associated with the defect category identifier are extracted from the fused feature vector. The feature dimension components are then combined into a feature subset. Calculate the mean and variance of each feature value within the feature subset, and combine the mean and variance into a numerical distribution statistic; The numerical distribution statistics and spatial influence factors are subjected to tensor product operation. The result is linearly transformed and the horizontal component is extracted as the size correlation quantity, and the vertical component is extracted as the depth correlation quantity. Based on the defect category identifier, the corresponding size conversion coefficient and depth conversion coefficient are determined. The size correlation quantity is scaled according to the size conversion coefficient to obtain the physical size. The depth correlation quantity is scaled according to the depth conversion coefficient to obtain the damage depth. The physical size and damage depth are combined into a quantitative evaluation parameter.

[0059] The quantitative assessment module receives the fused feature vector and defect category identifier from the defect identification module, as well as the spatial influence factor from the spatial association module. In the assessment of building exterior wall defects, the defect identification module outputs a fused feature vector with 256 dimensions using image processing and deep learning methods, containing multimodal information such as image texture features, edge features, color features, and depth features. The defect category identifier includes five types: cracks, peeling, leakage, hollowness, and weathering. The spatial influence factor output by the spatial association module is a scalar value ranging from 0 to 1, representing the degree of clustering of defects in spatial distribution. Taking a crack defect in a building exterior wall as an example, the fused feature vector contains information such as the crack's width, length, depth, shape, and surface texture, with a spatial influence factor of 0.78, indicating that the crack forms a high degree of clustering with other surrounding defects.

[0060] Semantic segmentation is performed on the fused feature vector based on the defect category identifier. This semantic segmentation process is implemented using a feature mapping table, which predefines the semantic association dimensions of various defects in the fused feature vector. For crack-type defects, dimensions 10-65 are extracted from the fused feature vector. These features mainly include information such as linear structure, continuity, and directionality. For spalling-type defects, dimensions 70-120 are extracted. These features mainly include information such as material loss, surface irregularity, and texture variation. For leakage-type defects, dimensions 125-170 are extracted. These features mainly include information such as moisture distribution, color variation, and humidity gradient. For hollow-type defects, dimensions 175-210 are extracted. These features mainly include information such as sound wave reflection, surface vibration, and material density. For weathering-type defects, dimensions 215-250 are extracted. These features mainly include information such as material aging, granulation, and surface roughness. In the crack defect example above, dimensions 10-65 of the feature vector are extracted, forming a 56-dimensional feature subset.

[0061] For the crack feature subset, the arithmetic mean of the 56 feature values ​​is calculated, yielding 0.68, representing the overall significance of the cracks; the variance of these 56 feature values ​​is calculated, yielding 0.12, representing the internal consistency of the crack features. The mean and variance are combined to form a two-dimensional vector [0.68, 0.12], which serves as the numerical distribution statistic. For the spalling feature subset, the numerical distribution statistic is [0.52, 0.25]; for the leakage feature subset, the numerical distribution statistic is [0.81, 0.09]; for the hollow feature subset, the numerical distribution statistic is [0.43, 0.31]; and for the weathering feature subset, the numerical distribution statistic is [0.37, 0.28].

[0062] The numerical distribution statistics and the spatial influence factor are multiplied by a tensor product. This tensor product multiplies the two-dimensional numerical distribution statistics by the scalar spatial influence factor, forming a new two-dimensional tensor. In the crack example, [0.68, 0.12] × 0.78 yields [0.5304, 0.0936]. This result is then linearly transformed using a pre-defined 2×2 matrix, determined based on the building material type and structural characteristics. The transformed result is a two-dimensional vector [0.8125, 0.2751]. The horizontal component 0.8125 serves as the size correlation quantity, representing the crack's size information; the vertical component 0.2751 serves as the depth correlation quantity, representing the crack's depth information.

[0063] Different defect types have different conversion coefficients, which are determined by the characteristics of building materials and historical data analysis. For crack defects in concrete exterior walls, the size conversion coefficient is 5.0 mm, representing the mapping of size correlation quantities to physical length units; the depth conversion coefficient is 20.0 mm, representing the mapping of depth correlation quantities to physical depth units. In the crack example, the size correlation quantity 0.8125 × the size conversion coefficient 5.0 mm yields a physical size of 4.0625 mm, representing the crack width; the depth correlation quantity 0.2751 × the depth conversion coefficient 20.0 mm yields a damage depth of 5.502 mm, representing the crack depth. The combination of physical size and damage depth forms the quantitative evaluation parameters [4.0625 mm, 5.502 mm].

[0064] The conversion factors differ depending on the type of defect. For spalling defects, the size conversion factor is 30.0 mm and the depth conversion factor is 15.0 mm; for leakage defects, the size conversion factor is 50.0 mm and the depth conversion factor is 10.0 mm; for hollow defects, the size conversion factor is 100.0 mm and the depth conversion factor is 40.0 mm; and for weathering defects, the size conversion factor is 80.0 mm and the depth conversion factor is 5.0 mm. These factors were determined based on extensive experimental data and professional engineering experience to ensure a high degree of consistency between the quantitative assessment results and the actual physical measurement results.

[0065] The quantitative assessment results can be further used for defect repair decisions and building safety assessments. Based on physical dimensions and damage depth, combined with principles of materials mechanics and structural engineering, the impact of defects on structural strength can be calculated. For example, a crack with a width of 4.0625 mm and a depth of 5.502 mm on a concrete exterior wall is classified as moderately dangerous and requires immediate repair; while a spalling defect with a diameter of 52 mm and a depth of 12 mm on a masonry exterior wall is classified as highly dangerous and requires immediate repair and reinforcement.

[0066] This invention achieves precise quantification of building exterior wall defects through semantic segmentation and statistical analysis of multimodal fusion features, providing a scientific basis for defect severity assessment and repair plan formulation. This technology organically combines visual features with spatial correlation characteristics, overcoming the limitations of traditional assessment methods that rely solely on single visual features, and comprehensively reflects the physical characteristics and potential risks of defects. Through a pre-set feature mapping table and conversion coefficient system, the quantitative assessment process considers the specificities of different types of defects while maintaining the consistency and comparability of assessment standards. It effectively integrates the individual characteristics and group distribution characteristics of defects into a unified assessment system, improving the accuracy and reliability of the assessment results.

[0067] The specific operation process of the level determination module is as follows: Receive the quantitative evaluation parameters output by the quantitative evaluation module and the defect category identifier output by the defect identification module; The defect category identifier is converted into a category code value, and the quantitative evaluation parameter is vector-concatenated with the category code value to form an evaluation vector. Extract the standard vectors of each level sequentially from the preset set of level standards, calculate the Euclidean distance between the evaluation vector and each level standard vector, and establish a mapping relationship between each Euclidean distance and the corresponding level standard vector; In the mapping relationship, the Euclidean distance with the smallest value is selected, the level corresponding to the smallest Euclidean distance is extracted, and the level is used as the defect level determination result.

[0068] The grade determination module receives the quantitative assessment parameters output by the quantitative evaluation module and the defect category identifier output by the defect identification module. The quantitative assessment parameter is a two-dimensional vector containing two components: physical size and damage depth. The defect category identifier is a textual description of the defect type. Taking a building exterior wall crack defect as an example, the quantitative assessment parameters are [4.0625mm, 5.502mm], representing the crack width and depth, respectively; the defect category identifier is "structural crack". Different types of defects have different ranges of quantitative assessment parameters. For example, the width of a concrete exterior wall structural crack is typically between 0.1mm and 10.0mm, and the depth is between 1.0mm and 30.0mm; while the area of ​​weathering and spalling is typically around 10.0mm. 2 Up to 500.0mm 2 The depth is between 0.5mm and 15.0mm.

[0069] Defect category identifiers are converted into category codes, which are implemented using a mapping table that predefines numerical codes corresponding to various defect types. The category code value for structural cracks is 1.0; for surface cracks, it is 2.0; for weathering and spalling, it is 3.0; for hollowing and detachment, it is 4.0; and for water seepage and leakage, it is 5.0. The category code values ​​are set considering the degree of danger and structural impact of different types of defects; the smaller the code value, the greater the potential danger. The quantitative evaluation parameters [4.0625mm, 5.502mm] are concatenated with the category code value 1.0 to obtain a three-dimensional evaluation vector [4.0625mm, 5.502mm, 1.0]. This evaluation vector comprehensively expresses the type, size, and depth information of the defect, providing a complete feature description for subsequent grade determination.

[0070] The preset grading standard set is determined based on building structure safety codes and engineering practice experience, and includes four levels: Level 1 (Severe), Level 2 (Moderate), Level 3 (Slight), and Level 4 (Safe). Each level corresponds to a standard vector, which has the same dimension and unit as the evaluation vector. For defects in concrete exterior walls, the standard vector for Level 1 (Severe) is [5.0mm, 10.0mm, 1.0], indicating that a structural crack width of 5.0mm and a depth of 10.0mm is considered severe; the standard vector for Level 2 (Moderate) is [3.0mm, 6.0mm, 1.0], indicating that a structural crack width of 3.0mm and a depth of 6.0mm is considered moderate; the standard vector for Level 3 (Slight) is [1.0mm, 3.0mm, 1.0], indicating that a structural crack width of 1.0mm and a depth of 3.0mm is considered slight; and the standard vector for Level 4 (Safe) is [0.1mm, 1.0mm, 1.0], indicating that a structural crack width of less than 0.1mm and a depth of less than 1.0mm is considered safe.

[0071] When calculating Euclidean distance, the units and weights of each dimension are considered to ensure comparability between different physical quantities. For the evaluation vector [4.0625mm, 5.502mm, 1.0], its Euclidean distance to the first-order standard vector [5.0mm, 10.0mm, 1.0] is calculated to be 4.65; its Euclidean distance to the second-order standard vector [3.0mm, 6.0mm, 1.0] is calculated to be 1.53; its Euclidean distance to the third-order standard vector [1.0mm, 3.0mm, 1.0] is calculated to be 3.83; and its Euclidean distance to the fourth-order standard vector [0.1mm, 1.0mm, 1.0] is calculated to be 6.51. The Euclidean distance calculation process first requires standardizing the vector elements to eliminate the influence of dimensions on the calculation results. Standardization can be achieved using normalization or standard fraction conversion methods to map different physical quantities to the same numerical range.

[0072] In the example above, the Euclidean distances between the evaluation vector and the four standard vectors are 4.65, 1.53, 3.83, and 6.51, respectively, with the minimum value of 1.53 corresponding to a level two (moderate) standard vector. Therefore, the defect level of the structural crack is determined to be level two (moderate), indicating that the defect poses a moderate level of safety risk and requires repair in the near future.

[0073] The standard vectors in the preset grade standard set differ for different types of defects. For surface cracks, the first-level standard vector is [2.0mm, 5.0mm, 2.0], the second-level standard vector is [1.0mm, 3.0mm, 2.0], the third-level standard vector is [0.5mm, 1.5mm, 2.0], and the fourth-level standard vector is [0.1mm, 0.5mm, 2.0]. The first-level standard vector for weathering and flaking is [300.0mm]. 2 [10.0mm, 3.0], the second-order standard vector is [200.0mm]. 2 [7.0mm, 3.0], the third-order standard vector is [100.0mm]. 2 [4.0mm, 3.0], the fourth-order standard vector is [50.0mm]. 2 [2.0mm, 3.0]. The first-order standard vector for hollow detachment is [500.0mm]. 2 [15.0mm, 4.0], the second-order standard vector is [300.0mm]. 2 [10.0mm, 4.0], the third-order standard vector is [150.0mm]. 2 [5.0mm, 4.0], the fourth-order standard vector is [50.0mm]. 2 [2.0mm, 4.0]. The first-order standard vector for water seepage is [200.0mm]. 2 [5.0mm, 5.0], the second-order standard vector is [100.0mm]. 2 [3.0mm, 5.0], the third-level standard vector is [50.0mm]. 2 [1.5mm, 5.0], the fourth-order standard vector is [10.0mm] 2 , 0.5mm, 5.0].

[0074] The grading results can be used for building exterior wall maintenance decisions and safety assessments. Level 1 (Severe) defects indicate significant structural safety hazards requiring immediate repair and reinforcement; Level 2 (Moderate) defects indicate some safety risk requiring near-term repair; Level 3 (Minor) defects indicate low safety risk and can be addressed during routine maintenance; Level 4 (Safe) defects indicate no impact on building safety and can be monitored without immediate action. The grading results are considered in conjunction with factors such as the building's usage, age, and importance to form the final maintenance recommendations.

[0075] For specific building materials and usage environments, the preset set of grade standards can be customized. For historical buildings or special structures, the thresholds in the standard vector can be fine-tuned based on the experience of professional engineers and the results of on-site surveys, making the grade determination more in line with the actual situation. For example, for concrete structures in high-humidity environments, the grade threshold for water seepage defects can be appropriately lowered to reflect the additional impact of humidity on structural materials; for buildings in seismic zones, the grade threshold for structural crack defects can be increased to reflect the potential threat of seismic activity to structural safety.

[0076] This invention achieves an objective and quantitative assessment of the severity of defects in building exterior walls through a vector space metric method, overcoming the subjectivity and inconsistency of traditional experience-based judgments. This technology organically integrates multi-dimensional features such as defect type, physical dimensions, and damage depth, establishing a unified evaluation standard system to ensure the comparability of different types of defects and the consistency of evaluation results. By using preset grade standard vectors and Euclidean distance calculations, it can accurately capture the differences between defect characteristics and standard thresholds, providing fine-grained grade classification.

[0077] The specific embodiments described above are preferred embodiments of the present invention and are not intended to limit the specific scope of the present invention. The scope of the present invention includes, but is not limited to, these specific embodiments. All equivalent changes made in accordance with the shape and structure of the present invention are within the protection scope of the present invention.

Claims

1. A system for identifying and quantitatively assessing building exterior wall defects based on multimodal fusion, characterized in that, include: The feature extraction module is used to acquire multimodal detection data of building exterior walls and extract features from the multimodal detection data to obtain feature representations of each modality. The feature fusion module is used to project the modal feature representations onto the same feature dimension space, calculate the response consistency of each modal feature representation, assign fusion weights according to the response consistency and perform weighted concatenation to obtain a fused feature vector; The defect identification module is used to identify defects based on the fused feature vector and obtain defect category identifiers and spatial location information; The spatial association module is used to calculate the spatial proximity of multiple defects based on the spatial location information, divide the defects whose spatial proximity satisfies the clustering condition into a defect set, calculate the spatial distribution density of the defect set and generate a spatial influence factor. The quantitative evaluation module is used to extract a feature subset corresponding to the defect category identifier from the fused feature vector, calculate the numerical distribution statistics of the feature subset, and weight the numerical distribution statistics with the spatial influence factor and map them to physical size and damage depth to obtain quantitative evaluation parameters. The level determination module is used to combine the quantitative evaluation parameters with the defect category identifier into a judgment vector, calculate the distance between the judgment vector and each level standard vector in the preset level standard set, and select the level corresponding to the smallest distance as the defect level determination result.

2. The system according to claim 1, characterized in that, The specific operation process of the feature extraction module is as follows: Acquire multimodal detection data of building exterior walls, including visual detection data of the building exterior wall surface and physical detection data of the building exterior wall interior; Noise filtering is performed on each modal data in the multimodal detection data to obtain filtered modal data. Multi-level encoding operations are performed on the filtered modal data, feature responses are extracted at each encoding level, and the feature responses at each encoding level are concatenated to obtain a multi-level feature representation. The multi-level feature representations are subjected to scale normalization to obtain normalized feature representations; The normalized feature representation is compared with the preset feature template by dimension-wise similarity calculation. Feature dimensions with similarity higher than the preset screening threshold are selected. Feature components corresponding to the feature dimensions are extracted from the normalized feature representation to form the feature representations of each modality.

3. The system according to claim 1, characterized in that, The specific operation process of the feature fusion module is as follows: Receive modal feature representations output by the feature extraction module; Each modal feature representation is subjected to dimensional transformation, and each modal feature representation is projected onto the same feature dimension space to obtain the projected modal feature representations. Calculate the feature response differences between the modal feature representations after projection in the corresponding regions of spatial locations, identify the sensitivity of each mode to building exterior wall defects based on the feature response differences, and determine the response consistency of each modal feature representation based on the sensitivity. After projection, feature regions whose feature activation intensity exceeds a preset activation threshold are extracted from each modal feature representation. The feature numerical statistics within the feature region are calculated as signal intensity. The response consistency and the signal intensity are normalized and weighted to obtain the fusion weight corresponding to each modal feature representation. The modal feature representations after projection are weighted element-wise according to the fusion weights, and the weighted modal feature representations are concatenated to obtain the fused feature vector.

4. The system according to claim 1, characterized in that, The specific operation process of the defect identification module is as follows: Receive the fused feature vector output by the feature fusion module; The fused feature vector is matched with the standard features of each defect category in the preset defect category feature library. The defect category corresponding to the standard feature with the highest similarity is selected as the defect category identifier. Spatial structure analysis is performed on the fused feature vector to identify spatial regions in which the feature activation values ​​in the fused feature vector are continuously distributed, and the boundary coordinates and geometric center coordinates of the spatial regions are calculated. Convert the boundary coordinates to the boundary position in the physical coordinate system of the building's exterior wall surface, and convert the geometric center coordinates to the center position in the physical coordinate system of the building's exterior wall surface. Output the boundary and center positions as spatial location information.

5. The system according to claim 1, characterized in that, The specific operation process of the spatial association module is as follows: Receive spatial location information of multiple defects output by the defect identification module; Calculate the Euclidean distance between the spatial location information of any two defects among multiple defects, use the Euclidean distance as the spatial proximity, filter defect pairs with spatial proximity less than a preset distance threshold, and mark the defects in the defect pairs as defects that satisfy the clustering condition. For the defects that meet the aggregation conditions, a connectivity analysis is performed, and the interconnected defects are grouped into the same defect set. The number of defects in the defect set and the area of ​​the spatial region covered by the defect set are counted, and the ratio of the number of defects to the area of ​​the spatial region is calculated as the spatial distribution density. The difference between the spatial distribution density and the preset benchmark density value is calculated, and the difference is exponentially calculated and normalized to obtain the spatial influence factor.

6. The system according to claim 1, characterized in that, The specific operation process of the quantitative evaluation module is as follows: Receive the fused feature vector and defect category identifier output by the defect identification module and the spatial influence factor output by the spatial association module; Semantic segmentation is performed on the fused feature vector based on the defect category identifier, and the feature dimension components that are semantically associated with the defect category identifier are extracted from the fused feature vector. The feature dimension components are then combined into a feature subset. Calculate the mean and variance of each feature value within the feature subset, and combine the mean and variance into a numerical distribution statistic; The numerical distribution statistics and spatial influence factors are subjected to tensor product operation. The result is linearly transformed and the horizontal component is extracted as the size correlation quantity, and the vertical component is extracted as the depth correlation quantity. Based on the defect category identifier, the corresponding size conversion coefficient and depth conversion coefficient are determined. The size correlation quantity is scaled according to the size conversion coefficient to obtain the physical size. The depth correlation quantity is scaled according to the depth conversion coefficient to obtain the damage depth. The physical size and damage depth are combined into a quantitative evaluation parameter.

7. The system according to claim 1, characterized in that, The specific operation process of the level determination module is as follows: Receive the quantitative evaluation parameters output by the quantitative evaluation module and the defect category identifier output by the defect identification module; The defect category identifier is converted into a category code value, and the quantitative evaluation parameter is vector-concatenated with the category code value to form an evaluation vector. Extract the standard vectors of each level sequentially from the preset set of level standards, calculate the Euclidean distance between the evaluation vector and each level standard vector, and establish a mapping relationship between each Euclidean distance and the corresponding level standard vector; In the mapping relationship, the Euclidean distance with the smallest value is selected, the level corresponding to the smallest Euclidean distance is extracted, and the level is used as the defect level determination result.