External wall visual inspection method and system based on deep learning

By acquiring a collection of exterior wall images under multiple angles and lighting conditions, and performing image preprocessing and feature recognition, the problems of low efficiency and poor accuracy in existing exterior wall detection methods are solved, and efficient and accurate exterior wall defect detection and report generation are achieved.

CN120495942BActive Publication Date: 2025-09-16WEIPAI CONSTR TECH (SHANGHAI) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510985799.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-09-16
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

Existing exterior wall inspection methods rely on manual inspection, which is inefficient and easily affected by subjective factors, resulting in uneven image quality and inaccurate defect identification, making it difficult to accurately distinguish different types of defects and determine the location of defective areas.

Method used

By acquiring a collection of exterior wall images under different angles and lighting conditions, image feature preprocessing is performed to generate standardized feature images. The pre-trained exterior wall defect detection model is called to identify defect features. The defect area is determined by combining confidence and spatial position correlation, and finally a detection report is generated.

Benefits of technology

It improves the accuracy and efficiency of exterior wall inspection, reduces maintenance costs and risks, and provides reliable defect information to trigger subsequent maintenance operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495942B_ABST
    Figure CN120495942B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for exterior wall visual inspection based on deep learning. First, an initial visual image set including multiple exterior wall image units with pixel coordinate marks collected under different shooting angles and lighting conditions is obtained, and then feature preprocessing is performed on the initial visual image set to generate a standardized feature image set. Then, a pre-trained exterior wall defect detection model is called to identify defect features on the standardized feature image set to obtain a preliminary detection result. Then, based on the confidence distribution of the defect type candidate set and the spatial position correlation of the candidate area in the preliminary detection result, a final detection result is determined, which includes a unique defect type label and the precise position coordinates of the defect area in the image coordinate system. Finally, an exterior wall defect detection report including defect type statistical information and a position distribution map is generated based on the final detection result, and sent to the target terminal device to trigger subsequent maintenance operations, thereby improving the accuracy and efficiency of exterior wall inspection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a method and system for visual inspection of exterior walls based on deep learning. Background Art

[0002] In the construction industry, exterior walls are a crucial component of buildings, and their quality is directly related to their safety and aesthetics. With the continuous development of the construction industry and the increasing scale of buildings, exterior wall structures are becoming more complex and diverse. Traditional exterior wall inspection methods rely primarily on manual visual inspection, which is not only inefficient but also subject to subjective factors, making it prone to missed inspections and false positives.

[0003] With the development of computer vision technology, image-based exterior wall inspection methods have gradually gained popularity. However, existing image-based exterior wall inspection methods have many shortcomings. First, captured exterior wall images are often affected by shooting angle and lighting conditions, resulting in varying image quality, making subsequent image processing and defect identification difficult. Second, existing inspection methods are inaccurate in defect feature extraction and identification, making it difficult to accurately distinguish different types of defects. Furthermore, they cannot precisely determine the location of defective areas, making them ineffective in providing a reliable basis for subsequent maintenance operations. Summary of the Invention

[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a method for visual inspection of exterior walls based on deep learning, the method comprising:

[0005] Acquire an initial visual image set of the exterior wall, wherein the initial visual image set includes a plurality of exterior wall image units with pixel coordinate marks collected at different shooting angles and under different lighting conditions;

[0006] Performing image feature preprocessing on the initial visual image set to generate a standardized feature image set including exterior wall surface texture features and structural contour features;

[0007] Calling a pre-trained exterior wall defect detection model to perform defect feature recognition processing on the standardized feature image set, and generating a preliminary detection result of the defect area in each exterior wall image unit, the preliminary detection result including a defect type candidate set and contour edge coordinate information of the corresponding candidate area;

[0008] Determine a final detection result of the defect area in each exterior wall image unit based on the confidence distribution of the defect type candidate set and the spatial position correlation of the candidate areas in the preliminary detection results, wherein the final detection result includes a unique defect type label and the precise position coordinates of the defect area in the image coordinate system;

[0009] An exterior wall defect detection report including defect type statistics and a location distribution map is generated based on the final detection result, and the exterior wall defect detection report is sent to a target terminal device to trigger subsequent maintenance operations.

[0010] On the other hand, an embodiment of the present invention also provides a deep learning-based exterior wall visual inspection system, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0011] Based on the above aspects, the embodiment of the present invention fully considers the complex factors in the actual detection environment by obtaining an initial visual image set including multiple exterior wall image units with pixel coordinate marks collected under different shooting angles and lighting conditions, and then performs image feature preprocessing on the initial visual image set to generate a standardized feature image set including exterior wall surface texture features and structural contour features, thereby effectively improving the image quality. The pre-trained exterior wall defect detection model is called to perform defect feature recognition processing on the standardized feature image set, and the preliminary detection results of the defect area in each exterior wall image unit can be quickly and accurately identified. According to the confidence distribution of the defect type candidate set in the preliminary detection results and the spatial position correlation of the candidate area, the final detection result of the defect area in each exterior wall image unit is determined, thereby further improving the accuracy and reliability of defect detection. Finally, based on the final detection result, an exterior wall defect detection report including defect type statistical information and a location distribution map is generated, and the exterior wall defect detection report is sent to the target terminal device to trigger subsequent maintenance operations, thereby significantly improving the efficiency and quality of exterior wall detection and reducing maintenance costs and risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a schematic diagram of the execution flow of the deep learning-based exterior wall visual detection method provided in an embodiment of the present invention.

[0013] Figure 2 Schematic diagram of exemplary hardware and software components of a deep learning-based exterior wall visual inspection system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0014] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a deep learning-based exterior wall visual inspection method provided by an embodiment of the present invention. The deep learning-based exterior wall visual inspection method is introduced in detail below.

[0015] Step S110: obtaining an initial visual image set of the exterior wall, wherein the initial visual image set includes a plurality of exterior wall image units with pixel coordinate markings collected under different shooting angles and lighting conditions.

[0016] In the application scenario of this embodiment, in order to comprehensively detect defects on the exterior wall of a building, an image acquisition device can be used to obtain an initial set of visual images of the exterior wall. For example, a drone equipped with a high-precision camera can be used as an image acquisition tool. The drone flies around the exterior wall of the building according to a preset flight path, and during the flight, it photographs the exterior wall from different heights and angles. At the same time, in order to obtain the exterior wall features under different lighting conditions, the photography work is scheduled at different time periods, such as early morning, noon, and evening. Because the light intensity and angle vary in different time periods, the collected images can contain richer information.

[0017] Each captured exterior wall image unit is labeled with pixel coordinates. This is achieved through the image acquisition device's internal system. As the image is captured, the position of each pixel in the image is automatically recorded and associated with the corresponding image data. These pixel-labeled exterior wall image units collectively constitute the initial visual image set. Each image unit in this initial visual image set contains rich pixel information that can reflect the appearance characteristics of different parts of the exterior wall.

[0018] Step S120: performing image feature preprocessing on the initial visual image set to generate a standardized feature image set including exterior wall surface texture features and structural contour features.

[0019] To improve the accuracy and reliability of subsequent exterior wall defect detection, the initial visual image set requires image feature preprocessing. The goal of image preprocessing is to enhance useful information within the image, remove noise and interference, and make the image more suitable for subsequent analysis and processing. The first step is color space conversion, converting the original RGB color space to the HSV color space. This is because in the RGB color space, color and brightness information are intertwined, making it difficult to independently process brightness and color. The HSV color space, on the other hand, separates the brightness channel from the color channel, facilitating subsequent targeted adjustments to brightness and color.

[0020] Step S121: performing color space conversion processing on each exterior wall image unit in the initial visual image set, converting the original RGB color space into the HSV color space to separate the brightness channel and the color channel.

[0021] For each exterior wall image unit in the initial visual image set, a color space conversion algorithm is used to convert the RGB color space to the HSV color space. This conversion process involves calculating and converting the RGB value of each pixel in the image, mapping it to the corresponding value in the HSV color space. Specifically, the hue, saturation, and brightness information of the color are calculated from the RGB values, thereby separating the brightness channel and color channel. During the conversion process, the accuracy and stability of the conversion algorithm must be guaranteed to ensure that the separated brightness channel and color channel accurately reflect the brightness and color characteristics of the image.

[0022] Step S122: performing histogram equalization processing on the brightness channel, enhancing the contrast between bright and dark areas in the image by redistributing pixel brightness values, and generating a brightness-enhanced brightness feature channel.

[0023] After obtaining the separated brightness channel, histogram equalization is used to enhance the contrast between light and dark areas in the image. The basic principle of histogram equalization is to redistribute the brightness values ​​of the pixels in the brightness channel to achieve a more uniform brightness distribution in the image. Specifically, the frequency of occurrence of each brightness value in the brightness channel is counted. Then, a transformation function is calculated based on the statistical results. The original brightness values ​​are mapped to a new brightness range using this transformation function. This brightens darker areas and darkens brighter areas, thereby enhancing the contrast of the image. After histogram equalization, a brightness-enhanced brightness feature channel is generated. The pixel information in this brightness feature channel can more clearly demonstrate the brightness changes on the exterior wall surface.

[0024] Step S123: performing noise filtering on the color channel, performing a convolution operation on the color channel using a Gaussian filter kernel of a preset size, and generating a color feature channel after noise suppression.

[0025] For the separated color channels, since they may be interfered with by various noises during the image acquisition process, such as electronic noise, ambient light noise, etc., these noises will affect the subsequent analysis of color features. Therefore, it is necessary to perform noise filtering on the color channels. Here, a Gaussian filter kernel of a preset size is used for convolution operation. The Gaussian filter kernel is a filter with a specific weight distribution, and its weight value is calculated according to the Gaussian function. When performing the convolution operation, the Gaussian filter kernel is covered on each pixel point of the color channel, and the pixel values ​​in its neighborhood are weighted and summed to smooth out the influence of the noise. Through multiple convolution operations, the noise in the color channel can be effectively suppressed, and the color feature channel after noise suppression can be generated, making the color information more accurate and clear.

[0026] Step S124: performing channel fusion processing on the brightness feature channel after brightness enhancement and the color feature channel after noise suppression, and generating a composite feature image containing brightness information and color information by superimposing them according to channel dimensions.

[0027] After obtaining the brightness-enhanced luminance feature channel and the noise-suppressed color feature channel, these two channels are fused. This fusion is performed by overlaying them according to the channel dimension. This combines the information at corresponding pixel positions of the luminance and color feature channels to form a composite feature image containing both luminance and color information. This combined image, which incorporates both luminance and color characteristics, can more comprehensively reflect the exterior wall's surface appearance. During the overlay process, it is necessary to ensure that the pixel sizes and positions of the two channels correspond one-to-one to ensure the accuracy of the fused image information.

[0028] Step S125: performing edge detection processing on the composite feature image to generate an edge feature map containing the structural contour information of the exterior wall surface.

[0029] To extract the structural outline of the exterior wall surface, edge detection is performed on the composite feature image. The goal of edge detection is to identify areas within the image where drastic changes in brightness or color occur. These areas typically correspond to object edges, or in other words, the structural outline of the exterior wall surface. The edge detection process consists of multiple steps. First, horizontal and vertical gradient operator templates are defined. Then, convolution operations are performed on the composite feature image to generate horizontal and vertical gradient images, respectively.

[0030] Step S1251: defining gradient operator templates in the horizontal direction and the vertical direction, and performing convolution operations on the composite feature image to obtain a horizontal gradient image and a vertical gradient image.

[0031] The horizontal and vertical gradient operator templates are pre-designed matrices that calculate the rate of change in brightness or color for each pixel in the image in the horizontal and vertical directions. Common gradient operator templates include the Sobel operator. During the convolution operation, the gradient operator template is overlaid on each pixel in the composite feature image, and the weighted sum of the pixel values ​​in its neighborhood is taken to obtain the horizontal or vertical gradient value for that pixel. By performing this convolution operation on the entire composite feature image, a horizontal gradient image and a vertical gradient image are generated. These two images respectively reflect the brightness or color changes of the composite feature image in the horizontal and vertical directions.

[0032] Step S1252: Calculate the gradient amplitude and gradient direction of each pixel point based on the horizontal gradient image and the vertical gradient image, where the gradient amplitude is the square root of the sum of the squares of the horizontal gradient value and the vertical gradient value, and the gradient direction is the inverse tangent value of the horizontal gradient value and the vertical gradient value.

[0033] After obtaining the horizontal gradient image and the vertical gradient image, it is necessary to calculate the gradient magnitude and gradient direction of each pixel. The gradient magnitude reflects the severity of the brightness or color change at the pixel, while the gradient direction indicates the direction of the change. When calculating the gradient magnitude, for each pixel, the square sum of its gradient values ​​in the horizontal and vertical directions is performed, and then the square root is taken to obtain the gradient magnitude. When calculating the gradient direction, it is obtained by calculating the inverse tangent of the horizontal gradient value and the vertical gradient value. By performing the above calculations for each pixel in the horizontal gradient image and the vertical gradient image, the gradient magnitude and gradient direction information of each pixel can be obtained. The above information constitutes a new image data set.

[0034] Step S1253: Perform non-maximum suppression on the gradient amplitude. By comparing the gradient values ​​of adjacent pixels in the gradient direction for each pixel, only the local maximum value is retained to refine the edge lines, thereby obtaining a gradient image after non-maximum suppression.

[0035] In order to refine the edge lines, the calculated gradient amplitude is subjected to non-maximum suppression. The basic idea of ​​non-maximum suppression is to compare the gradient amplitude of each pixel with that of its adjacent pixels in the gradient direction. If the gradient amplitude of the pixel is not a local maximum, its gradient amplitude is set to zero, and only the local maximum is retained. This can remove redundant information on the edge and make the edge lines more refined. In the specific operation, according to the gradient direction of each pixel, its adjacent pixels in the gradient direction are found, and then their gradient amplitudes are compared. Only the pixel with the largest gradient amplitude is retained, and the gradient amplitudes of other pixels are set to zero. By performing the above processing on the entire gradient amplitude image, a gradient image after non-maximum suppression can be obtained, and the edge lines in this gradient image are clearer and refined.

[0036] Step S1254: Binarize the gradient image after non-maximum suppression, mark pixels with gradient amplitudes greater than a gradient amplitude threshold as edge points, and generate a binary edge image containing the edge of the exterior wall surface structure contour.

[0037] After obtaining the non-maximum suppressed gradient image, it is binarized to further clarify edge points. The binarization process involves setting a gradient amplitude threshold. Pixels with gradient amplitudes greater than this threshold are marked as edge points, with their pixel values ​​set to a fixed value (e.g., maximum). Pixels with gradient amplitudes less than or equal to this threshold are marked as non-edge points, with their pixel values ​​set to another fixed value (e.g., minimum). This process converts the non-maximum suppressed gradient image into a binary edge image. This binary edge image contains only two pixel values, representing edge points and non-edge points, more clearly displaying the structural contour edges of the exterior wall surface.

[0038] Step S1255: Perform connected region analysis on the binary edge image, and mark the contour coordinate sequence of each connected edge region by detecting continuous edge points. Each contour coordinate sequence corresponds to an independent structural edge unit on the exterior wall surface, and an edge feature map containing the structural contour information of the exterior wall surface is generated.

[0039] Finally, the binary edge image is subjected to connected region analysis. The purpose of connected region analysis is to find connected regions consisting of continuous edge points in the image and mark the contour coordinate sequence of each connected region. The specific operation is to start from a certain edge point in the binary edge image, traverse its adjacent edge points, find all edge points connected to it, and form a connected region. Then, the contour coordinate sequence of the connected region is recorded. The contour coordinate sequence contains the boundary information of the connected region. By performing the above analysis on the entire binary edge image, the contour coordinate sequence of each independent structural edge unit can be obtained. These coordinate sequences together constitute the edge feature map containing the structural contour information of the exterior wall surface. This edge feature map accurately reflects the structural contour information of the exterior wall surface.

[0040] Step S126: performing texture analysis processing on the composite feature image to generate a texture feature matrix containing texture information of the exterior wall surface.

[0041] In addition to extracting the structural contours of the exterior wall surface, texture analysis is also required on the composite feature image to obtain the texture features of the exterior wall surface. The goal of texture analysis is to describe the texture information of the exterior wall surface, such as roughness and uniformity. These texture features are crucial for detecting exterior wall defects. The texture analysis process begins by setting the sliding window size and step size to cover all pixels in the composite feature image, with overlapping areas between adjacent sliding windows.

[0042] Step S1261: Set the size of the sliding window and the moving step so that the sliding window can cover all pixel areas of the composite feature image and there is an overlapping area between adjacent sliding windows, and calculate the gray level co-occurrence matrix for the pixel area in each sliding window.

[0043] The size of the sliding window and the moving step are set according to the characteristics of the composite feature image and the needs of texture analysis. The size of the sliding window determines the scope of the local area of ​​analysis, while the moving step determines the interval of window movement. By setting these two parameters reasonably, it can be ensured that the sliding window can cover all pixel areas of the composite feature image, and there is a certain overlapping area between adjacent sliding windows, thereby ensuring the continuity and integrity of the texture information. For each pixel area within the sliding window, its grayscale co-occurrence matrix is ​​calculated. The grayscale co-occurrence matrix is ​​a matrix used to describe the texture characteristics of an image. It counts the frequency of simultaneous occurrence of two pixel grayscale values ​​at a certain distance and direction. By calculating the grayscale co-occurrence matrix for the pixel area within each sliding window, a series of matrix data reflecting the texture characteristics of different local areas can be obtained.

[0044] Step S1262: For each grayscale co-occurrence matrix, the texture feature parameters are calculated respectively, and the texture feature parameters include contrast, correlation, energy and entropy, wherein the contrast measures the roughness of the texture by accumulating the product of the square of the grayscale difference and the corresponding probability, the correlation measures the spatial correlation of the texture by accumulating the joint probability of the grayscale products, the energy measures the uniformity of the texture by the sum of the squares of the grayscale joint probabilities, and the entropy measures the complexity of the texture by accumulating the logarithm of the grayscale joint probabilities.

[0045] After obtaining the gray level co-occurrence matrix of each sliding window, some texture feature parameters need to be calculated to further describe the texture information. These texture feature parameters include contrast, correlation, energy and entropy. Contrast is used to measure the roughness of the texture. It is calculated by accumulating the product of the square of the gray level difference and the corresponding probability. The larger the difference, the rougher the texture. Correlation is used to measure the spatial correlation of the texture. It is calculated by accumulating the joint probability of the gray level product, reflecting the degree of spatial correlation of the texture. Energy is used to measure the uniformity of the texture. It is calculated by the sum of the squares of the gray level joint probabilities. The larger the energy value, the more uniform the texture. Entropy is used to measure the complexity of the texture. It is calculated by accumulating the logarithm of the gray level joint probabilities. The larger the entropy value, the more complex the texture. By calculating these texture feature parameters for each gray level co-occurrence matrix, the texture feature description of each sliding window can be obtained.

[0046] Step S1263: performing averaging processing on the texture feature parameters extracted from each sliding window at multiple directions and distances to generate a texture feature vector for each pixel.

[0047] Since the texture feature parameters extracted at different directions and distances may vary, in order to obtain a more stable and accurate texture feature description, it is necessary to average the texture feature parameters extracted from each sliding window at multiple directions and distances. The specific operation is to average the texture feature parameters obtained at different directions and distances for each sliding window to obtain a comprehensive texture feature value. For each pixel, the comprehensive texture feature values ​​of the sliding window in which it is located are combined into the texture feature vector of that pixel. Through this averaging process, the fluctuation of texture features caused by changes in direction and distance can be reduced, and the accuracy of texture feature description can be improved.

[0048] Step S1264: Arrange the texture feature vectors of all pixel points in the order of pixel coordinates of the original image to generate a texture feature matrix containing the texture information of the exterior wall surface.

[0049] Finally, the texture feature vectors of all pixels are arranged in the order of the pixel coordinates of the original image to form a texture feature matrix containing the texture information of the exterior wall surface. During the arrangement process, it is necessary to ensure that the order of the texture feature vectors corresponds to the pixel coordinates of the original image one-to-one. This ensures that the texture information of the exterior wall surface is accurately reflected. This texture feature matrix contains rich texture information of the exterior wall surface.

[0050] Step S127: performing a standardization conversion and feature splicing operation on the edge feature map and the texture feature matrix to generate a standardized feature image set including the exterior wall surface texture features and the structural contour features.

[0051] After obtaining the edge feature map and texture feature matrix, normalization and feature concatenation are required to integrate these two different types of feature information. The purpose of normalization is to uniformly scale the data of the edge feature map and texture feature matrix, ensuring consistency in their numerical range and distribution, facilitating subsequent feature fusion and processing. The specific normalization method can be selected based on the characteristics of the data, such as normalization or standardization.

[0052] After the normalization conversion is complete, feature stitching is performed. This involves combining the standardized edge feature map and texture feature matrix according to predefined rules to form a new feature image set. This feature image set contains both the texture features and structural contour features of the exterior wall surface. During the stitching process, it is crucial to ensure that the correspondence between the edge feature map and the texture feature matrix is ​​correct, ensuring that the resulting stitched feature image set accurately reflects the characteristic information of the exterior wall surface.

[0053] Step S130: Calling a pre-trained exterior wall defect detection model to perform defect feature recognition processing on the standardized feature image set, and generating a preliminary detection result of the defect area in each exterior wall image unit. The preliminary detection result includes a defect type candidate set and contour edge coordinate information of the corresponding candidate area.

[0054] After obtaining a standardized feature image set containing exterior wall surface texture features and structural contour features, a pre-trained exterior wall defect detection model is used to identify these defects. This pre-trained exterior wall defect detection model is trained on a large amount of exterior wall image data. It learns the characteristic patterns of different types of exterior wall defects and can then identify defects based on the input standardized feature image set.

[0055] Step S131: Input the standardized feature image set into the feature extraction network of the exterior wall defect detection model to generate abstract feature maps of different levels. The feature extraction network includes multiple convolutional layers and pooling layers to extract local features in the standardized feature image layer by layer through the sliding window operation of the convolution kernel, and perform dimensionality reduction processing on the local features through the pooling layer.

[0056] First, the standardized feature image set is input into the feature extraction network of the exterior wall defect detection model. The feature extraction network is a crucial component of the exterior wall defect detection model and consists of multiple convolutional and pooling layers. The convolutional layer extracts local features from the input standardized feature image using a sliding window operation of the convolution kernel. A convolution kernel is a matrix with specific weights. During the sliding operation, it convolves with a local region of the input image to obtain the feature values ​​of that local region. By stacking multiple convolutional layers, local features at different levels can be gradually extracted, from basic features at the bottom layer to semantic features at the higher layer.

[0057] The pooling layer performs dimensionality reduction on the local features output by the convolutional layer. By aggregating the feature values ​​within a local region, such as taking the maximum or average value, the pooling layer reduces the dimensionality of the feature data while retaining important feature information. The pooling layer reduces the computational complexity of the model, improving its training efficiency and generalization capabilities. The feature extraction network's multiple convolutional and pooling layers ultimately generates abstract feature maps at different levels, corresponding to the underlying texture features, mid-level structural features, and high-level semantic features of the standardized feature image.

[0058] For example, step S1311: the first convolution layer uses a convolution kernel of a preset size to perform a convolution operation on the input standardized feature image, extracts the underlying basic features, and generates a first-level feature map with a preset number of channels.

[0059] In the first convolutional layer of the feature extraction network, a convolution operation is performed on the input standardized feature image using a convolution kernel of a preset size. The size and number of convolution kernels are pre-set based on the model design and actual needs. Through the convolution operation, the underlying basic features in the input image, such as simple edges, corner points, and other information, are extracted. During the convolution operation, the convolution kernel slides across the input standardized feature image. Each time it slides to a position, it multiplies the corresponding elements with the local image area at that position and sums the results to obtain an output value. As the convolution kernel continues to slide, a series of output values ​​can be generated. These output values ​​constitute the first-level feature map with a preset number of channels. Each channel in this first-level feature map represents a specific underlying basic feature. The number of channels depends on the preset number of convolution kernels. Different convolution kernels can extract different types of underlying features.

[0060] Step S1312: The first pooling layer uses the maximum pooling operation to reduce the dimensionality of the first-level feature map, retains the main features by selecting the maximum value in the pooling area, and generates a first-level pooling feature map with a size that is a set ratio of the input image.

[0061] After obtaining the first-level feature map, it is input into the first pooling layer for dimensionality reduction. The maximum pooling operation is used here. The principle of maximum pooling is to divide the first-level feature map into several non-overlapping pooling regions. For each pooling region, the maximum value is selected as the output value of that region. In this way, while reducing the size of the feature map, the main feature information within each pooling region is retained. After the maximum pooling operation, the size of the generated first-level pooling feature map is a set ratio of the size of the input first-level feature map. This can effectively reduce the amount of data and the complexity of subsequent calculations, while highlighting important feature information.

[0062] Step S1313: The second convolution layer performs a convolution operation on the first-level pooling feature map, extracts the structural features of the middle layer, and generates a second-level feature map with a channel number greater than the preset channel number.

[0063] The second convolutional layer receives the first-level pooled feature map as input and continues the convolution operation. The purpose of the convolution operation at this time is to extract mid-level structural features. Compared with the underlying basic features extracted by the first convolutional layer, mid-level structural features are more complex and abstract, and can reflect some local structural information in the exterior wall image, such as larger contours and texture patterns. During the convolution process, the number of convolution kernels used will be greater than the first layer, which makes the number of channels in the generated second-level feature map greater than the preset number of channels in the first-level feature map. More channels means that a wider variety of mid-level structural features can be extracted, thereby more comprehensively describing the characteristic information of the exterior wall image.

[0064] Step S1314: The second pooling layer performs dimensionality reduction processing on the second-level feature map and expands the receptive field to generate a second-level pooling feature map.

[0065] The second pooling layer processes the second-level feature map, not only reducing its dimensionality but also expanding its receptive field. Dimensionality reduction follows a similar approach to the first pooling layer, reducing the size of the feature map by dividing the second-level feature map into pooling regions and selecting the maximum value. Expanding the receptive field allows the pooling layer to capture a wider range of image information, allowing each output value to reflect a wider range of input features. This allows the model to consider more macroscopic structural information when processing features, resulting in a second-level pooled feature map that is smaller in size while still containing richer mid-level feature information.

[0066] Step S1315: The third convolutional layer performs a convolution operation on the second-level pooling feature map to extract high-level semantic features and generate a third-level feature map.

[0067] The third convolutional layer uses the second-level pooled feature map as input and focuses on extracting high-level semantic features. High-level semantic features provide a more abstract and meaningful description of the image, reflecting more holistic and conceptual information about the exterior wall image, such as whether the wall has certain defects. During the convolution operation, the convolution kernel further explores and integrates mid-level structural features to generate a third-level feature map containing high-level semantic information.

[0068] Step S1316: The third pooling layer performs dimensionality reduction processing on the third-level feature map to generate a third-level pooling feature map.

[0069] The third pooling layer performs dimensionality reduction on the third-level feature map. Its principles and operation are similar to those of the previous pooling layers. This dimensionality reduction reduces the amount of feature map data while retaining key information from high-level semantic features, generating a third-level pooled feature map. This pooled feature map is further reduced in size, but the feature information contained within it is more refined, helping to improve the model's computational efficiency and recognition accuracy.

[0070] Step S1317: The fourth convolutional layer performs a convolution operation on the third-level pooling feature map to extract further high-level semantic features and generate a fourth-level feature map.

[0071] The fourth convolutional layer performs convolution operations on the third-level pooled feature map, further deepening the extraction of high-level semantic features. As the number of convolutional layers increases, the model is able to achieve a deeper abstraction and understanding of image features. In this layer, more detailed and discriminative high-level semantic features are extracted from the third-level pooled feature map to generate the fourth-level feature map.

[0072] Step S1318: Output the feature maps of each level as abstract feature maps of different levels, corresponding to the bottom-level texture features, middle-level structural features and high-level semantic features in the standardized feature image.

[0073] Finally, the first-level feature map, first-level pooled feature map, second-level feature map, second-level pooled feature map, third-level feature map, third-level pooled feature map, and fourth-level feature map are output as abstract feature maps at different levels. These feature maps correspond to the low-level texture features, mid-level structural features, and high-level semantic features in the standardized feature image, respectively. The low-level texture feature map contains simple texture and basic feature information of the exterior wall surface, the mid-level structural feature map reflects the local structure and more complex texture patterns of the exterior wall, and the high-level semantic feature map provides abstract information about the overall exterior wall condition and possible defect types.

[0074] Step S132: Input the abstract feature maps of different levels into the feature fusion module, and use a top-down and bottom-up combined path to perform feature fusion to generate a fused feature map containing multi-scale context information. The top-down path combines high-level semantic features with low-level detail features, and the bottom-up path gradually abstracts low-level features into high-level features.

[0075] The abstract feature maps of different levels generated above are input into the feature fusion module. The feature fusion module adopts a unique fusion method, namely, a combination of top-down and bottom-up paths for feature fusion. In the top-down path, high-level semantic features can be combined with low-level detail features. High-level semantic features have more macroscopic and abstract information but lack details, while low-level detail features contain rich specific information but lack overall semantic understanding. By combining the two, high-level semantic features can obtain more detailed support, while low-level detail features can have clearer semantic meaning. For example, when identifying exterior wall defects, high-level semantic features may indicate the presence of a certain type of defect, but cannot accurately point out the specific location and boundaries of the defect, while low-level detail features can provide this specific information. The combination of the two can more accurately locate and describe the defect.

[0076] The bottom-up approach gradually abstracts low-level features into high-level features. Starting with bottom-level texture features and mid-level structural features, these low-level features are continuously abstracted and refined through a series of processing and integration, gradually acquiring the properties of high-level semantic features. During the feature fusion process, feature maps at different levels are processed and matched accordingly, then fused according to predefined rules to ultimately generate a fused feature map containing multi-scale contextual information. This fused feature map integrates feature information from different levels and scales, providing a more comprehensive and accurate description of the exterior wall image.

[0077] Step S133: Input the fused feature map into a region proposal network to generate multiple candidate regions that may contain defects, each candidate region containing position coordinates and size parameters in the image coordinate system.

[0078] The fused feature map is input into the region proposal network. The main function of the region proposal network is to find areas that may contain defects in the fused feature map. It scans and analyzes the fused feature map and predicts multiple candidate regions that may contain defects based on the information in the feature map. In the process of generating candidate regions, the region proposal network comprehensively evaluates the features in the fused feature map, taking into account factors such as the intensity and distribution of the features. For each predicted candidate region, its position coordinates and size parameters in the image coordinate system can be recorded. The position coordinates are used to determine the specific position of the candidate region in the image, and the size parameters describe the size of the candidate region. This information is very important for subsequent further processing and analysis of the candidate region. Through the processing of the region proposal network, multiple candidate regions that may contain defects are obtained, which narrows the scope that requires further analysis and improves the efficiency of defect detection.

[0079] Step S134: perform feature cropping and scaling processing on each candidate region, adjust the feature map of the candidate region to a fixed size of a preset size, and generate a feature map of the candidate region.

[0080] For each candidate region generated by the region proposal network, feature cropping and scaling are required. Feature cropping refers to extracting the feature part of the corresponding candidate region from the fused feature map, retaining only the feature information related to the candidate region, and removing other irrelevant background information. Then, the cropped feature map is scaled to adjust it to a fixed size of a preset size. The preset size is pre-set according to the design of the model and the requirements of subsequent processing. By uniformly adjusting the feature maps of the candidate regions to a fixed size, the feature maps of different candidate regions can be made consistent in size, which facilitates unified processing and analysis by the subsequent classification and regression subnetwork. After feature cropping and scaling, feature maps of each candidate region are generated. These feature maps are the same in size and contain the key feature information of the corresponding candidate region.

[0081] Step S135: Input the candidate region feature map into the classification regression subnetwork, predict the probability distribution of the candidate region belonging to different defect types through the classification branch, and predict the offset of the candidate region position coordinates and the size adjustment parameters through the regression branch.

[0082] The candidate region feature map is input into the classification and regression subnetwork. This subnetwork consists of two parts: a classification branch and a regression branch. The classification branch predicts the probability distribution of candidate regions belonging to different defect types. It analyzes and determines the features in the candidate region feature map. Based on the learned characteristic patterns of different defect types, it calculates the probability of the candidate region belonging to each defect type. For example, a candidate region may have a certain probability of being a crack defect or a certain probability of being a detachment defect. The classification branch will then provide specific probability values.

[0083] The regression branch is responsible for predicting the offset and size adjustment parameters for the candidate region's location coordinates. Because the candidate regions generated by the region proposal network may have certain positional deviations and inaccuracies in size, the regression branch uses the information in the candidate region's feature map to predict the required offset and size adjustment parameters for the candidate region's location coordinates. These predicted values ​​can be used to correct the candidate region's location and size to more accurately correspond to the actual defect area.

[0084] Step S136: Based on the probability distribution of the classification branch and the adjustment parameters of the regression branch, the candidate areas are subjected to position correction and non-maximum suppression processing, and the candidate areas with a confidence greater than a first set threshold and an overlap less than a second set threshold are retained to generate preliminary detection results of defective areas in each exterior wall image unit.

[0085] Based on the probability distribution obtained by the classification branch and the adjustment parameters obtained by the regression branch, the candidate regions are further processed. First, position correction is performed. Based on the position coordinate offset and size adjustment parameters predicted by the regression branch, the position and size of the candidate regions are adjusted to more accurately correspond to the actual defect area. Non-maximum suppression is then performed. The purpose of non-maximum suppression is to remove candidate regions with high overlap and low confidence. For each set of overlapping candidate regions, their confidence levels (i.e., the probability of a certain defect type predicted by the classification branch) are compared. Only the candidate regions with the highest confidence level are retained, while other overlapping candidate regions with lower confidence levels are removed.

[0086] During this process, two thresholds can be set: a first threshold and a second threshold. Only candidate regions with a confidence level greater than the first threshold and a degree of overlap less than the second threshold are retained. After the aforementioned screening and processing, a preliminary detection result for defect areas in each exterior wall image unit is generated. This preliminary detection result includes a set of candidate defect types and the coordinates of the corresponding candidate region's contour edges.

[0087] Step S140: Determine the final detection result of the defect area in each exterior wall image unit based on the confidence distribution of the defect type candidate set in the preliminary detection result and the spatial position correlation of the candidate area. The final detection result includes a unique defect type label and the precise position coordinates of the defect area in the image coordinate system.

[0088] Based on the preliminary detection results generated above, the final detection results for the defect areas within each exterior wall image unit need to be determined. This requires a comprehensive consideration of the confidence distribution of the defect type candidate set in the preliminary detection results and the spatial correlation of the candidate areas. The confidence distribution of the defect type candidate set reflects the likelihood that each candidate area belongs to a different defect type, while the spatial correlation of the candidate areas helps determine whether different candidate areas belong to the same defect.

[0089] Step S141: For each candidate defect area in each exterior wall image unit, extract the confidence value of each defect type in its defect type candidate set to generate a confidence vector, which reflects the confidence level of the exterior wall defect detection model that the candidate area belongs to different defect types.

[0090] For each candidate defect area in each exterior wall image unit, the confidence values ​​for each defect type in the candidate defect type set are extracted from the preliminary detection results. These confidence values ​​are arranged in a predefined order to generate a confidence vector. Each element of this confidence vector corresponds to the confidence level of a defect type, reflecting the exterior wall defect detection model's confidence that the candidate area belongs to a different defect type. For example, if the value of an element in the confidence vector is high, it indicates that the model believes the candidate area is highly likely to belong to the corresponding defect type.

[0091] Step S142: Calculate the difference between the maximum value and the second largest value in the confidence vector as a confidence interval parameter. When the confidence interval parameter is greater than a preset interval threshold, directly select the defect type as the preliminary defect type of the candidate area.

[0092] The difference between the maximum and second-largest values ​​in the confidence vector is calculated and used as the confidence interval parameter. This parameter can measure the certainty of the model's judgment on a certain defect type. If the confidence interval parameter is greater than the preset interval threshold, it means that the model's judgment on a certain defect type is relatively clear and has a high degree of certainty. In the above case, the defect type with the highest confidence is directly selected as the preliminary defect type for the candidate area. For example, if the confidence value of a defect type in the confidence vector is much greater than the confidence values ​​of other defect types, and the confidence interval parameter exceeds the preset interval threshold, then it can be considered that the candidate area is likely to belong to the defect type with the highest confidence.

[0093] Step S143: When the confidence interval parameter is less than or equal to the preset interval threshold, a secondary judgment is performed in combination with the spatial position information of the candidate area, the contour edge coordinate information of the candidate area is extracted, and the area size parameter and shape compactness parameter of the candidate area are calculated. The shape compactness parameter is calculated by the ratio of the square of the perimeter of the candidate area to the area, and is used to measure the shape regularity of the defect area.

[0094] When the confidence interval parameter is less than or equal to the preset interval threshold, the exterior wall defect detection model is uncertain about the defect type of the candidate area. Further judgment is required based on the candidate area's spatial location information. First, the candidate area's outline edge coordinates are extracted from the preliminary detection results, detailing the candidate area's boundary location.

[0095] Step S1431: For the contour edge coordinate sequence of the candidate region, the area of ​​the candidate region is calculated using a geometric integration method. The geometric integration method accurately obtains the area of ​​the region by integrating the contour coordinates, reflecting the actual size of the defect.

[0096] Based on the extracted candidate region's contour edge coordinate sequence, the geometric integration method is used to calculate the candidate region's area. Geometric integration involves performing a series of integral operations on the contour coordinates. This process divides the candidate region into multiple smaller regions based on the distribution of the contour coordinates. The area of ​​each smaller region is calculated and accumulated to accurately determine the total area of ​​the candidate region. This total area accurately reflects the actual size of the defect.

[0097] Step S1432: Use a contour tracking algorithm to calculate the distances between adjacent coordinate points along the contour edge and sum them up to obtain the perimeter of the candidate area, which is used to measure the boundary length of the defect area.

[0098] To determine the perimeter of the candidate region, a contour tracing algorithm is used. Starting from a specific coordinate point on the edge of the candidate region's contour, the algorithm sequentially calculates the distances between adjacent coordinate points along the contour edge. By analyzing the positional information of these adjacent coordinate points and using a specific distance calculation method, the distances between each segment of adjacent coordinate points are calculated. The sum of the distances between all adjacent coordinate points represents the perimeter of the candidate region. The perimeter of the candidate region can be used to measure the length of the defect boundary and reflect its size.

[0099] Step S1433: Calculate the shape compactness parameter based on the calculated area and perimeter. The shape compactness parameter is obtained by dividing the square of the perimeter by the area. A larger shape compactness indicates a narrower and longer area shape, and a smaller shape compactness indicates a closer-to-circular area shape.

[0100] After obtaining the area and perimeter of the candidate region, the shape compactness parameter is calculated. Specifically, the perimeter is squared and then divided by the area to obtain the shape compactness parameter. This shape compactness parameter is important as it measures the regularity of the defect region's shape. A large shape compactness parameter indicates that the candidate region's perimeter is large relative to its area, and the region's shape may be narrow, elongated, and irregular. A small shape compactness parameter indicates that the candidate region's shape is more circular and relatively regular.

[0101] Step S1434: normalize the area size parameter and the shape compactness parameter to facilitate matching with the parameter range in the association rule library, wherein the association rule library is established by analyzing historical detection data and contains the area size parameter range and shape compactness parameter range corresponding to different defect types.

[0102] To ensure that the calculated area size and shape compactness parameters effectively match the parameter ranges in the association rule library, these two parameters must be normalized. Normalization involves uniformly adjusting the data ranges of these two parameters to fit within an appropriate range. The association rule library, established by analyzing a large amount of historical inspection data, contains the area size and shape compactness parameter ranges corresponding to different defect types. After normalization, the area size and shape compactness parameters can be more accurately compared with the parameter ranges in the association rule library.

[0103] Step S144: Establish an association rule base between defect types and area size parameters and shape compactness parameters. The association rule base is established based on statistical analysis of historical detection data and contains parameter ranges corresponding to different defect types. The rule base is queried based on the parameters of the current candidate area to determine the preliminary defect type.

[0104] A rule library is established to associate defect types with area size parameters and shape compactness parameters. This association rule library is built based on statistical analysis of a large amount of historical inspection data. By analyzing and summarizing the area size parameters and shape compactness parameters of candidate regions for different defect types in the historical data, the corresponding parameter ranges for different defect types are determined. For example, a specific defect type may have a rough range in area size and certain characteristics in shape compactness.

[0105] The association rule library is queried based on the calculated area size and shape compactness parameters for the current candidate region. If the parameters for the current candidate region fall within the parameter range corresponding to a defect type, that defect type is determined as the preliminary defect type for the candidate region. This allows historical data experience to be leveraged to further improve the accuracy of defect type judgment.

[0106] Step S145: Perform spatial position correlation analysis on the preliminary defect types of all candidate defect areas, and calculate the ratio of the overlapping area of ​​any two candidate areas to the area of ​​the smaller candidate area as the overlap parameter to determine whether the candidate areas belong to the same defect.

[0107] Perform a spatial correlation analysis of the preliminary defect types for all candidate defect areas. Calculate the ratio of the overlapping area of ​​any two candidate areas to the area of ​​the smaller candidate area, and use this as the overlap parameter. This parameter measures the degree of overlap between two candidate areas. A large overlap parameter indicates that the two candidate areas overlap significantly in space and may represent the same defect. By calculating the overlap parameter, the spatial relationship between different candidate areas can be determined.

[0108] Step S146: When the overlap parameter is greater than a preset overlap threshold and the defect types are consistent, the two candidate regions are merged into one defect region, and its position coordinates are the union of the two candidate regions.

[0109] If the calculated overlap parameter exceeds the preset overlap threshold and the preliminary defect types of the two candidate regions are the same, it indicates that the two candidate regions are likely to belong to the same defect. In this case, the two candidate regions are merged into a single defect region. The location coordinates of the merged defect region are the union of the two candidate regions, that is, the areas covered by the two candidate regions are combined. This avoids repeated detection and recording of the same defect, improving the accuracy and simplicity of the detection results.

[0110] Step S147: When the overlap parameter is greater than the preset overlap threshold but the defect type is inconsistent, the candidate areas with confidence greater than the confidence threshold are retained, thereby generating the final detection result of the defect area in each exterior wall image unit, and each defect area contains a unique defect type label and position coordinates in the image coordinate system.

[0111] When the overlap parameter is greater than the preset overlap threshold, but the defect types of the two candidate areas are inconsistent, screening is required. Candidate areas with confidence greater than the confidence threshold are retained. The confidence threshold is a pre-set standard used to determine whether the model's judgment that the candidate area belongs to a certain defect type is reliable. By retaining candidate areas with higher confidence and removing candidate areas with lower confidence, the reliability of the final detection results can be improved. After the above series of processing and screening, the final detection results of the defect areas in each exterior wall image unit are generated. Each defect area contains a unique defect type label and position coordinates in the image coordinate system, thereby accurately describing the defect situation in each exterior wall image unit.

[0112] Step S150: generating an exterior wall defect detection report including defect type statistical information and a location distribution map based on the final detection result, and sending the exterior wall defect detection report to a target terminal device to trigger subsequent maintenance operations.

[0113] After obtaining the final detection results of the defect areas in each exterior wall image unit, a detailed exterior wall defect detection report needs to be generated based on these results. This report should include defect type statistics and location distribution maps, and the report should be sent to the target terminal device to trigger subsequent maintenance operations.

[0114] Step S151: Summarize the final detection results in all exterior wall image units, count the number of occurrences of each defect type and the total area of ​​the corresponding defect area, and generate a defect type statistical information table, which is used to display the occurrence frequency and impact range of different defect types.

[0115] First, the final inspection results for all exterior wall image units are comprehensively summarized. For each defect type, the number of times it appears across all image units is counted, which reflects the frequency of occurrence of that defect type. The total area of ​​all defect regions corresponding to each defect type is calculated, which reflects the impact of that defect type on the exterior wall. The counted number of occurrences of each defect type and the corresponding total area of ​​the defect regions are organized according to a predefined format to generate a defect type statistical information table. This defect type statistical information table allows for intuitive determination of the frequency and impact of different defect types.

[0116] Step S152: Establish a mapping relationship between the image coordinate system and the actual exterior wall physical coordinate system, and convert the position coordinates of the defect area in each exterior wall image unit into the coordinates in the actual exterior wall physical coordinate system by obtaining the intrinsic parameter matrix and extrinsic parameter matrix of the shooting device.

[0117] In order to accurately locate the location of defects on the actual exterior wall, it is necessary to establish a mapping relationship between the image coordinate system and the actual physical coordinate system of the exterior wall. This process relies on obtaining the intrinsic and extrinsic parameter matrices of the shooting device. The coordinates of the defect area position in the image coordinate system are converted into the coordinates in the actual physical coordinate system of the exterior wall through the intrinsic and extrinsic parameter matrices.

[0118] Step S1521: Obtain the intrinsic parameter matrix and extrinsic parameter matrix of the shooting device. The intrinsic parameter matrix contains the focal length, principal point coordinates and distortion parameters, which are used to describe the internal optical characteristics and imaging geometric relationship of the device. The extrinsic parameter matrix contains the rotation matrix and translation vector, which are used to describe the position and posture of the device in the world coordinate system.

[0119] First, we need to obtain the intrinsic and extrinsic parameter matrices of the camera. The intrinsic parameter matrix is ​​a crucial description of the camera's internal optical properties and imaging geometry. It contains information such as focal length, principal point coordinates, and distortion parameters. The focal length determines the camera's imaging range and magnification, the principal point coordinates determine the image center, and the distortion parameters are used to correct for any image distortion that may occur during the capture process. The extrinsic parameter matrix contains a rotation matrix and a translation vector. The rotation matrix describes the camera's rotation angle in the world coordinate system, while the translation vector describes the camera's translation position in the world coordinate system. By obtaining the intrinsic and extrinsic parameter matrices, we can accurately understand the camera's status and imaging characteristics.

[0120] Step S1522: define the actual physical coordinate system of the exterior wall, take the lower left corner of the wall as the origin, and establish a two-dimensional plane coordinate system with the horizontal rightward X axis and the vertical upward Y axis.

[0121] Next, define the physical coordinate system of the actual exterior wall. Typically, a two-dimensional coordinate system is constructed, with the lower-left corner of the wall as the origin, the horizontal rightward direction defined as the X-axis, and the vertical upward direction defined as the Y-axis. This coordinate system definition aligns with how people typically describe wall locations, making it easier to accurately map the coordinates in the image to the actual exterior wall.

[0122] Step S1523: performing distortion correction processing on each defect position coordinate in the image coordinate system using the distortion parameters in the intrinsic parameter matrix to obtain corrected coordinates.

[0123] Because the camera may have optical distortion, the coordinates in the image coordinate system may deviate from the actual physical location. Therefore, for each defect location coordinate in the image coordinate system, distortion correction is performed using the distortion parameters in the intrinsic parameter matrix. Specifically, based on the characteristics of the distortion parameters, the defect location coordinates are adjusted accordingly to eliminate the error caused by distortion, resulting in the corrected coordinates. This correction process ensures that the coordinates more accurately reflect the actual physical location.

[0124] Step S1524: converting the corrected coordinates into coordinates in the actual exterior wall physical coordinate system through an external parameter matrix.

[0125] After obtaining the corrected coordinates, they are converted to the physical coordinate system of the actual exterior wall using an extrinsic matrix. The rotation matrix in the extrinsic matrix adjusts the orientation of the corrected coordinates to align with the physical coordinate system of the actual exterior wall. The translation vector in the extrinsic matrix translates the coordinates, converting them from the relative position of the camera to the absolute position in the physical coordinate system of the actual exterior wall. Through these operations, the corrected coordinates are accurately converted to the physical coordinate system of the actual exterior wall.

[0126] Step S1525: Perform the above-mentioned conversion processing on the position coordinates of all defect areas in each exterior wall image unit to obtain a coordinate set in the physical coordinate system of the actual exterior wall. Each coordinate point in the coordinate set corresponds to an actual position on the exterior wall surface, which is used to mark the specific position of the defect on the actual exterior wall.

[0127] Finally, the aforementioned distortion correction and coordinate conversion process is performed on the coordinates of all defective areas within each exterior wall image unit. This process yields a coordinate set in the physical coordinate system of the actual exterior wall. Each coordinate point in this coordinate set corresponds to an actual location on the exterior wall surface. These coordinate points can be used to accurately identify the specific location of defects on the actual exterior wall, thus completing the mapping relationship between the image coordinate system and the physical coordinate system of the actual exterior wall.

[0128] Step S153: According to the coordinates in the physical coordinate system of the actual exterior wall, the location information of all defect areas is marked on the two-dimensional plane map to generate a location distribution map containing defect location marks. Each defect location mark includes a defect type icon and a coordinate annotation, which is used to display the distribution of defects on the actual exterior wall.

[0129] Based on the coordinates obtained previously in the physical coordinate system of the actual exterior wall, the location information of all defective areas is annotated on a two-dimensional plane map. This two-dimensional plane map is a planar representation of the actual exterior wall. During the annotation process, a mark is added to each defect location, and each mark contains a defect type icon and a coordinate annotation. The defect type icon can intuitively indicate the type of defect, for example, using different shapes or colors to distinguish different types of defects such as cracks and detachments. The coordinate annotation clearly displays the specific location coordinates of the defect on the actual exterior wall. The location distribution map generated in this way can clearly show the distribution of defects on the actual exterior wall, helping relevant personnel quickly understand the overall distribution pattern of exterior wall defects.

[0130] Step S154: performing cluster analysis on the defect positions in the position distribution map to generate defect cluster region identifiers including region boundary coordinates and defect density parameters.

[0131] Perform cluster analysis on the defect locations in the location distribution map. The purpose of cluster analysis is to identify areas where defect locations are relatively concentrated, i.e., defect clustering areas. Using a clustering algorithm, adjacent and closely spaced defect points are divided into clusters based on the coordinate information of the defect locations. For each cluster, the region boundary coordinates are determined; these coordinates describe the scope of the defect clustering area. Simultaneously, the defect density parameter within the area is calculated. The defect density parameter can be calculated as the ratio of the number of defects within the area to the area, reflecting the density of defects within the area. The region boundary coordinates and defect density parameter are combined to generate defect clustering area identifiers. These identifiers can help relevant personnel quickly identify which areas of the exterior wall have a higher concentration of defects, allowing for focused attention and treatment.

[0132] Step S155: Integrate the defect type statistical information table, location distribution map and defect cluster area identification into the preset report template to generate an exterior wall defect detection report containing text description and graphic display. The text description part is used to explain the detection results and analysis conclusions, and the graphic display part is used to present the distribution characteristics of the defects.

[0133] Integrate the previously generated defect type statistical information table, location distribution map, and defect cluster area identification into the preset report template. This report template is pre-designed and has a unified format and layout. During the integration process, the defect type statistical information table is inserted into the corresponding position of the report in the form of a table, and the location distribution map and defect cluster area identification are displayed in the report in the form of a graphic. At the same time, add a text description section to the report. The text description section details the detection results, such as the frequency of occurrence of various defect types, the scope of impact, etc., as well as the analysis conclusions based on these results, such as which defect types need to be handled first. The graphical display section uses the location distribution map and defect cluster area identification to intuitively present the distribution characteristics of the defects, making the report clearer and easier to understand.

[0134] Step S156: Send the exterior wall defect detection report to the target terminal device to trigger subsequent maintenance operations.

[0135] Finally, the generated exterior wall defect detection report is sent to the target terminal device. This target terminal device can be a mobile phone, computer, or other device used by relevant management personnel. This report promptly notifies relevant personnel of exterior wall defects, triggering subsequent maintenance actions. Based on the information in the report, relevant personnel can develop a detailed maintenance plan and arrange for repair personnel to repair the exterior wall defects, thereby ensuring the safety and normal use of the exterior wall.

[0136] In addition, the training process of the exterior wall defect detection model also includes the following steps:

[0137] First, a large amount of exterior wall image data is collected as training samples, covering different types of exterior wall defects and normal exterior walls, and accurate defect type labels and defect area location information are annotated for each training sample.

[0138] To train an accurate and effective exterior wall defect detection model, a large amount of exterior wall image data must be collected. This data should be as rich and diverse as possible, covering different types of exterior wall defects such as cracks, peeling, and hollowing, while also including image samples of normal exterior walls. When collecting data, image quality and clarity must be ensured, and images should be collected from different angles and lighting conditions to improve the model's generalization capabilities. Each collected sample must be accurately labeled with a defect type label and defect area location information. The defect type label clearly indicates the defect type present in the sample, while the defect area location information provides accurate supervision information for model training by marking the specific location coordinates of the defect in the image.

[0139] Then, the collected training samples are subjected to data augmentation processing, including random flipping, rotation, scaling and other operations, to increase the diversity and quantity of training data.

[0140] In other words, to further improve the model's generalization and robustness, data augmentation is performed on the collected training samples. There are various data augmentation methods, such as random flipping (flipping images horizontally or vertically to generate new image samples); random rotation (rotating images by a certain angle); and random scaling (enlarging or reducing images). These data augmentations increase the diversity and quantity of training data, enabling the model to learn more exterior wall features under different circumstances, thereby improving model performance.

[0141] Then, the training samples after data augmentation are divided into training set, validation set and test set. The training set is used to train the model, the validation set is used to adjust the hyperparameters of the model, and the test set is used to evaluate the final performance of the model.

[0142] After data augmentation, the training samples are divided into training, validation, and test sets according to a specific ratio. The training set is the primary data source for model training. The model iteratively learns on the training set, continuously adjusting its parameters to fit the characteristic patterns in the training data. The validation set is used to adjust the model's hyperparameters, such as the learning rate and batch size. The model's performance is evaluated on the validation set, and hyperparameters are adjusted based on the evaluation results to achieve better training results. The test set is used to evaluate the model's final performance. After model training and hyperparameter adjustment are completed, the model is tested on the test set to obtain performance metrics such as accuracy and recall on unknown data, thereby comprehensively evaluating the model's practical application capabilities.

[0143] On this basis, an exterior wall defect detection model is constructed, which includes modules such as feature extraction network, feature fusion module, region proposal network and classification regression sub-network, and the hierarchical structure and connection relationship of each module are determined.

[0144] Construct an exterior wall defect detection model, which primarily consists of modules such as a feature extraction network, a feature fusion module, a region proposal network, and a classification and regression subnetwork. The feature extraction network typically consists of multiple convolutional and pooling layers. Its hierarchical structure and convolution kernel settings should be designed based on actual needs to effectively extract features at different levels. The feature fusion module is responsible for fusing features from different levels, and its internal connections must ensure a combined top-down and bottom-up feature fusion path. The region proposal network is used to generate candidate regions that may contain defects. Its structure and parameters must be able to accurately predict the location and size of candidate regions. The classification and regression subnetwork consists of a classification branch that predicts the defect type of candidate regions, and a regression branch that corrects the location and size of candidate regions. Determine the hierarchical structure and connectivity of each module to ensure the model operates properly and achieves its intended functionality.

[0145] Next, the constructed exterior wall defect detection model is trained using the training set, and parameters such as the number of training iterations and learning rate are set. During the training process, the model parameters are continuously adjusted to minimize the loss function.

[0146] For example, a constructed exterior wall defect detection model is trained using a training set. During the training process, important training parameters must be set, such as the number of iterations and the learning rate. The number of iterations determines the number of rounds the model learns on the training set, while the learning rate controls the step size for updating the model parameters. In each iteration, the model predicts the samples in the training set based on the current parameters and then calculates the loss function between the predicted results and the labeled information. The loss function measures the difference between the model's predictions and the true labels. The model's goal is to minimize the loss function by continuously adjusting its parameters. Through multiple iterative training, the model will gradually learn the characteristic patterns of exterior wall defects, improving its defect detection accuracy.

[0147] During the training process, the validation set is used to evaluate the exterior wall defect detection model regularly. Based on the evaluation results, the hyperparameters of the exterior wall defect detection model, such as learning rate and batch size, are adjusted to improve the performance of the model.

[0148] During training, regularly evaluate the model using the validation set. Evaluation metrics can include accuracy, recall, and F1 score. Based on the evaluation results, adjust model hyperparameters, such as the learning rate and batch size. If the model's performance on the validation set does not meet expectations, you may need to lower the learning rate to make the model's parameter updates slower and more stable, or adjust the batch size to optimize model training efficiency. By continuously evaluating and adjusting hyperparameters on the validation set, you can ensure that the model performs well across different datasets and improve its generalization ability.

[0149] When the training reaches the preset number of iterations or the loss function converges to a certain degree, the test set is used to perform a final evaluation of the exterior wall defect detection model to obtain the final performance indicators of the exterior wall defect detection model, such as accuracy and recall rate.

[0150] When training reaches the preset number of iterations or the loss function converges to a certain level, the training of the exterior wall defect detection model is essentially complete. At this point, the model is finally evaluated using the test set. This test set contains data that the exterior wall defect detection model has never trained on before and can truly reflect its performance in real-world applications. By calculating metrics such as the precision and recall between the model's predictions on the test set and the true labels, the model's final performance metrics are obtained. These final performance metrics serve as an important basis for measuring model performance and determining whether the model meets the requirements of real-world applications. If the model's final performance metrics meet the expected standards, the model is considered successfully trained and can be used in real-world exterior wall defect detection tasks.

[0151] Figure 2A schematic diagram illustrates exemplary hardware and software components of a deep learning-based exterior wall visual inspection system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application. For example, a processor 120 can be used in the deep learning-based exterior wall visual inspection system 100 to perform the functions described in the present application.

[0152] The deep learning-based exterior wall visual inspection system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the deep learning-based exterior wall visual inspection method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0153] For example, the exterior wall visual inspection system 100 based on deep learning may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in different forms, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the exterior wall visual inspection system 100 based on deep learning may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The exterior wall visual inspection system 100 based on deep learning also includes an I / O interface 150 between the computer and other input and output devices.

[0154] For ease of explanation, only one processor is described in the deep learning-based exterior wall visual inspection system 100. However, it should be noted that the deep learning-based exterior wall visual inspection system 100 in this application may also include multiple processors, so the steps performed by one processor described in this application may also be performed jointly or individually by multiple processors. For example, if the processor of the deep learning-based exterior wall visual inspection system 100 performs steps A and B, it should be understood that steps A and B may also be performed jointly by two different processors or individually in one processor. For example, the first processor performs step A, the second processor performs step B, or the first processor and the second processor perform steps A and B together.

[0155] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned exterior wall visual detection method based on deep learning is implemented.

[0156] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. A method for visual inspection of exterior walls based on deep learning, characterized in that: The method comprises: Acquire an initial visual image set of the exterior wall, wherein the initial visual image set includes a plurality of exterior wall image units with pixel coordinate marks collected at different shooting angles and under different lighting conditions; Performing image feature preprocessing on the initial visual image set to generate a standardized feature image set including exterior wall surface texture features and structural contour features; Calling a pre-trained exterior wall defect detection model to perform defect feature recognition processing on the standardized feature image set, and generating a preliminary detection result of the defect area in each exterior wall image unit, the preliminary detection result including a defect type candidate set and contour edge coordinate information of the corresponding candidate area; Determine a final detection result of the defect area in each exterior wall image unit based on the confidence distribution of the defect type candidate set and the spatial position correlation of the candidate areas in the preliminary detection results, wherein the final detection result includes a unique defect type label and the precise position coordinates of the defect area in the image coordinate system; generating an exterior wall defect detection report including defect type statistics and a location distribution map based on the final detection result, and sending the exterior wall defect detection report to a target terminal device to trigger subsequent maintenance operations; The calling of the pre-trained exterior wall defect detection model to perform defect feature recognition processing on the standardized feature image set to generate preliminary detection results of defect areas in each exterior wall image unit includes: Inputting the standardized feature image set into the feature extraction network of the exterior wall defect detection model to generate abstract feature maps at different levels, the feature extraction network comprising multiple convolutional layers and pooling layers to extract local features from the standardized feature image layer by layer through a sliding window operation of the convolution kernel, and performing dimensionality reduction processing on the local features through the pooling layer; The abstract feature maps of different levels are input into the feature fusion module, and feature fusion is performed using a combination of top-down and bottom-up paths to generate a fused feature map containing multi-scale context information. The top-down path combines high-level semantic features with low-level detail features, while the bottom-up path gradually abstracts low-level features into high-level features. Inputting the fused feature map into a region proposal network to generate multiple candidate regions that may contain defects, each candidate region including position coordinates and size parameters in the image coordinate system; Perform feature cropping and scaling on each candidate region, adjust the feature map of the candidate region to a fixed size of a preset size, and generate a feature map of the candidate region; Input the candidate region feature map into the classification and regression subnetwork, predict the probability distribution of the candidate region belonging to different defect types through the classification branch, and predict the offset of the candidate region position coordinates and the adjustment parameters of the size through the regression branch; According to the probability distribution of the classification branch and the adjustment parameters of the regression branch, the candidate areas are corrected in position and non-maximum suppression is performed. The candidate areas with a confidence greater than a first set threshold and an overlap less than a second set threshold are retained to generate preliminary detection results of defective areas in each exterior wall image unit.

2. The exterior wall visual inspection method based on deep learning according to claim 1 is characterized in that: The performing image feature preprocessing on the initial visual image set to generate a standardized feature image set including exterior wall surface texture features and structural contour features includes: Performing color space conversion processing on each exterior wall image unit in the initial visual image set, converting the original RGB color space into the HSV color space to separate the brightness channel and the color channel; Performing histogram equalization on the brightness channel to enhance the contrast between bright and dark areas in the image by redistributing pixel brightness values, thereby generating a brightness-enhanced brightness feature channel; Performing noise filtering on the color channel, performing a convolution operation on the color channel using a Gaussian filter kernel of a preset size, and generating a color feature channel after noise suppression; Perform channel fusion processing on the brightness feature channel after brightness enhancement and the color feature channel after noise suppression, and generate a composite feature image containing brightness information and color information by superimposing them according to channel dimensions; Performing edge detection processing on the composite feature image to generate an edge feature map containing the structural contour information of the exterior wall surface; Performing texture analysis on the composite feature image to generate a texture feature matrix containing texture information of the exterior wall surface; The edge feature map and the texture feature matrix are subjected to standardized transformation and feature splicing operations to generate a standardized feature image set including the exterior wall surface texture features and the structural contour features.

3. The exterior wall visual inspection method based on deep learning according to claim 2 is characterized in that: The edge detection processing is performed on the composite feature image to generate an edge feature map containing the exterior wall surface structure contour information, including: Defining gradient operator templates in the horizontal direction and the vertical direction, and performing convolution operations on the composite feature image to obtain a horizontal gradient image and a vertical gradient image; Calculate the gradient magnitude and gradient direction of each pixel based on the horizontal gradient image and the vertical gradient image, where the gradient magnitude is the square root of the sum of the squares of the horizontal gradient value and the vertical gradient value, and the gradient direction is the inverse tangent of the horizontal gradient value and the vertical gradient value; Perform non-maximum suppression on the gradient amplitude. By comparing the gradient values ​​of adjacent pixels in the gradient direction of each pixel, only the local maximum value is retained to refine the edge lines, and the gradient image after non-maximum suppression is obtained. Binarization is performed on the gradient image after non-maximum suppression, and pixels with gradient amplitudes greater than the gradient amplitude threshold are marked as edge points, generating a binary edge image containing the edge of the exterior wall surface structure contour; Connected region analysis is performed on the binary edge image. A contour coordinate sequence of each connected edge region is marked by detecting continuous edge points. Each contour coordinate sequence corresponds to an independent structural edge unit on the exterior wall surface, and an edge feature map containing the structural contour information of the exterior wall surface is generated.

4. The exterior wall visual inspection method based on deep learning according to claim 2 is characterized in that: The performing texture analysis on the composite feature image to generate a texture feature matrix containing the texture information of the exterior wall surface includes: The size and moving step of the sliding window are set so that the sliding window can cover all pixel areas of the composite feature image and there is an overlapping area between adjacent sliding windows. The gray level co-occurrence matrix is ​​calculated for the pixel area within each sliding window. For each gray-level co-occurrence matrix, the texture feature parameters are calculated respectively, and the texture feature parameters include contrast, correlation, energy and entropy. Among them, contrast measures the roughness of texture by accumulating the product of squared gray-level difference and corresponding probability, correlation measures the spatial correlation of texture by accumulating the joint probability of gray-level products, energy measures the uniformity of texture by accumulating the square of gray-level joint probability, and entropy measures the complexity of texture by accumulating the logarithm of gray-level joint probability; The texture feature parameters extracted by each sliding window at multiple directions and distances are averaged to generate a texture feature vector for each pixel; The texture feature vectors of all pixel points are arranged in the order of the pixel coordinates of the original image to generate a texture feature matrix containing the texture information of the exterior wall surface.

5. The exterior wall visual inspection method based on deep learning according to claim 1 is characterized in that: Determining the final detection result of the defect area in each exterior wall image unit according to the confidence distribution of the defect type candidate set and the spatial position correlation of the candidate area in the preliminary detection result includes: For each candidate defect area in each exterior wall image unit, extract the confidence value of each defect type in its defect type candidate set to generate a confidence vector, which reflects the degree of confidence of the exterior wall defect detection model that the candidate area belongs to different defect types; Calculating the difference between the maximum value and the second largest value in the confidence vector as a confidence interval parameter, and when the confidence interval parameter is greater than a preset interval threshold, directly selecting the defect type as the preliminary defect type of the candidate area; When the confidence interval parameter is less than or equal to the preset interval threshold, a secondary judgment is performed in combination with the spatial position information of the candidate area, the contour edge coordinate information of the candidate area is extracted, and the area size parameter and shape compactness parameter of the candidate area are calculated. The shape compactness parameter is calculated by the ratio of the square of the perimeter of the candidate area to the area, and is used to measure the shape regularity of the defect area; Establish an association rule base between defect types and area size parameters and shape compactness parameters. This association rule base is established based on statistical analysis of historical inspection data and contains parameter ranges corresponding to different defect types. The rule base is queried based on the parameters of the current candidate area to determine the preliminary defect type. Perform spatial position correlation analysis on the preliminary defect types of all candidate defect areas, and calculate the ratio of the overlapping area of ​​any two candidate areas to the area of ​​the smaller candidate area as the overlap parameter to determine whether the candidate areas belong to the same defect; When the overlap parameter is greater than the preset overlap threshold and the defect type is consistent, the two candidate areas are merged into one defect area, and its position coordinates are the union of the two candidate areas; When the overlap parameter is greater than the preset overlap threshold but the defect type is inconsistent, the candidate areas with confidence greater than the confidence threshold are retained, thereby generating the final detection result of the defect area in each exterior wall image unit. Each defect area contains a unique defect type label and position coordinates in the image coordinate system.

6. The exterior wall visual inspection method based on deep learning according to claim 5 is characterized in that: The calculation of the area size parameter and shape compactness parameter of the candidate region includes: For the contour edge coordinate sequence of the candidate area, the geometric integration method is used to calculate the area of ​​the candidate area. This geometric integration method accurately obtains the area of ​​the area by integrating the contour coordinates, reflecting the actual size of the defect. The contour tracking algorithm is used to calculate the distances between adjacent coordinate points along the contour edge and sum them up to obtain the perimeter of the candidate area, which is used to measure the boundary length of the defect area. Based on the calculated area and perimeter, a shape compactness parameter is calculated. The shape compactness parameter is obtained by dividing the square of the perimeter by the area. A larger shape compactness indicates a narrower and longer area, while a smaller shape compactness indicates a closer circular area. The area size parameter and the shape compactness parameter are normalized to facilitate matching with the parameter ranges in the association rule library, wherein the association rule library is established by analyzing historical inspection data and includes area size parameter ranges and shape compactness parameter ranges corresponding to different defect types.

7. The exterior wall visual inspection method based on deep learning according to claim 1 is characterized in that: The generating of an exterior wall defect detection report including defect type statistical information and a location distribution map based on the final detection result, and sending the exterior wall defect detection report to a target terminal device to trigger subsequent maintenance operations, includes: Summarize the final detection results of all exterior wall image units, count the number of occurrences of each defect type and the total area of ​​the corresponding defect area, and generate a defect type statistical information table, which is used to show the occurrence frequency and impact range of different defect types; Establish a mapping relationship between the image coordinate system and the actual physical coordinate system of the exterior wall. By obtaining the intrinsic and extrinsic parameter matrices of the shooting device, convert the position coordinates of the defect area in each exterior wall image unit into the coordinates of the actual physical coordinate system of the exterior wall. Marking the location information of all defect areas on a two-dimensional plane map according to the coordinates in the physical coordinate system of the actual exterior wall to generate a location distribution map including defect location marks, each defect location mark including a defect type icon and coordinate annotations for displaying the distribution of defects on the actual exterior wall; Performing cluster analysis on the defect locations in the location distribution map to generate defect cluster area identifiers including area boundary coordinates and defect density parameters; Integrate the defect type statistics table, location distribution map, and defect cluster area identification into the preset report template to generate an exterior wall defect inspection report that includes text descriptions and graphical displays. The text description part is used to explain the inspection results and analysis conclusions, and the graphical display part is used to present the distribution characteristics of the defects. The exterior wall defect detection report is sent to a target terminal device to trigger subsequent maintenance operations.

8. The exterior wall visual inspection method based on deep learning according to claim 7 is characterized in that: The establishing of the mapping relationship between the image coordinate system and the actual exterior wall physical coordinate system includes: Obtain the intrinsic and extrinsic parameter matrices of the camera. The intrinsic parameter matrix includes the focal length, principal point coordinates, and distortion parameters, and is used to describe the internal optical characteristics and imaging geometry of the device. The extrinsic parameter matrix includes the rotation matrix and translation vector, and is used to describe the position and posture of the device in the world coordinate system. Define the actual physical coordinate system of the exterior wall. Take the lower left corner of the wall as the origin and establish a two-dimensional plane coordinate system with the horizontal rightward direction as the X axis and the vertical upward direction as the Y axis. Performing distortion correction processing on each defect position coordinate in the image coordinate system using the distortion parameters in the intrinsic parameter matrix to obtain corrected coordinates; Converting the corrected coordinates into coordinates in the actual exterior wall physical coordinate system through an external parameter matrix; The above conversion process is performed on the position coordinates of all defect areas in each exterior wall image unit to obtain a coordinate set in the physical coordinate system of the actual exterior wall. Each coordinate point in the coordinate set corresponds to an actual position on the exterior wall surface and is used to mark the specific position of the defect on the actual exterior wall.

9. A deep learning-based exterior wall visual inspection system, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the deep learning-based exterior wall visual inspection method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Concrete crack detection method and device

    CN118781093A

  • Concrete building outer wall detection method and system, computer equipment and storage medium

    CN119360219A

  • External wall quality defect identification method and system

    CN119579544A