Outer wall multi-type disease detection method and system based on unmanned aerial vehicle inspection
By combining drones with visible light and infrared thermal imaging image fusion and deep learning algorithms, the high cost and low efficiency of detecting multiple types of defects on building exterior walls have been solved, achieving efficient and accurate detection of multiple types of defects.
Patent Information
- Application Number
- CN202510742745.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-10-03
Smart Images

Figure CN120747779A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-type disease detection of building exterior walls, and in particular to a method and system for detecting multi-type disease of exterior walls based on drone inspection. Background Art
[0002] Typical manifestations of building facade defects include cracks, peeling, leakage and hollowing of external and internal materials. Detection and assessment of these deteriorations are generally carried out using inefficient and costly manual inspection methods, relying on professional inspectors to conduct inspections according to regular and standardized procedures.
[0003] Combining drones with advanced image capture instruments (such as high-resolution RGB cameras and infrared cameras) to obtain visual data about building facades can significantly reduce inspection costs and improve inspection accuracy. Using drones, infrared radiation can be used to obtain red thermal images of building facades. Differences in the facade's infrared radiation energy can be used to detect defects, making this inspection method widely used in non-destructive testing.
[0004] However, these methods all operate from a single-modality perspective and are only applicable to a limited variety of defects. To develop more broadly applicable detection methods, dual-modality detection methods using infrared and visible image fusion (IVIF) have become a research focus. This method leverages the complementary information of visible and infrared imaging to generate a more informative and advantageous fused image, enabling intuitive and accurate detection of building facade defects. This approach is of great significance for advancing the field of automated building facade defect detection. Summary of the Invention
[0005] The purpose of the present invention is to address the deficiencies in the existing technology and provide a method and system for detecting multiple types of exterior wall defects based on drone inspections.
[0006] The object of the present invention is achieved through the following technical solutions: In a first aspect, an embodiment of the present invention provides a method for detecting multiple types of defects on exterior walls based on drone inspection, comprising the following steps:
[0007] (1) Plan the drone's shooting route based on the surrounding environment of the inspected building and the size information of the building facade, and use the drone to collect visible light and infrared thermal imaging images of the building facade;
[0008] (2) Preprocess the visible light and infrared thermal imaging images to construct a dataset, and randomly divide it into a training set, a validation set, and a test set in proportion;
[0009] (3) An image fusion network is constructed based on a generative adversarial network. The image fusion network includes a generator, a visible light image discriminator, and an infrared thermal imaging discriminator. The image fusion network is subjected to adversarial training using a training set. During the training process, the parameters of the image fusion network are adjusted with the goal of minimizing the loss function of the image fusion network to obtain a trained generator.
[0010] (4) constructing a disease detection model based on the trained generator, the disease detection model including the generator trained in step (3) and the open source Mask R-CNN network; using the training set to train the disease detection model, during the training process, minimizing the loss function of the disease detection model is used as the optimization goal, only adjusting the parameters of the Mask R-CNN network, freezing the parameters of the generator, to obtain the trained disease detection model;
[0011] (5) The trained defect detection model is used to obtain the local defect map of the visible light and infrared thermal imaging images to be detected. The SURF algorithm is used to extract feature points, match, register, and splice the local defect map. Feature processing is performed on the overlapping areas of the image to achieve splicing of the local defect map and obtain the global defect detection map of the facade of the building to be tested.
[0012] Furthermore, in step (1), when the drone collects visible light and infrared thermal imaging images of the building facade, the drone should shoot according to the recommended shooting angle, shooting distance and shooting time to obtain visible light and infrared thermal imaging images of the building facade to be inspected;
[0013] Among them, the shooting angle, shooting distance and shooting time of the drone shall comply with the following regulations:
[0014] ① Shooting angle: The elevation angle is less than or equal to 45°, and the horizontal tilt angle is less than or equal to 30°; the angle between the camera lens axis and the normal vector of the exterior wall is less than or equal to 55°; if the distance between the lens and the exterior wall is greater than 4.00m, the tilt angle is less than or equal to 45°;
[0015] ② Camera shooting distance: should be kept between 1.50m and 4.75m;
[0016] ③ Shooting time: The best time to collect infrared thermal imaging images is from 9 am to 4 pm.
[0017] Furthermore, the preprocessing of the visible light and infrared thermal imaging images to construct a data set specifically includes:
[0018] (2.1) Performing data enhancement processing on visible light and infrared thermal imaging images using a variety of data enhancement methods; wherein the data enhancement methods include horizontal flipping, vertical flipping, rotation, shifting, cropping, and deformation and scaling;
[0019] (2.2) Labeling the visible light and infrared thermal imaging images after data enhancement processing with corresponding defect types to obtain pre-processed visible light and infrared thermal imaging images; wherein the defect types include cracks, spalling, water seepage, and hollowing;
[0020] (2.3) Construct a dataset based on the preprocessed visible light and infrared thermal imaging images.
[0021] Furthermore, in the image fusion network, the generator is built based on ResNet. The generator first uses the ResNet-50 network to extract the depth features of the original visible light and infrared thermal imaging images; then normalizes the depth features of the visible light and infrared thermal imaging images to generate initialization weights; then uses a softmax operation combined with the initialization weights to obtain the final weights; finally, based on the final weights, a weighted average strategy is used to fuse the image features of the original input visible light and infrared thermal imaging images to generate a fused image;
[0022] The visible light image discriminator and the infrared thermal imaging discriminator are essentially image classification neural networks. The visible light image discriminator is responsible for distinguishing texture details in visible light images, and the infrared thermal imaging discriminator is responsible for distinguishing thermal targets in infrared thermal imaging images.
[0023] The visible light and infrared thermal imaging images are input into the image fusion network, first entering the generator to generate a fused image; then the fused image and the original visible light image are input into the visible light image discriminator to obtain the probability that the fused image comes from the real visible light image; the fused image and the original infrared thermal imaging image are input into the infrared thermal imaging discriminator to obtain the probability that the fused image comes from the real infrared thermal imaging image.
[0024] Furthermore, the calculation formula of the loss function of the image fusion network is:
[0025]
[0026] in, represents the loss function of the image fusion network, E(*) represents the expected value of the distribution function, P data (x) represents the true sample distribution, P noise (x) is a noise distribution defined in low dimension, D(*) represents the probability that the fused image comes from a real visible light image or infrared thermal imaging image, and G(*) represents the fused image generated using visible light and infrared thermal imaging images.
[0027] Furthermore, the Mask R-CNN network first obtains the feature map of the input fusion image; then sets candidate regions of interest for the feature map, performs classification and regression on the candidate regions of interest, and then performs region of interest alignment operations. Subsequently, the regions of interest are classified, bounding box regression and mask generation are performed to achieve image instance segmentation. Different types of exterior wall diseases can be labeled with different colors and represented with clear outlines, obtaining the category, bounding box and mask of each instance.
[0028] Furthermore, the calculation formula of the loss function of the disease detection model is:
[0029] L=L cls +L box +L mask
[0030] Among them, L represents the loss function of the disease detection model, L cls represents the classification loss, L box represents the regression box loss, L mask represents the generated mask loss.
[0031] Furthermore, the step (5) includes the following sub-steps:
[0032] (5.1) Obtaining local defect maps: Using the trained defect detection model, obtain local defect maps of the visible light and infrared thermal imaging images to be detected;
[0033] (5.2) Extract feature points from the local defect image: Use the SURF algorithm to detect the feature points of the local defect image using the integral image and Hessian matrix approximation, which serve as the matching point set to be registered;
[0034] (5.3) Image registration: Based on the matching point set to be registered obtained in step (5.2), the two local defect images are converted to the same coordinate system and then image registration is performed;
[0035] (5.4) Image copying: copy the local defect map to the registration map obtained in step (5.3) to obtain a spliced image;
[0036] (5.5) Image crack removal processing: In the overlapping areas of the image spliced in step (5.4), the pixel value weights of the overlapping areas are given according to the reliability of the local defect map, and the pixel values of the overlapping areas of multiple local defect maps are fused by weighted averaging. The pixel value weights of the overlapping areas are then added together to form a new image as the global defect detection map of the building facade to be tested.
[0037] Furthermore, the method of using the SURF algorithm to detect the feature points of the local defect image using the integral image and the Hessian matrix approximation specifically includes the following sub-steps:
[0038] (5.2.1) Obtain the approximate Hessian matrix corresponding to the local defect map: First, search for the local defect map in all scale spaces, perform Gaussian filtering on the local defect map, and construct the Hessian matrix to obtain the filtered Hessian matrix, which is expressed as:
[0039]
[0040] Where H(x,y,σ) represents the filtered Hessian matrix, (x,y) represents the position of the pixel in the local defect map; L(x,y,σ)=G(σ)*I(x,y) represents the Gaussian scale space of the local defect map; G(σ) represents Gaussian convolution, σ represents the standard deviation of Gaussian convolution, and I(x,y) represents the pixel value of the pixel at the position (x,y) in the local defect map; L xx () means taking the derivative twice of the independent variable x, L xy () indicates that the derivative of the independent variable x is taken first and then the derivative of the independent variable y. yy () means taking the derivative of the independent variable y twice; then the discriminant of the Hessian matrix is obtained according to the filtered Hessian matrix, and its expression is:
[0041] Det(H)=L xx *L yy -L xy *L xy
[0042] Where Det(H) represents the discriminant value of the Hessian matrix. Then, a box filter is used to replace the Gaussian filter, and the approximate expression of the Hessian matrix discriminant is obtained by multiplying the weighting coefficient λ:
[0043] Det'(H)=L xx *L yy -(λ*L xy ) 2
[0044] Where Det'(H) represents the Hessian matrix approximation;
[0045] (5.2.2) Constructing scale space: The scale space of the SURF algorithm consists of O groups of S layers, each with different sizes;
[0046] (5.2.3) Feature point filtering and precise positioning: The Hessian matrix approximation of each pixel point processed by the Hessian matrix is compared with the Hessian matrix approximations of all its neighboring points in the image domain and scale domain. If the Hessian matrix approximation of a pixel point in the local defect image is greater than or less than the Hessian matrix approximations of all its neighboring points, the pixel point is considered an extreme point and is used as a feature point. After the feature points are initially located through the above steps, weak energy feature points and incorrectly located feature points are filtered out to select the final feature points.
[0047] (5.2.4) Calculate the main direction of the feature point: The SURF algorithm determines the main direction of the feature point by counting the Harr wavelet features in the circular neighborhood of the feature point. Specifically, within the circular neighborhood of the feature point, the sum of the horizontal and vertical Harr wavelet features of all points in a 60° sector is counted as the Harr wavelet eigenvalue within the sector area; the sector is rotated at intervals of 0.2 radians, and the Harr wavelet eigenvalues within the sector area are counted again; the above process is repeated to obtain the Harr wavelet eigenvalues of all sector areas in the circular neighborhood of the feature point; finally, the direction of the sector corresponding to the largest Harr wavelet eigenvalue is taken as the main direction of the feature point;
[0048] (5.2.5) Generate feature description: Use the SURF algorithm to extract a 4×4 rectangular area around the feature point. In each sub-area, calculate the Haar wavelet features of 25 pixels in the horizontal and vertical directions. The Haar wavelet features are the sum of the horizontal values, the sum of the vertical values, the sum of the horizontal absolute values, and the sum of the vertical absolute values.
[0049] (5.2.6) Obtain the set of matching points to be registered: Construct the set of matching points to be registered based on the filtered feature points, main directions of feature points, and feature descriptions.
[0050] A second aspect of an embodiment of the present invention provides a system for implementing the above-mentioned method for detecting multiple types of exterior wall defects based on drone inspection, comprising:
[0051] The image acquisition module is used to plan the drone's shooting route based on the surrounding environment of the inspected building and the size information of the building facade, and to collect visible light and infrared thermal imaging images of the building facade through the drone;
[0052] The data preprocessing and dataset construction module is used to preprocess visible light and infrared thermal imaging images to construct a dataset and randomly divide it into training set, validation set and test set in proportion;
[0053] The generator acquisition module is used to construct an image fusion network based on a generative adversarial network. The image fusion network includes a generator, a visible light image discriminator, and an infrared thermal imaging discriminator. The image fusion network is adversarially trained using a training set. During the training process, the parameters of the image fusion network are adjusted with the goal of minimizing the loss function of the image fusion network to obtain a trained generator.
[0054] A disease detection model acquisition module is used to build a disease detection model based on the trained generator. The disease detection model includes the trained generator and the open-source Mask R-CNN network. The disease detection model is trained using the training set. During the training process, the loss function of the disease detection model is minimized as the optimization goal. Only the parameters of the Mask R-CNN network are adjusted, and the parameters of the generator are frozen to obtain the trained disease detection model.
[0055] The global defect detection map acquisition module is used to use the trained disease detection model to obtain the local defect map of the visible light and infrared thermal imaging images to be detected. The SURF algorithm is used to extract feature points, match, register, and splice the local defect map, perform feature processing on the overlapping areas of the image, and realize the splicing of the local defect map to obtain the global defect detection map of the facade of the building to be tested.
[0056] The beneficial effects of the present invention are: with the help of drones, the present invention can obtain more intuitive and accurate fusion images of building exterior wall diseases through deep learning algorithms. It is highly practical and overcomes the limitations of single-modality disease detection. It can identify multiple exterior wall diseases and improve the efficiency of exterior wall disease detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is a flow chart of the method for detecting multiple types of defects on exterior walls based on drone inspections of the present invention;
[0058] Figure 2 Schematic diagram of the network structure of the image fusion network of the present invention;
[0059] Figure 3 is a flow chart of the generator of the present invention for generating a fused image;
[0060] Figure 4 It is a structural schematic diagram of the exterior wall multi-type disease detection system based on drone inspection of the present invention. DETAILED DESCRIPTION
[0061] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims. It should be understood that the foregoing general description and the detailed description that follows are exemplary and illustrative only and do not limit the present application.
[0062] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0063] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of..." or "when..." or "in response to determination." Moreover, the term "comprises," "comprising," or any other variant thereof is intended to cover non-exclusive inclusion, so that the process or method comprising a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process or method. In the absence of further restrictions, the elements defined by the statement "comprising a..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0064] The present invention will be described in detail below with reference to the accompanying drawings. Unless there is any conflict, the features of the following embodiments and implementations may be combined with each other.
[0065] See also Figure 1 The method for detecting multiple types of defects on exterior walls based on drone inspection of the present invention specifically includes the following steps:
[0066] (1) Plan the drone's shooting route based on the surrounding environment of the inspected building and the size information of the building facade to ensure that the drone can collect valid and qualified visible light and infrared thermal imaging images of the building facade, and collect visible light and infrared thermal imaging images of the building facade through the drone.
[0067] Furthermore, when drones collect visible light and infrared thermal imaging images of building facades, they should follow the recommended shooting angles, distances, and times to obtain visible light and infrared thermal images of the building facades to be inspected. This ensures that the collected visible light and infrared thermal images provide useful information for subsequent image data processing. The following regulations apply to drone shooting angles, distances, and times: ① Shooting angle: Because frontal radiation energy is greatest, when using drones equipped with infrared imaging equipment to photograph high-rise buildings, the standard stipulates that the elevation angle should be ≤45° and the horizontal tilt angle should be ≤30°. The angle between the camera's main axis and the normal vector of the exterior wall should be ≤55°. If the distance between the lens and the exterior wall is greater than 4.00m, the tilt angle should be appropriately reduced to ≤45°. ② Shooting distance: In most cases, the shooting distance should be maintained between 1.50m and 4.75m. A longer distance may be used, but it should not be too close (no less than 1.50m). ③ Shooting Time: The optimal time for infrared thermal imaging image acquisition is between 9:00 AM and 4:00 PM. Drone photography should be performed within this timeframe. The optimal shooting time varies for different facades, specifically 8:00 AM to 9:00 AM for the east facade, 11:00 AM to 1:00 PM for the south facade, and 3:00 PM to 4:00 PM for the west facade. For information on image acquisition times for different regions and building facades, refer to Appendix B, "Suitable Testing Time Periods for Infrared Thermography Testing of the Adhesion Quality of Building Exterior Wall Finishes in Selected Cities Nationwide," in JGJ T 277-2012, "Technical Specifications for Testing the Adhesion Quality of Building Exterior Wall Finishes Using Infrared Thermography."
[0068] (2) The visible light and infrared thermal imaging images are preprocessed to construct a dataset, which is then randomly divided into a training set, a validation set, and a test set in proportion.
[0069] It should be understood that the dataset is randomly divided into training set, validation set and test set in a certain ratio (such as 7:2:1, 8:1:1, etc.), and a well-performing building facade disease detection model is obtained through subsequent training, validation and testing.
[0070] Furthermore, the visible light and infrared thermal imaging images are preprocessed to construct a data set, specifically including:
[0071] (2.1) Data augmentation methods are used to enhance visible light and infrared thermal imaging images. These methods include, but are not limited to, horizontal flipping, vertical flipping, rotation (randomly selecting a specified angle for the image), shifting (moving the image in the X or Y direction, or both), cropping (randomly cropping the image), and scaling (stretching and scaling the image). This provides data diversity and enriches the dataset, effectively preventing overfitting of subsequent models and improving model accuracy.
[0072] (2.2) The visible light and infrared thermal imaging images after data enhancement are labeled with corresponding defect types to obtain pre-processed visible light and infrared thermal imaging images. Defect types include cracks, spalling, water seepage, and hollowing.
[0073] It should be noted that cracks are a common type of defect that can be identified through visible light images and infrared thermal imaging images; spalling refers to the phenomenon of material falling off or peeling off from the building surface, which can be more clearly identified in visible light images; water seepage refers to damage caused by prolonged or excessive exposure to a humid environment, which can be identified in both visible light images and infrared thermal imaging images; hollowing refers to the separation or peeling between the gypsum layer and the base layer or surface, which is generally identified by infrared thermal imaging images.
[0074] (2.3) Construct a dataset based on the preprocessed visible light and infrared thermal imaging images.
[0075] (3) An image fusion network is constructed based on Generative Adversarial Networks (GAN), which includes a generator, a visible light image discriminator, and an infrared thermal imaging discriminator, such as Figure 2 As shown in the figure, the image fusion network is adversarially trained using the training set. During the training process, the parameters of the image fusion network are adjusted with the optimization goal of minimizing the loss function of the image fusion network to obtain a trained generator.
[0076] In this embodiment, the image fusion network includes a generator, a visible light image discriminator, and an infrared thermal imaging discriminator. Figure 2 As shown in the figure, the generator is built based on ResNet. The generator first uses the ResNet-50 network to extract the depth features of the original visible light and infrared thermal imaging images. This step can obtain the image depth features under the two modalities. Then, the depth features of the visible light and infrared thermal imaging images are normalized to generate initialization weights. The softmax operation is then combined with the initialization weights to obtain the final weights. Finally, based on the final weights, the image features of the original input visible light and infrared thermal imaging images are fused using a weighted average strategy to generate a fused image, thereby achieving the purpose of image reconstruction. Figure 3 As shown in Figure 2, the visible light image discriminator and the infrared thermal imaging discriminator are essentially image classification neural networks. The visible light image discriminator is responsible for distinguishing texture details in visible light images, while the infrared thermal imaging discriminator is responsible for distinguishing thermal targets in infrared thermal imaging images.
[0077] Specifically, the visible light and infrared thermal imaging images are input into the image fusion network, first entering the generator to generate a fused image; then the fused image and the original visible light image are input into the visible light image discriminator to obtain the probability that the fused image comes from the real visible light image; the fused image and the original infrared thermal imaging image are input into the infrared thermal imaging discriminator to obtain the probability that the fused image comes from the real infrared thermal imaging image, as shown in Figure 2 shown.
[0078] It should be understood that the image fusion network includes three steps: image feature extraction, feature fusion, and image reconstruction, ultimately achieving the discriminant result of the fused image source. The image fusion network demonstrates the ability to prioritize the fusion of complementary information, focusing on the differences between infrared thermal imaging and visible light imagery. It focuses on conveying the structural properties of the target and the subtle texture differences of the environmental background, which are crucial for distinguishing the various features that characterize the target for target detection.
[0079] In this embodiment, adversarial training is performed on an image fusion network using a training set. During the training process, the parameters of the image fusion network are adjusted with the optimization objective of minimizing the loss function of the image fusion network to obtain a trained generator. Specifically, the training includes: inputting visible light and infrared thermal imaging images from the training set into the image fusion network to obtain the probability that the fused image is derived from a true visible light image and the probability that the fused image is derived from a true infrared thermal imaging image. The loss function of the image fusion network is calculated based on the probabilities that the fused image is derived from a true visible light image and the probabilities that the fused image is derived from a true infrared thermal imaging image. With the optimization objective of minimizing the loss function of the image fusion network, the parameters of the image fusion network, namely, the parameters of the generator, visible light image discriminator, and infrared thermal imaging discriminator, are adjusted until a preset number of training rounds is reached or the loss is less than a preset loss threshold, thereby obtaining a trained image fusion network and retaining the trained generator for subsequent steps.
[0080] It should be understood that the generative adversarial network means that during the training process, the goal of the generator is to generate realistic fused images as much as possible to deceive the discriminator, so that the discriminator cannot distinguish that the input is a fused image; while the goal of the discriminator is to try to distinguish the fused image generated by the generator from the real infrared thermal imaging image or visible light image. In this way, the generator and the discriminator constitute a dynamic game process.
[0081] Furthermore, the loss function of the image fusion network is calculated as:
[0082]
[0083] in, represents the loss function of the image fusion network, E(*) represents the expected value of the distribution function, P data (x) represents the true sample distribution, P noise (x) is a noise distribution defined in low dimension, D(*) represents the probability that the fused image comes from a real visible light image or infrared thermal imaging image, and G(*) represents the fused image generated using visible light and infrared thermal imaging images.
[0084] (4) Constructing a disease detection model based on the trained generator, the disease detection model includes the generator trained in step (3) and the open source Mask R-CNN network; using the training set to train the disease detection model, during the training process, minimizing the loss function of the disease detection model is used as the optimization goal, only adjusting the parameters of the Mask R-CNN network, freezing the parameters of the generator, to obtain the trained disease detection model.
[0085] In this embodiment, the Mask R-CNN network first obtains the feature map of the input fused image; then sets candidate regions of interest (RoIs) on the feature map, performs classification and regression on the candidate RoIs, and filters out some RoIs; then performs a RoIAlign operation, and then classifies the RoIs, regresses the bounding boxes, and generates masks to achieve image instance segmentation. It can label different types of exterior wall defects with different colors and represent them with clear outlines, intuitively displaying the type of defect, with high accuracy, fast detection speed, and stable performance in complex scenes, and obtain the category, bounding box, and mask for each instance. Therefore, the method described in the present invention can detect defects on various complex building exterior walls.
[0086] It should be understood that the Mask RCNN network is an existing compact and flexible general-purpose object instance segmentation framework that not only detects objects in an image but also provides high-quality segmentation results for each object. The Mask RCNN network is capable of multi-task joint training: ① Identify the category of each instance, such as wall, ancillary facilities, and damage; ② Accurately predict the bounding box of each instance; ③ Generate a pixel-level segmentation mask for each instance, which is a black and white mask.
[0087] In this embodiment, the training set is used to train the disease detection model. During the training process, the loss function of the disease detection model is minimized as the optimization goal. Only the parameters of the MaskR-CNN network are adjusted, and the parameters of the generator are frozen to obtain a trained disease detection model. Specifically, the visible light and infrared thermal imaging images in the training set are input into the disease detection model, and the corresponding fusion image is first entered into the generator trained in step (3) to generate the corresponding fusion image. The fusion image is then passed through the MaskR-CNN network to obtain the category, bounding box and mask of each instance. The loss function of the disease detection model is calculated based on the category, bounding box and mask of the instance and the corresponding defect type annotation in the training set. With the loss function of the disease detection model minimized as the optimization goal, only the parameters of the MaskR-CNN network are adjusted, and the parameters of the generator trained in step (3) are frozen until a preset training round is reached or the loss function threshold is less than a preset value, thereby obtaining a trained disease detection model.
[0088] Furthermore, the loss function of the disease detection model is calculated as follows:
[0089] L=L cls +L box +L mask
[0090] Among them, L represents the loss function of the disease detection model, L cls represents the classification loss, L box represents the regression box loss, L mask Represents the generated mask loss.
[0091] It should be noted that the Mask RCNN network can identify the four types of defects in the fused image and distinguish different types of exterior wall defects with different colors and display the defects with clear outlines, thereby separating the defects from the facade background, producing accurate segmentation results, and intuitively displaying the defects in the building facade.
[0092] (5) The trained defect detection model is used to obtain the local defect map of the visible light and infrared thermal imaging images to be detected. The SURF (Speeded Up Robust Features) algorithm is used to perform feature point extraction, matching, registration, and splicing on the local defect map. Feature processing is performed on the overlapping areas of the image to achieve splicing of the local defect map, thereby obtaining a global defect detection map of the facade of the building to be tested.
[0093] (5.1) Obtaining local defect maps: Use the trained defect detection model to obtain local defect maps of the visible light and infrared thermal imaging images to be detected.
[0094] (5.2) Feature point extraction from the local defect map: The SURF algorithm uses the integral image and Hessian matrix approximation to quickly detect feature points in the local defect map, which serve as the matching point set for registration. The integral image is a pre-computed image representation that can quickly calculate the sum of the image region; the Hessian matrix approximation is used to detect points in the image with large changes in curvature, which are usually feature points.
[0095] Furthermore, the SURF algorithm is used to quickly detect the feature points of the local defect map using the integral image and the Hessian matrix approximation, which specifically includes the following sub-steps:
[0096] (5.2.1) Obtain the approximate Hessian matrix corresponding to the local defect map: First, search for the local defect map in all scale spaces, perform Gaussian filtering on the local defect map, and construct the Hessian matrix to obtain the filtered Hessian matrix, which is expressed as:
[0097]
[0098] Where H(x,y,σ) represents the filtered Hessian matrix, (x,y) represents the position of the pixel in the local defect map; L(x,y,σ)=G(σ)*I(x,y) represents the Gaussian scale space of the local defect map, which is obtained by convolving the local defect map with different Gaussians; G(σ) represents Gaussian convolution, σ represents the standard deviation of Gaussian convolution, and I(x,y) represents the pixel value of the pixel at the position (x,y) in the local defect map; L xx () means taking the derivative twice of the independent variable x, L xy () indicates that the derivative of the independent variable x is taken first and then the derivative of the independent variable y. yy () means taking the derivative twice of the independent variable y. Then, the discriminant of the Hessian matrix is obtained based on the filtered Hessian matrix, and its expression is:
[0099] Det(H)=L xx *L yy -L xy *L xy
[0100] Wherein, Det(H) represents the discriminant value of the Hessian matrix. Since L in the discriminant of the Hessian matrix is the Gaussian convolution of the original local defect image, and since the Gaussian kernel obeys the normal distribution, the coefficient becomes smaller and smaller from the center point outward. In order to improve the operation speed, a box filter is used to replace the Gaussian filter, so it is multiplied by the weighting coefficient λ. For example, in this embodiment, λ = 0.9, the purpose is to balance the error caused by the use of the box filter approximation. Then the discriminant of the Hessian matrix is approximately expressed as:
[0101] Det'(H)=L xx *L yy -(λ*L xy ) 2
[0102] Where Det'(H) represents the Hessian matrix approximation, which can be used to identify potential points of interest that are invariant to scale and selection.
[0103] It should be understood that the Hessian matrix is a square matrix composed of the second-order partial derivatives of a multivariate function, describing the function's local curvature. The purpose of constructing the Hessian matrix is to generate stable edge points (i.e., mutation points) in the image, paving the way for feature point extraction. This method allows the calculation of the H-determinant Hessian matrix approximation for each pixel in the local defect image, and uses this Hessian matrix approximation to identify the feature points of the local defect image; the H-determinant refers to the Hessian matrix with H rows and H columns.
[0104] (5.2.2) Constructing scale space: The scale space of the SURF algorithm consists of O groups and S layers, each with different sizes. The sizes of images in different groups are consistent, and the template size of the box filter used in different groups gradually increases. The same size of box filter is used for different layers of images in the same group, and the scale space factor of the box filter gradually increases.
[0105] (5.2.3) Feature point filtering and precise positioning: The Hessian matrix approximation of each pixel point processed by the Hessian matrix (i.e., the discriminant value of the Hessian matrix of each pixel point) is compared with the Hessian matrix approximations of all adjacent points in its image domain (image of the same size) and scale domain (adjacent scale space). If the Hessian matrix approximation of a pixel point in the local defect image is greater than or less than the Hessian matrix approximations of all adjacent points, the pixel point is considered an extreme point and is used as a feature point. After the feature points are initially located through the above steps, feature points with relatively weak energy and incorrectly located feature points are filtered out to select the final stable feature points. Among them, if the brightness of a feature point is darker than that of other feature points, the feature point is considered to have relatively weak energy.
[0106] (5.2.4) Calculate the main direction of the feature point: In the SURF algorithm, the main direction of the feature point is determined by counting the Harr wavelet features in the circular neighborhood of the feature point. Specifically, in the circular neighborhood of the feature point, the sum of the horizontal and vertical Harr wavelet features of all points in a 60° sector is counted as the Harr wavelet eigenvalue in the sector area; the sector is rotated at intervals of 0.2 radians, and the Harr wavelet eigenvalue in the sector area is counted again; the above process is repeated to obtain the Harr wavelet eigenvalues in all sector areas in the circular neighborhood of the feature point; finally, the direction of the sector corresponding to the largest Harr wavelet eigenvalue is taken as the main direction of the feature point.
[0107] (5.2.5) Generate feature description: The SURF algorithm extracts 4×4 rectangular area blocks around the feature points. In each sub-area, the Haar wavelet features of 25 pixels in the horizontal and vertical directions are counted. The Haar wavelet features are the sum of the horizontal values, the sum of the vertical values, the sum of the horizontal absolute values, and the sum of the vertical absolute values.
[0108] (5.2.6) Obtain the set of matching points to be registered: Construct the set of matching points to be registered based on the filtered feature points, main directions of feature points, and feature descriptions.
[0109] (5.3) Image registration: Based on the matching point set to be registered obtained in step (5.2), the two local defect images are converted to the same coordinate system. In this embodiment, the findHomograpy function is used to calculate the coordinate transformation matrix, and then image registration is performed.
[0110] (5.4) Image copying: copy the local defect map to the registration map obtained in step (5.3) to obtain a spliced image.
[0111] (5.5) Image crack removal: For the overlapping areas of the images stitched together in step (5.4), weights are assigned to the pixel values of the overlapping areas based on the reliability of the local defect maps. The pixel values of the overlapping areas of multiple local defect maps are then fused using a weighted average. The weighted pixel values of the overlapping areas are then added together to create a new image, which serves as the global defect detection map for the building facade under test. This effectively avoids duplicate detection results. By stitching together the local defect maps described above, a global defect detection map for the building facade under test can be obtained, thereby achieving intuitive and accurate detection results for exterior wall defects.
[0112] It should be noted that after step (5.4), the stitching of the two images is not natural and the transition at the junction of the stitched image is poor due to the different lighting colors of the two images. Therefore, specific processing is required to solve the unnatural transition phenomenon.
[0113] It should be understood that when multiple images are stitched together, such as when two images are stitched together, there will be some overlap in the scenes of the two images. In this case, the overlapping parts must be aligned when stitching. The SURF algorithm can be understood as finding the overlapping parts of the scenes in the two images, and taking the feature points in these overlapping scenes as representatives. These feature points are aligned when stitching, so it is considered that the overlapping parts of the two images are successfully matched when stitching.
[0114] After the above steps, the method of the present invention can obtain a global defect detection map of the building facade to be inspected, achieving intuitive and accurate detection results for building exterior wall defects. The above-mentioned order of implementation of the present invention is for illustrative purposes only and does not represent the superiority or inferiority of the embodiments. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0115] It is worth mentioning that the embodiment of the present invention also provides a system for detecting multiple types of defects on exterior walls based on drone inspection, which is used to implement the method for detecting multiple types of defects on exterior walls based on drone inspection in the above embodiment. Figure 4 As shown, the system includes an image acquisition module, a data preprocessing and dataset construction module, a generator acquisition module, a disease detection model acquisition module and a global defect detection map acquisition module.
[0116] In this embodiment, the image acquisition module is used to plan the shooting route of the drone according to the surrounding environment of the detected building and the size information of the building facade, and to collect visible light and infrared thermal imaging images of the building facade through the drone.
[0117] In this embodiment, the data preprocessing and dataset construction module is used to preprocess visible light and infrared thermal imaging images to construct a dataset, and randomly divide it into a training set, a validation set, and a test set in proportion.
[0118] In this embodiment, the generator acquisition module is used to construct an image fusion network based on a generative adversarial network, which includes a generator, a visible light image discriminator, and an infrared thermal imaging discriminator; the image fusion network is adversarially trained using a training set. During the training process, the parameters of the image fusion network are adjusted with the optimization goal of minimizing the loss function of the image fusion network to obtain a trained generator.
[0119] In this embodiment, the disease detection model acquisition module is used to build a disease detection model based on the trained generator, which includes the trained generator and the open source Mask R-CNN network; the disease detection model is trained using the training set. During the training process, the loss function of the disease detection model is minimized as the optimization goal. Only the parameters of the Mask R-CNN network are adjusted, and the parameters of the generator are frozen to obtain the trained disease detection model.
[0120] In this embodiment, the global defect detection image acquisition module is used to use the trained disease detection model to obtain the local defect map of the visible light and infrared thermal imaging images to be detected, use the SURF algorithm to perform feature point extraction, matching, registration, splicing and other operations on the local defect map, perform feature processing on the overlapping areas of the image, realize the splicing of the local defect map, and obtain the global defect detection map of the facade of the building to be tested.
[0121] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for detecting multiple types of exterior wall defects based on drone inspection, characterized in that: The following steps are involved: (1) Plan the drone's shooting route based on the surrounding environment of the inspected building and the size information of the building facade, and use the drone to collect visible light and infrared thermal imaging images of the building facade; (2) Preprocess the visible light and infrared thermal imaging images to construct a dataset, and randomly divide it into a training set, a validation set, and a test set in proportion; (3) An image fusion network is constructed based on a generative adversarial network. The image fusion network includes a generator, a visible light image discriminator, and an infrared thermal imaging discriminator. The image fusion network is subjected to adversarial training using a training set. During the training process, the parameters of the image fusion network are adjusted with the goal of minimizing the loss function of the image fusion network to obtain a trained generator. (4) constructing a disease detection model based on the trained generator, the disease detection model including the generator trained in step (3) and the open source Mask R-CNN network; using the training set to train the disease detection model, during the training process, minimizing the loss function of the disease detection model is used as the optimization goal, only adjusting the parameters of the Mask R-CNN network, freezing the parameters of the generator, to obtain the trained disease detection model; (5) The trained defect detection model is used to obtain the local defect map of the visible light and infrared thermal imaging images to be detected. The SURF algorithm is used to extract feature points, match, register, and splice the local defect map. Feature processing is performed on the overlapping areas of the image to achieve splicing of the local defect map and obtain the global defect detection map of the facade of the building to be tested.
2. The method for detecting multiple types of exterior wall defects based on drone inspection according to claim 1 is characterized in that: In step (1), when the drone collects visible light and infrared thermal imaging images of the building facade, the drone should shoot according to the recommended shooting angle, shooting distance and shooting time to obtain visible light and infrared thermal imaging images of the building facade to be inspected; Among them, the shooting angle, shooting distance and shooting time of the drone shall comply with the following regulations: ① Shooting angle: The elevation angle is less than or equal to 45°, and the horizontal tilt angle is less than or equal to 30°; the angle between the camera lens axis and the normal vector of the exterior wall is less than or equal to 55°; if the distance between the lens and the exterior wall is greater than 4.00m, the tilt angle is less than or equal to 45°; ② Camera shooting distance: should be kept between 1.50m and 4.75m; ③ Shooting time: The best time to collect infrared thermal imaging images is from 9 am to 4 pm.
3. The method for detecting multiple types of exterior wall defects based on drone inspection according to claim 1 is characterized in that: The preprocessing of the visible light and infrared thermal imaging images to construct a data set specifically includes: (2.1) Performing data enhancement processing on visible light and infrared thermal imaging images using a variety of data enhancement methods; wherein the data enhancement methods include horizontal flipping, vertical flipping, rotation, shifting, cropping, and deformation and scaling; (2.2) Labeling the visible light and infrared thermal imaging images after data enhancement processing with corresponding defect types to obtain pre-processed visible light and infrared thermal imaging images; wherein the defect types include cracks, spalling, water seepage, and hollowing; (2.3) Construct a dataset based on the preprocessed visible light and infrared thermal imaging images.
4. The method for detecting multiple types of exterior wall defects based on drone inspection according to claim 1 is characterized in that: In the image fusion network, the generator is built based on ResNet. The generator first uses the ResNet-50 network to extract the depth features of the original visible light and infrared thermal imaging images; The depth features of the visible light and infrared thermal imaging images are then normalized to generate initial weights. A softmax operation is then used in combination with the initial weights to obtain the final weights. Finally, based on the final weights, the image features of the original input visible light and infrared thermal imaging images are fused using a weighted average strategy to generate a fused image. The visible light image discriminator and the infrared thermal imaging discriminator are essentially image classification neural networks. The visible light image discriminator is responsible for distinguishing texture details in visible light images, and the infrared thermal imaging discriminator is responsible for distinguishing thermal targets in infrared thermal imaging images. The visible light and infrared thermal imaging images are input into the image fusion network, first entering the generator to generate a fused image; then the fused image and the original visible light image are input into the visible light image discriminator to obtain the probability that the fused image comes from the real visible light image; the fused image and the original infrared thermal imaging image are input into the infrared thermal imaging discriminator to obtain the probability that the fused image comes from the real infrared thermal imaging image.
5. The method for detecting multiple types of exterior wall defects based on drone inspection according to claim 1 is characterized in that: The calculation formula of the loss function of the image fusion network is: in, represents the loss function of the image fusion network, E(*) represents the expected value of the distribution function, P data (x) represents the true sample distribution, P noise (x) is a noise distribution defined in low dimension, D(*) represents the probability that the fused image comes from a real visible light image or infrared thermal imaging image, and G(*) represents the fused image generated using visible light and infrared thermal imaging images.
6. The method for detecting multiple types of exterior wall defects based on drone inspection according to claim 1 is characterized in that: The Mask R-CNN network first obtains the feature map of the input fusion image; then sets candidate regions of interest for the feature map, performs classification regression on the candidate regions of interest, and then performs region of interest alignment. Subsequently, the regions of interest are classified, bounding box regression, and mask generation are performed to achieve image instance segmentation. Different types of exterior wall defects can be labeled with different colors and represented with clear outlines, obtaining the category, bounding box, and mask of each instance.
7. The method for detecting multiple types of exterior wall defects based on drone inspection according to claim 1 is characterized in that: The calculation formula of the loss function of the disease detection model is: L=L cls +L box +L mask Among them, L represents the loss function of the disease detection model, L cls represents the classification loss, L box represents the regression box loss, L mask represents the generated mask loss.
8. The method for detecting multiple types of exterior wall defects based on drone inspection according to claim 1 is characterized in that: The step (5) includes the following sub-steps: (5.1) Obtaining local defect maps: Using the trained defect detection model, obtain local defect maps of the visible light and infrared thermal imaging images to be detected; (5.2) Extract feature points from the local defect image: Use the SURF algorithm to detect the feature points of the local defect image using the integral image and Hessian matrix approximation, which serve as the matching point set to be registered; (5.3) Image registration: Based on the matching point set to be registered obtained in step (5.2), the two local defect images are converted to the same coordinate system and then image registration is performed; (5.4) Image copying: copy the local defect map to the registration map obtained in step (5.3) to obtain a spliced image; (5.5) Image crack removal processing: In the overlapping areas of the image spliced in step (5.4), the pixel value weights of the overlapping areas are given according to the reliability of the local defect map, and the pixel values of the overlapping areas of multiple local defect maps are fused by weighted averaging. The pixel value weights of the overlapping areas are then added together to form a new image as the global defect detection map of the building facade to be tested.
9. The method for detecting multiple types of exterior wall defects based on drone inspection according to claim 8 is characterized in that: The SURF algorithm is used to detect the feature points of the local defect map using the integral image and the Hessian matrix approximation, which specifically includes the following sub-steps: (5.2.1) Obtain the approximate Hessian matrix corresponding to the local defect map: First, search for the local defect map in all scale spaces, perform Gaussian filtering on the local defect map, and construct the Hessian matrix to obtain the filtered Hessian matrix, which is expressed as: Where H(x,y,σ) represents the filtered Hessian matrix, (x,y) represents the position of the pixel in the local defect map; L(x,y,σ)=G(σ)*I(x,y) represents the Gaussian scale space of the local defect map; G(σ) represents Gaussian convolution, σ represents the standard deviation of Gaussian convolution, and I(x,y) represents the pixel value of the pixel at the position (x,y) in the local defect map; L xx () means taking the derivative twice of the independent variable x, L xy () indicates that the derivative of the independent variable x is taken first and then the derivative of the independent variable y. yy () means taking the derivative of the independent variable y twice; then the discriminant of the Hessian matrix is obtained according to the filtered Hessian matrix, and its expression is: Det(H)=L xx *L yy -L xy *L xy Where Det(H) represents the discriminant value of the Hessian matrix. Then, a box filter is used to replace the Gaussian filter, and the approximate expression of the Hessian matrix discriminant is obtained by multiplying the weighting coefficient λ: Det'(H)=L xx *L yy -(λ*L xy ) 2 Where Det'(H) represents the Hessian matrix approximation; (5.2.2) Constructing scale space: The scale space of the SURF algorithm consists of O groups of S layers, each with different sizes; (5.2.3) Feature point filtering and precise positioning: The Hessian matrix approximation of each pixel point processed by the Hessian matrix is compared with the Hessian matrix approximations of all its neighboring points in the image domain and scale domain. If the Hessian matrix approximation of a pixel point in the local defect image is greater than or less than the Hessian matrix approximations of all its neighboring points, the pixel point is considered an extreme point and is used as a feature point. After the feature points are initially located through the above steps, weak energy feature points and incorrectly located feature points are filtered out to select the final feature points. (5.2.4) Calculate the main direction of the feature point: The SURF algorithm determines the main direction of the feature point by counting the Harr wavelet features in the circular neighborhood of the feature point. Specifically, within the circular neighborhood of the feature point, the sum of the horizontal and vertical Harr wavelet features of all points in a 60° sector is counted as the Harr wavelet eigenvalue within the sector area; the sector is rotated at intervals of 0.2 radians, and the Harr wavelet eigenvalues within the sector area are counted again; the above process is repeated to obtain the Harr wavelet eigenvalues of all sector areas in the circular neighborhood of the feature point; finally, the direction of the sector corresponding to the largest Harr wavelet eigenvalue is taken as the main direction of the feature point; (5.2.5) Generate feature description: Use the SURF algorithm to extract a 4×4 rectangular area around the feature point. In each sub-area, calculate the Haar wavelet features of 25 pixels in the horizontal and vertical directions. The Haar wavelet features are the sum of the horizontal values, the sum of the vertical values, the sum of the horizontal absolute values, and the sum of the vertical absolute values. (5.2.6) Obtain the set of matching points to be registered: Construct the set of matching points to be registered based on the filtered feature points, main directions of feature points, and feature descriptions.
10. A system for implementing the method for detecting multiple types of exterior wall defects based on drone inspection according to any one of claims 1 to 9, characterized in that: include: The image acquisition module is used to plan the drone's shooting route based on the surrounding environment of the inspected building and the size information of the building facade, and to collect visible light and infrared thermal imaging images of the building facade through the drone; The data preprocessing and dataset construction module is used to preprocess visible light and infrared thermal imaging images to construct a dataset and randomly divide it into training set, validation set and test set in proportion; The generator acquisition module is used to construct an image fusion network based on a generative adversarial network. The image fusion network includes a generator, a visible light image discriminator, and an infrared thermal imaging discriminator. The image fusion network is adversarially trained using a training set. During the training process, the parameters of the image fusion network are adjusted with the goal of minimizing the loss function of the image fusion network to obtain a trained generator. A disease detection model acquisition module is used to build a disease detection model based on the trained generator. The disease detection model includes the trained generator and the open-source Mask R-CNN network. The disease detection model is trained using the training set. During the training process, the loss function of the disease detection model is minimized as the optimization goal. Only the parameters of the Mask R-CNN network are adjusted, and the parameters of the generator are frozen to obtain the trained disease detection model. The global defect detection map acquisition module is used to use the trained disease detection model to obtain the local defect map of the visible light and infrared thermal imaging images to be detected. The SURF algorithm is used to extract feature points, match, register, and splice the local defect map, perform feature processing on the overlapping areas of the image, and realize the splicing of the local defect map to obtain the global defect detection map of the facade of the building to be tested.
Citation Information
Patent Citations
Beam bottom crack detection method, system and device based on image processing and medium
CN111008956A
Power equipment fault diagnosis method based on combination of image fusion and target detection
CN112733950A
Multi-granularity academic emotion recognition method combined with context information
CN115346259A
Building exterior wall defect detection method based on deep learning multi-modal image fusion
CN116091477A
Forest fire detection method based on unmanned aerial vehicle dual-mode image fusion
CN118379650A
Cited By
Intelligent detection system and method for bonding quality of building external wall thermal insulation layer
CN122335865A
An intelligent detection system and method for the bonding quality of the thermal insulation layer of a house building outer wall
CN122335865B