Image recognition and positioning method for building facade diseases

By taking pictures of building facade images and using image recognition network models for disease identification, the problems of low recognition accuracy and efficiency in the prior art are solved, and the rapid and accurate identification and positioning of building facade diseases are achieved, replacing traditional manual detection and reducing the risk of detection.

CN119992374AActive Publication Date: 2025-05-13CHINA ACAD OF BUILDING RES +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411973686.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The identification accuracy and efficiency of building facade diseases in the prior art need to be improved. Traditional manual detection methods are subjective, dangerous and inefficient. New methods such as laser scanning and 3D thermal imaging are safer, but time-consuming and have lower recognition accuracy.

Method used

Drones are used to take partial images of the building facade, and disease recognition and positioning are performed through image recognition network models. The network model includes a skeleton network, a neck network and a detection head network. It is trained using SIoU loss function, and combined with SIFT and RANSAC algorithms to perform image stitching to achieve rapid and accurate identification of diseases.

Benefits of technology

It realizes rapid and accurate identification of building facade diseases, improves detection efficiency, reduces calculation costs, and replaces traditional manual testing, greatly reducing the detection risk.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992374A_ABST
    Figure CN119992374A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition and positioning method for building facade diseases. The method comprises the following steps: shooting by an unmanned aerial vehicle at an angle that a camera lens directly faces a building facade to obtain a plurality of local images of the building facade; obtaining a data set from the positive image sample and the negative image sample; constructing a building facade disease image recognition network model, wherein the building facade disease image recognition network model comprises a skeleton network, a neck network and a detection head network; an optimal building facade disease image recognition network model is obtained; and carrying out splicing by using a scale invariant feature transform (SIFT) algorithm and a random sample consensus (RANSAC) algorithm to obtain a final building facade overall identification image. According to the method, rapid and accurate recognition of the building facade diseases can be realized, the detection efficiency can be improved, and the method has a relatively accurate recognition effect and relatively high practicability and engineering significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an image recognition and positioning method for building facade defects, and belongs to the technical field of image recognition. Background Art

[0002] The presence of building facade defects is a pressing issue during the building operation phase, and is usually attributed to mechanical and environmental factors. Typical defects are manifested as cracks, exterior wall peeling, or water seepage on the exterior wall. These defects affect the appearance and reduce the service life of the building. More seriously, the exterior wall peeling may cause safety accidents and irreparable losses. Structural damage detection is an integral part of structural health monitoring and is essential to ensure the safe operation of buildings. As an integral part of structural damage detection, the detection of building facade defects can enable the government and management to accurately understand the comprehensive condition of the building's exterior wall, thereby helping to formulate a reasonable maintenance plan. This is an effective way to reduce building maintenance costs, extend the service life of buildings, and mitigate the impact of exterior wall damage. Many countries and regions are formulating policies for regular standardized visual inspections. The detection of building facade defects has become a key component of building maintenance.

[0003] At present, visual inspection is a simple and reliable method to evaluate the condition of building appearance. Traditional building appearance inspection usually requires professionals to arrive at the inspection site with special tools and use visual observation, hammering and other techniques for evaluation. These methods rely on the inspector's expertise and experience, which is subjective, dangerous and inefficient. Due to the increase in the number and scale of buildings, manual visual inspection methods are no longer sufficient to meet the requirements of large-scale inspections. With the advancement of technology, many new methods (such as laser scanning, 3D thermal imaging and SLAM) are being used for exterior wall disease detection through drones and robotic platforms. Compared with traditional technologies, these new methods are more convenient and safer, but they are time-consuming, and the recognition accuracy and detection efficiency are low, which faces challenges in meeting the needs of large-scale inspections. Therefore, it is necessary to develop a more accurate and effective method for building facade disease detection to improve detection efficiency and reduce computational costs. Summary of the invention

[0004] The purpose of the present invention is to provide an image recognition and positioning method for building facade defects to solve the problem that the recognition accuracy and efficiency need to be improved in the prior art.

[0005] The technical solution of the present invention is:

[0006] A method for image recognition and positioning of building facade defects comprises the following steps:

[0007] S1, using a drone to take photos with a camera lens facing the building facade to obtain multiple partial images of the building facade;

[0008] S2. Screen the acquired partial images of building facades, use the partial images of building facades with building facade diseases as positive image samples, and annotate the types of diseases on the partial images of building facades, use the partial images of building facades without building facade diseases as negative image samples, obtain a data set from the positive image samples and the negative image samples, and divide the data set into a training set and a test set according to a set ratio;

[0009] S3. Construct a network model for image recognition of building facade defects. The network model for image recognition of building facade defects includes a skeleton network, a neck network, and a detection head network. The skeleton network performs convolution, pooling, and vector concatenation on the input local image of the building facade and outputs feature vectors of three scales. The feature vectors of three scales are input into the neck network and convolution, up-sampling, and vector concatenation are performed to obtain a feature map. The feature map is input into the detection head network to obtain a defect prediction image, including a prediction image with a prediction box or an image without a prediction box, i.e., an image without a defect.

[0010] S4, using the training set and the SIoU loss function that considers angle loss, distance loss and shape loss, the building facade disease image recognition network model constructed in step S3 is trained to obtain the trained building facade disease image recognition network model, the test set is input into the trained building facade disease image recognition network model, and the trained building facade disease image recognition network model is hyperparameter optimized according to the test set results to obtain the optimal building facade disease image recognition network model;

[0011] S5. Input multiple partial images of building facades to be identified into the optimal building facade disease image recognition network model to obtain multiple disease prediction images, and use the scale-invariant feature transformation algorithm SIFT and the random sampling consensus algorithm RANSAC to splice them to obtain the final building facade overall recognition image.

[0012] Furthermore, in step S1, the building facade defects include cracks, exterior wall peeling and / or exterior wall water seepage.

[0013] Furthermore, in the skeleton network, after the input local image of the building facade enters the skeleton network, it passes through 4 convolutional normalization activation modules CBS, the first high-efficiency layer aggregation network module ELAN1, the first maximum pooling module MP1, the second high-efficiency layer aggregation network module ELAN2, the second maximum pooling module MP2, the third high-efficiency layer aggregation network module ELAN3, the third maximum pooling module MP3 and the fourth high-efficiency layer aggregation network module ELAN4 in sequence, and the second high-efficiency layer aggregation network module ELAN2 outputs the feature vector T1, the third high-efficiency layer aggregation network module ELAN3 outputs the feature vector T2, and the fourth high-efficiency layer aggregation network module ELAN4 outputs the feature vector T3 to the neck network respectively.

[0014] Furthermore, in the neck network, the feature vector T3 is outputted as the feature vector T4 through the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network. The feature vector T4 is outputted to the first vector concatenation module Concat1 after passing through the first convolution normalization activation module CBS1 and the first upsampling module UPsample1. At the same time, the feature vector T4 is inputted into the fourth vector concatenation module Concat4. The feature vector T2 is outputted to the first vector concatenation module Concat1 after passing through the second convolution normalization activation module CBS2. The first vector concatenation module Concat1 performs tensor concatenation and outputs the feature vector T5.

[0015] The feature vector T5 is output to the second vector concatenation module Concat2 after passing through the first weighted efficient layer aggregation network module ELAN-W1, the third convolution normalization activation module CBS3 and the second upsampling module UPsample2. At the same time, the feature vector T5 is output to the third vector concatenation module Concat3 after passing through the first weighted efficient layer aggregation network module ELAN-W1, and the feature vector T1 is output to the second vector concatenation module Concat2 after passing through the fourth convolution normalization activation module CBS4. The second vector concatenation module Concat2 outputs the feature vector T6 after concatenation.

[0016] After the feature vector T6 passes through the second weighted efficient layer aggregation network module ELAN-W2, the feature vector T7 is output. The feature vector T7 is directly input into the detection head network. At the same time, the feature vector T7 passes through the fourth maximum pooling module MP4 and is output to the third vector concatenation module Concat3. The third vector concatenation module Concat3 performs tensor concatenation to obtain the feature vector T8.

[0017] The feature vector T8 passes through the third weighted efficient layer aggregation network module ELAN-W3 to output the feature vector T9, and the feature vector T9 is directly input into the detection head network. At the same time, the feature vector T9 passes through the fifth maximum pooling module MP5 and is output to the fourth vector concatenation module Concat4; the fourth vector concatenation module Concat4 performs tensor concatenation to obtain the feature vector T10, and the feature vector T10 passes through the fourth weighted efficient layer aggregation network module ELAN-W4 to output the feature vector T11 to the detection head network.

[0018] Further, in the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network, the feature vector T3 is input into the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network, and then passes through the fifth convolution normalization activation module CBS5 to output the feature vector t1. The feature vector t1 is directly input into the fifth concatenation module Concat5, and the feature vector t1 is input into the sixth maximum pooling module MP6. The sixth maximum pooling module MP6 outputs the feature vector t2. The feature vector t2 is directly input into the fifth concatenation module Concat5, and the feature vector t2 is input into the seventh maximum pooling module MP7. The seventh maximum pooling module MP7 outputs the feature vector t3. The feature vector t3 is directly input into the fifth concatenation module Concat5, and the feature vector t3 is input into the eighth maximum pooling module MP8. The eighth maximum pooling module MP8 outputs the feature vector t4. The feature vector t4 is directly input into the fifth concatenation module Concat5. The fifth concatenation module Concat5 outputs the feature vector t5. After the feature vector t5 passes through the sixth convolution normalization activation module CBS6, the feature vector T4 is obtained.

[0019] Furthermore, in the detection head network, the feature vector T7 passes through the first residual convolution normalization activation module REP+CBM1 and is input into the first classification module Classification1; the feature vector T9 passes through the second residual convolution normalization activation module REP+CBM2 and is input into the second classification module Classification2; the feature vector T11 passes through the third residual convolution normalization activation module REP+CBM3 and is input into the third classification module Classification3. Finally, the detection head network outputs a disease prediction image including a prediction image with a prediction box or an image without a prediction box, that is, an image without a disease.

[0020] Furthermore, in step S5, the scale-invariant feature transformation algorithm SIFT and the random sampling consensus algorithm RANSAC are used to splice the recognized partial images of the building facade to obtain the final overall recognition image of the building facade, specifically,

[0021] S51, the size of the image with the recognition result prediction box is H×W, and the coordinates of the center point of the prediction box of the recognition result prediction box (x i ,y i ), prediction box width w i , prediction box height h i And the predicted category p i ;

[0022] S52, calling the scale-invariant feature transform function, i.e., the SIFT function, in the cross-platform computer vision library OpenCV to perform extreme value detection, feature point location, feature point assignment, and feature description in the scale space of the defect prediction image output by the building facade defect image recognition network model, and then calling the random sampling consensus function, i.e., the RANSAC function, to complete feature detection and extraction, feature matching, perspective transformation, and image fusion of the image, to obtain a complete building facade image;

[0023] S53, the prediction box center point coordinates (x i ,y i ) is used for conversion: first mark the coordinate origin (x0, y0) of the local image before stitching, then find the coordinates (X0′, Y0′) of (x0, y0) in the stitched image, and calculate the coordinates (X0′, Y0′) of the center point of the prediction box in the complete building facade image. i ,Y i ):

[0024]

[0025] S53, according to the coordinates (X i ,Y i )Draw width w i , height h i After splicing the prediction frame to the complete building facade image, the final complete recognition image with the prediction frame is obtained.

[0026] The beneficial effects of the present invention are as follows: compared with the prior art, the image recognition and positioning method of building facade defects can realize rapid and accurate recognition of building facade defects, can improve detection efficiency, has a more accurate recognition effect, has strong practicality and engineering significance, can replace traditional manual detection, and greatly reduce the risk of detection. The method can realize the display of the defect location in the complete building facade image, the visualization effect of the defect location is better, and can provide more data support for building facade detection work. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of the flow of the image recognition and positioning method of building facade defects according to an embodiment of the present invention;

[0028] Figure 2 Schematic diagram for explaining the network model for identifying building facade disease images in the embodiment;

[0029] Figure 3 It is a schematic diagram for explaining the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network in the embodiment;

[0030] Figure 4 Schematic diagrams for marking the types of damage in the embodiments, (a) is a schematic diagram for marking cracks, (b) is a schematic diagram for marking water seepage on the exterior wall, and (c) is a schematic diagram for marking falling off of the exterior wall;

[0031] Figure 5 A schematic diagram of an image recognition result in an embodiment;

[0032] Figure 6 A schematic diagram of image stitching in an embodiment;

[0033] Figure 7 Schematic diagram of the final recognition image in the embodiment. DETAILED DESCRIPTION

[0034] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0035] The embodiment provides a method for image recognition and location of building facade defects, such as Figure 1 , including the following steps,

[0036] S1. A drone takes a plurality of partial images of the building facade by using a camera lens facing the building facade.

[0037] In step S1, the building facade defects include cracks, exterior wall peeling and / or exterior wall water seepage. The drone is designed with an automatic navigation route and shooting position to collect images and obtain the building facade image.

[0038] S2. Screen the acquired partial images of building facades, take the partial images of building facades with building facade diseases as positive image samples, and mark the types of diseases on the partial images of building facades, take the partial images of building facades without building facade diseases as negative image samples, obtain a data set from the positive image samples and the negative image samples, and divide the data set into a training set and a test set according to a set ratio.

[0039] In step S2, the data set is divided into a training set and a test set in a ratio of 8:2. The training set is used to train the model, and the test set is used to test the model.

[0040] S3, build a network model for building facade disease image recognition, such as Figure 2The network model for building facade defect image recognition includes a skeleton network, a neck network and a detection head network. The skeleton network performs convolution, pooling and vector splicing on the input local image of the building facade and outputs feature vectors of three scales. The feature vectors of three scales are input into the neck network and a feature map is obtained after convolution, upsampling and vector splicing. The feature map is input into the detection head network to obtain a defect prediction image, including a prediction image with a prediction box or an image without a prediction box, that is, an image without a defect.

[0041] In step S3, the backbone network Backbone includes four sequentially arranged convolution normalization activation modules CBS, the first high-efficiency layer aggregation network module ELAN1, the first maximum pooling module MP1, the second high-efficiency layer aggregation network module ELAN2, the second maximum pooling module MP2, the third high-efficiency layer aggregation network module ELAN3, the third maximum pooling module MP3 and the fourth high-efficiency layer aggregation network module ELAN4. In the backbone network Backbone, after the input building facade partial image enters the backbone network, it passes through four convolution normalization activation modules CBS, the first high-efficiency layer aggregation network module ELAN1, the first maximum pooling module MP1, the second high-efficiency layer aggregation network module ELAN2, the second maximum pooling module MP2, the third high-efficiency layer aggregation network module ELAN3, the third maximum pooling module MP3 and the fourth high-efficiency layer aggregation network module ELAN4 in sequence, and the second high-efficiency layer aggregation network module ELAN2 outputs the feature vector T1, the third high-efficiency layer aggregation network module ELAN3 outputs the feature vector T2, and the fourth high-efficiency layer aggregation network module ELAN4 outputs the feature vector T3 to the neck network Neck. Among them, eigenvector T1 is the eigenvector with the largest scale and the least eigenvalue among eigenvectors T1, T2, and T3, eigenvector T2 is the eigenvector with medium scale and medium eigenvalue, and eigenvector T3 is the eigenvector with the smallest scale and the most eigenvalue.

[0042] In step S3, the neck network Neck includes a spatial pyramid pooling module SPPELAN of an efficient layer aggregation network, 4 convolutional normalization activation modules CBS, 2 upsampling modules UPsample, 4 vector concatenation modules Concat, 2 maximum pooling modules MP and 4 weighted efficient layer aggregation network modules ELAN-W.

[0043] In the neck network Neck, the feature vector T3 passes through the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network to output the feature vector T4, and the feature vector T4 passes through the first convolution normalization activation module CBS1 and the first upsampling module UPsample1 and is output to the first vector splicing module Concat1. At the same time, the feature vector T4 is input to the fourth vector splicing module Concat4; the feature vector T2 passes through the second convolution normalization activation module CBS2 and is output to the first vector splicing module Concat1. The first vector splicing module Concat1 performs tensor splicing and outputs the feature vector T5;

[0044] The feature vector T5 is output to the second vector concatenation module Concat2 after passing through the first weighted efficient layer aggregation network module ELAN-W1, the third convolution normalization activation module CBS3 and the second upsampling module UPsample2. At the same time, the feature vector T5 is output to the third vector concatenation module Concat3 after passing through the first weighted efficient layer aggregation network module ELAN-W1, and the feature vector T1 is output to the second vector concatenation module Concat2 after passing through the fourth convolution normalization activation module CBS4. The second vector concatenation module Concat2 outputs the feature vector T6 after concatenation.

[0045] After the feature vector T6 passes through the second weighted efficient layer aggregation network module ELAN-W2, the feature vector T7 is output. The feature vector T7 can be directly input into the detection head network. At the same time, the feature vector T7 passes through the fourth maximum pooling module MP4 and is output to the third vector concatenation module Concat3. The third vector concatenation module Concat3 performs tensor concatenation to obtain the feature vector T8.

[0046] After the feature vector T8 passes through the third weighted efficient layer aggregation network module ELAN-W3, the feature vector T9 is output. The feature vector T9 can be directly input into the detection head network. At the same time, the feature vector T9 passes through the fifth maximum pooling module MP5 and is output to the fourth vector concatenation module Concat4; the fourth vector concatenation module Concat4 performs tensor concatenation to obtain the feature vector T10. After the feature vector T10 passes through the fourth weighted efficient layer aggregation network module ELAN-W4, the feature vector T11 is output to the detection head network Head.

[0047] The detection head network Head includes the first residual convolution normalization activation module REP+CBM1, the first classification module Classification1, the second residual convolution normalization activation module REP+CBM2, the second classification module Classification2, the third residual convolution normalization activation module REP+CBM3 and the third classification module Classification3. In the detection head network, the feature vector T7 is input into the first classification module Classification1 after passing through the first residual convolution normalization activation module REP+CBM1; the feature vector T9 is input into the second classification module Classification2 after passing through the second residual convolution normalization activation module REP+CBM2; the feature vector T11 is input into the third classification module Classification3 after passing through the third residual convolution normalization activation module REP+CBM3. Finally, the detection head network outputs a disease prediction image including a prediction image with a prediction box or an image without a prediction box, that is, an image without a disease.

[0048] like Figure 3 The spatial pyramid pooling module SPPELAN of the efficient layer aggregation network includes the fifth convolution normalization activation module CBS5, the sixth maximum pooling module MP6, the seventh maximum pooling module MP7, the eighth maximum pooling module MP8, the fifth splicing module Concat5 and the sixth convolution normalization activation module CBS6. In the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network, the feature vector T3 is input into the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network, and then passes through the fifth convolution normalization activation module CBS5 to output the feature vector t1. The feature vector t1 is directly input into the fifth concatenation module Concat5, and the feature vector t1 is input into the sixth maximum pooling module MP6. The sixth maximum pooling module MP6 outputs the feature vector t2. The feature vector t2 is directly input into the fifth concatenation module Concat5, and the feature vector t2 is input into the seventh maximum pooling module MP7. The seventh maximum pooling module MP7 outputs the feature vector t3. The feature vector t3 is directly input into the fifth concatenation module Concat5, and the feature vector t3 is input into the eighth maximum pooling module MP8. The eighth maximum pooling module MP8 outputs the feature vector t4. The feature vector t4 is directly input into the fifth concatenation module Concat5. The fifth concatenation module Concat5 outputs the feature vector t5. After the feature vector t5 passes through the sixth convolution normalization activation module CBS6, the feature vector T4 is obtained.

[0049] S4, using the training set and the SIoU loss function that considers angle loss, distance loss and shape loss, the building facade disease image recognition network model constructed in step S3 is trained to obtain the trained building facade disease image recognition network model, the test set is input into the trained building facade disease image recognition network model, and the trained building facade disease image recognition network model is hyperparameter optimized according to the test set results to obtain the optimal building facade disease image recognition network model;

[0050] In step S4, the experimental GPU uses RTX3090 and the CPU uses i7-9700. During the initial training, the hyperparameters are: Ir=0.01, Irf=0.1, batchsize=16, epoch=300. After training the building facade disease image recognition network model, the test set is input into the trained building facade disease image recognition network model, and the hyperparameters are optimized according to the test set results to obtain the optimal building facade disease image recognition network model.

[0051] S5. Input multiple partial images of building facades to be identified into the optimal building facade disease image recognition network model to obtain multiple disease prediction images, and use the scale-invariant feature transformation algorithm SIFT and the random sampling consensus algorithm RANSAC to splice them to obtain the final building facade overall recognition image.

[0052] In step S5, the scale-invariant feature transformation algorithm SIFT and the random sampling consensus algorithm RANSAC are used for splicing to obtain the final overall recognition image of the building facade, specifically,

[0053] S51, the size of the image with the recognition result prediction box is H×W, and the coordinates of the center point of the prediction box of the recognition result prediction box (x i ,y i ), prediction box width w i , prediction box height h i And the predicted category p i ;

[0054] S52, calling the scale-invariant feature transform function, i.e., the SIFT function, in the cross-platform computer vision library OpenCV to perform extreme value detection, feature point location, feature point assignment, and feature description in the scale space of the defect prediction image output by the building facade defect image recognition network model, and then calling the random sampling consensus function, i.e., the RANSAC function, to complete feature detection and extraction, feature matching, perspective transformation, and image fusion of the image, to obtain a complete building facade image;

[0055] S53, the prediction box center point coordinates (x i ,yi ) is used for conversion: first mark the coordinate origin (x0, y0) of the local image before stitching, then find the coordinates (X0′, Y0′) of (x0, y0) in the stitched image, and calculate the coordinates (X0′, Y0′) of the center point of the prediction box in the complete building facade image. i ,Y i ):

[0056]

[0057] S54, according to the coordinates (X i ,Y i )Draw width w i , height h i After splicing the prediction frame to the complete building facade image, the final complete recognition image with the prediction frame is obtained.

[0058] In step S5, the image with the recognition result prediction box obtained by the optimal building facade disease image recognition network model is a local image, so the scale-invariant feature transformation algorithm SIFT and the random sampling consensus algorithm RANSAC are used to splice the image to obtain a complete building facade image.

[0059] Compared with the existing technology, this image recognition and positioning method for building facade defects can realize rapid, accurate and automatic recognition of building facade defects, improve detection efficiency, have a more accurate recognition effect, have strong practicality and engineering significance, can replace traditional manual detection, and greatly reduce the risk of detection. This method has many advantages such as high recognition accuracy and good adaptability. It can realize the display of the defect location in the complete building facade image, and the visualization effect of the defect location is better, which can provide more data support for building facade detection.

[0060] A specific example of the image recognition and positioning method of the building facade disease of the embodiment is as follows:

[0061] S1. The drone takes photos of the building facade with its camera lens facing the building facade directly to obtain the building facade image.

[0062] S2. Screen the acquired building facade images, take the building facade images with building facade defects as positive image samples, and annotate the building facade images with defect types, including cracks, etc. Figure 4 (a) The exterior wall falls off. Figure 4 (b) and / or water seepage from external walls Figure 4 (c) , taking the building facade images without building facade diseases as negative image samples, obtaining a data set from the positive image samples and the negative image samples, and dividing the data set into a training set and a test set according to a set ratio.

[0063] S3. Construct a network model for image recognition of building facade defects. The network model for image recognition of building facade defects includes a skeleton network, a neck network and a detection head network.

[0064] S4. Use the training set and the SIoU loss function that considers angle loss, distance loss and shape loss to train the building facade disease image recognition network model constructed in step S3 to obtain the trained building facade disease image recognition network model. Input the test set into the trained building facade disease image recognition network model, and optimize the hyperparameters of the trained building facade disease image recognition network model according to the test set results to obtain the optimal building facade disease image recognition network model.

[0065] S5. Input the building facade image to be identified into the optimal building facade disease image recognition network model to obtain an image with a recognition result prediction frame, and use the scale-invariant feature transformation algorithm SIFT and the random sampling consensus algorithm RANSAC to splice the image with the recognition result prediction frame and the complete building facade image to obtain the final recognition image.

[0066] S51, such as Figure 5 , the image size with the recognition result prediction box is 800×600, the prediction box center coordinates (x1, y1) = (749, 324), the prediction box width w1 = 120, and the prediction box height h1 = 216;

[0067] S52. Calling SIFT in OpenCV can complete the extreme value detection, feature point location, feature point assignment and feature description of the image scale space, and then calling the RANSAC function to complete the image feature detection and extraction, feature matching, perspective transformation and image fusion to obtain a complete building facade image such as Figure 6 ;

[0068] S53, transform the coordinates of the center point of the prediction box (749, 324): first mark the coordinate origin (0, 0) of the local image before stitching, then find out the coordinates (1200, 0) of (0, 0) in the stitched image, and calculate the coordinates of the center point of the prediction box in the complete building facade image (1949, 324).

[0069] S54, the prediction frame obtained by local image recognition is finally displayed in a unified manner to form a complete building facade image after splicing. Figure 7 .

[0070] This image recognition and location method for building facade defects uses a trained deep learning model to identify building facade defects, and reflects the defect coordinates on the complete building facade image through image stitching and coordinate conversion. This method can quickly and accurately identify building facade defects and realize visual viewing of defects on the entire building facade.

[0071] The above embodiments are only used to help understand the method and core idea of ​​the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A method for image recognition and location of building facade defects, characterized by: The following steps are included: S1, using a drone to shoot a plurality of partial images of the building facade with a camera lens facing the building facade; S2. Screen the acquired partial images of building facades, use the partial images of building facades with building facade diseases as positive image samples, and annotate the types of diseases on the partial images of building facades, use the partial images of building facades without building facade diseases as negative image samples, obtain a data set from the positive image samples and the negative image samples, and divide the data set into a training set and a test set according to a set ratio; S3. Construct a network model for image recognition of building facade defects. The network model for image recognition of building facade defects includes a skeleton network, a neck network, and a detection head network. The skeleton network performs convolution, pooling, and vector concatenation on the input local image of the building facade and outputs feature vectors of three scales. The feature vectors of three scales are input into the neck network and convolution, up-sampling, and vector concatenation are performed to obtain a feature map. The feature map is input into the detection head network to obtain a defect prediction image, including a prediction image with a prediction box or an image without a prediction box, i.e., an image without a defect. S4, using the training set and the SIoU loss function that considers angle loss, distance loss and shape loss, the building facade disease image recognition network model constructed in step S3 is trained to obtain the trained building facade disease image recognition network model, the test set is input into the trained building facade disease image recognition network model, and the trained building facade disease image recognition network model is hyperparameter optimized according to the test set results to obtain the optimal building facade disease image recognition network model; S5. Input multiple partial images of building facades to be identified into the optimal building facade disease image recognition network model to obtain multiple disease prediction images, and use the scale-invariant feature transformation algorithm SIFT and the random sampling consensus algorithm RANSAC to splice them to obtain the final building facade overall recognition image.

2. The image recognition and positioning method for building facade defects according to claim 1, characterized in that: In step S1, the building facade defects include cracks, exterior wall peeling and / or exterior wall water seepage.

3. The image recognition and positioning method for building facade defects according to claim 1, characterized in that: In the skeleton network, after the input local image of the building facade enters the skeleton network, it passes through 4 convolutional normalization activation modules CBS, the first high-efficiency layer aggregation network module ELAN1, the first maximum pooling module MP1, the second high-efficiency layer aggregation network module ELAN2, the second maximum pooling module MP2, the third high-efficiency layer aggregation network module ELAN3, the third maximum pooling module MP3 and the fourth high-efficiency layer aggregation network module ELAN4 in sequence, and the second high-efficiency layer aggregation network module ELAN2 outputs the feature vector T1, the third high-efficiency layer aggregation network module ELAN3 outputs the feature vector T2, and the fourth high-efficiency layer aggregation network module ELAN4 outputs the feature vector T3 to the neck network.

4. The image recognition and location method of building facade defects according to claim 1, characterized in that: In the neck network, the feature vector T3 is outputted as the feature vector T4 through the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network. The feature vector T4 is outputted to the first vector concatenation module Concat1 after passing through the first convolution normalization activation module CBS1 and the first upsampling module UPsample1. At the same time, the feature vector T4 is inputted into the fourth vector concatenation module Concat4. The feature vector T2 is passed through the second convolution normalization activation module CBS2 and then output to the first vector concatenation module Concat1. The first vector concatenation module Concat1 performs tensor concatenation and outputs the feature vector T5. The feature vector T5 is output to the second vector concatenation module Concat2 after passing through the first weighted efficient layer aggregation network module ELAN-W1, the third convolution normalization activation module CBS3 and the second upsampling module UPsample2. At the same time, the feature vector T5 is output to the third vector concatenation module Concat3 after passing through the first weighted efficient layer aggregation network module ELAN-W1, and the feature vector T1 is output to the second vector concatenation module Concat2 after passing through the fourth convolution normalization activation module CBS4. The second vector concatenation module Concat2 outputs the feature vector T6 after concatenation. After the feature vector T6 passes through the second weighted efficient layer aggregation network module ELAN-W2, the feature vector T7 is output. The feature vector T7 is directly input into the detection head network. At the same time, the feature vector T7 passes through the fourth maximum pooling module MP4 and is output to the third vector concatenation module Concat3. The third vector concatenation module Concat3 performs tensor concatenation to obtain the feature vector T8. The feature vector T8 passes through the third weighted efficient layer aggregation network module ELAN-W3 to output the feature vector T9, and the feature vector T9 is directly input into the detection head network. At the same time, the feature vector T9 passes through the fifth maximum pooling module MP5 and is output to the fourth vector concatenation module Concat4; the fourth vector concatenation module Concat4 performs tensor concatenation to obtain the feature vector T10, and the feature vector T10 passes through the fourth weighted efficient layer aggregation network module ELAN-W4 to output the feature vector T11 to the detection head network.

5. The image recognition and positioning method for building facade defects according to claim 1, characterized in that: In the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network, the feature vector T3 is input into the spatial pyramid pooling module SPPELAN of the efficient layer aggregation network, and then passes through the fifth convolution normalization activation module CBS5 to output the feature vector t1. The feature vector t1 is directly input into the fifth concatenation module Concat5, and the feature vector t1 is input into the sixth maximum pooling module MP6. The sixth maximum pooling module MP6 outputs the feature vector t2. The feature vector t2 is directly input into the fifth concatenation module Concat5, and the feature vector t2 is input into the seventh maximum pooling module MP7. The seventh maximum pooling module MP7 outputs the feature vector t3. The feature vector t3 is directly input into the fifth concatenation module Concat5, and the feature vector t3 is input into the eighth maximum pooling module MP8. The eighth maximum pooling module MP8 outputs the feature vector t4. The feature vector t4 is directly input into the fifth concatenation module Concat5. The fifth concatenation module Concat5 outputs the feature vector t5. After the feature vector t5 passes through the sixth convolution normalization activation module CBS6, the feature vector T4 is obtained.

6. The image recognition and location method of building facade defects according to claim 1, characterized in that: In the detection head network, the feature vector T7 passes through the first residual convolution normalization activation module REP+CBM1 and is input into the first classification module Classification1; the feature vector T9 passes through the second residual convolution normalization activation module REP+CBM2 and is input into the second classification module Classification2; the feature vector T11 passes through the third residual convolution normalization activation module REP+CBM3 and is input into the third classification module Classification3. Finally, the detection head network outputs a disease prediction image including a prediction image with a prediction box or an image without a prediction box, that is, an image without a disease.

7. The image recognition and location method of building facade defects according to claim 1, characterized in that: In step S5, the scale-invariant feature transformation algorithm SIFT and the random sampling consensus algorithm RANSAC are used for splicing to obtain the final overall recognition image of the building facade, specifically, S51, the size of the image with the recognition result prediction box is H×W, and the coordinates of the center point of the prediction box of the recognition result prediction box (x i ,y i ), prediction box width w i , prediction box height h i And the predicted category p i ; S52, calling the scale-invariant feature transform function, i.e., the SIFT function, in the cross-platform computer vision library OpenCV to perform extreme value detection, feature point location, feature point assignment, and feature description in the scale space of the defect prediction image output by the building facade defect image recognition network model, and then calling the random sampling consensus function, i.e., the RANSAC function, to complete feature detection and extraction, feature matching, perspective transformation, and image fusion of the image, to obtain a complete building facade image; S53, the prediction box center point coordinates (x i ,y i ) is used for conversion: first mark the coordinate origin (x0, y0) of the local image before stitching, then find the coordinates (X0′, Y0′) of (x0, y0) in the stitched image, and calculate the coordinates (X0′, Y0′) of the center point of the prediction box in the complete building facade image. i ,Y i ): S53, according to the coordinates (X i ,Y i )Draw width w i , height h i After splicing the prediction frame to the complete building facade image, the final complete recognition image with the prediction frame is obtained.

Citation Information

Patent Citations

  • Concrete bridge apparent crack identification method based on novel attention mechanism

    CN115205230A

  • Grape leaf disease detection method based on deep learning

    CN117315648A

  • Road defect image detection method, equipment, medium and product

    CN118587500A

  • Harm prevention monitoring system and method

    WO2023164782A1