A method for identifying and locating damages of masonry building heritage based on computer vision
By improving the YOLOv5n network and introducing a variety of technical means, an Improved-YOLOv5n target detection network and multi-threshold image segmentation algorithm for masonry building heritage damage detection were developed, which solved the problem of automatic detection and quantification of masonry building heritage damage under complex backgrounds, and achieved high-precision and high-efficiency detection.
Patent Information
- Application Number
- CN202410826154.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-06-25
AI Technical Summary
It is difficult for the existing technology to automatically detect and quantify the damage of masonry building heritage in complex contexts. Traditional methods have problems such as low detection efficiency, poor accuracy, high risk and high difficulty.
Using a computer vision-based method, high-precision pictures were taken by drones, a large-view masonry building heritage damage data set was established, the YOLOv5n network was improved, SEAttention, Focal Loss and pruning technology was introduced, and the Improved-YOLOv5n object detection network was developed in combination with model distillation technology, and the loss positioning and quantization was used using a multi-threshold image segmentation algorithm based on genetic algorithm.
It has achieved rapid and accurate detection of damage to the heritage of masonry building, with a detection accuracy of 85.5%, and a detection speed of 824.3 FPS. It supports rapid detection, secondary development and reuse, and has the characteristics of automation and intelligence.
Smart Images

Figure CN118865163B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of civil engineering, and particularly relates to a method for identifying and positioning damages of masonry architectural heritage based on computer vision. Background Art
[0002] Architectural heritage has high historical value, artistic value and scientific value, and is an important material carrier of the country's "cultural confidence" and "cultural power". Masonry architectural heritage is the most important part of architectural heritage, with a large number and wide distribution. Due to the superposition of natural factors such as long-term weathering and rain erosion, as well as human factors such as traffic and construction, many masonry architectural heritages have damage diseases such as damage, weathering, cracking, deformation, and plant growth. Once these damage diseases accumulate to a certain extent, irreversible damage to the masonry architectural heritage will be caused, and in severe cases, the destruction of the cultural relic itself will be caused. Therefore, it is urgent to carry out preventive protection for masonry architectural heritage, predict risks in advance, and minimize the probability of risk occurrence. The structural safety monitoring of masonry architectural heritage is the main content of preventive protection, and an important damage index of masonry architectural heritage is damage. How to automatically detect and quantify the damage based on the limited collected image data in the case of complex backgrounds in the images is a difficult problem currently faced. Therefore, it is urgent to develop a damage detection method based on computer vision for the protection of masonry architectural heritage.
[0003] Currently, traditional detection methods rely on manual visual inspection supplemented by measurement or detection equipment, and have problems such as low detection efficiency, poor accuracy, high risk, and great difficulty. While computer vision methods are gradually showing great vitality in the quality inspection fields of the civil engineering industry and manufacturing, and their detection accuracy far exceeds that of traditional visual inspections. However, currently, there are still the following problems in the automatic detection and quantification of damages of masonry architectural heritage: (a) The establishment of datasets related to damage diseases of architectural heritage mostly involves artificial screening factors. For example, the images in the dataset of damage diseases of masonry architectural heritage usually only contain the surface of the damaged part, rather than the whole building. The noise in the images is small, the proportion of target damage diseases is large, and the generalization ability is weak; (b) In the research related to the target detection of architectural heritage, the network has weak feature extraction ability for aerial images with cluttered backgrounds, and does not pay attention to the importance of detection speed in the daily inspection of large-scale architectural heritage; (c) The image segmentation algorithms for cracks and spalling damages on the structural surface have developed relatively comprehensively, but the images usually have small noise, and the background and the target are on the same plane. The applicable image segmentation algorithms are difficult to segment large-angle images, and for masonry architectural heritage located in the wild, multiple image segmentation algorithms need to be combined to improve the segmentation accuracy. Summary of the Invention
[0004] To solve the above problems, the present invention discloses a method for identifying and locating damages of masonry architectural heritages based on computer vision, which supports rapid detection, secondary development, and reuse, has the characteristics of automation and intelligence, and ensures the accuracy of damage detection through computer vision methods.
[0005] To achieve the above object, the technical solution of the present invention is as follows:
[0006] A method for identifying and locating damages of masonry architectural heritages based on computer vision, comprising the following steps:
[0007] a) Set up an automatic damage identification module for masonry architectural heritages:
[0008] First step, form a data set. Use a drone to take pictures of masonry architectural heritages (including the main body of the architectural heritage and the environment) to obtain initial high-precision picture data, analyze the typical damage disease characteristics of masonry architectural heritages, and use the deep learning annotation software Labelme to annotate the pictures to obtain a data set of deficiencies and damages of masonry architectural heritages.
[0009] Second step, improve the YOLOv5n network according to the characteristics of UAV images to obtain an Improved-YOLOv5n automatic identification network. Since the background occupies a relatively large proportion in UAV aerial pictures, while the target damage disease area is relatively small, and the number of effective channels of the target object is small in multi-channel detection, SEAttention is introduced to assign different weights to different positions of the image from the channel perspective, so as to focus on more important feature information. The background part usually occupies a relatively large proportion in UAV aerial pictures, which is likely to cause an imbalance in the classification of the target and the background during training. Therefore, the Focal Loss idea is introduced to make the model pay more attention to difficult-to-classify samples. Using Focal-CIoU as the Bounding Box Regression Loss of the model has good results. Since masonry architectural heritages, such as the Great Wall, city walls, bridges, ancient towers, building complexes, etc., often have a large volume, high requirements are placed on the size, inference speed, and timeliness of the target detection model. By introducing a pruning method, the model size can be reduced without affecting the model performance, and the inference speed can be improved. After reducing the model size through pruning, it is still necessary to maintain the model performance. Therefore, the model distillation technology is introduced to train a smaller target model by learning the output of a complex model, so that its performance is close to that of the complex model. Finally, an Improved-YOLOv5n target detection network with a detection accuracy (mAP@0.5) reaching 85.5% and a detection speed (FPS) reaching 824.3 is obtained.
[0010] b) Set up an automatic damage location and quantification module for masonry architectural heritages:
[0011] First step: Set targets on the surface of the masonry building heritage to be predicted. According to the positions and sizes of the targets with known physical information presented in the captured images, obtain the corresponding perspective transformation matrix and pixel calibration values of the images. The perspective transformation method is as follows:
[0012]
[0013] In the formula, [X, Y, Z] T is the target point after transformation, [x, y, 1] T is the source point before transformation, both of which can be obtained according to the targets with known physical information, and M is the perspective transformation matrix.
[0014] The pixel calibration method is as follows:
[0015] W A = kω i
[0016] In the formula, W A is the true physical length; ω i is the number of pixels occupied by the target in the image; k is the conversion coefficient between the pixel length and the physical distance.
[0017] Second step: Identify the target plane where the damage is located and obtain the positions of the corner points of the target plane relative to the original image.
[0018] Third step: Send the image to be predicted into the Improved-YOLOv5n detection network for detection. After obtaining the predicted bounding boxes of the damage, perform cropping. Then, perform multi-threshold image segmentation based on the genetic algorithm on the cropped target damaged area. First, convert the image to a grayscale image. The grayscale range of the pixel points is from 0 to L - 1, the number of image pixels is N, and the number of pixels with the i-level grayscale value is N i , P i represents the probability that the pixels with the i-level grayscale value appear:
[0019]
[0020] In the formula, i represents the i-th level of grayscale, with a range of 0 to L - 1, and L - 1 is the maximum value of the grayscale.
[0021] Use the threshold T to segment the image into the target C0 (target) and C1 (background). ω 0 (T) and ω 1 (T) respectively represent the probabilities of C0 and C1 occurring when the threshold is T. ω 0 (T) and ω 1 (T) can be obtained through the following formula:
[0022]
[0023] ω 1 (T) = 1 - ω 0 (T)
[0024] The grayscale means of CO and C1 are μ 0 (T), μ 1 (T), and the grayscale mean of the whole image is μ, which can be obtained by the following formula:
[0025]
[0026] μ = ω 0 μ 0 + ω 1 μ 1
[0027] The between-class variance with threshold T is defined as follows:
[0028] σ 2 (T) = ω 0 (μ 0 - μ) 2 + ω 1 (μ 1 - μ) 2
[0029] Extend the single threshold to multiple thresholds, i.e., threshold = [T 1 , T 2 , …, T n . The cumulative between-class variance within each threshold interval is shown in the following formula:
[0030]
[0031] In the formula, n represents the number of thresholds; ω k represents the probability that the grayscale value is within the interval [T k-1 , T k .
[0032] At this time, the maximum between-class variance is defined as follows:
[0033]
[0034] When the between-class variance σ 2 (T 1 , T 2 , …, T n ) reaches the maximum value, the optimal threshold set for OTSU multi-threshold segmentation can be obtained This process uses a genetic algorithm to search for the optimal solution. Binary segmentation of the image is performed according to the optimal threshold set, and then opening and closing operations are further used to eliminate image noise and obtain clear damaged edges, so as to obtain the position of the damaged edge relative to the original image.
[0035] In the fourth step, the target corner points and damaged edge points are transformed and quantified by using the image pixel calibration values obtained in the first step and the perspective transformation matrix, so as to obtain the position of the damage on the plane where it is located and its true physical information such as length.
[0036] The beneficial effects of the present invention are as follows:
[0037] During the use of the present invention, first, in the damaged automatic recognition module part, a dataset of damaged diseases of masonry building heritage under a large view including the building heritage body and natural background is established; then, an Improved-YOLOv5n target detection network is proposed for the characteristics of large-view aerial images, which has higher detection accuracy, with mAP@0.5 reaching 0.855, and faster detection speed, with FPS reaching 824.3. Then, in the damaged automatic positioning and quantification module part, during the shooting process of daily inspection, by setting a target on the surface of the masonry building heritage to be photographed, the corresponding pixel calibration values and perspective transformation matrix of the image can be obtained; by identifying the target plane, the position of the corner points of the target plane relative to the original image can be obtained; the photographed image is detected by Improved-YOLOv5n to obtain the recognized image, and the main image including the target damage is obtained after cropping the target detection frame. The image is segmented by using a multi-threshold image segmentation algorithm based on the genetic algorithm, and binary segmentation is carried out under the guidance of the optimal threshold set, and then opening and closing operations are carried out to eliminate image noise, judge the clear damaged edge, and obtain the position of the damaged edge relative to the original image; finally, according to the pixel calibration values and the perspective transformation matrix, the corner points of the target plane and the position of the damaged edge relative to the original image are calibrated and transformed to obtain the position of the damage on the plane where it is located and its true physical information such as length. The damaged detection method of the present invention has the characteristics of supporting rapid detection, supporting secondary development, and supporting repeated utilization, adopts an automatic and information-based intelligent judgment method, and ensures the accuracy of damaged detection through computer vision methods. Description of the Drawings
[0038] Figure 1 It is a framework diagram of the automatic recognition and positioning system for damaged masonry building heritage based on computer vision proposed by the present invention. Detailed Embodiments
[0039] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0040] As shown in the figure, a method for automatically recognizing and positioning damaged masonry building heritage based on computer vision according to the present invention includes the following steps:
[0041] a) Set the damaged automatic recognition module 2 of the damaged automatic recognition and positioning system 1 of the masonry building heritage:
[0042] b) Set up the damage automatic positioning and quantification module 15 of the damaged automatic identification and positioning system 1 for masonry building heritage:
[0043] In the specific implementation process, the setting method of the damage automatic identification module 2 described in step a) includes:
[0044] First step, use the drone 3 to take pictures of the masonry building heritage (including the building body and the environment) 4 to obtain initial high-precision picture data, analyze the damaged disease characteristics of typical masonry building heritage, and use the deep learning annotation software Labelme to annotate the pictures to obtain the damaged disease dataset 6 of the masonry building heritage.
[0045] Second step, improve the YOLOv5n network 7 according to the characteristics of UAV images to obtain the Improved-YOLOv5n automatic identification network 12. Since the background occupies a large proportion in the UAV aerial pictures, while the target damaged disease area is small, and the number of effective channels of the target object is small in multi-channel detection, therefore, SEAttention 8 is introduced to assign different weights to different positions of the image from the perspective of channels, so as to pay attention to more important feature information. Usually, the background part occupies a large proportion in the UAV aerial pictures, which is likely to cause the problem of unbalanced classification of the target and the background during training. Therefore, the Focal Loss idea is introduced to make the model pay more attention to difficult-to-classify samples. Using Focal-CIoU (label 9) as the Bounding Box Regression Loss of the model has good results. Since the masonry building heritage, such as the Great Wall, city walls, bridges, ancient pagodas, building complexes, etc., often has a large volume, there are high requirements for the size, inference speed, and timeliness of the target detection model. By introducing the pruning method 10, the model size can be reduced without affecting the model performance, and the inference speed can be improved. After reducing the size of the model through pruning, it is still necessary to maintain the performance of the model. Therefore, the model distillation technology 11 is introduced to train a smaller target model by learning the output of the complex model, so that its performance is close to that of the complex model. Finally, the Improved-YOLOv5n target detection network 12 with a detection accuracy (mAP@0.5) reaching 85.5% and a detection speed (FPS) reaching 824.3 is obtained.
[0046] In the specific implementation process, the setting method of the damage automatic positioning and quantification module 15 described in step b) includes:
[0047] First step, set up a target 16 on the surface of the masonry building heritage to be predicted, and obtain the corresponding perspective transformation matrix 18 and pixel calibration value 19 of the image according to the position and size of the target 17 with known physical information presented on the captured image 13. The perspective transformation method is as follows:
[0048]
[0049] wherein, [X, Y, Z] T is the transformed target point, and [x, y, 1] T is the source point before transformation, both of which can be obtained from the target with known physical information, and M is the perspective transformation matrix.
[0050] The pixel calibration method is as follows:
[0051] W A = kω i
[0052] wherein, W A is the true physical length; ω i is the number of pixels occupied by the target in the image; k is the conversion coefficient between the pixel length and the physical distance.
[0053] In the second step, identify the target plane 20 where the damage is located, and obtain the position 21 of the corner points of the target plane relative to the original image.
[0054] In the third step, send the image to be predicted into the Improved-YOLOv5n detection network for detection 14. After obtaining the predicted bounding box 14 of the damage, perform cropping 22, and then perform multi-threshold image segmentation based on the genetic algorithm on the cropped target damaged area. First, convert the image into a grayscale image. The grayscale range of the pixel points is from 0 to L - 1, the number of image pixels is N, and the number of pixels with the i-level grayscale value is N i , P i represents the probability that the pixels with the i-level grayscale value appear:
[0055]
[0056] wherein, i represents the i-th level of grayscale, with a range of 0 to L - 1, and L - 1 is the maximum value of the grayscale.
[0057] Use the threshold T to segment the image into the target C0(target) and C1(background). ω 0 (T) and ω 1 (T) respectively represent the probabilities of C0 and C1 occurring when the threshold is T. ω 0 (T) and ω 1 (T) can be obtained through the following formula:
[0058]
[0059] ω 1 (T) = 1 - ω 0 (T)
[0060] The grayscale means of C0 and C1 are μ 0(T), μ 1 (T), the average gray value of the whole image is μ, which can be obtained by the following formula:
[0061]
[0062] μ = ω 0 μ 0 + ω 1 μ 1
[0063] The between-class variance with threshold T is defined as follows:
[0064] σ 2 (T) = ω 0 (μ 0 - μ) 2 + ω 1 (μ 1 - μ) 2
[0065] Extend the single threshold to multiple thresholds, i.e.: threshold = [T 1 , T 2 , …, T n . The cumulative between-class variance within each threshold interval is shown in the following formula:
[0066]
[0067] In the formula, n represents the number of thresholds; ω k represents the probability that the gray value is in the interval [T k-1 , T k .
[0068] At this time, the maximum between-class variance is defined as follows:
[0069]
[0070] When the between-class variance σ 2 (T 1 , T 2 , …, T n ) reaches the maximum value, the optimal threshold set for OTSU multi-threshold segmentation can be obtained This process uses a genetic algorithm to search for the optimal solution. According to the optimal threshold set, the image is binarized and segmented 24, and then opening and closing operations 25 are further used to eliminate image noise and obtain a clear damaged edge 26, and the position of the damaged edge relative to the original image can be obtained 27.
[0071] In the fourth step, the target corner points and damaged edge points are transformed and quantified using the image pixel calibration values and perspective transformation matrix obtained in the first step to obtain the position of the damage in the plane where it is located, as well as its length and other real physical information 28.
[0072] It should be noted that the above content only illustrates the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A computer vision-based method for identifying and locating damaged masonry building heritage, characterized in that: The following steps are involved: a) Set up an automatic identification module for damaged masonry architectural heritage; Step a1), forming a data set Use drones to photograph masonry architectural heritage to obtain initial high-precision image data, analyze the damage characteristics of typical masonry architectural heritage, and use deep learning annotation software Labelme to annotate the images to obtain a masonry architectural heritage damage data set; Step a2), improve the YOLOv5n network according to the characteristics of UAV images to obtain the Improved-YOLOv5n automatic recognition network Introducing SEAttention; Introduce the idea of FocalLoss; Introducing pruning methods; Introducing model distillation technology; b) Set up a module for automatic location and quantification of masonry building heritage damage Step b1), setting a target on the surface of the masonry building heritage to be predicted, and obtaining the corresponding perspective transformation matrix and pixel calibration value of the image according to the position and size of the target with known physical information on the captured image, Step b2), identifying the target plane where the damage is located, and obtaining the position of the corner point of the target plane relative to the original image; Step b3), the image to be predicted is sent to the Improved-YOLOv5n detection network for detection, and the damaged prediction frame is cropped after being obtained. Then, the cropped target defective area is subjected to multi-threshold image segmentation based on the genetic algorithm. First, the image is converted into a grayscale image, the grayscale range of the pixel is 0 to L-1, the number of image pixels is N, and the number of pixels of the i-level grayscale value is N i , P i Represents the probability of occurrence of a pixel with i-level gray value: In the formula, i represents the i-th grayscale, ranging from 0 to L-1, and L-1 is the maximum grayscale value; Use the threshold T to segment the image into targets C0 and C1. ω0(T) and ω1(T) represent the probability of C0 and C1 occurring when the threshold is T. ω0(T) and ω1(T) are obtained by the following formula: ω1(T)=1-ω0(T) The grayscale mean values of C0 and C1 are μ0(T) and μ1(T) respectively, and the grayscale mean value of the whole image is μ, which is obtained by the following formula: μ=ω0μ0+ω1μ1 The between-class variance with threshold T is defined as follows: s 2 (T)=ω0(μ0-μ) 2 +ω1(μ1-μ) 2 Expand the single threshold to multiple thresholds, that is, threshold = [T1, T2, ..., T n ], the cumulative inter-class variance within each threshold interval is as follows: Where n represents the number of thresholds; ω k Indicates that the gray value is between [T k-1 ,T k ] probability in the interval; At this time, the maximum between-class variance is defined as follows: When the between-class variance σ 2 (T1,T2,…,T n ) reaches the maximum value, the optimal threshold set of OTSU multi-threshold segmentation is obtained This process uses a genetic algorithm to search for the best result; the image is binarized according to the optimal threshold set, and then the image noise is further eliminated using opening and closing operations to obtain clear damaged edges and the position of the damaged edges relative to the original image; Step b4) uses the image pixel calibration values obtained in step b1) and the perspective transformation matrix to transform and quantize the target corner points and the damaged edge points to obtain the position of the damage on the plane and other real physical information.
2. According to the computer vision-based method for identifying and locating damaged masonry building heritage, the method is characterized in that: Step a2), improve the YOLOv5n network according to the characteristics of UAV images to obtain the Improved-YOLOv5n automatic recognition network, as follows: SEAttention is introduced to enable the model to assign different weight coefficients to different positions of the image from the perspective of the channel, thereby focusing on more important feature information; The introduction of the Focal Loss idea makes the model pay more attention to difficult-to-classify samples, and using Focal-CIoU as the Bounding Box Regression Loss of the model has a good effect; Introducing pruning methods to reduce model size and improve inference speed without affecting model performance; The model distillation technology is introduced to train a smaller target model by learning the output of a complex model, making its performance close to that of the complex model. Ultimately, the Improved-YOLOv5n target detection network is obtained with detection accuracy and speed that meet the standards.
3. The computer vision-based method for identifying and locating damaged masonry building heritage according to claim 1 is characterized in that: Step b1), setting a target on the surface of the masonry building heritage to be predicted, and obtaining the corresponding perspective transformation matrix and pixel calibration value of the image according to the position and size of the target with known physical information on the captured image, The perspective transformation method is as follows: Where [X,Y,Z] T is the target point after transformation, [x,y,1] T is the source point before transformation, which can be obtained according to the target with known physical information, and M is the perspective transformation matrix; The pixel calibration method is as follows: W A =kω i Where W A is the real physical length; ω i is the number of pixels occupied by the target in the image; k is the conversion coefficient between pixel length and physical distance.
Citation Information
Patent Citations
Patrol video image preprocessing method
CN114881869A
Container damage detection method based on machine vision and deep learning
CN115222697A
Cited By
A city wall targeted reinforcement method based on environment perception and disease identification
CN122757965A