Deep learning building earthquake damage recognition method and system combining multiple texture features
By performing multiple texture feature extraction and deep learning on SAR images, the low accuracy of single-polar SAR images in building earthquake damage recognition is solved, and high-precision building damage recognition is achieved, supporting post-disaster assessment and rescue.
Patent Information
- Application Number
- CN202410288367.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-03-14
AI Technical Summary
The prior art lacks a high-accuracy building earthquake damage recognition method for monopolar SAR images, and it is difficult to effectively judge the damaged area of buildings in earthquake disasters.
By obtaining synthetic aperture radar SAR images of earthquake-destructed areas, preprocessing, extracting the mean characteristics of the spatial domain, the median characteristics of the spatial domain and the Gabor characteristics of the frequency domain, combining deep convolutional neural networks for building seismic damage recognition, and using a variety of texture features for deep learning.
The accuracy of building earthquake damage identification and classification is achieved at 80.98%, which can provide data reference for the judgment of the damage degree of earthquake-damaged areas, support the formulation of rescue strategies and disaster loss assessment, and improve the accuracy and efficiency of identification.
Smart Images

Figure CN118038275B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of building earthquake damage feature recognition, and in particular to a deep learning building earthquake damage recognition method and system combining multiple texture features. Background Art
[0002] Major earthquake disasters often cause a large number of man-made features to be damaged and casualties. Therefore, it is extremely important to judge the disaster situation in the first time after the earthquake for disaster relief and rescue. Synthetic Aperture Radar (SAR) has the advantages of all-weather and all-day operation and is not easily affected by sunlight. Using SAR images for earthquake damage detection and loss estimation has gradually attracted widespread attention around the world. How to use the abstract information of SAR images to judge the damaged areas of buildings is an important research direction of SAR in the field of earthquake disasters.
[0003] However, the existing technology lacks a method for accurately identifying building earthquake damage in single-polarization SAR images, which is an urgent problem to be solved in the existing technology. Summary of the Invention
[0004] The present invention provides a deep learning building earthquake damage recognition method combining multiple texture features, which can overcome certain or some defects of the existing technology.
[0005] According to a deep learning building earthquake damage recognition method combining multiple texture features of the present invention, the method includes the following steps:
[0006] Step S1, obtaining a synthetic aperture radar (SAR) image of an earthquake-damaged area;
[0007] Step S2, preprocessing the original image of the SAR image to obtain a corrected single-polarization SAR image;
[0008] Subsequently, a sample set is selected from the single-polarization SAR image; the sample set includes images of each type of feature to be recognized and classified;
[0009] Step S3, through calculating the original SAR image processing, obtaining three preferred texture features applicable to building earthquake damage recognition, obtaining three preferred texture feature images, and adding the three preferred texture feature images corresponding to each sample in the sample set to obtain a multi-feature sample set;
[0010] The three preferred texture features include: the mean feature in the spatial domain, the median feature in the spatial domain, and the Gabor feature in the frequency domain;
[0011] Step S4, simultaneously superimposing the three preferred texture features and the original SAR image as multi-dimensional data inputs for deep convolutional neural network operations;
[0012] Step S5: Use the multi-feature sample set to train the deep convolutional neural network model to obtain a trained deep convolutional neural network model;
[0013] Step S6: Use the trained deep convolutional neural network model to perform ground object classification on the SAR image of the earthquake-damaged area.
[0014] Preferably, the ground object types in Step S2 include collapsed buildings, uncollapsed buildings, vegetation, and bare land.
[0015] Preferably, Step S2 specifically includes:
[0016] Step S21: Register the single-polarization SAR image with the optical image;
[0017] Step S22: On the registered single-polarization SAR image, extract the labeled ground object samples to form a SAR image sample set.
[0018] Preferably, in Step S22: Directly label the bare land and vegetation samples on the SAR image; Label the collapsed and uncollapsed building samples on the optical image.
[0019] Preferably, in Step S4, the formula for the operation is as follows:
[0020] Input(SAR∪Mean∪Median∪Gabor)→DCNN→Output(DB∪UDB)
[0021] In the formula, SAR, Mean, Median, and Gabor respectively represent the original SAR image, mean feature, median feature, and Gabor feature input into the convolutional neural network;
[0022] DCNN represents the deep convolutional neural network, and DB and UDB respectively represent the collapsed buildings and uncollapsed buildings in the classification results.
[0023] Preferably, in Step S3, the mathematical expression of the mean filtering corresponding to the mean feature is:
[0024]
[0025] In the formula, R is the gray value of the pixel after smoothing, N is the window smoothing size, and I ij is the initial gray value of each pixel in the smoothing window at (i,j);
[0026] The mathematical expression of the median filtering corresponding to the median feature is:
[0027] Y k =med{Xi+r,j+s ; r, s ∈ A}
[0028] Where A is the window for intercepting the image, X is the pixel value of the corresponding point. Y k represents the median value at point k, (i, j) is the upper left corner coordinate of the window, and r and s represent the number of pixels intercepted by the window.
[0029] The two-dimensional Gabor filter corresponding to the Gabor feature is defined as follows:
[0030]
[0031] Where (x, y) represents the two-dimensional spatial coordinate; u = xcosθ + ysinθ; y = -xsinθ + ycosθ; θ is the Gabor filter direction; σ x and σ y are the spatial scale factors for describing the frequency function; f is the frequency that determines the spatial scale factor, and usually σ x = σ y = 1 / f.
[0032] The present invention also provides a deep learning building earthquake damage recognition system that combines multiple texture features, including a building earthquake damage recognition device. The building earthquake damage recognition device includes a memory and a processor. The memory stores a computer program. It is characterized in that when the processor executes the computer program, the steps of the described method are realized.
[0033] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the described method are realized.
[0034] Compared with the prior art, the present invention has the following remarkable improvements:
[0035] (1) By calculating the original SAR image processing, the present invention obtains 3 preferred texture features applicable to building earthquake damage recognition, obtains 3 preferred texture feature images, and then combines them with the original SAR image to form a multi-dimensional data input for deep convolutional neural network operations;; Then, by adjusting the learning rate and the number of training times, the accuracy of the model is gradually improved, and finally a classification accuracy of 80.98% can be achieved.
[0036] (2) The relatively accurate extraction results of collapsed buildings and non-collapsed buildings obtained by the present invention through SAR image processing and analysis can, on the one hand, provide data reference for judging the damage degree of buildings in the earthquake-damaged area; on the other hand, it can also provide technical support for rescue strategy formulation, disaster loss assessment and subsequent reconstruction planning. Description of the Drawings
[0037] Figure 1Schematic diagram of the basic structure of the deep convolutional neural network in the present invention;
[0038] Figure 2 Flow chart of the classification and recognition of ground objects in the earthquake damage area in the present invention;
[0039] Figure 3 Single-polarization SAR image after calibration of the SAR image in the study area in the present invention;
[0040] Figure 4 is Figure 3 Marked samples of collapsed and non-collapsed buildings on the SAR image in;
[0041] Figure 5 Confusion matrix of the classification accuracy of the DCNN model after superimposing three preferred texture features in step (2) of Example 1;
[0042] Figure 6 Loss function of the DCNN after superimposing three preferred texture features in step (2) of Example 1;
[0043] Figure 7 Classification result map of the entire scene SAR image of the study area in step (3) of Example 1;
[0044] Figure 8 Classification result map of non-collapsed buildings in the sample area in step (3) of Example 1;
[0045] Figure 9 Classification result map of collapsed buildings in the sample area in step (3) of Example 1. Detailed implementation manners
[0046] Example 1
[0047] This example provides a deep learning building earthquake damage recognition method combining multiple texture features. In this example, the ground object types in the study area (earthquake damage area) can be roughly divided into four categories: collapsed buildings, non-collapsed buildings, vegetation, and bare land.
[0048] In order to improve the accuracy of ground object recognition and classification, this example selects and adopts:
[0049] On the basis of the original SAR image, preferred texture features in the spatial domain and frequency domain that are helpful for classification are added to the data input of the deep convolutional neural network, and texture features that are irrelevant or negatively correlated are excluded to improve the classification accuracy of the model.
[0050] This example will use a deep convolutional neural network (DCNN) to process the SAR image classification task.
[0051] The DCNN model will combine median features, mean features, and Gabor features to better capture the ground object texture information in SAR images.
[0052] As Figure 1 shown, there is a pooling layer after each convolutional layer. The first convolutional layer has 4 input channels (original image, Gabor, mean, median), 16 convolutional kernels with a size of 3×3, and the pooling layer has a size of 2×2; the second convolutional layer has 32 convolutional kernels with a size of 3×3, and the pooling layer has a size of 2×2; the third convolutional layer has 64 convolutional kernels with a size of 3×3, and the pooling layer has a size of 2×2. The model has a total of 112 convolutional kernels, and each convolutional kernel is a small learnable filter for extracting different features of the input image.
[0053] Specifically, in combination with Figure 2 the flowchart, the recognition and classification method in this embodiment mainly includes the following steps:
[0054] (1) Preprocessing of the image;
[0055] (2) Adding texture features in the spatial domain and frequency domain to the deep learning network classification model and training the deep learning network classification model;
[0056] (3) Using the trained deep learning network classification model for the earthquake-damaged area, classifying and extracting collapsed buildings.
[0057] For step (1): After a series of preprocessing on the original SAR image of the earthquake-damaged area obtained, a corrected single-polarization SAR image is obtained, as Figure 3 shown;
[0058] However, selecting a sample set for deep learning on the image is a very complex task. Compared with multi-spectral data or fully polarized SAR data, the single-polarization SAR image has only one polarization channel and cannot provide the auxiliary information required for visual interpretation through band operations, making the visual interpretation of SAR data extremely difficult.
[0059] In this embodiment, in order to obtain a more accurate sample area, we use the post-earthquake optical image of Google Earth for auxiliary annotation. By registering the SAR image with the optical image, the ground object types can be annotated on the optical image, as Figure 4 shown. On the registered SAR image, use remote sensing image processing software and corresponding programs to automatically extract the annotated sample areas to form a sample set.
[0060] We conducted tests on sample sets with various pixel sizes. The step size was set to half of the sample size. The individual sample sizes were 8*8, 16*16, 32*32, 64*64, and 128*128 respectively. After testing, it was found that smaller samples contained less information and had poorer classification effects. On the other hand, overly large sample sizes would result in excessive redundant information, increasing the learning burden of the model and also reducing the number of selectable regions. Eventually, a sample size of 64*64 was chosen, and this size had a higher accuracy than other sample sizes when classifying using only the original image. Nevertheless, the best accuracy when classifying using only the original SAR image was only 47.84%, and it was unable to accurately identify the types of ground objects in the earthquake-damaged areas. The vegetation areas were not identified as a whole, as shown in Table 1.
[0061] Since there are inevitably some processing errors after the image is corrected and registered. Therefore, in the optical image, only the areas of collapsed and non-collapsed buildings that are difficult to distinguish are marked, while the bare land and vegetation areas are directly marked on the SAR image. The marked samples of collapsed buildings and non-collapsed buildings are as shown in the SAR image Figure 4 shown. The experimental sample set was divided into four categories: collapsed buildings, non-collapsed buildings, vegetation, and bare land.
[0062] For step (2):
[0063] SAR images usually contain complex texture information, which is very important for image classification tasks. Therefore, it is necessary to extract these texture features from SAR images. In this embodiment, not only a deep convolutional neural network is used to improve the SAR image classification performance, but also three important texture feature extraction methods are combined to capture the texture information in the image. The specific process is as follows:
[0064] Add the three preferred texture feature images corresponding to each sample in the sample set to obtain a multi-feature sample set.
[0065] Overlay the feature image processed by spatial domain and frequency domain filtering with the original SAR image as the multi-dimensional data input of the DCNN, and then perform deep convolutional neural network operations.
[0066] In this embodiment, the three preferred texture features added include the mean feature in the spatial domain, the median feature in the spatial domain, and the Gabor feature in the frequency domain.
[0067] In this embodiment, the mean feature, median feature, and Gabor feature can significantly improve the accuracy of ground object recognition, as shown in Table 1.
[0068] Table 1 Classification accuracy using different texture features
[0069]
[0070] Specifically, the classification accuracy of adding median feature data input to the DCNN model reached 72.38%, the classification accuracy of adding mean feature data input to the DCNN model was 73.30%, and the classification accuracy of adding Gabor feature data input to the DCNN model was 73.42%. Therefore, these three preferred texture features are jointly used as the multi-dimensional data input of the DCNN model.
[0071] For the mean feature, the mean feature can represent the average level of the data set and is obtained by calculating the mean of the pixel values in the local window of the image. The mean feature is usually used to describe the overall brightness and uniformity of the image and helps to capture the texture information of the ground surface. The mean feature is often used to describe the central tendency of the data set to understand the average level of the data. In image processing, the mean filter is used to smooth the image and reduce noise.
[0072] Mean filtering calculates the average of the gray values of all pixels within the smoothing window and then assigns it to the central pixel of the smoothing window. Its mathematical expression is
[0073]
[0074] where R is the gray value of the pixel after smoothing, N is the size of the window for smoothing, and I ij is the initial gray value of each pixel within the smoothing window at (i, j).
[0075] For the median feature, the median feature focuses on the median in the data set or signal. The median is the value at the middle position after arranging the data in ascending order. The median feature is extracted by calculating the median of the pixel values in the local window of the image. This method helps to capture the brightness distribution and contrast features of the image. Especially in SAR images, it can help to distinguish different ground object types, such as buildings and natural features.
[0076] The median feature is usually used to describe the central tendency of the data set. For data sets with a large influence from outliers, the median is more robust than the mean. In image processing, the median filter is used to remove impulse noise or salt-and-pepper noise in the image. The mean feature provides the average level of the data, while the median feature provides the central value of the data, which is more robust for data sets disturbed by outliers.
[0077] Median filtering is to sort the values in the window, and the sequence sample value at the middle position is the output of the median filtering, which is expressed by the formula:
[0078] Y k =med{X i+r,j+s ;r,s∈A}
[0079] where A is the window for intercepting the image, and X is the pixel value at the corresponding point. Yk Denote the median at point k, where (i, j) is the upper left coordinate of the window, and r and s represent the number of pixels intercepted by the window.
[0080] Gabor features can effectively capture the texture structure in images, including directional and frequency information. In SAR image classification, Gabor features perform excellently in identifying ground object textures.
[0081] Gabor first proposed a two-dimensional Gabor filtering method suitable for texture feature extraction and recognition in 1946. Due to its high sensitivity to direction and scale, it can identify texture feature information with specific direction frequencies from images. The two-dimensional Gabor filter is defined as follows:
[0082]
[0083] In the formula, (x, y) represents two-dimensional spatial coordinates; u = xcosθ + ysinθ; y = -xsinθ + ycosθ; θ is the direction of the Gabor filter; σ x and σ y are spatial scale factors describing the frequency function; f is the frequency determining the spatial scale factor, and usually σ x = σ y = 1 / f.
[0084] In this embodiment, combining these texture feature extraction methods with DCNN can provide multi-level and multi-scale information for the classification model, which helps to more accurately distinguish different ground object types.
[0085] Specifically, in this embodiment, the selected 3 preferred texture features and the original SAR image are simultaneously superimposed into multi-dimensional data for input, and deep convolutional neural network operations are performed. By adjusting the learning rate and the number of training times, the accuracy of the model is gradually improved. Finally, a classification accuracy of 80.98% is achieved. The confusion matrix of the classification accuracy of the DCNN model is as Figure 5 shown. The loss function of the training is shown in Figure 6 . The classification accuracy is shown in Table 1.
[0086] During the training process, the number of samples in each class set is between 3000 and 5000, maintaining relative balance. Randomly select 30% of the samples from each class as the validation set, and the remaining 70% as the training set. Use the training set to train the model, and then use the trained model to classify and test the validation set, so as to judge the accuracy of building earthquake damage recognition of each model.
[0087] For step (3):
[0088] Use the trained deep convolutional neural network model to classify the SAR data of the entire study area.
[0089] The classification results are as follows Figure 7 shown, where the red area represents the collapsed building area, the yellow area represents the non-collapsed building area, the gray area represents the bare land, and the green area represents the vegetation area. Through this step, the comprehensive identification and classification of various types of ground objects in the earthquake-damaged area are completed.
[0090] Next, the area where the samples were selected is used to mask the classification map, as shown in Figure 8 、 Figure 9 shown, where the classification accuracy of the non-collapsed building area reaches 79.32%, and the classification accuracy of the collapsed building area also reaches 77.43%.
[0091] It is easy to understand that those skilled in the art can combine, split, reorganize, etc. the embodiments of the present application based on one or several embodiments provided by the present application to obtain other embodiments, and these embodiments do not exceed the protection scope of the present application.
[0092] The present invention and its implementation manners are schematically described above. This description is not restrictive. What is shown in the embodiments is only part of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative work without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A deep learning method for identifying building earthquake damage by combining multiple texture features, characterized in that It includes the following steps: The method includes the following steps: Step S1: Obtain a synthetic aperture radar (SAR) image of the earthquake-damaged area; Step S2: Preprocess the original image of the SAR image to obtain a corrected single-polarization SAR image; Subsequently, select a sample set from the single-polarization SAR image; the sample set includes images of each type of ground object to be identified and classified; the types of ground objects in Step S2 include collapsed buildings, non-collapsed buildings, vegetation, and bare land; Step S2 specifically includes: Step S21: Register the single-polarization SAR image with an optical image; Step S22: On the registered single-polarization SAR image, extract the labeled ground object samples to form a SAR image sample set; In Step S22: Directly label the bare land and vegetation samples on the SAR image; label the collapsed and non-collapsed building samples on the optical image; Step S3: Through calculations on the original SAR image processing, obtain three texture features for building earthquake damage identification, obtain three texture feature images, and add the three texture feature images corresponding to each sample in the sample set to obtain a multi-feature sample set; The three texture features include: the mean feature in the spatial domain, the median feature in the spatial domain, and the Gabor feature in the frequency domain; In Step S3, the mathematical expression of the mean filter corresponding to the mean feature is: where R is the smoothed pixel gray value, N is the window smoothing size, and I ij is the initial gray value of each pixel within the smoothing window at (i, j); The mathematical expression of the median filter corresponding to the median feature is: Y k = med{X i+r,j+s ; r, s ∈ A} Where A is the window for intercepting the image, X is the pixel value of the corresponding point; Y k represents the median value at the k-th point, (i, j) is the coordinate of the upper left corner of the window, and r and s represent the number of pixels intercepted by the window; The definition of the two-dimensional Gabor filter corresponding to the Gabor feature is as follows: Where (x, y) represents two-dimensional spatial coordinates; u = x cosθ + y sinθ; v = -x sinθ + y cosθ; θ is the direction of the Gabor filter; σ x and σ y are spatial scale factors for describing the frequency function; f is the frequency that determines the spatial scale factor, σ x = σ y = 1 / f; Step S4: Simultaneously stack the three texture features and the original SAR image into a multi-dimensional data input for deep convolutional neural network operations; Step S5: Use the multi-feature sample set to train the deep convolutional neural network model to obtain a trained deep convolutional neural network model; Step S6: Use the trained deep convolutional neural network model to classify the ground objects in the SAR image of the earthquake-damaged area; In Step S4, the formula for the operation is as follows: Input(SAR∪Mean∪Median∪Gabor)→DCNN→Output(DB∪UDB) In the formula, SAR, Mean, Median, and Gabor respectively represent the original SAR image, mean feature, median feature, and Gabor feature input into the deep convolutional neural network; DCNN represents the deep convolutional neural network, and DB and UDB respectively represent the collapsed buildings and non-collapsed buildings in the classification results.
2. A deep learning-based building earthquake damage recognition system that combines multiple texture features, including a building earthquake damage recognition device. The building earthquake damage recognition device includes a memory and a processor, and a computer program is stored in the memory. It is characterized in that: When the processor executes the computer program, it implements the steps of the method described in Claim 1.
3. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, it implements the steps of the method described in Claim 1.