Earthquake damaged building prediction method and related equipment

By combining ensemble learning algorithms and the SHAP framework, a building damage prediction model was constructed, which solved the problem of rapid and accurate identification and prediction of building damage after an earthquake, and achieved efficient emergency response and in-depth mechanistic analysis.

CN121598221APending Publication Date: 2026-03-03CHANGAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies cannot quickly identify and predict building damage after an earthquake, and lack understanding of the relationship between ground motion intensity, terrain features, and building properties, resulting in long assessment cycles, low accuracy, and a lack of quantitative analysis methods.

Method used

An ensemble learning algorithm is used to construct a building damage prediction model. The model is trained using multi-source influencing factor data and analyzed using the SHAP framework to quantify the contribution of each factor to the prediction results, thereby achieving fast and accurate damage prediction.

Benefits of technology

It enables rapid and automatic identification and high-precision prediction of post-earthquake building damage, improves emergency response efficiency, provides an interpretable understanding of damage mechanisms, and overcomes the timeliness and accuracy deficiencies of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598221A_ABST
    Figure CN121598221A_ABST
Patent Text Reader

Abstract

The invention discloses an earthquake damaged building prediction method and related equipment, and aims to solve the technical problems that the building damage condition cannot be quickly identified and predicted after an earthquake and damage mechanism analysis is lacked in the prior art. The method comprises the following steps: constructing a building damage prediction model by using a gradient lifting algorithm, and realizing rapid prediction by inputting near-real-time multi-source influence factor data which can be easily acquired after an earthquake; and meanwhile, an SHAP framework is adopted to carry out interpretability analysis on the model, the contribution degree of each influence factor is quantified, and a building damage mechanism is disclosed. According to the method, the problems of low efficiency, insufficient prediction precision and lack of mechanism analysis of a traditional evaluation method are effectively solved, and reliable technical support is provided for earthquake emergency response and post-disaster rescue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of disaster prevention and mitigation and earthquake engineering technology, specifically to a method for predicting earthquake-damaged buildings. Background Technology

[0002] In actual earthquake disaster emergency response and urban earthquake-resistant disaster prevention projects, especially during the golden rescue window after a strong earthquake, the rapid identification and prediction of post-earthquake damage to buildings is a key link in formulating rescue plans and resource allocation plans.

[0003] In terms of damage identification, current methods mainly rely on two types: on-site investigation and optical remote sensing interpretation. On-site investigation is inefficient and time-consuming, making it difficult to meet the needs of rapid post-earthquake response; optical remote sensing interpretation is easily affected by cloud cover and lighting conditions, making it difficult to quickly obtain high-quality image data in some seismic areas. Regarding damage prediction, existing empirical regression or statistical models often rely on only one or a few influencing factors, failing to fully consider the complex nonlinear relationships between seismic motion characteristics, topographic conditions, and building attributes, resulting in limited prediction accuracy.

[0004] This reveals several problems in the current prediction of earthquake-damaged buildings: relying on manual or remote sensing imagery for post-earthquake damage assessment is relatively time-consuming and cannot quickly obtain assessment results of damaged buildings to support emergency rescue and resource allocation; there is a lack of understanding of the relationship between ground motion intensity, terrain features, and building attributes and building damage; and there is a lack of a building damage prediction system that integrates multi-source earthquake and environmental data.

[0005] To address the aforementioned issues, there is an urgent need for a method for predicting earthquake-damaged buildings that can integrate seismic motion parameters, topographic factors, and building attributes, and possess high-precision prediction and interpretable analysis capabilities. Summary of the Invention

[0006] The purpose of this invention is to provide a method for predicting earthquake-damaged buildings, so as to overcome the technical problems of existing technologies being unable to quickly identify and predict building damage after an earthquake and lacking damage mechanism analysis.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for predicting earthquake-damaged buildings, comprising: Acquire near-real-time multi-source impact factor data for the target area after the earthquake, input it into the building damage prediction model, and output the prediction results for the target area; the building damage prediction model is established through the following steps: Obtain a multi-source impact factor dataset as training data; Training data is input into an ensemble learning model for training, and a building damage prediction model that reflects the mapping relationship between multi-source influencing factors and building damage density is established. Based on the SHAP framework, the process of outputting prediction results through the building damage prediction model is analyzed, the contribution of multi-source impact factor data to the prediction results is quantified, and the building damage mechanism under real-time multi-source impact factor data is obtained.

[0008] Multi-source impact factor data include peak ground velocity, peak ground acceleration, digital elevation model, epicentral distance, fault distance, lithology, building density, and building height.

[0009] After obtaining the multi-source impact factor dataset, the process also includes data preprocessing: Unify all raster data in the acquired multi-source impact factor dataset to the same spatial reference frame and resolution; Standardize the continuous variables in the multi-source impact factor dataset; One-hot encoding is used to numerically represent the categorical variables in the multi-source influence factor dataset.

[0010] Training data is input into the ensemble learning model for training, including: The hyperparameters of the ensemble learning model are optimized by combining random grid search with cross-validation. Hyperparameters include at least one of the following: maximum tree depth, number of base learners, learning rate, subsampling rate, and feature column sampling rate.

[0011] After establishing the building damage prediction model, the process also includes performance evaluation of the building damage prediction model: The performance of the building damage prediction model is evaluated using at least one of the following indicators: mean square error, root mean square error, mean absolute error, and coefficient of determination.

[0012] Analysis of building damage prediction models based on the SHAP framework includes: An interpreter is built using the training data as background data, and the SHAP value of the test set samples is calculated. The importance of each feature value is obtained by sorting the average absolute value of the SHAP values ​​and then integrating them to form the global feature importance. By using summary point plots to analyze the importance of global features, the directional relationship between feature values ​​and prediction contributions can be obtained.

[0013] After outputting the prediction results for the target region, the following is also included: The prediction results are combined with geographic coordinate information to generate a spatial distribution map of building damage density; Identify building damage hotspots based on spatial distribution maps of building damage density.

[0014] Secondly, the present invention provides a system for predicting earthquake-damaged buildings, comprising: The prediction module is used to acquire near-real-time multi-source impact factor data of the target area after an earthquake, input it into the building damage prediction model, and output the prediction results for the target area. The building damage prediction model is established through the following steps: Obtain a multi-source impact factor dataset as training data; Training data is input into an ensemble learning model for training, and a building damage prediction model that reflects the mapping relationship between multi-source influencing factors and building damage density is established. The damage mechanism analysis module is used to analyze the process of outputting prediction results through the building damage prediction model based on the SHAP framework, quantify the contribution of multi-source impact factor data to the prediction results, and obtain the building damage mechanism under real-time multi-source impact factor data.

[0015] Thirdly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the earthquake-damaged building prediction method described above.

[0016] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the earthquake-damaged building prediction method described above.

[0017] Compared with the prior art, the present invention has the following beneficial technical effects: Firstly, this invention provides a method for predicting earthquake-damaged buildings. By constructing a building damage prediction model based on an ensemble learning machine learning algorithm, it effectively overcomes the limitations of existing technologies. In terms of rapid identification, this method directly utilizes readily available near-real-time multi-source influencing factor data after the earthquake to drive the trained model for prediction, avoiding the time bottleneck of relying on manual surveys or optical remote sensing, thus achieving a substitute for assessment efficiency. Regarding prediction accuracy, the ensemble learning algorithm possesses strong nonlinear fitting capabilities, accurately learning the intrinsic mapping relationship between the complex features composed of seismic motion parameters, topographic factors, and building attributes and the building damage density, solving the prediction bias caused by the simplification problem in traditional models. In terms of mechanism analysis, the integrated SHAP interpretability analysis framework can quantitatively reveal the contribution of each influencing factor to the prediction results, transforming the model from a "black box" into an interpretable and verifiable decision-making system, thereby providing a data-driven understanding of damage mechanisms lacking in traditional methods.

[0018] Secondly, this invention provides an earthquake-damaged building prediction system. Through the collaborative operation of a prediction module and a damage mechanism analysis module, it effectively solves the problems faced by existing technologies. The prediction module utilizes a pre-trained ensemble learning model to process near-real-time multi-source influencing factor data that can be easily obtained after an earthquake, achieving rapid and automatic identification and prediction of building damage, significantly improving emergency response efficiency. The ensemble learning model can accurately capture the complex nonlinear relationship between multi-source influencing factors and building damage density, ensuring the accuracy of the prediction results. The damage mechanism analysis module, based on the SHAP framework, quantitatively analyzes the model's decision-making process, transforming the traditional black-box model into an interpretable system. This not only enhances the credibility of the results but also provides data support for a deeper understanding of earthquake damage mechanisms, solving the problem of the lack of quantitative analysis methods in existing technologies.

[0019] Thirdly, the present invention provides a computer device that, through a processor executing a specific computer program, can efficiently implement the steps of the method of the present invention. When performing data processing tasks, the computer device can accurately perform numerical calculations and logical judgments, avoiding errors caused by human factors. At the same time, since the computer program has high stability and reliability, it can ensure the accuracy and consistency of the data processing results.

[0020] Fourthly, the present invention provides a computer-readable storage medium. By programming the steps of the method of the present invention into a computer program and storing it on the computer-readable storage medium, users can easily load these programs onto any compatible computer device and execute them without rewriting or converting the code, which greatly improves the convenience and flexibility of program execution. Attached Figure Description

[0021] Figure 1 This is a flowchart of a method for predicting earthquake-damaged buildings in an embodiment of the present invention.

[0022] Figure 2 The following is a detailed map of the study area in this embodiment of the invention; wherein, (a) is the overall tectonic background of the study area; (b) is detailed information about the study area; (c) is a map of the seismic intensity distribution of the study area; (d) is a photograph of the earthquake site in the Besni region; and (e) is a photograph of the Kahramanmara region. (f) is a photo of the earthquake in the Hatay region.

[0023] Figure 3This is a spatial distribution map of the building damage influencing factors in the study area in this embodiment of the invention; where (a) is peak ground acceleration (PGA), (b) is peak ground velocity (PGV), (c) is the Copernicus digital elevation model, (d) is the distance from each point in the study area to the epicenter (e) is the distance from each point in the study area to the fault (f) is lithological data, (g) is building density, and (h) is building height.

[0024] Figure 4 This is a map showing the distribution of damaged buildings in the study area obtained by remote sensing in this embodiment of the invention; wherein, (a) is a map showing the building damage density distribution used for predictive model training and effect evaluation. (b) to (h) are building damage proxy maps (BDPM) of high building damage density areas at different locations in (a) used for comparative analysis of prediction results.

[0025] Figure 5 This is a building damage density map in an embodiment of the present invention; wherein, (a) is the predicted result of building damage density in the target area, and (b) to (g) are detailed maps of building damage density in different representative cities.

[0026] Figure 6 The following is a comparison and evaluation chart of the prediction results in the embodiments of the present invention; wherein, (a) is a spatial residual chart of the predicted value and the actual building damage density, (b) is a scatter plot of the predicted density and the actual density of the target area, (c) is a residual histogram and density curve, and (d) is a spatial hotspot comparison chart based on the Getis-OrdGi* statistic.

[0027] Figure 7 The diagrams shown are quantitative analysis diagrams of nonlinear effects in this embodiment of the invention; where (a) is a scatter plot of predicted building damage density and actual density on the test set, (b) is a residual histogram, (c) is a feature importance diagram of the ensemble learning model based on SHAP values, and (d) is a SHAP summary diagram of the ensemble learning model.

[0028] Figure 8 This is a schematic diagram of an earthquake-damaged building prediction system according to an embodiment of the present invention. Detailed Implementation

[0029] The key to earthquake disaster emergency response lies in the ability to quickly and accurately assess building damage after a strong earthquake. However, existing technologies have significant bottlenecks: at the identification level, reliance on inefficient manual surveys or weather-sensitive optical remote sensing makes it difficult to meet the timeliness requirements of the golden rescue period; at the prediction level, traditional models fail to fully consider the complex nonlinear relationships between ground motion parameters, topographic factors, and building attributes, resulting in limited prediction accuracy and a lack of quantitative analysis methods to reveal the underlying coupling mechanisms. These factors collectively lead to three core problems: excessively long assessment cycles, insufficient understanding of the underlying mechanisms, and a lack of reliable prediction systems. Therefore, a new method that can integrate multi-source data and possess both high-precision prediction and interpretable analysis capabilities is urgently needed.

[0030] Based on the above background, this invention proposes a method and related equipment for predicting earthquake-damaged buildings. The aim is to achieve rapid identification and accurate prediction of post-earthquake building damage by constructing a building damage prediction model based on ensemble learning. Utilizing readily available near-real-time multi-source influencing factor data after an earthquake, time-consuming manual investigations and remote sensing image interpretation are avoided, improving assessment efficiency. The ensemble learning model can effectively learn the complex nonlinear relationships between seismic motion parameters, topographic factors, and building attributes, overcoming the shortcomings of insufficient prediction accuracy in traditional models. Furthermore, the trained model is analyzed using the SHAP framework to quantify the contribution of each influencing factor, transforming the model's black-box decision-making process into an interpretable damage mechanism, thus addressing the lack of quantitative mechanistic analysis methods in existing technologies.

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] To facilitate understanding of the embodiments of the present invention, several key terms involved in this application will be defined below. Unless otherwise specified, these terms have the following meanings in this application: Peak Ground Velocity (PGV); Peak ground acceleration (PGA); Digital Elevation Model (DEM); Distance to Epicenter (DE); Distance to Fault (DF); Lithology; Copernicus Digital Elevation Model (COP-DEM); Building density (BD); Building height; United States Geological Survey (USGS) Global Building Atlas (GBA) Global Lithological Map (GLiM); Building Damage Proxy Map (BDPM); One-Hot Encoding; Gradient Boosting Decision Tree (GBDT); Randomized Search Cross-Validation (RandomizedSearchCV); Mean Squared Error (MSE); Root Mean Squared Error (RMSE); Mean Absolute Error (MAE); Coefficient of determination (R²); SHAP (SHapley Additive exPlanations); Hotspot analysis (Getis-Ord Gi*).

[0033] Reference Figure 1 The image shows a specific embodiment of the earthquake-damaged building prediction method provided by the present invention, comprising: Near-real-time multi-source impact factor data of the target area after the earthquake is acquired and input into the building damage prediction model, which then outputs the prediction results for the target area. The building damage prediction model is established through the following steps: Obtain a multi-source impact factor dataset as training data; Training data is input into an ensemble learning model for training, and a building damage prediction model that reflects the mapping relationship between multi-source influencing factors and building damage density is established. Based on the SHAP framework, the process of outputting prediction results through the building damage prediction model is analyzed, the contribution of multi-source impact factor data to the prediction results is quantified, and the building damage mechanism under real-time multi-source impact factor data is obtained.

[0034] In this specific implementation, a pre-trained building damage prediction model is constructed and applied. This model can automatically output prediction results for the target area immediately after an earthquake, based on readily available near-real-time multi-source influencing factor data. This design cleverly avoids the necessity of waiting for high-resolution remote sensing images or conducting time-consuming on-site investigations after an earthquake, thus gaining valuable time for a rapid post-earthquake emergency response.

[0035] The process of establishing the building damage prediction model ensures the comprehensiveness of the model's learning data by acquiring historical datasets covering multiple dimensions such as earthquake motion, topography, and building attributes as training data. The training data is then input into the ensemble learning model for training. This machine learning algorithm efficiently learns a deep, non-linear mapping relationship between complex multi-source influencing factors and building damage density. The final model is essentially a computational engine capable of accurately simulating the complex mechanisms of earthquake-induced damage.

[0036] Furthermore, the trained building damage prediction model is analyzed based on the SHAP framework to enhance technical transparency and credibility. This step is not a simple application of the model, but rather an exploration of its internal decision-making logic. By quantifying the contribution of each multi-source influencing factor to the final prediction result, the model's "black box" decision-making process is transformed into understandable and verifiable quantitative indicators. The resulting building damage mechanism not only verifies the consistency between the model's prediction patterns and physical reality and engineering experience, but also provides a solid theoretical basis for optimizing seismic design of buildings and guiding seismic risk zoning.

[0037] The three components mentioned above support each other and work closely together: high-quality training data is a prerequisite for model accuracy, ensemble learning algorithms are the core tool for achieving high-precision nonlinear mapping, and SHAP interpretability analysis ensures the scientific rigor and credibility of the model's decision-making process. Together, these three elements enable this technical solution to ultimately achieve the goal of rapid, accurate, and mechanistically sound prediction of earthquake-damaged buildings.

[0038] A specific embodiment of the present invention also provides a method for predicting earthquake-damaged buildings, comprising the following steps: S1, Data Collection and Preprocessing. This includes acquiring a multi-source impact factor dataset, which contains ground motion parameters, topographic factors, and building attribute data; and performing spatial registration and standardization on the multi-source impact factor dataset to form standardized training data. This step solves the technical challenges of multi-source data fusion and standardization.

[0039] In this specific implementation, a typical earthquake-affected area is selected as the study area. Existing multi-source impact factor data within the study area are directly used as prior information to train the building damage prediction model. The building damage prediction model is used to quickly carry out post-earthquake building damage prediction analysis, rather than relying on remote sensing images obtained after the earthquake, thus breaking through the timeliness limitations of traditional assessment methods.

[0040] Specifically, the selected influencing factors include: peak ground velocity (PGV), peak ground acceleration (PGA), digital elevation model (DEM), epicentral distance (DE), fault distance (DF), lithology, building density (BD), and building height.

[0041] The PGV and PGA data were obtained from the U.S. Geological Survey (USGS) website; DE and DF were obtained through spatial buffer analysis based on epicenter location and fault data provided by the USGS, with a buffer distance interval set at 15 m. Building density was obtained through spatial point density analysis of building distribution data in the Global Building Atlas (GBA); building height data were directly obtained from the GBA dataset, in meters.

[0042] In addition, it also includes collecting historical building damage results within the study area and using point density analysis to generate a density map of damaged buildings, which is used as training samples for the damage prediction model, providing reliable data support for subsequent model construction and validation.

[0043] S2, Feature Construction and Model Training. By inputting the standardized training data into an ensemble learning-based machine learning model, it can capture complex nonlinear relationships and feature interaction effects from the multi-source influencing factors, establishing a nonlinear mapping relationship from the multi-source influencing factors to building damage density. This solves the technical problem of traditional models failing to capture complex nonlinear relationships, resulting in low prediction accuracy.

[0044] Specifically, this includes feature construction and spatial matching of various influencing factors after completing the collection and preprocessing of multi-source data.

[0045] First, all acquired raster data are standardized to the same spatial reference frame and resolution to ensure consistency in spatial scale across different data types. In this specific implementation, the spatial reference frame is WGS84, and the resolution is 15m.

[0046] Subsequently, continuous variables, i.e., influencing factors other than lithology, were standardized to eliminate the impact of dimensional differences on model training; categorical variables, i.e., lithology, were numerically represented using one-hot encoding, as follows: 1: Water body, 2: Unconsolidated sediments, 3: Mixed sedimentary rocks, 4: Basic volcanic rocks, 5: Carbonate sedimentary rocks, 6: Intermediate volcanic rocks, 7: Intermediate plutonic rocks, 8: Evaporites, 9: Acidic plutonic rocks, 10: Basic plutonic rocks, 11: Metamorphic rocks, 12: Siliceous clastic sedimentary rocks, 13: Volcanic clastic rocks.

[0047] After feature construction was completed, 6 million building points were randomly selected within the study area as training samples for the building damage prediction model. The building points included 3 million points in the damaged building area and 3 million points in the non-damaged building area to ensure the representativeness of the samples.

[0048] Using the Multivalue Extract to Point tool in ArcGIS software, the raster values ​​of eight influencing factors and building damage density at the locations of 6 million points were extracted into a point attribute table. The field names are PGA, PGV, DEM, DE, Lithology, DF, BD, Height, and Damage, corresponding to Peak Surface Velocity (PGV), Peak Surface Acceleration (PGA), Digital Elevation Model (DEM), Epicentral Distance (DE), Fault Distance (DF), Lithology, Building Density (BD), and Building Height (Height), respectively. Next, the attribute table was converted into a .csv file for input into the building damage prediction model.

[0049] This specific implementation uses an ensemble learning algorithm based on a gradient boosting framework to establish a building damage prediction model. As an efficient ensemble learning model based on gradient boosting trees (GBDT), it can effectively handle nonlinear relationships and improve prediction accuracy.

[0050] The density of damaged buildings was used as the dependent variable, and each influencing factor was used as an independent variable. The data was divided into training data and test data in a ratio of 8:2.

[0051] To improve the performance of the building damage prediction model, this specific implementation combines Randomized SearchCV (Randomized SearchCV) and 5-fold cross-validation to optimize the main hyperparameters of the building damage prediction model. Specifically, the maximum tree depth (max_depth) is set between 4 and 10, with an interval of 2; the number of base learners (n_estimators) is set between 100 and 300, with an interval of 50; the learning rate (learning_rate) is set between 0.05 and 0.2, with an interval of 0.05; and the subsample rate and feature column sampling rate (colsample_bytree) are set between 0.6 and 1.0, with an interval of 0.2.

[0052] This specific implementation ensures computational accuracy and model stability by dividing the training and test sets and using a unified data type, resulting in the optimized building damage prediction model exhibiting good fitting performance on the training data. Furthermore, the performance of the building damage prediction model is evaluated using metrics such as mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²), guaranteeing the reliability and robustness of the prediction results and providing a reliable foundation for accurate prediction of building damage density.

[0053] S3, Damaged Building Density Prediction. After an earthquake, using the earthquake-stricken area as the target region, near-real-time multi-source impact factor data for the target region is acquired. This preprocessed data is then input into a trained building damage prediction model, outputting a predicted building damage density map of the target region as the prediction result. This addresses the technical bottleneck of low efficiency and inability to meet the critical rescue window during post-earthquake manual assessments.

[0054] After the building damage model is trained and the optimal gradient boosting model is obtained, the building damage model can be applied to the target area where an earthquake occurs. The dataset collected in the target area is called the target dataset. The building damage prediction model can predict the building damage density based on the target dataset.

[0055] The collected target dataset contains 21 million points of damaged buildings to be predicted. Each point's attributes include latitude and longitude information and eight damage impact factors (PGA, PGV, DEM, DE, Lithology, DF, BD, Height). To ensure the consistency between the prediction results and the training results, the data in the target dataset were preprocessed, including replacing or deleting missing values ​​and outliers (-9999, infinity), and ensuring that the feature column order was consistent with the training data.

[0056] After processing, the preprocessed target dataset is input into the building damage prediction model for prediction. The prediction results are then merged with the original data and output to generate a new data file containing latitude and longitude coordinates and predicted damage density. The predicted damage density is compared with the actual damage density values, and goodness-of-fit indices such as mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²) are calculated to verify the prediction performance.

[0057] Furthermore, this specific embodiment also includes a visualization analysis of the prediction results, including outputting scatter plots of the actual damage values ​​and predicted values, as well as residual histograms, to intuitively display the model's prediction accuracy and error distribution. The results show that the building damage prediction model can reproduce the spatial distribution of building damage well, and the residuals are mainly concentrated near zero, indicating that the prediction bias is small and the prediction results are reliable.

[0058] S4, Importance Analysis of Influencing Factors. Based on the SHAP framework, the trained building damage prediction model is analyzed to quantify the contribution of each influencing factor to the prediction results, thereby revealing the mechanism of building damage. This addresses the technical problem that existing methods lack mechanistic understanding and that existing prediction models are "black boxes."

[0059] To further explain the output of the building damage prediction model and reveal the relative importance of various factors affecting earthquake-damaged buildings, this specific implementation method applies the SHAP framework to analyze the trained building damage prediction model.

[0060] An interpreter was constructed using a set of 4.8 million points as background data, and the SHAP framework analysis results for 1.2 million test point samples were calculated to obtain global and local contributions. The global feature importance was determined by ranking the average absolute values ​​of the SHAP framework analysis results, and a summary point plot was used to observe the directional and nonlinear relationships between the eigenvalues, i.e., each influencing factor (PGA, PGV, DEM, DE, Lithology, DF, BD, Height), and the predicted contribution. The ranking of the average absolute values ​​of the SHAP framework analysis results quantifies the overall contribution of each influencing factor to the building damage prediction results, thereby identifying the main control factors. The summary point plot analysis not only reveals the nonlinear and directional relationships between each factor and the prediction results but also verifies whether the model output conforms to the laws of earthquake engineering. These analysis results provide important basis for the explanation of earthquake damage mechanisms, model optimization, and post-earthquake risk assessment.

[0061] To make the earthquake-damaged building prediction method and its beneficial effects provided by this invention clearer, the following explanation is based on examples. (Refer to...) Figure 2As shown, this specific embodiment uses the Kahramanmarash earthquake zone in Türkiye in 2023 as the study area. Figure 2 (a) in the figure represents the overall tectonic background of the study area. The white rectangular dashed box indicates... Figure 2 The range of (b) in the text. Where NAF is the North Anatolian Fault and EAF is the East Anatolian Fault. Figure 2 (b) in the figure provides more detailed information about the study area. The white dashed box indicates the extent of the study area, and at the same time, it corresponds to... Figure 2 The range shown in (c) is as follows. Figure 2 (c) in the diagram shows the seismic intensity distribution of the study area with administrative boundaries. Seismic intensity data are from the United States Geological Survey (USGS). Figure 2 (d)-(f) in the image are photos of the earthquake scene; Figure 2 (d) in the image is a photograph of the earthquake scene in the Besni area; Figure 2 (e) in the text refers to Kahramanmara. Earthquake scene photos of the area; Figure 2 (f) in the image is a photograph of the earthquake scene in the Hatay region.

[0062] Reference Figure 3 As shown, this represents the spatial distribution of factors influencing building damage in the study area. Figure 3 In the figure, (a) represents the peak ground acceleration (PGA). Figure 3 (b) in the figure represents the peak ground speed (PGV). Figure 3 (c) in the figure represents the Copernicus Digital Elevation Model, specifically with a resolution of 30m. Figure 3 In the figure, (d) represents the distance from each point in the study area to the epicenter, i.e., the epicentral distance. Figure 3 In the figure, (e) represents the distance from each point in the study area to the fault, i.e., the fault distance. Figure 3 (f) represents lithological data, sourced from the Global Lithological Map. The lithology is numerically represented using unique thermal coding, specifically as follows: 1: Water body, 2: Unconsolidated sediments, 3: Mixed sedimentary rocks, 4: Basic volcanic rocks, 5: Carbonate sedimentary rocks, 6: Intermediate volcanic rocks, 7: Intermediate plutonic rocks, 8: Evaporites, 9: Acidic plutonic rocks, 10: Basic plutonic rocks, 11: Metamorphic rocks, 12: Siliceous clastic sedimentary rocks, 13: Volcanic clastic rocks.

[0063] Figure 3 In this context, (g) represents building density, which is obtained by performing spatial point density analysis on building distribution data in the Global Buildings Dataset (GBA). Figure 3 In this context, (h) represents the building height, which is directly derived from the GBA dataset and is expressed in meters (m).

[0064] Reference Figure 4 The image shows a map illustrating the distribution of damaged buildings in the study area. Figure 4 (a) in the figure is a map showing the building damage density distribution used for training and evaluating the building damage prediction model. Figure 4 (b)-(h) in the diagram represent the Building Damage Proxy Map (BDPM) for areas with high building damage density, and their spatial locations correspond to... Figure 4 The white dashed box in (a) shows the difference in coseismic coherence. Larger coherence differences indicate more severe building damage, indicated by a color gradient from light yellow to dark red.

[0065] Reference Figure 5 The image shows a map of building damage density. Figure 5 (a) in the figure represents the predicted building damage density in the study area obtained using this method. Figure 5 (b)-(g) in the map are detailed maps of building damage density in representative cities. For example... Figure 5 As shown in (a), the spatial prediction of building damage density highlights obvious hotspot areas, including Kahramanmara. Osmaniye, Nurdagi, Malatya, Adiyaman, and Elbistan. These hotspots highly overlap with areas that experienced strong earthquake tremors and are consistent with surface rupture lines, showing a building damage density consistent with that revealed in this specific embodiment, i.e. Figure 4 The spatial distribution characteristics are consistent with (a) in the data. Furthermore, building damage density maps of several major cities, i.e. Figure 5 (b)-(g) clearly present the spatial distribution of building damage, providing an important reference for disaster relief planning. The preliminary results show that the building damage prediction model provided in this specific implementation can effectively capture the geophysical and structural driving factors of building damage and can reproduce the heterogeneous damage patterns in urban and rural areas.

[0066] Reference Figure 6 The image shown is a comparison and evaluation chart of the prediction results. Among them, Figure 6 (a) in the figure is a spatial residual diagram of the predicted value and the actual building damage density. Figure 6 (b) in the figure is a scatter plot comparing the predicted density and the actual density of the study area. Figure 6 (c) in the figure represents the residual histogram and density curve. Figure 6 (d) in the figure is a spatial hotspot comparison chart based on Getis-Ord Gi* statistics.

[0067] To verify the effectiveness of the model's predictions, this specific implementation method evaluated the predicted building damage density from multiple perspectives. The residual analysis results are as follows: Figure 6 As shown in (a), approximately 90% of the residuals are close to zero, with only a few larger deviations, indicating that the building damage prediction model performs well spatially. A scatter plot of predicted values ​​versus actual density is shown below. Figure 6 As shown in (b) of the diagram, R 2 = 0.8083, indicating a high degree of consistency between the two. See the residual distribution plot for reference. Figure 6 As shown in (c), the bias is concentrated near zero, indicating a small deviation. The model evaluation metrics are RMSE = 488.10 and MAE = 309.36, further validating the overall reliability of the building damage prediction model.

[0068] Spatial hotspot analysis was performed using Getis-Ord Gi* statistics, and the results are as follows: Figure 6 As shown in (d), the predicted hotspots highly overlap with the actual hotspots, with a precision of 0.762, recall of 0.730, F1-score of 0.745, and hotspot overlap rate of 0.730. This demonstrates the powerful ability of the building damage prediction model to reproduce the spatial distribution pattern of building damage. In summary, the building damage prediction model used in this specific implementation can effectively identify earthquake-induced changes in building density, providing a new method for timely acquisition of emergency response information in the future.

[0069] Reference Figure 7 The image shows the analytical results of the building damage prediction model used in this specific embodiment under the SHAP framework. Figure 7 (a) is a scatter plot of the predicted building damage density versus the actual density on the test set. Figure 7 (b) in the figure is the residual histogram obtained by subtracting the predicted damage density from the actual damage density. Figure 7 (c) in the figure is a feature importance diagram in the building damage prediction model based on the SHAP framework analysis results. Figure 7 Figure (d) in the figure summarizes the analysis results of the SHAP framework for the building damage prediction model. In this specific embodiment, a gradient boosting machine learning model was used for training. The resulting building damage prediction model, combined with the SHAP framework, quantitatively analyzed the nonlinear impact of various influencing factors on building damage. The evaluation metrics (RMSE = 556.83, MAE = 354.35, R² = 0.85) indicate that the building damage prediction model has good performance. Furthermore, the scatter plot of the test set and prediction results... Figure 7 (a) and the residual plot, i.e. Figure 7(b) further demonstrates the strong predictive power of the building damage prediction model, thus ensuring the reliability of the factor importance analysis. The interpretation results of the SHAP framework analysis show that BD, DEM, and PGA are the main factors influencing building damage, and all are positively correlated with the degree of damage. Figure 7 As shown in (c) to (d) in the diagram.

[0070] In a specific embodiment of the present invention, an earthquake-damaged building prediction system is also provided, referring to... Figure 8 As shown, it includes: The prediction module is used to acquire near-real-time multi-source impact factor data of the target area after an earthquake, input it into the building damage prediction model, and output the prediction results for the target area. The building damage prediction model is established through the following steps: Obtain a multi-source impact factor dataset as training data; The training data is input into the gradient boosting model for training, and a building damage prediction model that reflects the mapping relationship between multi-source influencing factors and building damage density is established. The damage mechanism analysis module is used to analyze the process of outputting prediction results through the building damage prediction model based on the SHAP framework, quantify the contribution of multi-source impact factor data to the prediction results, and obtain the building damage mechanism under real-time multi-source impact factor data.

[0071] In this specific implementation, a modular architecture is used to realize the materialization and automation of the methodology. The system consists of two core functional units: a prediction module and a damage mechanism analysis module. Each module performs its own function while being interconnected, jointly ensuring the efficient execution of the technical solution and the output of results.

[0072] The prediction module, as the system's front-end execution unit, is responsible for rapid post-earthquake response. This module processes near-real-time multi-source influencing factor data of the target area by calling a pre-trained building damage prediction model and directly outputs intuitive prediction results. This design eliminates the need for manual intervention in complex model calculations during post-earthquake assessment, achieving one-click automated generation from data to conclusions, significantly improving the efficiency of emergency response.

[0073] The damage mechanism analysis module serves as the system's backend analysis unit, providing scientific evidence and theoretical support for the prediction results. This module specializes in in-depth analysis of building damage prediction models based on the SHAP framework. By quantifying the contribution of various multi-source influencing factors to the prediction results, it transforms the model's internal decision-making logic into quantifiable mechanistic interpretations. This function not only enhances the reliability and credibility of the entire system's output but also elevates the prediction system from a mere tool to a scientific platform for understanding the mechanisms of earthquake damage.

[0074] The prediction module and the damage mechanism analysis module have a close functional synergy. The prediction module is responsible for quickly generating conclusions, while the damage mechanism analysis module delves into the causes. Both modules work together on the same building damage prediction model; the former uses the model for rapid prediction, while the latter analyzes the model to verify and deepen understanding. This collaborative working mode enables the system to not only provide timely disaster assessments but also output long-term valuable mechanistic knowledge, achieving a unification of the dual goals of emergency response and scientific research. Through this modular division of labor and collaboration, the entire system constitutes a complete technical solution that meets both practical needs and possesses scientific depth.

[0075] In a specific embodiment of the present invention, a computer device is also provided. Specifically, the computer device includes a processor and a memory. The memory is used to store a computer program, which includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to realize a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used to acquire near-real-time multi-source influence factor data of the target area after an earthquake, input it into a building damage prediction model, and output the prediction results of the target area. The building damage prediction model is established through the following steps: Obtain a multi-source impact factor dataset as training data; The training data is input into the gradient boosting model for training, and a building damage prediction model that reflects the mapping relationship between multi-source influencing factors and building damage density is established. Based on the SHAP framework, the process of outputting prediction results through the building damage prediction model is analyzed, the contribution of multi-source impact factor data to the prediction results is quantified, and the building damage mechanism under real-time multi-source impact factor data is obtained.

[0076] In a specific embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and extended storage media supported by the terminal device. The computer-readable storage medium provides storage space, which stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the methods in the above embodiments; the one or more instructions in the computer-readable storage medium are loaded and executed by the processor in the following steps: acquiring near-real-time multi-source influence factor data of the target area after the earthquake, inputting it into the building damage prediction model, and outputting the prediction results of the target area; the building damage prediction model is established through the following steps: Obtain a multi-source impact factor dataset as training data; The training data is input into the gradient boosting model for training, and a building damage prediction model that reflects the mapping relationship between multi-source influencing factors and building damage density is established. Based on the SHAP framework, the process of outputting prediction results through the building damage prediction model is analyzed, the contribution of multi-source impact factor data to the prediction results is quantified, and the building damage mechanism under real-time multi-source impact factor data is obtained.

[0077] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0078] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0079] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0080] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0081] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the scope of the invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0082] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be appropriately combined to form other embodiments that can be understood by those skilled in the art. The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.

Claims

1. A method for predicting earthquake-damaged buildings, characterized in that, include: Near-real-time multi-source impact factor data of the target area after the earthquake is acquired and input into the building damage prediction model, which then outputs the prediction results for the target area. The building damage prediction model is established through the following steps: Obtain a multi-source impact factor dataset as training data; Training data is input into an ensemble learning model for training, and a building damage prediction model that reflects the mapping relationship between multi-source influencing factors and building damage density is established. Based on the SHAP framework, the process of outputting prediction results through the building damage prediction model is analyzed, the contribution of multi-source impact factor data to the prediction results is quantified, and the building damage mechanism under real-time multi-source impact factor data is obtained.

2. The method for predicting earthquake-damaged buildings according to claim 1, characterized in that, The multi-source impact factor data includes peak ground velocity, peak ground acceleration, digital elevation model, epicenter distance, fault distance, lithology, building density, and building height.

3. The method for predicting earthquake-damaged buildings according to claim 1, characterized in that, After obtaining the multi-source impact factor dataset, the process also includes data preprocessing: Unify all raster data in the acquired multi-source impact factor dataset to the same spatial reference frame and resolution; Standardize the continuous variables in the multi-source impact factor dataset; One-hot encoding is used to numerically represent the categorical variables in the multi-source influence factor dataset.

4. The method for predicting earthquake-damaged buildings according to claim 1, characterized in that, The step of inputting training data into the ensemble learning model for training includes: The hyperparameters of the ensemble learning model are optimized by combining random grid search with cross-validation. The hyperparameters include at least one of the following: maximum tree depth, number of base learners, learning rate, subsampling rate, and feature column sampling rate.

5. The method for predicting earthquake-damaged buildings according to claim 1, characterized in that, After establishing the building damage prediction model, the process also includes performance evaluation of the building damage prediction model: The performance of the building damage prediction model is evaluated using at least one of the following indicators: mean square error, root mean square error, mean absolute error, and coefficient of determination.

6. The method for predicting earthquake-damaged buildings according to claim 1, characterized in that, The analysis of the building damage prediction model based on the SHAP framework includes: An interpreter is built using the training data as background data, and the SHAP value of the test set samples is calculated. The importance of each feature value is obtained by sorting the average absolute value of the SHAP values ​​and then integrating them to form the global feature importance. By using summary point plots to analyze the importance of global features, the directional relationship between feature values ​​and prediction contributions can be obtained.

7. The method for predicting earthquake-damaged buildings according to claim 1, characterized in that, After outputting the prediction result for the target region, the following are also included: The prediction results are combined with geographic coordinate information to generate a spatial distribution map of building damage density; Identify building damage hotspots based on spatial distribution maps of building damage density.

8. A system for predicting earthquake-damaged buildings, characterized in that, include: The prediction module is used to acquire near-real-time multi-source impact factor data of the target area after an earthquake, input it into the building damage prediction model, and output the prediction results for the target area. The building damage prediction model is established through the following steps: Obtain a multi-source impact factor dataset as training data; Training data is input into an ensemble learning model for training, and a building damage prediction model that reflects the mapping relationship between multi-source influencing factors and building damage density is established. The damage mechanism analysis module is used to analyze the process of outputting prediction results through the building damage prediction model based on the SHAP framework, quantify the contribution of multi-source impact factor data to the prediction results, and obtain the building damage mechanism under real-time multi-source impact factor data.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the earthquake-damaged building prediction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the earthquake-damaged building prediction method as described in any one of claims 1 to 7.