A method for predicting seismic sand liquefaction deformation based on machine learning

By constructing a multi-source coupled seismic dataset and combining XGBoost and Bayesian optimization techniques, the problems of dataset integrity and model optimization in earthquake sand liquefaction deformation prediction were solved, achieving high-precision prediction and convenient engineering applications.

CN121389736BActive Publication Date: 2026-04-17CHINA INST OF WATER RESOURCES & HYDROPOWER RES
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA INST OF WATER RESOURCES & HYDROPOWER RES
Filing Date
2025-10-13
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies lack effective integration of multi-source information in predicting liquefaction deformation of sandy soil during earthquakes. The datasets are incomplete and lack representativeness, making it difficult to handle missing and outlier values. Model hyperparameter adjustments rely on empirical settings, and the deployment methods are limited, affecting prediction accuracy and applicability.

Method used

A multi-source coupled seismic dataset was constructed. Missing values ​​were handled by mean imputation, outliers were removed by IQR, key features were selected by Spearman correlation coefficient, and Z-score standardization was performed. A prediction model was built by combining the XGBoost algorithm and Bayesian optimization techniques. The model was then encapsulated into a web service interface using Flask to improve data quality and optimize the model.

Benefits of technology

It significantly improves the data quality and model accuracy for earthquake sand liquefaction deformation prediction, supports cross-system applications and real-time prediction, adapts to complex geological and seismic conditions, and provides reliable technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121389736B_ABST
    Figure CN121389736B_ABST
Patent Text Reader

Abstract

This invention discloses a machine learning-based method for predicting earthquake-induced sand liquefaction deformation, relating to the fields of earthquake engineering and machine learning. The method constructs a multi-source coupled earthquake dataset using engineering survey reports and monitoring data from historical earthquake cases. This dataset undergoes missing value processing, outlier detection, feature selection, and standardization to obtain an earthquake-induced sand liquefaction deformation dataset. Based on this dataset, a prediction model is constructed. The root mean square error, mean absolute error, and coefficient of determination of the model are calculated. Based on corresponding thresholds, the model is adjusted for hyperparameters using Bayesian optimization or encapsulated using Flask. New earthquake-induced sand liquefaction feature data is then processed and input into the encapsulated model to obtain the predicted earthquake-induced sand liquefaction deformation results. This method can effectively predict earthquake-induced sand liquefaction deformation, providing support for earthquake disaster prevention and engineering design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of machine learning and geological engineering, and specifically to a machine learning-based method for predicting earthquake sand liquefaction deformation. Background Technology

[0002] In the fields of earthquake engineering and geological engineering, earthquake-induced sand liquefaction poses a significant threat to the safety of buildings and geological structures. Traditional sand liquefaction prediction methods are mainly based on empirical formulas and simple physical models, relying on limited field data and theoretical assumptions. They are difficult to fully consider complex geological conditions, seismic characteristics, and the interaction between geological conditions and seismic characteristics. Faced with complex and ever-changing realities, traditional methods have limited prediction accuracy and applicability.

[0003] With the rapid development of machine learning technology, it has demonstrated powerful predictive and analytical capabilities in numerous fields. In earthquake engineering, machine learning provides new ideas and methods for predicting earthquake sand liquefaction deformation. However, existing methods still have limitations. Earthquake sand liquefaction is influenced by multiple factors, including groundwater level, soil layer distribution, and seismic motion parameters. Existing methods lack effective integration of multi-source information, resulting in insufficient completeness and representativeness of the dataset. Their handling of missing and outlier values, as well as feature selection methods, are relatively simplistic, easily overlooking key features or introducing noise, thus affecting model training performance. Furthermore, existing methods rely on empirical settings for hyperparameter adjustment of machine learning models, making it difficult to automatically search for optimal parameter combinations. The models also have limited deployment options and lack convenient interfaces to support real-time prediction and cross-system applications. Therefore, how to effectively integrate multi-source data, optimize data processing, improve model performance, and facilitate practical applications in earthquake sand liquefaction deformation prediction has become an urgent problem to solve. To address this, a machine learning-based method for predicting earthquake sand liquefaction deformation is proposed. Summary of the Invention

[0004] The purpose of this invention is to provide a machine learning-based method for predicting earthquake sand liquefaction deformation, in order to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A machine learning-based method for predicting earthquake-induced liquefaction deformation of sandy soil includes the following steps:

[0007] S1. By using engineering survey reports and monitoring data of historical earthquake cases, obtain earthquake sand liquefaction characteristic data and sand liquefaction deformation amount of earthquake cases, and construct a multi-source coupled earthquake dataset accordingly;

[0008] S2. The missing value processing of the multi-source coupled seismic dataset is performed. The multi-source coupled seismic dataset with missing value processing is then subjected to anomaly detection using IQR. Feature selection processing is then performed on the multi-source coupled seismic dataset with anomaly detection. At the same time, Z-score standardization is used to perform standardization processing to obtain the seismic sand liquefaction deformation dataset, which is then stored in the database.

[0009] S3. Based on the earthquake sand liquefaction deformation dataset, construct an earthquake sand liquefaction deformation prediction model, calculate the root mean square error, mean absolute error, and coefficient of determination for the earthquake sand liquefaction deformation prediction model, judge the constructed earthquake sand liquefaction deformation prediction model according to the root mean square error threshold, mean absolute error threshold, and coefficient of determination threshold, and output the judgment result.

[0010] S4. Based on the judgment results, the hyperparameters of the earthquake sand liquefaction deformation prediction model are adjusted or the model is encapsulated. The earthquake sand liquefaction deformation prediction model after hyperparameter adjustment is retrained and updated using the earthquake sand liquefaction deformation dataset.

[0011] S5. Obtain new earthquake sand liquefaction characteristic data, perform missing value processing, outlier detection, feature selection processing, and standardization processing on the new earthquake sand liquefaction characteristic data, and input it into the encapsulated earthquake sand liquefaction deformation prediction model to obtain the earthquake sand liquefaction deformation prediction results.

[0012] Preferably, the method for constructing a multi-source coupled seismic dataset is as follows:

[0013] Based on engineering survey reports and monitoring data of historical earthquake cases, the following parameters are obtained: groundwater level depth, soil layer depth, soil layer distribution, relative density, water content, effective overburden stress, clay content, SPT blow count, soil particle specific gravity, void ratio, peak ground acceleration, earthquake duration, dominant period, velocity response spectrum, magnitude, epicentral distance, and surface subsidence. Groundwater level depth, soil layer depth, and soil layer distribution reflect the geological conditions of the earthquake-affected area. Relative density, water content, effective overburden stress, clay content, SPT blow count, soil particle specific gravity, and void ratio reflect the physical and mechanical properties of the soil in the earthquake-affected area. Peak ground acceleration, earthquake duration, dominant period, velocity response spectrum, magnitude, and epicentral distance reflect the earthquake intensity and characteristics of the earthquake case.

[0014] For each earthquake case, geographical location information and a unique ID number are manually labeled. Based on the geographical location information and unique ID number of the earthquake case, the earthquake sand liquefaction characteristic data and sand liquefaction deformation amount corresponding to the earthquake case are matched with the geographical location information and unique ID number of the earthquake case to establish a correlation and construct a multi-source coupled earthquake dataset. The multi-source coupled earthquake dataset consists of the geographical location information and unique ID number of the earthquake case, as well as the earthquake sand liquefaction characteristic data and sand liquefaction deformation amount determined by the geographical location information and unique ID number of the earthquake case. The earthquake sand liquefaction characteristic data includes groundwater level depth, soil layer burial depth, soil layer distribution, relative density, water content, effective overlying stress, clay content, SPT blow count, soil particle specific gravity, void ratio, peak ground acceleration, earthquake duration, dominant period, velocity response spectrum value, magnitude, and epicentral distance. The sand liquefaction deformation amount is the surface subsidence amount, which is used to measure the degree of earthquake sand liquefaction deformation.

[0015] Preferably, the method for handling missing values ​​and detecting anomalies in the multi-source coupled seismic dataset, and for standardizing the multi-source coupled seismic dataset after feature selection processing, is as follows:

[0016] For multi-source coupled seismic datasets, the mean imputation method is used to fill missing values ​​by calculating the average value of existing data. This method can maintain the overall characteristics of the data to a certain extent and avoid data bias caused by missing values.

[0017] The formula for the mean-filling method is:

[0018]

[0019] in, This is the result of filling in missing values ​​after processing with the mean imputation method. The weight of each data point when calculating the average. To perform the summation operation on all data in dataset S, Let i be the i-th data in the data set S;

[0020] For a multi-source coupled seismic dataset that has undergone missing value processing, the first quartile Q1 and the third quartile Q3 of the dataset are calculated to obtain IQR = Q3 - Q1. Based on this, a reasonable range is determined. Generally, data points that are less than Q1 - 1.5 * IQR or greater than Q3 + 1.5 * IQR are considered outliers. Once outliers are identified in a multi-source coupled seismic dataset that has undergone missing value processing, they are discarded to avoid interference from outlier data with subsequent analysis and model training, thus ensuring the reliability and accuracy of the data.

[0021] The source-coupled seismic dataset after feature selection is standardized using the Z-score standardization formula. After standardization, the mean of the data in the dataset becomes 0, the standard deviation becomes 1, and all data are mapped to a distribution with a uniform scale.

[0022] The Z-score standardization formula is as follows:

[0023] ;

[0024] in, The original data values, The mean of the data. denoted as the standard deviation of the data.

[0025] Preferably, the method for performing feature selection processing on the multi-source coupled seismic dataset after anomaly detection and obtaining the seismic sand liquefaction deformation dataset is as follows:

[0026] The Spearman correlation coefficients of all feature elements in the multi-source coupled seismic dataset after anomaly detection were calculated using the Spearman correlation coefficient formula. The correlation threshold is set to 0.1. The absolute values ​​of the Spearman correlation coefficients of all feature elements are compared with the correlation threshold. Feature elements with Spearman correlation coefficients less than or equal to the correlation threshold are deleted from the multi-source coupled seismic dataset. This is how feature selection processing of the multi-source coupled seismic dataset is achieved.

[0027] The formula for the Spearman correlation coefficient is:

[0028]

[0029] in, It is the Spearman correlation coefficient of the characteristic elements. It is the rank difference between the characteristic element and the i-th observed value of sand liquefaction deformation. It is the sample size;

[0030] The multi-source coupled seismic dataset that has undergone feature selection processing is standardized to obtain the seismic sand liquefaction deformation dataset. The seismic sand liquefaction deformation dataset consists of the sand liquefaction deformation amount and the corresponding seismic sand liquefaction feature data that has undergone feature selection processing.

[0031] The characteristic elements are the groundwater level depth, soil layer depth, soil layer distribution, relative density, water content, effective overlying stress, clay content, SPT blow count, soil particle specific gravity, void ratio, peak ground acceleration, earthquake duration, dominant period, velocity response spectrum value, magnitude or epicentral distance of the earthquake sand liquefaction characteristic data.

[0032] Preferably, the method for constructing an earthquake sand liquefaction deformation prediction model based on an earthquake sand liquefaction deformation dataset is as follows:

[0033] The earthquake sand liquefaction deformation dataset was obtained from the database and divided into training and testing sets in an 8:2 ratio. An XGBoost model was built using the xgboost library, and hyperparameters of the XGBoost model, including the number of trees, learning rate, maximum tree depth, subsampling ratio, and column sampling ratio, were set. The xgboost library is a gradient boosting library, and the XGBoost model is a machine learning algorithm.

[0034] The XGBoost model is trained using the training set to obtain a trained XGBoost model, which is a prediction model for earthquake sand liquefaction deformation.

[0035] Preferably, the method for judging the constructed earthquake sand liquefaction deformation prediction model based on the root mean square error threshold, the mean absolute error threshold, and the coefficient of determination threshold, and outputting the judgment result:

[0036] The test set is input into the earthquake sand liquefaction deformation prediction model. Based on the input data and the patterns learned during model training, the model predicts the earthquake sand liquefaction deformation and outputs the prediction result, i.e., the predicted amount of sand liquefaction deformation. Based on the predicted sand liquefaction deformation and the actual sand liquefaction deformation in the test set The root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (COP) of the earthquake sand liquefaction deformation prediction model were obtained using the root mean square error formula, the mean absolute error formula, and the COP formula, respectively. ;

[0037] The root mean square error (RMSE) and mean absolute error (MAE) are respectively:

[0038]

[0039]

[0040] Where N is the number of samples, It is the predicted liquefaction deformation of sand in the i-th sample. This is the actual amount of sand liquefaction deformation for the i-th sample. and These are the root mean square error and mean absolute error of the earthquake sand liquefaction deformation prediction model, respectively.

[0041] The formula for the coefficient of determination is:

[0042]

[0043] in, It is the sample size. It is the predicted liquefaction deformation of sand in the i-th sample. This is the actual amount of sand liquefaction deformation for the i-th sample. It is the coefficient of determination for the earthquake sand liquefaction deformation prediction model. It is the average of the actual sand liquefaction deformation of all samples;

[0044] The root mean square error, mean absolute error, and coefficient of determination of the earthquake sand liquefaction deformation prediction model are compared with the root mean square error threshold, mean absolute error threshold, and coefficient of determination threshold, respectively.

[0045] If the root mean square error of the earthquake sand liquefaction deformation prediction model is less than the root mean square error threshold of 1, the mean absolute error is less than the mean absolute error threshold of 0.8, and the coefficient of determination is greater than the coefficient of determination threshold of 0.7, then the judgment result is to deploy the earthquake sand liquefaction deformation prediction model.

[0046] Otherwise, the hyperparameters of the earthquake sand liquefaction deformation prediction model will be adjusted.

[0047] Preferably, the method for adjusting the hyperparameters of the earthquake sand liquefaction deformation prediction model is as follows:

[0048] The hyperparameters of the earthquake sand liquefaction deformation prediction model were adjusted using Bayesian optimization. Bayesian optimization is a hyperparameter tuning method based on Bayes' theorem, which can efficiently search for the optimal combination of hyperparameters in the hyperparameter search space. The hyperparameters of the earthquake sand liquefaction deformation prediction model are the number of trees, learning rate, maximum tree depth, subsampling ratio, and column sampling ratio.

[0049] The hyperparameter search space is the set of value ranges set for hyperparameters during the machine learning model tuning process.

[0050] Preferably, the method for encapsulating the earthquake sand liquefaction deformation prediction model:

[0051] This paper describes how to encapsulate an earthquake sand liquefaction deformation prediction model into a web service interface using Flask, a lightweight Python web framework.

[0052] Preferably, the method for obtaining the prediction results of earthquake sand liquefaction deformation is as follows:

[0053] The new seismic sand liquefaction feature data, after missing value processing, outlier detection, feature selection, and standardization, is packaged into a data packet in JSON format. The data packet is then sent to the server's application programming interface (API) via HTTP protocol. The server obtains the data packet through the API, calls the seismic sand liquefaction deformation prediction model encapsulated as a web service interface, and outputs the seismic sand liquefaction deformation prediction result based on the data packet. The prediction result is then returned via HTTP protocol. JSON format is a lightweight data exchange format, and HTTP protocol is a network protocol.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] 1. This invention realizes the construction of a multi-source data integration and high-quality data processing system. By using engineering survey reports and monitoring data of historical earthquake cases, it obtains multi-source coupled feature data of geological conditions, soil physical and mechanical properties, and seismic characteristics. By establishing associations through geographical location information and unique ID numbers, it solves the problem of insufficient dataset completeness in existing technologies and constructs a more representative multi-source coupled earthquake dataset. It adopts the mean imputation method to handle missing values, the IQR method to remove outliers, the Spearman correlation coefficient to screen key features, and Z-score standardization to form a full-process data processing system of missing value repair, outlier filtering, feature selection, and scale unification. It effectively retains core features and filters noise, significantly improves the quality of input model data, and lays a data foundation for high-precision prediction.

[0056] 2. This invention achieves intelligent model optimization and convenient engineering application support. It constructs a prediction model based on the XGBoost algorithm and introduces Bayesian optimization technology to automatically tune the model's hyperparameters, overcoming the limitations of existing technologies that rely on experience-based settings. This efficiently searches for the optimal parameter combination, improving the model's adaptability to complex geological and seismic conditions and its prediction accuracy. Furthermore, it encapsulates the model into a standardized web service interface using the lightweight web framework Flask, supporting JSON data interaction based on the HTTP protocol. This enables cross-system calls and real-time prediction, solving the problems of single deployment and missing interfaces in traditional methods. It facilitates integration with earthquake monitoring and engineering design platforms. Simultaneously, it establishes a multi-dimensional model evaluation mechanism, forming a closed loop of training, evaluation, optimization, and deployment, ensuring continuous optimization of model performance and providing reliable technical support for earthquake disaster prevention and infrastructure design. Attached Figure Description

[0057] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This is a flowchart of the steps of the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be described in detail below. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other implementation methods obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0060] Examples, such as Figure 1 As shown, a machine learning-based method for predicting earthquake sand liquefaction deformation includes the following steps:

[0061] S1. By using engineering survey reports and monitoring data of historical earthquake cases, obtain earthquake sand liquefaction characteristic data and sand liquefaction deformation amount of earthquake cases, and construct a multi-source coupled earthquake dataset accordingly;

[0062] S2. The missing value processing of the multi-source coupled seismic dataset is performed. The multi-source coupled seismic dataset with missing value processing is then subjected to anomaly detection using IQR. Feature selection processing is then performed on the multi-source coupled seismic dataset with anomaly detection. At the same time, Z-score standardization is used to perform standardization processing to obtain the seismic sand liquefaction deformation dataset, which is then stored in the database.

[0063] S3. Based on the earthquake sand liquefaction deformation dataset, construct an earthquake sand liquefaction deformation prediction model, calculate the root mean square error, mean absolute error, and coefficient of determination for the earthquake sand liquefaction deformation prediction model, judge the constructed earthquake sand liquefaction deformation prediction model according to the root mean square error threshold, mean absolute error threshold, and coefficient of determination threshold, and output the judgment result.

[0064] S4. Based on the judgment results, the hyperparameters of the earthquake sand liquefaction deformation prediction model are adjusted or the model is encapsulated. The earthquake sand liquefaction deformation prediction model after hyperparameter adjustment is retrained and updated using the earthquake sand liquefaction deformation dataset.

[0065] S5. Obtain new earthquake sand liquefaction characteristic data, perform missing value processing, outlier detection, feature selection processing, and standardization processing on the new earthquake sand liquefaction characteristic data, and input it into the encapsulated earthquake sand liquefaction deformation prediction model to obtain the earthquake sand liquefaction deformation prediction results.

[0066] Furthermore, the working principle of the present invention will be illustrated below through embodiments:

[0067] This experiment focuses on a coastal city prone to earthquakes, which has complex geological conditions, a high groundwater level, and widespread sandy soil.

[0068] By collecting engineering survey reports and monitoring data of historical earthquake cases, a total of 120 historical earthquake cases from 1980 to 2022 were collected in a seismically active area. Each case contains 20 sand liquefaction characteristic data and surface subsidence. Each case is labeled with a unique ID and geographical coordinates. The characteristic data and deformation are associated with the ID and geographical coordinates to form a structured dataset and construct a multi-source coupled earthquake dataset.

[0069] For multi-source coupled seismic datasets, mean imputation was employed. For example, when clay content features were missing in sand liquefaction characteristic data, the mean of the non-missing clay content features was calculated, assuming a mean of 15.2%, and the missing data was imputed. Outlier detection was then performed on the multi-source coupled seismic datasets after missing value processing. Taking SPT blow counts in sand liquefaction characteristic data as an example, the IQR was calculated to be 4, and outliers ranged from 0 to 16. Cases with SPT blow counts less than 0 or greater than 16 were deleted, resulting in the detection and discarding of 5 outlier samples. Feature selection was then performed on the outlier-detected multi-source coupled seismic datasets, and the Spearman correlation coefficient between each feature and surface subsidence was calculated. Features with correlation coefficients less than or equal to 0.1 were removed. For example, the correlation coefficient of soil particle weight in the sand liquefaction feature data was 0.08, so it was removed. Finally, 14 key features were retained: groundwater level depth, soil layer depth, soil layer distribution, relative density, water content, effective overlying stress, clay content, SPT blow count, soil particle weight, void ratio, peak ground acceleration, earthquake duration, dominant period, velocity response spectrum, magnitude, and epicentral distance. The retained features were standardized using Z-scores. For example, the standardized value of a sample with a groundwater level depth of 3.2m was -0.375.

[0070] The 120 samples were divided into a training set (96 samples) and a test set (24 samples) in an 8:2 ratio. An XGBoost model was used with the following initial hyperparameters: 100 trees, a learning rate of 0.1, a maximum tree depth of 6, a subsampling ratio of 0.8, and a column sampling ratio of 0.8. The validation set loss was recorded during training, and the model converged after 100 iterations, resulting in the earthquake sand liquefaction deformation prediction model. The test set was then input into the trained model to calculate the predicted earthquake sand liquefaction deformation. The model's root mean square error (RMSE) was 2.15 cm, mean absolute error (MAE) was 1.82 cm, and coefficient of determination (R²) was 0.72. Comparing the RMSE, MAE, and R² of the earthquake sand liquefaction deformation prediction model with their respective thresholds, it was found that the RMSE of the earthquake sand liquefaction deformation prediction model was greater than 1, the MAE was greater than 0.8, and the R² was greater than 0.7. Since the RMSE and MAE did not meet the standards, hyperparameter adjustment was triggered.

[0071] A hyperparameter search space is defined, in which the number of trees ranges from 50 to 200, the learning rate ranges from 0.01 to 0.2, the maximum tree depth ranges from 3 to 10, the subsampling ratio ranges from 0.6 to 1.0, and the column sampling ratio ranges from 0.6 to 1.0. After 20 iterations, the optimal hyperparameters are: the optimal number of trees is 150, the optimal learning rate is 0.05, the optimal maximum tree depth is 8, and the optimal subsampling ratio and column sampling ratio are both 0.7. The hyperparameters of the earthquake sand liquefaction deformation prediction model are adjusted according to the optimal parameters, and the model is retrained. The root mean square error (RMSE) of the retrained earthquake sand liquefaction deformation prediction model is calculated to be 0.98 cm, the mean absolute error (MAE) is 0.75 cm, and the coefficient of determination (R²) is 0.85. The results show that RMSE is less than 1, MAE is less than 0.8, and R² is greater than 0.7, satisfying the deployment conditions.

[0072] The earthquake sand liquefaction deformation prediction model was encapsulated as a Web service interface using Flask. New site feature data, which had already undergone missing value processing, outlier detection, feature selection, and standardization, was acquired. This site feature data was packaged into a data packet and sent to the server's application programming interface (API) via an HTTP POST request. The server received the data packet through the API and called the earthquake sand liquefaction deformation prediction model encapsulated as a Web service interface. Based on the data packet, the earthquake sand liquefaction deformation prediction model output the earthquake sand liquefaction deformation prediction result, returning a predicted surface settlement of 12.3 cm. Field monitoring verified that the error was 0.9 cm, meeting the accuracy requirements for engineering applications.

[0073] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0074] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0075] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; modifications to the technical solutions described in the foregoing embodiments, or equivalent substitutions of some of the technical features, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for predicting seismic sand liquefaction deformation based on machine learning, characterized by, Includes the following steps: S1. By using engineering survey reports and monitoring data of historical earthquake cases, obtain earthquake sand liquefaction characteristic data and sand liquefaction deformation amount of earthquake cases, and construct a multi-source coupled earthquake dataset accordingly; S2. The missing value processing of the multi-source coupled seismic dataset is performed. The multi-source coupled seismic dataset with missing value processing is then subjected to anomaly detection using IQR. Feature selection processing is then performed on the multi-source coupled seismic dataset with anomaly detection. At the same time, Z-score standardization is used to perform standardization processing to obtain the seismic sand liquefaction deformation dataset, which is then stored in the database. S3. Based on the earthquake sand liquefaction deformation dataset, construct an earthquake sand liquefaction deformation prediction model, calculate the root mean square error, mean absolute error, and coefficient of determination for the earthquake sand liquefaction deformation prediction model, judge the constructed earthquake sand liquefaction deformation prediction model according to the root mean square error threshold, mean absolute error threshold, and coefficient of determination threshold, and output the judgment result. S4. Based on the judgment results, the hyperparameters of the earthquake sand liquefaction deformation prediction model are adjusted or the model is encapsulated. The earthquake sand liquefaction deformation prediction model after hyperparameter adjustment is retrained and updated using the earthquake sand liquefaction deformation dataset. S5. Obtain new earthquake sand liquefaction characteristic data, perform missing value processing, outlier detection, feature selection processing and standardization processing on the new earthquake sand liquefaction characteristic data, and input it into the encapsulated earthquake sand liquefaction deformation prediction model to obtain the earthquake sand liquefaction deformation prediction results. The method for constructing the earthquake-induced sand liquefaction deformation prediction model is as follows: The earthquake sand liquefaction deformation dataset was obtained from the database and divided into training and testing sets in an 8:2 ratio. An XGBoost model was built using the xgboost library, and the hyperparameters of the XGBoost model were set, including the number of trees, learning rate, maximum tree depth, subsampling ratio, and column sampling ratio. The XGBoost model was then trained using the training set to obtain the trained XGBoost model, which is the earthquake sand liquefaction deformation prediction model. The method for judging and outputting the judgment result of the constructed earthquake sand liquefaction deformation prediction model is as follows: The test set is input into the earthquake sand liquefaction deformation prediction model. Based on the input data and the rules learned during the model training process, the model predicts the earthquake sand liquefaction deformation and outputs the prediction result, that is, the predicted sand liquefaction deformation amount. Based on the predicted sand liquefaction deformation amount and the actual sand liquefaction deformation amount in the test set, the root mean square error, mean absolute error and coefficient of determination of the earthquake sand liquefaction deformation prediction model are obtained by using the root mean square error formula, the mean absolute error formula and the coefficient of determination formula, respectively. The root mean square error, mean absolute error, and coefficient of determination of the earthquake sand liquefaction deformation prediction model are compared with the root mean square error threshold, mean absolute error threshold, and coefficient of determination threshold, respectively. If the root mean square error of the earthquake sand liquefaction deformation prediction model is less than the root mean square error threshold, the mean absolute error is less than the mean absolute error threshold, and the coefficient of determination is greater than the coefficient of determination threshold, then the judgment result is to deploy the earthquake sand liquefaction deformation prediction model. Otherwise, the hyperparameters of the earthquake sand liquefaction deformation prediction model will be adjusted. The root mean square error threshold, the mean absolute error threshold, and the coefficient of determination threshold are 1, 0.8, and 0.7, respectively.

2. The method of claim 1, wherein the method of predicting the seismic sand liquefaction deformation based on the machine learning is characterized by, The method for constructing a multi-source coupled seismic dataset: Based on engineering survey reports and testing data of historical earthquake cases, the following parameters were obtained: groundwater level depth, soil layer depth, soil layer distribution, relative density, water content, effective overburden stress, clay content, SPT blow count, soil particle specific gravity, void ratio, peak ground acceleration, earthquake duration, dominant period, velocity response spectrum, magnitude, epicentral distance, and surface subsidence. For each earthquake case, the geographical location information and unique ID number are manually labeled. Based on the geographical location information and unique ID number of the earthquake case, the earthquake sand liquefaction characteristic data and sand liquefaction deformation amount corresponding to the earthquake case are matched with the geographical location information and unique ID number of the earthquake case to establish the association and construct a multi-source coupled earthquake dataset. The earthquake-induced sand liquefaction characteristic data include groundwater level depth, soil layer burial depth, soil layer distribution, relative density, water content, effective overlying stress, clay content, SPT blow count, soil particle specific gravity, void ratio, peak ground acceleration, earthquake duration, dominant period, velocity response spectrum value, magnitude, and epicentral distance. The amount of sand liquefaction deformation is the amount of surface subsidence.

3. The method of claim 2, wherein the method of predicting the seismic sand liquefaction deformation based on the machine learning is characterized by, The method described above involves handling missing values ​​and detecting anomalies in a multi-source coupled seismic dataset, as well as standardizing the multi-source coupled seismic dataset after feature selection. Missing values ​​were filled using the mean imputation method for multi-source coupled seismic datasets; Outliers in the multi-source coupled seismic dataset after missing value processing are filtered using the IQR method and then discarded. Z-score normalization was performed on the source-coupled seismic dataset after feature selection.

4. According to claim 3, a method for predicting earthquake sand liquefaction deformation based on machine learning, characterized in that, The method described above involves feature selection processing of the multi-source coupled seismic dataset after anomaly detection and obtaining the seismic sand liquefaction deformation dataset: The Spearman correlation coefficient of all feature elements in the multi-source coupled seismic dataset after anomaly detection is calculated using the Spearman correlation coefficient formula. The correlation threshold is set to 0.

1. The absolute value of the Spearman correlation coefficient of all feature elements is compared with the correlation threshold. Feature elements with Spearman correlation coefficients less than or equal to the correlation threshold are deleted from the multi-source coupled seismic dataset. This is how feature selection processing of the multi-source coupled seismic dataset is achieved. The multi-source coupled seismic dataset that has undergone feature selection processing is standardized to obtain the seismic sand liquefaction deformation dataset. The characteristic elements are the groundwater level depth, soil layer depth, soil layer distribution, relative density, water content, effective overlying stress, clay content, SPT blow count, soil particle specific gravity, void ratio, peak ground acceleration, earthquake duration, dominant period, velocity response spectrum value, magnitude or epicentral distance of the earthquake sand liquefaction characteristic data.

5. A method for predicting earthquake sand liquefaction deformation based on machine learning, as described in claim 1, characterized in that, The method for adjusting the hyperparameters of the earthquake sand liquefaction deformation prediction model: The hyperparameters of the earthquake sand liquefaction deformation prediction model were adjusted using Bayesian optimization. The Bayesian optimization is a hyperparameter tuning method based on Bayes' theorem.

6. According to claim 5, a method for predicting earthquake sand liquefaction deformation based on machine learning, characterized in that, The method for encapsulating the earthquake sand liquefaction deformation prediction model: The earthquake sand liquefaction deformation prediction model is encapsulated into a web service interface using Flask, thereby realizing the model encapsulation of the earthquake sand liquefaction deformation prediction model. Flask is a lightweight Python web framework.

7. A method for predicting earthquake sand liquefaction deformation based on machine learning, as described in claim 6, is characterized in that... The method for obtaining the prediction results of earthquake-induced sand liquefaction deformation: The new seismic sand liquefaction feature data, which has undergone missing value processing, outlier detection, feature selection processing, and standardization processing, is packaged into a data packet in JSON format. The data packet is then sent to the seismic sand liquefaction deformation prediction model, which is encapsulated as a web service interface, via the HTTP protocol. At the same time, the seismic sand liquefaction deformation prediction model outputs the seismic sand liquefaction deformation prediction results and returns them via the HTTP protocol. The JSON format is a lightweight data exchange format; The HTTP protocol is a network protocol.

Citation Information

Patent Citations

  • Deep saturated sand earthquake-induced liquefaction judgment method

    CN106408211A

  • Testing device for simulating sand liquefaction phenomenon when different earthquake magnitudes of earthquakes occur

    CN108122473A