Hydrate area submarine landslide risk assessment method based on cascade machine learning

By employing cascaded machine learning methods and loss functions based on physical constraints, combined with a dynamic weighting integration strategy, the problems of data sparsity and reliability in the risk assessment of submarine landslides in hydrate areas were solved. This approach enabled comprehensive and high-precision landslide risk assessment and interpretable analysis, supporting disaster prevention and control decision-making.

CN121743977APending Publication Date: 2026-03-27崂山国家实验室
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies for risk assessment of submarine landslides in hydrate areas suffer from problems such as data sparsity, poor reliability of prediction results, and lack of physical rationality and interpretability of models, making it difficult to achieve full coverage, real-time and efficient risk assessment.

Method used

Using a cascaded machine learning approach, a hydrate prediction model and a landslide assessment model are constructed. By combining a loss function with physical constraints and a dynamic weighting integration strategy, a hydrate occurrence depth distribution map is generated and a landslide susceptibility zoning is carried out. Dynamic risk simulation is performed using environmental features such as earthquakes and faults.

Benefits of technology

It effectively overcomes the problem of data scarcity, improves the physical rationality and predictive robustness of the model, achieves full coverage of landslide risk assessment, and provides high-precision and highly interpretable risk assessment results, providing a scientific basis for disaster prevention and control decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121743977A_ABST
    Figure CN121743977A_ABST
Patent Text Reader

Abstract

The invention discloses a hydrate area submarine landslide risk assessment method based on cascade machine learning, and belongs to the technical field of marine geological disaster prediction. The method comprises the following steps: firstly, acquiring historical landslide data, hydrate occurrence data and various environmental characteristics of a research area, and preprocessing the historical landslide data, the hydrate occurrence data and the various environmental characteristics; constructing a hydrate occurrence data set based on local hydrate data and environmental characteristics, training a neural network model fused with physical constraints, predicting the hydrate occurrence depth of the whole region, and generating a distribution diagram; and finally, fusing the prediction depth as a key feature with an environment feature, constructing an enhanced landslide data set, and training an ensemble learning model by adopting a dynamic weight integration strategy based on a geological landscape unit to complete landslide risk prediction. According to the method, the prediction bottleneck caused by hydrate data scarcity is effectively solved through the cascade architecture, the rationality and robustness of the model are improved through physical constraints, and high-precision and interpretable evaluation of the submarine landslide risk of the hydrate region is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of marine geological disaster prediction, and particularly relates to a hydrate zone submarine landslide risk assessment method based on cascading machine learning. BACKGROUND

[0002] Natural gas hydrate is widely distributed in the deep-sea shallow sediments of continental margins. Its phase change can significantly change the mechanical properties of sediments and induce abnormal pore pressure, which is one of the key geological factors triggering submarine landslides. Studies have shown that more than 60% of large submarine landslides in the world have significant spatial correlation with the boundaries of hydrate stability zones. Therefore, how to accurately assess the risk of submarine landslides in hydrate zones has become a core problem that needs to be solved in the fields of marine engineering safety, submarine resource development, and geological disaster prevention.

[0003] In recent years, some studies have attempted to apply machine learning algorithms such as decision trees, random forests, and neural networks to submarine landslide prediction. However, these methods mostly directly incorporate hydrates as one of the ordinary environmental factors into the model, which leads to the model's inability to fully learn the internal physical mechanisms of hydrate's influence on landslides. More importantly, existing machine learning models generally face two fundamental difficulties: 1. Data sparsity dilemma: If PHOD data is used directly, the model cannot make predictions in areas without PHOD due to the limited coverage of "golden data". If PHOD is abandoned, the key indicators describing the influence of hydrates are lost, leading to a significant decrease in the model's evaluation accuracy in hydrate zones. 2. Poor reliability of prediction results: Standard machine learning models, as purely data-driven tools, rely entirely on statistical correlations for prediction, which can produce a large number of results that violate physical common sense, such as predicting hydrate occurrence in shallow water or predicting PHOD greater than the local water depth, severely damaging the model's practicality and credibility.

[0004] With the advancement of marine observation technology and the development of machine learning methods, submarine landslide risk assessment technology has gradually evolved from traditional qualitative analysis to quantitative and intelligent direction. Several representative technical solutions have emerged, but there are still the following key bottlenecks: The invention patent with application number 202511078184.8 discloses a submarine landslide stability evaluation method based on multi-source geophysical data. This method fuses multi-beam sounding, submarine seismograph, pore water pressure monitoring and other multi-source data to construct a three-dimensional geological model and realize landslide stability visualization analysis. However, this method also has obvious limitations: first, it highly depends on intensive and high-precision in-situ geophysical data, which has high implementation cost and limited coverage in wide sea areas or data-scarce areas; second, the model is essentially a “static evaluation” and does not fully consider the dynamic influence of external triggering events such as earthquakes and hydrate rapid decomposition; third, it ignores the core mechanism of the weakening effect of hydrate decomposition on sediment strength; fourth, the model decision-making process is close to a “black box” and lacks explainable analysis of feature contribution, which is not conducive to disaster cause analysis and prevention and control decision-making.

[0005] In addition, the invention patent with application number 202310000068.9 discloses a method for evaluating submarine landslide risk during hydrate exploitation. This method establishes a heat-flow-force-phase change coupled numerical model and quantitatively evaluates the submarine stability during exploitation by combining the Mohr-Coulomb criterion. Although this method has a clear physical mechanism basis, it still faces significant defects: first, it requires high precision and integrity of geological parameters, especially in “golden data” sparse areas such as PHOD (potential hydrate occurrence depth), making it difficult to implement reliable modeling; second, it has high computational complexity and long simulation period, making it difficult to achieve large-scale, near-real-time risk assessment; third, it lacks rapid response capability to external dynamic events such as earthquakes and storm waves. SUMMARY

[0006] To address the limitations of existing technologies in data utilization, prediction gaps, and poor model explainability in submarine landslide risk assessment in hydrate areas, the present invention proposes a method for evaluating submarine landslide risk in hydrate areas based on cascading machine learning.

[0007] The objectives of the present invention can be achieved through the following technical solutions: A method for evaluating submarine landslide risk in hydrate areas based on cascading machine learning, comprising the following steps: S1, data acquisition and preprocessing: acquire historical landslide data, hydrate occurrence data and various environmental feature data of the study area, and perform outlier and missing value processing on the data; S2, hydrate prediction model construction: construct a hydrate occurrence dataset based on local hydrate data and various environmental features, train a physical information neural network model, predict the hydrate occurrence depth of the entire area and generate a distribution map; S3. Landslide Prediction Model Construction: The predicted hydrate occurrence depth is used as a key feature and combined with environmental feature data to construct an enhanced landslide dataset. Based on the geological landscape unit division and dynamic weight integration strategy, an ensemble learning model is trained to construct a landslide prediction model.

[0008] Furthermore, it also includes the following steps: S4. Landslide Prediction: Based on the landslide prediction model, the contribution of each environmental feature is calculated and a landslide susceptibility zoning map is generated. The feature weights are dynamically adjusted through an event feature mapping table to achieve dynamic simulation of landslide risk for specific events.

[0009] Furthermore, in step S3, the dynamic weight integration strategy based on geological landscape units includes: S301. The optimal number of clusters is determined by the contour coefficient, and the study area is divided into multiple geological landscape units. S302. Train multiple machine learning models within each geological landscape unit, select the optimal model based on the comprehensive score, and assign weights accordingly. S303. Based on the geological unit to which the point to be predicted belongs, the corresponding weights are called for dynamic weighted integration to generate a landslide susceptibility zoning map of the whole region.

[0010] Furthermore, the total loss function of the physical information neural network model is: , in, For data fitting loss, As the water depth constraint weight, For the burial depth constraint weight, This is a physical constraint penalty term for the lower limit of water depth. To deepen the physical constraint penalty term for rationality, and: , in, This represents the batch sample size. It is a safety margin. It is a rectified linear function; in, For the first The depth of hydrate distribution at each sample location For the first The actual water depth at each sample location For the upper boundary safety margin, This is the lower boundary scaling factor.

[0011] Further, in the step S4, the event feature mapping table is a pre-constructed mapping relationship table reflecting the weight adjustment coefficients of each environmental feature under different event types; in simulating a specific event, the input features are real-time weighted and preprocessed according to the mapping table, and then input into the landslide prediction model for dynamic risk assessment.

[0012] Further, in the step S1, the environmental feature data includes a peak ground acceleration (PGA) determined based on a seismic event, and: , wherein, represents a seismic weight; a, b, and c are regional geological attenuation coefficients; is a magnitude; is the distance of the target point from the earthquake.

[0013] Compared with the prior art, the advantages and positive effects of the present application are as follows: (1) Effectively overcomes the bottleneck of modeling and prediction caused by the scarcity of hydrate "golden data". The present application adopts a "hydrate prediction-landslide assessment" cascading machine learning architecture, which only needs a small amount of measured hydrate data of known points, and can generate a hydrate occurrence depth distribution map of the entire study area with high precision by using easily accessible environmental features (such as earthquakes, faults, water depth, etc.) through an improved neural network model. This method fundamentally solves the predicament that traditional evaluation techniques are difficult to predict in areas without measured data due to insufficient data coverage, and provides a full-coverage key data foundation for subsequent landslide risk assessment in a low-cost and efficient manner.

[0014] (2) Significantly improves the physical rationality and prediction robustness of the model. The present application introduces a water depth lower limit constraint and a burial depth rationality constraint based on physical laws during the neural network training process, which are embedded in the total loss function in the form of penalty terms, to ensure that the prediction results strictly meet the basic physical conditions of the hydrate stability domain. This method effectively avoids the lack of physical constraints and low prediction accuracy problems that are prone to occur in traditional pure data-driven models, and enhances the generalization ability and spatial continuity of the model in data sparse areas.

[0015] (3) Realizes the explainability of the landslide risk assessment process and the transparency of the decision support. The present application integrates SHAP and other explainable machine learning techniques to quantitatively analyze the contribution of hydrate occurrence depth, seismic intensity, fault distance, and other features to the model prediction results, which not only accurately identifies the main controlling factors of landslides, but also reveals the nonlinear interaction effects between multiple factors. This method changes the black box decision of traditional machine learning into an explainable glass box analysis, making the risk assessment conclusion have high precision and strong explainability, and providing a scientific basis for disaster prevention and control decisions. BRIEF DESCRIPTION OF DRAWINGS

[0016] The invention will now be further described with reference to the accompanying drawings.

[0017] Figure 1 A flowchart illustrating the landslide susceptibility assessment method based on cascaded machine learning and SHAP interpretability analysis provided in this embodiment of the invention. Figure 2 for Figure 1 The data processing results shown in the method are as follows; Figure 3 To provide complete data for hydrates predicted by the improved neural network model; Figure 4 The data clustering results are used to divide the study area; Figure 5 For application Figure 1 The diagram shows the global feature importance ranking of SHAP obtained by the method shown. Figure 6 For application Figure 1 A schematic diagram illustrating the impact of key features identified by the method and their interactions on landslide susceptibility; Figure 7 Based on Figure 1 The method shown ultimately generates a regional landslide susceptibility zoning map. Figure 8 This is a result diagram illustrating the simulation of regional landslide susceptibility under a new earthquake scenario using a feature mapping table. Figure 9 This is a diagram showing the results of simulating regional landslide susceptibility under a new earthquake scenario using simple parameter replacement. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Before the embodiment is described, the following terms are first explained: PGA: ground motion acceleration; Df: distance to fault; BSR: bottom simulating reflector; PHOD: potential hydrate hosting depth; GeoAge: stratigraphic age; Lithology: bottom lithology; AUC: area under curve; ROC: receiver operating characteristic curve; PR: precision-recall curve; XGBoost: extreme gradient boosting algorithm; LightGBM: light gradient boosting machine algorithm; CatBoost: categorical boosting algorithm; KNN: K-nearest neighbors algorithm; SVC: support vector machine classification algorithm; MLP: multi-layer perceptron algorithm; ExtraTrees: extremely random trees algorithm; GBDT: gradient boosting decision tree algorithm; Bagging: bootstrap aggregating algorithm.

[0020] The application provides a hydrate zone submarine landslide risk assessment method based on cascading machine learning, comprising: S1, data acquisition and preprocessing: in the data acquisition and preprocessing stage, the embodiment constructs a complete data set for supporting subsequent hydrate distribution prediction and landslide risk assessment through systematic data integration and feature construction. This step mainly collects historical landslide event data, hydrate hosting evidence data and multi-source environmental feature data including water depth, earthquake, fault, stratigraphic lithology, and performs abnormal value identification and missing value processing on the data.

[0021] S101, according to the difference of landslide performance form, the landslide is divided into unstable surface, unstable line and unstable point. According to the difference of different types of geological mechanism and influence range, 5 kilometers, 3 kilometers and 10 kilometers buffer zones are set. The purpose of this step is to realize the differentiation and quantification of the spatial influence range of landslide geological disasters, and solve the problems of rough spatial expression and insufficient physical basis caused by the use of single and fixed buffer zone radius in traditional evaluation. Provide analysis unit with clear geological significance and spatial logic for subsequent feature extraction and model training, and improve the spatial resolution and physical interpretability of landslide risk assessment.

[0022] S102, combine hydrate hosting data and water depth data to convert into a numerical parameter that can be directly used for model training, potential hydrate hosting depth; S103, ground motion key index quantification, based on multiple earthquake data to determine the key index of slope stability, and propose a peak ground acceleration PGA formula that comprehensively reflects the cumulative effect of multiple historical earthquakes: Wherein, represents the weight of the earthquake; a, b, c are regional geological attenuation coefficients; is the magnitude of the earthquake; is the distance from the target point to the earthquake.

[0023] The peak ground acceleration solution simulates the cumulative effect of regional long-term seismic load on slope stability by weighted superposition, ensures that the output is a scalar value that can be directly input into a machine learning model, and solves the adaptation problem between seismic load parameters and model input.

[0024] S104, based on the water depth data of the study area, calculate the slope, roughness and curvature topographic features; S105, obtain the distance to faults (Distance to faults, Df), stratum age (GeoAge) and lithology (Lithology) characteristics of the study area.

[0025] S2, hydrate occurrence depth prediction: based on local hydrate data and various environmental characteristics including seismic, fault, water depth, construct a hydrate occurrence dataset; use the dataset to train an improved neural network model, establish a hydrate occurrence depth prediction model; based on the model and the environmental characteristics of the entire study area, generate complete data of hydrate occurrence in the study area.

[0026] S201, randomly select several sample points in the hydrate occurrence data area, extract the hydrate occurrence depth of each point as the label, and extract the environmental feature data of the corresponding point to construct a hydrate occurrence dataset; divide the dataset according to the ratio of training set to test set 4:1.

[0027] S202, train the improved neural network regression model, and through iterative optimization, make the model fit, and construct a hydrate occurrence depth prediction model-physical information neural network model.

[0028] In the prior art, when predicting the occurrence depth of natural gas hydrate using a machine learning model, a standard multilayer perceptron model is usually used as a pure data-driven black box tool. The model only aims to find the statistical correlation between the input environmental characteristic variables and the output label, and completely ignores the physical boundary conditions that must be followed by the natural gas hydrate system, which are strictly defined by geophysical and thermodynamic theory. Such pure data-driven methods have significant limitations: ① The predicted results are physically inconsistent: the model may produce a large number of predicted values that violate basic physical common sense. For example, predicting hydrate occurrence in shallow water areas (water depth much less than 300 meters), or predicting hydrate occurrence depth greater than the local total water depth (i.e. hydrate "suspended" in seawater), these prediction results are physically infeasible, which seriously damages the practical value of the model. ② Poor model generalization ability and divergent results: due to the lack of guidance of physical laws, the model performs very unstably in areas where the training data coverage is insufficient or there is noise, and the prediction results show high divergence and discontinuity in spatial distribution.

[0029] This embodiment fundamentally modifies the standard machine learning model, proposing a loss function based on physical constraints. Prior knowledge of the hydrate stability domain is embedded into the training process as a penalty term, achieving a deep integration of data-driven and physics-driven approaches. Specifically, based on the mean squared error loss function, two physical constraint penalty terms are introduced to form a more targeted total loss function: , in, For data fitting loss, Water depth constraint weights: used to adjust the physical constraint penalty term for the lower limit of water depth. The larger the contribution of the model to the total loss function, the more strictly the model adheres to the physical rule that "the bottom boundary of the hydrate stability zone should not exceed the actual water depth" during training. Weights for hydrate burial depth constraints: Used to adjust the physical constraint penalty for burial depth rationality. The proportion of contribution to the total loss function. The larger the value, the more stringent the model's requirement for fitting the prior physical knowledge that "the possible distribution depth of hydrates should be within a reasonable burial depth range". This is a physical constraint penalty term for the lower limit of water depth. Its purpose is to force all prediction points in the model to be located at the critical water depth where hydrates are stable. The specific calculation formula is as follows: , in, This represents the number of samples in the batch. This is a safety margin, used to provide a buffer zone to enhance the robustness of the model and prevent predicted values ​​from getting too close to the theoretical boundary. Physically, it allows a certain safety margin above the actual water depth at the bottom of the hydrate stability zone, providing a safety space when the model predicts the water depth at the location. Below When this occurs, a penalty proportional to the square of the violation will be generated, thus strongly guiding the model's predicted output away from the physically invalid region. It is a rectified linear function that ensures a positive penalty value is only generated when the constraint is violated.

[0030] This is a physical constraint penalty term for the reasonable burial depth. It ensures that the predicted PHOD value does not exceed the local water depth, nor does it exceed a reasonable upper limit determined by geological common sense (e.g., set to α times the water depth, where α can be 0.3). The specific calculation formula is as follows: The first part of the formula penalizes cases where "PHOD is greater than water depth"; the second part penalizes unrealistic cases where "PHOD is too deep". The model predicts the first The hydrate distribution depth at each sample location is typically defined as the top boundary of the hydrate stability zone or a characteristic depth with a high probability of occurrence, and is one of the core outputs of the model. For the first The actual water depth at each sample location is a known input feature of the model. As a safety margin for the upper boundary, taking into account seabed topographic relief, data errors, and geological uncertainties, an absolute safety offset lower than the actual water depth is set for the possible distribution depth of hydrates. The lower boundary proportionality coefficient is usually set to 0-1. Based on the prior knowledge that hydrates are difficult to form in very shallow sediments, the minimum starting position of the possible distribution depth of hydrates is defined as a proportion of the actual water depth.

[0031] The above two penalties force all PHOD predictions of the model to satisfy: , This constraint essentially embeds the key geophysical fact that "hydrates are found within a specific depth window beneath the seabed," ensuring that the model's predictions not only statistically fit the training samples but also have reasonable spatial distribution characteristics in a physical sense, thereby significantly improving the geological reliability and engineering applicability of the model output.

[0032] In this embodiment, the neural network model not only learns the statistical laws of the data during the training process, but is also explicitly guided by the physical laws, thereby ensuring that the output results not only fit the observed data, but also strictly conform to the physical feasible region of the hydrate system, significantly improving the model's predictive rationality, spatial continuity and generalization ability in unknown regions.

[0033] S203. Based on the environmental characteristics of the entire region, the physical information neural network model is used to predict the hydrate occurrence depth of the entire study area and generate a distribution map.

[0034] S3. Using the predicted PHOD as a key feature, it is fused with other environmental features. A landslide dataset is constructed by combining this with historical landslide location data. A set of base prediction models is trained using multiple ensemble learning algorithms. Finally, a final landslide prediction model is constructed through an ensemble strategy based on dynamic weighting of geological landscape units. This method effectively overcomes the shortcomings of traditional static ensembles in failing to consider spatial heterogeneity. The specific steps are as follows: S301. Randomly select sample points within the study area, extract 9 features including PHOD for each sample point, construct a landslide dataset, and divide the training set and test set in a 4:1 ratio.

[0035] S302, initialize various machine learning models in Python, use the compute_sample_weight function to give higher weights to landslide samples, and use the sigmoid function to calibrate the model output probability to improve the reliability of probability prediction.

[0036] S303, dynamic weight integrated modeling based on geological landscape units. To solve the problem that the traditional integrated method uses global unified weight and cannot reflect the difference of landslide dominant mechanism in different geological units, the two-stage dynamic integrated modeling strategy of the embodiment is adopted. First, the study area is divided into geological landscape units with internal homogeneity and unit heterogeneity according to the geological environment characteristics, and then the optimal model and its integrated weight are determined independently in each unit.

[0037] S304, data-driven geological zoning of the study area based on the landslide data set, which provides a key basis for selecting the most suitable prediction model for different regions. In order to objectively determine the classification of the geological units of the study area and ensure that the divided units have high internal homogeneity and unit heterogeneity (which is crucial for subsequent allocation of different optimal models for different units), the second-order clustering method is adopted, and the silhouette coefficient is selected as the criterion for determining the optimal clustering number K. All environmental characteristics of all grid points in the study area are used as clustering variables. By calculating the average silhouette coefficient corresponding to different preset K values, the K value that makes the average silhouette coefficient maximum is selected as the optimal clustering number, so that the study area is scientifically divided into K geological landscape units . This method ensures that the samples in each unit are highly similar in environmental characteristics, while the sample characteristics of different units are significantly different, laying an ideal spatial foundation for subsequent differentiated model integration.

[0038] S305, unit-specific model weight calculation: in each geological landscape unit U k , train KNN, SVM, DecisionTree, ExtraTrees, RandomForest, GBDT, XGBoost, LightGBM, CatBoost and Bagging, a total of 10 common machine learning models. The following five key indicators are selected, and a comprehensive score is calculated for each model, as shown in Table 1: Table 1: Model evaluation indicators and corresponding weights selected by the embodiment of the present application

[0039] Note: Brier score needs to be processed in reverse: here we use 1-Brier to normalize to the same dimension. According to the above table, the comprehensive score formula is: , The three best-performing base models are selected based on the overall score, and the integration weights are allocated according to the score ratio, thereby ensuring that the integrated model can adaptively focus on the landslide mechanism that dominates the unit in different geological units.

[0040] S306, Dynamic Weighted Prediction Across the Entire Region: When predicting landslide susceptibility across the entire region, the system first determines the pixel points to be predicted. Geological landscape unit to which it belongs Then call the dedicated weight set corresponding to that unit. The predicted probabilities of the four base models are dynamically weighted and integrated to obtain the final landslide probability at that point. :

[0041] By traversing all pixels in the entire area, a seamless landslide susceptibility zoning map that fully considers the heterogeneity of geological space is finally generated.

[0042] S4. Landslide Susceptibility Prediction and Key Factor Analysis: Based on the landslide prediction model, the SHAP values ​​of each influencing factor are calculated to determine the main controlling factors of landslide deformation and their interaction effects. Simultaneously, based on various environmental characteristics of hydrate-bearing data, the landslide susceptibility distribution of the entire study area is predicted. The model can also simulate changes in landslide susceptibility after a specific event by adjusting environmental parameters. Step S4 includes: S401. Calculate the contribution of each feature to the model prediction results through SHAP analysis, and use the magnitude and positive / negative of the SHAP value to represent the degree and direction of influence. Visualize the importance and interaction of features with the help of summary plots and dependency plots.

[0043] S402. A Dynamic Risk Simulation Method Based on Event-Triggered Feature Weighting. Traditional static parameter replacement methods (when simulating the risk following a specific event (such as an earthquake), this method only replaces the feature values ​​corresponding to the new event in the model's input data, while the model's structure, parameters, and the weighting relationships of all features remain completely unchanged) have a fundamental limitation: the model's weighting for all features is fixed. For example, the model's sensitivity to the "slope" feature is the same regardless of whether an earthquake occurs. This clearly does not conform to geological principles—after an earthquake, the model should focus more on features related to earthquake dynamics (such as PGA and fault distance), relatively weakening the influence of some static terrain features. This invention introduces a dynamic feature weighting mechanism during the model inference stage, enabling the model to adaptively adjust its decision logic according to the event type, thereby achieving dynamic risk simulation that is more consistent with physical principles.

[0044] S4021. After the model training is completed, we construct an event type-feature weight mapping table offline. This table defines the weight adjustment coefficients of each input feature when different types of events (such as earthquakes, large-scale decomposition of hydrates) are triggered.

[0045] S4022. When simulating a specific event (such as an earthquake), recalculate the entire region's PGA field according to method S103. Then, look up the weight coefficients of each feature corresponding to the "strong earthquake" event from the mapping table. Before inputting the feature data into the final integrated model for prediction, perform real-time weighted preprocessing on the input feature vector.

[0046] S4023. Input the weighted new feature vector into the original model. Because the scale of the input features has changed, the decision-maker embedded in the model will naturally change the priority of its split points, thus effectively amplifying the effect of features strongly correlated with the event and suppressing the influence of weakly correlated features. The final output is a dynamic risk map that conforms to the event triggering mechanism.

[0047] The landslide prediction model of this invention has the ability to simulate changes in susceptibility after a specific event by calculating corresponding feature mapping tables and adjusting environmental parameters. For example, new earthquake event parameters can be input to quickly simulate and predict changes in seabed slope stability after the earthquake, generating a new landslide susceptibility zoning map. This functionality enables the method to provide strong technical support and decision support for safety monitoring, early disaster warning, and emergency response planning for marine engineering projects such as subsea pipelines and drilling platforms.

[0048] The present invention will be described below with specific examples. The present invention proposes a risk assessment method for submarine landslides in hydrate areas based on cascaded machine learning, and its implementation method is as follows: S1. Data Acquisition and Preprocessing. Multiple types of data were acquired from the North Atlantic study area, including historical landslide data, hydrate occurrence data, and environmental characteristic data such as earthquakes, faults, and water depth. Outliers and missing values ​​were cleaned and imputed. Specific steps are as follows: S101. Convert all collected data to the WGS1984_UTM_Zone0N coordinate system. Use the "Data Management Tools—Projection and Transformation—Define Projection" function in ArcGIS to complete the coordinate system unification.

[0049] S102. Import the landslide data into ArcGIS and set buffer zones of 5 km, 3 km, and 10 km according to the instability type (area, line, point) to represent areas with potential instability risk. This study area is located in the Western European waters of the North Atlantic, geographically ranging from longitude 36°W to 13°E and latitude 52°N to 80°N. The geological structure of this area is active and the sedimentary system is complex. Historically, several large-scale submarine landslides have occurred, including typical landslides such as Storegga, Sklimnadupet, Trendjupet, and Feni Ridge. With the advancement of marine geological survey and monitoring technologies, considerable research has been conducted on the formation mechanisms, triggering factors (such as seismic activity, natural gas hydrate decomposition, and rapid sedimentary loading) and potential geological hazard risks (such as tsunami potential) of landslides in this area.

[0050] S103. Based on seismic data, calculate the peak ground acceleration using the following formula:

[0051] in, Indicates the earthquake weight; a, b, and c are regional geological attenuation coefficients (the North Atlantic attenuation coefficients are taken as a=0.5, b=1.2, c=1); Magnitude; This represents the great circle distance between the target point and the seismic point. PGA calculations are performed point-by-point using Python, and the final results are as follows. Figure 2 As shown.

[0052] S104. Use the Euclidean distance tool in ArcGIS to calculate the distance Df from each point to the fault.

[0053] S105. Based on water depth data, calculate three terrain indicators: slope, roughness, and curvature in ArcGIS.

[0054] S106. The geological era is divided into 13 units (Quaternary to Precambrian) and numbered sequentially from 1 to 13.

[0055] S107. Lithology is divided into four categories: sedimentary rocks, sedimentary rocks, metamorphic rocks, and igneous rocks, and numbered 1 to 4 respectively.

[0056] S2. Hydrate Prediction Model Construction: To address the modeling challenges caused by the lack of PHOD data, a hydrate occurrence dataset is constructed based on local hydrate data and various environmental features, including earthquakes, faults, and water depth. An improved neural network model is trained using this dataset to establish a hydrate occurrence depth prediction model. Based on this model and the environmental characteristics of the entire study area, complete hydrate occurrence data for the study area is generated, thus providing comprehensive PHOD features for subsequent landslide prediction. The specific steps are as follows: S201. Select 2000 random points in the area with BSR data, and extract the hydrate occurrence depth as a label based on the water depth.

[0057] S202. Extract 8 environmental features from each random point and construct a hydrate prediction dataset.

[0058] S203. Randomly divide the dataset into training and test sets in a 4:1 ratio.

[0059] S204. Initialize the neural network regression model using the MLPRegressor function of sklearn.neural_network in Python, adopt the Adam optimization algorithm (initial learning rate 0.001), the maximum number of iterations is 200, and 5-fold cross-validation is used to optimize the model performance.

[0060] S205. The model validation results show that the coefficient of determination is 0.90 and the explained variance is 0.91. The predicted PHOD has good consistency with the measured data, indicating that the model has high accuracy.

[0061] S206. Based on the trained model, predict the PHOD of the entire study area, such as... Figure 3 As shown.

[0062] S207. It should be noted that PHOD is a theoretical index constructed to quantify the potential impact of hydrate stability zones on submarine slope stability. This parameter is introduced into the model as a dummy variable to assess the association between hydrate systems and landslide risk.

[0063] S301. 10,000 sample points were randomly selected in the study area, of which 2,258 were located within the historical landslide body (positive sample) and 7,742 were located outside the landslide body (negative sample). The positive sample accounted for 22.6%, which is a typical imbalanced data.

[0064] S302. Extract 9 features, including PHOD, for each sample point to construct a landslide dataset. Use a second-order clustering method for this landslide dataset, and select the silhouette coefficient as the optimal number of clusters, which is 2.

[0065] S303. After determining the optimal number of clusters to be 2, this study used K-means clustering to cluster the created landslide dataset. The clustering results are displayed in the PCA space. Figure 4 .

[0066] S304. Two datasets obtained based on clustering are Cluster_0 and Cluster_1, which are divided into training and test sets in a 4:1 ratio. Various machine learning models are initialized in Python. The `compute_sample_weight` function is used to assign higher weights to landslide samples, and the sigmoid function is used to calibrate the output probabilities. The results obtained based on training on Cluster_0 are shown in Table 2 below: Table 2 shows the training results of all models based on the research sub-region Cluster_0.

[0067] The results obtained from training based on Cluster_1 are shown in Table 3 below: Table 3 shows the training results of all models based on the research sub-region Cluster_1.

[0068] For each of the two clustering datasets, the top three scores are selected to construct a weighted ensemble model. The weights can be allocated according to the proportion of the overall score, as shown in Table 4. Table 4 shows the optimal model and weights selected for the two training sub-regions.

[0069] To verify the effectiveness of the proposed optimal geological landscape unit division method based on contour coefficient, this embodiment trains the model on the original dataset, and the results are shown in Table 5 below: Table 5 shows the training results of all models based on the entire study area.

[0070] Based on the comparative analysis of the experimental results in this embodiment, the model performance is significantly improved after adopting the optimal geological landscape unit division method based on the contour coefficient. On the two geological unit datasets obtained by clustering, the ensemble model shows a clear advantage over the traditional single model in all evaluation indicators.

[0071] Specifically, in terms of classification accuracy, the ensemble model trained on clustered data achieved an F1 score of 0.8796, an improvement of approximately 8.9 percentage points compared to the traditional method's 0.791, indicating that the model achieved a better balance between precision and recall. Regarding the model's discriminative ability, the area under the ROC curve reached a maximum of 0.986, an improvement of approximately 3.4% compared to the traditional best result of 0.954, demonstrating a further enhancement in the model's ability to distinguish between positive and negative samples.

[0072] Of particular note is the Brier score achieved by the clustering ensemble model, which reached a minimum of 0.0371 in probability calibration accuracy, a 42.3% reduction compared to the traditional method's 0.077. This indicates that the model's output probability values ​​are closer to actual observations, significantly improving prediction reliability. On the key business metric TRR@30, the clustering ensemble model achieved a perfect score of 1.000 on the Cluster_0 dataset, outperforming the traditional method's 0.921, demonstrating superior identification capabilities in predicting the top 30% of high-risk cases.

[0073] Furthermore, on the key metric of the area under the precision-recall curve in imbalanced classification tasks, the clustering ensemble model achieved a score of 0.932, an improvement of approximately 8.5% compared to the traditional method's 0.845, indicating a significant advantage in identifying positive samples. In terms of overall evaluation score, the clustering ensemble model reached a maximum of 0.946, an improvement of approximately 6.1% compared to the traditional method's 0.879, thus validating the effectiveness of the proposed method.

[0074] The above experimental results fully demonstrate that the optimal geological landscape unit division method based on the contour coefficient can effectively improve the predictive performance of the landslide susceptibility assessment model and provide a more reliable technical means for geological disaster risk assessment.

[0075] Furthermore, to verify the effectiveness of the PGA calculation method proposed in this invention, a comparison was made without considering seismic data. The evaluation metrics of the integrated model on all datasets are shown in Table 6 below: Table 6 shows the accuracy comparison of all models considering PGA.

[0076] Based on model training results and feature weight analysis, the peak ground acceleration formula constructed in this invention demonstrates significant effectiveness in improving the accuracy of submarine landslide risk assessment. From the model performance comparison, the comprehensive evaluation indicators of all machine learning models showed a systematic improvement after introducing the PGA feature: the average F1-Score increased from 0.657 to 0.681, an increase of 3.7%, indicating enhanced classification ability under imbalanced positive and negative samples; the average BrierScore decreased from 0.157 to 0.146, a decrease of 7.0%, reflecting a significant improvement in the model's probability prediction calibration; and the average TRAR@30 increased from 0.804 to 0.827, an increase of 2.9%, demonstrating that the model's ability to identify real landslides in the top 30% high-risk area was strengthened.

[0077] More importantly, in feature importance analysis, PGA demonstrated a dominant influence in several key models. In the random forest model, PGA became the feature with the highest weight, significantly exceeding traditional geological factors such as water depth and slope. In gradient boosting models such as XGBoost and LightGBM, PGA also ranked among the top three in feature importance. This weight distribution pattern confirms the core contribution of PGA to landslide prediction from the perspective of the inherent decision-making mechanism of machine learning, indicating that the model does indeed use the earthquake triggering mechanism as a key basis for judging landslide risk. This feature importance pattern is highly consistent with the geological mechanism—earthquakes, as the main triggering factor for submarine landslides, should dominate the risk assessment process through their dynamic effects quantified by PGA.

[0078] Validation results show that the ROC-AUC of landslide prediction improved by an average of 11.7% after the introduction of PHOD, confirming its effectiveness as a predictor and demonstrating the advantages of cascaded machine learning methods in this evaluation task.

[0079] The cascaded architecture employed in this invention fundamentally solves the modeling dilemma caused by sparse PHOD data. Traditional direct modeling methods face a choice: if only regions with PHOD data are used to train the model, the sample size is severely insufficient and it cannot predict PHOD-free regions; if PHOD features are abandoned, the key mechanism by which hydrates influence landslides is lost. This invention, through a cascaded approach that first predicts PHOD and then uses it for landslide assessment, fully utilizes the value of local PHOD data while achieving landslide risk assessment across the entire region.

[0080] S4. Landslide Susceptibility Mapping and Causal Interpretation. Based on the trained model, landslide susceptibility is predicted and mapped for the entire region. SHAP values ​​are used to analyze the contribution and interaction effects of various features on landslide risk, enhancing model interpretability. Dynamic changes in landslide susceptibility are simulated by adjusting specific environmental parameters (such as seismic activity). Specific steps are as follows: S401. Calculate the SHAP value of each feature using TreeExplainer from the SHAP library in Python, and analyze the feature contribution map (...). Figure 5 This demonstrates the weight and direction of influence of various environmental characteristics on landslide prediction.

[0081] S402. Using shap.dependence_plot, the contribution trends of PGA, fault, PHOD, water depth and slope to landslides under the influence of the most relevant variables were plotted, revealing the coupling effect between features. Figure 6 In the figure, a is the peak ground acceleration PGA, b is the distance from the fault Df, c is the potential depth of hydrates PHOD, d is the water depth, and e is the slope.

[0082] S403. Based on nine environmental characteristics of the entire region, a landslide susceptibility distribution map was calculated using a landslide prediction model. Figure 7 ).

[0083] S404. Based on the landslide probability (P), the susceptibility is divided into five levels: extremely high risk area (P>0.8): mainly distributed in the Storegga landslide area, FeniDrift sedimentary fan and Knipovich mid-ocean ridge, with a total area of ​​about 130,000 km²; high risk area (0.6≤P≤0.8): distributed in a ring around the extremely high risk area, with an area of ​​about 190,000 km², and the disaster chain expansion trend is significant; medium risk area (0.4≤P<0.6): scattered distribution, with a total area of ​​210,000 km²; (4) low risk area (0.2≤P<0.4): concentrated in Iceland and near-shore areas, with a total area of ​​340,000 km²; (5) extremely low risk area (P<0.2): widely developed in deep-sea basins and ocean floor, with a total area of ​​1,420,000 km², and the terrain in this area is gentle and the sedimentary environment is stable.

[0084] S405. Taking an earthquake event as an example, parameter control simulation is performed: The typical location of the JanMayen fault zone (70.931°N, 6.657°E) is selected as the epicenter to simulate an earthquake with a moment magnitude of 6.0 and a focal depth of 20km.

[0085] S406. Prepare 9-dimensional environmental test data to simulate events such as strong earthquakes (e.g., amplify the "earthquake" feature value by 2 times). Use the basic model to predict the event scenario results. Analyze the influence of features under the events by ranking their importance, calculate the event weight coefficients and limit them to the range of 0.3-3.0. Integrate the background state and event weights to generate a mapping table as shown in Table 7. Table 7 shows the dynamic characteristic weight mapping table for strong earthquake triggering.

[0086] Add the new earthquake data to the earthquake database and recalculate the PGA (step S103).

[0087] S407. Perform real-time weighted preprocessing on the input feature vector: Weighted feature vector = [water depth * 0.3, slope * 3, roughness * 0.3, curvature * 0.52, Df * 3, PGA * 0.53, Lithology * 0.3, GeoAge * 0.3, PHOD * 0.3]. Input the weighted feature vector into the model and recalculate the landslide susceptibility of the entire area (step S403).

[0088] S408, Comparison with the original ( Figure 7 ) and post-earthquake susceptibility map ( Figure 8 The study found that the high-risk zone induced by the earthquake (P>0.6) was mainly distributed south of the epicenter. This distribution was related to multiple factors, including distance from the fault (Df<100km), pHOD, water depth, and slope. This southward deviation phenomenon has important reference value for the seismic design of submarine engineering.

[0089] S409. To verify the advantages of using feature mapping tables in landslide susceptibility prediction, this embodiment conducted relevant verification work. Specifically, a landslide susceptibility map predicted through simple parameter substitution was generated (see...). Figure 9 ).

[0090] From a physical mechanism perspective, landslides are the result of the interaction of multiple geological factors. While earthquakes have a significant impact on landslides, they are not the sole determining factor. The slope determines the magnitude of the gravity component of the soil and rock mass along the slope; the steeper the slope, the stronger the tendency of the gravity component to cause the soil and rock mass to slide. Parameters such as Df may represent certain physical properties of the soil and rock mass, such as the internal friction coefficient and cohesion, which affect the soil and rock mass's ability to resist sliding. In actual geological environments, these factors interact and constrain each other.

[0091] Will Figure 9 The prediction map generated by introducing a feature mapping table (see...) Figure 8 By comparing them, it is clear that... Figure 9 Focusing solely on seismic variations while largely neglecting other crucial parameters such as slope and depth (Df) leads to an overestimation of landslide sensitivity in high-slope areas and near faults. This is because simple parameter substitution fails to fully consider the coupling relationships between various physical factors, merely amplifying the seismic effect and disrupting the equilibrium of factors in the actual geological environment.

[0092] The prediction method that introduces feature mapping tables can reasonably adjust the weights of each input feature based on different event types (such as earthquakes, large-scale decomposition of hydrates, etc.). It fully considers the changes in the actual influence of each physical factor under different geological events. For example, in earthquake events, it appropriately increases the weight of earthquake-related features to reflect the impact of earthquakes, while also taking into account the role of other factors, so that the weights of each factor conform to the interrelationships under the actual physical mechanism.

[0093] Further comparison of the results of the two prediction methods with historical landslide data in the area clearly reveals that the prediction method incorporating the feature mapping table shows a higher similarity to historical landslide areas. Therefore, this example fully demonstrates that the prediction method incorporating the feature mapping table, by adhering to physical mechanisms and comprehensively considering multiple geological factors, is significantly effective in predicting landslide susceptibility.

[0094] The embodiments of the present invention have been described in detail above, but the content described is only a preferred embodiment of the present invention and should not be considered as limiting the scope of the present invention. All equivalent changes and improvements made in accordance with the scope of the present invention should still fall within the patent coverage of the present invention.

Claims

1. A method for risk assessment of submarine landslides in hydrate zones based on cascaded machine learning, characterized in that, Includes the following steps: S1. Data Acquisition and Preprocessing: Acquire historical landslide data, hydrate occurrence data, and various environmental characteristic data of the study area, and process outliers and missing values ​​in the data; S2. Construction of hydrate prediction model: Based on local hydrate data and various environmental features, a hydrate occurrence dataset is constructed, a physical information neural network model is trained, the hydrate occurrence depth of the whole region is predicted, and a distribution map is generated. S3. Landslide Prediction Model Construction: The predicted hydrate occurrence depth is used as a key feature and combined with environmental feature data to construct an enhanced landslide dataset. Based on the geological landscape unit division and dynamic weight integration strategy, an ensemble learning model is trained to construct a landslide prediction model.

2. The method for risk assessment of submarine landslides in hydrate zones based on cascaded machine learning according to claim 1, characterized in that, It also includes the following steps: S4. Landslide Prediction: Based on the landslide prediction model, the contribution of each environmental feature is calculated and a landslide susceptibility zoning map is generated. The feature weights are dynamically adjusted through an event feature mapping table to achieve dynamic simulation of landslide risk for specific events.

3. The method for risk assessment of submarine landslides in hydrate zones based on cascaded machine learning according to claim 1, characterized in that, In step S3, the dynamic weight integration strategy based on geological landscape units includes: S301. The optimal number of clusters is determined by the contour coefficient, and the study area is divided into multiple geological landscape units. S302. Train multiple machine learning models within each geological landscape unit, select the optimal model based on the comprehensive score, and assign weights accordingly. S303. Based on the geological unit to which the point to be predicted belongs, the corresponding weights are called for dynamic weighted integration to generate a landslide susceptibility zoning map of the whole region.

4. The method for risk assessment of submarine landslides in hydrate zones based on cascaded machine learning according to claim 1, characterized in that, The total loss function of the physical information neural network model is: , in, For data fitting loss, As the water depth constraint weight, For the burial depth constraint weight, This is a physical constraint penalty term for the lower limit of water depth. To deepen the physical constraint penalty term for rationality, and: , in, This represents the batch sample size. It is a safety margin. It is a rectified linear function; , in, For the first The depth of hydrate distribution at each sample location For the first The actual water depth at each sample location For the upper boundary safety margin, This is the lower boundary scaling factor.

5. The method for risk assessment of submarine landslides in hydrate zones based on cascaded machine learning according to claim 2, characterized in that, In step S4, the event feature mapping table is a pre-constructed mapping relationship table that reflects the weight adjustment coefficients of various environmental features under different event types. When simulating a specific event, the input features are pre-processed in real time according to the mapping table and then input into the landslide prediction model for dynamic risk assessment.

6. The method for risk assessment of submarine landslides in hydrate zones based on cascaded machine learning according to claim 1, characterized in that, In step S1, the environmental characteristic data includes peak ground acceleration (PGA) determined based on seismic events, and: , in, Indicates earthquake weight; a, b, and c are regional geological attenuation coefficients; Magnitude; This represents the distance from the target point to the earthquake.

Citation Information

Patent Citations

  • A method for assessing the risk of submarine landslides during the exploitation of hydrates

    CN116258368B

  • Seabed landslide mass stability evaluation method based on multi-source geophysical data

    CN120820132A