Landslide susceptibility prediction method for reservoir areas without samples based on domain adaptive transfer learning

Through the domain adaptive transfer learning method, the source domain data feature alignment and unsupervised clustering are used to solve the landslide susceptibility prediction problem in the area without landslide sample library, and achieve more accurate and comprehensive prediction results.

CN115630336BActive Publication Date: 2025-09-19FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211343626.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-30
Publication Date
2025-09-19
Estimated Expiration
2042-10-30

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively predict landslide susceptibility in remote reservoir areas without landslide samples, mainly because the differences in data sets from different regions are not fully considered, resulting in the model's insignificant prediction ability in areas without samples.

Method used

A method based on domain adaptive transfer learning is adopted to conduct cross-regional susceptibility evaluation by utilizing historical landslide information in the source domain through feature alignment and unsupervised clustering, establish a feature transformation subspace, and generate a new dataset for prediction.

Benefits of technology

It improves the accuracy and comprehensiveness of landslide susceptibility prediction in sample-free reservoir areas, reduces modeling errors, overcomes the difficulties of data calibration and repeated modeling, and realizes unsupervised landslide susceptibility prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115630336B_ABST
    Figure CN115630336B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for predicting landslide susceptibility in sample-free reservoir areas based on domain adaptive transfer learning. The method comprises the following steps: S1, collecting multi-source data, determining source areas with sufficient samples, selecting universal evaluation indicators suitable for thematic susceptibility analysis, and performing indicator analysis; S2, determining unlabeled samples in the target domain, ensuring that the selected samples are representative, classifying them using a clustering method, and extracting the same number of samples from different classes; S3, using a feature-based domain adaptive method, adjusting the adaptive factor, and aligning the source domain data with the target domain unlabeled data; S4, selecting an appropriate machine learning model, using labeled samples in the source domain as a training set, predicting the susceptibility results of the target domain, and partitioning the susceptibility index using the natural break point method. The present invention addresses the difficulty of traditional methods in implementing landslide susceptibility evaluation in remote reservoir areas without samples, and provides a new approach for landslide susceptibility prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geological disaster prediction, and in particular to a method for predicting landslide susceptibility in a sample-free reservoir area based on domain adaptive transfer learning. Background Art

[0002] Since the 1980s, landslide susceptibility assessment has become a crucial tool for promoting landslide prevention, control, and disaster reduction efforts both domestically and internationally. However, when landslides occur in newly constructed reservoirs or in remote, small- or medium-sized reservoirs with no landslide history, environmental changes and data scarcity hinder the effectiveness of susceptibility assessment. Once a landslide occurs near the reservoir bank, the secondary damage can be irreparable. Providing landslide susceptibility predictions for remote reservoirs without landslide samples and assessing their potential is of vital importance.

[0003] Few susceptibility models can show satisfactory prediction results in sample-free areas. The root cause of this problem is that the susceptibility models based on machine learning ignore the differences between data sets in different regions based on the assumption that data are independent and identically distributed. The disaster-causing characteristics belonging to different study areas generally vary greatly. Finding the intrinsic connection between disaster-causing characteristics between different regions and overcoming the problem of data differences in factor sets between different regions are the fundamental ways to solve the problem that the susceptibility model has no significant predictive ability in sample-free areas. This invention combines the current hot spots in the development of artificial intelligence and innovatively introduces the concept of "transfer learning" to solve this problem.

[0004] Domain adaptation (DA), a branch of transfer learning, studies the problem of predicting target domain labels using a source domain dataset with the same characteristics but a different data distribution as the target domain, while maintaining the same task labels between two domains. This approach seeks to identify data feature connections between the source and target domains and leverages feature intersections for cross-domain learning, also known as feature-based domain adaptation transfer learning. Through feature-based domain adaptation transfer learning, a cross-regional landslide susceptibility assessment of the target domain is conducted using only historical landslide information from the source domain study area, providing a solution for unsupervised susceptibility prediction in areas without sample data. Summary of the Invention

[0005] The purpose of the present invention is to provide a landslide susceptibility prediction method for reservoir areas without samples based on domain adaptive transfer learning, which can provide landslide susceptibility prediction for remote reservoir areas without landslide samples.

[0006] To achieve the above object, the technical solution of the present invention is: a method for predicting landslide susceptibility in reservoir areas without samples based on domain adaptive transfer learning, comprising the following steps:

[0007] S1. Determine the scope of the study area and use the GIS platform to analyze the inundation area and catchment area upstream of the reservoir. The reservoir area with a predetermined amount of historical landslide data within the analysis range is used as the source area, and the reservoir area with less than the predetermined amount or no historical landslide data is used as the target area.

[0008] S2. Use remote sensing, field surveys, and spatial analysis to acquire multi-source data and identify universal disaster-causing indicators applicable to different study areas. Conduct factor analysis on the selected indicators, eliminate redundant and low-importance factors, and establish a thematic landslide prediction indicator system.

[0009] S3. Establish a representative sample set for the study area and perform preliminary susceptibility zoning of the study area using unsupervised clustering. Randomly extract non-landslide samples in the low and very low susceptibility zones in the source domain, which are equal to the number of historical landslide data. Randomly extract equal number of sample points in each susceptibility zone in the target domain.

[0010] S4. Analyze the factors that induce landslides in different regions, perform distribution adaptive adjustment using a feature-based domain adaptive transfer learning method, align the features of source domain data with the target domain unlabeled data, establish a feature transformation subspace, and generate new source and target domain datasets;

[0011] S5. Select a machine learning model, use source domain samples as training sets, and predict and divide the landslide susceptibility of the target domain.

[0012] In one embodiment of the present invention, step S1 takes the catchment area within the normal water level inundation range upstream of the barrage as the research scope, resamples the research area grid to ensure consistent resolution, and uses the grid unit as the basic unit for susceptibility evaluation.

[0013] In one embodiment of the present invention, the universal disaster-causing indicator factors of step S2 include: reservoir area terrain humidity index, normalized vegetation cover index, distance from reservoir water area, distance from reservoir area road, distance from geological boundary, reservoir bank stratum lithology, land use type and reservoir water inundation landslide elevation ratio; the remaining indicators are determined according to the special landslide type, and when the selected indicators are non-numerical variables, they need to be converted into virtual variables according to predetermined rules; the factor analysis adopts the Pearson correlation coefficient method combined with variance inflation factor and tolerance for analysis.

[0014] In one embodiment of the present invention, the clustering method of step S3 is K-PSO clustering. First, the improved PSO algorithm is used to find the optimal five initial cluster center points, and then the K-means algorithm is used to find the clustering results to preliminarily generate a susceptibility zoning map.

[0015] In one embodiment of the present invention, the feature-based domain adaptive transfer learning method in step S4 is a balanced adaptive distribution algorithm, and the specific steps of generating new source domain and target domain datasets are as follows:

[0016] S41, importing the source domain and target domain prediction indicator system obtained in step S2 into the source domain and target domain sample set established in step S3, and substituting into the feature-based domain adaptive transfer learning method;

[0017] S42. Calculate the initial maximum mean difference (MMD) of different data sets, adjust the data dimension, and find the optimal subspace that can align the data of the two domains. The MMD distance between the source domain and the target domain samples is expressed as:

[0018]

[0019] S43. Adjust the distribution adaptation factor to find the best fit ratio between the marginal distribution and the joint distribution of the data. Output the new source and target domain datasets after the distribution alignment. Verify that the MMD of the new dataset reaches the minimum value. The final optimization function obtained by simplifying the kernel method is:

[0020]

[0021] stA T XHX T A=I,0≤u≤1 (3)

[0022] Combine the above formulas to calculate the transformation matrix A, and finally obtain the new source domain and target domain samples after mapping.

[0023] In one embodiment of the present invention, the susceptibility interval in step S5 is divided into intervals by combining the natural break point method and the susceptibility index distribution law in step S3 using a fixed threshold method. The specific steps of the machine learning model for evaluating the susceptibility of the target domain area are as follows:

[0024] S51, using the principal component hazard factors after dimensionality reduction in the mapped new source domain sample data as model input features, and the known landslide and non-landslide classification results as output, and training a machine learning model classifier based on the mapped new source domain sample data;

[0025] S52, mapping the target domain global grid cell data to the aligned optimal subspace to generate a new mapped target domain global grid data set; using the principal component hazard factor after dimensionality reduction as a model input feature, and outputting a target domain landslide susceptibility index based on grid cells;

[0026] S53. Divide the landslide susceptibility index of the entire target area into five susceptibility intervals of extremely high, high, medium, low, and extremely low using the natural break point method, generate a susceptibility zoning map, and verify it with the susceptibility zoning map obtained in step S3, and use the fixed threshold method to delineate the final susceptibility interval.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] 1. Combine hydrological analysis methods to determine the impact range of reservoir bank landslides, avoid the generation of non-reservoir bank landslide samples, reduce susceptibility modeling errors, and improve sample accuracy.

[0029] 2. Using unsupervised clustering to extract landslide samples can more comprehensively and accurately reflect the essential characteristics of regional slopes with samples instead of the whole, reducing the sample uncertainty of susceptibility prediction.

[0030] 3. Using domain adaptive transfer learning methods to align sample features, using source domain sample modeling with a large amount of labeled data to achieve susceptibility prediction for areas without sample libraries, overcoming difficult problems such as landslide data calibration and repeated modeling, is a new technical concept for landslide susceptibility assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a technical roadmap for an embodiment of the present invention.

[0032] Figure 2 This is the visualization result of the sample feature subset distribution before and after the domain adaptive transfer learning method is adopted in the embodiment of the present invention.

[0033] Figure 3 This is a landslide susceptibility prediction map for the target area according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.

[0035] A specific embodiment of the present invention discloses a method for predicting landslide susceptibility in reservoir areas without samples based on domain adaptive transfer learning. The technical roadmap is as follows: Figure 1 As shown, the method includes the following steps:

[0036] S1. Taking Chitan Reservoir and Mianhuatan Reservoir in Fujian Province as the research objects, the inundation area and catchment area of ​​the reservoir upstream were analyzed using the GIS platform. The Chitan Reservoir area, which has sufficient historical landslide data within the analysis range, was used as the source area, and the Mianhuatan Reservoir area, which has only a small amount of historical landslide data or no historical landslide data, was used as the target area.

[0037] Specifically, the catchment area within the inundation range of the normal water level upstream of the reservoir dam was taken as the study area, and the study areas with different resolutions were resampled to ensure the same raster resolution. The raster unit was selected as the basic unit for susceptibility evaluation.

[0038] S2. Use remote sensing identification, field investigation and spatial analysis to obtain multi-source data and identify universal disaster-causing indicator factors belonging to different study areas; conduct factor analysis on the selected indicators, eliminate redundant and low-importance factors, and establish a special landslide prediction indicator system.

[0039] Specifically, based on the results of the factor survey, the universal disaster-causing indicator factors used in this example are as follows: reservoir slope gradient, reservoir slope aspect, terrain curvature, reservoir terrain wetness index, normalized vegetation cover index, distance from reservoir roads, distance from reservoir waters, distance from geological boundaries, reservoir bank lithology, land use type, and the ratio of landslide elevations inundated by reservoir water. Reservoir bank lithology and land use type are non-numerical variables. The contribution of each category to landslide development was calculated and converted into dummy variables in the form of an exponential range of 0-1. A factor collinearity analysis was conducted using the Pearson correlation coefficient combined with the variance inflation factor and tolerance. The two highly collinear factors, terrain curvature and normalized vegetation cover index, were eliminated. Table 1 shows the contribution of landslide susceptibility in each interval determined using frequency ratio analysis.

[0040] Table 1 Analysis of the contribution of frequency ratio of source domain disaster-causing index factors

[0041]

[0042]

[0043] S3. Establish a representative sample set for the study area and perform preliminary susceptibility zoning of the study area using unsupervised clustering. Randomly extract non-landslide samples equal to the landslide data in the low and very low susceptibility zones in the source domain. Randomly extract equal sample points in each susceptibility zone in the target domain.

[0044] Specifically, this example divides susceptibility into five intervals: extremely high, high, medium, low, and extremely low. The unsupervised K-PSO clustering algorithm is applied to the source domain and target domain respectively. The improved PSO algorithm is first used to find the optimal five initial cluster centers. Then the K-means algorithm is used to find the clustering results. The susceptibility zoning maps of the source domain and target domain are generated respectively. On this basis, sample points are extracted to establish a landslide dataset.

[0045] S4. Analyze the factors that induce landslides in different regions, perform distribution adaptive adjustment using a feature-based domain adaptive transfer learning method, align the features of the source domain data with the target domain unlabeled data, establish a feature transformation subspace, and generate new source domain and target domain datasets.

[0046] Specifically, the feature-based domain adaptive transfer learning method is a balanced adaptive distribution algorithm, which generates new source domain and target domain datasets in the following steps:

[0047] S41, importing the source domain and target domain prediction indicator system obtained in step S2 into the source domain and target domain sample set established in step S3, and substituting it into the feature-based domain adaptive transfer learning method;

[0048] S42. Calculate the initial maximum mean discrepancy (MMD) of different data sets, adjust the data dimension, and find the best subspace that can align the data of the two domains. The MMD distance between the source domain and the target domain samples can be expressed as

[0049]

[0050] S43. Adjust the distribution adaptation factor to find the best fit ratio between the marginal distribution and the joint distribution of the data. Output the new source and target domain datasets after the distribution alignment. Verify that the MMD of the new dataset reaches the minimum value. The final optimization function obtained by simplifying the kernel method is:

[0051]

[0052] stA T XHX T A=I,0≤u≤1 (3)

[0053] Combine the above formula to calculate the transformation matrix A, and finally get the new source domain and target domain samples after mapping. Figure 2 shown.

[0054] S5: Select a suitable machine learning model, use source domain samples as training sets, and predict and divide the landslide susceptibility of the target domain.

[0055] Specifically, the susceptibility interval is divided by combining the natural break point method with the susceptibility index distribution law in step 3, and the fixed threshold method is used. The specific steps of the selected machine learning model to evaluate the susceptibility of the target domain area are as follows:

[0056] S51, using the principal component hazard factors after dimensionality reduction in the mapped new source domain sample data as model input features, and the known landslide and non-landslide classification results as output, and training a machine learning model classifier based on the mapped new source domain sample data;

[0057] S52. Map the target domain global grid cell data to the aligned optimal subspace to generate a new mapped target domain global grid data set. Use the principal component hazard factors after dimensionality reduction as model input features and output a target domain landslide susceptibility index based on grid cells.

[0058] S53. Divide the landslide susceptibility index of the entire target area into five susceptibility intervals: very high, high, medium, low, and very low, using the natural break point method. Generate a susceptibility zoning map, and verify it with the susceptibility zoning map obtained in step 3. Use the fixed threshold method to delineate the final susceptibility interval. The results are as follows: Figure 3 shown.

[0059] The above description is merely a preferred embodiment of the present invention and is intended only to help understand the method and core concept of the present invention. The scope of protection of the present invention is not limited to the above embodiment. It is understood that other improvements and variations directly derived or imagined by those skilled in the art without departing from the spirit and concept of the present invention should be considered to be included in the scope of protection of the present invention.

Claims

1. A landslide susceptibility prediction method for reservoir areas without samples based on domain adaptive transfer learning, characterized by: The steps include: S1. Determine the scope of the study area and use the GIS platform to analyze the inundation area and catchment area upstream of the reservoir. The reservoir area with a predetermined amount of historical landslide data within the analysis range is used as the source area, and the reservoir area with less than the predetermined amount or no historical landslide data is used as the target area. S2. Use methods including remote sensing identification, field surveys, and spatial analysis to obtain multi-source data and identify universal disaster-causing indicators belonging to different study areas. These include: reservoir terrain moisture index, normalized vegetation cover index, distance to reservoir waters, distance to reservoir roads, distance to geological boundaries, reservoir bank stratum lithology, land use type, and the ratio of reservoir water-inundated landslide elevation. Other indicators are determined based on the thematic landslide type. If the selected indicators are non-numerical variables, they must be converted into dummy variables according to predetermined rules. The selected indicators are analyzed using the Pearson correlation coefficient method combined with the variance inflation factor and tolerance, and redundant and low-importance factors are eliminated to establish a thematic landslide prediction indicator system. S3. Establish a representative sample set for the study area and perform preliminary susceptibility zoning of the study area using unsupervised clustering. Randomly extract non-landslide samples in the low and very low susceptibility zones in the source domain, which are equal to the number of historical landslide data. Randomly extract equal number of sample points in each susceptibility zone in the target domain. S4. Analyze the factors that induce landslides in different regions, perform distribution adaptive adjustment using a feature-based domain adaptive transfer learning method, align the features of the source domain data with the target domain unlabeled data, establish a feature transformation subspace, and generate new source and target domain datasets. The specific steps are as follows: S41, importing the source domain and target domain prediction indicator system obtained in step S2 into the source domain and target domain sample set established in step S3, and substituting into the feature-based domain adaptive transfer learning method; S42. Calculate the initial maximum mean difference (MMD) of different data sets, adjust the data dimension, and find the optimal subspace that can align the data of the two domains. The MMD distance between the source domain and the target domain samples is expressed as: S43. Adjust the distribution adaptation factor to find the best fit ratio between the marginal distribution and joint distribution of the data. Output the new source and target domain datasets after distribution alignment. Verify that the MMD of the new dataset reaches the minimum value. The final optimization function obtained by simplifying the kernel method is: s.t. A T XHX T A=I, 0≤u≤1 (3) Combine the above formula to calculate the transformation matrix A, and finally obtain the new source domain and target domain samples after mapping; S5. Select a machine learning model, use source domain samples as a training set, and predict and divide the landslide susceptibility of the target domain. The susceptibility interval is divided into intervals using a fixed threshold method, combining the natural break point method and the susceptibility index distribution law of step S3. The specific steps of the machine learning model in evaluating the susceptibility of the target domain are as follows: S51, using the principal component hazard factors after dimensionality reduction in the mapped new source domain sample data as model input features, and the known landslide and non-landslide classification results as output, and training a machine learning model classifier based on the mapped new source domain sample data; S52, mapping the target domain global grid cell data to the aligned optimal subspace to generate a new mapped target domain global grid data set; using the principal component hazard factor after dimensionality reduction as a model input feature, and outputting a target domain landslide susceptibility index based on grid cells; S53. Divide the landslide susceptibility index of the entire target area into five susceptibility intervals of extremely high, high, medium, low, and extremely low using the natural break point method, generate a susceptibility zoning map, and verify it with the susceptibility zoning map obtained in step S3, and use the fixed threshold method to delineate the final susceptibility interval.

2. The method for predicting landslide susceptibility in reservoir areas without samples based on domain adaptive transfer learning according to claim 1 is characterized in that: In step S1, the catchment area within the normal water level inundation range upstream of the barrage is used as the research scope, the grid of the research area is resampled to ensure consistent resolution, and the grid unit is used as the basic unit for susceptibility evaluation.

3. The method for predicting landslide susceptibility in reservoir areas without samples based on domain adaptive transfer learning according to claim 1 is characterized in that: The clustering method in step S3 is K-PSO clustering. First, the improved PSO algorithm is used to find the optimal five initial cluster center points, and then the K-means algorithm is used to find the clustering results to preliminarily generate a susceptibility zoning map.