Landslide risk prediction method and system based on landslide data set expansion

By augmenting landslide samples, a few types of samples are expanded using BFOA-KSMOTE oversampling module and generative adversarial network model, the problem of sample imbalance in landslide problems is solved, and the training efficiency and prediction accuracy of landslide prediction models are improved.

CN120032169APending Publication Date: 2025-05-23NORTHWEST RES INST CO LTD OF C R E C +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510110855.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art encounters the problem of sample imbalance and small sample size in landslide problems, resulting in too long training time for machine learning and deep learning models, low prediction accuracy, and unstable convergence.

Method used

By augmenting the landslide samples, augmenting a few class samples using the BFOA-KSMOTE oversampling module and a generative adversarial network model, and a new minority class samples are generated to balance the data set.

Benefits of technology

Through data expansion technology, the problem of sample imbalance is solved, the training efficiency and prediction accuracy of the landslide prediction model are improved, and more reliable landslide risk prediction is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032169A_ABST
    Figure CN120032169A_ABST
Patent Text Reader

Abstract

The invention provides a landslide risk prediction method and system based on landslide data set expansion, and relates to the technical field of geological disaster recognition, and the method comprises the steps: obtaining an original landslide sample; screening landslide influence factors, and reclassifying the original landslide samples; selecting sample points; making one-dimensional array data according to the sample points, and constructing a data set; expanding the data of the data set by adopting a data expansion network; training a landslide prediction model by adopting the expanded data set; deploying the trained landslide prediction model, and carrying out landslide risk prediction; wherein the data expansion network comprises a BFOA-KSMOTE oversampling module, a K-means clustering algorithm is used for clustering minority classes, edge samples are removed, and then new minority class samples are generated through the synthesis minority class oversampling technology so as to balance a data set. The problem of unbalanced landslide samples can be solved, and the landslide prediction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geological disaster identification, and in particular to a landslide risk prediction method and system based on landslide data set expansion. Background Art

[0002] Landslide is a process in which geological structures undergo severe deformation. It is a geological process in which a geological structure slides downward in the direction of a slope under the influence of multiple common factors such as rainfall. As one of the main geological disasters, landslides can cause a large number of casualties and property losses in severe cases. With the development of computer technology, fully automatic and semi-automatic extraction of landslide information through remote sensing technology after landslide disasters has gradually become a new research direction. The image is segmented into landslide or semi-landslide areas based on color, texture, etc. With the development of computer science, various machine learning methods such as artificial neural networks (ANN), decision trees (DT), support vector machines (SVM) and random forests (RF) have been applied to the field of landslides. With the support of sufficient learning samples, artificial intelligence algorithms can be used to establish relationships between identified landslides and influencing factors to predict the spatial probability of potential landslides.

[0003] At present, machine learning and deep learning algorithms will encounter problems such as sample imbalance and small sample size when dealing with landslide problems. Landslide samples usually only occupy a small part of the study area, and most of the study area is non-landslide area. The number of landslide data and non-landslide data is unbalanced, and the number of landslide samples is small. This situation will lead to long training time for machine learning and deep learning models, low prediction accuracy, unstable convergence, etc. Summary of the invention

[0004] In order to address the deficiencies of the prior art, the purpose of the present invention is to provide a landslide risk prediction method and system based on landslide data set expansion, which solves the problem of landslide sample imbalance by expanding landslide samples, thereby improving the accuracy of landslide prediction.

[0005] To achieve the above object, according to some embodiments, a first aspect of the present invention provides a landslide risk prediction method based on landslide data set expansion, comprising:

[0006] Obtaining original landslide samples;

[0007] Screen the landslide influencing factors and reclassify the original landslide samples;

[0008] Selecting sample points, wherein the sample points include positive sample points and negative sample points, and both the positive sample points and the negative sample points include all landslide influencing factors after reclassification;

[0009] Produce one-dimensional array data according to sample points and construct a data set;

[0010] Use data expansion network to expand the data set data;

[0011] The expanded dataset is used to train the landslide prediction model;

[0012] Deploy the trained landslide prediction model to predict landslide risks;

[0013] The data augmentation network includes a BFOA-KSMOTE oversampling module, which uses the K-means clustering algorithm to cluster the minority classes and remove marginal samples, and then generates new minority class samples through the synthetic minority class oversampling technology to balance the data set.

[0014] A second aspect of the present invention provides a landslide risk prediction system based on landslide data set expansion, comprising:

[0015] A data acquisition module, configured to acquire original landslide samples;

[0016] The landslide impact factor screening module is configured to screen the landslide impact factors and reclassify the original landslide samples;

[0017] A sample point selection module is configured to select sample points, wherein the sample points include positive sample points and negative sample points, and both the positive sample points and the negative sample points include all landslide influencing factors after reclassification;

[0018] A data set construction module is configured to produce one-dimensional array data according to sample points to construct a data set;

[0019] A data augmentation module is configured to augment the data set data using a data augmentation network; wherein the data augmentation network includes a BFOA-KSMOTE oversampling module, the BFOA-KSMOTE oversampling module uses a K-means clustering algorithm to cluster the minority class and remove marginal samples, and then generates new minority class samples through a synthetic minority class oversampling technique to balance the data set;

[0020] A training module configured to train a landslide prediction model using the expanded data set;

[0021] The prediction module is configured to deploy the trained landslide prediction model to perform landslide risk prediction.

[0022] A third aspect of the present invention provides an electronic device, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to complete the steps of the above-mentioned landslide risk prediction method based on landslide data set expansion.

[0023] According to a fourth aspect of the present invention, a computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, complete the steps of the above-mentioned landslide risk prediction method based on landslide data set expansion.

[0024] A fifth aspect of the present invention provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned landslide risk prediction method based on landslide data set expansion.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] The present invention provides a landslide risk prediction method and system based on landslide data set expansion. A BFOA-KSMOTE oversampling module is added to the network construction. A K-means clustering algorithm and a synthetic minority class oversampling technique are used to generate new minority class samples to balance the data set. The K-means clustering algorithm is used to cluster the minority classes. At the same time, an abnormal tolerance strategy is adopted to remove marginal samples, thereby eliminating the influence of marginal samples and realizing data expansion of minority class landslide samples.

[0027] The present invention screens landslide influencing factors, reclassifies the original data of several landslide influencing factors, and takes into account the limited existing landslide inventory data, the imbalance between classes in the data set, and the sensitivity to the distribution of marginal samples, which leads to the problems of low quality of generated samples and low target recognition accuracy. A generative adversarial network model is used to expand the minority class samples to solve the problem of sample imbalance.

[0028] The present invention adds a mixed extraction module of high and low prone areas in network construction, which improves the success rate of landslide samples being classified as very high landslide susceptibility, compensates for the subjectivity of landslide delineation, compensates for the randomness of non-landslide samples, provides a more reliable reference for difficult-to-identify landslides, solves the problem of low recognition accuracy, and improves the accuracy of landslide prediction.

[0029] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0031] Figure 1 is a flow chart of the method in Embodiment 1 of the present invention;

[0032] Figure 2 It is a schematic diagram of the processing flow of the high and low prone area mixed extraction module;

[0033] Figure 3 Schematic diagram of the processing flow of the BFOA-KSMOTE oversampling module. DETAILED DESCRIPTION

[0034] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0035] Embodiment 1

[0036] Embodiment 1 of the present invention, as Figures 1 - 3 As shown, a landslide risk prediction method based on landslide dataset expansion is provided, including:

[0037] S1. Obtain the original landslide sample;

[0038] S2, screening landslide influencing factors and reclassifying the original landslide samples;

[0039] S3, selecting sample points, wherein the sample points include positive sample points and negative sample points, and both the positive sample points and the negative sample points include all landslide influencing factors after reclassification;

[0040] S4, making one-dimensional array data according to the sample points and constructing a data set;

[0041] S5, using data expansion network to expand the data set data;

[0042] S6, using the expanded data set to train the landslide prediction model;

[0043] S7, deploy the trained landslide prediction model to predict landslide risks;

[0044] The data augmentation network includes a BFOA-KSMOTE oversampling module, which uses the K-means clustering algorithm to cluster the minority classes and remove marginal samples, and then generates new minority class samples through the synthetic minority class oversampling technology to balance the data set.

[0045] The present invention combines the adversarial network data expansion strategy, the high- and low-prone area mixed extraction strategy, and the BFOA-KSMOTE oversampling technology to propose a landslide risk prediction method based on the expansion of the landslide data set. The landslide influencing factors are screened, and the original data of the screened landslide influencing factors are reclassified to ensure that there is no or low correlation between the various landslide influencing factors. After selecting the indicators, positive samples are extracted near the disaster point in the landslide historical disaster map and the upper section of the landslide historical disaster coordinate map, and negative samples are extracted near the place where the landslide did not occur. Then the reclassified landslide influencing factors and landslide labels of the sample points are input into ENVI5.3 for layer superposition to obtain an image cube, and the pixels at the same position of the image cube are extracted as one-dimensional array data. According to the amplification requirements of the minority class samples, the one-dimensional array data is divided into a training set X and a training set according to a ratio of 6:4. 1 And the test set T 1 . The training set X 1 The landslide data in the image cube is input into the network. The landslide features are one-dimensional array data generated by the image cube. The optimized generative adversarial network model B-KSmo-HV-WGAN model based on the convolutional neural network is constructed. The generative adversarial network (WGAN) data expansion strategy, high and low prone area mixed extraction strategy, and BFOA-KSMOTE oversampling technology are used to create additional landslide samples. Finally, the balanced data set is input into the back-end machine learning classification model for training, and the test set is tested using the optimal model generated by the training to obtain the landslide prediction results.

[0046] In step S2, the pre-processed multi-source data of the study area needs to be input into the ArcMap 10.6 platform to extract landslide influencing factors. The landslide influencing factors include: extracting topographic and hydrological factors based on the global digital elevation model DEM, including elevation, slope, aspect, topography, slope position, slope variability, aspect variability, profile curvature, plane curvature, land use type, terrain undulation, surface cutting depth, terrain moisture index, elevation variation coefficient, surface roughness, and distance from the river; extracting geomorphic factors based on the regional geological map, including distance from the fault, lithology, etc.

[0047] The landslide impact factors were input into the R language platform to calculate the Pearson correlation index, and the factors with correlation index exceeding 0.7 were selectively eliminated; the landslide impact factors were input into SPSS for multicollinearity analysis to determine the collinearity between factors; the factors with Pearson correlation index exceeding 0.7 and the factors with collinearity problems in SPSS were merged, and then these factors were eliminated. Finally, the original data of several screened landslide impact factors were updated and classified.

[0048] Specifically, in this embodiment, first, the Pearson correlation index is calculated in R language and a preliminary screening is performed to find factor pairs with a correlation index exceeding 0.7, and the screened factor pairs with a correlation index exceeding 0.7 are preliminarily eliminated; secondly, multicollinearity analysis is performed in SPSS, and after the factors are imported into SPSS, collinearity diagnosis is performed using the software's built-in function, and the "tolerance" and "variance inflation factor (VIF)" in the results are used to determine whether there is a collinearity problem. Generally speaking, a tolerance less than 0.1 or a VIF greater than 10 indicates a serious collinearity problem and the factor needs to be eliminated; combined with the results of the Pearson correlation index calculation and collinearity analysis, the screened factors are eliminated to obtain the screened data.

[0049] In step S3, the positive samples are samples where landslides occur, and the negative samples are samples where landslides do not occur, and both samples contain all landslide influencing factors after the reclassification.

[0050] In this embodiment, the sample point selection method includes: obtaining the landslide historical disaster map and the landslide historical disaster coordinate map of the study area through ArcGis software, and selecting to extract positive samples near the landslide site in a section of the map, and extracting negative samples near the site where the landslide did not occur.

[0051] In step S4, the reclassified landslide impact factors and landslide labels of the sample points are input into ENVI 5.3 for layer overlay to generate an h×w×c image cube, where h is the length of each landslide impact factor, w is the width, and c is the number of channels of the image cube, specifically the number of landslide impact factors. Then, the pixels at the same position of the image cube are extracted as one-dimensional array data, and the one-dimensional array data is divided into a training set X and a training set X according to the requirements of minority class amplification in a ratio of 6:4. 1 And the test set T 1 , the amount of data is less at this time.

[0052] In step S5, the training set X 1 The landslide data in the image cube are input into the data expansion network, and the landslide features are one-dimensional array data generated by the image cube. The optimized generative adversarial network model B-KSmo-HV-WGAN model based on convolutional neural network is constructed, that is, the data expansion network, to achieve data expansion.

[0053] The B-KSmo-HV-WGAN model uses the generative adversarial network (GAN) data augmentation strategy, the mixed extraction strategy of high and low prone areas, and the BFOA-KSMOTE oversampling technique to create additional landslide samples.

[0054] The optimized generative adversarial network model B-KSmo-HV-WGAN based on convolutional neural network is constructed according to the landslide characteristics, including: high and low prone area mixed extraction module, BFOA-KSMOTE oversampling module and WGAN module. First, the landslide data is extracted using the high and low prone area mixed extraction module to extract landslide samples and non-landslide samples, and the samples are added to the training set X in a ratio of 6:4. 1 And the test set T 1 Then, the BFOA-KSMOTE oversampling module uses the K-means clustering algorithm and the synthetic minority class oversampling technique to generate new minority class samples to balance the data set. In data expansion, there is usually a problem of being sensitive to the distribution of marginal samples, resulting in low quality of generated samples. The K-means clustering algorithm is used to cluster the minority classes first, and the abnormal tolerance strategy is used to remove marginal samples to eliminate the influence of marginal samples. The WGAN module includes a data generator G that simulates real data and a discriminator D that determines the authenticity of the data.

[0055] In the mixed extraction module of high and low prone areas, all data were imported into the Scikit-Learn library in Python, and the data were analyzed using RF (Random Forest) to establish a random forest model. After the random forest model was established, the raster image of the study area was converted into vector points using the "raster to point" GIS tool; the information value of the landslide influencing factor was extracted using the "extract multi-value to point" GIS tool. The data was imported into the established model to calculate the landslide susceptibility score. Finally, the landslide susceptibility map was obtained using the "point to raster" GIS tool. The results were imported into GIS, and the landslide susceptibility was divided into 4 levels using the natural breakpoint method: low, medium, high, and very high. Then, samples that were not determined to be landslides in the original data set were randomly selected in the very high and high landslide susceptibility areas, and non-landslide sample data were also extracted from the low landslide susceptibility area, and the samples were added to the training set and the test set in a ratio of 6:4.

[0056] In the BFOA-KSMOTE oversampling module, an improved fruit fly optimization algorithm is used to optimize the K-means clustering algorithm, solving the problem of the initial cluster center being sensitive and easy to fall into local optimality in the algorithm; an improved clustering tolerance strategy is adopted to optimize the SMOTE oversampling technology, solving the marginalization problem of synthesizing new samples.

[0057] As an alternative implementation, the basic process of BFOA-KSMOTE is as follows:

[0058] S51: Initialize the fruit fly population size sizepop, the maximum number of dimension iterations maxgen, use F distribution to define the fruit fly flight radius L and the location of the initialized fruit fly population, and give the number of clusters k, sample abnormality tolerance Endure, sample abnormality degree Degree, and oversampling multiple N;

[0059] S52: Use formula (1.1) to redefine the position information of individual fruit flies, compare their optimal values ​​to define the elite group, and at the same time, each individual fruit fly starts to update its own position according to the assigned flight radius and direction;

[0060]

[0061] Where x represents the sample point, C i represents the i-th cluster, k represents the number of clusters divided, Cluster C i The cluster center is calculated as follows:

[0062]

[0063] S53: Since the direction and distance of the initial food source are unknown, the distance of each fruit fly from the origin is calculated first, and the concentration value of the food source is determined according to formula (1.6), and the taste concentration value Smell of each fruit fly is calculated according to formula (1.7). i ;

[0064]

[0065]

[0066]

[0067]

[0068] Smell i =fitness(S i ) (1.7)

[0069] Among them, X in the above formula axis , Y axis represents the initial population position information, rand represents the random distribution number, X i , Y i is the updated coordinate information, i represents the i-th fruit fly individual, and fitness represents the judgment function of the taste concentration value, that is, the fitness function.

[0070] S54: Record the position information of the fruit fly with the best smell concentration value in the current group, and store its specific information in the optimal group Smellbest ;

[0071] S55: According to the recorded position information of the best individual bestSmell, use it as the initial cluster center of clustering, and perform the following operations: Calculate the cluster center to which each object in the data set belongs according to the Euclidean distance using formula (1.8); Recalculate the cluster center of each cluster according to formula (1.2) based on the objects in the assigned cluster; Record the distance between each sample in the cluster and the cluster center. If the distance is greater than the average distance between the sample and the cluster center, Degree++; If Degree>Endure, remove the sample point and no longer participate in clustering; Record the position information and clustering results of each cluster center, and re-use the cluster center position as an individual of the fruit fly population. According to the optimal taste concentration value and the recorded position information, each individual of the fruit fly population flies in this direction;

[0072]

[0073] Among them, dist(x, y) represents the Euclidean distance of each sample, X i , Y i are the coordinates of the object, and n represents the total number of objects in each cluster;

[0074] S56: Determine whether the number of iterations maxgen is reached, if not, repeat steps S52 to S56;

[0075] S57: For the clustering results of the minority class, for the minority class samples in the cluster, the distance of the sample set is calculated using the Euclidean distance to obtain the k nearest neighbors;

[0076] S58: According to the oversampling multiple N, for the minority class sample x, randomly select samples from its k nearest neighbors;

[0077] S59: For the selected k nearest neighbor samples, use formula (1.9) and the original samples to construct new samples;

[0078]

[0079] Where: X i ∈R d , X new represents the synthesized new sample, X i,att is the i-th sample in the minority class sample, att∈(1, 2, ..., d), represents the attribute value, It represents a random number between 0 and 1. ij Represents sample X i The jth nearest neighbor of .

[0080] S510: Use an anomaly classifier to perform anomaly detection on the oversampled data set.

[0081] In some embodiments, in the WGAN module, the Wassertein distance, a probability distribution difference indicator, is used to construct a loss function Loss that can judge true and false data, and calculate the distance between the pseudo data generated by the generator G and the real data distribution. This indicator can effectively calculate the difference between true and false data, and add a gradient penalty function to solve the gradient vanishing problem.

[0082] The training of generator G includes constructing a normally distributed noise matrix Noise according to the network batch and landslide, and inputting it into generator G. Generator G is constructed with a convolutional neural network as the underlying network. Discriminator D inputs the pseudo data generated by generator G and the real landslide data into discriminator D. The discriminator uses a one-dimensional convolution layer and a maximum pooling layer to learn the distribution of true and false data, and adds a Softmax function to distinguish true and false labels; then, according to the Loss value, the parameters of generator G and discriminator D are updated alternately using the random gradient descent method until the maximum number of training times is completed and the generative adversarial network training is ended. Then, the noise matrix Noise with the same feature dimension as the landslide is input into generator G, and the landslide amplification data X is generated according to the difference in the number of landslide samples and non-landslide samples. k , and add it to the original training set X 1 Construct a balanced dataset X p (i.e. the expanded training set) for model training.

[0083] In step S6, the balanced data set is input into a backend machine learning classification model for training, and the test set is tested using the optimal model generated by the training to obtain a landslide prediction result. In some embodiments, the backend machine learning classification model is a support vector machine.

[0084] After obtaining the landslide prediction results, the results were compared with the field-verified landslide labels, and the model prediction performance was evaluated based on some accuracy indicators (AUC, G-mean, F1 score, Kappa coefficient, and Matthews correlation coefficient).

[0085] Embodiment 2

[0086] This embodiment provides a landslide risk prediction system based on landslide data set expansion, including:

[0087] A data acquisition module, configured to acquire original landslide samples;

[0088] The landslide impact factor screening module is configured to screen the landslide impact factors and reclassify the original landslide samples;

[0089] A sample point selection module is configured to select sample points, wherein the sample points include positive sample points and negative sample points, and both the positive sample points and the negative sample points include all landslide influencing factors after reclassification;

[0090] A data set construction module is configured to make one-dimensional array data according to sample points and divide it into a training set and a test set;

[0091] A data augmentation module is configured to perform data augmentation on the training set data using a data augmentation network; wherein the data augmentation network includes a BFOA-KSMOTE oversampling module, the BFOA-KSMOTE oversampling module uses a K-means clustering algorithm to cluster the minority class and remove marginal samples, and then generates new minority class samples through a synthetic minority class oversampling technique to balance the data set;

[0092] A training module is configured to train the landslide prediction model using the expanded training set and evaluate the landslide prediction model using the test set data;

[0093] The prediction module is configured to deploy the trained landslide prediction model to perform landslide risk prediction.

[0094] It should be noted here that the various modules in this embodiment correspond one-to-one to the steps of the method in Example 1, and the specific implementation process is the same, which will not be repeated here.

[0095] Embodiment 3

[0096] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to complete the steps of the method in the first embodiment.

[0097] The processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and LA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0098] The memory may include one or more computer-readable media, which may be non-transitory. The memory may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable medium in the memory is used to store at least one computer program, which is used to be executed by the processor to implement a landslide risk prediction method based on landslide data set expansion provided in an embodiment of the present invention.

[0099] Those skilled in the art can understand that the electronic device provided in this embodiment can include more or fewer components, or combine certain components, or adopt different component arrangements.

[0100] Embodiment 4

[0101] This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the method in the first embodiment are completed.

[0102] Embodiment 5

[0103] This embodiment provides a computer program product, including a computer program / instruction, which implements the steps of the method in the first embodiment when executed by a processor.

[0104] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A landslide risk prediction method based on landslide data set expansion, characterized in that: include: Obtaining original landslide samples; Screen the landslide influencing factors and reclassify the original landslide samples; Selecting sample points, wherein the sample points include positive sample points and negative sample points, and both the positive sample points and the negative sample points include all landslide influencing factors after reclassification; Produce one-dimensional array data according to sample points and construct a data set; Use data expansion network to expand the data set data; The expanded dataset is used to train the landslide prediction model; Deploy the trained landslide prediction model to predict landslide risks; The data augmentation network includes a BFOA-KSMOTE oversampling module, which uses the K-means clustering algorithm to cluster the minority classes and remove marginal samples, and then generates new minority class samples through the synthetic minority class oversampling technology to balance the data set.

2. A landslide risk prediction method based on landslide data set expansion as claimed in claim 1, characterized in that: The landslide influencing factors include: topography and hydrological factors, including elevation, slope, slope direction, topography, slope position, slope variability, slope direction variability, profile curvature, plane curvature, land use type, terrain undulation, surface cutting depth, terrain moisture index, elevation variation coefficient, surface roughness, and distance from the river; geomorphological factors include distance from faults and lithology.

3. The landslide risk prediction method based on landslide data set expansion according to claim 1, characterized in that: The data expansion network includes a high- and low-prone area mixed extraction module, a BFOA-KSMOTE oversampling module and a WGAN module. The high- and low-prone area mixed extraction module extracts landslide samples and non-landslide samples and adds them to the training set and the test set; the BFOA-KSMOTE oversampling module expands the data set data to generate new minority class samples; the WGAN module further amplifies the training set data through adversarial learning.

4. A landslide risk prediction method based on landslide data set expansion as claimed in claim 3, characterized in that: The high- and low-prone areas mixed extraction module calculates the landslide sensitivity score of each sample, divides the landslide sensitivity level according to the landslide sensitivity score, extracts samples with high and very high landslide levels that are not determined to be landslides and adds them to the training set and the test set, and extracts non-landslide samples with low landslide levels and adds them to the training set and the test set.

5. The landslide risk prediction method based on landslide data set expansion according to claim 3, characterized in that: The BFOA-KSMOTE oversampling module performs clustering through an improved fruit fly optimization algorithm. For the clustered minority class samples, samples are randomly selected from their k nearest neighbors, and new samples are constructed by combining the minority class samples and the randomly selected k nearest neighbors of the samples.

6. A landslide risk prediction method based on landslide data set expansion as claimed in claim 5, characterized in that: The improved fruit fly optimization algorithm calculates the flavor concentration value of each fruit fly individual and uses the position of the individual with the optimal flavor concentration value as the initial clustering center for clustering.

7. A landslide risk prediction system based on landslide data set expansion, characterized in that: include: A data acquisition module, configured to acquire original landslide samples; The landslide impact factor screening module is configured to screen the landslide impact factors and reclassify the original landslide samples; A sample point selection module is configured to select sample points, wherein the sample points include positive sample points and negative sample points, and both the positive sample points and the negative sample points include all landslide influencing factors after reclassification; A data set construction module is configured to produce one-dimensional array data according to sample points to construct a data set; A data augmentation module is configured to augment the data set data using a data augmentation network; wherein the data augmentation network includes a BFOA-KSMOTE oversampling module, the BFOA-KSMOTE oversampling module uses a K-means clustering algorithm to cluster the minority class and remove marginal samples, and then generates new minority class samples through a synthetic minority class oversampling technique to balance the data set; A training module configured to train a landslide prediction model using the expanded data set; The prediction module is configured to deploy the trained landslide prediction model to perform landslide risk prediction.

8. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to complete the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the steps of the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.