Reservoir landslide underground water level fluctuation main control factor identification method based on data mining
By constructing multi-timescale feature windows and discretizing feature variables, and combining them with the Apriori algorithm to identify the main controlling factors of groundwater level fluctuations in reservoir landslides, the problems of lag and poor mobility of response variables in existing technologies are solved, and early warning and efficient identification of landslide disasters are realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YANTAI UNIV
- Filing Date
- 2025-12-22
- Publication Date
- 2026-05-01
AI Technical Summary
In the early warning of reservoir landslide disasters, the high-frequency and continuous change characteristics of groundwater level data are not fully utilized by existing technologies, resulting in lagging response variables, insufficient precursor identification capabilities, and poor method transferability, making it difficult to meet the real-time monitoring needs in new scenarios.
Using a data mining-based approach, a two-step clustering algorithm and the Apriori algorithm are employed, combined with groundwater level monitoring data, to construct multi-timescale feature windows. Feature variables are then discretized and association rules are mined to identify the main controlling factors of groundwater level fluctuations and their combinations.
It enables sensitive identification of early potential changes in landslides, improves the timeliness and accuracy of early warning, supports data processing at different time scales, and enhances the intelligence and applicability of landslide disaster early warning.
Smart Images

Figure CN121959079A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of groundwater monitoring and analysis technology, specifically relating to a method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining. This method utilizes time-series monitoring data, combined with a two-step clustering algorithm and Apriori association rule mining technology, to systematically identify the main controlling factors and their combinations in reservoir landslide areas, providing technical support for screening early warning factors and analyzing triggering mechanisms for reservoir landslides. Background Technology
[0002] Groundwater level fluctuations are a key internal factor in landslide deformation and failure. The spatiotemporal dynamic evolution of reservoir landslides is influenced by a combination of external environmental factors, including rainfall, reservoir water level fluctuations, and engineering activities. Traditional analytical methods have limitations in handling the nonlinear characteristics of data and factor coupling correlations, making it difficult to quantitatively identify the complex response relationship between groundwater level fluctuations and external environmental factors.
[0003] Existing research has attempted to apply two-step clustering and the Apriori algorithm to identify the main controlling factors of landslide surface displacement, and has achieved certain results. For example, the paper "Beibei Yang, Zhongqiang Liu, Suzanne Lacasse, XinLiang. Spatiotemporal deformation characteristics of Outang landslide and identification of triggering factors using data mining. Journal of RockMechanics and Geotechnical Engineering, 2024, 16(10):4088-4104. https: / / doi.org / 10.1016 / j.jrmge.2023.09.030" discloses a method for identifying the main controlling factors of landslide displacement based on two-step clustering and the Apriori algorithm. The method mainly includes the following implementation steps: Data acquisition and preprocessing: Collect GNSS displacement data of landslide monitoring points and external environmental data such as reservoir water level and rainfall, and perform time-series alignment, missing value imputation and normalization on the data; Two-step clustering analysis: Discretization of landslide displacement rate and candidate triggering factors. The algorithm first performs pre-clustering to generate sub-clusters, then automatically determines the optimal number of clusters based on the BIC criterion, and outputs the clustering interval and label for each factor; Association rule mining: Based on the Apriori algorithm, candidate frequent itemsets are constructed, and minimum support and confidence are set. Typical association rules between landslide displacement change rate and factors such as reservoir water level and rainfall are extracted from them to form the main control factor identification logic. Results output: Representative association rules were selected and interpreted to identify the main driving factors at different landslide locations.
[0004] However, this type of method only uses landslide surface displacement as the response variable, ignoring groundwater level, an important indicator reflecting the internal response of landslides. Its applicability in new application scenarios is significantly limited, mainly in the following ways: First, the response variable is lagging, resulting in insufficient early warning identification ability: Surface displacement usually occurs in the middle and late stages of landslides, with a significant lag, making it difficult to capture early disaster-causing signals of landslides in a timely manner. Groundwater level, as an important internal factor inducing landslides, often shows abnormal fluctuations before surface deformation. Existing methods ignore the characteristics of the landslide evolution process: "external environmental factors (such as rainfall, reservoir water level changes, etc.) → transformation into slope groundwater → interaction between slope groundwater and soil mass → landslide", resulting in poor early warning capability. Second, groundwater level data has the characteristics of high frequency and continuous change, and its data structure is significantly different from long-term displacement data. Existing methods lack the ability to remodel or adjust the feature factor system for groundwater level data, limiting in-depth analysis of landslide causation mechanisms. Furthermore, they suffer from poor method transferability, failing to generalize the analysis scale and parameter system: landslide displacement analysis often uses monthly data, while groundwater level monitoring can reach daily, hourly, or even minute-level scales. Existing methods lack mechanisms for processing multi-timescale inputs and do not consider the need to adjust variable discretization, threshold division, and other parameter systems due to changes in the response object, making it difficult to meet the demands for algorithm flexibility and practicality in new scenarios. Therefore, a novel method is urgently needed that can automatically extract the main controlling factors of groundwater level fluctuations based on real-time monitoring data to improve the early warning capability of reservoir landslides and the accuracy of causation mechanism analysis. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention proposes a data mining-based method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides. The method reconstructs and expands the original method in terms of response mechanism, data structure, and analysis scale, effectively solving the problems of lag in response variables, incomplete mechanism identification, and poor data scale adaptability in existing methods for identifying the main controlling factors of landslide disasters.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: A method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining, specifically including the following steps: S1. Data Acquisition and Preprocessing: Based on the reservoir landslide automatic monitoring system, high-frequency continuous monitoring data of groundwater level at different spatial locations of the landslide are automatically collected, as well as multi-source external environmental factor data with the same time resolution as the groundwater level data. The external environmental factor data includes at least reservoir water level data and rainfall data. Missing value processing, outlier removal, and time series alignment are performed on the collected raw data to construct a multivariate time series dataset with a unified time benchmark. S2. Feature window construction and indicator system reconstruction: Taking groundwater or groundwater level fluctuation characteristics as the core response variable, multi-time scale feature variables are extracted based on the sliding time window, and combined with the temporal response characteristics of groundwater level fluctuation, the traditional environmental factor indicator system is expanded and reconstructed to form a feature indicator system for the dynamic evolution process of groundwater. S3. Variable discretization based on two-step clustering: The continuous variables in the feature index system are discretized using a two-step clustering algorithm. By automatically determining the number of clusters and the corresponding category boundaries, the continuous variables are converted into discrete category variables, thereby improving the adaptability of variables to association rule mining and enhancing the stability of the data structure. S4. Identification of controlling factors based on Apriori algorithm: Taking groundwater level or groundwater as the fluctuation feature as the target term, the Apriori algorithm is used to perform frequent itemset mining and strong association rule analysis on the discretized feature variables to identify the combination of controlling factors that cause significant fluctuations in groundwater level and their interaction relationships. S5. Results Output and Rule Interpretation: Output the identified combination of main control factors and their corresponding statistical indicators such as support and confidence. Combine the hydrogeological conditions and engineering background of the landslide area to interpret and analyze the association rules, providing technical basis for the selection of landslide groundwater level early warning factors and the identification of triggering mechanisms.
[0007] As a further technical solution of the present invention, the automatic monitoring system for reservoir landslides in step S1 includes a groundwater level monitoring device, a reservoir water level monitoring device, and a rainfall monitoring device. The automatic monitoring system for reservoir landslides collects continuous monitoring data of groundwater levels at different spatial locations of the landslide, and simultaneously acquires reservoir water level data and rainfall data with the same time resolution as the groundwater level monitoring data. The monitoring data can be daily, hourly, or higher frequency data.
[0008] As a further technical solution of the present invention, the specific process of step S2 is as follows: S201. Determination of core response variables: Groundwater level or groundwater level fluctuation characteristics are used as core response variables to measure changes in groundwater dynamic processes in the landslide area. S202. Construction of multi-timescale feature windows: In view of the characteristic that groundwater level has different time lag responses to external environmental factors, a multi-timescale feature window is constructed, including short timescale, medium timescale and long timescale. Each scale is used alone or in combination to characterize the rapid response, cumulative response and lag response of groundwater level. The specific length of the timescale is set according to the time resolution of monitoring data, landslide type and hydrogeological conditions. S203. Extraction of feature variables corresponding to time scales: In feature windows at different time scales, feature variables that match the groundwater level response characteristics are extracted respectively. The types, statistical forms and combination methods of the feature variables constructed at different time scales are different to achieve multi-level representation. The feature variables include rainfall and its statistical characteristics within the window, previous cumulative rainfall, reservoir water level elevation, reservoir water level change amplitude and change rate, and combined feature variables related to groundwater dynamic processes. S204. Based on the response characteristics of groundwater level fluctuations, the traditional environmental factor index system is expanded and adjusted to form a characteristic index system for the dynamic evolution of groundwater, providing a characteristic basis for subsequent discretization processing and identification of main control factors.
[0009] As a further technical solution of the present invention, the specific process of step S3 is as follows: S301. Data Preprocessing: For the original continuous variable dataset X={x1, x2, ..., x...} n Preprocessing is performed to reduce the impact of differences in the units and numerical ranges of different variables on the distance calculation results. The data preprocessing methods include standardization, normalization, or interval scaling. When using standardization, the calculation formula is as follows: ,in, The mean of the variable. Let z be the standard deviation of the variable. i These are the standardized variable values; S302, Pre-clustering: Preprocessed data samples are sequentially input into the clustering model using a scanning method. The similarity between data points is calculated using a distance metric function to construct several initial sub-clusters. The distance metric function includes Euclidean distance. When Euclidean distance is used, its calculation method is as follows: , where x i and x j Let m represent two samples, where m is the dimension of the variable. When the distance between a new sample and the center of an existing sub-cluster is less than a preset threshold, the sample is merged into the corresponding sub-cluster; otherwise, a new sub-cluster is generated, thus forming a cluster feature tree structure. S303. Automatic Determination of Clustering and Number of Categories: Based on the pre-clustering results, hierarchical clustering is used to merge the sub-clusters, and the optimal number of categories, K, is automatically determined through model selection criteria. The Bayesian Information Criterion (BIC) is used as the evaluation index, and its expression is: Where L is the likelihood function value of the clustering model, p is the number of model parameters, and n is the number of samples. When BIC reaches its minimum value, the corresponding number of clusters K is determined as the optimal number of categories. S304. Category Interval Determination and Boundary Inverse Transformation: Based on the final clustering results, the value range of samples within each category is statistically analyzed in the feature space to obtain the interval boundaries corresponding to each category. When the clustering analysis is performed based on the standardized feature space, the interval boundaries are inversely transformed according to the standardization parameters to map them back to the original variable space. The inverse transformation relationship is as follows: Where z is the boundary value in the standardized space, and x is the boundary value of the original variable after restoration. The mean of the variable. The standard deviation of the variable is used to determine the upper and lower boundaries of each discrete category in the original variable space. k b k ], where k = 1, 2, ..., K; S305. Variable Discretization Mapping: Based on the determined upper and lower boundaries, continuous variables are mapped to corresponding discrete categorical variables. The mapping rules are as follows: ⇒x i →C k This allows for the transformation of continuous variables into discrete categorical variables, and the discretization results are used as input data for subsequent association rule mining.
[0010] As a further technical solution of the present invention, the specific process of step S4 is as follows: S401. Determining the Target Item and Input Data: Taking the groundwater level fluctuation rate or groundwater level change state as the target item Y, the feature variable set {X1, X2, ..., X...} discretized in step S3 is used to determine the target item Y. n} as the candidate set input to the Apriori algorithm, expressed as: , where X i For the discretized environmental characteristic variables, Y is the target variable for groundwater level fluctuations; S402. Frequent Itemset Mining: Based on the minimum support threshold min_sup, the Apriori algorithm is used to iteratively generate candidate itemsets, and frequent itemsets that meet the following conditions are selected: , where I k ⊆I is the candidate set, count(I) k ) represents the number of samples containing this itemset, N represents the total number of samples, and freq(I) represents the number of samples containing this itemset. kSupport is represented by a number; frequent itemsets are combinations of feature variables that occur simultaneously with a high probability in historical data. S403. Strong Association Rule Generation: Generating Association Rules R:X Based on Frequent Itemsets a ⇒Y, calculate rule confidence: Only rules that satisfy the minimum confidence threshold min_conf are retained, i.e., strong association rules; S404, Identification of Controlling Factor Combinations and Interaction Relationships: In strong association rules, the set of antecedent variables X containing the target term Y is included. a For combinations of controlling factors, the interaction effects of multiple variables appearing simultaneously in the antecedents can be measured using lift. ; S405. Identification Result Output: Output the identification result, which includes: Main Control Factor Combination X a This output contains frequent itemsets and strong association rules, as well as the interaction relationships between various combinations. This is the raw mining result and does not involve direction or threshold discrimination. It is used for subsequent rule interpretation.
[0011] As a further technical solution of the present invention, the specific process of step S5 is as follows: S501, Rule Judgment: Based on the output of S4, significance judgment is performed, including support, confidence, and lift threshold screening; the combination of major control factors that significantly affect groundwater level fluctuations is selected. S502, Factor Influence Direction and Threshold Analysis: For each combination of main controlling factors, analyze its influence direction and degree on groundwater level, and identify key threshold conditions: X a Significant fluctuations in Y within a specific interval can be identified, thus revealing synergistic or interactive effects between different factors. S503. Output of Interpretation Results: The rules and analysis results after discrimination will be visualized or tabulated, including the combination of main control factors and discrimination thresholds; the direction of influence of groundwater level fluctuations; the explanation of the interaction between factors; and the output results will serve as the technical basis for the selection of landslide groundwater level early warning factors and the analysis of triggering mechanisms.
[0012] Compared with existing technologies, the present invention has significant advantages in terms of technical methods, applicable objects, and engineering value, mainly reflected in the following aspects: (1) This invention uses groundwater level as the core internal indicator of landslide response, which has higher sensitivity and predictive power compared to surface displacement. By systematically identifying the coupling characteristics of groundwater level and external environmental factors, this method can effectively capture landslide precursor signals, realize early identification and early warning, and significantly improve the response speed and prediction accuracy of intelligent early warning of landslide disasters.
[0013] (2) This invention supports input at different time scales, can process high-frequency groundwater level monitoring data, and can be extended to factor identification tasks of other endogenous-driven geological disasters. By constructing a multi-scale data adaptation mechanism, the method has good adaptability and generalization ability in terms of response objects, time resolution and factor system, which improves the versatility and promotion potential of the method.
[0014] (3) This invention provides a theoretical basis and technical support for the development of a refined landslide risk management and monitoring and early warning system. The proposed identification framework can be used as a factor screening module in an intelligent early warning model to realize dynamic identification and intelligent management of landslide risks across multiple factors, scales, and stages, providing technical support for disaster prevention and control in reservoir areas and other important engineering areas. Attached Figure Description
[0015] Figure 1 A flowchart of association rule mining provided by this invention.
[0016] Figure 2 The flowchart of the method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining provided by the present invention. Detailed Implementation
[0017] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the described embodiments are only some embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Equivalent substitutions or modifications made to the embodiments by those skilled in the art without departing from the technical concept of the present invention should all fall within the scope of protection of the present invention. The accompanying drawings are only used to assist in illustrating the technical solution of this embodiment and do not constitute a specific limitation of the present invention.
[0018] This embodiment provides a method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining technology. It is used to identify and extract the main driving factors affecting groundwater level changes in reservoir landslides under multi-source monitoring data conditions. Through systematic analysis of automated monitoring data, this method can characterize the groundwater level change features of different spatial locations of reservoir landslides under multi-time scale conditions, and on this basis, effectively identify the main controlling factors of groundwater level fluctuations.
[0019] like Figure 1 and Figure 2 As shown in the figure, this embodiment provides a method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining, specifically including the following steps: S1. Data Acquisition and Preprocessing: In this embodiment, a landslide on the bank of a reservoir is used as the application object. An automated monitoring system is deployed in the landslide area. The automated monitoring system includes at least a groundwater level monitoring device, a reservoir water level monitoring device, and a rainfall monitoring device. Through the automated monitoring system, continuous groundwater level monitoring data is collected from different spatial locations of the landslide. Simultaneously, reservoir water level data and rainfall data with the same time resolution as the groundwater level monitoring data are acquired. The monitoring data can be daily, hourly, or higher frequency data. The collected raw data is then preprocessed, and the preprocessing includes, but is not limited to: Interpolate or remove missing data; Abnormal data can be identified and removed using thresholding, statistical testing, or time-series-based methods. Time alignment is performed on data from different sources to construct a multivariate time series dataset with a unified time base.
[0020] S2. Feature window construction and indicator system reconstruction: In this embodiment, groundwater level or groundwater level fluctuation characteristics are used as the core response variable.
[0021] To address the time-lag response of groundwater levels to rainfall and reservoir water level changes, a sliding time window approach is employed to construct multi-timescale feature windows. These timescales can be set based on the temporal resolution of the monitoring data, landslide type, and hydrogeological conditions, and can be used individually or in combination. Within feature windows at different timescales, feature variables matching the groundwater level response characteristics are extracted. These feature variables include, but are not limited to: rainfall amount and its statistical characteristics within the window; previous cumulative rainfall; reservoir water level elevation; reservoir water level change amplitude and rate of change; and combined feature variables characterizing groundwater dynamic processes.
[0022] By using the above methods, the traditional environmental factor index system is expanded and reconstructed to form a characteristic index system oriented towards the dynamic evolution of groundwater, thereby enhancing the ability to characterize the groundwater level response mechanism.
[0023] S3. Variable discretization based on two-step clustering: In this embodiment, the continuous feature variables constructed in step S2 are discretized using a two-step clustering algorithm.
[0024] The two-step clustering process includes a pre-clustering stage and a clustering stage: In the pre-clustering stage, several initial sub-clusters are constructed by sequential scanning. During the clustering phase, the optimal number of clusters is automatically determined based on the model evaluation criteria, and the value range corresponding to each cluster is obtained.
[0025] The discretization process described above transforms continuous feature variables into discrete categorical variables, thereby improving the adaptability of the variables to subsequent association rule mining algorithms and enhancing the stability of the data structure.
[0026] S4. Identification of master control factors based on the Apriori algorithm: In this embodiment, the discretized feature variables are used as candidate itemsets, and the groundwater level or groundwater level fluctuation status is used as the target item. The Apriori algorithm is used to perform frequent itemset mining and strong association rule analysis.
[0027] Support and confidence thresholds are set, with support ≥ 1.5% and confidence ≥ 85%. Association rules that meet the conditions are screened to identify the main control factor combination that causes significant fluctuations in groundwater level. Furthermore, the interaction relationship under the combined effect of multiple factors is analyzed through the lift index to reveal the synergistic influence characteristics of different environmental factors on groundwater level fluctuations.
[0028] S5. Results Output and Rule Explanation: In this embodiment, the output identification results include the combination of master control factors and their corresponding statistical indicators such as support, confidence and lift.
[0029] Based on the hydrogeological conditions and engineering background of the landslide area, the association rules are interpreted and analyzed to identify the direction, degree of influence and potential threshold range of each combination of main control factors on groundwater level fluctuations. This forms a discrimination rule that can be used for the selection of landslide groundwater level early warning factors and the analysis of triggering mechanisms, providing a technical basis for the dynamic identification and intelligent management of landslide risks in reservoir areas and other engineering areas.
[0030] In this embodiment, the relevant algorithm parameters and time window settings can be flexibly adjusted according to the characteristics of the monitoring data and the identification target, and the present invention does not impose specific limitations on them.
[0031] The core of the technical solution of this invention is as follows: 1) This embodiment constructs a response system with groundwater level as the core, enhances the sensitivity and timeliness of risk identification, and innovatively uses groundwater level fluctuation as the target variable for landslide disaster response. By analyzing its correlation with external factors such as rainfall and reservoir water level, a framework for identifying the main control factors that reflect the early potential changes of landslides is constructed, which significantly improves the advance and effectiveness of risk perception.
[0032] 2) By reconstructing the analysis process to adapt to high-frequency time-series data and new disaster-causing factors, the mechanism identification is deepened. In view of the continuity and high temporal resolution of groundwater level data, the attribute discretization method and factor window design in the original method are optimized. Strongly correlated variables such as short-term rainfall and sudden changes in reservoir water level are introduced to establish a factor index system that is more in line with the groundwater dynamic process, thereby accurately revealing the internal driving mechanism of landslide disaster.
[0033] 3) A multi-scale data mining mechanism was designed to improve the applicability and scenario transferability of the method. It supports time window input from hourly to monthly scales, and can dynamically adjust the analysis granularity and parameter settings. It constructs a multi-scale main control factor identification model suitable for different landslide response stages, which significantly improves the generalization ability and engineering application value of the method under different data sources, monitoring accuracy and response objects.
[0034] The models, algorithms, etc. not described in detail in this invention are all general technologies in the field, and therefore will not be described in detail here.
Claims
1. A method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining, characterized in that, Specifically, the following steps are included: S1. Data Acquisition and Preprocessing: Based on the reservoir landslide automatic monitoring system, high-frequency continuous monitoring data of groundwater level in different spatial locations of the landslide is automatically collected, as well as multi-source external environmental factor data with the same time resolution as the groundwater level data. The external environmental factor data includes at least reservoir water level data and rainfall data. The collected raw data were processed for missing values, outlier removal, and time series alignment to construct a multivariate time series dataset with a unified time base. S2. Feature window construction and indicator system reconstruction: Taking groundwater or groundwater level fluctuation characteristics as the core response variable, multi-time scale feature variables are extracted based on the sliding time window, and combined with the temporal response characteristics of groundwater level fluctuation, the traditional environmental factor indicator system is expanded and reconstructed to form a feature indicator system for the dynamic evolution process of groundwater. S3. Variable discretization based on two-step clustering: The continuous variables in the feature index system are discretized using a two-step clustering algorithm. By automatically determining the number of clusters and the corresponding category boundaries, the continuous variables are converted into discrete category variables. S4. Identification of controlling factors based on Apriori algorithm: Taking groundwater level or groundwater as the fluctuation feature as the target term, the Apriori algorithm is used to perform frequent itemset mining and strong association rule analysis on the discretized feature variables to identify the combination of controlling factors that cause significant fluctuations in groundwater level and their interaction relationships. S5. Results Output and Rule Interpretation: Output the identified combination of main control factors and their corresponding statistical indicators such as support and confidence. Combine the hydrogeological conditions and engineering background of the landslide area to interpret and analyze the association rules, providing technical basis for the selection of landslide groundwater level early warning factors and the identification of triggering mechanisms.
2. The method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining as described in claim 1, characterized in that, The automatic monitoring system for reservoir landslides described in step S1 includes a groundwater level monitoring device, a reservoir water level monitoring device, and a rainfall monitoring device. The system collects continuous groundwater level monitoring data from different spatial locations of the landslide and simultaneously acquires reservoir water level data and rainfall data with the same time resolution as the groundwater level monitoring data. The monitoring data can be daily, hourly, or higher frequency data.
3. The method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining as described in claim 2, characterized in that, The specific process of step S2 is as follows: S201. Determination of core response variables: Groundwater level or groundwater level fluctuation characteristics are used as core response variables to measure changes in groundwater dynamic processes in the landslide area. S202. Construction of multi-timescale feature windows: In view of the characteristic that groundwater level has different time lag responses to external environmental factors, a multi-timescale feature window is constructed, including short timescale, medium timescale and long timescale. Each scale is used alone or in combination to characterize the rapid response, cumulative response and lag response of groundwater level. The specific length of the timescale is set according to the time resolution of monitoring data, landslide type and hydrogeological conditions. S203. Extraction of feature variables corresponding to time scales: In feature windows at different time scales, feature variables that match the groundwater level response characteristics are extracted respectively. The types, statistical forms and combination methods of the feature variables constructed at different time scales are different to achieve multi-level representation. The feature variables include rainfall and its statistical characteristics within the window, previous cumulative rainfall, reservoir water level elevation, reservoir water level change amplitude and change rate, and combined feature variables related to groundwater dynamic processes. S204. Based on the response characteristics of groundwater level fluctuations, the traditional environmental factor index system is expanded and adjusted to form a characteristic index system for the dynamic evolution of groundwater, providing a characteristic basis for subsequent discretization processing and identification of main control factors.
4. The method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining as described in claim 3, characterized in that, The specific process of step S3 is as follows: S301. Data Preprocessing: For the original continuous variable dataset X={x1, x2, ..., x...} n Preprocessing is performed to reduce the impact of differences in the units and numerical ranges of different variables on the distance calculation results. The data preprocessing methods include standardization, normalization, or interval scaling. When using standardization, the calculation formula is as follows: , in, The mean of the variable. Let z be the standard deviation of the variable. i These are the standardized variable values; S302, Pre-clustering: Preprocessed data samples are sequentially input into the clustering model using a scanning method. The similarity between data points is calculated using a distance metric function to construct several initial sub-clusters. The distance metric function includes Euclidean distance. When Euclidean distance is used, its calculation method is as follows: , where x i and x j Let m represent two samples, where m is the dimension of the variable. When the distance between a new sample and the center of an existing sub-cluster is less than a preset threshold, the sample is merged into the corresponding sub-cluster; otherwise, a new sub-cluster is generated, thus forming a cluster feature tree structure. S303. Automatic Determination of Clustering and Number of Categories: Based on the pre-clustering results, hierarchical clustering is used to merge the sub-clusters, and the optimal number of categories, K, is automatically determined through model selection criteria. The Bayesian Information Criterion (BIC) is used as the evaluation index, and its expression is: Where L is the likelihood function value of the clustering model, p is the number of model parameters, and n is the number of samples. When BIC reaches its minimum value, the corresponding number of clusters K is determined as the optimal number of categories. S304. Category Interval Determination and Boundary Inverse Transformation: Based on the final clustering results, the value range of samples within each category is statistically analyzed in the feature space to obtain the interval boundaries corresponding to each category. When the clustering analysis is performed based on the standardized feature space, the interval boundaries are inversely transformed according to the standardization parameters to map them back to the original variable space. The inverse transformation relationship is as follows: Where z is the boundary value in the standardized space, and x is the boundary value of the original variable after restoration. The mean of the variable. The standard deviation of the variable is used to determine the upper and lower boundaries of each discrete category in the original variable space. k b k ], where k = 1, 2, ..., K; S305. Variable Discretization Mapping: Based on the determined upper and lower boundaries, continuous variables are mapped to corresponding discrete categorical variables. The mapping rules are as follows: ⇒x i →C k This allows for the transformation of continuous variables into discrete categorical variables, and the discretization results are used as input data for subsequent association rule mining.
5. The method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining according to claim 4, characterized in that, The specific process of step S4 is as follows: S401. Determining the Target Item and Input Data: Taking the groundwater level fluctuation rate or groundwater level change state as the target item Y, the feature variable set {X1, X2, ..., X...} discretized in step S3 is used to determine the target item Y. n } as the candidate set input to the Apriori algorithm, expressed as: , where X i For the discretized environmental characteristic variables, Y is the target variable for groundwater level fluctuations; S402. Frequent Itemset Mining: Based on the minimum support threshold min_sup, the Apriori algorithm is used to iteratively generate candidate itemsets, and frequent itemsets that meet the following conditions are selected: , where I k ⊆I is the candidate set, count(I) k ) represents the number of samples containing this itemset, N represents the total number of samples, and freq(I) represents the number of samples containing this itemset. k Support is represented by a number; frequent itemsets are combinations of feature variables that occur simultaneously with a high probability in historical data. S403. Strong Association Rule Generation: Generating Association Rules R:X Based on Frequent Itemsets a ⇒Y, calculate rule confidence: Only rules that satisfy the minimum confidence threshold min_conf are retained, i.e., strong association rules; S404, Identification of Controlling Factor Combinations and Interaction Relationships: In strong association rules, the set of antecedent variables X containing the target term Y is included. a For combinations of controlling factors, the interaction effects of multiple variables appearing simultaneously in the antecedents can be measured using lift. ; S405. Identification Result Output: Output the identification result, which includes: Main Control Factor Combination X a This output contains frequent itemsets and strong association rules, as well as the interaction relationships between various combinations. This is the raw mining result and does not involve direction or threshold discrimination. It is used for subsequent rule interpretation.
6. The method for identifying the main controlling factors of groundwater level fluctuations in reservoir landslides based on data mining as described in claim 5, characterized in that, The specific process of step S5 is as follows: S501, Rule Judgment: Based on the output of S4, significance judgment is performed, including support, confidence, and lift threshold screening; the combination of major control factors that significantly affect groundwater level fluctuations is selected. S502, Factor Influence Direction and Threshold Analysis: For each combination of main controlling factors, analyze its influence direction and degree on groundwater level, and identify key threshold conditions: X a Significant fluctuations in Y within a specific interval can be identified, thus revealing synergistic or interactive effects between different factors. S503. Output of Interpretation Results: The rules and analysis results after discrimination will be visualized or tabulated, including the combination of main control factors and discrimination thresholds; the direction of influence of groundwater level fluctuations; the explanation of the interaction between factors; and the output results will serve as the technical basis for the selection of landslide groundwater level early warning factors and the analysis of triggering mechanisms.