Model combination-based karst basin sediment transport space change influence factor decoupling method and system

By employing a model combination method, correlation analysis and random forest models are used to decouple the multi-factor influences on sediment transport in karst watersheds, thus solving the complexity of sediment transport variation in karst watersheds. This approach enables accurate identification and contribution rate analysis of key factors, thereby enhancing the robustness and explanatory power of the model.

CN121456444AActive Publication Date: 2026-02-03INSTITUTE OF SUBTROPICAL AGRICULTURE CHINESE ACADEMY OF SCIENCES
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511796689.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-02-03
Estimated Expiration
2045-12-02

AI Technical Summary

Technical Problem

Existing research has difficulty in systematically quantifying the comprehensive impact of multi-factor coupling on sediment transport changes in karst basins. Traditional methods are also unable to reliably identify key driving factors under the hydrogeological structure and landscape characteristics of karst regions, resulting in scattered conclusions and poor reproducibility.

Method used

A model-based approach was adopted, combining correlation analysis, random forest, and partial least squares structural equation modeling to construct a method for decoupling the influencing factors of spatial variation of sediment transport in karst watersheds. This method includes basic data construction, latent variable mapping, correlation screening, random forest model, and structural equation modeling, thereby achieving quantitative decoupling of sediment transport.

Benefits of technology

It significantly improves the identification accuracy and model stability of the multi-factor effects of sediment transport in karst watersheds, enabling accurate identification and contribution rate analysis of the main controlling factors under complex geomorphological and multi-source data conditions, and providing scientific support for watershed soil and water conservation and ecological restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456444A_ABST
    Figure CN121456444A_ABST
Patent Text Reader

Abstract

The invention provides a karst basin sediment transport space change influence factor decoupling method and system based on model combination, and relates to the technical field of hydrologic landform modeling. The method comprises the following steps: establishing a basic watershed data set, establishing a relationship between potential variables and observation variables under a five-class factor framework of climate, lithology, soil, terrain and landscape, screening variables by adopting correlation analysis, evaluating importance in combination with a random forest, and selecting representative variables. And calculating a path coefficient and effect intensity by using a partial least square structure equation model, and quantitatively analyzing the action direction and contribution rate of each potential factor to the sediment transport space change. According to the method, quantitative decoupling of multi-factor action is realized, and the main control factor identification precision, the action mechanism analysis depth and the model stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of hydrological geomorphology modeling, in particular to a method and system for decoupling influencing factors of spatial variation of sediment discharge in a karst basin based on model combination. BACKGROUND

[0002] The sediment transport process in karst areas is affected by multiple factors such as climate, lithology, soil, topography and landscape pattern. Most existing researches only analyze a single or a few factors, which is difficult to systematically quantify the comprehensive influence of multi-factor coupling on the variation of sediment discharge at the spatial scale. In particular, in karst basins where carbonate rocks are widely distributed, surface and underground hydrological processes are highly connected, and landscape structure is complex and diverse, the interaction of precipitation input, topography and landscape pattern makes the sediment discharge show significant spatial heterogeneity. Traditional single-model analysis or simple correlation analysis method is difficult to reveal the nonlinear coupling relationship between factors and their contribution to the sediment transport mechanism, and it is difficult to realize quantitative decoupling of the driving mechanism.

[0003] The existing methods are generally disconnected in the aspects of variable screening and causal identification. On the one hand, the variable system affecting the sediment discharge is large and has serious collinearity, and the traditional screening method based on linear assumption is difficult to identify the key driving factors stably; on the other hand, although the structural equation model can reflect the direct and indirect effects between latent variables, it is highly dependent on the representativeness of observation indicators, and if the variable screening is insufficient in the early stage, the reliability of path estimation is easily affected. The special hydrogeological structure in karst areas, the discontinuity of soil distribution and the highly fragmented characteristics of landscape further amplify the above problems, leading to scattered conclusions and poor reproducibility in revealing the dominant factors, contribution rate and interaction of the spatial variation of sediment discharge, and there is still a lack of a comprehensive analysis method that can realize stable decoupling under the conditions of complex geomorphology and multi-source data. SUMMARY

[0004] In order to overcome the shortcomings of the prior art, the purpose of the present application is to provide a method and system for decoupling influencing factors of spatial variation of sediment discharge in a karst basin based on model combination, which realizes quantitative decoupling of multi-factor action of sediment discharge in a karst basin by combining correlation analysis, random forest and partial least squares structural equation model, and significantly improves the identification accuracy of dominant factors, the depth of mechanism analysis and the stability of the model.

[0005] To achieve the above purpose, the present application provides the following scheme:

[0006] A method for decoupling influencing factors of spatial variation of sediment discharge in a karst basin based on model combination, comprising:

[0007] constructing a basic data set of the target karst basin;

[0008] The correspondence between potential variables and observation variables is established according to the basic data set under the framework of five factors of climate, lithology, soil, topography and landscape;

[0009] Pearson correlation analysis is performed with the observation variables as input and the sediment discharge as output, and the observation variables with significant correlation are selected to generate a candidate factor set;

[0010] A random forest model is established with the candidate factor set as input, the importance of variables is evaluated according to the out-of-bag error (OBB Error), and representative observation variables are selected;

[0011] A partial least squares structural equation model is constructed with climate, lithology, soil, topography and landscape as potential variables, combined with the representative observation variables, the path coefficient and the effect strength are calculated, and the model estimation result is obtained;

[0012] Based on the model estimation result, the action direction and contribution rate of each potential variable on the spatial variation of sediment discharge are analyzed, the dominant influencing factor is determined, and the decoupling result is output.

[0013] A model combination-based karst watershed sediment discharge spatial variation influencing factor decoupling system, comprising:

[0014] A basic data construction unit is configured to construct a basic data set of a target karst watershed;

[0015] A potential variable mapping unit is configured to establish a correspondence between potential variables and observation variables according to the basic data set under the framework of five factors of climate, lithology, soil, topography and landscape;

[0016] A correlation screening unit is configured to perform Pearson correlation analysis with the observation variables as input and the sediment discharge as output, select the observation variables with significant correlation, and generate a candidate factor set;

[0017] A variable optimization unit is configured to establish a random forest model with the candidate factor set as input, evaluate the importance of variables according to the out-of-bag error, and select representative observation variables;

[0018] A structural equation modeling unit is configured to construct a partial least squares structural equation model with climate, lithology, soil, topography and landscape as potential variables, combined with the representative observation variables, calculate the path coefficient and the effect strength, and obtain the model estimation result;

[0019] A decoupling analysis unit is configured to analyze the action direction and contribution rate of each potential variable on the spatial variation of sediment discharge based on the model estimation result, determine the dominant influencing factor, and output the decoupling result.

[0020] The present application discloses the following technical effects:

[0021] The application effectively overcomes the identification limitations of traditional linear models under the conditions of multiple collinearity and high-dimensional heterogeneous data by incorporating five types of potential variables, climate, lithology, soil, topography and landscape, into a unified framework, using correlation analysis for preliminary screening and random forest optimization, and realizing the robust screening of key observation variables affecting the spatial variation of sediment discharge. Compared with relying only on single statistical correlation or principal component method, the method can significantly improve the identification accuracy of multi-factor synergistic effect and the distinction degree of dominant factors.

[0022] The application introduces a PLS-SEM model based on the results of random forest optimization, uses path coefficient decomposition to directly realize the effects of climate, lithology, soil, topography and landscape on sediment discharge, and realizes the quantitative decoupling of the mechanism. The model can simultaneously evaluate the causal chain and interaction direction between potential variables, making the positive and negative influence relationship in the complex system more explicit, and the explanation rate is significantly higher than that of a single model, which is suitable for high-heterogeneity characteristic analysis in karst basins.

[0023] The application forms a closed-loop mechanism from variable screening to structural equation verification by connecting multiple models in series, avoids the dependence on large samples or normal distribution, and improves the robustness of the model under complex topography and limited sample conditions. The decoupling results not only reveal the dominant role of landscape and topographic factors, but also serve as a quantitative decision basis for watershed soil and water conservation and ecological restoration, providing scientific support for water and sediment process research and management in karst areas. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0025] Figure 1 The method flowchart provided for the embodiments of the present application;

[0026] Figure 2 The 1950-2015 typical karst basin annual sediment discharge time variation schematic diagram provided for the embodiments of the present application;

[0027] Figure 3 The 2003-2015 monthly sediment discharge time variation schematic diagram of three typical karst basins provided for the embodiments of the present application;

[0028] Figure 4 The research basin runoff spatiotemporal variation schematic diagram provided for the embodiments of the present application;

[0029] Figure 5A schematic diagram of the research on the spatial and temporal variation of sediment discharge in a basin is provided for the embodiment of the present application.

[0030] Figure 6 A schematic diagram of Pearson correlation analysis between runoff, sediment discharge and climate, soil, terrain, land use and landscape factors is provided for the embodiment of the present application.

[0031] Figure 7 A schematic diagram of random forest algorithm analysis on the relative importance of different variables to runoff is provided for the embodiment of the present application.

[0032] Figure 8 A schematic diagram of the first PLS-SEM model analysis result is provided for the embodiment of the present application.

[0033] Figure 9 A schematic diagram of the second PLS-SEM model analysis result is provided for the embodiment of the present application.

[0034] Figure 10 A schematic diagram of the research on the spatial and temporal variation of RDs, R25 and PT in a basin is provided for the embodiment of the present application.

[0035] Figure 11 A schematic diagram of correlation analysis between climate factors and runoff and between landscape factors and sediment discharge is provided for the embodiment of the present application.

[0036] Figure 12 A schematic diagram of the relationship between R25 and runoff is provided for the embodiment of the present application.

[0037] Figure 13 A schematic diagram of the relationship between RX3 and sediment discharge is provided for the embodiment of the present application.

[0038] Figure 14 A schematic diagram of the variation trend of water area (W), weighted average shape index (SHAPE-AM) and landscape shape index (LSI) factors in landscape factors is provided for the embodiment of the present application. DETAILED DESCRIPTION

[0039] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0040] The application aims to provide a model combination-based karst watershed sediment discharge spatial variation influencing factor decoupling method and system.

[0041] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the application will be further described in detail below in combination with the drawings and specific embodiments.

[0042] Figure 1 The method flowchart provided by the embodiment of the application is shown in Figure 1 The application provides a model combination-based karst watershed sediment discharge spatial variation influencing factor decoupling method, which comprises the following steps:

[0043] Step 100: constructing a basic data set of a target karst watershed;

[0044] Step 200: under the framework of five factors of climate, lithology, soil, topography and landscape, establishing a corresponding relationship between potential variables and observed variables according to the basic data set;

[0045] Step 300: performing Pearson correlation analysis with the observed variables as input and the sediment discharge as output, selecting the observed variables with significant correlation, and generating a candidate factor set;

[0046] Step 400: establishing a random forest model with the candidate factor set as input, evaluating the importance of the variables according to the change of out-of-bag error, and screening out representative observed variables;

[0047] Step 500: constructing a partial least squares structural equation model with the climate, lithology, soil, topography and landscape as potential variables and combining the representative observed variables, calculating the path coefficient and effect strength, and obtaining the model estimation result;

[0048] Step 600: analyzing the action direction and contribution rate of each potential variable on the spatial variation of the sediment discharge based on the model estimation result, determining the dominant influencing factor, and outputting the decoupling result.

[0049] Specifically, the present embodiment first collects and processes basic data. Meteorological data (precipitation, temperature, potential evapotranspiration data) are obtained from 52 weather stations in the basin (target karst basin) and surrounding areas where the study area is located, provided by China Meteorological Information Network. Potential evapotranspiration (PET) is calculated using the Penman formula, and precipitation and potential evapotranspiration are calculated using the CoKriging interpolation algorithm in ArcGIS 10.8 software to obtain the spatial average of each basin. The lithology map comes from the Institute of Geochemistry, Chinese Academy of Sciences, to calculate the carbonate rock coverage of the karst basin. Normalized Difference Vegetation Index (NDVI) and Enhanced Vegetation Index (EVI) are retrieved from the Google Earth Engine platform. Land use data is selected from the 300 million land cover data set in China from 1990 to 2021, which includes nine land use types: farmland, forest, shrub, grassland, water, ice and snow, bare land, impervious surface and wetland. Soil data is extracted from the World Soil Database (HWSD) to obtain soil particle size distribution, soil bulk density (SBD), electrical conductivity (EC), calcium carbonate content (CAC), pH and organic carbon content (SOC) and other data.

[0050] Topographic data are extracted from DEM data, and 30-meter resolution digital elevation model (DEM) data are downloaded from the National Aeronautics and Space Administration. The downloaded NASA DEM data set is preprocessed by format conversion, projection conversion and mask cropping, and the corresponding 40 basin boundaries are generated from DEM using ArcGIS 10.8. Landscape data selects commonly used landscape indices (edge index, shape index, diversity index, etc.) that can reflect changes in landscape pattern, including patch, class, and landscape. It is obtained from the land use map by using Fragstats 4.2 software. Runoff and sediment data sets come from the basin hydrological station and “China Sediment Bulletin”, which are measured at the outlet of the basin. The above data sets are checked before publication to ensure their reliability and consistency.

[0051] Optionally, the present embodiment provides a partial least squares-structural equation model (PLS-SEM), which is a causal modeling method that combines principal component analysis with multiple regression, aiming to maximize the explained variance of the relevant structure. Based on certain assumptions, the complex relationship between latent variables can be measured through corresponding observed variables. The model can be divided into an internal model and an external model. The internal model (structural model) solves the complex relationship between the interaction of latent variables, while the external model (measurement model) considers the relationship between each latent variable and the corresponding observed variable. The PLS-SEM model combines the measurement model and the structural model to establish a conceptual model of the relationship between independent variables and dependent variables. The model uses an iterative algorithm to solve the components of the measurement model, and predicts the relevant path coefficients in the structural model through the partial least squares method. The relationship (ξ j ) between latent variables (structural model) can be represented as: ;

[0052] wherein ξ j (j=1,...,j) refers to a general endogenous latent variable; β ji is the path coefficient between the i-th exogenous latent variable and the j-th endogenous latent variable; ξ j is the internal relationship error of the model. The relationship between latent variables (ξ j ) and observed variables (X jk ) can be expressed as: ;

[0053] λ jk is whether the j-th explicit variable is related in the k-th block; the error term ε jk represents the uncertainty error in measurement. The goodness-of-fit index (GoF) is selected to validate the model to determine the predictive ability of the model. That is, the geometric mean of the average community index and the average R 2 value product. It can be expressed as: ;

[0054] The PLS-SEM model goodness-of-fit index is mainly used to evaluate the overall predictive ability of the model. In this system, the total effect between two variables is the sum of the direct effect and the indirect effect, where the direct effect is determined by the corresponding path coefficient, and the indirect effect refers to the path involving intermediate variables. The model quantifies the direct and indirect effects between multiple factors by constructing a causal relationship network between latent variables and observed variables, and is suitable for analyzing small sample data and complex system relationships. This method has less data limitations and is mainly applied to theoretical development and result prediction. The present embodiment uses R and the "PLSP" software package, and uses a component-based PLS-SEM model.

[0055] Random forest (RF) is a non-parametric regression machine learning algorithm that combines random classification and regression trees, mainly used for classification and prediction. This method mainly prioritizes the relative importance of influencing factors, which helps the classification and prediction paradigm. The algorithm uses a bootstrap resampling technique to randomly select N training sets from the original data set and builds classification regression trees, generating an RF consisting of N classification regression trees, and finally the most repeated tree is the final result. For each iteration, each sample of the uncaptured data part generates an out-of-bag (OOB) sample to guide the exclusion of data points in the sample, and finally evaluates the model performance. The mean square error (MSE) calculated from the OOB data can be used as an indicator to quantify the importance of the predicted variables. The larger the mean square error, the more important the corresponding variable, and the greater the contribution to the model. This algorithm requires fewer parameters to adjust and can handle small sample sizes and complex data structures. The RF algorithm is widely used for its skilled management of extensive and complex data sets. The RF algorithm implementation is based on the "randomForest" package in the R statistical environment, and the calculation is performed in the Matlab 2016a processing framework.

[0056] As an optional implementation, the embodiment introduces "core factor priority probability" and "redundancy suppression weighted split criterion" in the random forest training process. The core factor priority probability refers to the statistical correlation strength between each observation variable and the sediment discharge obtained in the previous step, and the prediction importance reflected by the error increment of the observation variable on the out-of-bag sample, which is fused into the probability for node feature sampling; The higher the probability, the greater the frequency of being drawn into the candidate set. The redundancy suppression weighted split criterion is to consider the error reduction amplitude of the current node, the core factor priority probability of the variable in the global range, and the maximum correlation between the variable and the variables already used in the tree path when selecting the division variable of the node, to balance the local precision improvement and global variable independence of node division, and to preferentially select the candidate that can significantly reduce the error without redundancy with the variables already used. The embodiment faces a total of 103 observation variables and 5 potential variable categories, and at least one used variable is involved in the redundancy judgment at each node to ensure the effectiveness and interpretability of the path constraint.

[0057] The embodiment combines two types of evidence into a single, comparable sampling probability: one is the correlation strength with the sediment discharge (taking the absolute value to ensure that the direction does not affect the weight), and the other is the mean square error increment on the out-of-bag sample after shuffling the variable, which is used to measure the degree of dependence of the model on the variable. In order to obtain a stable probability distribution, the embodiment uses exponential normalization for all candidate variables, which compresses the two types of evidence with different dimensions to the same scale of zero to one, and makes the sum of the probabilities of all candidate variables strictly equal to one. In order to avoid excessive concentration, the embodiment uses the default setting of smoothing temperature 1, does not introduce additional weights to be adjusted, the ratio of the number of features of each node participating in the candidate to the total number of features is not less than 0.2, in order to maintain the balance between exploration and computational cost; the significance determination follows the previous screening criteria, with a threshold of 0.05, which is used to filter variables with low evidence strength into the sampling pool.

[0058] After completing the node-level feature sampling, the embodiment calculates a weighted score for each candidate variable, which is composed of three parts: the error reduction of the current node (reflecting the immediate benefit), the core factor priority probability (reflecting the global importance), and the redundancy penalty (reflecting the maximum correlation with the previous path variables). When multiple candidates have similar scores, the candidate with greater error reduction is preferred to ensure the improvement of local fitting; the lower limit of the redundancy penalty is fixed at 1, corresponding to the baseline strength without redundancy; the upper limit of the correlation is 1, and when it approaches the upper limit, the penalty term increases, inhibiting the repeated selection of variables highly correlated with the variables already used in the path. The embodiment completes the above decision without introducing any new adjustable weights, ensuring that the selection logic is stable and consistent with the previous evidence.

[0059] Within the entire forest range, the embodiment accumulates the weighted scores of the selected variables at each node to form a forest-level variable importance sequence, and outputs a representative observation variable set from high to low based on this, which is directly used for subsequent structural equation modeling. To reduce sampling accidents, the training uses bootstrap resampling, with a sample size of about 63% of the original data set and an out-of-bag sample ratio of about 37%; the forest size is determined after comprehensive verification from four alternative sizes: 50, 100, 200, and 300; the empirical upper limit of the representative observation variable is not higher than 30, in order to control the dimension and stability of the subsequent measurement model.

[0060] The embodiment breaks through the conventional practice of traditional random forests, which completely rely on random selection or fixed parameters in variable sampling and node division, by introducing an integrated mechanism of core factor priority probability and redundancy suppression, achieving dynamic self-calibration of variable importance without increasing artificial adjustment and hyperparameters.

[0061] The core of the embodiment is that the statistical evidence provided by the pre-order correlation analysis and the out-of-bag error increment is used to generate variable sampling weights with probabilistic significance, so that high-contribution variables are more likely to be selected by nodes; at the same time, the redundancy penalty is used to control the repeated selection of collinear variables at the node level, so as to improve the interpretability and sparsity while ensuring the stability of the model. The structure makes the variable selection process have both global correlation guidance and local error minimization characteristics, forming an adaptive decision tree integration mechanism. Compared with the traditional random forest which only selects the division variable based on information gain or pure error index, the scheme realizes the robust screening of high-dimensional features in the karst basin multi-factor coupling scene through the joint design of probabilistic feature sampling and redundancy self-inhibition, and significantly improves the interpretability of the model and the reliability of the variable importance evaluation.

[0062] To investigate the spatial differences in runoff and sediment yield and the influencing factors, the present embodiment takes the 40 sub-basins of the Xijiang and Wujiang rivers in the whole study area as the object. In recent years, to control serious soil and water loss and restore the degraded environment, the long-term, policy-driven green food project has been implemented in the region since the late 1990s / early 2000s, which changes farmland into forest or grassland under the restoration of natural vegetation. The sediment yield has a significant downward trend in the early 21st century and becomes relatively stable during the study period (as shown in Figure 2 ). Therefore, the data records of the period 2009-2012 are selected (as shown in Figure 3 ).

[0063] The PLS-SEM model can maximize the explained variance of the relevant structure and has lower requirements for sample size and data distribution. The present embodiment aims to use the advantages of the PLS-SEM model to provide a comprehensive and detailed understanding of the factors controlling IC in the karst basin. To determine the main factors controlling the changes in runoff and sediment yield in the 40 karst basins during 2009-2012, Pearson correlation analysis is first used to quantify the relationship between runoff, sediment yield, and their potential variables. Then, random forest (RF) analysis is performed using the factors that have a significant impact on runoff and sediment yield changes. Finally, the main influencing factors of runoff and sediment yield changes selected by the RF algorithm are input into the PLS-SEM model to quantitatively analyze the effects of climate, lithology, soil, topography, and landscape on the changes in runoff and sediment yield in the basin.

[0064] The study selected 103 potential variables that affect runoff and sediment yield, which can be divided into five types: climate (17), lithology (1), soil (9), topography (27), and landscape (49). To decouple the complex relationships between runoff, sediment yield, and potential variables, PLS-SEM model analysis was conducted using R and the "PLSPM" package. This appendix systematically reviews the various potential factors that affect sediment yield in karst basins. Among them, climate factors include consecutive dry days (CDD), consecutive wet days (CWD), potential evapotranspiration (PE), different intensity rainfall indicators (such as R25, R50, RX1, etc.), and temperature-related variables; topographic factors involve elevation (E), slope (S), aspect (A), topographic wetness index (TWI), stream power index (SPI), and basin morphological characteristics (such as area, shape, river network density, etc.); soil factors include the composition of each particle size (GC, SC, STC, CC), bulk density (SBD), organic carbon (SOC), pH value, and other physical and chemical properties; landscape indices describe landscape structure from multiple angles, including patch scale (NP, PD, LPI), shape (LSI, SHAPE), spatial configuration (CONTIG, COHESION, AI), and diversity (SHDI, SIDI); and lithology focuses on the coverage of carbonate rocks (CBC). These factors collectively form a comprehensive parameter system for analyzing sediment transport mechanisms in karst regions. This example selects the PLS-SEM model to explore the interactive conceptual model between multiple variables, assuming that the potential variables (climate, lithology, soil, topography, and landscape factors) have a significant impact on the dependent variables (runoff and sediment yield).

[0065] This example selects 40 basins in the karst region of Southwest China as the research object to identify the main controlling factors of basin runoff and sediment yield. Table 5-1 shows the basic statistical information of the location, basin area, average runoff, and sediment yield of the 40 basins. The basin area of the 40 basins ranges from 420 km 2 to 339175 km 2 , with the smallest basin area being No. 17 Fuyang and the largest being No. 19 Gaoyao. The average runoff and sediment yield of the basins are based on four years (2009-2012) of average data. During the study period, the runoff of the 40 basins varied from 146 to 1445 mm, with the smallest and largest basins being No. 37 Xiqiao and No. 20 Guilin, respectively. The sediment yield varied from 2.0 to 246 t km -2 , as shown in Tables 1 and 2. Overall, with the increase of basin area, both runoff and sediment yield tend to decrease. For example, No. 20 Guilin (2527 km 2 ), No. 25 Lipu (892 km 2 ), and No. 33 Taiping (2744 km 2) were relatively small, while the runoff was 1445, 1038 and 1201 mm, and the sediment discharge was 102, 170 and 164 t km –2 , respectively. It can be seen that the runoff and the annual average sediment discharge were relatively large. Correspondingly, the 34th Tian'e basin (104772 km 2 ) and the 35th Wuzhou basin (314326 km 2 ) were relatively large, but the runoff (347, 514 mm) and the sediment discharge (2, 28 t km -2 ) were relatively small.

[0066] Table 1 Basic statistical information table of 20 hydrological stations Site Longitude (E) Latitude (N) Area (km 2 )]]> Mean runoff (mm) Average sediment discharge (t km -2 )]]> 1 Bajialou 108°19′ 29°27′ 3903 731 210 2 Hefeng 106°55′ 26°56′ 420 446 89 3 Lijutang 107°16′ 27°26′ 4821 376 16 4 Qixiangguan 104°57′ 27°9′ 2880 281 148 5 Shipantang 106°12′ 27°9′ 1441 325 53 6 Shiqian 108°13′ 27°31′ 758 479 34 7 Sinan 108°15′ 27°56′ 50352 398 9 8 Wulong 107°44′ 29°20′ 83035 475 29 9 Yanhe 108°30′ 28°34′ 54412 387 16 10 Yangchang 105°11′ 26°39′ 2438 392 122 11 Changba 107°41′ 28°48′ 5499 459 132 12 Baiben 107°51′ 25°59′ 1442 640 86 13 Caotouping 104°57′ 25°52′ 4957 998 128 14 Dadukou 104°43′ 26°17′ 8104 341 246 15 Dahuangjiangkou 110°12′ 23°35′ 276582 504 31 16 Duangting 109°40′ 24°26′ 7456 860 80 17 Fuyang 111°16′ 24°51′ 482 578 31 18 Gaoche 105°40′ 25°52′ 2146 566 76 19 Gaoyao 112°28′ 23°30′ 339175 521 40 20 Guilin 110°19′ 25°14′ 2527 1445 102

[0067] Table 2 Basic statistical information table of another 20 hydrological stations Site Longitude (E) Latitude (N) Area (km 2 )]]> Mean runoff (mm) Average sediment discharge (t km -2 )]]> 21 Jinji 110°50′ 23°13′ 9094 753 163 22 Laocun 110°57′ 24°15′ 1582 659 145 23 Leigongtan 106°35′ 25°25′ 5439 589 24 24 Libo 107°52′ 25°25′ 1283 624 146 25 Lipu 110°24′ 24°3′ 892 1038 170 26 Liuzhou 109°24′ 24°20′ 45469 719 98 27 Maling 104°55′ 25°11′ 2143 469 156 28 Malong 108°19′ 24°14′ 3092 645 74 29 Pingle 110°40′ 24°37′ 11883 920 53 30 Pinglihe 107°3′ 25°50′ 1410 545 35 31 Rongshui 109°15′ 25°4′ 23516 792 96 32 Sancha 108°57′ 24°28′ 16488 639 81 33 Taiping 110°37′ 23°43′ 2744 1201 164 34 Tian'e 107°10′ 24°59′ 104772 347 2 35 Wuzhou 111°20′ 23°28′ 314326 514 28 36 Wuxuan 109°39′ 23°35′ 195730 502 34 37 Xiqiao 103°38′ 25°1′ 3042 146 23 38 Yongwei 109°17′ 25°42′ 12883 546 213 39 Zanyi 103°50′ 25°35′ 572 156 8 40 Zouwei 108°53′ 23°23′ 1867 814 38

[0068] The time variation of runoff ranged from 83 to 1781 mm. The minimum runoff occurred in the 37th Xiqiao basin in 2011, and the maximum runoff occurred in the 33rd Taiping basin in 2010 (as shown in Figure 4 ). The variation of sediment discharge ranged from 1 to 427 t km -2 , with the minimum value in the 39th Zhenyi basin in 2011 and the maximum value in the 33rd Taiping basin in 2012 (as shown in Figure 5 ).

[0069] Among the 103 selected potential variables, most of the factors were significantly correlated with runoff or sediment discharge (P < 0.05) (as shown in Figure 6 ). Among them, * and ** indicate that the potential variables are significantly correlated with sediment discharge at P < 0.05 and P < 0.01; ^ and ^^ indicate that the potential variables are significantly correlated with runoff at P < 0.05 and P < 0.01. It is shown that 60 factors are significantly correlated with runoff, and 23 factors are significantly correlated with sediment discharge. The potential variables of runoff are mainly landscape and climate variables, while the factors that have more influence on sediment discharge are topographic factors. Specifically, the factors that are significantly correlated with runoff are 16 climate factors, 1 lithology factor, 5 soil factors, 8 topographic factors, and 30 landscape factors. The factors that are significantly correlated with sediment discharge are 2 climate factors, 1 lithology factor, 3 soil factors, 10 topographic factors, and 7 landscape factors.

[0070] The descriptive statistics of the selected factors affecting runoff or sediment transport, including mean, median, maximum, minimum, standard deviation (S.D.), and coefficient of variation (CV), are shown in Tables 3 and 4. The terrain factors have stronger variability (CV > 100%) than other factors. The CV of climate factors ranges from 8% to 48%, and the highest daily temperature (Tx) is relatively stable with lower variability (CV < 10%). The CV of soil properties ranges from 2% to 60%, and SBD and pH have low variability (CV < 10%). The CV of terrain factors ranges from 4% to 217%, and the terrain wetness index (TWI), drainage density (DD), and sediment transport capacity index (LS) are relatively stable with lower variability (CV < 10%). The catchment area (AREA), catchment perimeter (CP), stream length (L), stream length (SL), fragment shock level (SM), stream power index (SPI), minimum elevation (Emin), and outlet height (EO) have strong variability (CV > 100%). The CV of landscape factors ranges from 1% to 212%, and the mean shape index (SHAPE_MN), mean fractal dimension index (FRAC_MN), mean perimeter-area ratio (PARA_MN), and perimeter-area fractal dimension (PAFRAC) have weak variability (CV < 10%). The patch number (NP) and landscape shape index (LSI) have strong variability (CV > 100%). Other landscape factor variables show moderate variability (10% < CV < 100%).

[0071] Table 3. The first selected potential variables affecting runoff or sediment transport Variable Mean Median Max Min S.D. CV Climate RDs 102.2 103.4 118.8 70.3 11.7 11% RX3 116.4 108.3 178.1 66.4 28.2 24% CDD 30.7 29.5 58.5 19.3 9.6 31% CWD 7.2 7.2 11.0 4.8 1.5 21% PT 1106.4 1067.8 1621.1 655.3 273.5 25% RD25 12.3 12.1 957.3 5.5 4.4 36% R25 520.9 532.6 957.3 209.6 199.6 38% RD50 3.3 2.8 458.0 0.5 1.6 47% R50 221.5 190.9 458.0 27.3 105.7 48% RI 7.3 7.0 178.1 4.0 1.9 25% RX1 76.9 76.7 94.6 46.6 8.6 11% RX2 100.5 95.6 145.1 56.4 20.5 20% RX4 131.0 121.6 203.4 70.4 32.5 25% RX5 140.2 133.5 212.0 75.4 32.6 23% T 17.2 16.5 21.6 12.7 2.5 14% Tx 28.7 29.1 50.0 23.9 2.4 8% Lithology CBC 54.4 57.2 93.2 0.8 28.9 53% Soil pH 5.7 5.6 6.5 5.1 0.3 6% CAC 1.2 1.0 3.2 0.2 0.7 60% EC 0.1 0.1 0.2 0.1 0.0 11% STC 32.0 31.3 44.3 25.5 3.7 11% CC 34.5 34.2 44.3 23.2 5.5 16% SBD 1.3 1.3 1.6 1.3 0.0 2% SOC 1.3 1.3 6.5 1.1 0.2 12% Topography S 17.8 18.0 23.8 11.0 2.7 15% TWI 6.0 6.0 6.7 5.5 0.3 4% AREA 40277 4365 339175 420.0 86340 214% CP 1136 428.9 6579 102.5 1637 144% L 237.3 114.2 1082 30.3 278.3 117% FF 0.3 0.3 0.5 0.2 0.1 22% SL 19854 1998 169083 207.2 42986 217% DD 0.5 0.5 0.6 0.4 0.0 7% SM 4025 425.5 34268 38.0 8734 217% LS 64.2 63.1 76.8 56.0 5.3 8% SPI 49170 18642 236604 1721 69597 142%

[0072] Table 3. The second selected potential variables affecting runoff or sediment transport Variable Mean Median Max Min S.D. CV Emin 442.3 200.0 2888.0 2.0 492.3 111% Emax 2144.3 2014.5 2888.0 1114.0 532.9 25% E 984.0 875.9 2076.0 294.9 524.0 53% EO 461.7 237.0 1877.0 35.0 488.3 106% HI 0.3 0.3 93.2 0.1 0.1 35% DF 522.3 454.9 1183.5 104.2 262.1 50% Landscape SHDI 1.0 1.0 1.3 0.6 0.2 16% W 0.6 0.4 2.3 0.0 0.6 100% NP 3150 316.5 26072 57.0 6687 212% LSI 28.0 15.0 108.4 6.1 29.9 107% SHAPE AM 14.0 8.5 54.7 3.7 13.4 95% PR 5.3 5.0 6.0 3.0 0.7 14% RPR 87.5 83.3 100.0 50.0 12.4 14% FL 23.8 22.1 83.6 7.6 8.2 34% WL 57.6 58.3 83.9 29.1 13.8 24% GL 17.0 14.5 47.6 2.5 10.6 63% LPI 50.4 52.9 83.6 11.9 18.8 37% ED 7.9 8.1 10.3 4.1 1.6 20% GYRATE MN 927.4 913.2 1155.4 754.5 101.0 11% GYRATE CV 179.4 186.8 234.0 117.1 32.9 14.1% SHAPE MN 1.3 1.3 1.5 1.2 0.1 6% SHAPE CV 69.1 70.1 91.1 49.1 10.4 15% FRAC MN 1.0 1.0 22.6 1.0 0.0 1% FRAC CV 3.9 3.9 4.8 3.0 0.4 11% PARA MN 34.8 34.7 36.8 32.9 1.0 3% PARA AM 17.1 17.3 22.6 9.4 3.4 20% PARA CV 19.7 19.9 23.4 15.8 1.9 10% CONTIG MN 0.1 0.1 0.2 0.1 0.0 19% CONTIG AM 0.6 0.6 1.7 0.4 0.1 15% CONTIG CV 131.3 125.1 189.1 104.0 19.8 15% PAFRAC 1.6 1.7 76.5 1.6 0.0 2% ENN MN 2801.6 2768.2 3750.8 2253.9 313.5 11% CONTAG 42.9 42.8 66.3 9.1 10.9 25% PLADJ 57.3 56.8 76.5 43.5 8.5 15% DIVISION 0.7 0.7 1.0 0.3 0.2 25% SPLIT 5.2 3.5 22.9 1.4 4.8 92% SIDI 0.6 0.6 1.2 0.3 0.1 19% MSIDI 0.8 0.9 1.2 0.3 0.2 26% SHEI 0.6 0.6 1.0 0.4 0.1 19% SIEI 0.7 0.7 78.39 0.4 0.1 19% MSIEI 0.5 0.5 0.91 0.2 0.2 29% AI 58.8 58.2 78.4 45.3 8.3 14%

[0073] In Tables 3 and 4, S.D. represents the standard deviation, and CV represents the coefficient of variation. The italic font represents the variables significantly related to runoff, the normal font represents the variables significantly related to sediment transport, and the bold font represents the variables significantly related to both runoff and sediment transport.

[0074] Specifically, Pearson correlation analysis was used to select factors significantly related to runoff (P < 0.05). Subsequently, the importance of each factor to runoff was obtained by the RF algorithm (e.g., Fig. 2a and Fig. 2b). Figure 7The first three variables with the highest relative importance were selected as the observed variables of the PLS-SEM framework structure according to the normalized RF importance values. Rainfall (R25), rainy days (RDs), and annual total precipitation (PT) were determined as the observed variables of the climate factor, and pH, SBD, and SOC were determined as the observed variables of the soil factor. EO, elevation (E), and height integral (HI) were determined as the main observed variables of the terrain factor, and farmland (FL), edge density (ED), and weighted average perimeter area ratio (PARA-AM) were determined as the observed variables of the landscape factor.

[0075] The selected important factors were used as the observed variables of the latent variables according to the RF algorithm analysis, and the PLS-SEM model was used to explore the relative importance of climate, lithology, soil, terrain, and landscape in decoupling runoff. The GOF result of the PLS-SEM model was 0.76, which was greater than 0.5, indicating that the model was meaningful. Climate, lithology, soil, terrain, and landscape could collectively explain 79% of the total variation of runoff (as shown in FIG. 6). Figure 8 The results of the path coefficients (β) were in the order of climate > lithology > terrain > soil > landscape. Among them, the climate factors (P, T, and PET) had the greatest impact on runoff, with a β value of 0.589, and had a very significant positive effect (P < 0.01). The climate factor played a dominant role in the change of runoff, while the other factors had a smaller impact. Lithology, soil, terrain, and landscape had a non-significant negative impact on runoff.

[0076] Alternatively, the climate factors (RDs and maximum 3-day precipitation (RX3)) and soil factors (EC, CAC, and pH) that affect sediment discharge were determined through Pearson correlation analysis. In addition, the terrain and landscape factors that control sediment discharge were further obtained through the RF algorithm to finally form the PLS-SEM framework. However, the introduction of the terrain and landscape observed variables determined by the RF algorithm into the PLS-SEM model resulted in a poor final prediction result. Therefore, by simulating various possible PLS-SEM framework types, the reselected terrain and landscape factors were used to obtain the optimal solution of the results. Among them, the form factor (FF) and slope (S) were selected as the main terrain factors, and the landscape shape index (LSI), weighted average shape index (SHAPE-AM), and water area percentage (W) were selected as the main landscape factors.

[0077] The PLS-SEM model was used to explore the relative importance of climate, lithology, soil, terrain, and landscape in decoupling sediment discharge. The GOF result of the model was 0.51, indicating that the model was meaningful. From Figure 9It can be seen that climate, lithology, soil, topography and landscape can jointly explain 59% of the variability of sediment discharge. In addition to the climate factor, lithology, soil, topography and landscape factors have a significant impact on the change of sediment discharge (as shown in Table 5). Among them, the landscape factor has the greatest impact on the change of sediment discharge (P < 0.01, β =-0.458), followed by lithology (P < 0.05, β =-0.337), topography (P < 0.05, β = 0.246), soil (P < 0.1, β =-0.198) and climate (β =-0.005) factors. The topographic factor has a significant positive correlation with sediment discharge, while the lithology, soil and landscape factors have a significant negative correlation with sediment discharge (as shown in Figure 9

[0078] Table 5 Standardized path coefficient effect analysis in PLS-SEM model Variable Direct effect P Direct effect P Runoff Sediment discharge Geomorphology -0.106 0.409 Geomorphology 0.246 0.044 Landscape -0.050 0.734 Landscape -0.458 0.000 Climate 0.589 0.000 Climate -0.005 0.971 Soil -0.079 0.543 Soil -0.198 0.099 Lithology -0.181 0.106 Lithology -0.337 0.023

[0079] With the increase of precipitation, the runoff trend increased significantly, which is consistent with the results of many previous studies. According to the results of RF algorithm, PT, RDs and R25 were selected as the main climate variables, which can reflect the intensity of precipitation and have a significant positive impact on runoff. R25 factor is usually used to represent extreme climate events, so frequent extreme climate may greatly affect the runoff change in the climate-sensitive southwest karst area. The spatiotemporal variation of R25, RDs and PT factors is basically consistent with that of runoff (as shown in Figure 10 The results show that R25, RDs and PT are significantly correlated with runoff (P < 0.01; as shown in Figure 11 In addition, it is found that the larger the proportion of karst area, the lower the runoff and R25 (as shown in Figure 12 Figure 12 , where r represents the significance level of Pearson coefficient; r CBC ​​The results of the partial correlation analysis between R25 and runoff in the case of controlling the proportion of karst area are shown in Table 4. This can be attributed to the complex hydrogeological conditions in karst areas, where precipitation enters the groundwater system through numerous surface pores and cracks. In general, surface runoff occurs on karst slopes only when the precipitation is greater than 60 mm. However, extreme rainfall can quickly saturate the surface karst area infiltration, leading to the occurrence of surface runoff. The high hydrological connectivity in karst areas further promotes the influence of climate factors on runoff. Due to the high rainfall intensity in this region, the early soil moisture in the basin is relatively high. This reduces the subsequent water infiltration capacity, playing a key role in the generation and intensity of runoff. In the karst areas of the Mediterranean, extreme rainfall can also cause significant variations in river flow. For non-karst basins, previous studies have also shown that climate is the main factor affecting runoff changes. Overall, the runoff in karst areas is mainly influenced by climate, especially extreme climate, while the influence of lithology, soil, topography, and landscape on runoff is not significant.

[0080] Climate, lithology, soil, topography, and landscape factors can collectively explain 59% of the total variation in sediment yield. It is worth noting that the sediment yield in karst areas is more susceptible to special geological, soil properties, topography, and highly heterogeneous landscapes, while the influence of climate factors on sediment yield changes in karst basins is smaller. Overall, normal precipitation events generally have no significant impact on soil erosion in this region. In extreme rainfall or erosive precipitation events, erosive sediment is generally transported to the underground system through seepage structures such as cracks and sinkholes. Therefore, the eroded sediment may block the sinkhole outlets in karst depressions, leading to flooding disasters. In addition, a large amount of eroded sediment is deposited in karst depressions or transported to another basin through underground channels, which will reduce the impact of precipitation on sediment yield changes. Compared with runoff, sediment yield is less sensitive to extreme rainfall factors (r = 0.33; as shown in Table 4, where r represents the significance level of the Pearson coefficient; r Figure 13 Figure 13 CBC The results of the partial correlation analysis between RX3 and sediment yield in the case of controlling the proportion of karst area are shown in Table 4. In addition, the partial correlation analysis shows that the correlation between runoff and extreme rainfall is significant when considering the proportion of karst area in the basin as a control variable, while the correlation between sediment yield and extreme rainfall is not significant. Previous studies have also shown that the impact of extreme rainfall on runoff is greater than that on sediment yield. Factors affecting sediment yield changes need to consider the complex topography and highly heterogeneous landscape of karst systems, and the impact of precipitation and other external conditions on sediment yield changes in different karst basins is smaller.

[0081] ​​Generally, sediment yield is closely related to lithological characteristics. The higher the coverage of soluble carbonate rock, the less surface runoff and sediment yield. In addition, the soil thickness in karst areas is much smaller than that in non-karst areas, and the spatial distribution of soil is extremely uneven and discontinuous. Eroded sediment is mainly deposited in low-lying areas such as karst depressions, resulting in a continuous decrease in erodible soil in steep areas. Topographic factors can provide kinetic energy for the transport of surface sediment, affecting not only soil thickness but also the occurrence and intensity of sediment transport. These unique topographic features strongly influence the generation and transport of sediment to rivers in karst areas.

[0082] It is worth emphasizing that landscape factors have the greatest impact on sediment yield. As expected, the spatial variations of W, SHAPE-AM, and LSI are consistent with sediment yield, and there is a significant correlation (P < 0.05, as shown in Figure 11 and Figure 14 Overall, landscape is an important factor affecting the process of soil and water loss in a watershed. From the perspective of landscape ecology, the landscape structure of a watershed determines the surface nutrients, soil erosion, and sediment transport. In karst areas, the widespread distribution of soluble carbonate rocks, the discontinuous distribution of soil, and the spatial fragmentation of plant growth lead to high landscape heterogeneity. In recent years, China has launched a series of ecological restoration projects (such as forest conversion and karst rock desertification). This ecological engineering reduces the risk of rock desertification in karst areas by improving the stability of the ecosystem, which is conducive to the sustainable development of the ecological environment in karst areas.

[0083] Forest and grassland have a good impact on soil and water conservation, while the large-scale expansion of farmland leads to massive soil erosion. Prior to the 21st century, changes in land cover types in the karst areas of Southwest China exacerbated the processes of runoff and sediment storage and transport. After 2000, various ecological restoration engineering measures, including vegetation restoration and reconstruction, soil and water conservation, water resource development, and soil quality improvement, reduced the risk of soil erosion and rock desertification in the karst areas of Southwest China. However, the soil loss tolerance in karst areas is only 30-68 t km -2 , and once the soil is eroded, it is difficult to recover. In the future, on the one hand, we must consider the distribution and quantity of cultivated land and pay more attention to controlling forest cutting and other behaviors to further reduce the occurrence of soil erosion. On the other hand, in the karst areas of Southwest China, we should maintain a highly heterogeneous and fragmented landscape to reduce surface runoff and soil erosion.

[0084] It is worth noting that the main factors affecting sediment yield vary between karst and non-karst regions. Topography and land use are the main factors affecting sediment yield in the Loess Plateau, one of the most severely eroded regions in China and globally. In recent years, a series of soil and water conservation measures have been implemented in the Loess Plateau, including engineering (dams, reservoirs), biological (afforestation and grass planting), and agricultural measures (no-tillage and crop rotation). The overall trend of vegetation coverage is increasing, and the main factor of sediment yield reduction gradually shifts from terrace construction, dam enclosure, and precipitation change to vegetation restoration. Through the analysis of the factors controlling sediment yield in the Godavari Basin of the Indian Peninsula, it is concluded that topography, lithology, and land use have a great influence on sediment yield variation. In the hilly red soil region of southern China, land use composition and pattern, and topography have a significant impact on specific sediment yield. However, the relative importance of different factors on the variability of sediment yield in heterogeneous basins has not been quantified in previous studies.

[0085] In this example, the average sediment yield of 40 karst basins ranges from 2 to 246 t km -2 As shown in Tables 1 and 2, it is much lower than that of other non-karst basins. The average sediment yield in the red soil region of the Loess Plateau in southern China is much higher than that in karst regions. In karst regions, most of the soil covering the rock surface is eroded. However, the soil formation rate is much lower than the erosion rate. The high proportion of carbonate rocks in karst regions and the widespread distribution of cracks and pipes not only allow precipitation to penetrate into the underground karst system and store it underground, but also affect the formation of surface runoff and reduce sediment transport driven by surface runoff. Some eroded soil fills the karst pipes, blocking the drainage outlets of karst depressions. Although the sediment yield in karst basins is relatively low, the risk of soil erosion is still high.

[0086] For karst regions, such as the Mediterranean region of Europe, terrace farming and sustainable agriculture systems based on irrigation are widely used. Adopting new vegetation management patterns and soil conservation plant species is an important measure for soil resource protection in karst regions. Karst and karst development in the southwest of China are more severe than in the Mediterranean region, and more attention should be paid to it. Embedding farmland into karst basins with a variety of landscape types can effectively reduce the risk of soil erosion through appropriate soil and water conservation measures. In addition, attention should also be paid to soil erosion on low-slope land with low forest and grass coverage, and reducing the impact of underground pores through engineering and plant measures may be an effective method. In addition, developing rich tourism resources related to karst landforms may be a way to reduce land demand. This approach can improve the fragile ecological environment. This example highlights the advantages of PLS-SEM models in decoupling the relationship between runoff, sediment yield, and its potential variables. The research results provide a scientific basis for effective management of soil and water resources in karst basins.

[0087] Specifically, the embodiment explores the variation characteristics of runoff and sediment discharge in different karst basins, quantifies the coupling effect of multiple environmental factors and water and sediment processes in 40 karst basins by using a PLS-SEM model, and reveals the dominant role of landscape heterogeneity in sediment discharge. The research results help to deepen the understanding of the characteristics of runoff and sediment in karst basins, and provide scientific basis for effective control of water and soil loss and sustainable development of ecological environment. The main results are as follows:

[0088] (1) The Pearson correlation analysis results show that among the 103 variables of climate, lithology, soil, topography and landscape factors, most of the factors are significantly correlated with runoff or sediment discharge. Among them, 60 factors are significantly correlated with runoff, mainly landscape and climate factors; 23 factors are significantly correlated with sediment discharge, and the factors that affect sediment discharge more are topography and landscape factors.

[0089] (2) The PLS-SEM model is used to quantitatively distinguish the driving factors of the variation of runoff and sediment discharge in karst basins. Climate, lithology, soil, topography and landscape factors can jointly explain 79% of the variation of runoff. Climate factors have the greatest influence on runoff, with a significant positive effect, while lithology, soil, topography and landscape have a non-significant negative effect on runoff; runoff, climate, lithology, soil, topography and landscape can jointly explain 59% of the variation of sediment discharge. Except for climate factors, lithology, soil, topography and landscape factors have a significant influence on sediment discharge, with landscape factors having the greatest influence on sediment discharge, followed by lithology, topography, soil and climate. Topography factors have a significant positive correlation with sediment discharge, while lithology, soil and landscape factors have a significant negative correlation with sediment discharge.

[0090] Corresponding to the above method, the embodiment also provides a system for decoupling influencing factors of spatial variation of sediment discharge in a karst basin based on model combination, comprising:

[0091] A basic data construction unit is configured to construct a basic data set of a target karst basin;

[0092] A latent variable mapping unit is configured to establish a correspondence between latent variables and observed variables under the framework of five types of factors, i.e. climate, lithology, soil, topography and landscape, according to the basic data set;

[0093] A correlation screening unit is configured to perform Pearson correlation analysis with the observed variables as input and the sediment discharge as output, select the observed variables with significant correlation, and generate a candidate factor set;

[0094] A variable optimization unit is configured to establish a random forest model with the candidate factor set as input, evaluate the importance of variables according to the variation of out-of-bag error, and select representative observed variables;

[0095] A structural equation modeling unit is configured to construct a partial least squares structural equation model by taking climate, lithology, soil, terrain and landscape as latent variables and combining the representative observation variables, to calculate path coefficients and effect strengths, and to obtain model estimation results;

[0096] A decoupling analysis unit is configured to analyze the action direction and contribution rate of each latent variable on the spatial variation of sediment discharge based on the model estimation results, to determine a dominant influencing factor, and to output decoupling results.

[0097] The principles and implementation manners of the present application are described herein by using specific examples, and the above examples are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, the specific implementation manners and application ranges will be changed according to the idea of the present application. In summary, the content of the present specification should not be understood as a limitation on the present application.

Claims

1. A method for decoupling factors influencing spatial variation of sediment transport in karst watersheds based on model combination, characterized in that, include: Construct the basic dataset for the target karst watershed; Within the framework of five factors—climate, lithology, soil, topography, and landscape—a correspondence between latent variables and observed variables is established based on the aforementioned basic dataset. Using the observed variables as input and sediment transport as output, Pearson correlation analysis was performed to select the observed variables with significant correlations and generate a candidate factor set. A random forest model is established using the candidate factor set as input, and representative observed variables are selected based on the importance of the evaluation variables for changes in out-of-package error. Using climate, lithology, soil, topography, and landscape as latent variables, a partial least squares structural equation model was constructed in conjunction with the aforementioned representative observed variables. The path coefficients and effect strengths were calculated to obtain the model estimation results. Based on the model estimation results, the influence direction and contribution rate of each potential variable on the spatial variation of sediment transport are analyzed, the dominant influencing factors are determined, and the decoupling results are output.

2. The method for decoupling the influencing factors of spatial variation of sediment transport in karst watersheds based on model combination as described in claim 1, characterized in that, The basic dataset for the watershed in the study area is constructed, including: The target karst basin is divided into several sub-basins as spatial analysis units, and the research period for data processing is determined to form a benchmark framework for data registration and summarization. Precipitation and temperature observation data were obtained from the target karst basin and surrounding meteorological stations. Potential evapotranspiration was calculated based on the Penman formula. Temporal and spatial registration was completed in the sub-basin and study period dimensions to obtain a set of meteorological variables. The carbonate rock cover of each sub-basin was determined based on the lithological map, and the lithological variables were obtained. Particle size distribution, bulk density, electrical conductivity, calcium carbonate content, pH, and organic carbon content were extracted from the World Soil Database and registered in the dimensions of sub-watershed and study period to obtain a set of soil variables. Elevation and slope were extracted based on the digital elevation model, and the sub-basin scale was summarized to obtain a set of topographic variables. Landscape indices were calculated based on land use remote sensing and annual land cover data, and a set of landscape variables was obtained by summarizing at the sub-basin scale. Observational sequences of runoff and sediment transport were obtained from watershed hydrological stations and published annual reports, and registered at the sub-watershed and study period dimensions as response variables. Missing term processing and standardization are performed on the meteorological variable set, the lithological variable set, the soil variable set, the topographic variable set, the landscape variable set, and the response variable to obtain the basic dataset.

3. The method for decoupling the influencing factors of spatial variation of sediment transport in karst watersheds based on model combination as described in claim 1, characterized in that, Within the framework of five factors—climate, lithology, soil, topography, and landscape—the correspondence between latent variables and observed variables is established based on the aforementioned basic dataset, including: Climate, lithology, soil, topography, and landscape are used as five categories of potential variables, which serve as the superordinate variables for the subsequent measurement model. Based on the aforementioned basic dataset, a list of observed variables is extracted according to the latent variable categories to form a candidate observed variable library; Each observed variable in the candidate observed variable library is mapped to the corresponding latent variable category according to its physical meaning and data source, forming a hierarchical mapping of "latent variable - observed variable" one-to-one attribution; Under the unified caliber of sub-basins and study period, the corresponding entries of the basic dataset are called to verify and solidify the hierarchical mapping, thereby obtaining the correspondence between latent variables and observed variables.

4. The method for decoupling the influencing factors of spatial variation of sediment transport in karst watersheds based on model combination as described in claim 3, characterized in that, The candidate observation variable pool consists of 103 observation variables: 17 climate variables, 1 lithology variable, 9 soil variables, 27 topographic variables, and 49 landscape variables.

5. The method for decoupling the influencing factors of spatial variation of sediment transport in karst watersheds based on model combination as described in claim 1, characterized in that, Using the observed variables as input and sediment transport as output, a Pearson correlation analysis was performed. Observed variables with significant correlations were selected to generate a candidate factor set, including: The observed variable matrix and the corresponding sediment transport sequence are retrieved from the aforementioned basic dataset as input data for correlation analysis; The Pearson correlation coefficient and significance level between each observed variable and sediment transport were calculated one by one to obtain a set of paired results for correlation coefficient and P-value. Based on the significance determination criteria, observed variables with significant correlation were selected to form a candidate factor list; the significance determination was based on a p-value < 0.05 as the inclusion criterion. The list of candidate factors is organized into the candidate factor set.

6. The method for decoupling the influencing factors of spatial variation of sediment transport in karst watersheds based on model combination as described in claim 1, characterized in that, A random forest model is built using the candidate factor set as input. The importance of variables is evaluated based on the change in out-of-package error, and representative observed variables are selected, including: Using the observed variable matrix corresponding to the candidate factor set as the independent variable and the sediment transport sequence as the dependent variable, a dataset for regression modeling is constructed. Based on the dataset used for regression modeling, a training subset is generated by bootstrapping resampling, and the random forest model consisting of multiple regression decision trees is trained according to the training subset, while retaining the corresponding out-of-bag samples; For each of the observed variables, a perturbation is applied to the out-of-bag samples, and the increment of the out-of-bag mean square error before and after the perturbation is calculated to obtain the variable importance score of each of the observed variables; The variables are sorted from highest to lowest importance score, and the representative observed variables are determined based on the sorting results.

7. The method for decoupling the influencing factors of spatial variation of sediment transport in karst watersheds based on model combination as described in claim 1, characterized in that, Using climate, lithology, soil, topography, and landscape as latent variables, and combining these with representative observed variables, a partial least squares structural equation model was constructed. Path coefficients and effect strengths were calculated to obtain model estimation results, including: The representative set of observation variables is called, and the measurement blocks are divided according to climate, lithology, soil, topography and landscape according to the correspondence. The observation variable matrix for modeling is generated, and the sediment transport is used as the dependent variable to complete the setting of the measurement model and the structural model. Climate, lithology, soil, topography, and landscape are used as potential variables in the structural model. A partial least squares structural equation model is established, and the measurement model and the structural model are jointly configured. The external weights and path coefficients were estimated using a partial least squares iterative algorithm to obtain the direct, indirect, and total effects of each latent variable on sediment transport, and the significance level was calculated. The goodness of fit and the explanatory power of sediment transport are calculated, and the path coefficients, effect decomposition results and evaluation indicators are output as the basis for determining the dominant factors and ranking their contribution rates, and the model estimation results are obtained.

8. The method for decoupling the influencing factors of spatial variation of sediment transport in karst watersheds based on model combination as described in claim 1, characterized in that, Based on the model estimation results, the influence direction and contribution rate of each latent variable on the spatial variation of sediment transport are analyzed, the dominant influencing factors are identified, and the decoupling results are output, including: The direction of action is determined based on the sign of the path coefficient corresponding to the sediment transport volume as the dependent variable, and the corresponding significance level is used as the basis for validity. Using the total effect of each latent variable on the sediment transport as the basis for relative contribution, the contribution rates of the five categories of latent variables—climate, lithology, soil, topography, and landscape—are ranked from high to low, and their significance is marked. The leading and significant latent variables are identified as dominant influencing factors, and together with their direction of action, contribution rate, significance, goodness of fit, and explanatory value, they form the decoupling results for use in watershed soil and water conservation and landscape optimization decisions.

9. The method for decoupling the influencing factors of spatial variation of sediment transport in karst watersheds based on model combination as described in claim 1, characterized in that, The random forest model employs a self-calibrating structure with core factor priority probability and redundancy suppression for each candidate observed variable. First, a priority probability is constructed based on Pearson correlation analysis and out-of-package error increment. Then, feature sampling is performed at each tree node using this priority probability. Finally, the partitioning variable is selected according to the redundancy suppression weighted partitioning criterion, and the importance weight of the variable in the forest is determined. The expression for the redundancy suppression weighted partitioning criterion is: ;in, From the formula Determine, and make on each node Take the largest value as the partitioning variable for the current node, and simultaneously set the values ​​of each node... Accumulation serves as a forest-level measure of variable importance; among which, To use variables in the current node The decrease in mean square error resulting from the partitioning; For variables The priority probability of the core factor; This is the set of variables that have been selected as partitioning variables along the path from the root to this node in the current tree; For variables variables in the path set The Pearson correlation coefficient; For variables Pearson correlation coefficient between sediment transport and sediment transport volume; To test variables on out-of-package samples The increment of mean square error caused by scrambling.

10. A decoupling system for factors influencing spatial variation of sediment transport in karst watersheds based on model combination, characterized in that, include: Basic data construction unit, used to build the basic dataset for the target karst watershed; The latent variable mapping unit is used to establish the correspondence between latent variables and observed variables based on the basic dataset within the framework of five factors: climate, lithology, soil, topography, and landscape. The correlation screening unit is used to perform Pearson correlation analysis with the observed variable as input and sediment transport as output, select the observed variable with significant correlation, and generate a candidate factor set. The variable selection unit is used to build a random forest model with the candidate factor set as input, evaluate the importance of variables based on the change of out-of-package error, and select representative observed variables. The structural equation modeling unit is used to construct a partial least squares structural equation model using climate, lithology, soil, topography and landscape as latent variables, combined with the representative observed variables, to calculate path coefficients and effect strengths, and obtain model estimation results. The decoupling analysis unit is used to analyze the direction and contribution rate of each potential variable on the spatial variation of sediment transport based on the model estimation results, determine the dominant influencing factors, and output the decoupling results.

Citation Information

Patent Citations

  • Flood seasonal driving mechanism identification and attribution method based on partial differential equation

    CN115034548A

  • Domain speech recognition method and system based on RAG

    CN119296516A

  • Karst forest soil carbon cycle feature extraction method based on multi-modal data fusion

    CN120164550A