A method for identifying and analyzing driving mechanism of river-lake reverse inflow based on interpretable machine learning

By dividing the long-sequence hydrological data of the river-lake system into characteristic periods and constructing models, the backflow of rivers and lakes is identified and the driving factors are quantified. This solves the shortcomings of existing technologies in terms of identification accuracy and driving mechanism analysis, and achieves high-precision backflow identification and driving mechanism analysis.

CN121901900BActive Publication Date: 2026-06-19HOHAI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HOHAI UNIV
Filing Date
2026-03-25
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify and quantify key driving factors and their physical thresholds for backflow in rivers and lakes, and also find it difficult to analyze the evolution of the driving mechanisms.

Method used

Long-series hydrological data from the Jianghu system were collected, and input feature sets were generated by dividing the data into characteristic periods. A backflow identification model was constructed using a gradient boosting decision tree model combined with class weights and Bayesian optimization. SHAP interpretation analysis was then performed to determine key driving factors and physical thresholds.

Benefits of technology

It improves the accuracy and stability of river and lake backflow identification, can quantify key driving factors and their physical thresholds, analyze the impact of engineering regulation on backflow mechanism, and compare changes in driving mechanism across time periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901900B_ABST
    Figure CN121901900B_ABST
Patent Text Reader

Abstract

This invention discloses a method for identifying and analyzing the driving mechanism of river and lake backflow based on interpretable machine learning, belonging to the fields of hydrological and water resources analysis and artificial intelligence application technology. The method includes: collecting long-sequence hydrological data of the river and lake system; dividing the time series into different characteristic periods using the operational nodes of large-scale water conservancy projects as boundaries; constructing a multi-dimensional input feature set; training a backflow identification model based on a gradient boosting decision tree using a class weight balancing strategy and a Bayesian optimization algorithm; calculating the marginal contribution value of each hydrological driving factor using the SHAP interpretability method; identifying the physical threshold inducing backflow by combining a SHAP dependency graph with a binary box plot; and comparing the analysis results of different periods to reveal the evolution law of the backflow driving mechanism. This invention effectively solves the problems of traditional methods' difficulty in capturing nonlinear hydrological responses and the lack of physical mechanism explanation in black-box models, and can accurately identify backflow events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of hydrological and water resources analysis and artificial intelligence applications, and in particular to a method for identifying and analyzing the driving mechanism of river and lake backflow based on interpretable machine learning. Background Technology

[0002] There is a continuous water exchange relationship between large rivers and connected lakes. Under natural hydrological conditions, lake water is mostly discharged from the lake area into the main stream. When the river water level is higher than the lake water level, the direction of water flow in the connecting channel reverses, forming a backflow process from the river into the lake. This process directly affects lake replenishment, water level processes, sediment transport, and aquatic ecological patterns. Existing research on river-lake backflow usually uses cross-sectional flow monitoring, correlation statistical analysis, or numerical hydrodynamic models for judgment. Cross-sectional flow monitoring relies on existing station deployments and has limited historical coverage. Correlation statistical analysis is difficult to handle nonlinear responses under multi-factor coupling conditions. Although numerical hydrodynamic models have a strong mechanistic basis, parameter calibration is labor-intensive, and cross-period comparative analysis is complex. With the introduction of machine learning methods into the hydrological field, the accuracy of multi-factor classification analysis has improved, but most models only provide identification results and cannot further provide the contribution of driving factors, key thresholds, and their migration patterns before and after engineering operation.

[0003] For example, CN112229771A discloses a remote sensing monitoring method for river water backflow into lakes based on suspended matter tracing. This method inverts the suspended matter concentration distribution through satellite remote sensing reflectance and determines whether river water has backflowed into the lake based on the movement of the concentration change boundary line. The key points are remote sensing monitoring and spatial boundary identification. Although this method can determine whether backflow has occurred, it does not establish a phased identification model based on long-sequence hydrological data, nor does it interpret and analyze the marginal contribution and physical threshold of backflow driving factors.

[0004] For example, CN115728463A discloses an interpretable water quality prediction method based on semi-embedded feature selection. This scheme interprets the importance of watershed characteristics based on machine learning regression algorithms and SHAP values, focusing on water quality prediction and explanation of driving factors. Although this scheme introduces the SHAP interpretation approach, its object is water quality prediction and does not involve the identification of river and lake backflow, phased modeling, identification of backflow physical thresholds, and evolution analysis of driving mechanisms before and after the operation of large-scale water conservancy projects.

[0005] Therefore, existing technologies lack a technical solution that can not only accurately identify river and lake backflow, but also quantify key driving factors and their physical thresholds in stages before and after the operation of large-scale water conservancy projects, and further provide the results of the evolution of the driving mechanism. Summary of the Invention

[0006] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract and title of the invention. Such simplifications or omissions shall not be used to limit the scope of the present invention.

[0007] In view of the aforementioned existing problems, the present invention is proposed.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0009] As a preferred embodiment of the method for identifying and analyzing the driving mechanism of river and lake backflow based on interpretable machine learning described in this invention, the method involves: collecting long-sequence hydrological data of the river and lake system, and dividing the hydrological data into characteristic periods according to the construction and operation time nodes of large-scale water conservancy projects in the upper reaches of the river, thereby obtaining a phased hydrological dataset.

[0010] An input feature set is generated based on the phased hydrological dataset, and the input feature set includes hydraulic gradient features, flow characteristics, flow ratio features, trend features, lag features, time features, and topographic features.

[0011] The input feature set is divided into a training sample set and a test sample set. The input features are normalized based on the training sample set, and the normalization parameters are used for feature transformation of the test sample set.

[0012] The training sample set is input into the gradient boosting decision tree model, and the backflow samples and non-backflow samples are trained by combining class weights and balance constraints. The backflow identification model is obtained through Bayesian optimization.

[0013] Input the hydrological data to be analyzed for the time period into the backflow identification model to obtain the backflow identification results;

[0014] The backflow identification model is subjected to SHAP interpretation analysis. Based on the marginal contribution value of each input feature, SHAP dependency graph and binary box plot, the key driving factors and physical thresholds are determined. The evolution results of the driving mechanism are generated based on the key driving factors and physical thresholds at different times.

[0015] As a preferred embodiment of the method for identifying and analyzing the driving mechanism of river and lake backflow based on interpretable machine learning described in this invention, the river and lake system includes rivers and lakes connected to the rivers, and a backflow event is defined as a flow value at the connection channel between the river and the lake being less than zero; the characteristic period includes the pre-operation period, the transition period, and the post-operation period, wherein the transition period is the initial water storage operation stage of the large-scale water conservancy project, and the post-operation period is the normal water storage operation stage of the large-scale water conservancy project.

[0016] As a preferred solution of the method for identifying and analyzing the driving mechanism of river-lake backflow based on interpretable machine learning according to the present invention, wherein: the hydraulic gradient features include the external hydraulic gradient at the connection between the river and the lake, the hydraulic gradient in the outlet section of the lake, and the hydraulic gradient inside the lake; the flow features include the main river flow and the total inflow of tributaries into the lake; the flow ratio feature is the ratio of the main river flow to the total inflow of tributaries into the lake; the trend feature is the moving average of the hydraulic gradient; the lag feature is the lag value of the key variable; the time feature is the day count; the terrain feature is the water level change value caused by the change in the riverbed elevation.

[0017] As a preferred solution of the method for identifying and analyzing the driving mechanism of river-lake backflow based on interpretable machine learning according to the present invention, wherein: the moving average duration of the trend feature is 7 days, and the lag durations of the lag feature include 1 day, 3 days, 7 days, and 10 days.

[0018] As a preferred solution of the method for identifying and analyzing the driving mechanism of river-lake backflow based on interpretable machine learning according to the present invention, wherein: the input feature set is divided according to a preset ratio to obtain a training sample set and a test sample set, and the division ratio of the training sample set to the test sample set is 8:2;

[0019] The normalization process uses Z-score standardization, and its calculation formula is:

[0020]

[0021] wherein, represents the original feature value, represents the mean of the corresponding feature, represents the standard deviation of the corresponding feature; represents the standardized feature value.

[0022] As a preferred solution of the method for identifying and analyzing the driving mechanism of river-lake backflow based on interpretable machine learning according to the present invention, wherein: the gradient boosting decision tree model is the CatBoost model, and the class weights are calculated according to the number of backflow samples and the number of non-backflow samples, and the calculation formula is:

[0023]

[0024] wherein, represents the backflow sample weight, represents the number of non-backflow samples, represents the number of backflow samples.

[0025] As a preferred embodiment of the method for identifying and analyzing the driving mechanism of inverted flow in rivers and lakes based on interpretable machine learning described in this invention, the Bayesian optimization includes: setting the value ranges of tree depth, learning rate, regularization coefficient, number of leaf nodes, and number of iterations; randomly selecting initial hyperparameter samples; constructing a Gaussian process surrogate model; selecting the configuration of hyperparameters to be evaluated based on the expected improvement; and performing hierarchical cross-validation. The Gaussian process surrogate model is updated with scores until the iteration termination condition is met, at which point the optimized hyperparameter configuration is output.

[0026] As a preferred embodiment of the method for identifying and analyzing the driving mechanism of Jianghu backflow based on interpretable machine learning described in this invention, wherein: the expected improvement amount and The fractions are calculated using the following formulas:

[0027]

[0028]

[0029]

[0030]

[0031] in, This indicates the hyperparameter configuration to be evaluated; This represents the objective function value corresponding to the hyperparameter configuration to be evaluated; This indicates that the optimal objective function value in the hyperparameter configuration has been evaluated; Indicates accuracy; Indicates recall rate; Indicates the number of days of backflow that were correctly identified; This indicates the number of days that were mistakenly identified as backflow but were not; This indicates the actual number of days of backflow that were missed in the initial assessment; This represents the expected improvement in the hyperparameter configuration h to be evaluated.

[0032] As a preferred embodiment of the method for identifying and analyzing the driving mechanism of Jianghu backflow based on interpretable machine learning described in this invention, the SHAP interpretation analysis includes: calculating the SHAP value of each input feature of each sample, aggregating the mean absolute values ​​of all samples to obtain a global importance ranking, drawing a SHAP dependency graph of key driving factors, and cross-comparing the inflection points in the SHAP dependency graph with the sample boundary regions in the corresponding binary box plots to obtain the physical thresholds corresponding to the key driving factors.

[0033] As a preferred embodiment of the method for identifying and analyzing the driving mechanism of inverted currents in the Jianghu (river / lake) based on interpretable machine learning as described in this invention, the SHAP value is calculated according to the following formula:

[0034]

[0035] in, Indicates the first The SHAP value of each feature; Represents the total set of features; Indicates only the first A single-element set of features; Indicates from the total set of features Remove the first The set of features remaining after considering all features; express It is any feature subset of the remaining feature set, including the empty set, a subset consisting of a single feature, a subset consisting of multiple features, and a subset consisting of all remaining features; This represents the number of features in the feature subset; Indicates the total number of features; Indicates containing the first Model prediction output for each feature; Indicates that it does not include the first The model predicts output for each feature.

[0036] The beneficial effects of this invention are:

[0037] 1. This invention organizes long-sequence hydrological data in stages and incorporates hydraulic gradient features, flow characteristics, flow ratio features, trend features, lag features, time features and topographic features into the input feature set to form a multi-dimensional representation system corresponding to the hydrological exchange process between rivers and lakes. This solves the problem that it is difficult to cover the conditions for backflow formation under the condition of single variable analysis and improves the completeness of the hydrological information corresponding to backflow identification.

[0038] 2. This invention corrects the sample distribution bias under the condition of scarce backflow samples by introducing class weights and Bayesian optimization into the CatBoost modeling process, and matches the tree depth, learning rate, regularization coefficient, number of leaf nodes and number of iterations with the sample structure. This solves the problems of insufficient recognition ability of backflow minority class samples and fluctuation caused by artificial setting of hyperparameters, and improves the stability and accuracy of backflow recognition results.

[0039] 3. This invention uses SHAP interpretation analysis on the backflow identification model and combines SHAP dependency graphs and binary box plots to determine the physical thresholds corresponding to key driving factors. This solves the problem that existing black box models can only give classification results but cannot give the contribution of driving factors and threshold ranges, thus establishing a correspondence between backflow identification results and specific hydrological conditions.

[0040] 4. This invention solves the problem of difficulty in comparing the mechanism of river and lake backflow under engineering regulation conditions across different periods by ranking the key driving factors and the migration of physical thresholds in the pre-operation, transition, and post-operation periods of the project. It can provide the phased change results of the driving mechanism from hydrological dominance to the combined effect of engineering regulation and topographic change. Attached Figure Description

[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0042] Figure 1 This is a flowchart illustrating the method for identifying and analyzing the driving mechanism of Jianghu backflow based on interpretable machine learning, as shown in this invention.

[0043] Figure 2 This is a schematic diagram showing the performance comparison results of the various models shown in this invention across three periods;

[0044] Figure 3 This is a SHAP summary diagram of the backflow driving factors at different stages as shown in this invention.

[0045] Figure 4 This invention presents the SHAP dependency diagram and binary box plot for the Three Gorges Dam before its operation.

[0046] Figure 5 This invention presents the transitional SHAP dependency graph and binary box plot.

[0047] Figure 6 This invention presents the SHAP dependency graph and binary box plot for the period after the Three Gorges Dam's operation. Detailed Implementation

[0048] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0049] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort should fall within the scope of protection of this invention.

[0050] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0051] According to an embodiment of the present invention, in combination Figure 1 The flowchart shown illustrates a method for identifying and analyzing the driving mechanism of backflow in the Yangtze River basin based on interpretable machine learning, which specifically includes the following steps:

[0052] S1. Collect long-sequence hydrological data from the river-lake system, and divide the hydrological data into characteristic periods based on the construction and operation timelines of large-scale water conservancy projects in the upper reaches of the river, obtaining a phased hydrological dataset. Note that the following points should be noted in this step:

[0053] In a preferred embodiment, the river-lake system includes rivers and connected lakes, and a backflow event is defined as a flow value at the connection channel between the river and the lake being less than zero. The characteristic period includes the pre-operation period, the transition period, and the post-operation period. The transition period is the initial water storage and operation stage of the large-scale water conservancy project, and the post-operation period is the normal water storage and operation stage of the large-scale water conservancy project.

[0054] Specifically, the method for dividing characteristic periods based on time nodes includes the following steps:

[0055] S11. Obtain long-term hydrological monitoring data: Collect long-term daily basic hydrological monitoring data such as water level and flow rate in the study area, and perform preprocessing such as missing value imputation and outlier removal.

[0056] S12. Determine the project operation time nodes: Obtain the key time nodes of large-scale water conservancy projects in the upper reaches of the river, including the start time before reservoir impoundment, the start time of the initial impoundment stage, and the start time of normal impoundment operation.

[0057] S13. Divide the hydrological sequence: Using key time nodes as boundaries, divide the long sequence of hydrological monitoring data into three sub-sequences in chronological order: the period before project operation, the transition period, and the period after project operation.

[0058] S14. Extracting staged backflow labels: In each subsequence, extract the daily flow observation values ​​at the connection between the river and the lake, and mark the dates with flow less than zero as backflow samples (positive class) and the dates with flow greater than or equal to zero as non-backflow samples (negative class), thus obtaining a staged hydrological dataset with attached category labels.

[0059] S2. Generate an input feature set based on the phased hydrological dataset. The input feature set includes hydraulic gradient features, flow rate features, flow ratio features, trend features, lag features, temporal features, and topographic features. Note the following in this step:

[0060] In a preferred embodiment, the hydraulic gradient characteristics include the external hydraulic gradient at the river-lake junction, the hydraulic gradient at the lake outlet, and the internal hydraulic gradient of the lake; the flow characteristics include the main river flow and the total flow of tributaries flowing into the lake; the flow ratio characteristic is the ratio of the main river flow to the total flow of tributaries flowing into the lake; the trend characteristic is the moving average of the hydraulic gradient; the lag characteristic is the lagged value of the key variable; the time characteristic is the daily count; and the topographic characteristic is the change in riverbed elevation and the resulting change in water level.

[0061] Specifically, the method for generating input feature sets based on phased hydrological datasets includes the following steps:

[0062] S21. Calculate the spatial hydraulic gradient and flow characteristics: For the data of each day, calculate the water level difference between the main river station and the lake control station as the external hydraulic gradient, calculate the water level difference between the lake control station and the monitoring station inside the lake as the hydraulic gradient of the lake outlet section, calculate the water level difference between each station inside the lake as the hydraulic gradient inside the lake, and simultaneously extract the flow of the main river and the total flow of the tributaries flowing into the lake on that day.

[0063] S22. Calculate the hydrological ratio and time / topographic features: Calculate the dimensionless ratio of the main stream flow to the total tributary flow; extract the date information of the day and convert it into a day count (1 to 365 / 366) as a time feature; and calculate the corresponding water level change value as a topographic feature based on the measured riverbed elevation data.

[0064] S23. Extracting lag features: Using the current day as the baseline, trace back to the historical direction and extract the hydraulic gradient and flow observations of the previous 1 day, 3 days, 7 days and 10 days as lag features.

[0065] S24. Calculate trend characteristics: Using the current day and past continuous historical data (e.g., the previous 7 days) as a window, calculate the moving average of the hydraulic gradient as a trend characteristic to characterize changes in low-frequency hydrological background.

[0066] S25. Feature Fusion and Alignment: All features of the same date calculated in the above steps are horizontally concatenated to obtain a complete input feature set.

[0067] S3. Divide the input feature set into a training sample set and a test sample set; normalize the input features based on the training sample set, and use the normalization parameters for feature transformation of the test sample set. Note the following in this step:

[0068] Specifically, the input feature set is divided into a training sample set and a test sample set in an 8:2 ratio.

[0069] In a preferred embodiment, when backflow samples are relatively scarce compared to non-backflow samples, a stratified sampling method can be used to partition the samples in order to maintain the relative stability of the positive and negative sample distribution in the training sample set and the test sample set. Specifically, this includes: first, dividing the input feature set into a backflow sample subset and a non-backflow sample subset according to the sample labels; then, extracting training samples and test samples from the backflow sample subset and the non-backflow sample subset according to the same or a preset ratio; finally, merging the extracted training samples to form a training sample set, and merging the extracted test samples to form a test sample set, so that the ratio of backflow samples to non-backflow samples in the training sample set and the test sample set is basically consistent with that of the original sample set.

[0070] Subsequently, normalization parameters for each input feature are calculated based solely on the training sample set, and feature transformations are performed on the training and test sample sets using these normalization parameters to ensure that the training and test samples are in the same feature scale space.

[0071] As an example, the normalization process uses Z-score standardization, which is calculated as follows:

[0072]

[0073] in, Represents the original feature values. This represents the mean of the corresponding feature. This represents the standard deviation of the corresponding feature; This represents the standardized feature value.

[0074] S4. Input the training sample set into the gradient boosting decision tree model, and train the model on both inverted and non-inverted samples by combining class weights and balance constraints. Then, obtain the inverted sample identification model through Bayesian optimization. Note the following in this step:

[0075] Specifically, the machine learning model based on gradient boosting decision trees is the CatBoost model. The CatBoost model uses ordered boosting and a symmetric tree structure, which can effectively handle class features and reduce overfitting.

[0076] It should be noted that the class weighting used in this embodiment is to address the class imbalance problem where the number of backflow event samples is less than the number of non-backflow event samples. By setting class weights, backflow samples receive a greater weight in the loss function. The class weights are calculated based on the number of backflow samples and the number of non-backflow samples, and the calculation formula is as follows:

[0077]

[0078] in, Indicates the weight of the backflow sample. Indicates the number of non-backflow samples. This indicates the number of samples that were reversed.

[0079] In a preferred embodiment, obtaining the backflow identification model through Bayesian optimization is a process of optimizing the model's hyperparameters using the Bayesian optimization algorithm, including the following steps:

[0080] S41. Definition of Hyperparameter Space: Define the range of values ​​for hyperparameters, including tree depth, learning rate, regularization coefficient, number of leaf nodes, and number of iterations.

[0081] S42. Selection of initial samples: Randomly select a set of initial samples from the given hyperparameter space as the starting point for optimization;

[0082] S43. Construction of Gaussian process surrogate model: Based on the evaluation results of the initial sample set, a Gaussian process surrogate model is constructed to approximately simulate the distribution of the objective function in the hyperparameter space;

[0083] S44. Calculation of the acquisition function: Based on the Gaussian process surrogate model, the acquisition function is calculated using the expected improvement amount to balance exploration and utilization;

[0084] S45. Selection of the next evaluation point: Based on the evaluation results of the acquired function, identify the hyperparameter configuration that is expected to achieve the most significant improvement in the performance of the objective function;

[0085] S46. Evaluation of the objective function: The identified hyperparameter configurations are applied to actual model training and k-fold hierarchical cross-validation. The score is used as the objective function for evaluation; when the backflow sample is relatively scarce, k-fold stratified cross-validation is preferred to maintain the relative consistency of the ratio of backflow sample to non-backflow sample in each fold.

[0086] S47. Update of Gaussian process proxy model: Update the Gaussian process proxy model by taking the newly evaluated points into consideration.

[0087] S48. Evaluation of convergence criteria: Monitor the optimization process to evaluate whether it has converged. The convergence criterion is to reach a preset number of iterations. If the convergence criterion is not met, proceed to S44; if the convergence criterion is met, proceed to S49.

[0088] S49. Optimize hyperparameter output: Output the optimized hyperparameter configuration to complete the construction of the backflow identification model.

[0089] As an example, k is preferably 5.

[0090] As an example, the expected improvement amount and The fractions are calculated using the following formulas:

[0091]

[0092]

[0093]

[0094]

[0095] in, This indicates the hyperparameter configuration to be evaluated; This represents the objective function value corresponding to the hyperparameter configuration to be evaluated; This indicates that the optimal objective function value in the hyperparameter configuration has been evaluated; Indicates accuracy; Indicates recall rate; Indicates the number of days of backflow that were correctly identified; This indicates the number of days that were mistakenly identified as backflow but were not; This indicates the actual number of days of backflow that were missed in the initial assessment; This represents the expected improvement in the hyperparameter configuration h to be evaluated.

[0096] S5. Input the normalized test sample set from step S3 into the backflow identification model to obtain the backflow identification result.

[0097] Specifically, the standardized input feature vectors corresponding to a single test sample are first input into the trained backflow identification model in sequence; then, the backflow identification model classifies and predicts the sample based on the decision rules formed during the training phase to obtain the corresponding backflow probability value; the above process is repeated for all samples in the test sample set to obtain the backflow probability prediction result for the entire test sample set.

[0098] Furthermore, based on a preset discrimination threshold (e.g., 0.5), the probability value of backflow occurrence is converted into a classification result; dates with a probability greater than or equal to the threshold are determined as backflow occurrence days, and dates with a probability less than the threshold are determined as non-backflow days; the final backflow identification result includes a binary classification judgment result (occurred or not occurred) and a backflow event sequence composed of the results of each day sorted by time, which is used for subsequent SHAP interpretation analysis, key driving factor identification, and physical threshold determination.

[0099] In a preferred embodiment, the backflow identification model is a CatBoost model; for each test sample, the CatBoost model determines the output value of the leaf node in each tree according to the splitting path of each input feature in multiple decision trees, integrates the output results of each tree to obtain the comprehensive prediction score of the test sample, and then converts the comprehensive prediction score into a backflow probability value.

[0100] Furthermore, in this embodiment, the method for classifying and predicting the sample based on the decision rules formed during the training phase by the backflow identification model to obtain the corresponding backflow probability value specifically includes: inputting the standardized input feature vector corresponding to a single test sample into the backflow identification model in the same feature order as in the training phase, wherein the feature order is consistent with the feature arrangement order in the training sample set; wherein, the standardized input feature vector includes at least the external hydraulic gradient, the hydraulic gradient at the lake outlet, the hydraulic gradient inside the lake, the flow rate of the main river, the total flow rate of the tributaries flowing into the lake, the flow ratio feature, the trend feature, the lag feature, the time feature, and the topographic feature; the decision rule in this embodiment is that the backflow identification model, during the training phase, classifies and predicts the sample based on the training sample set. The tree splitting rule set is formed dynamically. The tree splitting rule set consists of node decision rules of multiple decision trees. Each node decision rule corresponds to an input feature, a splitting threshold, and a branch direction. When an input feature value satisfies the corresponding threshold comparison relationship, the sample is passed down along the first branch of that node. When an input feature value does not satisfy the corresponding threshold comparison relationship, the sample is passed down along the second branch of that node until it enters a leaf node. Each leaf node has a pre-recorded output value corresponding to the path. The splitting threshold and leaf node output value in the tree splitting rule set are derived from the iterative optimization results of the loss function during the training phase and are fixed under the target hyperparameter configuration determined by Bayesian optimization.

[0101] Specifically, when performing classification prediction on any test sample, the backflow identification model sequentially substitutes the standardized input feature vector into the node determination rules of each decision tree, determines the transmission path of the test sample in each decision tree according to the comparison results of each input feature and the corresponding split threshold, and obtains the leaf node output value of the test sample in each decision tree; then, the leaf node output values ​​corresponding to all decision trees are accumulated to obtain the comprehensive prediction score of the test sample; the comprehensive prediction score is the original backflow identification score corresponding to the test sample.

[0102] Furthermore, the comprehensive prediction score is input into the probability mapping function (Sigmoid function) to obtain the backflow probability value corresponding to the test sample; the backflow probability value ranges from 0 to 1; the closer the value is to 1, the higher the probability that the test sample belongs to the backflow sample; the closer the value is to 0, the higher the probability that the test sample belongs to the non-backflow sample.

[0103] The backflow probability value is compared with the discrimination threshold to obtain the classification result of the test sample; when the backflow probability value is greater than or equal to the discrimination threshold, the test sample is judged as a backflow sample; when the backflow probability value is less than the discrimination threshold, the test sample is judged as a non-backflow sample.

[0104] It should be noted that the classification prediction process also includes: using a symmetric tree structure in each decision tree to determine the path of the test samples, that is, using the same splitting features and splitting thresholds to branch the samples at the same level; thus, each test sample follows a consistent comparison rule at the same level, thereby obtaining stable leaf node output values; the comprehensive prediction score formed by accumulating the leaf node output values ​​of multiple decision trees is then converted into a backflow probability value through a probability mapping function, thereby completing the classification prediction process from input features to backflow probability values.

[0105] S6. Perform SHAP interpretation analysis on the backflow identification model. Determine key driving factors and physical thresholds based on the marginal contribution values ​​of each input feature, SHAP dependency graph, and binary box plot. Generate the evolution results of the driving mechanism based on the key driving factors and physical thresholds at different times. Note that the following should be noted in this step:

[0106] S61. Calculation of SHAP value: For each input feature of each sample, calculate its SHAP value, which represents the marginal contribution of that feature to the model's prediction result. The SHAP value is calculated according to the following formula:

[0107]

[0108] in, Indicates the first The SHAP value of each feature, Represents the total set of features. Indicates only the first A single-element set of features; Indicates from the total set of features Remove the first The set of features remaining after considering all features; express It is any feature subset of the remaining feature set, including the empty set, a subset consisting of a single feature, a subset consisting of multiple features, and a subset consisting of all remaining features. Indicates the number of features in the feature subset. Represents the total number of features. Indicates containing the first Model prediction output for each feature Indicates that it does not include the first The model prediction output for each feature; for the j-th feature of the i-th sample, its feature value is denoted as... The corresponding SHAP value is denoted as .

[0109] S62. Global Importance Analysis: By aggregating the mean absolute values ​​of all samples' SHAP (Shape Allocation Per Second) values, the global importance ranking of each feature is obtained, identifying the dominant driving factors at different times. The specific implementation method includes the following sub-steps:

[0110] S621. Calculate the global SHAP importance of each feature: For each input feature j, calculate the average of the absolute values ​​of its SHAP values ​​across all test samples, which is taken as the global importance score of that feature. The calculation formula is as follows:

[0111]

[0112] in, Let M represent the global importance score of the j-th feature, and M represent the total number of test samples. Represents the SHAP value of the j-th feature of the i-th sample;

[0113] S622. Rank the features by global importance: Rank all input features according to their global importance scores. Sort the features from highest to lowest importance to obtain a ranking list.

[0114] S623. Identify the main driving factors: Select the features with the highest importance scores in the ranking list (e.g., the top 5) as the main driving factors;

[0115] S624. Phased comparison of main driving factors: Repeat steps S621 to S623 for the test samples in the pre-operation, transition, and post-operation periods to obtain a ranking list of main driving factors for each period. By comparing the ranking changes of main driving factors in different periods, the temporal evolution characteristics of the driving mechanism can be identified.

[0116] S63. Local Dependency Analysis: Plot the SHAP dependency graph of key driving factors and analyze the nonlinear response relationship between eigenvalues ​​and SHAP values. Wherein:

[0117] S631. Select key driving factors: From the main driving factors identified in step S62, select one or more features with the highest global importance as key driving factors for local dependency analysis.

[0118] S632. Extracting Feature Values ​​and Corresponding SHAP Values: For the selected key driving factor j, extract the feature values ​​of that feature from all test samples. and its corresponding SHAP value , where i=1,2,...,M, and M is the total number of test samples;

[0119] S633. Draw the SHAP dependency graph: using eigenvalues The x-axis represents the SHAP value. Using the ordinate as the vertical axis, all sample points , Plotted on a two-dimensional scatter plot, forming a SHAP dependency graph;

[0120] S634. Analyze the nonlinear response relationship: Observe the distribution pattern of scatter points in the SHAP dependency graph and identify eigenvalues. With SHAP value Nonlinear response relationship between them;

[0121] Preferably, the nonlinear response relationship includes, but is not limited to:

[0122] Monotonically increasing relationship: eigenvalues When it increases, the SHAP value The continuous increase indicates that the larger the characteristic value, the stronger its positive contribution to the occurrence of backflow;

[0123] Monotonically decreasing relationship: eigenvalues When it increases, the SHAP value The continuous decrease indicates that the larger the characteristic value, the stronger its negative contribution to the occurrence of backflow;

[0124] Threshold effect: SHAP value At a certain characteristic value If the value remains stable or close to zero within the range, and increases or decreases sharply after exceeding a certain threshold, it indicates that there are critical conditions for the influence of this feature on backflow.

[0125] Non-monotonic relationship: SHAP value With eigenvalue The changes show a trend of first increasing and then decreasing or first decreasing and then increasing, indicating that the influence of this feature on backflow varies in direction or intensity within different value ranges.

[0126] S635. Extract candidate physical thresholds: Based on the SHAP values ​​in the SHAP dependency graph. Extract the corresponding feature values ​​at the inflection points where the value changes from negative to positive and from positive to negative. This serves as a candidate physical threshold for subsequent cross-validation.

[0127] S636. Draw binary box plots and compare distribution differences: Extract the feature value distributions of the backflow samples and non-backflow samples on the key driving factor, draw binary box plots, and compare the distribution differences between the two types of samples; the comparison includes the difference in median and the degree of overlap of the quartile intervals; if there is a significant difference in the median between the backflow samples and the non-backflow samples, it indicates that the key driving factor has a good ability to distinguish between the two types of samples; if the overlap of the quartile intervals of the two types of samples is small, it indicates that the key driving factor has a strong discriminative ability within this value range.

[0128] S637, Candidate Threshold Cross-Validation: The candidate threshold extracted in step S635 is cross-validated with the difference in the distribution of the two types of samples obtained in step S636. The specific method includes: determining whether the candidate threshold falls into or is close to the boundary between the backflow sample and the non-backflow sample; if the candidate threshold is between the medians of the two types of samples, it is considered that the candidate threshold is supported by both the model response features and the sample statistical distribution.

[0129] S638. Determine the physical threshold and output the results: Determine the candidate thresholds that have passed cross-validation as hydrological thresholds with physical significance, and output the physical threshold results of the key driving factor; the results shall include at least the name of the key driving factor, the location of the candidate threshold, the corresponding SHAP response change characteristics, the distribution difference between backflow samples and non-backflow samples, and the finally confirmed physical threshold.

[0130] S64. Generate the evolution results of the driving mechanism: Based on the main driving factors and their corresponding physical thresholds identified in the pre-operation, transition, and post-operation periods of the project, conduct a phased comparative analysis to generate the evolution results of the driving mechanism; the evolution results of the driving mechanism shall include at least the following: the ranking changes of the main driving factors in different periods, the changes of the physical thresholds of each key driving factor, and the changes in the nonlinear response relationship between the eigenvalues ​​and the SHAP values.

[0131] The following is combined with Figures 2-6 The schematic diagrams shown, along with some preferred or optional examples of the present invention, more specifically describe the implementation process and / or effects of certain embodiments of the present invention.

[0132] This embodiment takes the Yangtze-Poyang Lake system in China as the research object, and collects long-term hydrological data from 1970 to 2020. The data comes from the Jiangxi Provincial Hydrological Monitoring Center and includes: daily water level data of five water level stations in the Poyang Lake area (Hukou, Xingzi, Duchang, Tangyin, and Kangshan); daily flow data of the Hankou station on the main stream of the Yangtze River; daily inflow data of the five rivers (Ganjiang, Fuhe, Xinjiang, Raohe, and Xiushui) into the lake; and daily flow data of the Hukou station, which is used to define backflow events.

[0133] Based on the construction and operation timeline of the Three Gorges Dam (TGD), the research period is divided into three characteristic periods: Pre-TGD (1970-2002), a total of 33 years; Transition (2003-2008), the initial impoundment and operation phase of the Three Gorges Dam, a total of 6 years; and Post-TGD (2009-2020), the experimental and normal impoundment and operation phase at 175m of the Three Gorges Dam, a total of 12 years. Backflow events are defined as the daily flow at the Hukou station. This means that the water flows from the main stream of the Yangtze River to Poyang Lake.

[0134] Based on the physical mechanism of river-lake hydrological interaction, the following input feature set is constructed, as shown in Table 1:

[0135] Table 1. Input characteristics of the Yangtze River backflow into Poyang Lake

[0136]

[0137] The data was divided into training and test sets in an 8:2 ratio using stratified sampling to ensure that the ratio of backflow samples to non-backflow samples was consistent in both sets. Z-score normalization was used for normalization.

[0138] Furthermore, this embodiment selects eight mainstream machine learning models for comparative experiments, including: CatBoost, XGBoost, LightGBM, Random Forest (RF), Support Vector Machine (SVC), K Nearest Neighbors (KNN), Logistic Regression (LR), and Naive Bayes (NB).

[0139] To address the class imbalance problem where the number of backflow samples is far less than the number of non-backflow samples, a class weight balancing strategy is adopted, and the weight calculation formula is as follows:

[0140]

[0141] in, For the weights of the backflow samples, This represents the number of non-backflow samples. This represents the number of samples collected during backflow.

[0142] Reference Figure 2 This describes the process of Bayesian hyperparameter optimization, and the steps involved are shown in S41-S49 above. The optimized hyperparameter configuration is shown in Table 2.

[0143] Table 2. Hyperparameters and search space of the CatBoost model for Bayesian optimization.

[0144]

[0145] The performance comparison results of each model in the three periods are as follows: Figure 2 As shown, the results indicate that the CatBoost model performs best in all three periods, especially in the Post-TGD period when backflow samples are most scarce, where it still maintains a performance level of 0.76. The scores demonstrate extremely strong robustness.

[0146] Preferably, the SHAP method is used to interpret the prediction results of the CatBoost model. SHAP summary plots for three periods are generated. By aggregating the mean absolute SHAP values ​​of all samples, the global importance ranking of each feature is obtained. The top 10 driving factors for each period are as follows: Figure 3 As shown.

[0147] Reference Figure 4 , Figure 4 This invention presents the SHAP dependency plot and binary box plot of the top 5 key driving factors during the Pre-TGD period of the Three Gorges Dam. The SHAP dependency plot shows the nonlinear contribution and physical threshold of each driving factor to backflow occurrence, while the binary box plot verifies the threshold by comparing the distribution differences between backflow and non-backflow samples. Figure 4 As shown, during this period, backflow is mainly controlled by natural hydrological rhythms, with the difference in water levels between rivers and lakes and the flow rate of the Yangtze River main stream being the dominant factors.

[0148] Specifically, the water level difference between the Yangtze River and its tributaries exhibits a significant nonlinear abrupt change, with a physical threshold of approximately 0.71 m for promoting backflow; exceeding this value leads to a sharp increase in the probability of backflow. The flow rate of the Yangtze River main stream shows an exponential growth response, with a threshold of approximately 2.78 × 10⁻⁶ m. 4 m 3 / s, and the flow rate of the Yangtze River main stream the previous day also showed approximately 2.62 × 10⁻⁶. 4 m 3 The threshold effect of / s indicates that the inertial backwater effect of the Yangtze River flow is significant; the hydraulic gradient in the northern part of the lake is negatively correlated with the probability of backflow, and a smaller hydraulic gradient is more likely to induce backflow; the flow ratio (the ratio of Yangtze River flow to lake flow) shows a nearly linear response trend, with a threshold of about 11.68.

[0149] Reference Figure 5 , Figure 5 This invention presents the SHAP dependency plot and binary box plot of the top 5 key driving factors during the transition period; through comparison with... Figure 4 The comparison revealed that the backflow driving mechanism and hydrological thresholds have undergone significant changes; for example... Figure 5 As shown, the flow rate of the Yangtze River main stream and the previous day's flow rate remained the dominant factors, but their physical thresholds for inducing backflow significantly increased to 3.52 × 10⁻⁶. 4 m3 / s and 3.08×10 4 m 3 / s; Meanwhile, although the threshold for the difference between river and lake water levels is around 0.69m, the box plot shows that the distribution of backflow and non-backflow samples on this factor has a large overlap, reducing the sensitivity of differentiation.

[0150] It is noteworthy that the importance of seasonal factors (such as day count, DoY) increased significantly, and a threshold effect of increasing the probability of backflow was observed on day 198 (around late June). In addition, the influence of the hydraulic gradient in the northern part of the lake, which reflects the topography and drainage capacity, was also significantly enhanced, with a threshold of about 0.13m. These characteristics indicate that the interaction between the lake and the river during the transition period changed from being mainly controlled by natural hydrology to being jointly influenced by reservoir scheduling, seasonal patterns, and topographic evolution.

[0151] Reference Figure 6 , Figure 6 This invention presents the SHAP dependency plot and binary box plot of the top 5 key driving factors in the post-TGD period of the Three Gorges Dam; as shown in the present invention. Figure 6 As shown, during the stable operation period of the Three Gorges Dam, the river-lake system reached a new equilibrium state, with a significant decrease in backflow frequency and volume. The SHAP dependency plot indicates that the Yangtze River main stream flow remains a key driving factor, but under the long-term regulation of the reservoir, its physical threshold for inducing backflow has decreased to approximately 2.62 × 10⁻⁶. 4 m 3 / s, the previous day's Yangtze River main stream flow threshold also dropped to 2.53×10 4 m 3 / s.

[0152] Furthermore, a crucial shift has occurred in the driving mechanism: the persistent difference between river and lake water levels (such as the 7-day moving average difference) has replaced the instantaneous difference as the more critical factor, with a threshold of 0.11m. Below this threshold, the probability of backflow increases significantly. Meanwhile, the distinguishing ability of the instantaneous difference between river and lake water levels has further weakened. The box plot shows that the medians of the two types of samples (0.62m and 0.59m) are extremely close, lacking an effective dividing line. This reveals that the normalized operation of the Three Gorges Dam has weakened the impact of instantaneous flood peaks, making the persistent backwater effect of the river water replace the instantaneous difference as the key condition for inducing backflow.

[0153] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for identifying and analyzing the driving mechanism of indirect influence in online platforms based on interpretable machine learning, characterized in that, include: Long-sequence hydrological data of the river and lake system are collected, and the hydrological data are divided into characteristic periods according to the construction and operation time nodes of large-scale water conservancy projects in the upper reaches of the river to obtain a phased hydrological dataset. An input feature set is generated based on the phased hydrological dataset, and the input feature set includes hydraulic gradient features, flow characteristics, flow ratio features, trend features, lag features, time features, and topographic features. The input feature set is divided into a training sample set and a test sample set. The input features are normalized based on the training sample set, and the normalization parameters are used for feature transformation of the test sample set. The training sample set is input into the gradient boosting decision tree model, and the backflow samples and non-backflow samples are trained by combining class weights and balance constraints. The backflow identification model is obtained through Bayesian optimization. Input the hydrological data to be analyzed for the time period into the backflow identification model to obtain the backflow identification results; The backflow identification model is subjected to SHAP interpretation analysis. Key driving factors and physical thresholds are determined based on the marginal contribution values ​​of each input feature, the SHAP dependency graph, and the binary box plot. Furthermore, the evolution results of the driving mechanism are generated based on the key driving factors and physical thresholds at different time periods. Specifically, this includes: Global importance analysis: By aggregating the mean absolute values ​​of all samples, the global importance ranking of each feature is obtained, and the main driving factors at different times are identified. Local dependency analysis: Plot the SHAP dependency graph of key driving factors and analyze the nonlinear response relationship between eigenvalues ​​and SHAP values; Plot binary box plots and compare distribution differences: Extract the feature value distributions of backflow samples and non-backflow samples on key driving factors, plot binary box plots, and compare the distribution differences between the two types of samples; the comparison includes the difference in median and the degree of overlap of interquartile ranges; Candidate threshold cross-validation: The extracted candidate thresholds are cross-validated with the results of the differences in the distributions of the two classes of samples; Determine the physical threshold and output the results: The candidate thresholds that pass cross-validation are determined as hydrological thresholds with physical significance, and the physical threshold results of the key driving factor are output; the results should include at least the name of the key driving factor, the location of the candidate threshold, the corresponding SHAP response change characteristics, the distribution difference between backflow samples and non-backflow samples, and the finally confirmed physical threshold. Evolution results of the driving mechanism are generated: Based on the main driving factors and their corresponding physical thresholds identified in the pre-operation, transition, and post-operation periods, a phased comparative analysis is conducted to generate the evolution results of the driving mechanism. The evolution results of the driving mechanism include at least the changes in the ranking of the main driving factors in different periods, the changes in the physical thresholds of each key driving factor, and the changes in the nonlinear response relationship between the eigenvalues ​​and SHAP values.

2. The method for identifying and analyzing the driving mechanism of Jianghu reverse flow based on interpretable machine learning according to claim 1, characterized in that, The river-lake system includes rivers and lakes connected to the rivers. A backflow event is defined as a flow rate at the connection between a river and a lake that is less than zero. The characteristic periods include the pre-operation period, the transition period, and the post-operation period. The transition period is the initial water storage and operation stage of the large-scale water conservancy project, and the post-operation period is the normal water storage and operation stage of the large-scale water conservancy project.

3. The method for identifying and analyzing the driving mechanism of Jianghu reverse flow based on interpretable machine learning according to claim 1, characterized in that, The hydraulic gradient features include the external hydraulic gradient at the river-lake junction, the hydraulic gradient at the lake outlet, and the internal hydraulic gradient of the lake; the flow characteristics include the main river flow and the total flow of tributaries flowing into the lake; the flow ratio feature is the ratio of the main river flow to the total flow of tributaries flowing into the lake; the trend feature is the moving average of the hydraulic gradient; the lag feature is the lagged value of the key variable; the time feature is the daily count; and the topographic feature is the water level change caused by changes in riverbed elevation.

4. The method for identifying and analyzing the driving mechanism of Jianghu reverse flow based on interpretable machine learning according to claim 1, characterized in that, The moving average duration of the trend feature is 7 days, and the lag duration of the lag feature includes 1 day, 3 days, 7 days, and 10 days.

5. The method for identifying and analyzing the driving mechanism of river-lake backflow based on interpretable machine learning according to claim 1, wherein The input feature set is divided according to a preset ratio to obtain a training sample set and a test sample set, wherein the ratio of the training sample set to the test sample set is 8:

2. The normalization process uses Z-score standardization, and its calculation formula is as follows: in, Represents the original feature values. This represents the mean of the corresponding feature. This represents the standard deviation of the corresponding feature; This represents the standardized feature value.

6. The method for identifying and analyzing the driving mechanism of Jianghu reverse flow based on interpretable machine learning according to claim 1, characterized in that, The gradient boosting decision tree model is a CatBoost model, and the class weights are calculated based on the number of backflow samples and the number of non-backflow samples, using the following formula: in, Indicates the weight of the backflow sample. Indicates the number of non-backflow samples. This indicates the number of samples that were reversed.

7. The method for identifying and analyzing the driving mechanism of river-lake backflow based on interpretable machine learning according to claim 1, wherein The Bayesian optimization includes: setting the range of values ​​for tree depth, learning rate, regularization coefficient, number of leaf nodes, and number of iterations; randomly selecting initial hyperparameter samples; constructing a Gaussian process surrogate model; selecting the hyperparameter configuration to be evaluated based on the expected improvement amount; updating the Gaussian process surrogate model based on the F1 score of hierarchical cross-validation; and outputting the optimized hyperparameter configuration after the iteration termination condition is met.

8. The method for identifying and analyzing the driving mechanism of Jianghu reverse flow based on interpretable machine learning according to claim 7, characterized in that, The expected improvement and F1 score are calculated according to the following formulas: in, This indicates the hyperparameter configuration to be evaluated; This represents the objective function value corresponding to the hyperparameter configuration to be evaluated; This indicates that the optimal objective function value in the hyperparameter configuration has been evaluated; Indicates accuracy; Indicates recall rate; Indicates the number of days of backflow that were correctly identified; This indicates the number of days that were mistakenly identified as backflow but were not; This indicates the actual number of days of backflow that were missed in the initial assessment; This represents the expected improvement in the hyperparameter configuration h to be evaluated.

9. The method for identifying and analyzing the driving mechanism of Jianghu reverse flow based on interpretable machine learning according to claim 1, characterized in that, The SHAP interpretation analysis includes: calculating the SHAP value of each input feature for each sample, aggregating the mean absolute SHAP values ​​of all samples to obtain a global importance ranking, drawing a SHAP dependency graph of key driving factors, and cross-comparing the inflection points in the SHAP dependency graph with the sample boundary regions in the corresponding binary box plots to obtain the physical thresholds corresponding to the key driving factors.

10. The method for identifying and analyzing the driving mechanism of Jianghu reverse flow based on interpretable machine learning according to claim 9, characterized in that, The SHAP value is calculated according to the following formula: in, Indicates the first The SHAP value of each feature; Represents the total set of features; Indicates only the first A single-element set of features; Indicates from the total set of features Remove the first The set of features remaining after considering all features; express It is any feature subset of the remaining feature set, including the empty set, a subset consisting of a single feature, a subset consisting of multiple features, and a subset consisting of all remaining features; This represents the number of features in the feature subset; Indicates the total number of features; Indicates containing the first Model prediction output for each feature; Indicates that it does not include the first The model predicts output for each feature.