A method and device for calculating base flow in conductivity deficient areas
By constructing a feature contribution identification model and a conductivity regression model in areas with scarce conductivity, and using machine learning algorithms to fill the missing EC data, the problem of the CMB method being inapplicable was solved, high-precision base flow calculation was achieved, and the scope of application was broadened.
Patent Information
- Application Number
- CN202411428904.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-10-14
AI Technical Summary
In existing technologies, the conductivity mass balance method (CMB method) cannot be used to calculate baseflow in areas with poor conductivity, mainly because the regression relationship cannot be generalized due to the lack of EC data and regional specificity, which affects the accuracy and wide application of baseflow segmentation.
By constructing a characteristic contribution identification model and a conductivity regression model, environmental parameters are used to fill the gaps in EC time series data, and a machine learning algorithm is combined to calculate baseflow in conductivity-deficient areas. This involves obtaining hydrological characteristic data of the target site, performing contribution identification and conductivity regression, and finally using the conductivity mass balance method to segment baseflow.
It achieves high-precision baseflow calculation in areas with poor electrical conductivity, expands the application scope of the CMB method, and improves the accuracy and applicability of baseflow segmentation.
Smart Images

Figure CN119415818B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of hydrological research, and relates to a base flow calculation method and device for areas lacking electrical conductivity. Background Art
[0002] During a rainfall event in a watershed, the hydrograph formed at the outlet section is composed of different water sources. Distinguishing the different runoff components and their respective proportions requires segmenting the runoff into surface runoff and baseflow, a process known as baseflow segmentation. Baseflow segmentation is a fundamental issue in engineering and applied hydrology. Its results are not only crucial for runoff calculations and hydrological simulations, but also have significant implications for industrial and agricultural water supply, water security, water resource assessment, pollutant transport, and riverine ecological and environmental protection. Baseflow, composed of regional deep groundwater, shallow riverbank groundwater, intersoil flow, soil water, and riverbank return flow, represents the minimum volume of water a river can maintain during its dry season. As a crucial component of river flow, baseflow is crucial for our understanding of hydrological processes and for the protection and management of surface water resources. Accurately segmenting baseflow is crucial. However, baseflow cannot be measured directly and can only be estimated indirectly from river flow using various methods.
[0003] Currently, the main baseflow segmentation methods include non-tracer and tracer methods. Non-tracer methods, such as the ECK (Asseparation method developed by Eckhardt), separate high-frequency signals from low-frequency signals in the water flow, thereby dividing runoff into direct runoff and subsurface runoff. However, due to the subjective nature of parameter determination, accurate baseflow estimation is impossible. Tracer methods, on the other hand, assume that the water flow is composed of different components, each with characteristic concentrations of one or more conservative chemical components. Baseflow can be segmented based on the characteristic concentrations of these components. Compared to other chemical components, electric conductivity (EC) is relatively easy to measure and correlates with river flow. Therefore, the conductivity mass balance (CMB) method, which uses EC as a tracer, is currently the most accurate baseflow segmentation method and is widely used for high-precision baseflow calculations in areas with high conductivity. However, this method requires that stations have contemporaneous, daily-scale flow and EC data. However, EC data from stations in most parts of the world are measured only monthly or even quarterly, resulting in a high degree of intermittent and low resolution in EC monitoring. This makes it impossible to use the CMB method for baseflow calculations, limiting its widespread application.
[0004] To address the difficulty of using the CMB method in areas with scarce EC data, existing approaches typically construct a functional relationship between discharge and EC, regressing it to obtain time-series EC data. This regressed EC is then combined with the CMB formula to calculate baseflow. While this method is simple and easy to use, due to the region-specific nature of discharge variations, the constructed regression equations are difficult to generalize to other areas. Furthermore, these equations fail to simultaneously reflect the complex influence of multiple parameters, such as climate and watershed characteristics, on EC. Therefore, they are unsuitable for accurate EC regression of large numbers of stations with long sequences. Consequently, the CMB method cannot be applied to baseflow calculations in areas with insufficient conductivity.
[0005] Therefore, it is necessary to design a baseflow calculation scheme for areas with a lack of conductivity data, taking into account regional environmental parameters and filling the gaps in EC time series data measurements, so that the CMB method can be used to calculate baseflow; and solve the problem in the existing technology that the CMB method cannot be applied to areas with a lack of conductivity for baseflow calculation. Summary of the Invention
[0006] The present invention aims to provide a method and apparatus for calculating baseflow in areas with poor conductivity, which fills the gaps between EC time series data measurements in areas with poor conductivity based on environmental parameters, such as watershed properties and climatic conditions. This method is suitable for regressing EC at large quantities of long-sequence stations, fills the gaps in EC time series data measurements, and thus enables the use of the CMB method for baseflow calculation, solving the problem in the prior art that the CMB method cannot be applied to baseflow calculation in areas with poor conductivity.
[0007] In order to achieve the above object, the present invention provides the following technical solutions:
[0008] In a first aspect, the present invention provides a method for calculating baseflow in conductivity-deficient areas, which may include:
[0009] Obtaining target hydrological characteristic data of a target site, the target hydrological characteristic data including at least dynamic environmental parameters and static environmental parameters related to baseflow generation; the target site is an area with similar environmental characteristics to the preset site and poor conductivity;
[0010] Using a feature contribution recognition model to identify the contribution of the target hydrological feature data, and obtain target parameter combination data for conductivity regression of the target site; the target parameter combination data represents parameter combination data whose conductivity regression contribution is greater than a preset threshold;
[0011] Based on the target parameter combination data, according to the target time scale, a conductivity regression model is used to perform prediction to obtain target conductivity regression data for the target site;
[0012] Based on the target hydrological characteristic data and the target conductivity regression data, base flow of the target site is calculated by using the conductivity mass balance method, so as to obtain the base flow of the target site.
[0013] In a second aspect, the application provides a base flow calculation device for conductivity-deficient areas, which can include:
[0014] The acquisition module is configured to acquire target hydrological characteristic data of a target site, the target hydrological characteristic data including at least dynamic and static environmental parameters related to base flow generation; the target site is an area similar to a preset site in environmental characteristics and deficient in conductivity;
[0015] The contribution degree identification module is configured to identify the contribution degree of the target hydrological characteristic data by using a feature contribution degree identification model, so as to obtain target parameter combination data for conductivity regression of the target site; the target parameter combination data represents parameter combination data with a conductivity regression contribution degree greater than a preset threshold;
[0016] The conductivity regression prediction module is configured to predict by using a conductivity regression model according to a target time scale based on the target parameter combination data, so as to obtain target conductivity regression data for the target site;
[0017] The calculation module is configured to calculate the base flow of the target site by using the conductivity mass balance method based on the target hydrological characteristic data and the target conductivity regression data, so as to obtain the base flow of the target site.
[0018] Compared with the prior art, the present invention provides a method for calculating baseflow in areas with a lack of electrical conductivity. The method obtains target hydrological characteristic data of a target site, wherein the target hydrological characteristic data includes at least dynamic environmental parameters and static environmental parameters related to baseflow generation. The target site is an area with similar environmental characteristics to a preset site and lacks electrical conductivity. First, a characteristic contribution recognition model is used to identify the contribution of the target hydrological characteristic data to obtain target parameter combination data for electrical conductivity regression of the target site. The target parameter combination data represents parameter combination data whose conductivity regression contribution is greater than a preset threshold. Further, based on the target time scale, the conductivity regression model is used to identify the target parameter combination data. Conductivity regression prediction is performed to obtain target conductivity regression data for the target site; finally, based on the target hydrological characteristic data and the target conductivity regression data, the CMB method is used to perform baseflow segmentation calculation on the target site to obtain the baseflow of the target site; based on this, the pre-established feature contribution recognition model and the conductivity regression model can be used to obtain the target conductivity regression data for conductivity-deficient areas, fill the missing EC time series data measurement values, and obtain the regressed EC of a large number of long-sequence sites at the target time scale, so that the CMB method can be used to calculate the baseflow, realizing the use of the CMB method to calculate the baseflow in conductivity-deficient areas and expanding the application scope of the CMB method. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0020] Figure 1 The main flow chart of a method for calculating base flow in conductivity-deficient areas provided by the present invention;
[0021] Figure 2 A comparison chart of the model-regressed EC and the measured EC for a baseflow calculation method for conductivity-deficient areas provided by the present invention;
[0022] Figure 3 A comparison chart of the base flow calculation method for conductivity-deficient areas provided by the present invention and the base flow calculated by the ECK empirical value method;
[0023] Figure 4 This is a structural schematic diagram of a base flow calculation device for areas with low electrical conductivity provided by the present invention. DETAILED DESCRIPTION
[0024] To facilitate a clear description of the technical solutions of the embodiments of the present invention, the embodiments of the present invention use terms such as "first" and "second" to distinguish between identical or similar items with substantially the same functions and effects. For example, the first threshold and the second threshold are merely used to distinguish between different thresholds and do not limit their order of precedence. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that terms such as "first" and "second" do not necessarily define differences.
[0025] It should be noted that, in the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present invention should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0026] In the present invention, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects in the preceding time are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b, c can be single or multiple.
[0027] Existing high-precision CMB methods, when performing baseflow segmentation calculations, assume that (a) the characteristic concentrations of different stream components are distinct; (b) tracers do not react during transport through the medium of the catchment; and (c) the different components are always completely mixed in the river. This method requires that stations have contemporaneous, daily-scale flow and EC data. However, EC data for stations in most parts of the world are measured monthly or even quarterly, and EC monitoring is generally highly intermittent (monitoring is conducted only during flood or drought seasons or specific events) and has low resolution. This limits the widespread application of the CMB method and makes it difficult to apply the CMB method to baseflow segmentation calculations in areas lacking EC data. To address the difficulty of using the CMB method in areas lacking EC data, most existing solutions construct a functional relationship between flow and EC, regressing to obtain EC time series data, and then combining the regressed EC with the CMB method formula to calculate baseflow. While this method is simple and easy to use, the constructed regression equations are difficult to generalize to other regions due to the region-specific nature of flow variations. Furthermore, these equations cannot simultaneously reflect the complex influence of multiple parameters, such as climate and watershed characteristics, on EC. Therefore, they are not suitable for accurately regressing EC across large numbers of stations with long time series, limiting the accuracy of the CMB method. Machine learning, which can integrate and extract patterns from multi-metric data, has been widely applied in various hydrological fields. However, until now, no research has attempted to apply machine learning algorithms to fill in the gaps between EC time series measurements based on watershed attributes and climate conditions.
[0028] Based on this, the present invention proposes a method and device for calculating baseflow in areas with insufficient EC data, combined with a machine learning algorithm. By pre-constructing a feature contribution recognition model and an EC regression model based on areas with sufficient EC data and similar environments, EC regression data that meets the requirements of the CMB method is obtained. This allows the use of the high-precision CMB method to perform baseflow segmentation calculations in areas with insufficient EC data without affecting the calculation accuracy of the CMB method. This solves the problem in the prior art that the CMB method cannot be applied to baseflow calculations in areas with insufficient EC data.
[0029] Next, the technical solution of the present invention is described in detail with reference to the accompanying drawings:
[0030] See also Figure 1 , Figure 1 This is a main flow chart of a method for calculating baseflow in conductivity-deficient areas provided by the present invention. It should be noted that this method for calculating baseflow in conductivity-deficient areas is applicable to areas with environmental characteristics similar to those of a preset site and lacking conductivity. The method is executed by a server or terminal device equipped with the technical solution disclosed in the embodiments of the present invention, such as a computing service platform or handheld computing device.
[0031] exist Figure 1 In the method, the method may include:
[0032] Step 110: Obtain target hydrological characteristic data of a target site, wherein the target hydrological characteristic data includes at least dynamic environmental parameters and static environmental parameters related to baseflow generation; the target site is an area with similar environmental characteristics to the preset site and poor conductivity.
[0033] In step 110, the preset site is an area with sufficient conductivity and meets the requirements for baseflow segmentation calculation using the high-precision CMB method, such as the Missouri River in the United States; the target site is an area with environmental characteristics similar to those of the preset site and lacking conductivity; thus, the EC regression model constructed using the environmental characteristics of the preset site can be introduced into the environmental characteristic information of the target site to perform conductivity prediction, thereby compensating for the missing EC time series data measurements in the conductivity-deficient area and obtaining the EC time series data required for baseflow calculation using the CMB method.
[0034] Step 120: Using a feature contribution recognition model to identify the contribution of the target hydrological feature data, obtain target parameter combination data for conductivity regression of the target site; the target parameter combination data represents parameter combination data whose conductivity regression contribution is greater than a preset threshold.
[0035] Step 130: Based on the target parameter combination data and according to the target time scale, a conductivity regression model is used to perform prediction to obtain target conductivity regression data for the target site.
[0036] Step 140: Based on the target hydrological characteristic data and the target conductivity regression data, a CMB method is used to perform base flow segmentation calculation on the target site to obtain the base flow of the target site.
[0037] In steps 120 to 140, the characteristic contribution identification model and the conductivity regression model are network models constructed based on the hydrological characteristic data of the preset site and are pre-set in the server. When performing baseflow segmentation calculations in an area with environmental characteristics similar to those of the preset site and lacking conductivity, the above two models can be called to perform identification and prediction, thereby obtaining the conductivity regression data required for the CMB method to perform baseflow segmentation calculations. Therefore, based on the conductivity regression data and the target hydrological characteristic data, the CMB method can be used to perform baseflow segmentation calculations for the target site to obtain the baseflow of the target site. The lack of EC in the area will not prevent the high-precision CMB method from being used for baseflow segmentation calculations.
[0038] Based on this, the present invention provides a baseflow calculation method for conductivity-deficient areas. By utilizing a feature contribution recognition model and a conductivity regression model, target conductivity regression data for the conductivity-deficient areas is obtained, the missing EC time series data measurement values are filled, and the regression EC of a large number of long-sequence stations at the target time scale, such as the regression EC at the daily scale, is obtained, so that the CMB method can be used to calculate baseflow. This realizes the use of the CMB method to calculate baseflow in conductivity-deficient areas, broadening the application scope of the CMB method.
[0039] Preferably, in step 110, the dynamic environmental parameters may include at least daily precipitation time series data, temperature time series data, evapotranspiration time series data, flow time series data, and electrical conductivity. The static environmental parameters may include at least altitude, slope, sand content, clay content, annual average precipitation, annual average temperature, and annual average potential evaporation.
[0040] It should be noted that, during the specific implementation process, a variety of data related to these processes can be collected, not limited to the above data; because the electrical conductivity EC in the dynamic environmental parameters is not daily-scale data, it is impossible to directly use the CMB method to calculate the baseflow.
[0041] For example, dynamic environmental parameters, including daily precipitation and temperature time series, can be downloaded from websites such as ERA5-Land. Static environmental parameters, including elevation, slope, and sand grain ratio, can also be downloaded. For example, if a study of a specific site in the Missouri River Basin determines that the minimum training set data set is 156, then at least 156 weekly data points related to EC generation can be collected for this target site to facilitate subsequent research.
[0042] Preferably, the step of using a feature contribution recognition model to identify the contribution of the target hydrological feature data to obtain target parameter combination data for conductivity regression of the target site may include:
[0043] Obtaining hydrological characteristic data for a preset site, wherein the preset site is a site whose conductivity satisfies base flow segmentation calculation using the CMB method; dividing the hydrological characteristic data into ecological zones to obtain a plurality of ecological zones; the ecological zones are regions with similar environmental characteristics, including climate, vegetation, soil type, and geological conditions; and establishing the characteristic contribution identification model and the conductivity regression model based on the hydrological characteristic data for the plurality of ecological zones.
[0044] As an example, the EC data for the Missouri River basin in the United States meets the requirements for baseflow segmentation calculation using the CMB method, as it has daily-scale EC data. Therefore, this paper uses the Missouri River basin as a study area, constructs a feature contribution identification model and a conductivity regression model, and stores these models on a server. When baseflow segmentation calculation is required for areas with similar environmental characteristics to the pre-set site and low conductivity, the contribution identification model and conductivity regression model provided by this invention can be used to obtain a conductivity data sequence that meets the requirements of the CMB method, thereby enabling high-precision baseflow calculations in conductivity-deficient areas using the CMB method.
[0045] First, the dynamic and static environmental parameters of the stations in the study area are collected; the dynamic environmental parameters can include daily precipitation time series, temperature time series, evapotranspiration time series, flow time series and EC, etc.; the static environmental parameters can include altitude, slope, sand particle ratio, clay particle ratio, annual average precipitation, annual average temperature, annual average potential evaporation, etc.
[0046] Furthermore, areas within the study area with similar environmental characteristics, such as climate, vegetation, soil type, and geological conditions, are divided into ecological zones. The preset sites of the present invention can be divided into the Western Mountains (WM) ecological zone, the Northern Plains (NP) ecological zone, the Southern Plains (SP) ecological zone, and the Temperate Plains (TP) ecological zone. The following uses the Southern Plains (SP) ecological zone as an example to construct a feature contribution recognition model and a conductivity regression model. The construction methods for the other three ecological zones are the same as those for the Southern Plains (SP) ecological zone.
[0047] The main dynamic parameters of the SP ecoregion in the Missouri River Basin can be used. These parameters include the vegetation foliage index and flow, and the main dynamic parameters include elevation, drought index, annual average precipitation, annual average potential evaporation, snowfall percentage, and annual average temperature.
[0048] Preferably, establishing a feature contribution recognition model may include: performing data preprocessing on the hydrological feature data in the plurality of said ecological zones to obtain a sample data set including a first training set and a first test set; performing model training on a network model using the first training set to obtain a first feature contribution recognition model; and performing recognition accuracy testing on the first feature contribution recognition model using the first test set.
[0049] If the recognition accuracy of the first feature contribution recognition model meets the preset recognition accuracy requirements, the first feature contribution recognition model will be used as the feature contribution recognition model; if the recognition accuracy of the first feature contribution recognition model does not meet the recognition accuracy requirements, the network model will continue to be trained until a feature contribution recognition model that meets the recognition accuracy requirements is obtained.
[0050] Preferably, the first feature contribution recognition model includes determining multiple algorithm models; the multiple algorithm models include the formula:
[0051] FI(X j )=∑ 树 ∑ 节点t ΔMSE j (t) (1)
[0052]
[0053] In formulas (1) to (3), FI is the feature contribution, N is the number of samples, and y i is the true value of the i-th sample, is the model calculation value of the i-th sample, t is the t-th node, N L is the number of samples of the adjacent nodes on the left side of the t-th node, N T is the number of samples of the adjacent nodes on the right side of the t-th node, Nt is the number of samples of the t-th node, t R is the right adjacent node of the t-th node, t L is the left adjacent node of the t-th node, and MSE is the mean square error reduced by the feature when splitting the node.
[0054] Specifically, various data from the Southern Plains Ecological Region (SP) are first integrated. 80% of the integrated data constitutes the first training set for the feature recognition model, and the remaining 20% serves as the first test set. The network model is trained using the first training set, and the recognition accuracy of the feature recognition model is evaluated using the first test set. The network model preferably uses a random forest regression model (RF-Regressor) as the feature recognition model in the present invention. In the algorithm, n_estimators can be set to 100, and random_state can be set to 42.
[0055] The importance of the feature in the EC regression is measured by the accumulated value of the mean squared error (ΔMSE) reduced by the feature at the split node, the more the split times, the greater the error reduction, the higher the importance of the feature. Considering the influence of dynamic and static parameters on the regression results of the model, the importance of each feature FI is determined by formulas (1) to (3); that is, the contribution degree FI of each feature is obtained. Specifically, the importance of each feature can be output by the RF-Regressor model attribute feature_importances, and the features with a total mean squared error greater than or equal to 80% are selected as the main control parameters of the EC regression of the ecological region.
[0056] Based on this, the RF-Regressor model can be trained using the training set, and a first feature contribution degree identification model can be obtained by self-learning or semi-supervised learning. Then, the first feature contribution degree identification model is verified using the test set to determine whether the difference between the main control parameters output by the first feature contribution degree identification model and the actual main control parameters is large. If the main control parameters output by the first feature contribution degree identification model are the same as or not much different from the actual main control parameters, it means that the first feature contribution degree identification model meets the requirements, and the first feature contribution degree identification model can be applied to other basins with similar environmental characteristics for feature identification. If the difference between the main control parameters output by the first feature contribution degree identification model and the actual main control parameters is large, the RF-Regressor model is further trained using the first training set or newly collected data until the main control parameters output by the trained model are the same as or not much different from the actual main control parameters of the Missouri River basin.
[0057] Further, the establishment of the conductivity regression model can include: using the feature contribution degree identification model to identify the hydrological feature data of multiple ecological regions to obtain multiple main control parameters for the conductivity regression of the multiple ecological regions; the multiple main control parameters are parameters with a total mean squared error greater than or equal to 80%; and the multiple main control parameters are discretely sampled at a target number of intervals of weeks to obtain a second training set of a target time length.
[0058] The network model is trained based on the second training set to obtain a first conductivity regression model. Specifically, during the model training process, the daily scale dynamic and static environmental parameters in the second training set are used as inputs, and the daily scale EC is used as an output to perform machine learning or semi-supervised learning on the network model, thereby obtaining the first conductivity regression model.
[0059] The prediction accuracy of the first conductivity regression model is verified using a second test set, resulting in a model fit verification result for the first conductivity regression model. The second test set comprises hydrological characteristic data from the predetermined site, excluding conductivity. Specifically, during the model testing process, dynamic and static environmental parameters at the daily scale are used as input to determine whether the first conductivity regression model can output a daily-scale regressed EC, thereby determining whether the daily-scale regressed EC deviates from the measured daily-scale EC at the predetermined site.
[0060] If the verification result indicates that the prediction accuracy of the first conductivity regression model meets the prediction accuracy requirement, the first conductivity regression model is used as the conductivity regression model.
[0061] If the verification result indicates that the prediction accuracy of the first conductivity regression model does not meet the prediction accuracy requirement, the network model is continuously trained until a conductivity regression model that meets the prediction accuracy requirement is obtained.
[0062] It should be noted that the prediction accuracy represents the deviation between the obtained regression EC and the measured EC, and this deviation can be determined according to the specific application scenario.
[0063] The prediction accuracy verification of the first conductivity regression model using the second test set may include using the formula:
[0064]
[0065] Calculate the root mean square error, which is used to evaluate the prediction accuracy of the conductivity regression model; where RMSE is the root mean square error, EC mod,i is the regression EC obtained by the model at time i, is the average value obtained from the measured EC, and n is the number of data.
[0066] Specifically, the master control parameters determined using the contribution identification model can be discretely sampled at intervals of T = 7 days to simulate a hydrological station monitoring EC on a weekly basis. This time interval can be modified based on the specific conditions of the study site and monitoring conditions; for example, a hydrological station monitoring EC on a daily basis can be simulated. This description uses T = 7 days as an example.
[0067] It should be noted that in areas where EC data are scarce, the amount of data required to segment the base flow directly determines the economic cost and generalizability of the model. Therefore, in order to improve the economy and generalizability of the model, it is necessary to determine the optimal data length of the second training set.
[0068] Therefore, discrete sampling of the plurality of master control parameters for a target number of times at weekly intervals to obtain a second training set of a target time length may specifically include:
[0069] First, the weekly training data size was subjectively set to 52, 78, 104, 130, 156, 182, 208, and 260, respectively, to represent the data of sites with a time length of 12 months, 18 months, and 24 months. In other words, the target time length can include 12 months, 18 months, or 24 months.
[0070] Next, discrete sampling of the master control parameters is performed at weekly intervals for a target number of times, such as 52, 78, 104, 130, 156, 182, 208, and 260 times, to form a training set of corresponding data volume. That is, the target number of times may include 52, 78, 104, 130, 156, 182, 208, or 260.
[0071] It should be noted that, assuming that the sampling interval of the training set data is weekly, that is, one number is taken per week, when the amount of training set data is 52, 52 * 7 days / 30 days = 12 months, and so on.
[0072] The present invention prefers the RF-Regressor model as the regression model for continuous EC and uses hyperparameter tuning to determine the parameters of the RF model. GridSearchCV is used to optimize four parameters of the RF model: max_depth (the maximum depth of the decision tree); n_estimators (the number of decision trees in the random forest); min_samples_leaf (the minimum number of samples per leaf node); and min_samples_split (the minimum number of samples required to split a node). The optimal parameter combination for the RF model is output. For example, the remaining control parameters are processed to the same frequency as the EC data, and then the optimal parameter combination for the RF model for EC regression is identified, determining max_depth = 8, min_samples_leaf = 2, min_samples_split = 2, and n_estimators = 250.
[0073] The conductivity regression model can also be cross-validated and evaluated. For example, after determining the best parameter combination, the performance of the model can be evaluated using 5-fold cross validation (cv=5). 2 To reflect the model fitting situation.
[0074] Furthermore, the input data can be preprocessed. The continuous data of the main control parameters, excluding EC, is used as the test set, with EC as the target variable and the remaining parameters as feature variables. The algorithm first processes missing values and outliers in the imported data, and then uses StandardScaler to standardize the feature data.
[0075] Finally, the root mean square error (RMSE) can be calculated using formula (4).
[0076] Based on this, the above method can be used to perform five sampling and calculations to obtain the corresponding root mean square error (RMSE). Specifically, the training set consists of 52, 78, 104, 130, 156, 182, 208, and 260 samples from the discrete data, which are then used to train the model and obtain the RMSE of the model regression results. The obtained RMSE is subjected to a normal distribution and T-test. Among the target times with no significant difference (p>0.05), the shortest one is selected as the optimal acceptable target time. Experiments have shown that 156 samplings is the minimum acceptable monitoring period and the optimal data length for the training set. Therefore, in conductivity-deficient areas with similar environmental characteristics, the method of sampling hydrological data 156 times can be used for EC regression prediction, ensuring the accuracy of the prediction results while minimizing the cost.
[0077] In summary, the model can be trained using a training set with minimal data to improve its training efficiency. Furthermore, the test set can be used to input environmental characteristic variables other than EC to verify whether the target variable EC can be obtained. This will confirm the prediction accuracy of the conductivity regression model.
[0078] Furthermore, you can also combine EC regression results and visualization output. Output the time series table of regressed EC and visualize it with the measured EC for better data comparison.
[0079] Furthermore, based on the regressed EC and hydrological characteristic data, the CMB method can be used to perform base flow segmentation calculation on the preset site to obtain the base flow of the preset site; and the accuracy of the regressed EC can be further evaluated.
[0080] For details, please refer to Figures 2 to 3 , Figure 2 A comparison chart of the model-regressed EC and the measured EC for a baseflow calculation method for conductivity-deficient areas provided by the present invention; Figure 3 A comparison chart of the base flow calculation method for conductivity-deficient areas provided by the present invention and the base flow calculated by the ECK empirical value method.
[0081] exist Figure 2In the figure, the EC curve is constructed using the data from the 07096000 station as an example. The horizontal axis represents time in days; the vertical axis represents conductivity EC in μs / cm; the curve composed of black dots is the measured EC curve, and the curve composed of blue dots is the regression conductivity EC curve predicted by the model constructed using the method of the present invention. Figure 1 It can be obtained without a doubt: Choose R 2 The RMSE evaluation model regression results were used to evaluate the conductivity regression model constructed by the present invention, and the model regression fitting degree R 2 The EC obtained by the conductivity regression model is well matched with the measured EC data in terms of time variation, which shows that the feature contribution recognition model and the conductivity regression model constructed by the method of the present invention have good fit with the measured EC data and meet the design requirements.
[0082] exist Figure 3 In the figure, the horizontal axis represents time in days; the vertical axis represents base flow in m 3 / s; The curve composed of black dots is the measured data curve of base flow, the curve composed of blue diamonds is the model regression curve, and the curve composed of pink dots is the curve obtained by using the ECK method to perform base flow segmentation calculation. In order to highlight the superiority of the method provided by the present invention, the segmentation results of the commonly used ECK empirical value method are introduced for comparison, and the RMSE model is set to 4.70, RMSE ECK Set to 7.93. Figure 3 As can be seen, compared to the baseflow segmentation results of the ECK method, the regression results of the method provided in this specification are more consistent with the measured results and have a lower calculated RMSE. This also shows that the regressed EC obtained using the method provided in this specification meets the EC scale requirements for high-precision baseflow segmentation calculations using the CMB method. This high-precision baseflow segmentation method can be applied to data-scarce areas, and the resulting baseflow segmentation accuracy is higher than that of the commonly used ECK method.
[0083] Preferably, in step 120, the use of the feature contribution recognition model to identify the contribution of the target hydrological characteristic data to obtain target parameter combination data for conductivity regression of the target site may include: dividing the target hydrological characteristic data into ecological zones to obtain multiple target ecological zones; inputting the hydrological characteristic data of the multiple target ecological zones into the feature contribution recognition model to obtain the contribution of each feature; using feature parameters with a contribution greater than or equal to 80% as target parameters for the multiple target ecological zones; and performing data combination processing on the target parameters of the multiple target ecological zones to obtain target parameter combination data for conductivity regression of the target site.
[0084] It should be noted that the method used to divide the target site into ecological zones is the same as the method used to divide the preset site into ecological zones, and will not be repeated here. Therefore, the target site and the preset site have the same environmental characteristics, and the ecological zones are divided in the same way. Therefore, the feature contribution model established based on the hydrological characteristic data of the preset site can be used to identify the contribution of features in multiple ecological zones, obtain the contribution of each feature, further select features with a contribution greater than 80% as target parameter features, and finally combine all target parameter features of the multiple target ecological zones to obtain the target parameter combination data for the conductivity regression of the target site. Similarly, it can be understood that the conductivity regression model constructed for the preset site can also be used to predict the target conductivity regression data for the target site; thus, based on the target hydrological characteristic data and the target conductivity regression data, the CMB method can be used to perform baseflow segmentation calculations on the target site to obtain the baseflow of the target site.
[0085] Preferably, in step 140, performing base flow segmentation calculation on the target site using the CMB method based on the target hydrological characteristic data and the target conductivity regression data to obtain the base flow of the target site may include:
[0086] Using the formula:
[0087]
[0088] Calculate the base flow of the target site; where q BF is the base traffic of the target site, EC RO The minimum conductivity and EC are calculated year by year. BF is the maximum annual conductivity, Q is the target hydrological characteristic data, and EC is the target conductivity regression data.
[0089] Specifically, the parameter EC can be determined by the daily EC data obtained by model regression. RO 128.69μs / cm, EC BF The daily flow rate is 338.97 μs / cm. Substituting the determined parameters, electrical conductivity and collected Q data into formula (5), the daily flow rate is obtained.
[0090] In a second aspect, the present invention provides a base flow calculation device for areas with low conductivity. Figure 4 , Figure 4 This is a structural schematic diagram of a base flow calculation device for areas with low electrical conductivity provided by the present invention.
[0091] exist Figure 4 , the computing device may include:
[0092] The acquisition module 410 is used to obtain target hydrological characteristic data of a target site, wherein the target hydrological characteristic data includes at least dynamic environmental parameters and static environmental parameters related to baseflow generation; the target site is an area with similar environmental characteristics to the preset site and poor conductivity.
[0093] The contribution identification module 420 is used to identify the contribution of the target hydrological characteristic data using a feature contribution identification model to obtain target parameter combination data for conductivity regression of the target site; the target parameter combination data represents parameter combination data whose conductivity regression contribution is greater than a preset threshold.
[0094] The conductivity regression prediction module 430 is configured to perform prediction based on the target parameter combination data and according to a target time scale using a conductivity regression model to obtain target conductivity regression data for the target site.
[0095] The calculation module 440 is configured to perform baseflow segmentation calculation on the target site using a conductivity mass balance method based on the target hydrological characteristic data and the target conductivity regression data to obtain the baseflow of the target site.
[0096] Based on this, the present invention provides a baseflow calculation device for conductivity-deficient areas, which uses an acquisition module 410 to obtain target hydrological characteristic data of a target site, wherein the target hydrological characteristic data includes at least dynamic environmental parameters and static environmental parameters related to baseflow generation; the target site is an area with similar environmental characteristics to a preset site and lacking conductivity; further, a contribution identification module 420 is used to identify the contribution of the target hydrological characteristic data to obtain target parameter combination data for conductivity regression of the target site; the target parameter combination data represents parameter combination data whose conductivity regression contribution is greater than a preset threshold; and a conductivity regression prediction module 430 is used to perform prediction based on the target parameter combination data and a target time scale using a conductivity regression model to obtain target conductivity regression data for the target site; finally, a calculation module 440 is used to perform baseflow segmentation calculation on the target site using a conductivity mass balance method based on the target hydrological characteristic data and the target conductivity regression data to obtain the baseflow of the target site; thereby, the conductivity mass balance method can be used to perform baseflow segmentation calculation for conductivity-deficient areas.
[0097] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0098] Although the present invention has been described with reference to specific features and embodiments thereof, it will be apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the invention. Accordingly, this specification and drawings are merely illustrative of the invention as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the invention. It will be apparent that various modifications and variations may be made to the present invention by those skilled in the art without departing from the spirit and scope of the invention. Thus, the present invention is intended to include such modifications and variations as fall within the scope of the claims of the present invention and their equivalents.
Claims
1. A method for calculating baseflow in areas with low electrical conductivity, characterized in that: include: Obtaining target hydrological characteristic data of a target site, the target hydrological characteristic data including at least dynamic environmental parameters and static environmental parameters related to baseflow generation; the target site is an area with similar environmental characteristics to the preset site and poor conductivity; Using a feature contribution recognition model to identify the contribution of the target hydrological feature data, and obtain target parameter combination data for conductivity regression of the target site; the target parameter combination data represents parameter combination data whose conductivity regression contribution is greater than a preset threshold; Based on the target parameter combination data, according to the target time scale, a conductivity regression model is used to perform prediction to obtain target conductivity regression data for the target site; Based on the target hydrological characteristic data and the target conductivity regression data, a conductivity mass balance method is used to perform base flow segmentation calculation on the target site to obtain the base flow of the target site; The method of using a feature contribution recognition model to identify the contribution of the target hydrological feature data to obtain target parameter combination data for conductivity regression of the target site includes: Obtaining hydrological characteristic data of a preset site; the preset site is a site whose conductivity satisfies the base flow split calculation using the conductivity mass balance method; Dividing the hydrological characteristic data into ecological zones to obtain a plurality of ecological zones; the ecological zones are regions with similar environmental characteristics; the environmental characteristics include climate, vegetation, soil type, and geological conditions; Based on the hydrological characteristic data of the plurality of ecological zones, the characteristic contribution recognition model and the conductivity regression model are established.
2. The method according to claim 1, wherein The method of performing base flow segmentation calculation on the target site based on the target hydrological characteristic data and the target conductivity regression data using the conductivity mass balance method to obtain the base flow of the target site includes: Using the formula: Get the base flow of the target site; where q BF is the base traffic of the target site, EC RO The minimum conductivity and EC are calculated year by year. BF is the maximum annual conductivity, Q is the target hydrological characteristic data, and EC is the target conductivity regression data.
3. The method according to claim 1, wherein The step of establishing a feature contribution recognition model includes: Performing data preprocessing on the hydrological characteristic data in the plurality of ecological zones to obtain a sample data set including a first training set and a first test set; Using the first training set to train the network model, a first feature contribution recognition model is obtained; Using the first test set to perform a recognition accuracy test on the first feature contribution recognition model; If the recognition accuracy of the first feature contribution recognition model meets the preset recognition accuracy requirement, the first feature contribution recognition model is used as the feature contribution recognition model; If the recognition accuracy of the first feature contribution recognition model does not meet the recognition accuracy requirement, the network model continues to be trained until a feature contribution recognition model that meets the recognition accuracy requirement is obtained.
4. The method according to claim 3, wherein The first feature contribution recognition model includes determining a plurality of algorithm models; The algorithm models include the formula: Among them, FI is the feature contribution, N is the number of samples, y i is the true value of the i-th sample, is the model calculation value of the i-th sample, t is the t-th node, N L is the number of samples of the adjacent nodes on the left side of the t-th node, N R is the number of samples of the adjacent nodes on the right side of the t-th node, Nt is the number of samples of the t-th node, t R is the right adjacent node of the t-th node, t L is the left adjacent node of the t-th node, ΔMSE j (t) is the use of feature X at node t j The reduction in mean square error caused by splitting.
5. The method according to claim 1, wherein The establishment of the conductivity regression model comprises: Using the feature contribution recognition model to perform feature recognition on the hydrological feature data of the plurality of ecological zones, a plurality of main control parameters for conductivity regression of the plurality of ecological zones are obtained; the plurality of main control parameters are parameters whose sum of mean square errors is greater than or equal to 80%; discretely sampling the plurality of master control parameters for a target number of times at weekly intervals to obtain a second training set of a target time length; Performing model training on the network model based on the second training set to obtain a first conductivity regression model; Verifying the prediction accuracy of the first conductivity regression model using a second test set to obtain a verification result of the model fitting of the first conductivity regression model; the second test set is data other than conductivity in the hydrological characteristic data of the preset site; If the verification result indicates that the prediction accuracy of the first conductivity regression model meets the prediction accuracy requirement, the first conductivity regression model is used as the conductivity regression model; If the verification result indicates that the prediction accuracy of the first conductivity regression model does not meet the prediction accuracy requirement, the network model is continuously trained until a conductivity regression model that meets the prediction accuracy requirement is obtained.
6. The method according to claim 5, wherein The use of the second test set to verify the prediction accuracy of the first conductivity regression model includes using the formula: Calculate the root mean square error, which is used to evaluate the prediction accuracy of the conductivity regression model; where RMSE is the root mean square error, EC mod,i is the regression EC obtained by the model at time i, is the average value obtained from the measured EC, and n is the number of data.
7. The method according to claim 1, wherein The method of using a feature contribution recognition model to identify the contribution of the target hydrological feature data to obtain target parameter combination data for conductivity regression of the target site includes: Dividing the target hydrological characteristic data into ecological zones to obtain a plurality of target ecological zones; Inputting the hydrological characteristic data of the plurality of target ecological zones into the characteristic contribution recognition model to obtain the contribution of each characteristic; Taking characteristic parameters with a contribution greater than or equal to 80% as target parameters for the plurality of target ecological zones; The target parameters of the plurality of target ecological zones are subjected to data combination processing to obtain target parameter combination data for conductivity regression of the target site.
8. The method according to claim 1, wherein The dynamic environmental parameters include at least daily precipitation time series data, temperature time series data, evapotranspiration time series data and flow time series data, as well as electrical conductivity; the static environmental parameters include at least altitude, slope, sand particle ratio, clay particle ratio, annual average precipitation, annual average temperature and annual average potential evaporation.
9. A base flow calculation device for areas with low conductivity, characterized in that: include: An acquisition module, the acquisition module being configured to acquire target hydrological characteristic data of a target site, the target hydrological characteristic data comprising at least dynamic environmental parameters and static environmental parameters related to baseflow generation; the target site being an area having similar environmental characteristics to a preset site and lacking in electrical conductivity; a contribution recognition module, the contribution recognition module being configured to perform contribution recognition on the target hydrological characteristic data using a characteristic contribution recognition model to obtain target parameter combination data for conductivity regression of the target site; the target parameter combination data representing parameter combination data whose conductivity regression contribution is greater than a preset threshold; A conductivity regression prediction module is used to predict target conductivity regression data for the target site using a conductivity regression model based on the target parameter combination data and according to a target time scale; a calculation module, configured to perform baseflow segmentation calculation on the target site using a conductivity mass balance method based on the target hydrological characteristic data and target conductivity regression data, to obtain the baseflow of the target site; The method of using a feature contribution recognition model to identify the contribution of the target hydrological feature data to obtain target parameter combination data for conductivity regression of the target site includes: Obtaining hydrological characteristic data of a preset site; the preset site is a site whose conductivity satisfies the base flow split calculation using the conductivity mass balance method; Dividing the hydrological characteristic data into ecological zones to obtain a plurality of ecological zones; the ecological zones are regions with similar environmental characteristics; the environmental characteristics include climate, vegetation, soil type, and geological conditions; Based on the hydrological characteristic data of the plurality of ecological zones, the characteristic contribution recognition model and the conductivity regression model are established.
Citation Information
Patent Citations
Intelligent flow prediction method and system for information-deficient drainage basin
CN118261285A
KR1018726460000B1