A hydrological process model establishing method based on checking and verifying data distribution characteristics consistency
By using discretization sampling methods and MDUPLEX data allocation technology, the problem of inconsistent data distribution between hydrological process model verification and validation was solved, thereby improving the robustness and prediction accuracy of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2022-05-05
- Publication Date
- 2026-04-28
AI Technical Summary
The inconsistency in the distribution of verification and validation data in existing hydrological process models leads to large differences in model performance, making it difficult to maintain stability and effectiveness in practical applications.
A discretized sampling method is adopted, and the dataset is allocated to the verification and validation sets through the MDUPLEX method to ensure the consistency of data distribution characteristics. A hydrological process model based on the consistency of verification and validation data distribution characteristics is established.
It improves the performance consistency of the model during the verification and validation periods, enhances the transferability and prediction accuracy of the model, and reduces the performance differences of the model under different conditions.
Smart Images

Figure CN115034150B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of water conservancy, specifically involving the establishment and application technology of hydrological process models. Background Technology
[0002] Physical process-based hydrological models are widely used in watershed rainfall-runoff prediction, drought and flood forecasting, and other applications. These models are fundamentally "process-driven," interpreting runoff formation based on quantitative studies of the water cycle. Typically, before establishing such models, the model structure and physical boundary conditions need to be determined based on theoretical foundations. Furthermore, the control volume must obey basic laws of conservation of matter and energy, thus providing a relatively good explanation of the overall runoff formation process.
[0003] However, since process models typically require the introduction of various simplifying assumptions and parameter adjustments to better simulate the actual rainfall-runoff processes in the current watershed, model verification and validation are necessary before practical application. Verification involves adjusting model parameters using historical data from specific locations, while validation involves using the verified model to predict independent data to verify whether the model can provide good predictive performance under new, unseen conditions.
[0004] Because the simulation results of process-driven models are often influenced by their initial state, hydrological process models typically require continuous time-series observation data for model validation. This prevents these models from using short, random, or small-batch data (data types commonly used in traditional machine learning) for validation training. This is a key difference between hydrological process-driven models and data-driven models.
[0005] To address the requirement for continuous time data to validate hydrological process models, the traditional approach is to divide the available continuous observation data into two proportionally continuous time series: one for model validation and the other for performance verification. However, this approach often results in significantly worse model performance on the validation set compared to the validation set. This problem was identified as early as the 1980s, the root cause being the potential significant differences in hydrological conditions between the validation and verification datasets. For example, if the data in the validation dataset primarily represents relatively humid hydroclimatic conditions, while the data in the validation dataset represents relatively arid conditions, then the model validated under humid conditions will inevitably perform significantly worse when validating arid hydrological events, and vice versa. To address this issue, some researchers have suggested extending the validation dataset as much as possible, using a sufficiently large time span to ensure the model can utilize data features from the entire watershed during the validation period. However, this approach is subjective, and previous research has shown that its effectiveness is often unsatisfactory. Therefore, ensuring the consistency of the distribution of verification and validation data to improve the performance of hydrological process models is a difficult problem in the field of water conservancy, and also an important reason that hinders the widespread practical application of models.
[0006] To address this issue, this invention proposes a novel method for establishing hydrological process models based on the consistency of distribution characteristics between verification and validation data. This method completely abandons the traditional approach of selecting continuous time data for model verification, instead employing a discretization sampling method to ensure consistency in data distribution characteristics between the verification and validation datasets. In this method, the model runs continuously across the entire dataset from start to finish, and then selects discrete time data for model verification. To ensure consistency in distribution characteristics between the selected verification and validation data, this invention also proposes a new sampling method, MDUPLEX, to complete the dataset allocation. This invention is original, completely changing the traditional method for establishing hydrological dynamic models, and achieving robust model performance and good transferability, which is of great significance to model work in the field of water conservancy. Summary of the Invention
[0007] The technical problem to be solved by this invention is to propose a method for establishing a hydrological model based on the consistency of the distribution characteristics of verification and validation data. By discretizing sampling, the performance consistency of the hydrological process model during the verification and validation period is guaranteed, thereby improving the effectiveness of the hydrological process model and the stability of its engineering applications.
[0008] The specific technical solution adopted in this invention is as follows:
[0009] A method for establishing a hydrological process model based on the consistency of data distribution characteristics in verification and validation, comprising the following steps:
[0010] S1: Propose a data discretization verification approach for hydrological process models according to steps S11-S12.
[0011] S11: Set the overall data to the "startup" stage. This data will not be used for model verification and validation, but only for setting the initial parameters of the model to reduce initialization errors. The hydrological process model structure is specified by the user.
[0012] S12: Abandoning the traditional practice of using continuous time series data for modeling hydrological processes, and taking the consistency of data distribution characteristics between the calibration and validation periods as the goal, runoff data is discretely allocated to the calibration and validation datasets according to a discretized data allocation method;
[0013] S2: Following steps S21-28, use the MDUPLEX method to assign the original runoff dataset D to the check set C and the validation set E.
[0014] S21: First back up D and record it as D. b Determine the proportion of data allocated to C and E based on user needs. C ,P E ;
[0015] S22: Determine the size n of the basic sampling pool (n pairs of data), calculated using the following formula:
[0016]
[0017] S23: Determine the number of sample pairs n allocated from the basic sampling pool to C and E. C and n E :
[0018]
[0019]
[0020] S24: Find the pair of data x with the greatest Euclidean distance in D. i ,x j And allocate to C using sampling without replacement;
[0021] S25: Repeat step S24, assigning a pair of data to E;
[0022] S26: Find the next pair of data x in D. i ,x j The first data x i The single-linkage distance with C is the greatest, and the second data is x. j Next;
[0023] S27: Repeat step S26, assigning a pair of data to E, and repeat this sampling method until the allocation amount n determined by the basic sampling pool is satisfied. C and n E When one of the allocations reaches the required amount, all subsequent sampled data pairs are allocated to the other dataset.
[0024] S28: At this point, the sampling work of the first basic sampling pool is completed. Then, the next round of basic sampling pools will begin, and steps S26-27 will be repeated until all data in D is allocated to C and E.
[0025] S3: Verify and validate the model according to steps S31-32, determine the model parameters, and establish a hydrological process model.
[0026] S31: Model in D b The process runs continuously from start to finish, and the data in the verification set C is used for model parameter selection.
[0027] S32: Again in D b The model was run continuously and its predictive performance was validated using data from the validation set E. At this point, the model was successfully built.
[0028] Compared with the prior art, the present invention has the following advantages:
[0029] (1) Compared with the traditional hydrological process model verification method, the discretization verification method proposed in this invention does not require the time series relationship between data, which means that the available data allocation methods are greatly increased, which is conducive to maintaining the consistency of data distribution characteristics during the model verification period and the validation period.
[0030] (2) The MDUPLEX method proposed in this invention can effectively classify the original observation data into subsets with similar distribution characteristics, and has no requirements on the time span of the data.
[0031] (3) This invention proposes a new approach to the verification and validation of hydrological process models, which can effectively solve the problem of inconsistent data distribution characteristics allocated during the verification and validation periods of such models, thereby ensuring the robustness of model performance, enabling the model to have good transferability, improving the accuracy of its simulation and prediction, and enhancing the reliability of the model. Attached Figure Description
[0032] Figure 1 This is a roadmap for the specific implementation of the present invention.
[0033] Figure 2 This is a schematic diagram comparing the method of the present invention with traditional methods.
[0034] Figure 3This is a comparison chart of the overall performance distribution of models established by different methods in 163 watersheds in Example 163.
[0035] Figure 4 This is a comparison chart of the overall performance distribution of the model after grouping the 163 watersheds according to the data runoff skewness in Example 1.
[0036] Figure 5 This is a comparison chart of the performance changes of models established by different methods in 163 watersheds during the verification and validation periods.
[0037] Figure 6 This is a comparison chart of the performance changes of the model during the verification and validation periods after grouping the runoff skewness data of 163 watersheds in Example 163. Detailed Implementation
[0038] The present invention will now be described in detail with reference to the accompanying drawings and embodiments, so that those skilled in the art can better understand the essence of the present invention.
[0039] See Figure 1 A method for establishing a hydrological process model based on the consistency of data distribution characteristics in verification and validation is described below:
[0040] S1: Propose a data discretization verification approach for hydrological process models according to steps S11-S12.
[0041] S11: Set the overall data to the "startup" stage. This data will not be used for model verification and validation, but only for setting the initial parameters of the model to reduce initialization errors. The hydrological process model structure is specified by the user.
[0042] S12: Abandoning the traditional practice of using continuous time series data for modeling hydrological processes, and taking the consistency of data distribution characteristics between the calibration and validation periods as the goal, runoff data is discretely allocated to the calibration and validation datasets according to a discretized data allocation method;
[0043] S2: Following steps S21-28, use the MDUPLEX method to assign the original runoff dataset D to the check set C and the validation set E;
[0044] S21: First back up D and record it as D. b Determine the proportion of data allocated to C and E based on user needs. C ,P E ;
[0045] S22: Determine the size n of the basic sampling pool (n pairs of data), calculated using the following formula:
[0046]
[0047] S23: Determine the number of sample pairs n allocated from the basic sampling pool to C and E. C and n E :
[0048]
[0049]
[0050] S24: Find the pair of data x with the greatest Euclidean distance in D. i ,x j And allocate to C using sampling without replacement;
[0051] S25: Repeat step S24, assigning a pair of data to E;
[0052] S26: Find the next pair of data x in D. i ,x j The first data x i The single-linkage distance with C is the greatest, and the second data is x. j Next;
[0053] S27: Repeat step S26, assigning a pair of data to E, and repeat this sampling method until the allocation amount n determined by the basic sampling pool is satisfied. C and n E When one of the allocations reaches the required amount, all subsequent sampled data pairs are allocated to the other dataset.
[0054] S28: At this point, the sampling work of the first basic sampling pool is completed. Then, the next round of basic sampling pools will begin, and steps S26-27 will be repeated until all data in D is allocated to C and E.
[0055] S3: Verify and validate the model according to steps S31-32, determine the model parameters, and establish a hydrological process model;
[0056] S31: Model in D b The process runs continuously from start to finish, and the data in the verification set C is used for model parameter selection.
[0057] S32: Again in D b The model was run continuously and its predictive performance was validated using data from the validation set E. At this point, the model was successfully built.
[0058] Figure 2 This is a schematic diagram comparing the method of the present invention with traditional methods. It can be seen that the core idea of the present invention is to obtain better model performance through data discretization verification.
[0059] The following section combines this method with specific embodiments to demonstrate its specific technical effects; the specific steps of the method will not be repeated here.
[0060] Example
[0061] The method proposed in this invention, along with traditional continuous data verification methods, is applied to a conceptual rainfall-runoff model, a hydrological model based on real physical processes. The advantages of this invention are demonstrated through statistical testing on a large number of watersheds.
[0062] Three well-known process rainfall-runoff (CRR) models, namely GR4J, AWBM, and CMD, were selected for testing. The dataset consisted of 163 publicly available watersheds. These watershed data were all analyzed and processed by previous researchers, and the data time span exceeded 30 years, which met the data length requirements of the CRR model.
[0063] The KGE metric is used to evaluate model performance, with values ranging from negative infinity to 1. A KGE closer to 1 indicates a better fit. To facilitate the evaluation of the overall simulation performance and robustness across different time periods, KGE is defined here. ALL The overall KGE value of the model across the entire dataset is given by ΔKGE, which is defined as the difference between the KGE values during the model validation period and the calibration period.
[0064] Furthermore, the concept of Generative Adversarial Networks (GANs) was employed to evaluate the consistency of data distribution between the model's calibration and validation periods by training the classifier through adversarial validation. The evaluation metric used was AUC (Area Under Curve). An AUC value close to 0.5 indicates that the data distribution characteristics of the two datasets are consistent, and the classifier cannot distinguish the source of the data. If the AUC value is close to 1, it indicates that the distribution differences between the two datasets are extremely significant, and the classifier can accurately distinguish the source of the data.
[0065] Data from monitoring stations in watershed number 10 were selected, and the model structure used was GR4J to demonstrate the specific implementation process of the model construction method proposed in this invention:
[0066] (1) The dataset records data from January 1, 1970 to December 31, 2013, with precision in days and a total length of 16071. The first 365 days of data are used as the "startup" phase of the model to initialize the model parameters.
[0067] (2) Set the data ratio used during the verification period and the validation period to P. C =0.6,P E If the value is 0.4, then the amount of data required for C is 9643, and the amount of data required for E is 6428.
[0068] (3) Take a pair of data for C and E in accordance with step S24.
[0069] (4) The total number of sample pairs in the basic sampling pool is determined by formula 1-1, n = 3. Then the number of sample pairs allocated to C and E is calculated by formulas 1-2 and 1-3, n. C =2 and n E =1.
[0070] (5) Every 3 samplings constitute a basic sampling pool. In each sampling pool, a pair of data is sampled for C according to step S26, and then a pair of data is sampled for E. At this time, the number of samplings of E in this round of sampling pool has reached the requirement, and only one more pair of data needs to be sampled for C to end this round of sampling.
[0071] (6) If there is still data in D, then the next round of basic sampling pool sampling will be carried out.
[0072] (7) After all data has been distributed, modeling is performed using all data (i.e., backup of D). The model parameters are checked using data within C. A set of optimal parameters is found to maximize the KGE values of the simulated and observed values at each point in C. In the current No. 10 watershed, the KGE during the check period is 0.82.
[0073] (8) A model was built using the parameters obtained after verification. The data range used in the model is still the entire dataset (i.e., 16071 data points). The simulated and observed values of each point in E at the corresponding positions in the model were compared, and the KGE value was calculated. By comparing the KGE values of the two periods, the degree of change in the simulation performance of the model can be obtained. In the current No. 10 watershed, the KGE during the validation period was 0.81, and it can be seen that ΔKGE = -0.01. This indicates that the model established using the method of this invention in the No. 10 watershed has extremely high performance robustness.
[0074] To compare the effectiveness of the method of this invention with the traditional CRR modeling method, modeling was performed in 163 watersheds using the method of this invention and three traditional methods, thereby statistically comparing the impact of different modeling methods on model performance.
[0075] Figure 3 This is a comparison chart of the overall performance of models built using the method of this invention and conventional methods across all 163 watersheds, including AUC and KGE. ALL The distribution is shown in the graph, with each line containing 163 data points. From... Figure 3(a) It is evident that the distribution characteristics of the data selected using the MDUPLEX method through discretization maintain high consistency between the verification and validation periods, with the AUC distribution generally remaining between 0.5 and 0.6. In contrast, traditional modeling methods clearly exhibit inconsistent data characteristic distributions between the verification and validation periods, with AUC values ranging from 0.5 to 0.9, which is consistent with previous research findings. Figure 3 (b), (c), and (d) show the overall performance comparison results of the models established using the method of this invention and the traditional method in three different CRR models, respectively. As can be seen from the results, the model obtained by the method of this invention has a significant statistical advantage over the traditional CRR modeling method.
[0076] Figure 4 It is based on the skewness of runoff data. Figure 3 The results were grouped and compared to account for the influence of the watershed's own data characteristics (data distribution skewness). From Figure 4 The results show that when the data skewness is small, the KGE obtained by the method of this invention is comparable to that obtained by the traditional method. ALL The KGE values are all relatively high, with a median around 0.9. This is because when the data skewness is small, the difference in data distribution characteristics between the verification and validation periods is also relatively small, and traditional methods can also achieve good performance. However, when the data skewness of the watershed is large, the advantages of the method of this invention are significantly improved, and the KGE values of the method of this invention under the three CRR models are higher. ALL The median values are all above 0.85, while those obtained by traditional methods are reduced to around 0.8. This indicates that the overall performance of the model obtained by the method of this invention is significantly better than that of traditional methods in watersheds with high data skewness.
[0077] Figure 5 This chart compares the performance differences of models established across all 163 watersheds during the calibration and validation periods. The closer the ΔKGE value is to 0, the smaller the performance change, indicating a more robust model. Figure 5 The results show that the ΔKGE distribution of the method of this invention is largely concentrated between -0.1 and 0, while the ΔKGE distribution of the traditional method is distributed between -0.6 and 0. This indicates that the performance degradation of the model under the method of this invention during the validation period is significantly less than that of the three traditional continuous data verification modeling methods. This also demonstrates that the discretized data sampling method enables the model to learn the real hydrological processes within the watershed to a greater extent, and the parameters obtained by the model training during the verification period do not overfit the verification set data, thus maintaining good performance during the validation period.
[0078] Figure 6 It is based on the skewness of runoff data. Figure 5The results were grouped and compared to examine the changes in model performance between the proposed method and traditional methods under different data characteristics over two periods. Figure 6 The results show that when the skewness of runoff data increases, the performance difference of the model during the calibration and validation periods will become significantly larger. The performance difference of the method of this invention has a very obvious advantage in watersheds with high skewness. The ΔKGE distribution is mostly concentrated between -0.2 and 0, while the ΔKGE distribution of the traditional method is distributed from -0.6 to 0. Figure 6 The results show that the method of the present invention has significantly better robustness than the traditional method in different periods under high skewness watersheds.
[0079] The results above demonstrate that the discretized hydrological process model verification method proposed in this invention has significant advantages over traditional continuous data sampling modeling methods. It effectively improves the overall performance and predictive performance of conceptual rainfall-runoff models, and the performance difference between the verification and validation periods is significantly reduced. Furthermore, the applicability of the model construction method is not limited to the three CRR models used in this embodiment. Based on the principle of consistency in the distribution characteristics of verification and validation data, the proposed model construction method can theoretically be widely applied to any other process-based hydrological process model, exhibiting broad application prospects and significant potential for promotion and practical application.
Claims
1. A method for establishing a hydrological process model based on the consistency of data distribution characteristics in verification and validation, characterized in that, The steps are as follows: S1: Propose a data discretization verification approach for hydrological process models according to steps S11-S12; S11: Set the overall data to the "startup" stage. This data will not be used for model verification and validation, but only for setting the initial parameters of the model to reduce initialization errors. The hydrological process model structure is specified by the user. S12: Abandoning the traditional practice of using continuous time series data for modeling hydrological processes, and taking the consistency of data distribution characteristics between the calibration and validation periods as the goal, runoff data is discretely allocated to the calibration and validation datasets according to a discretized data allocation method; S2: Following steps S21-28, use the MDUPLEX method to assign the original runoff dataset D to the check set C and the validation set E; S21: First back up D and record it as D. b Determine the proportion of data allocated to C and E based on user needs. C ,P E ; S22: Determine the size n of the basic sampling pool, i.e., n pairs of data. The calculation formula is as follows: n =[1 / min( P C , P E )+0.5] 1-1 S23: Determine the number of sample pairs allocated from the basic sampling pool to C and E. and : n C =[ n × P C +0.5] 1-2 n E =[ n × P E +0.5] 1-3 S24: Find the pair of data x with the greatest Euclidean distance in D. i ,x j And allocate to C using sampling without replacement; S25: Repeat step S24, assigning a pair of data to E; S26: Find the next pair of data in D. , The first data The distance to the single-linkage of C is the greatest, the second data point. Next; S27: Repeat step S26, assigning a pair of data to E, and repeat this sampling method until the allocation determined by the basic sampling pool is satisfied. and When one of the allocations reaches the required amount, all subsequent sampled data pairs are allocated to the other dataset. S28: At this point, the sampling work of the first basic sampling pool is completed. Then, the next round of basic sampling pools will begin, and steps S26-27 will be repeated until all data in D is allocated to C and E. S3: Verify and validate the model according to steps S31-32, determine the model parameters, and establish a hydrological process model; S31: Model in D b The process runs continuously from start to finish, and the data in the verification set C is used for model parameter selection. S32: Again in D b The model was run continuously and its predictive performance was validated using data from the validation set E. At this point, the model was successfully built.
Citation Information
Patent Citations
Hydrological model parameter time-varying form construction method
CN112883558A
Hydro-meteorological time sequence-based runoff simulation method and system
WO2021217776A1