A Method and System for Error Analysis of Runoff Reanalysis Data Based on Panel Regression Analysis

By constructing a fixed-effects model using panel regression analysis and bootstrap clustering, the problem of diagnosing systematic errors in runoff reanalysis caused by meteorological errors in large-sample watersheds was solved, achieving error analysis with high accuracy and reliability.

CN119646546BActive Publication Date: 2025-10-31SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411755276.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-10-31
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively diagnose the qualitative effects of meteorological errors on the annual-scale systemic errors in runoff reanalysis in large sample watersheds, leading to systemic bias in global runoff reanalysis data.

Method used

A fixed-effects panel regression model for meteorological input bias and runoff simulation bias was constructed using panel regression analysis. The watershed was clustered and resampled using the bootstrap method to generate a coefficient distribution for measuring the spatial heterogeneity effect of meteorological bias.

Benefits of technology

Effectively diagnose the heterogeneity of meteorological errors in the annual-scale systematic error reanalysis of runoff in large sample watersheds, reduce the impact of omitted variables and cross-watershed heterogeneity, and improve the accuracy and reliability of error analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646546B_ABST
    Figure CN119646546B_ABST
Patent Text Reader

Abstract

This invention relates to the field of hydrological data analysis technology, and proposes a method and system for error analysis of runoff reanalysis data based on panel regression analysis. The method includes the following steps: collecting watershed hydrological and meteorological datasets, watershed runoff reanalysis data, and meteorological reanalysis data, and extracting watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data; constructing a fixed-effects panel regression model of meteorological input deviation and runoff simulation deviation based on the deviation between the reanalysis data and the observed values; clustering all watersheds according to the watershed station attribute data to obtain the spatial clusters of the watersheds and the attribute characteristics of each cluster, and resampling each cluster using a bootstrapping method to form the first panel data; performing fixed-effects panel regression on the first panel data to obtain the corresponding regression coefficients, and generating a first coefficient distribution for measuring the spatial heterogeneity effect of meteorological deviation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydrological data analysis technology, and more specifically, to a method and system for error analysis of runoff reanalysis data based on panel regression analysis. Background Technology

[0002] Global runoff reanalysis data, driven by meteorological reanalysis data, is the output of global hydrological models and provides continuous time-series runoff data, widely used to support water resource adaptability assessments and management decisions. Precipitation and evaporation are the main drivers of global runoff change; however, climate models are prone to bias in simulating precipitation and potential evapotranspiration variables. When meteorological reanalysis data is biased, its input into global hydrological models (GHMs) for runoff simulation leads to systematic bias in the global runoff reanalysis data. Due to the heterogeneity of the land surface and the complexity of the interaction between climate inputs and watershed characteristics, it is currently difficult to conduct effective and highly accurate diagnostic analysis of the qualitative effects of meteorological errors on the annual-scale systematic errors in runoff reanalysis in large-sample watersheds. Summary of the Invention

[0003] To overcome the shortcomings of existing technologies in effectively diagnosing the qualitative effects of meteorological errors on the annual-scale systemic errors in runoff reanalysis in large sample watersheds, this invention provides a method and system for error analysis of runoff reanalysis data based on panel regression analysis.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0005] A method for error analysis of runoff reanalysis data based on panel regression analysis includes the following steps:

[0006] Collect watershed hydrological and meteorological datasets, watershed runoff reanalysis data, and meteorological reanalysis data, and extract watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data;

[0007] Based on the deviation between reanalysis data and observed values, a fixed-effects panel regression model is constructed to compare meteorological input deviation with runoff simulation deviation.

[0008] All watersheds are clustered based on the watershed station attribute data to obtain the spatial clusters of the watersheds and the attribute characteristics of each cluster. Each cluster is then resampled using a bootstrap method to form the first panel data. Fixed-effects panel regression is performed on the first panel data to obtain the corresponding regression coefficients, generating the first coefficient distribution used to measure the spatial heterogeneity effect of meteorological deviations.

[0009] Furthermore, this invention also proposes a runoff reanalysis data error analysis system based on panel regression analysis, applying the runoff reanalysis data error analysis method proposed in this invention. The system includes:

[0010] The data processing module is used to collect watershed hydrological and meteorological datasets, watershed runoff reanalysis data, and meteorological reanalysis data, and to extract watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data.

[0011] The first sampling module is used to cluster all watersheds based on the watershed station attribute data to obtain the spatial clusters of the watersheds and the attribute features of each cluster, and to resample each cluster using the bootstrap method to form the first panel data.

[0012] The fixed-effects panel regression module is equipped with a fixed-effects panel regression model for meteorological input bias and runoff simulation bias. It is used to perform fixed-effects panel regression on panel data, obtain the corresponding regression coefficients, and generate the first coefficient distribution to measure the spatial heterogeneity effect of meteorological bias.

[0013] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0014] This invention utilizes a fixed-effects panel regression model to analyze errors in runoff reanalysis data. It introduces fixed effects in the individual and time dimensions, and captures the differences between individuals and / or between times in the watershed through the fixed-effects panel regression model. This reduces the bias caused by omitted variables, controls the heterogeneity of unobservable variables between different watersheds and over time, and effectively diagnoses the heterogeneous effects of meteorological errors on the annual-scale systematic errors in runoff reanalysis in large sample watersheds.

[0015] This invention uses a fixed-effects panel regression model to analyze the impact of meteorological input errors on the annual-scale systematic errors of runoff simulation. It solves the problems of traditional statistical analysis methods, which are easily affected by omitted variables, cross-basin heterogeneity, and time-varying factors. This makes it easier to assess the systematic errors and diagnose the influencing factors of runoff reanalysis data in large sample watersheds. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating an error analysis method for runoff reanalysis data based on panel regression analysis according to an embodiment of the present invention.

[0017] Figure 2 This is a schematic diagram illustrating the annual variation of the deviation between CAMELS watershed reanalysis data and observation data according to an embodiment of the present invention.

[0018] Figure 3 This is an annual average scatter plot of CAMELS watershed reanalysis data and observation data according to an embodiment of the present invention.

[0019] Figure 4 This is a schematic diagram illustrating the watershed attributes of various clusters in a CAMELS watershed according to an embodiment of the present invention. Figure 5 This is a schematic diagram illustrating the sensitivity of runoff quantile deviation to meteorological deviation under different watershed clusterings according to an embodiment of the present invention.

[0020] Figure 6 This is a schematic diagram illustrating the sensitivity of seasonal-scale runoff quantile deviation to meteorological deviation according to an embodiment of the present invention.

[0021] Figure 7 This is a schematic diagram illustrating the sensitivity of runoff quantile deviation to precipitation deviation in each watershed cluster according to an embodiment of the present invention.

[0022] Figure 8 This is a schematic diagram illustrating the sensitivity of runoff quantile deviation of each watershed cluster to potential evapotranspiration deviation according to an embodiment of the present invention.

[0023] Figure 9 This is a schematic diagram illustrating the influence of watershed attributes on precipitation deviation effects according to an embodiment of the present invention.

[0024] Figure 10 This is a schematic diagram illustrating the influence of watershed properties on potential evapotranspiration deviation effects according to an embodiment of the present invention.

[0025] Figure 11 This is an architectural diagram of a runoff reanalysis data error analysis system based on panel regression analysis, according to an embodiment of the present invention. Detailed Implementation

[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0027] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0028] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0029] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0030] Example 1

[0031] This embodiment proposes a method for error analysis of runoff reanalysis data based on panel regression analysis, such as... Figure 1 The diagram shown is a flowchart of the runoff reanalysis data error analysis method in this embodiment.

[0032] The runoff reanalysis data error analysis method proposed in this embodiment includes the following steps:

[0033] S1. Collect watershed hydrological and meteorological datasets, watershed runoff reanalysis data, and meteorological reanalysis data, and extract watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data;

[0034] S2. Based on the deviation between reanalysis data and observed values, a fixed-effects panel regression model of meteorological input deviation and runoff simulation deviation is constructed.

[0035] S3. Cluster all watersheds based on the watershed station attribute data to obtain the spatial clusters of the watersheds and the attribute characteristics of each cluster. Then, resample each cluster using the bootstrap method to form the first panel data. Perform fixed-effects panel regression on the first panel data to obtain the corresponding regression coefficients and generate the first coefficient distribution used to measure the spatial heterogeneity effect of meteorological deviations.

[0036] The fixed effects panel regression model used in this embodiment is a statistical model for analyzing panel data, which refers to observational data of the same group of individuals (i.e., watershed) at multiple time points. This model is particularly suitable for studying the impact of individual effects on the relationships between variables, and for controlling for individual effects when they may be related to explanatory variables.

[0037] This embodiment uses a fixed-effects panel regression model to analyze the error of runoff reanalysis data. It introduces fixed effects in the individual and time dimensions, and captures the differences between individuals in the watershed and / or between times through the fixed-effects panel regression model, reducing the bias caused by omitted variables, and controlling the heterogeneity of unobservable variables between different watersheds and over time, thereby improving the accuracy of error analysis.

[0038] This embodiment uses a fixed-effects panel regression model to analyze the impact of meteorological input errors on the annual-scale systematic errors of runoff simulation. This solves the problem that traditional statistical analysis methods are easily affected by omitted variables, cross-basin heterogeneity, and time-varying factors, making it easier to assess systematic errors and diagnose influencing factors in large-sample watersheds for runoff reanalysis data.

[0039] In an optional embodiment, step S1, extracting watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data, includes the following steps:

[0040] For watershed runoff reanalysis data, watershed runoff reanalysis time series data are extracted from the watershed runoff reanalysis data based on the latitude and longitude attributes of the watershed outlet hydrological station;

[0041] For watershed meteorological reanalysis data, the raster data in the meteorological reanalysis data is cropped according to the vector boundary of the watershed, and only the raster information inside the watershed is retained;

[0042] For the cropped raster data, calculate the actual area of ​​each raster within the watershed;

[0043] The cropped raster data is traversed to extract the meteorological value and its corresponding area weight for each pixel; the meteorological value includes precipitation, potential evapotranspiration, and temperature.

[0044] A time series of watershed meteorological reanalysis data is generated by calculating the area-weighted average based on the meteorological value and area weight of each pixel.

[0045] The expression for calculating the area-weighted average is as follows:

[0046]

[0047] For each watershed, the above calculations are repeated to finally generate time series data for watershed meteorological reanalysis.

[0048] It should be noted that the watershed hydrological and meteorological dataset collected in step S1 is observational data, while the watershed runoff reanalysis data and meteorological reanalysis data are simulated data.

[0049] Each watershed unit contains multiple grids with different values, and some grids do not fall completely within the watershed vector boundary, so extraction is required.

[0050] In this embodiment, the area-weighted average method is selected to extract meteorological reanalysis data, resulting in time-series data at the watershed level. Its functions include:

[0051] 1) The area-weighted average takes into account the relative size and position of each grid in the watershed, which can more accurately reflect the true situation of meteorological variables in the watershed and effectively improve the accuracy of error analysis;

[0052] 2) Area-weighted averaging can effectively integrate spatially heterogeneous meteorological conditions within a watershed;

[0053] 3) Use area-weighted averaging to eliminate biases caused by excessive or insufficient influence of meteorological data from certain small areas on the basin average, thus ensuring the reasonableness of the results;

[0054] 4) The average value of each watershed is calculated based on the same weighting principle, which enables reasonable comparison of meteorological data from different watersheds or different time periods, ensuring the comparability of the data.

[0055] In an optional embodiment, in step S2, a fixed-effects panel regression model of meteorological input bias and runoff simulation bias is constructed based on the bias between the reanalysis data and the observed values, including the following steps:

[0056] Construct a panel dataset; the panel dataset includes the annual runoff quantile deviation Q for each watershed. bias Bias P of annual meteorological data bias PET bias and T bias Its expression is:

[0057]

[0058] Wherein, the subscript bias indicates the watershed bias, the subscript reanalysis indicates the area-weighted time series of the watershed reanalysis data, and the subscript obs indicates the time series of the watershed observation data; specifically, P bias P represents the deviation of annual precipitation in the basin. reanalysis P represents annual precipitation in the time series data of watershed runoff reanalysis. obs This represents annual precipitation in time-series data of watershed observations; PET bias Indicates the annual average potential evapotranspiration deviation of the watershed, PET reanalysis PET represents the annual average potential evapotranspiration under watershed runoff reanalysis time series data. obsT represents the annual average potential evapotranspiration under time series data of watershed observation data; bias T represents the annual average temperature deviation of the watershed. reanalysis T represents the annual average temperature in the watershed runoff reanalysis time series data. obs Q represents the annual average temperature under time series data of watershed observation data; bias Q represents the annual runoff quantile deviation of the watershed. reanalysis Q represents the runoff quantile in the watershed runoff reanalysis time series data. obs This represents the runoff quantile in the time series data of watershed observation data;

[0059] Based on the aforementioned panel dataset, a fixed-effects panel regression model is constructed to compare meteorological input bias with runoff simulation bias; its expression is as follows:

[0060]

[0061] in, P represents the systematic bias of the runoff percentile q in watershed i over time period t. i,t PET represents the precipitation deviation in watershed i over time period t. i,t T represents the average potential evapotranspiration deviation of watershed i over time period t. i,t β1 represents the average temperature deviation in watershed i over time period t; β2 and β1 are the regression coefficients of precipitation deviation and potential evapotranspiration deviation, respectively; β3 is the regression coefficient of the coupling effect between temperature and precipitation deviation; μ i For the fixed effect of watershed i, γ t For time-fixed effects over time period t, ε i,t Let be the error term for watershed i during time period t.

[0062] In this embodiment, the deviation between the reanalysis data and the observed values ​​refers to the deviation obtained by subtracting the runoff reanalysis data and meteorological reanalysis data (used as simulated data) from the observed data extracted from the hydro-meteorological dataset within the same time period under any watershed. The annual-scale runoff quantile deviation and annual-scale meteorological data deviation for any watershed are then calculated. The expression is as follows:

[0063] bias i =sim i -obs i

[0064] Where, sim i It is the runoff reanalysis time series of watershed i, obs i These are observational data for watershed i, bias i This is the deviation value of watershed i. The deviation P of the annual precipitation in the watershed is... biasAnnual average potential evapotranspiration deviation of the basin (PET) bias and the annual average temperature deviation T of the basin bias This constitutes the bias in annual-scale meteorological data.

[0065] In an optional embodiment, step S3 involves resampling each cluster using a bootstrap method to form the first panel data, including the following steps:

[0066] The watersheds are classified using the self-organizing map clustering method based on their watershed attribute values, resulting in spatial clustering results and attribute characteristics of each cluster.

[0067] For each cluster, watershed subsets are sampled with replacement, where the number of samples drawn each time is the number of watersheds in that cluster. This constructs a panel dataset of watershed subsets, resulting in the first panel dataset.

[0068] In this embodiment, the watersheds are first classified according to their attributes, and those with similar attribute characteristics are grouped into one category, thereby obtaining the spatial clustering of the watersheds and the attribute characteristics of each cluster.

[0069] Furthermore, Bootstrap sampling with replacement was used, with the number of samples drawn each time corresponding to the number of watersheds in the cluster. A panel dataset of the subset was then constructed, and a fixed-effects panel regression model was used to obtain the corresponding regression coefficients. This step was repeated several times to obtain the coefficient distribution interval for the cluster. This process was repeated for all clusters to obtain the coefficient distribution for each cluster, thus obtaining the spatial effect distribution of meteorological deviations on the systematic deviations of different runoff quantiles.

[0070] For example, in one specific implementation, a bootstrap sampling with replacement is performed on a subset of the watershed in a cluster to construct a first panel dataset. The corresponding regression coefficients are then obtained through a fixed-effects panel regression model. This process is repeated 1000 times to obtain the coefficient distribution interval for that cluster. This process is repeated for each cluster to obtain the coefficient distribution for each cluster, thus obtaining the spatial effect distribution of meteorological deviations on the systematic deviations of different runoff quantiles.

[0071] Furthermore, in an optional embodiment, step S3 further includes the following step:

[0072] The watershed hydrological and meteorological dataset is resampled using the bootstrap method to form a second panel data. Fixed-effects panel regression is then performed on the second panel data to obtain the corresponding regression coefficients, generating a second coefficient distribution to measure the uncertainty of the regression coefficients.

[0073] In this embodiment, a bootstrap method is used to quantify the uncertainty of regression coefficients. Specifically, for each bootstrap, sampling with replacement is performed from a large sample watershed dataset, with the number of samples drawn being the number of watershed samples. Some watersheds may be sampled multiple times, while others may not be sampled at all. Each bootstrap generates a new panel dataset, and fixed-effects panel regression is performed on the new panel data to obtain the corresponding regression coefficients. This step is repeated several times to obtain the statistics of the coefficients. Furthermore, a preset confidence interval is taken for each coefficient to obtain the coefficient distribution confidence interval used to measure the uncertainty of the response coefficients.

[0074] For example, in one specific implementation, second panel data is obtained by sampling with replacement from a large sample watershed dataset. Fixed-effects panel regression is performed on the second panel data to obtain the corresponding regression coefficients. This process is repeated 1000 times to obtain the statistics of the corresponding coefficients. Then, a 95% confidence interval is taken for each coefficient to obtain a second coefficient distribution used to measure the uncertainty of the regression coefficients.

[0075] Furthermore, in an optional embodiment, step S3 further includes the following step:

[0076] All watersheds are divided into three categories based on the magnitude of the attribute values ​​of each watershed attribute;

[0077] The attributes of each watershed are resampled using the bootstrap method to form a third panel data. Fixed-effects panel regression is then performed on the third panel data to obtain the corresponding regression coefficients, generating a third coefficient distribution to measure the degree of influence of each watershed attribute value on the meteorological deviation effect.

[0078] In this embodiment, the influence of watershed attribute values ​​on meteorological deviation effects is analyzed by classifying watershed attributes according to their magnitude and then performing resampling and fixed-effects panel regression analysis.

[0079] For example, in one specific implementation process, all watersheds are divided into three categories—small, medium, and large—based on the magnitude of their watershed attribute values. Bootstrap resampling is then performed on each of these three categories to obtain third panel data. Fixed-effects panel regression is then performed on this third panel data to obtain the corresponding regression coefficients. This process is repeated 1000 times to obtain the distribution of the corresponding coefficients. Finally, a Mann-Whitney U test is performed on each of the three categories for each attribute value to determine whether statistically different attribute values ​​have a significant impact on meteorological deviations.

[0080] The watershed attributes include watershed area, average watershed slope, precipitation-to-snow water ratio, forest coverage, and aridity.

[0081] Example 2

[0082] This embodiment applies the runoff reanalysis data error analysis method proposed in Embodiment 1 to perform systematic error source analysis on GloFAS runoff reanalysis data in a large sample watershed of CAMELS.

[0083] Step 1: Collect watershed hydrological and meteorological datasets, watershed runoff reanalysis data, and meteorological reanalysis data, and extract watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data.

[0084] In a specific implementation process, the third-party Python libraries Pandas, Xarray, and Geopandas are used to read runoff reanalysis data (.nc) and meteorological reanalysis data (.nc) and store them in variables named Q_reanalysis, P_reanalysis, PET_reanalysis, and T_reanalysis, respectively. Watershed hydrological and meteorological observation data are read using Pandas and stored in variables named Q_obs, P_obs, PET_obs, and T_obs, respectively. Watershed attributes are stored in the attribute field.

[0085] Then, the meteorological data was extracted and reanalyzed at the watershed scale using the function nc_remapper() of the Python third-party library EASYMORE in combination with the watershed attribute vector boundary file .shp. The runoff reanalysis time series of the watershed stations was extracted from the runoff reanalysis grid data using the latitude and longitude information of the hydrological station at the watershed outlet.

[0086] Step 2: Compare the meteorological reanalysis and runoff reanalysis time series with the observation data to obtain the deviation, and construct a fixed-effects panel regression model of meteorological input deviation and runoff simulation deviation.

[0087] In a specific implementation process, Pandas is used to add timestamps to the reanalysis data time series and the observation data time series, calculate the daily scale mean and the daily scale mean deviation, and obtain the deviation between the reanalysis data and the observation values.

[0088] For example, such as Figure 2 The figure shows the intra-annual variation of reanalysis and observational data for 671 watersheds in the CAMELS dataset. As can be seen, the variability of runoff reanalysis is smaller than that of observed runoff. Precipitation reanalysis and observational data show large diurnal variations throughout the year, with the mean daily precipitation deviation fluctuating between -0.5 and 0.5 mm. Potential evapotranspiration reanalysis data is lower than observed values ​​during the summer when evapotranspiration is high, and higher than observed values ​​during the spring and winter when evapotranspiration is low, and its deviation exhibits strong seasonality. Temperature reanalysis data agrees well with observed values ​​compared to other hydrological variables, and temperature reanalysis data is higher than observed values ​​throughout the year. Figure 3 As shown, this is an annual average scatter plot of reanalysis and observation data for 671 watersheds in the CAMELS dataset, displaying the annual average observation and reanalysis scatter values ​​for runoff, precipitation, potential evapotranspiration, and air temperature.

[0089] Then, the PanelOLS function from the Python third-party library linearmodels is used to establish the panel regression equation.

[0090] Step 3: After classifying the large sample watersheds according to their watershed attributes, resample the samples to obtain panel data.

[0091] In one specific implementation process, the watershed attributes are clustered using the third-party libraries minisom and scipy.cluster.hierarchy to obtain the cluster number of each watershed, and the watersheds are divided into 10 categories.

[0092] like Figure 4 The diagram shown illustrates the watershed attributes of each cluster within the CAMELS watershed. Figure 4 It can be seen that: Cluster 1 has a small watershed area, a relatively steep slope, precipitation concentrated in winter, few heavy precipitation events, and a relatively large watershed runoff; Cluster 2 has a large watershed area, thick soil, gentle terrain, precipitation concentrated in summer, and a large baseflow; Cluster 3 has relatively low seasonality of precipitation and a perennially humid watershed; Cluster 4 has a small watershed area, high soil permeability, low groundwater porosity, large forest cover, and high runoff variability; Cluster 5 has low seasonality of precipitation, a low precipitation-to-snow water ratio, low soil permeability, and large forest cover; Cluster 6 has a steep slope, high aridity, and is greatly affected by snowfall; Cluster 7 has high aridity, steep slope, low forest cover, and a long period of low flow; Cluster 8 has flat terrain, thick soil, high permeability, and large forest cover; Cluster 9 has flat terrain, a large watershed area, a dry climate, and highly seasonal precipitation; Cluster 10 has thick soil, little forest cover, is affected by snowfall in winter, and is dry in summer.

[0093] Panel datasets were constructed by annually analyzing meteorological reanalysis biases and runoff reanalysis biases for all watersheds.

[0094] Step 4: Perform fixed-effects panel regression analysis to obtain the error analysis results.

[0095] In a specific implementation process, the panel dataset is sampled with replacement using the Python third-party library NumPy to obtain a new panel dataset. This process is repeated 1000 times to obtain 1000 subsets of panel data. Then, the fixed-effects panel regression equation function is called for each subset to obtain the corresponding regression coefficients.

[0096] For panel datasets obtained by clustering based on watershed station attributes and Bootstrap resampling, the distribution of regression coefficients obtained after fixed-effects panel regression can be used to measure the spatial heterogeneity effect of meteorological biases and visualize them as error analysis results.

[0097] For example, such as Figure 5 , 6 The diagram shows the sensitivity of runoff quantile deviation to meteorological deviation under different watershed clustering, and the sensitivity of runoff quantile deviation to meteorological deviation at the seasonal scale. Figure 5 As can be seen, runoff deviation becomes sensitive to deviations in meteorological elements starting from Q50. With increasing runoff quantiles, the sensitivity exhibits a non-linear increase and the confidence interval widens. Figure 6 The seasonality of the effect of meteorological bias on runoff quantile bias across different seasons is demonstrated. The seasonal variation of the precipitation bias effect is smaller than that of the potential evapotranspiration bias effect.

[0098] Furthermore, based on the watershed clusters obtained in step 3, each cluster uses the fixed effects model and resampling function from step 4 imported in Python to calculate the sensitivity of runoff quantile deviation to precipitation deviation and potential evapotranspiration deviation for different watershed clusters, resulting in the following: Figure 7 , 8 The diagram shows the sensitivity of runoff quantile deviation to precipitation deviation for each watershed cluster, and the diagram shows the sensitivity of runoff quantile deviation to potential evapotranspiration deviation for each watershed cluster. Figure 7 It can be seen that for low flow rates, clusters 2, 3, 4, 5, and 10 are more sensitive to low flow rates compared to other clusters. For high flow rates, clusters 8, 9, and 10 show greater effects and greater uncertainty; these clusters have flat watershed topography and gentle slopes. Clusters 2 and 8 show significant uncertainty in their effects on low and high flow rates, respectively; both clusters have thick soil layers, high underground porosity, large watershed water storage and release volumes, and are influenced by groundwater. Clusters 6 and 7 show relatively low uncertainty in their sensitivity to precipitation deviations regardless of whether the flow rate is high or low; these clusters have arid watersheds, precipitation is concentrated in winter, slopes are steep, and the uncertainty in land surface runoff generation and confluence processes is low. Cluster 6 is a snow-dominated watershed; precipitation entering the watershed is first covered by snow and snow cover, which buffers the conversion of precipitation into runoff. Figure 8 It can be seen that the median of the potential evapotranspiration bias for clusters 1, 2, 3, 4, and 8 in terms of the low-flow effect is negative, indicating that the potential evapotranspiration bias makes the runoff reanalysis data more accurate in simulating low flow rates. The main reason is... Figure 2The data shows that global runoff reanalysis data is characterized by high flow rates below observed values ​​and low flow rates above observed values. Cluster 4 exhibits negative uncertainty in its effect during high flow rates because its high soil permeability allows groundwater to replenish soil moisture lost through evaporation. Potential evapotranspiration bias in clusters 6, 7, 9, and 10 is near zero in low flow rates. These clusters have low forest cover, high aridity, and smaller maximum leaf area indices than other clusters, indicating minimal seasonal variation in crop transpiration. The uncertainty in the effect of clusters 6 and 7 is near zero for all runoff quantiles. Clusters 2 and 3 show large effects and uncertainties during high flow rates. Their moist watersheds and large maximum leaf area indices, coupled with forest transpiration, increase the sensitivity and uncertainty to bias.

[0099] For panel data formed by resampling the watershed hydrological and meteorological dataset, the distribution of regression coefficients obtained after fixed-effects panel regression can be used to measure the uncertainty of regression coefficients and visualize them as error analysis results.

[0100] Step 5: Diagnose the impact of watershed attribute values.

[0101] Specifically, watershed attributes such as watershed area, average watershed slope, precipitation-to-snow-water ratio, forest cover, and aridity were selected. Stations were categorized into three classes (small, medium, and large) based on the magnitude of these attribute values. By calling the resampling function from step 4, bootstrap resampling was performed on each class to obtain panel data on the magnitude of the watershed attribute values. These data were then subjected to fixed-effects panel regression. The resulting regression coefficient distribution can be used to measure the degree of influence of each watershed attribute value on meteorological bias. Furthermore, the Mann-Whitney U test was performed on the three categories of each attribute value to determine whether statistically significant differences in attribute value magnitude have a significant impact on meteorological bias.

[0102] For example, such as Figure 9 , 10 The diagram illustrates the influence of watershed attributes on precipitation deviation and the influence of watershed attributes on potential evapotranspiration deviation. In this embodiment, the CAMELS large-sample watersheds are categorized into low-attribute-value, medium-attribute-value, and high-attribute-value watersheds based on the magnitude of specific watershed attribute values. Figure 9 This study analyzes the impact of specific watershed attributes on the precipitation bias effect in runoff quantiles for low (Q5), medium (Q50), and high (Q95) flow rates. Overall, watershed attributes influence the magnitude and uncertainty of the precipitation bias effect, and the trend of watershed attribute values ​​affecting the low, medium, and high flow rate effects is consistent. The figures show that larger watershed areas, soil thickness, and groundwater porosity result in greater precipitation bias effects. Larger watershed slopes and higher precipitation-to-snow water ratios result in smaller effects and lower uncertainty. Furthermore, Figure 9 The data also shows that the impact of watershed attributes on low flow rates is smaller compared to high flow rates. The steeper the slope, the smaller the precipitation effect and the lower the uncertainty, especially for high flow rates, because the surface runoff confluence time is shorter. Precipitation seasonality also demonstrates that the greater the seasonality, the greater the bias effect and uncertainty; the more concentrated the precipitation, the greater the deviation in runoff generation and confluence patterns and the process described by hydrological models. Soil depth reflects the water storage capacity of the soil; for low flow rates, the attribute values ​​do not change with soil depth at medium levels, but for high flow rates, they continue to change with depth.

[0103] and Figure 10 The study shows inconsistent trends in the influence of watershed attribute values ​​on potential evapotranspiration at low, medium, and high flow rates. For low flow rates, when aridity is low or the forest coverage and maximum leaf area index (MLA) ratios are medium or high, the potential evapotranspiration deviation effect is negative. For medium flow rates, a higher forest coverage value has a greater impact on the potential evapotranspiration deviation effect. Greater soil thickness and permeability result in a smaller potential evapotranspiration deviation. The effect of the high MLA ratio in the watershed is negative. The trend of the forest coverage attribute's influence on the potential evapotranspiration deviation effect reveals that the effect increases with increasing attribute value for medium and high flow rates, while the effect decreases for low flow rates. The effect of forest coverage on low and high flow rates differs; low flow rates occur in autumn, and the leaf area index of forests varies greatly in autumn, thus increasing the uncertainty of watershed evaporation variation. Higher groundwater permeability reduces the effect and uncertainty on medium and high flow rates, as higher permeability allows groundwater to replenish evaporation, thus reducing the potential evapotranspiration deviation effect.

[0104] Alternatively, the above steps can be encapsulated into a function using the `def()` method, and the code saved as a `.py` file. When the class function needs to be called, it can be accessed using the `import` function.

[0105] Example 3

[0106] This embodiment applies the runoff reanalysis data error analysis method proposed in Embodiment 1, and proposes a runoff reanalysis data error analysis system based on panel regression analysis, such as... Figure 11 The diagram shown is an architecture diagram of the runoff reanalysis data error analysis system in this embodiment.

[0107] The runoff reanalysis data error analysis system proposed in this embodiment includes:

[0108] The data processing module is used to collect watershed hydrological and meteorological datasets, watershed runoff reanalysis data, and meteorological reanalysis data, and to extract watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data.

[0109] The first sampling module is used to cluster all watersheds based on the watershed station attribute data to obtain the spatial clusters of the watersheds and the attribute features of each cluster, and to resample each cluster using the bootstrap method to form the first panel data.

[0110] The fixed-effects panel regression module is equipped with a fixed-effects panel regression model for meteorological input bias and runoff simulation bias. It is used to perform fixed-effects panel regression on panel data, obtain the corresponding regression coefficients, and generate the first coefficient distribution to measure the spatial heterogeneity effect of meteorological bias.

[0111] Further, optionally, the system further includes:

[0112] The second sampling module is used to resample the watershed hydrological and meteorological dataset using a bootstrap method to form the second panel data.

[0113] The fixed-effects panel regression module is also used to perform fixed-effects panel regression on the second panel data to obtain the corresponding regression coefficients and generate a second coefficient distribution to measure the uncertainty of the regression coefficients.

[0114] Further, optionally, the system further includes:

[0115] The third sampling module is used to divide all watersheds into three categories according to the attribute values ​​of each watershed attribute; and to resample each watershed attribute using a bootstrap method to form the third panel data.

[0116] The fixed-effects panel regression module is also used to perform fixed-effects panel regression on the third panel data to obtain the corresponding regression coefficients and generate a third coefficient distribution to measure the degree of influence of each watershed attribute value on the meteorological deviation effect.

[0117] It is understood that the system in this embodiment corresponds to the method in Embodiment 1 above, and the options in Embodiment 1 above are also applicable to this embodiment, so they will not be described again here.

[0118] Example 4

[0119] This embodiment proposes a computer device, including a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor performs all or part of the steps of the runoff reanalysis data error analysis method based on panel regression analysis as proposed in Embodiment 1.

[0120] Example 5

[0121] This embodiment proposes a storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, all or part of the steps of the runoff reanalysis data error analysis method based on panel regression analysis proposed in Embodiment 1 are implemented.

[0122] By way of example, the storage medium includes, but is not limited to, USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks or optical disks, and other media capable of storing program code.

[0123] By way of example, the instructions, programs, code sets, or instruction sets may be implemented using conventional programming languages.

[0124] By way of example, the processor includes, but is not limited to, smartphones, personal computers, servers, network devices, etc., for performing all or part of the steps of the runoff reanalysis data error analysis method based on panel regression analysis described in Example 1.

[0125] The various embodiments in this invention are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely exemplary. The modules described as separate components may or may not be physically separate. When implementing the present invention, the functions of each module can be implemented in one or more software and / or hardware. Alternatively, some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0126] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for error analysis of runoff reanalysis data based on panel regression analysis, characterized in that, Includes the following steps: Collect watershed hydrological and meteorological datasets, watershed runoff reanalysis data, and meteorological reanalysis data, and extract watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data; Based on the deviation between reanalysis data and observed values, a fixed-effects panel regression model is constructed to compare meteorological input deviation with runoff simulation deviation. The fixed-effects panel regression model for constructing meteorological input bias and runoff simulation bias based on the bias between reanalysis data and observed values ​​includes the following steps: Construct a panel dataset; the panel dataset includes the annual runoff quantile deviation Q for each watershed. bias The deviation from annual-scale meteorological data; its expression is: Among them, P bias P represents the deviation of annual precipitation in the basin. reanalysis P represents annual precipitation in the time series data of watershed runoff reanalysis. obs This represents annual precipitation in time-series data of watershed observations; PET bias Indicates the annual average potential evapotranspiration deviation of the watershed, PET renalysis PET represents the annual average potential evapotranspiration under watershed runoff reanalysis time series data. obs T represents the annual average potential evapotranspiration under time series data of watershed observation data; bias T represents the annual average temperature deviation of the watershed. reanalysis T represents the annual average temperature in the watershed runoff reanalysis time series data. obs Q represents the annual average temperature under time series data of watershed observation data; bias Q represents the annual runoff quantile deviation of the watershed. reanalysis Q represents the runoff quantile in the watershed runoff reanalysis time series data. obs This represents the runoff quantile in the time series data of watershed observation data; Based on the aforementioned panel dataset, a fixed-effects panel regression model is constructed to compare meteorological input bias with runoff simulation bias; its expression is as follows: in, P represents the systematic bias of the runoff percentile q in watershed i over time period t. i,t PET represents the precipitation deviation in watershed i over time period t. i,t T represents the average potential evapotranspiration deviation of watershed i over time period t. i,t β1 represents the average temperature deviation in watershed i over time period t; β2 and β1 are the regression coefficients of precipitation deviation and potential evapotranspiration deviation, respectively; β3 is the regression coefficient of the coupling effect between temperature and precipitation deviation; μ i For the fixed effect of watershed i, γ t For time-fixed effects over time period t, ε i,t Let be the error term for watershed i during time period t; All watersheds are clustered based on the watershed station attribute data to obtain the spatial clusters of the watersheds and the attribute characteristics of each cluster. Each cluster is then resampled using a bootstrap method to form the first panel data. Fixed-effects panel regression is performed on the first panel data to obtain the corresponding regression coefficients, generating the first coefficient distribution used to measure the spatial heterogeneity effect of meteorological deviations.

2. The method for error analysis of runoff reanalysis data according to claim 1, characterized in that, The extraction of watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data includes the following steps: Based on the latitude and longitude attributes of the hydrological stations at the watershed outlet, watershed runoff reanalysis time series data are extracted from the watershed runoff reanalysis data. The raster data in the meteorological reanalysis data is cropped according to the vector boundary of the watershed; For the cropped raster data, calculate the actual area of ​​each raster within the watershed; The cropped raster data is traversed to extract the meteorological value of each pixel and its corresponding area weight; A time series of watershed meteorological reanalysis data is generated by calculating the area-weighted average based on the meteorological value and area weight of each pixel.

3. The method for error analysis of runoff reanalysis data according to claim 1, characterized in that, The process of resampling each cluster using a bootstrap method to form the first panel data includes the following steps: The watersheds are classified using the self-organizing map clustering method based on their watershed attribute values, resulting in spatial clustering results and attribute characteristics of each cluster. For each cluster, watershed subsets are sampled with replacement, where the number of samples drawn each time is the number of watersheds in that cluster. This constructs a panel dataset of watershed subsets, resulting in the first panel dataset.

4. The method for error analysis of runoff reanalysis data according to any one of claims 1 to 3, characterized in that, The method further includes: The watershed hydrological and meteorological dataset is resampled using the bootstrap method to form a second panel data. Fixed-effects panel regression is then performed on the second panel data to obtain the corresponding regression coefficients, generating a second coefficient distribution to measure the uncertainty of the regression coefficients.

5. The method for error analysis of runoff reanalysis data according to any one of claims 1 to 3, characterized in that, The method further includes: All watersheds are divided into three categories based on the magnitude of the attribute values ​​of each watershed attribute; The attributes of each watershed are resampled using the bootstrap method to form a third panel data. Fixed-effects panel regression is then performed on the third panel data to obtain the corresponding regression coefficients, generating a third coefficient distribution to measure the degree of influence of each watershed attribute value on the meteorological deviation effect.

6. The method for error analysis of runoff reanalysis data according to claim 5, characterized in that, The method further includes: Mann-Whitney U tests were performed on the three categories of each watershed attribute to obtain the significance analysis results of the influence of watershed attribute values ​​on meteorological deviation effects.

7. A system for error analysis of runoff reanalysis data based on panel regression analysis, employing the runoff reanalysis data error analysis method according to any one of claims 1 to 6, characterized in that, include: The data processing module is used to collect watershed hydrological and meteorological datasets, watershed runoff reanalysis data, and meteorological reanalysis data, and to extract watershed runoff reanalysis time series data and watershed meteorological reanalysis time series data. The first sampling module is used to cluster all watersheds based on the watershed station attribute data to obtain the spatial clusters of the watersheds and the attribute features of each cluster, and to resample each cluster using the bootstrap method to form the first panel data. The fixed-effects panel regression module is equipped with a fixed-effects panel regression model for meteorological input bias and runoff simulation bias. It is used to perform fixed-effects panel regression on panel data, obtain the corresponding regression coefficients, and generate the first coefficient distribution to measure the spatial heterogeneity effect of meteorological bias.

8. The runoff reanalysis data error analysis system according to claim 7, characterized in that, The system also includes: The second sampling module is used to resample the watershed hydrological and meteorological dataset using a bootstrap method to form the second panel data. The fixed-effects panel regression module is also used to perform fixed-effects panel regression on the second panel data to obtain the corresponding regression coefficients and generate a second coefficient distribution to measure the uncertainty of the regression coefficients.

9. The runoff reanalysis data error analysis system according to claim 8, characterized in that, The system also includes: The third sampling module is used to divide all watersheds into three categories according to the attribute values ​​of each watershed attribute; and to resample each watershed attribute using a bootstrap method to form the third panel data. The fixed-effects panel regression module is also used to perform fixed-effects panel regression on the third panel data to obtain the corresponding regression coefficients and generate a third coefficient distribution to measure the degree of influence of each watershed attribute value on the meteorological deviation effect.

Citation Information

Patent Citations

  • Physical mechanism and artificial intelligence fused runoff backtracking simulation method and system

    CN117493476A

  • Photovoltaic panel surface temperature prediction method

    CN117747021A