Mixing statistical framework-based evapotranspiration Bayesian model average fusion method

By employing a Bayesian model averaging method with a hybrid statistical framework, the problem of insufficient accuracy in estimating evapotranspiration in high-altitude and complex terrain regions using traditional Bayesian model averaging methods is solved, achieving higher accuracy and stability, especially with a significant improvement in prediction capability at alpine meadow sites.

CN121009508APending Publication Date: 2025-11-25UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511231486.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-31
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Traditional Bayesian model averaging assumes a normal distribution when fusing multi-source evapotranspiration products, which cannot adapt to the heteroscedasticity characteristics in hydrological applications. This results in insufficient accuracy of evapotranspiration estimation in high-altitude and complex terrain areas. Furthermore, the simple averaging method fails to optimize weights according to model applicability, leading to systematic bias and high uncertainty in the fusion results.

Method used

A Bayesian model averaging method with a hybrid statistical framework is adopted. The goodness of fit is evaluated by the Akaike information criterion and the Bayesian information criterion, the optimal statistical distribution is selected, and the weights are calculated by the pseudo-Bayesian model averaging method. The posterior distribution and weights are optimized by combining the NUTS algorithm and leave-one-out cross-validation to generate the final evapotranspiration estimate.

Benefits of technology

It significantly improved the accuracy and stability of evapotranspiration estimation, with the average correlation coefficient of the fusion results increasing by 12.25% and the root mean square error decreasing by 4.25 mm/month. The model's adaptability and uncertainty quantification capabilities were significantly enhanced, especially at alpine meadow sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009508A_ABST
    Figure CN121009508A_ABST
Patent Text Reader

Abstract

The invention discloses an evapotranspiration Bayesian model average fusion method based on a hybrid statistical framework in China, and relates to the technical field of crossing of ecological hydrological remote sensing and Bayesian statistical modeling. According to the method, the four types of ET products are selected to be fused; the core algorithm can avoid structural deviation caused by a single model; the space-time coverage is complete, the space-time coverage covers the research area in 2007-2021, and the data consistency requirement of the fusion model is met; the reliability of a public data set which is widely used at home and abroad is widely verified in multiple independent researches. And the fusion process is realized through weighted average of the prediction probability density function of each model, and finally monthly ET estimation is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of eco-hydrological remote sensing and Bayesian statistical modeling, specifically to a Bayesian model averaging fusion method based on a hybrid statistical framework, which is suitable for high-precision estimation of evapotranspiration in high-altitude and complex terrain areas. Background Technology

[0002] Terrestrial evapotranspiration is a core process of water vapor transport and energy exchange in the Earth's soil-vegetation-atmosphere system, and its accurate estimation is crucial for understanding the impact of the Tibetan Plateau on regional and global climate dynamics. Remote sensing technology, due to its advantages of large-scale, multi-temporal data acquisition, has become an important tool for quantifying regional evapotranspiration. Existing evapotranspiration products, generated based on different model structures and parameterization schemes (such as reanalysis, water balance, or remote sensing algorithms), provide reasonable estimates at local scales. However, at regional and global scales, due to the complexity of driving factors and environmental variables, significant differences and uncertainties exist in the spatial patterns, temporal variations, and evapotranspiration component classifications among the products.

[0003] To integrate multi-source evapotranspiration data, Bayesian model averaging, as a statistical method, effectively utilizes the strengths of different models by dynamically adjusting the weights of each model to fuse the predicted distribution. However, traditional Bayesian model averaging typically assumes that the predicted probability density function follows a normal distribution when fusing evapotranspiration products. In hydrological applications, residual errors often exhibit heteroscedasticity, limiting its accuracy in evapotranspiration estimation in complex terrains and leading to systematic bias and high uncertainty in the fusion results. Furthermore, simple averaging methods assign equal weights to all models, failing to optimize weights based on model applicability; traditional Bayesian model averaging also lacks flexible distribution assumptions and prior parameter initialization mechanisms. These shortcomings make existing technologies challenging in evapotranspiration estimation in high-altitude, complex terrain areas, and unable to effectively integrate multi-source evapotranspiration products to improve overall accuracy. Summary of the Invention

[0004] In existing technologies, traditional Bayesian average models, when fusing multi-source evapotranspiration products, generally assume that the predicted probability density function follows a normal distribution. This makes them unsuitable for the heteroscedasticity of residual errors in hydrological applications, resulting in insufficient accuracy in evapotranspiration estimation in high-altitude and complex terrain regions such as the Qinghai-Tibet Plateau.

[0005] In view of this, the purpose of this invention is to provide a Bayesian model averaging framework with flexible distribution assumptions and prior parameter initialization mechanisms to improve the fusion accuracy and stability of evapotranspiration data. To achieve the above objective, the technical solution adopted by this invention is: an evapotranspiration Bayesian model averaging fusion method based on a hybrid statistical framework, comprising the following steps:

[0006] Step 1: Obtain the dataset for the target area, including evapotranspiration products modeled based on different physical processes and surface observation station data; evapotranspiration products include:

[0007] Step 2: Data preprocessing: For multiple evapotranspiration products, unify them to the target spatiotemporal resolution; for latent heat flux data from surface observation stations, after quality control, convert them into evapotranspiration values ​​according to the following formula, and then aggregate them into monthly data;

[0008]

[0009] Where ET is the evapotranspiration value, LE represents the latent heat flux, and Ta represents the air temperature;

[0010] Step 3: Define the candidate distribution set: Based on the hydrological error characteristics, normal Gaussian, gamma, lognormal, Beta, Weibull and Gaussian mixture models are selected as candidate distributions;

[0011] Step 4: Denote each evaporation product as model M. i Then, the optimal statistical distribution is independently determined from the distributions in step 3. The goodness of fit is evaluated using the Akaike information criterion and the Bayesian information criterion, and the distribution with the smallest information criterion is selected as model M. i The optimal statistical distribution D i ;

[0012] Step 5: For each model M i From its optimal distribution D i Extracting the parameter vector θ i As model M i Step 6: Using the optimal parameters as prior knowledge, explicitly constrain the prediction process of the Bayesian model averaging; the prediction probability density function of the Bayesian model averaging is the weighted average of the prediction probability density functions of each model, and its specific expression is:

[0013]

[0014] Where p(y|M1,M2,…,M) N Let y represent the weighted prediction probability density function, and y represent the ET value to be estimated. T D represents the ET value calculated in step 2. i (y;θ i Let p(M) be the probability density function of the optimal distribution Di of model Mi at point y. i |y T ) indicates that in the measured data y T When model Mi is the weight w of the "optimal model", i The probability density of point prediction is calculated based on the expected logarithm using the pseudo-Bayesian model averaging method.

[0015] Step 7: Weight and fuse the predicted probability density functions of candidate evapotranspiration products to generate the final monthly evapotranspiration estimate.

[0016] Furthermore, the calculation methods for the Akaike information criterion and the Bayesian information criterion in step 4 are as follows:

[0017] AIC = 2k - 2ln(L)

[0018] BIC = kln(n) - 2ln(L)

[0019] Where AIC represents the goodness-of-fit evaluation of the Akaike Information Criterion, BIC represents the goodness-of-fit evaluation of the Bayesian Information Criterion, k represents the number of distribution parameters, L is the maximum value of the likelihood function of the corresponding model, and n is the sample size.

[0020] Furthermore, the method for calculating the fusion weight of each evaporation product in step 7 is as follows:

[0021] Computational model Mi for y j ELPD's predictive ability:

[0022]

[0023] in, For M i For y j The predicted probability density function is implemented using leave-one-out cross-validation (i.e., each time a data point y is...). j Excluding those, use the remaining data y -j Train model Mi, and then use the trained model to predict the excluded points y. j (To calculate its probability density); the larger the ELPD value, the stronger the predictive ability of model Mi. Normalize the ELPD value to obtain the model weights w. i ,Right now

[0024]

[0025] Among them, delp i =max j (ELPD j )-ELPD i .

[0026] Compared with existing technologies (single evapotranspiration product, Bayesian model averaging with normal distribution assumption, and simple averaging method SA), it shows significant advantages in evapotranspiration estimation accuracy, model adaptability, and uncertainty quantification, specifically:

[0027] Overall performance is superior to individual products: the average correlation coefficient (r) of the Hybrid-BMA fusion results is 12.25% higher than that of individual products such as ERA5-Land, GLEAM, PML-v2, and TerraClimate, and the average root mean square error (RMSE) is reduced by 4.25 mm / month (see...). Figure 3 (a)-(d)).

[0028] Stability is superior to the comparison methods: the fusion result of Hybrid-BMA closely matches the fitted line (red) of the measured data with the 1:1 reference line (gray), indicating that its predicted values ​​deviate less from the true values; while the SA method (simple averaging) suffers from persistent underestimation. Although Normal-BMA (Bayesian model averaging with normal distribution assumption) is better than SA, Hybrid-BMA further optimizes RMSE (only 0.25 mm / month lower) and MAE (only 0.35 mm / month lower), demonstrating the flexible adaptability of the hybrid distribution to different residual patterns (see...). Figure 3 (e)-(g)).

[0029] Enhanced adaptability to different land cover types: Taking the Qinghai-Tibet Plateau as an example, its land cover types are complex (mainly desert and alpine meadow). Hybrid-BMA performed particularly well at alpine meadow sites, with a significantly higher average correlation coefficient (r = 0.83) compared to single products (such as ERA5-Land 0.72 and GLEAM 0.75), and a lower standard deviation (26.05 mm / month) than SA (28.12 mm / month) and Normal-BMA (26.87 mm / month), indicating a stronger stable predictive ability for grassland evapotranspiration (see...). Figure 4 ). Attached Figure Description

[0030] Figure 1 This is a technical approach for the average fusion method of evapotranspiration Bayesian models based on a hybrid statistical framework.

[0031] Figure 2 The study area is defined by its geographical location, IGBP land cover type, and flux observation station locations. Land cover includes: alpine grassland (GRA), desert (BSV), evergreen coniferous forest (ENF), evergreen broad-leaved forest (EBF), deciduous broad-leaved forest (DBF), mixed forest (MF), closed shrubland (CSH), open shrubland (OSH), woody savanna (WSA), savanna (SAV), permanent wetland (WET), cultivated land (CRO), urban and built land (URB), cultivated land / natural vegetation mosaic (CVM), snow and ice (SNO), and water bodies (WAT).

[0032] Figure 3The quantitative evaluation graphs for the evapotranspiration estimation results mainly include a scatter regression plot (ag) of the estimated values ​​and observed values ​​and a posterior probability density function curve (h).

[0033] Figure 4 The results are specific to the main land cover types (desert and alpine meadow) in the study area.

[0034] Figure 5 This serves to verify the spatial consistency between single evapotranspiration products and fusion results in the study area. Detailed Implementation

[0035] The main reasons for selecting these four types of ET products for fusion in this invention are: (1) their core algorithm principles are different, which can avoid the structural bias caused by a single model; (2) their spatiotemporal coverage is complete, all covering the study area from 2007 to 2021, meeting the data consistency requirements of the fusion model; and (3) they are all publicly available datasets widely used at home and abroad, and their reliability has been extensively verified in many independent studies. The fusion process is achieved by weighted averaging of the prediction probability density functions of each model, and finally generating monthly ET estimates. The following describes in detail the implementation steps of the Hybrid-BMA method proposed in this invention, using a specific implementation example of the 0.1° resolution monthly evapotranspiration estimation of the Qinghai-Tibet Plateau from 2007 to 2021, and compares it with the specific implementation example. Figure 1 Explain the technical details of each step.

[0036] Step 1: Data Collection and Preprocessing

[0037] (1) Multi-source evaporation products;

[0038] The European Centre for Medium-Range Weather Forecasts (ECMWF) Generation 5 Land Surface Reanalysis Dataset (ERA5-Land), the Global Land Surface Evapotranspiration Amsterdam Model (GLEAM), the Penman-Montes-Lunin Version 2 Evapotranspiration Model (PML-v2), and the High-Resolution Land Climate and Hydrology Dataset (TerraClimate);

[0039] Four widely used evapotranspiration products based on different physical process algorithms were selected (Table 1), and their spatiotemporal resolution was standardized to 0.1° / month (2007-2021). For the PML-v2 product, since its original resolution is 500m / 8 days, this invention uses a (time) weighted summation and (spatial) arithmetic average method to aggregate it to the target spatiotemporal resolution. The TerraClimate product (4km / month) underwent the same spatial resolution processing.

[0040] Table 1. Example Data - Evaporation Product Introduction

[0041]

[0042] (2) Measured evapotranspiration data:

[0043] Downloaded from the National Tibetan Plateau Scientific Data Center (https: / / data.tpdc.ac.cn), a total of 12 eddy covariance flux towers were obtained. For example... Figure 2 As shown, the stations cover typical ecological areas of the Qinghai-Tibet Plateau (alpine meadows, deserts, etc.), and specific information is shown in Table 2 (station name, latitude and longitude, IGBP land cover type, and data availability time range).

[0044] Table 2. Example Data - Introduction to Eddywise Correlation Flux Station Data

[0045]

[0046] Invalid data, including erroneous data (QC=2) and missing data (QC=8), are excluded based on the quality control codes provided by the data center. Based on the principle of energy conservation, the latent heat flux in the eddy covariance flux data is converted into evapotranspiration according to formula (1).

[0047] To match the monthly scale resolution of evapotranspiration products, hourly flux data needs to be aggregated step by step into monthly evapotranspiration data:

[0048] For any day t, count the number of valid hourly data points actually recorded that day, denoted as N (N≤24, due to potential missing data caused by instrument malfunctions, data transmission interruptions, etc.). If N≤12 for a single day (i.e., at least 12 hours of valid data), then discard the data for that day. Sum the daily evapotranspiration and amplify it to the theoretical total over 24 hours using a standardization factor of 24 / N, i.e.

[0049]

[0050] Daily evapotranspiration is consolidated into monthly evapotranspiration. Similarly, for any month T, the number of valid days within that month is counted and denoted as M. If M ≤ 20 for a single month (i.e., at least 20 days of valid data), the data for that month is discarded. After summing the evapotranspiration data for valid days within the month, it is amplified to the theoretical total for the entire month using a standardization factor M0 / M (the total number of days in the month is denoted as M0).

[0051]

[0052] Step 2: Constructing the averaging framework of a Bayesian model based on mixed statistical distributions

[0053] The construction process of the average framework of the Bayesian model with mixed statistical distribution hypothesis proposed in this invention includes four core steps: model selection of mixed distribution hypothesis, initialization of Bayesian model average framework, posterior distribution sampling, and model weight calculation. The following operations are all completed in the Python 3.8 environment, and the dependent libraries include PyMC (Bayesian inference) and ArviZ (model diagnostics). (1) Model selection of mixed statistical distribution hypothesis

[0054] Define the candidate distribution set: Based on the skewness and heavy-tailed characteristics of hydrological errors, the following six statistical distributions are selected as candidate sets, including Gaussian distribution, gamma distribution, log-normal distribution, Beta distribution, Weibull distribution and Gaussian mixture model.

[0055] Model Fitting and Optimal Distribution Selection: For each evapotranspiration product Mi, calculate the residual between the predicted Mi value and the measured evapotranspiration data. Fit all distributions in the candidate set to the residuals and calculate the parameters of each distribution. Evaluate the fit of each distribution using the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), and select the distribution with the smallest information criterion as the optimal distribution Di for Mi. The calculation formula is:

[0056]

[0057] Where k is the number of distribution parameters, L is the maximum value of the likelihood function, and n is the sample size.

[0058] (2) Initialization of the average frame of the Bayesian model

[0059] This invention will determine the optimal distribution parameter θ for each Mi. opt,i As prior knowledge, the prediction process of the Bayesian model's average frame is explicitly constrained. The prediction probability density function of model Mi is determined by its optimal distribution D. opt,i Give it directly, that is

[0060] p(y|M i ) = D opt,i (y;θ opt,i )

[0061] Among them, D opt,i (y;θ opt,i ) is the probability density function of the optimal distribution of Mi.

[0062] The average prediction probability density function of the Bayesian model is a weighted average of the prediction probability density functions of each model, with weights w. i Based on the posterior model probability p(M) i |y T The decision was made by averaging the pseudo-Bayesian model.

[0063] (3) Posterior distribution sampling and weight calculation

[0064] This invention employs the No-U-Turn Sampler (NUTS) algorithm and Leave-One-Out Cross-Validation (LOO-CV) to optimize the posterior distribution and weights:

[0065] The NUTS algorithm is an adaptive Markov chain Monte Carlo (MCMC) method that efficiently samples high-dimensional posterior distributions by automatically adjusting the step size and number of iterations. This invention uses the NUTS implementation from the PyMC library, with specific parameters set as follows: a total of 2000 iterations (1000 during the warm-up period and 1000 during the sampling period), 4 independent chains, and a target acceptance rate of 0.95.

[0066] To avoid the high complexity of traditional Bayesian posterior calculations, this invention employs a pseudo-Bayesian model averaging method. It calculates the expected log-wise pointwise prediction density (ELPD) of each model using LOO-CV and normalizes it to the weights w. i LOO-CV excludes each observation y j Refit model Mi to the remaining data y -j And assess Mi's impact on y j The predictive power of ELPD. The formula for calculating ELPD is:

[0067]

[0068] in For M i For y j The predicted probability density function; the larger the ELPD value, the stronger the predictive ability of the model Mi. The model weights w are obtained by normalizing the ELPD value. i ,Right now

[0069]

[0070] Among them, delp i =max j (ELPD j )-ELPD i (Ensure that the weights are non-negative and sum to 1).

[0071] Step 3: Framework Validation and Output

[0072] This step details the output verification process of the invention framework, including three core components: visualization verification of model prediction results, quantitative evaluation of prediction accuracy, and land cover type specificity analysis. All operations below are based on a Python 3.8 environment, with dependencies including Matplotlib, Seaborn, and SciPy.

[0073] Visualization of the posterior probability density function curve (corresponding) Figure 3(h): Using the fusion prediction probability density function output by the invention framework (obtained by weighted average of the optimal distributions of each model), generate probability density curves for individual evaporation products and fusion results, and calculate key statistics (mean).

[0074] Prediction accuracy quantitative evaluation: This invention uses point estimation accuracy index and spatial consistency verification, and conducts a horizontal comparison with single product and comparative methods (simple average (SA), Normal-BMA based on normal distribution assumption) (corresponding to...). Figure 3 (ag)). Point estimation accuracy metrics include correlation coefficient (r), root mean square error (RMSE), and mean absolute error (MAE). Spatial consistency verification is performed by calculating the difference between the fusion result and the monthly average evapotranspiration values ​​of each evapotranspiration product and then spatially visualizing the result (corresponding to...). Figure 5 ),Right now

[0075] ΔET=ET Hybrid -ET product

[0076] Land cover type specificity analysis: The surface cover types on the Qinghai-Tibet Plateau are complex (mainly desert BSV and alpine meadow GRA), and the driving mechanisms of ET (Effective Transformation) differ significantly among different types. This invention targets these two dominant land cover types and further validates the adaptability of the framework using Taylor diagrams. Taylor diagrams comprehensively evaluate the consistency between model predictions and measured data using three indicators: standardized root mean square error, correlation coefficient, and standard deviation. Figure 4 ).

Claims

1. A method for averaging and fusion of evapotranspiration Bayesian models based on a hybrid statistical framework, comprising the following steps: Step 1: Obtain the dataset for the target area, including evapotranspiration products modeled based on different physical processes and data from surface observation stations; Evaporation products include: Step 2: Data preprocessing: For multiple evapotranspiration products, unify them to the target spatiotemporal resolution; For latent heat flux data from surface observation stations, after quality control, the data is converted into evapotranspiration values ​​according to the following formula, and then aggregated into monthly data. Where ET is the evapotranspiration value, LE represents the latent heat flux, and Ta represents the air temperature; Step 3: Define the candidate distribution set: Based on the hydrological error characteristics, normal Gaussian, gamma, lognormal, Beta, Weibull and Gaussian mixture models are selected as candidate distributions; Step 4: Denote each evaporation product as model M. i Then, the optimal statistical distribution is independently determined from the distributions in step 3. The goodness of fit is evaluated using the Akaike information criterion and the Bayesian information criterion, and the distribution with the smallest information criterion is selected as model M. i The optimal statistical distribution D i ; Step 5: For each model M i From its optimal distribution D i Extracting the parameter vector θ i As model M i The optimal parameters; Step 6: Using the optimal parameters as prior knowledge, explicitly constrain the prediction process of the Bayesian model averaging; the prediction probability density function of the Bayesian model averaging is the weighted average of the prediction probability density functions of each model. Step 7: Weight and fuse the predicted probability density functions of candidate evapotranspiration products to generate the final monthly evapotranspiration estimate.

2. The averaging fusion method for evapotranspiration Bayesian models based on a hybrid statistical framework as described in claim 1, characterized in that, The calculation methods for the goodness-of-fit evaluation using the Akaike information criterion and the Bayesian information criterion in step 4 are as follows: AIC = 2k - 2ln(L) BIC = kln(n) - 2ln(L) Where AIC represents the goodness-of-fit evaluation of the Akaike Information Criterion, BIC represents the goodness-of-fit evaluation of the Bayesian Information Criterion, k represents the number of distribution parameters, L is the maximum value of the likelihood function of the corresponding model, and n is the sample size.

3. The averaging fusion method for evapotranspiration Bayesian models based on a hybrid statistical framework as described in claim 1, characterized in that, The method for calculating the fusion weight of each evaporation product in step 7 of the evaluation is as follows: Computational model Mi for y j ELPD's predictive ability: in, For M i For y j The predicted probability density function is implemented using leave-one-out cross-validation (i.e., each time a data point y is...). j Excluding those, use the remaining data y -j Train model Mi, and then use the trained model to predict the excluded points y. j (To calculate its probability density); the larger the ELPD value, the stronger the predictive ability of model Mi. Normalize the ELPD value to obtain the model weights w. i ,Right now Among them, delp i =max j (ELPD j )-ELPD i .

4. The averaging fusion method for evapotranspiration Bayesian models based on a hybrid statistical framework as described in claim 1, characterized in that, The weighted average formula in step 6 is: Where p(y|M1,M2,…,M) N Let y represent the weighted prediction probability density function, and y represent the ET value to be estimated. T D represents the ET value calculated in step 2. i (y;θ i Let p(M) be the probability density function of the optimal distribution Di of model Mi at point y. i |y T ) indicates that in the measured data y T When model Mi is the weight w of the "optimal model", i The probability density of point prediction is calculated based on the expected logarithmic point prediction using the pseudo-Bayesian model averaging method.