Underground water multi-model average weight calculation method based on Bayesian Stacking

By using the Bayesian Stacking method, leave-one-out cross-validation and nonlinear optimization algorithm to calculate the optimal weights of multiple groundwater models, the problems of long calculation time and poor stability in existing technologies are solved, and more efficient and accurate multi-model average weight calculation is achieved, thereby improving the accuracy of groundwater simulation.

CN120724892APending Publication Date: 2025-09-30NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510823187.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

The existing multi-model weight calculation method is time-consuming, has poor result stability, and the model prior weight setting is subjective, which limits its application in groundwater resource management and pollution prevention.

Method used

The Bayesian Stacking method is used to calculate the optimal weight of each numerical model through leave-one-out cross-validation and nonlinear optimization algorithm, thereby improving the calculation accuracy and efficiency of the average weight of multiple models.

Benefits of technology

The calculation accuracy and efficiency of multi-model average weights are improved, the accuracy and reliability of groundwater simulation are enhanced, and the calculation cost is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724892A_ABST
    Figure CN120724892A_ABST
Patent Text Reader

Abstract

The invention discloses a groundwater multi-model average weight calculation method based on Bayes ian Stacking (Bayes ian Stacking). According to the method, a multi-structure model set M is constructed, and one-leaving cross validation probability distribution of each observation point in each numerical simulation model is calculated. Then, based on the one-leaving cross validation probability distribution of each observation point in each numerical simulation model, determining a one-leaving method mean square error corresponding to the corresponding numerical simulation model, based on a constraint condition, minimizing the one-leaving method mean square error by solving a nonlinear optimization problem, and obtaining an optimal weight of each numerical simulation model; and finally, based on the prediction distribution and the weight of each numerical simulation model, obtaining weighted comprehensive prediction distribution of the prediction points. According to the method, the calculation precision and calculation efficiency of the multi-model average weight are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of uncertainty quantification of groundwater numerical simulation, and specifically to a groundwater multi-model average weight calculation method based on Bayesian Stacking. Background Art

[0002] The application of numerical models in the field of groundwater environments has received widespread attention, providing an important scientific basis for the sustainable development, utilization, management, and pollution prevention of groundwater resources. Numerical models are generalizations and mathematical abstractions of actual groundwater systems, which are highly complex. Consequently, uncertainties in model parameters and structures often exist, directly impacting the accuracy and reliability of groundwater simulations.

[0003] Currently, research on model parameter uncertainty analysis is relatively extensive and in-depth, while research and analysis on model structural uncertainty is relatively lagging behind. Using multi-model simulation methods to construct a collection of multiple groundwater numerical models based on different numerical simulation theories or assumptions is an important strategy for reducing model structural uncertainty. The essence of multi-model simulation methods is to reduce uncertainty and improve simulation accuracy by integrating the simulation prediction results of multiple numerical models. Currently, the main methods for calculating multi-model weights include Bayesian Model Averaging (BMA) and Bayesian Model Selection (BMS).

[0004] However, these traditional methods still face some challenges in practical applications, such as long calculation time, poor result stability, and subjective setting of model prior weights, which directly limit their application effects. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for calculating the average weight of multiple groundwater models based on Bayesian Stacking. By traversing all observation points i=1,2,...,N for each numerical model, a leave-one-out cross-validation is performed to obtain the leave-one-out cross-validation probability distribution of the i-th observation; combined with scoring rules and constraints, a constrained nonlinear optimization algorithm is used to solve the optimal weight of each numerical model to improve the calculation accuracy and efficiency of the multi-model average weight.

[0006] In order to solve the above technical problems, the present invention provides the following technical solution: a method for calculating the average weight of groundwater multi-models based on Bayesian Stacking, which comprises the following steps:

[0007] Constructing a multi-structure model set M; the multi-structure model set M includes a plurality of numerical simulation models;

[0008] Calculate the leave-one-out cross-validation probability distribution of each observation point in each numerical simulation model;

[0009] Determine the leave-one-out mean square error corresponding to the corresponding numerical simulation model based on the leave-one-out cross-validation probability distribution of each observation point in each numerical simulation model;

[0010] Based on the constraints, the optimal weights of each numerical simulation model are obtained by solving the nonlinear optimization problem to minimize the leave-one-out mean square error.

[0011] Based on the prediction distribution and weights of each numerical simulation model, the weighted comprehensive prediction distribution of the prediction points is obtained.

[0012] According to the above technical solution, the step of calculating the leave-one-out cross-validation probability distribution of each observation point in each numerical simulation model includes:

[0013] Obtain the parameter range and corresponding prior distribution of each numerical simulation model;

[0014] For each numerical simulation model, Markov chain Monte Carlo sampling is performed based on the entire data set to obtain the posterior distribution of each numerical simulation model;

[0015] Based on the posterior distribution of each numerical simulation model, for n observation points, n-1 data points are taken each time to train the model, and the parameter posterior distribution is obtained by the Markov chain Monte Carlo algorithm. According to the parameter posterior distribution, the leave-one-out cross-validation probability distribution of each observation point is obtained by using the leave-one-out cross-validation.

[0016] According to the above technical solution, the constraints are:

[0017]

[0018] Where K represents the number of numerical simulation models in the multi-structure model set, w k Represents the numerical simulation model M k The average weight of .

[0019] According to the above technical solution, the optimal weight w of the numerical simulation model is:

[0020]

[0021] S represents the scoring rule; p k,-i (y i ) represents the numerical simulation model M k In the numerical simulation model, n-1 points other than the i-th point are used to train the numerical simulation model, and the probability distribution of the observation point i is obtained; y i Represents the numerical simulation model M k The corresponding predicted value; n represents the number of observation points; C1k Represents a set, and the weight of the solved numerical simulation model is defined in this set. The model weight is between 0 and 1, and the sum is 1; w is the optimal weight of the numerical simulation model and belongs to C1 k gather.

[0022] in,

[0023] p k,-i (y i )=∫p(y i |θ k ,M k )p(θ k |y -i ,M k )dθ k ;

[0024] In the formula, p(y i |θ k ,M k ) is the numerical simulation model M k and parameter θ k The joint likelihood function, p(θ k |y -i ,M k ) represents the model M k In data y other than point i -i Next, the parameter θ k The predicted distribution of .

[0025] According to the above technical solution, the weighted comprehensive prediction distribution of the prediction points is:

[0026]

[0027] Where w k is the numerical simulation model M k The average weight of p(y -i |y i ,M k ) is the observation point y i The probability density value corresponding to the true value of in the leave-one-out prediction distribution.

[0028] Compared with the prior art, the present invention has the following beneficial effects: the present invention introduces the Bayesian Stacking method, performs leave-one-out cross-validation (LOO) on all observation points of each numerical simulation model, and obtains the leave-one-out cross-validation (LOO) probability distribution of the i-th observation; through the constraint conditions, a constrained nonlinear optimization algorithm is used to solve the optimal weight of each numerical model, thereby improving the calculation accuracy and efficiency of the average weight of multiple models. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0030] Figure 1 It is a flowchart of the steps of the method for calculating the average weight of multiple groundwater models based on Bayesian Stacking of the present invention;

[0031] Figure 2 is the penetration curve of PFOA migration in saturated limestone;

[0032] Figure 3 is the posterior distribution of each model parameter;

[0033] Figure 4 is the marginal likelihood convergence trajectory diagram;

[0034] Figure 5 is a schematic diagram of the BMA simulation results and 95% confidence interval;

[0035] Figure 6 It is a schematic diagram of BS simulation results and 95% confidence interval. DETAILED DESCRIPTION

[0036] The following will combine the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work, 100 mesh stainless steel gauze is laid at both ends to support solid particles and disperse water flow. The inlet end of the column is connected to a peristaltic pump to realize sampling control, and the outlet end of the column is connected to an automatic collector to collect samples. The column is filled by wet method, and 10PV deionized water is injected by peristaltic pump to keep the column stable, followed by injection of 3PV tracer solution, and then 3PV deionized water is injected to flush the column, and the outflow liquid is collected regularly by automatic partial collector. The tracer migration process can be described by the convection diffusion equation, and the formula is as follows:

[0037]

[0038] Where C is the tracer concentration [ML -3 ], t is time [T], D is the hydrodynamic diffusion coefficient [L 2 T -1 ], z is the vertical space coordinate [L], v is the pore water velocity [LT -1 ].

[0039] The column was flushed with 25 PV of PFOA-free background solution to allow the chemical conditions in the column to reach equilibrium. Then, 3 PV of PFOA solution was injected using a peristaltic pump, and the column was immediately flushed with PFOA-free background solution until the effluent was free of PFOA. Figure 2 This is the penetration curve of PFOA in saturated limestone. A variety of equilibrium and non-equilibrium models are used to simulate the migration of PFOA in porous media.

[0040] According to the above process, numerical simulation models M1, M2 and M3 were constructed to describe the migration process of perfluorooctanoic acid in saturated porous media.

[0041] Model 1 M1:

[0042] The two-point kinetic model is used to describe the migration process of PFOA in saturated porous media, and its governing equation is:

[0043]

[0044]

[0045] Where C1 is the concentration of perfluorooctanoic acid in the liquid phase [ML -3 ], t is the time [T], ρ is the density of the porous medium [ML -3 ], θ is the volume saturation [-], S1 is the solid phase adsorption amount of PFOA at adsorption site 1 [MM -1 ], S2 is the solid phase adsorption amount of perfluorooctanoic acid at adsorption site 2 [MM -1 ], D is the hydrodynamic diffusion coefficient [L 2 T -1 ], z is the vertical space coordinate [L], v is the pore water velocity [LT -1 ], k1 is the first-order adsorption coefficient at adsorption site 1 [T -1 ], k2 is the first-order adsorption coefficient at adsorption site 2 [T -1 ], k d2 is the first-order desorption coefficient at adsorption site 2 [T -1 ].

[0046] Model 2 M2:

[0047] The linear two-point reversible adsorption model is used to describe the migration process of PFOA in saturated porous media. The governing equation is:

[0048]

[0049] S=K d C1

[0050] Where f is the proportion of equilibrium adsorption sites [-], K dis the solid phase adsorption coefficient [M -1 L 3 ], S is the solid phase adsorption capacity [MM -1 ], α represents the first-order kinetic adsorption coefficient [T -1 ].

[0051] Model 3 M3:

[0052] The Freundlich two-point reversible adsorption model is used to describe the migration process of PFOA in saturated porous media. The governing equation is:

[0053]

[0054] S=K d C1 β

[0055] Where β is the adsorption isotherm coefficient [-].

[0056] S2. Parameter uncertainty analysis was performed for numerical simulation models M1, M2, and M3. Three Markov chains were set in each Markov chain simulation, with the chain lengths for both the warm-up and the formal phases set to 1000. The 98 observation points were divided into two parts: the first 40 were used for model identification, and the last 58 were used for model validation. Figure 3 Obtain the posterior distribution of each model parameter for Markov chain identification. As can be seen, each model parameter has a peak. Therefore, a multi-model averaging method is selected to improve simulation accuracy.

[0057] S3. First, use the AME method to calculate the edge likelihood value, such as Figure 4 As shown in , when the number of samples reaches 200,000 (i.e., the model runs 200,000 times), the marginal likelihood value converges to a stable state. The marginal likelihood values ​​of each model are normalized to obtain the BMA weight of each model. According to the BMA weight of each model, the prediction results are obtained as follows Figure 5 (BMA) shown.

[0058] S4. Calculate the weight of each model using the groundwater multi-model average weight calculation method (BS) based on Bayesian Stacking. The specific steps are as follows:

[0059] S401, calculate the leave-one-out cross-validation probability distribution of the true value of each observation point in each numerical simulation model

[0060] S402, determining the leave-one-out cross-validation mean square error corresponding to the corresponding numerical simulation model based on the leave-one-out cross-validation probability distribution of each observation point in each numerical simulation model, specifically the following steps:

[0061] Obtain the parameter range and corresponding prior distribution of each numerical simulation model;

[0062] For each numerical simulation model, Markov chain Monte Carlo sampling is performed based on the entire data set to obtain the posterior distribution of each numerical simulation model;

[0063] Based on the posterior distribution of each numerical simulation model, for n observation points, n-1 data points are taken each time to train the numerical simulation model, and the parameter posterior distribution is obtained by the Markov chain Monte Carlo algorithm. According to the parameter posterior distribution, the leave-one-out cross-validation probability distribution of each observation point is obtained by using the leave-one-out cross-validation.

[0064] S403, based on constraints By solving the nonlinear optimization problem to minimize the leave-one-out mean square error, the optimal weight w of each numerical simulation model is obtained;

[0065]

[0066] Where K represents the number of numerical simulation models in the multi-structure model set, w k Represents the numerical simulation model M k The average weight of . S represents the scoring rule; the set C1 k Represents a set, and the weight of the solved numerical simulation model is defined in this set. The model weight is between 0 and 1, and the sum is 1; w is the optimal weight of the numerical simulation model and belongs to C1 k Collection.p k,-i (y i ) represents the numerical simulation model M k In the numerical simulation model, n-1 points other than the i-th point are used to train the numerical simulation model, and the probability distribution of the observation point i is obtained; y i Represents the numerical simulation model M k The corresponding predicted value, n represents the number of observation points.

[0067] S404, based on the prediction distribution and weight of each numerical simulation model, obtain the weighted comprehensive prediction distribution of the prediction point, and obtain the weighted comprehensive prediction distribution of the prediction point as follows: Figure 6 (BS) shown.

[0068] Among them, the weighted comprehensive prediction distribution of the prediction points is:

[0069]

[0070] Where w k is the numerical simulation model M k The average weight of p(y -i |y i ,M k ) is the observation point y iThe probability density value corresponding to the true value of in the leave-one-out prediction distribution.

[0071] Among them, when using leave-one-out cross-validation to obtain BS weights, the model needs to be run 120,000 times (number of observation points in the identification period * number of Markov chains * chain length), which improves the calculation efficiency.

[0072] The weight calculation results of the two multi-model methods in steps S3 and S4 are shown in Table 1. It can be seen that the order of weights assigned by the two methods is: M1>M2>M3.

[0073] Table 1 Comparison of weights between traditional multi-model averaging method and Bayesian Stacking (BS) model

[0074]

[0075]

[0076] At the same time from Figure 5 (BMA) Figure 6 (BS) It can be seen that both methods improve the simulation accuracy of PFOA migration in saturated porous media. The prediction evaluation indicators of the methods are further calculated, as shown in Table 2.

[0077] Table 2. Prediction indicators of traditional multi-model averaging method and Bayesian Stacking (BS) simulation

[0078]

[0079] And from Table 2, it can be found that the three simulation indicators of the BS comprehensive prediction results are better than those of BMA, which shows that the BS method can better handle cases and improve both simulation accuracy and efficiency.

[0080] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0081] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. The groundwater multi-model average weight calculation method based on Bayesian Stacking is characterized by: The steps include: Constructing a multi-structure model set M; the multi-structure model set M includes a plurality of numerical simulation models; Calculate the leave-one-out cross-validation probability distribution of each observation point in each numerical simulation model; Determine the leave-one-out mean square error corresponding to the corresponding numerical simulation model based on the leave-one-out cross-validation probability distribution of each observation point in each numerical simulation model; Based on the constraints, the optimal weights of each numerical simulation model are obtained by solving the nonlinear optimization problem to minimize the leave-one-out mean square error. Based on the prediction distribution and weights of each numerical simulation model, the weighted comprehensive prediction distribution of the prediction points is obtained.

2. The method for calculating groundwater multi-model average weights based on Bayesian Stacking according to claim 1, characterized in that: The step of calculating the leave-one-out cross-validation probability distribution of each observation point in each of the numerical simulation models comprises: Obtain the parameter range and corresponding prior distribution of each numerical simulation model; For each numerical simulation model, Markov chain Monte Carlo sampling is performed based on the entire data set to obtain the posterior distribution of each numerical simulation model; Based on the posterior distribution of each numerical simulation model, for n observation points, n-1 data points are taken each time to train the numerical simulation model, and the parameter posterior distribution is obtained by the Markov chain Monte Carlo algorithm. According to the parameter posterior distribution, the leave-one-out cross-validation probability distribution of each observation point is obtained by using the leave-one-out cross-validation.

3. The method for calculating groundwater multi-model average weights based on Bayesian Stacking according to claim 1, characterized in that: The constraints are: Where K represents the number of numerical simulation models in the multi-structure model set, w k Represents the numerical simulation model M k The average weight of .

4. The method for calculating groundwater multi-model average weights based on Bayesian Stacking according to claim 1, characterized in that: The optimal weight w of the numerical simulation model is: Where S represents the scoring rule; p k,-i (y i ) represents the numerical simulation model M k In the numerical simulation model, n-1 points other than the i-th point are used to train the numerical simulation model, and the probability distribution of the observation point i is obtained; y i Represents the numerical simulation model M k The corresponding predicted value; n represents the number of observation points; C1 k Represents a set, and the weight of the solved numerical simulation model is defined in this set. The model weight is between 0 and 1, and the sum is 1; w is the optimal weight of the numerical simulation model and belongs to C1 k gather.

5. The method for calculating groundwater multi-model average weights based on Bayesian Stacking according to claim 1, characterized in that: The weighted comprehensive forecast distribution of the forecast points: Where w k is the numerical simulation model M k The average weight of p(y -i |y i ,M k ) is the observation point y i The probability density value corresponding to the true value of in the leave-one-out prediction distribution.