A method and device for characterizing a coastal heterogeneous aquifer based on machine learning
Patent Information
- Application Number
- CN202310916614.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-25
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2043-07-25
AI Technical Summary
ESMDA算法在处理多数据源的情况下具有优势,但在非高斯系统中难以取得令人满意的结果
[0039] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention avoids the Gaussian assumption required by the Kalman gain formula by replacing the Kalman gain formula with a machine learning model; secondly, the machine learning model used in the present invention has the ability to extract complex (including non-Gaussian) features and learn nonlinear relationships from a large amount of training data. In the problem of characterizing coastal heterogeneous aquifers, the inversion effect obtained will be better than the ESMDA algorithm based on Kalman update.
Smart Images

Figure CN117057221B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of hydrogeology technology, specifically relating to a method and apparatus for characterizing coastal heterogeneous aquifers based on machine learning. Background Technology
[0002] Seawater intrusion is a phenomenon where the brackish water interface of coastal aquifers extends inland due to overexploitation of freshwater resources in coastal aquifers and land use changes in coastal areas. Seawater intrusion seriously threatens water security in coastal areas. To achieve scientific management of coastal aquifers, accurate prediction of seawater intrusion processes is necessary, and numerical models are a common method. Therefore, accurately characterizing heterogeneous coastal aquifers is a crucial step in achieving accurate prediction of seawater intrusion. However, obtaining accurate characterizations of coastal aquifers through drilling and other exploration activities requires substantial costs, making this method impractical. A feasible solution is to use data assimilation methods to characterize heterogeneous coastal aquifers using other relatively easily accessible observational data.
[0003] In the field of data assimilation, mainstream methods include Markov Chain Monte Carlo (MCMC), Kalman Filtering, Variational Methods, and ESMDA. MCMC is a Monte Carlo method based on Markov chains. It requires a large number of samples to obtain sufficient information, and the computational cost is enormous due to multiple forward model calculations during its execution. Furthermore, convergence issues may arise during sample extraction. Even though MCMC has a complete theoretical basis, its computational cost limits its ability to handle high-dimensional, nonlinear, and non-Gaussian problems. Kalman Filtering is a recursive filtering algorithm used to estimate system states from noisy measurements. Kalman Filtering excels in handling Gaussian noise and time-varying systems, but has limitations with nonlinear and non-Gaussian systems. Variational Methods are optimization methods used to compute the posterior and marginal distributions in probabilistic models. Variational Methods simplify complex calculations by approximating the true posterior distribution. Variational Methods are advantageous in handling large-scale data and model parameter learning, but cannot handle complex posterior distributions such as non-Gaussian distributions. The ESMDA algorithm is an ensemble learning-based method used to estimate parameters and latent variables in probabilistic models. ESMDA improves accuracy by integrating predictions from multiple models and multiple data sources. While it excels in handling multiple data sources, it struggles to achieve satisfactory results in non-Gaussian systems. Clearly, there is currently no effective data assimilation algorithm for characterizing high-dimensional, nonlinear, and non-Gaussian aquifers. Therefore, developing a robust data assimilation algorithm for characterizing heterogeneous coastal aquifers is a pressing issue. Summary of the Invention
[0004] Purpose of the invention: In the problem of characterizing heterogeneous aquifers in coastal areas, this invention provides a method and apparatus for characterizing heterogeneous aquifers in coastal areas based on machine learning, and the inversion effect obtained is better than that of the ESMDA algorithm based on Kalman update.
[0005] Technical Solution: This invention provides a method for characterizing heterogeneous coastal aquifers based on machine learning, comprising the following steps:
[0006] (1) The simulation data of several sets of hydraulic conductivity coefficients randomly generated from the prior parameter distribution and the output of the numerical model of seawater intrusion were obtained from the numerical model of seawater intrusion.
[0007] (2) By subtracting the samples pairwise, a large number of samples of parameter differences and simulated data differences are generated from the prior set;
[0008] (3) Use the above data to train a machine learning model and construct a mapping relationship between the simulated data difference and the parameter difference;
[0009] (4) Input the difference between the observed data and the prior simulation data set into the trained machine learning model, and add the output to the hydraulic conductivity coefficient set to obtain the update result of the prior hydraulic conductivity coefficient set.
[0010] (5) Repeat steps (2) to (4) until the iteration ends, and obtain the hydraulic conductivity coefficient after iteration.
[0011] Furthermore, the simulation data in step (1) includes water head, seawater concentration, and land-based pollutant concentration.
[0012] Furthermore, the implementation process of step (1) is as follows:
[0013] N is randomly selected from the prior distribution of hydraulic conductivity. e Group of hydraulic conductivity coefficient samples Substitute into the numerical model of seawater intrusion based on COMSOL Multiphysics Obtain the simulation data output by the corresponding model: in
[0014] Furthermore, the numerical model for seawater intrusion includes a flow model and a transport model;
[0015] The flow model:
[0016]
[0017]
[0018] Where u is Darcy velocity; K is hydraulic conductivity; p is pore pressure; ρ is fluid density; g is gravitational acceleration; ∈ P (-) represents porosity; Q m For source and sink items;
[0019] The transport model:
[0020]
[0021]
[0022] Among them, c i c represents the concentration of substance i in the liquid. P,i Indicates the amount of adsorption by solid particles; J i R is the mass flux diffusion flux vector; i Express the reaction rate of this substance; S i For any source term; D represents the effective diffusion coefficient of liquid i; F,i D is the fluid diffusion coefficient; D,i Let i represent the dispersion tensor of liquid i.
[0023] Furthermore, the implementation process of step (2) is as follows:
[0024] From {X (0) ,Y (0) If no two samples are taken from the given set, and the two sets are subtracted, we get N = N. e (N e -1) / 2 sets of training data, that is:
[0025]
[0026]
[0027] Where, ε ij A random sample of observation errors; This represents the difference between simulated data. This represents the parameter difference.
[0028] Furthermore, the implementation process of step (3) is as follows:
[0029] Will As input, The output is used to train a machine learning model, obtaining the mapping relationship between simulated data differences and parameter differences.
[0030] Furthermore, the implementation process of step (4) is as follows:
[0031]
[0032] in, This is the updated set of hydraulic conductivity coefficients. For observational data.
[0033] Furthermore, the implementation process of step (5) is as follows:
[0034] Set the number of iterations to N. iter First from {X (t-1) ,Y (t-1 Generate models for training machine learning models. Data, The collection is updated as follows: After N iter After this update, the final parameter set is:
[0035] The present invention also provides an apparatus comprising a memory and a processor, wherein:
[0036] Memory is used to store computer programs that can run on a processor;
[0037] A processor, configured to, while running the computer program, execute the steps of the machine learning-based coastal heterogeneous aquifer characterization method described above.
[0038] The present invention also provides a storage medium storing a computer program, which, when executed by at least one processor, implements the steps of the machine learning-based coastal heterogeneous aquifer characterization method described above.
[0039] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention avoids the Gaussian assumption required by the Kalman gain formula by replacing the Kalman gain formula with a machine learning model; secondly, the machine learning model used in the present invention has the ability to extract complex (including non-Gaussian) features and learn nonlinear relationships from a large amount of training data. In the problem of characterizing coastal heterogeneous aquifers, the inversion effect obtained will be better than the ESMDA algorithm based on Kalman update. Attached Figure Description
[0040] Figure 1 A schematic diagram of the method for characterizing coastal heterogeneous aquifers based on machine learning in this invention;
[0041] Figure 2 A schematic diagram showing the setup of the numerical model for seawater intrusion in the implementation case;
[0042] Figure 3 The following are comparison diagrams of the reference field and the characterization effect; where (a) is the reference field diagram of the hydraulic conductivity coefficient of the implementation case; (b) is the distribution diagram of the observation wells of the implementation case; (c) is the characterization effect diagram of ESMDA of the implementation case; and (d) is the characterization effect diagram of DA of the implementation case. ML Depicting the effect;
[0043] Figure 4 A schematic diagram of the machine learning model used in the implementation case;
[0044] Figure 5 In the implementation case, ESMDA and DA ML The root mean square error curve of the obtained mean hydraulic conductivity coefficient compared with the reference value. Detailed Implementation
[0045] The present invention will now be described in further detail with reference to the accompanying drawings.
[0046] This invention provides a method for characterizing coastal heterogeneous aquifers based on machine learning (DA). ML ),like Figure 1 As shown, it includes the following steps:
[0047] S1. Obtain several sets of parameter samples (i.e., hydraulic conductivity coefficients) randomly generated from the prior parameter distribution by the seawater intrusion numerical model, and their corresponding model output simulation data (i.e., water head, seawater concentration, and land-based pollutant concentration).
[0048] 500 sets of hydraulic conductivity samples were randomly selected from the prior distribution of hydraulic conductivity. Substitute into the model Obtain the simulation data output by the corresponding model: in
[0049] The seawater intrusion and land-based pollutant transport were simulated by coupling the “Darcy’s Law” and “The Transport of Diluted Species in Porous Media” modules in COMSOL Multiphysics. The flow and transport models were solved sequentially using a separate solver.
[0050] For flow models:
[0051]
[0052]
[0053] Where u (m / s) is Darcy velocity; K (m / s) is hydraulic conductivity; p (Pa) is pore pressure; ρ (kg / m 3 ) represents the fluid density; g (m / s) 2 ) represents gravitational acceleration; ∈ P (-) represents porosity; Q m (kg / (m 3 ·s)) represents source and sink items.
[0054] For the transport model:
[0055]
[0056]
[0057] Among them, c i (mol / m 3 ) represents the concentration of substance i in the liquid, c P,i (-) indicates the amount of solid particles adsorbed (i.e., the number of moles per unit dry weight of the solid); J i (mol / m 2 ·s) is the mass flux diffusion flux vector; R i (mol / m 3 ·s) represents the reaction rate expression for this substance; S i (kg / (m 3 ·s)) is any source term; D represents the effective diffusion coefficient of liquid i; F,i (m 2 / s) is the fluid diffusion coefficient; D D,i (m 2 / s) represents the dispersion tensor of liquid i.
[0058] S2. By subtracting the samples pairwise, a large number of samples are generated from the prior set to obtain the parameter difference and the difference of the simulated data.
[0059] From {X (0) ,Y (0) By subtracting two non-repeating samples from the given set, we obtain N = 500(500-1) / 2 = 124750 sets of training data. Where: ε ij A random sample of observation errors; This represents the difference between simulated data. This represents the parameter difference.
[0060] S3. Use the above data to train a machine learning model and construct a mapping relationship between the difference in simulated data and the difference in parameters.
[0061] Will As input, Substituting as output Figure 4 The example shows how a machine learning model called "DenseNet" is trained to obtain the mapping relationship between simulated data differences and parameter differences.
[0062] S4: Input the difference between the observed data (i.e., the actual observed values of water head, seawater concentration, and land-based pollutant concentration) and the prior simulated data set into the trained machine learning model. Add the output (the predicted difference between the prior parameters and the actual parameter values) to the prior parameter set (i.e., the hydraulic conductivity coefficient) to obtain the updated result of the prior hydraulic conductivity coefficient set.
[0063]
[0064] in, This is the updated set of hydraulic conductivity coefficients. For observational data.
[0065] S5. For strongly nonlinear problems, repeat steps S2 to S4 until the iteration ends to obtain the parameters (hydraulic conductivity coefficient) after iteration, which is the characterization of the coastal aquifer.
[0066] Set the number of iterations to N. iter First from {X (t-1) ,Y (t-1) Generate models for training machine learning models. Data, The collection is updated as follows: After N iter After this update, the final parameter set is:
[0067] Based on the same inventive concept, the present invention also provides an apparatus comprising a memory and a processor, wherein: the memory is used to store a computer program capable of running on the processor; and the processor is used to execute, when running the computer program, the steps of the method for characterizing coastal heterogeneous aquifers based on machine learning as described above.
[0068] The present invention also provides a storage medium storing a computer program, which, when executed by at least one processor, implements the steps of the machine learning-based coastal heterogeneous aquifer characterization method described above.
[0069] To verify that this invention can be used to characterize heterogeneous coastal aquifers, an example is considered: a two-dimensional vertical coastal aquifer. For example... Figure 2 As shown, the study area has dimensions of L×B=240×30m, with a constant freshwater head h on the left.f =31.6m, concentration is C f =0kg / m 3 The right side shows a constant seawater head h. s =31m, concentration is C s =35kg / m 3 The top and bottom edges are impermeable. This study simulates a total duration of 4500 days. Initially, the region is filled with fresh water with a head of 30 m. A pumping well with a diameter of 0.5 m located at (100, 25) pumps water at a rate of 2000 kg / d after 1000 days. A square region with a side length of 0.5 m located at the upper left corner (20, 25) pumps water at a rate of 35 kg / (m²) over 1500-2500 days. 3 •d) releases land-based pollutants outward at an intensity that is high.
[0070] In this numerical case, the observed data (i.e., the actual observed values of water head, seawater concentration, and land-based pollutant concentration) will be... Figure 3 The 91×241-dimensional hydraulic conductivity (K) reference field shown in Figure (a) is substituted into the numerical model for calculation and a normal distribution is added. The measurement error was obtained. The data for water head, seawater concentration, and land-based pollutant concentration were based on... Figure 3 The observation well shown in (b) was obtained at times t = [300, 600, ..., 4500]d, t = [300, 600, ..., 4500]d, and t = [1800, 2100, ..., 4500]d.
[0071] Figure 3 In diagram (c), the mean field of K obtained by the ESMDA algorithm is shown. Clearly, the ESMDA algorithm only characterizes a portion of the reference field using observed data, and its characterization of the channel features exhibited in the reference field is poor. However, as... Figure 3 As shown in (d), by DA ML The mean field of hydraulic conductivity coefficients obtained by the algorithm inversion is closer to the reference field than that obtained by ESMDA, demonstrating a better characterization of the connectivity of the reference field and capturing almost all channel characteristics of the reference field. Figure 5 It can be seen that, from DA ML The root mean squared error (RMSE) between the hydraulic conductivity coefficient obtained by the algorithm and the reference value is significantly smaller than the RMSE between the simulated head value and the observed head value obtained by the ESMDA algorithm. Considering various evaluation criteria, it can be considered that DA... ML The method outperforms ESMDA in the inversion of groundwater parameters under non-Gaussian assumptions.
[0072] Based on the above analysis, the machine learning-based method for characterizing coastal heterogeneous aquifers of the present invention has the ability to characterize complex coastal heterogeneous aquifers.
[0073] The above embodiments are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solutions based on the technical concept proposed in this invention shall fall within the scope of protection of this invention.
Claims
1. A method for characterizing coastal heterogeneous aquifers based on machine learning, characterized in that, Includes the following steps: (1) The simulation data of several sets of hydraulic conductivity coefficients randomly generated from the prior distribution of parameters and the output of the numerical model of seawater intrusion were obtained from the numerical model of seawater intrusion. (2) By subtracting the samples pairwise, a large number of samples of parameter differences and simulated data differences are generated from the prior set; (3) Use the above data to train a machine learning model and construct a mapping relationship between the difference in simulated data and the difference in parameters. ; The machine learning model employs a DenseNet convolutional neural network based on a dense connection mechanism. The DenseNet convolutional neural network includes an initial feature transformation module D1, four dense blocks, a transition module D3, and a two-dimensional result reconstruction module D4. The initial feature transformation module D1 extracts initial features through transposed convolution, activation layers, convolution, and max pooling. The four dense blocks respectively include two, three, three, and two dense units D2. Each dense unit D2 consists of batch normalization, activation layers, and convolution. The transition module D3 performs feature transformation using transposed convolution, convolution, and average pooling. The two-dimensional result reconstruction module D4 outputs a single-channel result through multi-level upsampling and convolution. (4) Input the difference between the observed data and the prior simulation data set into the trained machine learning model, and add the output to the hydraulic conductivity coefficient set to obtain the update result of the prior hydraulic conductivity coefficient set; (5) Repeat steps (2) to (4) until the end of the iteration to obtain the hydraulic conductivity coefficient after iteration; The simulation data in step (1) are water head, seawater concentration, and land-based pollutant concentration; The implementation process of step (1) is as follows: Randomly selected from the prior distribution of hydraulic conductivity coefficient Group of hydraulic conductivity coefficient samples Substitute into the numerical model of seawater intrusion based on COMSOL Multiphysics Obtain the simulation data output by the corresponding model: ,in ; The numerical model for seawater intrusion includes a flow model and a transport model; The flow model: in, For Darcy's speed; The hydraulic conductivity coefficient; Pore pressure; For fluid density; It is the acceleration due to gravity; Porosity; For source and sink items; The transport model: in, Indicates liquid Concentration of such substances Indicates the amount of solid particles adsorbed; The mass flux diffusion flux vector; Express the reaction rate of this substance; For any source term; Indicates liquid The effective diffusion coefficient; The fluid diffusion coefficient; Indicates liquid The diffusion tensor.
2. The method for characterizing coastal heterogeneous aquifers based on machine learning according to claim 1, characterized in that, The implementation process of step (2) is as follows: from Two groups of samples are taken without repetition and then subtracted to obtain the result. The training data, i.e.: in, A random sample of observation errors; This represents the difference between simulated data. This represents the parameter difference.
3. The method for characterizing coastal heterogeneous aquifers based on machine learning according to claim 1, characterized in that, The implementation process of step (4) is as follows: in, This is the updated set of hydraulic conductivity coefficients. For observational data.
4. The method for characterizing coastal heterogeneous aquifers based on machine learning according to claim 1, characterized in that, The implementation process of step (5) is as follows: Set iteration First from Generate models for training machine learning models Data, , , The collection is updated as follows: ,go through After this update, the final parameter set is: .
5. A device, characterized in that, Includes memory and processor, wherein: Memory is used to store computer programs that can run on a processor; A processor, configured to, while running the computer program, perform the steps of the machine learning-based method for characterizing coastal heterogeneous aquifers as described in any one of claims 1 to 4.
6. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by at least one processor, implements the steps of the machine learning-based method for characterizing coastal heterogeneous aquifers as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Adjusting method of technological parameters of sewage treatment and device
CN102616927A
Industrial process soft measurement modeling method based on block increment random configuration network
CN109635337A