Method and device for constructing accurate ML quantitative model for complex groundwater environment

By constructing a precise ML quantitative model for complex groundwater environments, selecting associated characteristic variables and establishing multi-parameter spatiotemporal evolution control equations, and combining machine learning and data enhancement techniques, the problems of insufficient model suitability and accuracy in existing technologies are solved, and efficient and accurate groundwater pollutant distribution prediction and real-time evaluation are achieved.

CN119783493BActive Publication Date: 2025-09-30UNIV OF CHINESE ACAD OF SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411543386.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-09-30
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

Existing technologies have insufficient model suitability and accuracy in complex groundwater pollution prediction, high cost, low efficiency, and poor real-time performance, making it difficult to achieve rapid and accurate pollution status quantification and risk assessment.

Method used

A precise ML quantitative model construction method for complex groundwater environment is adopted, associated characteristic variables are selected, and a multi-parameter spatiotemporal evolution control equation is established. Pollutant distribution is predicted through a machine learning prediction model, and a suitable machine learning model is constructed by combining data enhancement and feature dimensionality reduction techniques.

Benefits of technology

It achieves efficient and accurate prediction and real-time evaluation of groundwater pollutant distribution under a conditional framework, improves the prediction accuracy and practicality of complex groundwater environments, and provides rapid early warning support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783493B_ABST
    Figure CN119783493B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for constructing a precise ML quantitative model for a complex groundwater environment, including selecting associated characteristic variables related to the spatiotemporal evolution of a complex groundwater pollution environment, establishing a conceptual model of the spatiotemporal evolution of a complex groundwater environment for a typical conditional framework, establishing and selecting an appropriate multi-parameter complex groundwater environment spatiotemporal evolution control equation, obtaining a corresponding groundwater environment spatiotemporal evolution numerical simulation system, obtaining a typical scenario data set based on this, and establishing a precise ML quantitative model for a complex groundwater environment with different pertinences and application characteristics and a certain degree of versatility according to demand. The present invention can accurately and efficiently realize the reliable quantification and prediction of the soil and groundwater environmental status, evolution process and related impacts in different application scenarios under the premise of meeting a given conditional framework, significantly improving the accuracy, efficiency and practicality of prediction, prejudgment and real-time evaluation of the spatiotemporal evolution status of a complex groundwater pollution environment and the spatiotemporal distribution of pollutants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil-groundwater pollution prediction, and specifically to a method and device for constructing an accurate ML quantitative model for a complex groundwater environment. Background Art

[0002] At present, in the prediction, evaluation and risk assessment of groundwater pollution, the application of numerical simulation systems to evaluate and quantify complex groundwater pollution status and environmental risks for specific application scenarios is a commonly used research method. However, in order to achieve accurate numerical simulation, it is necessary to comprehensively consider and refine the numerous influencing factors and complex evolution processes of specific scenarios, which is time-consuming and highly professional. In addition, there is a large uncertainty in the modeling and prediction of complex groundwater systems. The use of Monte Carlo and other methods for uncertainty assessment requires hundreds or thousands of numerical simulation tests. However, the computational cost is huge or even unacceptable, and the efficiency is low. Therefore, there is an urgent need for a complex groundwater environment spatiotemporal evolution prediction and quantification method that is highly applicable, accurate, efficient, simple and practical, and has a certain degree of versatility for different specific conditions.

[0003] In response to the above problems, the present invention provides a method and device for constructing an accurate ML quantitative model for a complex groundwater environment, so as to solve the problems in the prior art of the suitability and accuracy of the models used in the prediction and quantification of characteristic variables of complex groundwater pollution environments under specific conditions, large uncertainty, high cost, low efficiency, and poor real-time performance. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for constructing an accurate ML quantitative model of a complex groundwater environment, including selecting associated characteristic variables related to the spatiotemporal evolution of a complex groundwater pollution environment, establishing a conceptual model of the spatiotemporal evolution of a complex groundwater environment based on a typical condition framework, establishing and selecting an appropriate multi-parameter complex groundwater environment spatiotemporal evolution control equation, obtaining a corresponding groundwater environment spatiotemporal evolution numerical simulation system, obtaining a typical scenario data set based on this, and establishing an accurate ML quantitative model of a complex groundwater environment with different pertinence and application characteristics and a certain degree of versatility according to needs. The specific steps of the method are as follows:

[0005] Step 1: Determine the common feature set of a complex groundwater pollution environment system to be studied, which is called the system condition framework. Corresponding to the system condition framework, give a conceptual model of the spatiotemporal evolution of the complex groundwater pollution environment, select the characteristic variables associated with the spatiotemporal evolution of groundwater pollution, obtain the mathematical control equations of the spatiotemporal evolution of groundwater pollution, and obtain the numerical simulation system of the spatiotemporal evolution of groundwater pollution;

[0006] Step 2: Determine the input parameter set V = [v1, v2, ...v n ], select the range of each model input parameter to form the model input parameter range set U = [u1, u2, ..., un ];

[0007] Step 3: Based on the range set of model input parameter variables, different typical scenarios consisting of different combinations of model input parameters for the groundwater pollution environment system are obtained through scenario sampling. Based on the numerical simulation system, corresponding simulation values ​​of characteristic variables associated with different typical scenarios are obtained. Based on different model input parameter combinations for different typical scenarios and their corresponding numerical simulation values, a basic modeling data set is obtained.

[0008] Step 4: When an enhanced modeling dataset is needed, the input parameters of each model are sampled within their variation range according to the given preset number of supplementary scenarios, and different supplementary scenarios consisting of different combinations of model input parameters of the groundwater environmental system are obtained. The corresponding associated characteristic variable simulation values ​​of the different supplementary scenarios are obtained based on the numerical simulation system. The enhanced modeling dataset is obtained by combining the model input parameter combinations of all different supplementary scenarios and their corresponding associated characteristic variable simulation values.

[0009] Step 5: Consider all model input parameters for the associated characteristic variables and complete the comprehensive identification of the main controlling factors;

[0010] Step 6: Partially or fully construct in advance machine learning prediction models with certain versatility and different suitability and application characteristics as needed, including ML model 1, ML model 2, ML model 3, ML model 4, and ML model 5;

[0011] Step 7: Based on the preset ML model accuracy value, determine whether the ML model accuracy meets the model target accuracy. If so, complete the modeling and obtain the optimized ML model. Otherwise, reduce the preset model accuracy, increase the scenario sampling, and change the machine learning algorithm until the ML model accuracy meets the model target accuracy.

[0012] Step 8: Given an actual application scenario, determine the framework that meets the given typical conditions, and determine the appropriate required machine learning prediction model mentioned above to further obtain the specific prediction value of the application scenario.

[0013] Furthermore, the selected characteristic variables associated with the spatiotemporal evolution of groundwater pollution in step one include groundwater pollutant concentration, saturated zone pollutant cross-section or interface flux, pollution plume migration and diffusion front position, pollution plume stabilization time, stable pollution plume distribution area, distance between the stable pollution plume front and the pollution source, stable pollution plume average concentration and pollution plume dissipation time, soil pollutant concentration, vadose zone gaseous pollutant concentration, vadose zone cross-section pollutant flux, vadose zone-saturated zone pollutant interface flux, soil-air interface pollutant flux and rock-air interface pollutant flux, and NAPL pollutant spatiotemporal distribution.

[0014] Furthermore, the set of common characteristics of a type of groundwater pollution environmental system to be studied in step one is called the system condition framework, which includes determining the characteristics of atmospheric precipitation infiltration and recharge, determining the characteristics of the vadose zone, determining the characteristics of the saturated zone, determining the characteristics of the pollution source, determining the characteristics of groundwater pollution prevention and control and remediation, determining the characteristics of the source and sink of the pumping and injection wells, determining the type of groundwater pollutant migration and transformation process, determining the characteristics of the boundary conditions and determining the characteristics of the initial conditions.

[0015] Furthermore, the comprehensive identification of the main controlling factors in step five includes the application of two or more main controlling factor identification methods, including different machine learning methods, hierarchical analysis, variance analysis, range analysis, and sensitivity analysis, and the "and" or "or" set of the main controlling factors obtained based on the adopted methods is used to obtain the final main controlling factor parameter variables.

[0016] Furthermore, the associated characteristic variables and model input parameter variables in steps one and two are processed using any one or any combination of preset data processing methods, wherein the preset data processing methods include retaining the value as it is, increasing the value by the same multiple, decreasing the value by the same multiple, logarithmic conversion of the value, unifying the positive and negative values ​​of the value, and normalizing the value; obtaining a preset number of scenario samples of the model input parameter variables includes obtaining a preset number of random scenario samples based on the random value of the model input parameter variables according to the variation range of the model input parameter variables, and including determining scenario samples formed by selecting the variation level of each of the model input parameter variables and then combining the model input parameter variables;

[0017] Furthermore, the machine learning algorithms in step seven include KAN and FCN;

[0018] Furthermore, ML model 1 in step six is ​​an efficient and accurate full-factor model, including all parameters required when the model is applied, and no random occlusion sampling data enhancement during modeling; ML model 2 is a simple and reliable master factor model, including all master factor parameters required when the model is applied, and no random occlusion sampling data enhancement during modeling; ML model 3 is a wide-adaptability model in which non-master factor parameters are optional, including all master factor parameters required when the model is applied, but non-master factor parameters can be missing, and the model has good wide adaptability; random occlusion sampling is implemented for data enhancement of non-master factor parameters during modeling; ML model 4 is a wide-adaptability model in which any factor parameter can be missing, including any factor parameter can be missing when the model is applied, and the model has good wide adaptability; random occlusion sampling is implemented for data enhancement of all parameters during modeling; ML model 5 is a customized input parameter specific model, including the need to pre-select parameters when the model is applied, and no random occlusion sampling data enhancement during modeling.

[0019] Furthermore, the characteristics of atmospheric precipitation infiltration recharge are determined, including whether there is atmospheric precipitation infiltration recharge and the spatiotemporal variation characteristics of atmospheric precipitation infiltration recharge intensity; the characteristics of the vadose zone are determined, including medium type, lithologic structure, medium heterogeneity and anisotropy; the characteristics of the saturated zone are determined, including medium type, water-bearing system structure, seepage dimension, flow state, medium heterogeneity and anisotropy; the characteristics of the pollution source are determined, focusing on the phase state of pollutants, the type of pollutants, the spatial distribution characteristics of the pollution source, the source strength variation and dynamic characteristics; the characteristics of groundwater pollution prevention, control and remediation are determined, including whether there are groundwater pollution prevention, control and remediation measures, the type of groundwater pollution prevention, control and remediation measures, and the underground The spatiotemporal distribution characteristics of water pollution prevention, control and remediation measures, determination of the source and sink characteristics of pumping and injection wells include the presence or absence of pumping and injection wells, the spatial distribution characteristics of pumping and injection wells, and the dynamic characteristics of water volume in pumping and injection wells; determination of the interaction characteristics between surface water and groundwater include the presence or absence of interaction between the two, the spatial distribution characteristics of surface water bodies, and the dynamic characteristics of surface water levels; determination of the types of groundwater pollutant migration and transformation processes include advection processes, diffusion processes, adsorption and desorption processes, degradation processes, volatilization processes, etc.; determination of boundary condition characteristics include the boundary condition characteristics of the hydrodynamic field and the boundary condition characteristics of the hydrochemical field; determination of initial condition characteristics include the initial condition characteristics of the hydrodynamic field and the initial condition characteristics of the hydrochemical field.

[0020] A device for constructing a precise ML quantitative model of a complex groundwater environment, comprising a module for constructing a conceptual model of the spatiotemporal evolution of a multi-parameter complex groundwater pollution environment, a module for comprehensive identification of main controlling factors, a module for constructing a combination of scenario sampling parameters and a scenario numerical simulation module, a data enhancement module, a feature dimensionality reduction module, and a precise ML model construction and application module; the data enhancement module comprises parameter random occlusion sampling data enhancement; the feature dimensionality reduction module comprises a convolutional neural network; and the precise ML model construction and application module comprises the construction and application of a precise ML model based on KAN.

[0021] The present invention has the following advantages: the present invention can accurately and efficiently realize the reliable quantification and prediction of the soil and groundwater environmental status, evolution process and related impacts in different application scenarios under the premise of meeting a given condition framework, significantly improving the accuracy, efficiency and practicality of the prediction and real-time evaluation of the spatiotemporal evolution status of complex groundwater pollution environment and the spatiotemporal distribution of pollutants, and providing strong support for the protection of water and soil environmental systems and the safe utilization of resources, especially the rapid prediction and early warning of complex groundwater environmental systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a conceptual model diagram of the present invention. DETAILED DESCRIPTION

[0023] In this embodiment, the model adopts a study area with a length, width and height dimension of 20*20*10.5m, which is covered by 1 to 5 layers of heterogeneous soil layers (the number of layers is one of the random variables) and a fractured rock structure below. The study area has groundwater at a certain depth, which serves as a pollution source with a constant concentration. There are periodic air pressure fluctuations and rainfall on the surface. The designed conceptual model is as follows: (a) The lower part of the study area is fractured rock, and three groups of orthogonal fractures are designed and their directions follow the Fisher distribution. The inclination of the fractures can be changed, and the UDFM algorithm is used to realize the coupling of fractures and rocks; the upper part of the study area is covered with layered heterogeneous soil of different thicknesses; (b) The surface thickness of 0.5m is regarded as the surface air layer, with a sinusoidally changing pressure boundary to simulate the periodic change of air pressure, and at the same time according to Different rainfall infiltration values ​​are assigned to the soil surface each quarter. The study area is located below a saturated aquifer, which flows from west to east and has different hydraulic gradients. (c) The saturated aquifer is a pollution source with a constant VOCs concentration. The initial concentrations in other areas are all 0. The surface air boundary pollutant concentration is assumed to be constant at 0, that is, the surface has the maximum diffusion flux. (e) The model considers the distribution between gas, liquid, and solid phases, but does not consider the attenuation effect of biochemical reactions on gas phase VOCs, and does not involve the transport of NAPL phase. (f) The model takes g = 9.8 m / s2 to consider the density effect and the influence of temperature. The time step varies throughout the simulation time range, the gas phase viscosity is constant, and it is assumed that there is a local chemical equilibrium when VOCs are distributed between different phases (gas, water, and solid).

[0024] In this embodiment, a numerical simulation of VOCs contamination in groundwater in a layered heterogeneous soil-fractured rock double-layer structure is established, fully considering various actual conditions such as surface pressure fluctuations and rainfall to achieve a realistic portrayal of the site. Through a detailed investigation of pollutant migration patterns, a VOCs migration and transformation model under complex media conditions is established by coupling multiple processes. Using supercomputing, numerical models with hundreds of sets of random parameter distributions are established and used as the basic training set. Variance analysis, data enhancement based on master control factors, and convolutional feature extraction are used to reduce the dimensionality of high-dimensional feature data. Using stable fluxes as labels, an efficient and reliable deep learning universal surrogate model is established. This solves the problems of dimensionality curse, low versatility, and insufficient data volume that exist in the construction of complex groundwater pollution surrogate models. Furthermore, a surrogate model solution with high prediction accuracy can be achieved even with partial parameter input. In particular, suitable ML models with different characteristics and advantages can be established according to actual needs, providing a reliable ML model for accurate and rapid VOCs surface flux prediction.

[0025] In this embodiment, the conceptual model construction and numerical simulation system selection are as follows: Figure 1As shown in the figure, the model adopts a study area with a length, width and height dimension of 20*20*10.5m, which is covered by 1 to 5 layers of heterogeneous soil layers (the number of layers is one of the random variables) and a fractured rock structure below. The study area has groundwater at a certain depth, which serves as a pollution source with a constant concentration. There are periodic air pressure fluctuations and rainfall on the surface. The designed conceptual model is as follows: (a) The lower part of the study area is fractured rock, and three groups of orthogonal fractures are designed and their directions follow the Fisher distribution. The inclination of the fractures can be changed, and the UDFM algorithm is used to realize the coupling of fractures and rocks; the upper part of the study area is covered with layered heterogeneous soil of different thicknesses; (b) The surface thickness of 0.5m is regarded as the surface air layer, with a sinusoidally changing pressure boundary to simulate the periodic change of air pressure, and at the same time according to the season. Different rainfall infiltration values ​​are assigned to the soil surface. The study area is located below a saturated aquifer, which flows from west to east and has different hydraulic gradients. (c) The saturated aquifer is a pollution source with a constant VOCs concentration. The initial concentrations in other areas are all 0. The surface air boundary pollutant concentration is assumed to be constant at 0, that is, the surface has the maximum diffusion flux. (e) The model considers the distribution between gas, liquid, and solid phases, does not consider the attenuation effect of biochemical reactions on gas phase VOCs, and does not involve the transmission of NAPL phase. (f) The model takes g = 9.8 m / s2 to consider the density effect and the influence of temperature. The time step varies throughout the simulation time range, the gas phase viscosity is constant, and it is assumed that there is a local chemical equilibrium when VOCs are distributed between different phases (gas, water, and solid).

[0026] In this embodiment, the governing equations for the spatiotemporal evolution of VOCs multiphases are described by the mass conservation equation (in concentration form). FEHM is selected as the numerical simulation system for this embodiment. This numerical simulation system can simulate the flow and transport of groundwater in fractures and porous media, and is suitable for systems with complex geometries. It uses the finite volume method and finite element method to solve the heat and mass conservation equations. FEHM has been widely used and rigorously verified in many fields, including non-isothermal multiphase flow, migration and transformation of gaseous pollutants, and other related fields.

[0027] In this embodiment, the numerical simulation parameters are selected as shown in Table 1. The study fully considers the migration and transformation processes of VOCs in loose pores and bedrock cracks to accurately quantify key indicators such as flux and concentration. A total of 76 parameters are incorporated into the model, including the total soil thickness h, the total number of soil layers l, the thickness of the first to fifth layers of soil h1-h5, the gas-solid partition coefficient K of the first to fifth layers of soil ssa1 -K ssa5 , the soil adsorption coefficient K from the first to the fifth layer sd1 -K sd5 , soil dispersivity α from the first to the fifth layer s1 -α s5, the soil permeability K from the first to the fifth layer s1 -K s5 , the soil porosity from the first to the fifth layer p s1 -p s5 , the residual saturation of soil from the first to the fifth layer S rr1 -S rr5 , the coefficients VG of the soil VG model from the first to the fifth layer a1 -VG a5 , the coefficients VG of the soil VG model from the first to the fifth layer n1 -VG n5 , gas phase diffusion coefficient D v , rock gas-solid partition coefficient K rsa , rock adsorption coefficient K rd , water phase diffusion coefficient D w , Henry's coefficient K H , initial concentration of pollutants C0, air pressure fluctuation period A fp , water table depth H, temperature T, rainfall in the first and fourth quarters R 14 , rainfall in the second and third quarters R 23 , hydraulic gradient i, bedrock dispersivity α rx (x direction), crack diffusion α fx (x direction), crack density F d , vertical crack angle F vephi , the inclination angle F of the vertical crack 1 vdip1 , the inclination angle F of vertical crack 2 vdip2 , horizontal crack inclination F hdip1 , horizontal crack inclination F hephi , ekappa, mean radius F of crack clusters 1 to 3 r1 -F r3 , the standard deviation F of the radius distribution of crack clusters 1 to 3 rv1 -F rv3 , mean crack width F a , standard deviation of crack width F av After fully investigating the distribution of each parameter under actual environmental information, each parameter is randomly sampled within a reasonable range, and the model simulation is continued until the surface gas phase pollutant diffusion flux reaches stability. This embodiment is designed to generate 400 sets of models with completely random parameters. In order to run numerical simulations of such a scale, the research uses a supercomputing platform to help realize the calculation and storage of large-scale numerical simulations. The grid is generated on the cloud supercomputing platform, and the FEHM solution is run on the cloud server. The total amount of data generated by the simulation exceeds 20T.

[0028] In this embodiment, basic simulation data acquisition and data enhancement: the 400 groups of numerical simulation data are divided into 350 training sets and 50 validation sets, and data enhancement is performed on each of them. First, ANOVA is used to extract the main controlling factors of stable flux, stable time, and cumulative amount before stability, and the original data is upsampled by random mask operation to ensure that the main controlling factor of the current label remains unchanged. 5%-50% of the remaining features are randomly deleted, and the 350 training sets are enhanced to 3500 groups, and the 50 validation sets are enhanced to 200 groups. On the basis of determining the main factors, the data enhancement effects of all 76 features participating in random masking and retaining the random mask of the main controlling factors are obtained.

[0029] In this embodiment, data preprocessing: The training of regression tasks using deep neural networks usually requires a dataset with sufficiently large known features and labels. The study generated 400 sets of datasets containing 76 features through numerical simulation, with stable flux as the label. In order to eliminate the influence of data dimension, the study used Ln to normalize the input and output. The ln1p normalization function was used for the feature parameters to ensure the validity of the data. For the stable flux f, due to its large order of magnitude difference and the fact that it is a negative number after using ln, the study took its inverse.

[0030] In this embodiment, the basic architecture of the ML model is constructed by combining KAN with weighted random upsampling. First, a CNN encoder is used for feature extraction, the original fully connected layer is deleted, and the convolution layer is used to extract local features. Features are extracted layer by layer through multi-layer convolution. The lower layers can learn simple local features, while the higher layers can combine these local features to form more complex global features. At the same time, the pooling layer is used to reduce the feature dimension, retain important information and reduce computational complexity, thereby improving the efficiency of the model. Finally, the KAN network is used to complete the regression task, greatly reducing the number of parameters. This network structure is abbreviated as AUCK (ANOVA-UPSAMPLING-CNN-KAN).

[0031] In this embodiment, the main controlling factors are identified as shown in Table 2.

[0032] In this embodiment, a corresponding model is established based on the basic data set: In order to discuss the performance of the network on the original data set, a CNN encoder is used to extract features from the original data, and then the data is fed into the KAN network for prediction. The data set contains 76 feature dimensions, labeled as stable flux. 350 sets of data are used for training and 50 sets of data are used for verification. A two-layer convolutional network is established based on the data scale. The first layer has an input channel input of 1, an input sequence length of 76, 16 convolution kernels, or 16 output channels, a convolution kernel size of 3, and padding and stride of 1. The second layer uses 32 convolution kernels of size 3, which are directly connected to the KAN network. The verification R of the stable flux is calculated. 2 are 0.96 respectively. See Table 3 for the specific results.

[0033] In this embodiment, two different models are obtained based on data enhancement: ML model 3 and ML model 4. The results are shown in Table 4. The model using weighted random upsampling is better than the model using full parameter random upsampling, which verifies R 2 The performance on stable flux labels has been improved, indicating that under the same network structure, weighted random upsampling based on variance analysis and hierarchical analysis has a good improvement on model accuracy.

[0034] In this embodiment, the specific conditions of the actual application scenario are obtained, and after comparative analysis, it is determined that it meets the given typical condition framework. Since the specific values ​​of all the main control factor parameters can be determined, and the values ​​of some non-main control factor parameters are highly uncertain or currently unknown, ML model 2 is selected, and based on the specific values ​​of the actual main control factor parameters, the specific prediction value of the application scenario is further obtained.

[0035] Table 1 Random sampling scheme for model parameters

[0036] Table 2 Results of variance analysis and hierarchical analysis weight discrimination of main control factors

[0037] Table 3. Deep learning network model based on KAN application raw data

[0038] Table 4. Different deep learning network models enhanced by KAN data

[0039] Although the specific embodiments of the present invention are described in detail in conjunction with the accompanying drawings, they should not be construed as limiting the scope of protection of the present invention. Various modifications and variations that can be made by those skilled in the art without creative work within the scope described in the claims are still within the scope of protection of the present invention.

[0040] Table 1

[0041]

[0042]

[0043]

[0044] Table 2

[0045]

[0046] Table 3

[0047]

[0048] Table 4

[0049]

Claims

1. A method for constructing a precise ML quantitative model for a complex groundwater environment, including selecting associated characteristic variables related to the spatiotemporal evolution of a complex groundwater contaminated environment, establishing a conceptual model of the spatiotemporal evolution of a complex groundwater environment based on a typical conditional framework, establishing and selecting appropriate multi-parameter control equations for the spatiotemporal evolution of a complex groundwater environment, obtaining a corresponding numerical simulation system for the spatiotemporal evolution of the groundwater environment, and obtaining a typical scenario dataset based on this. Based on this, a precise ML quantitative model for a complex groundwater environment with different targeted and application characteristics and a certain degree of versatility is established as required. The method is characterized by: The specific steps of the method are as follows: Step 1: Determine the common feature set of a complex groundwater pollution environment system to be studied, which is called the system condition framework. Corresponding to the system condition framework, give a conceptual model of the spatiotemporal evolution of the complex groundwater pollution environment, select the characteristic variables associated with the spatiotemporal evolution of groundwater pollution, obtain the mathematical control equations of the spatiotemporal evolution of groundwater pollution, and obtain the numerical simulation system of the spatiotemporal evolution of groundwater pollution; Step 2: Determine the input parameter set V = [v1, v2, ...v n ], select the range of each model input parameter to form the model input parameter range set U = [u1, u2, ..., u n ]; Step 3: Based on the range set of model input parameter variables, different typical scenarios consisting of different combinations of model input parameters for the groundwater pollution environment system are obtained through scenario sampling. Based on the numerical simulation system, corresponding simulation values ​​of characteristic variables associated with different typical scenarios are obtained. Based on different model input parameter combinations for different typical scenarios and their corresponding numerical simulation values, a basic modeling data set is obtained. Step 4: When an enhanced modeling dataset is needed, the input parameters of each model are sampled within their variation range according to the given preset number of supplementary scenarios, and different supplementary scenarios consisting of different combinations of model input parameters of the groundwater environmental system are obtained. The corresponding associated characteristic variable simulation values ​​of the different supplementary scenarios are obtained based on the numerical simulation system. The enhanced modeling dataset is obtained by combining the model input parameter combinations of all different supplementary scenarios and their corresponding associated characteristic variable simulation values. Step 5: Consider all model input parameters for the associated characteristic variables and complete the comprehensive identification of the main controlling factors; Step 6: Partially or fully construct in advance machine learning prediction models with certain versatility and different suitability and application characteristics as needed, including ML model 1, ML model 2, ML model 3, ML model 4, and ML model 5; Step 7: Based on the preset ML model accuracy value, determine whether the ML model accuracy meets the model target accuracy. If so, complete the modeling and obtain the optimized ML model. Otherwise, reduce the preset model accuracy, increase the scenario sampling, and change the machine learning algorithm until the ML model accuracy meets the model target accuracy. Step 8. Given an actual application scenario, determine the framework that meets the given typical conditions based on specific conditions, and determine the appropriate required machine learning prediction model mentioned above to further obtain the specific prediction value of the application scenario.

2. The method for constructing a precise ML quantitative model for a complex groundwater environment according to claim 1, characterized in that: The selected characteristic variables associated with the spatiotemporal evolution of groundwater pollution in the step one include groundwater pollutant concentration, saturated zone pollutant cross-section or interface flux, pollution plume migration and diffusion front position, pollution plume stabilization time, stable pollution plume distribution area, distance between the stable pollution plume front and the pollution source, stable pollution plume average concentration and pollution plume dissipation time, soil pollutant concentration, vadose zone gaseous pollutant concentration, vadose zone cross-section pollutant flux, vadose zone-saturated zone pollutant interface flux, soil-air interface pollutant flux and rock-air interface pollutant flux, and NAPL pollutant spatiotemporal distribution.

3. The method for constructing a precise ML quantitative model for a complex groundwater environment according to claim 1, characterized in that: The set of common characteristics of a type of groundwater pollution environment system to be studied in the step one is called the system condition framework, which includes determining the infiltration and recharge characteristics of atmospheric precipitation, determining the characteristics of the vadose zone, determining the characteristics of the saturated zone, determining the characteristics of the pollution source, determining the characteristics of the groundwater pollution prevention and control and remediation, determining the source and sink characteristics of the pumping and injection wells, determining the source and sink characteristics of the pumping and injection wells, determining the type of groundwater pollutant migration and transformation process, determining the boundary condition characteristics and determining the initial condition characteristics.

4. The method for constructing a precise ML quantitative model for a complex groundwater environment according to claim 1, characterized in that: The comprehensive identification of the main controlling factors in step five includes applying two or more main controlling factor identification methods, including different machine learning methods, hierarchical analysis, variance analysis, range analysis, and sensitivity analysis, and obtaining the final main controlling factor parameter variables based on the "union" or "intersection" set of the main controlling factors obtained by the adopted methods.

5. The method for constructing a precise ML quantitative model for a complex groundwater environment according to claim 1, characterized in that: The associated characteristic variables and model input parameter variables in step one and step two are processed using any one or any combination of preset data processing methods, and the preset data processing methods include retaining the values ​​as they are, increasing the values ​​by the same multiple, decreasing the values ​​by the same multiple, logarithmic conversion of the values, unifying the positive and negative values, and normalizing the values. The model input parameter variables are used to obtain a preset number of scenario samples, including random scenario samples obtained by randomly taking values ​​of the model input parameter variables according to the range of change of the model input parameter variables, and deterministic scenario samples formed by selecting the change level of each model input parameter variable and then combining the model input parameter variables.

6. The method for constructing a precise ML quantitative model for a complex groundwater environment according to claim 1, characterized in that: The machine learning algorithms in step seven include KAN and FCN.

7. The method for constructing an accurate ML quantitative model for a complex groundwater environment according to claim 1, characterized in that: The ML model 1 in step 6 is an efficient and accurate full-factor model, including all parameters required for model application, and no random occlusion sampling data enhancement during modeling; the ML model 2 is a simple and reliable master factor model, including all master factor parameters required for model application, and no random occlusion sampling data enhancement during modeling; the ML model 3 is a wide-adaptability model in which non-master factor parameters are optional, including all master factor parameters required for model application, but non-master factor parameters can be missing, and the model has good wide adaptability; random occlusion sampling is implemented for data enhancement of non-master factor parameters during modeling; the ML model 4 is a wide-adaptability model in which any factor parameter can be missing, including any factor parameter can be missing when the model is applied, and the model has good wide adaptability; During modeling, all parameters are subjected to random occlusion sampling for data enhancement; the ML model 5 is a customized input parameter specific model, including the need to pre-select parameters when the model is applied, and there is no random occlusion sampling data enhancement during modeling.

8. The method for constructing a precise ML quantitative model for a complex groundwater environment according to claim 3, characterized in that: The determination of atmospheric precipitation infiltration recharge characteristics includes whether there is atmospheric precipitation infiltration recharge and the spatiotemporal variation characteristics of atmospheric precipitation infiltration recharge intensity; the determination of vadose zone characteristics includes medium type, lithologic structure, medium heterogeneity and anisotropy characteristics; the determination of saturated zone characteristics includes medium type, water-bearing system structure, seepage dimension, flow state, medium heterogeneity and anisotropy characteristics; the determination of pollution source characteristics focuses on pollutant phase state, pollutant type, pollution source spatial distribution characteristics, source strength variation and dynamic characteristics; the determination of groundwater pollution prevention, control and remediation characteristics includes whether there are groundwater pollution prevention, control and remediation measures, groundwater pollution prevention, control and remediation measures type, underground The spatiotemporal distribution characteristics of water pollution prevention, control and remediation measures; the determination of the source and sink characteristics of pumping and injection wells includes the presence or absence of pumping and injection wells, the spatial distribution characteristics of pumping and injection wells, and the dynamic characteristics of the water volume of pumping and injection wells; the determination of the interaction characteristics between surface water and groundwater includes the presence or absence of interaction between the two, the spatial distribution characteristics of surface water bodies, and the dynamic characteristics of the water level of surface water bodies; the determination of the types of groundwater pollutant migration and transformation processes includes horizontal flow processes, diffusion processes, adsorption and desorption processes, degradation processes, volatilization processes, etc.; the determination of boundary condition characteristics includes the boundary condition characteristics of the hydrodynamic field and the boundary condition characteristics of the hydrochemical field; the determination of initial condition characteristics includes the initial condition characteristics of the hydrodynamic field and the initial condition characteristics of the hydrochemical field.

9. A device for constructing a precise ML quantitative model for complex groundwater environments, characterized by: The device includes a module for constructing a conceptual model of the spatiotemporal evolution of a complex groundwater pollution environment with multiple parameters, a module for comprehensively identifying main control factors, a module for constructing a combination of scenario sampling parameters and a scenario numerical simulation module, a data enhancement module, a feature dimensionality reduction module, and a module for constructing and applying a precise ML model; the data enhancement module includes data enhancement using random occlusion sampling of parameters; the feature dimensionality reduction module includes a convolutional neural network; and the precise ML model construction and application module includes construction and application of a precise ML model based on KAN.