A Method and Apparatus for Constructing a General Prediction Model for Groundwater Pollution Based on ANN and SVM
By using a general prediction model for groundwater pollution based on ANN and SVM, the problems of high prediction uncertainty, high cost, and low efficiency in groundwater pollution monitoring are solved. This model enables efficient and accurate prediction of the spatiotemporal distribution and evolution of groundwater pollutants, thereby improving the accuracy and practicality of risk assessment.
Patent Information
- Application Number
- CN202411543118.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-10-31
AI Technical Summary
Existing technologies for groundwater pollution monitoring and risk assessment suffer from problems such as high prediction uncertainty, high cost, and low efficiency. In particular, in the modeling of complex media groundwater systems, methods such as Montto Carlo are computationally expensive and inefficient.
A general prediction model for groundwater pollution based on ANN and SVM is adopted. By determining the conceptual model of the spatiotemporal evolution of groundwater pollution under typical conditions, the associated characteristic variables are obtained. The basic typical scenario sampling is obtained by using the orthogonal experimental method, a multivariate regression statistical prediction model is established, and the model is optimized by combining ANN and SVM to obtain the prediction confidence intervals of the associated characteristic variables.
It enables efficient and accurate prediction of the spatiotemporal distribution and evolution of groundwater pollutants under given conditions, improves the accuracy and practicality of complex groundwater environmental status assessment, and provides strong support for rapid early warning and risk management of soil and groundwater pollution.
Smart Images

Figure CN119441983B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of soil-groundwater pollution prevention and risk management technology, specifically involving a method and apparatus for constructing a general prediction model for groundwater pollution based on ANN and SVM. Background Technology
[0002] Currently, in the monitoring and risk assessment of groundwater pollution, numerical simulation systems are commonly used to evaluate and quantify the state of groundwater pollution and environmental risks. However, accurate numerical simulation requires comprehensive consideration of numerous influencing factors and complex evolution processes, which is time-consuming and highly specialized. Furthermore, the modeling and prediction of complex media groundwater systems inherently involve significant uncertainties. Uncertainty assessment using methods such as Montto Carlo requires hundreds or thousands of numerical simulation experiments, which are computationally expensive and even unacceptable, and inefficient. Therefore, there is an urgent need for a practical, efficient, and simple method to identify the value range of groundwater pollution-related characteristic variables. Summary of the Invention
[0003] The purpose of this invention is to provide a method and apparatus for constructing a general prediction model for groundwater pollution based on ANN and SVM, so as to solve the technical problems of high uncertainty, high cost and low efficiency in the prediction of environmental variables of groundwater pollution in the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] In a first aspect, this invention provides a method for constructing a general prediction model for groundwater pollution based on ANN and SVM, including:
[0006] Step A: Determine the conceptual model of the spatiotemporal evolution of groundwater pollution under typical conditions and obtain the corresponding set of typical features, i.e., the typical condition framework; select the characteristic variables associated with the spatiotemporal evolution of groundwater pollution; obtain the mathematical control equations for the migration and transformation of groundwater pollutants; and obtain the numerical simulation system for the spatiotemporal evolution of groundwater pollution.
[0007] Step B: Determine the input parameter set V = [v1, v2, ..., v] for the groundwater pollution spatiotemporal evolution prediction model. k Select the range of variation for each model input parameter to form the model input parameter range set U = [u1, u2, ..., u] k This section identifies all the parameters to be considered and their range of variation.
[0008] Step C: Select the change level of each model input parameter variable, obtain the basic typical scenario sampling of groundwater pollution spatiotemporal evolution under the given condition framework based on the orthogonal experimental method, and obtain the corresponding different model input parameter variable value combinations, i.e. parameter combinations, for each typical scenario;
[0009] Step D: Based on the numerical simulation system for the spatiotemporal evolution of groundwater pollution, obtain the simulated values of the associated feature variables corresponding to each basic typical scenario, and combine all parameter combinations with the corresponding simulated values of associated feature variables to form a basic training set for machine learning, which is also a basic dataset for statistical modeling.
[0010] Step E: Based on the preset number of scenario samples and input parameters, obtain a preset number of scenario samples to obtain a combination of data augmentation scenario parameters. Use the numerical simulation system to obtain the data augmentation scenario simulation value of the associated feature variable for each scenario. Use the combination of data augmentation scenario parameters and the set of data augmentation scenario simulation values as a machine learning augmentation dataset, which is also a statistical modeling augmentation dataset.
[0011] Step F: Use the basic training set of machine learning and the basic dataset of statistical modeling as the training set of machine learning and the basic dataset of statistical modeling, respectively. At the same time, use the enhanced dataset of machine learning and the enhanced dataset of statistical modeling as the test set of machine learning and the validation set of statistical modeling, respectively. Alternatively, use the set of basic training set of machine learning and the enhanced dataset of machine learning as the training set of machine learning, and the set of basic dataset of statistical modeling and the enhanced dataset of statistical modeling as the statistical modeling dataset. In this case, follow step E to obtain the enhanced dataset of machine learning, which is also the enhanced dataset of statistical modeling, and then obtain the test set of machine learning and the validation set of statistical modeling.
[0012] Step G: Based on the statistical modeling dataset, establish and optimize a multiple regression statistical prediction model for the associated feature variables and the model input parameters;
[0013] Step H: Obtain the prediction error value dataset of the associated feature variables through the multivariate regression statistical prediction model and the statistical modeling validation set; based on the prediction error value dataset, use statistical analysis methods to determine the upper and lower limits of the distribution interval of the prediction error random variable corresponding to the given preset confidence level.
[0014] Step 1: Apply a preset data processing method to the associated feature variables, the model input parameters, and each dataset;
[0015] Step J: Based on the simulated values of all associated feature variables corresponding to the obtained machine learning basic training set and machine learning basic test set, obtain their variation range [a,b]. Then, use M partition node values (boundary points) [y1,y2,…,y…] to represent this range. M Further divided into M+1 sub-intervals, namely: <y1、[y1,y2)、...、[y M-1 ,y M ), ≥y M ;
[0016] Step K: Based on the divided sub-intervals and the simulated values of the associated feature variables of the basic training set and the simulated values of the associated feature variables of the basic test set, obtain the output labels of each scenario of the basic training set and the corresponding output labels of each scenario of the basic test set;
[0017] Step L: Based on the machine learning basic training set and the machine learning basic test set, apply the ANN machine learning modeling principle and the optimized modeling method to obtain a preset number of optimized ANN prediction models. Based on the basic test set and the corresponding output labels of each scenario, obtain the test accuracy of the partition prediction model, and then obtain the optimized ANN prediction model with the highest possible test accuracy of the partition prediction model;
[0018] Step M: Based on the preset scenario sampling number and the model input parameter variables, obtain a preset number of scenario samplings, obtain the data enhancement scenario parameter combinations. Through the numerical simulation system, obtain the data enhancement scenario simulation values of the associated feature variables corresponding to each scenario. Based on the data enhancement scenario simulation values and the upper and lower limits of the distribution interval of the prediction error random variable corresponding to the given preset confidence level, determine the prediction confidence interval of the associated feature variables; Based on the model input parameter variables, obtain a preset number of scenario parameter combinations and the prediction confidence interval, and then merge them to obtain the statistical analysis enhanced modeling data set;
[0019] Step N: Based on the machine learning basic training set and the corresponding output labels of each scenario, the machine learning basic test set and the corresponding output labels of each scenario, and apply the statistical analysis enhanced data set and its corresponding output labels to supplement the two end intervals, including <y1 and ≥y M , obtain the SVM machine learning data set, apply the SVM machine learning modeling principle and the optimized modeling method, optimize the SVM partition prediction model, and obtain the accuracy of the optimized SVM partition prediction model;
[0020] Step O: Compare the accuracy of the ANN partition prediction model with the accuracy of the SVM partition prediction model, select the machine learning partition prediction model with higher accuracy, and determine whether the selected machine learning partition prediction model meets the preset accuracy target. If so, obtain the final machine learning partition prediction model that meets the model accuracy; If not, compare the accuracy of the ANN partition prediction model with the accuracy of the SVM partition prediction model, select the machine learning partition prediction model with higher accuracy, and determine whether the selected machine learning partition prediction model meets the preset accuracy target. If so, obtain the final machine learning partition prediction model that meets the model accuracy; If not, select any one or any combination of operations from increasing the basic typical scenario sampling number, the preset scenario sampling number, reducing the preset model accuracy, and reducing the partition node value, execute, and go back to Step E to Step N, complete the corresponding steps until the final machine learning partition prediction model that meets the model accuracy is obtained;
[0021] Step P: For a given specific application scenario, obtain the specific values of the input parameters of the specific model, determine that the specific conditions meet the requirements of the given typical condition framework, and obtain the reliable interval of the predicted values of the associated feature variables based on the obtained final machine learning partition prediction model that meets the accuracy of the model.
[0022] Furthermore, in step A, the conceptual model for determining the spatiotemporal evolution of groundwater pollution under typical conditions specifically includes: atmospheric precipitation infiltration recharge characteristics, vadose zone characteristics, saturated zone characteristics, pollution source characteristics, groundwater pollution prevention and remediation characteristics, water well source and sink characteristics, surface water and groundwater interaction characteristics, groundwater pollutant migration and transformation process types, boundary condition characteristics, and initial condition characteristics.
[0023] Further, in step A, the associated characteristic variables include: groundwater pollutant concentration, pollutant flux at the saturated zone cross section or interface, location of the pollutant plume migration and diffusion front, pollutant plume stabilization time, stable pollutant plume distribution area, distance of the stable pollutant plume front from the pollution source, average concentration of the stable pollutant plume, and pollutant plume dissipation time; as well as soil pollutant concentration, vadose zone gaseous pollutant concentration, vadose zone cross section pollutant flux, vadose zone-saturated zone pollutant interface flux, soil-gas interface pollutant flux, rock-gas interface pollutant flux, and spatiotemporal distribution of NAPL pollutants.
[0024] Furthermore, in step B, the set of parameters includes hydraulic conductivity coefficient, adsorption parameter, dispersion parameter, and degradation parameter.
[0025] Further, step E specifically includes: selecting the number of random samples, determining or selecting the applicable random distribution type and corresponding distribution parameters for each model input parameter, and completing the generation of random numbers for the model input parameters to form a new parameter combination scenario.
[0026] Furthermore, in step I, the preset data processing method includes: preserving the original value, increasing the value by the same multiple, decreasing the value by the same multiple, logarithmic transformation of the value, unifying the positive and negative values of the value, and normalizing the value, and any one or any combination of the methods is used for processing.
[0027] Furthermore, in step L, the application of ANN machine learning modeling principles and optimization modeling methods includes the optimization of ANN machine learning methods, the optimization of machine learning modeling search methods, and the optimization of modeling correlation parameters; the ANN machine learning methods include: BP neural networks and other machine learning methods that include ANN.
[0028] Furthermore, in step L, obtaining the optimized ANN prediction model with the highest possible test accuracy of the partition prediction model includes sorting the ANN prediction models from highest to lowest according to their test accuracy, and the first ANN prediction model that satisfies the ANN modeling optimization objective is the target model.
[0029] Furthermore, in step N, the optimization modeling method in the application of SVM machine learning modeling principles and optimization modeling methods includes: selection of SVM machine learning model modeling optimization method and selection of modeling kernel function type and its associated parameters.
[0030] Secondly, this invention provides a device for constructing a general prediction model for groundwater pollution based on ANN and SVM, comprising:
[0031] The typical condition framework construction module for the conceptual model of groundwater pollution is used to determine the conceptual model of the spatiotemporal evolution of groundwater pollution under typical conditions and to obtain the corresponding set of typical features, i.e., the typical condition framework.
[0032] The scenario construction and numerical simulation module obtains basic typical scenario samples of groundwater pollution spatiotemporal evolution under given conditions based on the orthogonal experimental method, generates different combinations of model input parameter values for each typical scenario, i.e., parameter combinations, and obtains simulated values of corresponding related feature variables for each scenario through the numerical simulation system; at the same time, based on the preset number of scenario samples and input parameters, a preset number of scenario samples are obtained to obtain a machine learning test set and a statistical modeling validation set.
[0033] The statistical model building and analysis module, based on the basic training set of machine learning and its output labels corresponding to each scenario, as well as the statistical analysis-enhanced dataset and its corresponding output labels, performs statistical model building and prediction error analysis.
[0034] The Machine Learning Model Building and Analysis Application Module is used to apply machine learning modeling principles and optimization methods to build and optimize machine learning prediction models, and provide prediction results for practical application scenarios.
[0035] Based on the above technical solution, the embodiments of the present invention can produce at least the following technical effects:
[0036] This invention provides a method for constructing a general prediction model for groundwater pollution based on ANN and SVM. It selects associated characteristic variables related to the spatiotemporal evolution of groundwater pollution and establishes and optimizes alternative prediction models for groundwater pollution associated characteristic variables based on numerical simulation and different machine learning algorithms for complex groundwater environmental systems under typical conditional frameworks. This method can accurately and efficiently predict the value range of soil and groundwater environmental states, evolution processes, and related impacts in different application scenarios under given conditional frameworks. It significantly improves the accuracy, efficiency, and practicality of assessing and predicting the spatiotemporal evolution of complex groundwater pollution and the spatiotemporal distribution of pollutants, providing strong support for the protection of soil and water environmental systems and the safe utilization of resources, especially for the rapid prediction, early warning, and reliable risk management of soil and groundwater pollution. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0038] Figure 1 This is a schematic diagram of the structure of an embodiment of the present invention;
[0039] Figure 2 These are the DMPS numerical simulation results from 54 representative scenarios in the embodiments of the present invention;
[0040] Figure 3 These are the T-DMPS numerical simulation results from 54 representative scenarios in the embodiments of the present invention. Detailed Implementation
[0041] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
[0042] This embodiment provides a method and apparatus for constructing a general prediction model for groundwater pollution based on ANN (MLB) and SVM. The following is a detailed description in conjunction with the appendix. Figure 1-3 They will be described together.
[0043] 1. Selection of related feature variables
[0044] First, a generalized analysis of the groundwater aquifer system was conducted based on the general characteristics of different typical study areas to determine the range of variation for each parameter. Orthogonal experimental design was used to set the parameters of each model at different levels. Fifty-four representative scenarios were selected according to the experimental design rules, and numerical simulations were performed using GMS software. For each simulation scenario, two characteristic factors of the spatiotemporal evolution of groundwater pollution under natural decay conditions were obtained: the plume stabilization distance (DMPS) and the stabilization time (T-DMPS) of the stable plume. After processing the 54 sets of orthogonal experimental numerical simulation results with the ln logarithm, a multiple regression model was statistically analyzed. Backpropagation (BP) and SVM machine learning algorithms were applied to establish and optimize a general predictive model. Combined with actual conditions, the specific values of the model input parameters were obtained. The established general predictive model was then used to predict the location of the groundwater plume stabilization time and migration distance.
[0045] 2. Generalization of the Spatiotemporal Evolution of Natural Decay of Groundwater Pollution
[0046] By referencing the on-site conditions of the study area, a conceptual model for the migration and transformation of groundwater pollution was established, such as... Figure 1 As shown, the groundwater environment system is generalized from top to bottom as follows: vadose zone, unconfined aquifer, weakly permeable layer, confined aquifer, and impermeable base. The study area has a length of L, a width of B, and a thickness of Hz. Appropriate grid division standards were adopted while meeting the accuracy requirements of numerical simulation calculations. It is assumed that all layers in the groundwater aquifer system are homogeneous and isotropic, with constant head boundaries set on the east and west sides, and a specific head difference between the boundaries. A constant pollution source exists in the unconfined aquifer, and the pollutants are easily degradable dissolved organic matter, represented by benzene series compounds. The organic pollutants in the groundwater decay naturally, and the migration and transformation processes of the pollutants include convection, hydrodynamic dispersion, adsorption, and degradation.
[0047] 3. Obtain the mathematical governing equations for the migration and transformation of groundwater pollutants, and obtain the spatiotemporal evolution data of groundwater pollution.
[0048] Value simulation system
[0049] (1) Establishment of the mathematical model (governing equation) for groundwater seepage:
[0050]
[0051] In the formula: Ω represents the seepage region; t represents time (T); x, y, z represent coordinates (L); K x K y K z Permeability parameters (LT) along the x, y, z directions -1 ); ω is the source / sink term, representing the volume of groundwater flowing out of or into a unit volume of aquifer per unit time (L).3 T -1 μ is the specific yield of the aquifer (dimensionless); H(x,y,z,t) is the groundwater level (L).
[0052] (2) Mathematical model (governing equation) for solute transport of organic pollutants:
[0053] Taking into account convection, dispersion, adsorption, and biodegradation, the mathematical governing equations for the transport of typical organic pollutant solutes are as follows:
[0054]
[0055] In the formula: C k The concentration (ML) of pollutants (BTEX) in groundwater -3 );q s The flow rate into a unit volume of aquifer per unit time is expressed as the source-sink term (T). -1 ); θ is porosity (dimensionless); D ij Hydrodynamic dispersion coefficient (L) 2 T -1 );ν i Groundwater seepage velocity (LT) -1 ).
[0056] 4. Determine the parameters and their range of variation, and design typical scenarios through orthogonal experiments.
[0057] The natural decay process of organic pollutants in groundwater involves numerous parameters. Scenario analysis is used to explore this process, simulating different results by changing the values of model parameters; each set of model values corresponds to a specific scenario. Orthogonal experiments, as a statistical method, are frequently used to solve multi-factor, multi-level experimental design problems. Their main idea is to understand the overall experimental situation by analyzing a representative subset of experimental results. Using regular orthogonal array design, representative level combinations are selected from all factor and level combinations, offering advantages such as high efficiency, speed, and cost-effectiveness.
[0058] This orthogonal experimental design selected 25 parameters that can influence the natural attenuation of groundwater pollution plumes. Due to the large number of parameters, according to the design rules of orthogonal experiments, an L54 (2^1, 3^24) orthogonal array was used to design the experimental scheme. Except for the effective porosity of weakly permeable water, which was set to 2 levels, all other parameters were set to 3 levels, generating a total of 54 orthogonal experimental groups corresponding to 54 simulation scenarios. The parameters were set according to their possible actual distribution range, from high to low. Table 1 lists the different value levels of the 25 parameters that may be involved in this groundwater pollution simulation model.
[0059] Table 1 shows different value levels of each parameter.
[0060]
[0061] Note: This model selected 25 parameters that may affect the numerical simulation experiment of groundwater pollution, including characteristic parameters related to hydrogeology, hydroculture, and pollution sources. ΔH h The difference in head between the upstream and downstream boundaries; ΔH v is the head difference between the unconfined aquifer and the confined aquifer; R is the precipitation recharge; Cs is the pollution source degradation coefficient; C0 is the pollution source concentration; A is the pollution source area; r is the ratio of the pollution source thickness to the unconfined aquifer thickness; P is the effective porosity; M is the average aquifer thickness; Kw is the permeability coefficient; Ka is the adsorption coefficient; Kd is the degradation coefficient; D is the dispersion.
[0062] 5. Obtain numerical simulation results for each typical scenario based on orthogonal experimental scenarios and numerical simulation system.
[0063] This study used GMS software to numerically simulate 54 representative groundwater pollution natural decay scenarios generated based on orthogonal experimental design results. Based on the simulation results, two types of data were selected as the spatiotemporal evolution correlation characteristics variables for each representative simulation scenario: the distance of maximum plume spreading (DMPS) and the time to reach the DMPS. DMPS was calculated by counting the number of grid points in the unconfined aquifer with a concentration exceeding 0.01 mg / L when the plume reached its maximum distance. T-DMPS represents the time required for the DMPS to reach its maximum distance. The simulation results for the 54 representative scenarios are summarized as follows: Figure 2 and Figure 3 As shown.
[0064] 6. Based on the range of parameter variation, set the scenario sampling number to 20, obtain 20 scenario samples of the model input parameter variables, obtain the data augmentation scenario parameter combination, and obtain the data augmentation scenario simulation value of the associated feature variable corresponding to each scenario through the numerical simulation system.
[0065] 7. Establish a multiple regression statistical prediction model for associated characteristic variables based on scenario numerical simulation results.
[0066] Based on 54 sets of orthogonal experimental scenario parameter combinations and 20 sets of data-enhanced sampling scenario parameter combinations, along with their numerical simulation results, the spatiotemporal evolution correlation characteristic variables DMPS and T-DMPS of the two pollution plumes were used as the model dependent variables (Y). Multivariate regression statistical models were established using full-factor model parameters (X). Multiple regression analysis was used to establish statistical prediction models for DMPS and T-DMPS, and these models were then used to predict the scenario values of the correlation characteristic variables of groundwater pollution spatiotemporal evolution under natural decay conditions.
[0067] After logarithmic processing of the data, the statistical model of the spatiotemporal evolution characteristic factors of groundwater pollution under natural decay conditions is as follows:
[0068] lnY=lnρ+α1lnΔH h +α2lnΔH v +α3lnP1+α4lnΔM1+…+α 25 lnr
[0069] After conversion, we get:
[0070]
[0071] Where Y represents the spatiotemporal evolution characteristics of groundwater pollution plumes, including DMPS and T-DMPS.
[0072] The DMPS (abbreviated as D) multiple regression statistical model with all factors of the model as independent variables X is expressed as follows:
[0073]
[0074] R 2 =0.939
[0075] The T-DMPS (abbreviated as T) multiple regression statistical model, with all factors of the model as independent variables X, is expressed as follows:
[0076]
[0077] R 2 =0.909
[0078] Based on the results of the comprehensive analysis of variance, a multivariate regression model with full factors was constructed, and its R-squared value was [value missing]. 2 The value reached above 0.9. This indicates a strong correlation between the spatiotemporal evolution correlation characteristic variables of groundwater pollution plumes under natural decay conditions and the various model parameters.
[0079] 8. Randomly sample to obtain the prediction error of the statistical model and determine the prediction confidence interval of the statistical prediction model.
[0080] Given 50 random samples, 50 scenario samples are obtained based on the model input parameters to obtain the data augmentation scenario parameter combination. The data augmentation scenario simulation value of the associated feature variable corresponding to each scenario is obtained through a numerical simulation system. Based on the data augmentation scenario simulation value, the model prediction value is statistically predicted to obtain the upper and lower limits of the random variable distribution interval of the prediction error corresponding to a 5% confidence level, and then the prediction confidence interval of the associated feature variable is determined.
[0081] 9. Random sampling to obtain SVM-based modeling and statistical analysis augmentation dataset
[0082] Given 1000 random samples, 1000 scenario samples are obtained based on the model input parameters, resulting in data augmentation scenario parameter combinations. Based on the statistical prediction model, the simulated data augmentation scenario values of the associated feature variables for each scenario are obtained. Based on the 1000 scenario parameter combinations and the predicted confidence intervals of the corresponding associated feature variables, the SVM modeling statistical analysis augmentation dataset is obtained.
[0083] 10. Construction of Groundwater Pollution Prediction Model Based on ANN and SVM
[0084] A prediction model was constructed using the Full Perception Layer (MLB) neural network in ANN, with 25 influencing factors in the natural decay of groundwater pollution as input parameters and DMPS and T-DMPS as output indicators.
[0085] For the construction of prediction models using ANN and SVM, different value ranges were defined for each associated feature variable of DMPS and T-DMPS, and different labels were assigned. Corresponding labels were then assigned to the ranges where the simulated values corresponded to different scenarios. Specifically, SVM modeling and statistical analysis were applied to both ends of the range to strengthen the dataset.
[0086] Both types of machine learning models have low prediction accuracy, which does not meet the requirement that the prediction accuracy must be above 0.8.
[0087] 11. The model's prediction accuracy is low. To improve this, random sampling should be used to obtain a dataset for strengthening machine learning modeling, and a corresponding machine learning prediction model should be established.
[0088] Machine learning models have low prediction accuracy, so it is necessary to increase the sampling scenarios.
[0089] Given 300 random samples, 300 scenario samples are obtained based on the model input parameters to obtain data augmentation scenario parameter combinations. The data augmentation scenario simulation values of the corresponding associated feature variables for each scenario are obtained through a numerical simulation system. Based on the 300 scenario parameter combinations and the corresponding associated feature variable simulation values, a machine learning modeling augmentation dataset is obtained.
[0090] After increasing the number of samples, the classification results of the SVM prediction model and the neural network prediction model are shown in Tables 2-5 below.
[0091] Table 2 SVM Model Training and Testing Results
[0092]
[0093] Table 3. SVM Model Training and Testing Results
[0094]
[0095] Table 4. SVM Model Training and Testing Results
[0096]
[0097] Table 5. Training and testing results of the MLB neural network model
[0098]
[0099]
[0100] 12. Further increase the random sampling to obtain a machine learning modeling reinforcement dataset (up to 600 sets), and establish corresponding machine learning prediction models.
[0101] The model was trained after the number of data sets was increased from 300 to 600.
[0102] After increasing the number of samples again, the classification results of the SVM prediction model and the neural network prediction model are shown in Tables 6-9 below.
[0103] Table 6. SVM Model Training and Testing Results
[0104]
[0105] Table 7. SVM Model Training and Testing Results
[0106]
[0107] Table 8. SVM Model Training and Testing Results
[0108]
[0109] Table 9. Training and Testing Results of the MLB Neural Network Model
[0110]
[0111]
[0112] 13. Further increase the random sampling to obtain a machine learning modeling reinforcement dataset (up to 1200 sets), establish corresponding machine learning prediction models, and optimize the machine learning prediction models.
[0113] The model was trained after the number of data sets was increased from 600 to 1200.
[0114] After increasing the number of samples again, the classification results of the SVM prediction model and the neural network prediction model are shown in Tables 10-13 below.
[0115] Table 10 SVM Model Training and Testing Results
[0116]
[0117] Table 11 SVM Model Training and Testing Results
[0118]
[0119] Table 12 SVM Model Training and Testing Results
[0120]
[0121] Table 13 MLB Neural Network Model Training and Testing Results
[0122]
[0123]
[0124] Based on the aforementioned research results, comparing the prediction accuracy of SVM-based and ANN-based models, the preferred machine learning prediction model is as follows: for plume stabilization time, the SVM-based prediction model is preferred; while for plume stabilization migration distance, the MLB neural network prediction model is preferred.
[0125] 14. Specific application scenarios
[0126] The specific application case is as follows: Taking a chemical industrial cluster in the Yellow River alluvial plain as a typical research area, a conceptual model capable of constructing different representative simulation scenarios is established. The average annual rainfall in this area is approximately 600 mm, and the average annual evaporation is approximately 1650 mm. The groundwater in the industrial park is mainly Class IV loose rock pore water, flowing from west to east. The overall terrain is flat, and the average annual groundwater level change does not exceed 2 meters. The pollutants exceeding standards in this area are widely distributed and diverse, mainly consisting of organic pollutants such as benzene and toluene.
[0127] Based on the given specific application scenario conditions, it is determined that the actual conditions are consistent with the given typical feature set, and the actual parameters fall within the variation range of each modeling parameter variable. According to the specific values of the input parameters of the specific model, the obtained preferred machine learning partition prediction model is applied to predict the specific partition where the relevant feature variables of the specific application scenario are located (the stable time interval is 10-15 years, and the stable distance is 500-600m), which may lead to the pollution of downstream sensitive protected receptors. Thus, a reliable prediction and assessment of the spatiotemporal evolution of the groundwater environment and its environmental impact is achieved.
[0128] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a general prediction model for groundwater pollution based on ANN and SVM, characterized in that, include: Step A: Determine the conceptual model of the spatiotemporal evolution of groundwater pollution under typical conditions and obtain the corresponding set of typical features, i.e., the typical condition framework; select the characteristic variables associated with the spatiotemporal evolution of groundwater pollution; obtain the mathematical control equations for the migration and transformation of groundwater pollutants; and obtain the numerical simulation system for the spatiotemporal evolution of groundwater pollution. Step B: Determine the input parameter set V = [v1, v2, ..., v] for the groundwater pollution spatiotemporal evolution prediction model. k Select the range of variation for each model input parameter to form the model input parameter range set U = [u1, u2, ..., u] k This section identifies all the parameters to be considered and their range of variation. Step C: Select the change level of each model input parameter variable, obtain the basic typical scenario sampling of groundwater pollution spatiotemporal evolution under the given condition framework based on the orthogonal experimental method, and obtain the corresponding different model input parameter variable value combinations, i.e. parameter combinations, for each typical scenario; Step D: Based on the numerical simulation system for the spatiotemporal evolution of groundwater pollution, obtain the simulated values of the associated feature variables corresponding to each basic typical scenario, and combine all parameter combinations with the corresponding simulated values of associated feature variables to form a basic training set for machine learning, which is also a basic dataset for statistical modeling. Step E: Based on the preset number of scenario samples and input parameters, obtain a preset number of scenario samples to obtain a combination of data augmentation scenario parameters. Use the numerical simulation system to obtain the data augmentation scenario simulation value of the associated feature variable for each scenario. Use the combination of data augmentation scenario parameters and the set of data augmentation scenario simulation values as a machine learning augmentation dataset, which is also a statistical modeling augmentation dataset. Step F: Use the basic training set of machine learning and the basic dataset of statistical modeling as the training set of machine learning and the basic dataset of statistical modeling, respectively. At the same time, use the enhanced dataset of machine learning and the enhanced dataset of statistical modeling as the test set of machine learning and the validation set of statistical modeling, respectively. Alternatively, use the set of basic training set of machine learning and the enhanced dataset of machine learning as the training set of machine learning, and the set of basic dataset of statistical modeling and the enhanced dataset of statistical modeling as the statistical modeling dataset. In this case, follow step E to obtain the enhanced dataset of machine learning, which is also the enhanced dataset of statistical modeling, and then obtain the test set of machine learning and the validation set of statistical modeling. Step G: Based on the statistical modeling dataset, establish and optimize a multiple regression statistical prediction model for the associated feature variables and the model input parameters; Step H: Obtain the prediction error value dataset of the associated feature variables through the multivariate regression statistical prediction model and the statistical modeling validation set; based on the prediction error value dataset, use statistical analysis methods to determine the upper and lower limits of the distribution interval of the prediction error random variable corresponding to the given preset confidence level. Step 1: Apply a preset data processing method to the associated feature variables, the model input parameters, and each dataset; Step J: Based on the simulated values of all related feature variables corresponding to the obtained machine learning basic training set and machine learning basic test set, obtain their variation range [a,b]. Then, use M partitioned node value sets [y1,y2,…,y…] to represent this range. M Further divided into M+1 sub-intervals, namely: <y1、[y1,y2)、...、[y M-1 ,y M ), ≥y M ; Step K: Based on the divided sub-intervals and the simulated values of the associated feature variables in the basic training set and the simulated values of the associated feature variables in the basic test set, obtain the output labels for each scenario in the basic training set and the corresponding output labels for each scenario in the basic test set. Step L: Based on the basic training set and basic test set of machine learning, apply the ANN machine learning modeling principle and optimization modeling method to obtain a preset number of optimized ANN prediction models. Based on the basic test set and the output labels corresponding to each scenario, obtain the test accuracy of the partition prediction model, and then obtain an optimized ANN prediction model with the highest possible test accuracy of the partition prediction model. Step M: Based on the preset number of scenario samples and the model input parameters, obtain a preset number of scenario samples to obtain a combination of data augmentation scenario parameters. Use the numerical simulation system to obtain the data augmentation scenario simulation value of the associated feature variable corresponding to each scenario. Determine the prediction confidence interval of the associated feature variable based on the data augmentation scenario simulation value and the upper and lower limits of the distribution interval of the random variable of the prediction error corresponding to the given preset confidence level. Based on the input parameters of the model, a preset number of scenario parameter combinations and the prediction confidence interval are obtained, and then merged to obtain a statistical analysis enhanced modeling dataset. Step N: Based on the machine learning basic training set and its corresponding output labels for each scenario, the machine learning basic test set and its corresponding output labels for each scenario, and the application statistical analysis reinforcement data set and its corresponding output labels, supplement the two end intervals, including <y1 and ≥y M , obtain the SVM machine learning data set, apply the SVM machine learning modeling principle and the optimization modeling method, optimize the SVM partition prediction model, and obtain the accuracy of the optimized SVM partition prediction model; Step O: Compare the accuracy of the ANN partition prediction model with that of the SVM partition prediction model, select the machine learning partition prediction model with higher accuracy, and determine whether the selected machine learning partition prediction model meets the preset accuracy target. If yes, obtain the final machine learning partition prediction model that meets the model accuracy target; if not, select any one or any combination of operations from increasing the number of basic typical scenario samples, the number of preset scenario samples, reducing the preset model accuracy, and reducing the partition node values, and return to step E, then to step N, to complete the corresponding steps until the final machine learning partition prediction model that meets the model accuracy target is obtained. Step P: For a given specific application scenario, obtain the specific values of the input parameters of the specific model, determine that the specific conditions meet the requirements of the given typical condition framework, and obtain the reliable interval of the predicted values of the associated feature variables based on the obtained final machine learning partition prediction model that meets the accuracy of the model.
2. The construction method according to claim 1, characterized in that, In step A, the conceptual model for determining the spatiotemporal evolution of groundwater pollution under typical conditions specifically includes: atmospheric precipitation infiltration recharge characteristics, vadose zone characteristics, saturated zone characteristics, pollution source characteristics, groundwater pollution prevention and remediation characteristics, water well source and sink characteristics, surface water and groundwater interaction characteristics, groundwater pollutant migration and transformation process types, boundary condition characteristics, and initial condition characteristics.
3. The construction method according to claim 1, characterized in that, In step A, the associated characteristic variables include: groundwater pollutant concentration, pollutant flux at the saturated zone cross section or interface, location of the pollutant plume migration and diffusion front, pollutant plume stabilization time, stable pollutant plume distribution area, distance of the stable pollutant plume front from the pollution source, average concentration of the stable pollutant plume, and pollutant plume dissipation time; as well as soil pollutant concentration, vadose zone gaseous pollutant concentration, vadose zone cross section pollutant flux, vadose zone-saturated zone pollutant interface flux, soil-gas interface pollutant flux, rock-gas interface pollutant flux, and spatiotemporal distribution of NAPL pollutants.
4. The construction method according to claim 1, characterized in that, Step E specifically includes: selecting the number of random samples, determining or selecting the applicable random distribution type and corresponding distribution parameters for each model input parameter, and completing the generation of random numbers for the model input parameters to form a new parameter combination scenario.
5. The construction method according to claim 1, characterized in that, In step I, the preset data processing method includes: keeping the value as is, increasing the value by the same amount, decreasing the value by the same amount, logarithmic transformation of the value, unifying the positive and negative values of the value, and normalizing the value, and any one or any combination of the methods is used for processing.
6. The construction method according to claim 1, characterized in that, In step L, the application of ANN machine learning modeling principles and optimization modeling methods includes the optimization of ANN machine learning methods, the optimization of machine learning modeling search methods, and the optimization of modeling correlation parameters. The machine learning methods mentioned include: BP neural networks and other machine learning methods that incorporate ANN.
7. The construction method according to claim 1, characterized in that, In step L, obtaining the optimized ANN prediction model with the highest possible test accuracy of the partition prediction model includes sorting the ANN prediction models from highest to lowest according to their test accuracy, and the first ANN prediction model that meets the ANN modeling optimization objective is the target model.
8. The construction method according to claim 1, characterized in that, In step N, the optimization modeling method in the application of SVM machine learning modeling principles and optimization modeling methods includes: selection of SVM machine learning model modeling optimization method and selection of modeling kernel function type and its associated parameters.
9. A device for constructing a general prediction model for groundwater pollution based on ANN and SVM, used to implement the construction method described in any one of claims 1-8, characterized in that, include: The typical condition framework construction module for the conceptual model of groundwater pollution is used to determine the conceptual model of the spatiotemporal evolution of groundwater pollution under typical conditions and to obtain the corresponding set of typical features, i.e., the typical condition framework. The scenario construction and numerical simulation module obtains basic typical scenario samples of groundwater pollution spatiotemporal evolution under given conditions based on the orthogonal experimental method, generates different combinations of model input parameter values for each typical scenario, i.e., parameter combinations, and obtains simulated values of corresponding related feature variables for each scenario through the numerical simulation system; at the same time, based on the preset number of scenario samples and input parameters, a preset number of scenario samples are obtained to obtain a machine learning test set and a statistical modeling validation set. The statistical model building and analysis module, based on the basic training set of machine learning and its output labels corresponding to each scenario, as well as the statistical analysis-enhanced dataset and its corresponding output labels, performs statistical model building and prediction error analysis. The Machine Learning Model Building and Analysis Application Module is used to apply machine learning modeling principles and optimization methods to build and optimize machine learning prediction models, and provide prediction results for practical application scenarios.
Citation Information
Patent Citations
Organic pollutant migration numerical model substitution method based on multi-core extreme learning machine
CN114492164A
Data preprocessing and storing method
CN114996769A