A water pollution source tracing method based on artificial intelligence
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI
- Filing Date
- 2024-08-30
- Publication Date
- 2026-05-26
Smart Images

Figure CN119089123B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water quality monitoring technology, and in particular to a method for tracing the source of water pollution based on artificial intelligence. Background Technology
[0002] Sudden water pollution incidents are among the most severe environmental pollution incidents, characterized by their diversity, suddenness, severity, difficulty in handling, randomness, and uncertainty. To minimize the losses from sudden water pollution incidents, it is necessary to quickly locate and identify the leakage source after the incident, simulate the spatial impact range and extent of pollutants after the incident, and assist in taking control measures.
[0003] Currently, water pollution source tracing research methods generally employ either on-site sampling and measurement or mathematical modeling. On-site sampling primarily utilizes source tracing technologies such as isotope tracing, water ripple identification, and ultraviolet spectroscopy to study pollutant sources. While these methods offer high stability and accuracy, they are labor-intensive, time-consuming, and difficult to use for timely pollution source identification, thus hindering timely and effective control of pollution incidents. Mathematical modeling, on the other hand, offers advantages such as flexibility, speed, and strong operability. It helps decision-makers understand the migration, diffusion, and temporal changes of pollutants in the aquatic environment, grasp the impact of pollutants on the aquatic environment, and thus make timely and accurate responses to emergencies. Summary of the Invention
[0004] This invention aims to address at least one of the technical problems existing in related technologies. Due to the covert and sudden nature of water pollution incidents, it is often impossible to know in advance the spatial location and intensity of the leakage source, making it urgent to develop methods for rapidly locating and identifying leakage sources. Therefore, this invention provides a water pollution source tracing method based on artificial intelligence.
[0005] This invention provides an artificial intelligence-based method for tracing the source of water pollution, comprising:
[0006] S1: Select a basic model to simulate water pollution events under various emission conditions and obtain a dataset of water quality diffusion relationships;
[0007] S2: Based on the aforementioned water quality diffusion relationship dataset, establish the first PCE model;
[0008] S3: Optimize the first PCE model based on the basic model to obtain the second PCE model;
[0009] S4: The water quality diffusion relationship dataset is inverted using the second PCE model, and the second PCE model is optimized based on the inversion results to obtain a third PCE model;
[0010] S5: Use the third PCE model as the source tracing model to trace the source of the water pollution event to be traced.
[0011] According to the water pollution source tracing method based on artificial intelligence provided by the present invention, the basic model selected in step S1 is the Delft3D model.
[0012] According to the water pollution source tracing method based on artificial intelligence provided by the present invention, step S1, which involves simulating water pollution events under multiple emission conditions, further includes:
[0013] S11: Define the simulation region, and divide the simulation region into multiple simulation meshes;
[0014] S12: Set simulation parameters and simulation variables, sample the simulation parameters and select the grid of the downstream region of the simulation area as the observation point location to obtain multiple emission conditions;
[0015] S13: Based on the distributed computing framework that integrates the basic model, various emission conditions are substituted to simulate water pollution events and obtain a water quality diffusion relationship dataset.
[0016] According to the artificial intelligence-based water pollution source tracing method provided by the present invention, the simulation parameters in step S12 include simulation duration, calculation time step, gravitational acceleration, and water density, wherein the gravitational acceleration and the water density are the default values of the basic model;
[0017] The simulation variables in step S12 include the start time, end time, and amount of water pollution event discharge.
[0018] According to the water pollution source tracing method based on artificial intelligence provided by the present invention, step S13 further includes:
[0019] S131: Select Hadoop MapReduce as the distributed computing framework, integrate the basic model, and build a simulated cluster computing system.
[0020] S132: Using the aforementioned simulation cluster computing system, water pollution events under various emission conditions are simulated to obtain simulation results;
[0021] S133: Obtain the raw binary data of the simulation results, and parse and store the raw binary data in a columnar database;
[0022] S134: Extract the record set of the aforementioned database using Spark in-memory grid technology to obtain a water quality diffusion relationship dataset.
[0023] According to the artificial intelligence-based water pollution source tracing method provided by the present invention, the expression of the control equation of the basic model in step S1 is as follows:
[0024]
[0025]
[0026]
[0027] in, For the water level at the observation point, for Water flow in the direction, for Water flow in the direction, For the flow rate at the observation point, for Flow velocity in direction for Flow velocity in direction The density of observation points, The Coriolis force coefficient, Bottom of the water body Shear stress in the direction, Bottom of the water body Shear stress in the direction, The horizontal viscosity coefficient, This refers to the water pressure.
[0028] According to the artificial intelligence-based water pollution source tracing method provided by the present invention, the expression of the diffusion equation of the basic model in step S1 is as follows:
[0029]
[0030] in, The concentration of pollutants, for Diffusion coefficient in the direction, for Diffusion coefficient in the direction, For source and sink items, This is the reaction term.
[0031] According to the water pollution source tracing method based on artificial intelligence provided by the present invention, step S2 further includes:
[0032] S21: Using the placement point method in the surrogate modeling cardinality of the PCE algorithm, fit the data in the water quality diffusion relationship dataset to obtain the fitted dataset;
[0033] S22: Input the fitted dataset into the preset model and train the model using the chaospy software library based on the Python programming language to obtain the first PCE model.
[0034] According to the water pollution source tracing method based on artificial intelligence provided by the present invention, step S3 further includes:
[0035] S31: Perform MC sampling on the simulation parameters to obtain the sampling parameters;
[0036] S32: Input the sampling parameters into the basic model for simulation to obtain the sampling dataset;
[0037] S33: Input the sampling parameters into the first PCE model to obtain the first PCE dataset;
[0038] S34: Based on the sampled dataset and the first PCE dataset, perform multi-objective optimization on the first PCE model to obtain the second PCE model.
[0039] According to the water pollution source tracing method based on artificial intelligence provided by the present invention, when performing multi-objective optimization on the first PCE model, the selected objectives include:
[0040] The accuracy index, namely the Nash efficiency coefficient;
[0041] The efficiency index is the time consumed in the simulation calculation;
[0042] The fidelity index is the Sobol parameter sensitivity index.
[0043] According to the water pollution source tracing method based on artificial intelligence provided by the present invention, step S4 further includes:
[0044] S41: Randomly select an inversion dataset from the sampled dataset;
[0045] S42: Use the output dataset in the inversion dataset as the observation value, input the observation value into the second PCE model, and obtain the inversion result;
[0046] S43: Using the input dataset in the inversion dataset as the true value, compare the inversion result with the true value, and optimize the second PCE model with the Nash efficiency coefficient as the objective function to obtain the third PCE model.
[0047] This invention provides an artificial intelligence-based method for tracing water pollution sources, aiming to infer the location of pollution sources, the timing and intensity of pollutant emissions, and to provide a scientific method for intelligent water environment management and decision-making. This invention utilizes the Delft3D model, Hadoop big data cluster computing, artificial neural networks, multi-objective Bayesian optimization, and Monte Carlo simulation (MC) technology. It employs surrogate modeling methods to approximate and replace the predictive capabilities of complex water quality models and supports simulation optimization. This provides a new method for solving the problem of reverse pollution source inversion. It not only has advantages such as flexibility, speed, and strong operability, but also enables the understanding of the impact of pollutants on the water environment, thereby allowing for timely and accurate responses and improving the accuracy of identifying the location sets of concealed and sudden water pollution events.
[0048] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0050] Figure 1 This is a schematic diagram of a water pollution source tracing method based on artificial intelligence provided in an embodiment of the present invention;
[0051] Figure 2 This is a schematic diagram of the method for simulating water pollution events under various emission conditions provided in an embodiment of the present invention;
[0052] Figure 3 This is a schematic diagram of the method for obtaining the first PCE model provided in an embodiment of the present invention;
[0053] Figure 4 This is a schematic diagram of the method for obtaining the second PCE model provided in an embodiment of the present invention;
[0054] Figure 5 This is a schematic diagram of the method for obtaining the third PCE model provided in an embodiment of the present invention. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. The following embodiments are used to illustrate this invention but should not be used to limit the scope of this invention.
[0056] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0057] The following is combined Figures 1 to 5 Description of embodiments of the present invention.
[0058] like Figure 1 As shown, this invention provides a water pollution source tracing method based on artificial intelligence, comprising:
[0059] S1: Select a basic model to simulate water pollution events under various emission conditions and obtain a dataset of water quality diffusion relationships.
[0060] In step S1, a hydrodynamic and water quality model for the study area is first constructed. Scenario simulations under different emission conditions are conducted through experimental design to obtain a large training dataset that can characterize the time-series response relationship between pollution source emission patterns and pollutant concentrations. This dataset inherently reflects the mechanistic model process.
[0061] In step S1, the base model selected is the Delft3D model.
[0062] Furthermore, the selected Delft3D model is a mature, advanced, and open-source hydrodynamic and water quality model. As the basic model in this step, it has the characteristics of fast calculation speed and high simulation accuracy. In particular, it is very suitable for boundary edge processing for curved surface mesh generation technology.
[0063] like Figure 2 As shown, step S1, which involves simulating water pollution events under various emission conditions, further includes:
[0064] S11: Define the simulation region, and divide the simulation region into multiple simulation meshes;
[0065] Step S11 aims to define the regional scope for the study. Based on factors such as typicality, representativeness, and data availability, the spatial modeling area for hydrodynamic water quality simulation is selected to ensure that the research on water quality prediction and pollution source inversion in this area is representative. The selected study area should have the characteristics of having a concentration of industrial enterprises within its scope, a high risk of sudden pollution accidents, and an urgent need for pollution source inversion.
[0066] Specifically, in step S11, the RGFGRID software is first used to perform planar two-dimensional orthogonal quadrilateral meshing on the study area, the QUICKIN software is used to interpolate the subsurface data of the study area, and then the QUICKPLOT software is used to perform post-processing operations such as visualization of the simulation results.
[0067] S12: Set simulation parameters and simulation variables, sample the simulation parameters and select the grid of the downstream region of the simulation area as the observation point location to obtain multiple emission conditions;
[0068] In addition, to highlight the feasibility of the method, the boundary conditions and model setting conditions are generalized. In the early stage of implementation, the actual pollution discharge along the route is not considered. Instead, a simulation test method is adopted to select the migration and diffusion process of conservative substances to replace the water quality process.
[0069] The simulation parameters in step S12 include simulation duration, calculation time step, gravitational acceleration, and water density, wherein the gravitational acceleration and water density are the default values of the basic model.
[0070] The simulation variables in step S12 include the start time, end time, and amount of water pollution event discharge.
[0071] When selecting simulation parameters and variables, considering the urgency of the water pollution incident and to conform to the actual situation, the simulation duration was set to 20 days, the calculation time step was generally set to 5 minutes, and other parameters, such as gravitational acceleration and water density, adopted the model default values.
[0072] Then, the grid covering a quarter of the upstream region of the entire study area was set as the optional emission point location. Along with different emission start times, end times, and different emission amounts as parameter variables, the Latin hypercube sampling method was used to sample the above four parameters. Finally, the grids at three equally spaced center locations within a quarter of the downstream region of the study area were selected as the observation point locations.
[0073] In order to simplify the complex water quality change process into a conservative material migration and diffusion process, after step S12, different emission scenarios were established by sampling using the above parameters.
[0074] S13: Based on the distributed computing framework that integrates the basic model, various emission conditions are substituted to simulate water pollution events and obtain a water quality diffusion relationship dataset.
[0075] Step S13 further includes:
[0076] S131: Select Hadoop MapReduce as the distributed computing framework, integrate the basic model, and build a simulated cluster computing system.
[0077] In step S131, a water environment simulation cluster computing system is first built based on the Hadoop MapReduce distributed computing framework and by integrating the Delft3D model computing engine. The hardware configuration performance and the roles of the computing nodes are shown in Table 1.
[0078] Table 1 Hardware and software environment of the water environment simulation cluster computing system
[0079] Machine name Hardware configuration Roles in the cluster Master HP ProLiant DL580 g7, 32 cores, 2.67 GHz, 64 GB RAM, 900 GB hard drive. NameNode, JobTracker Master2 HP ProLiant DL580 g7, 32 cores, 2.67 GHz, 64GB RAM, 900 GB hard drive. Secondary NameNode s201~s205 ThinkCentre M8411t, 8 cores, 3.4 GHz, 32 GB RAM, 2 TB hard drive. DataNode, TaskTracker s206~s212 HP ProLiant DL580 g7, 32 cores, 2.67 GHz, 64GB RAM, 2TB HDD DataNode, TaskTracker
[0080] S132: Using the aforementioned simulation cluster computing system, water pollution events under various emission conditions are simulated to obtain simulation results;
[0081] S133: Obtain the raw binary data of the simulation results, and parse and store the raw binary data in a columnar database;
[0082] S134: Extract the record set of the aforementioned database using Spark in-memory grid technology to obtain a water quality diffusion relationship dataset.
[0083] After step S131, the emission scenario characterized by the aforementioned Latin hypercube sampling is simulated using a cluster computing system. Then, the raw binary data of the simulation results is stored using the Hadoop HDFS file system, and the parsed records of the simulation results are stored using an HBase columnar database. Finally, the record set in HBase is quickly extracted using Spark memory grid technology to generate a "four-parameter" input-"three-observation-point" concentration time series output dataset. The four parameters are the simulation parameters in step S12, and the three observation points are the simulation variables in step S12. The generation of this dataset originates from the mechanistic model, and therefore inherently characterizes the mechanistic process of pollutant emission input-concentration output response in the study area.
[0084] The governing equations of the basic model in step S1 are expressed as follows:
[0085]
[0086]
[0087]
[0088] in, For the water level at the observation point, for Water flow in the direction, for Water flow in the direction, For the flow rate at the observation point, for Flow velocity in direction for Flow velocity in direction The density of observation points, The Coriolis force coefficient, Bottom of the water body Shear stress in the direction, Bottom of the water body Shear stress in the direction, The horizontal viscosity coefficient, This refers to the water pressure.
[0089] The governing equations described above are the two-dimensional Navier-Stokes equations used in the Delft3D two-dimensional hydrodynamic model.
[0090] The expression for the diffusion equation of the basic model in step S1 is as follows:
[0091]
[0092] in, The concentration of pollutants, for Diffusion coefficient in the direction, for Diffusion coefficient in the direction, For source and sink items, This is the reaction term.
[0093] Furthermore, the two-dimensional hydrodynamic and water quality model uses orthogonal curve meshes for mesh generation. Numerical solutions to multidimensional problems based on mesh generation have higher requirements for stability. To obtain a stable scheme, implicit schemes are more advantageous than explicit schemes.
[0094] To reduce the bandwidth of the coefficient matrix in the implicit scheme for solving the two-dimensional equation system, this invention employs the Alternating Direction Implicit (ADI) method to discretely solve the above-mentioned governing equation system. Except for the vertical viscous term and the convection term, which are solved using the central difference method, the rest are solved using the ADI method, i.e., the alternating direction implicit difference method.
[0095] S2: Based on the aforementioned water quality diffusion relationship dataset, establish the first PCE model.
[0096] Furthermore, the first PCE model is based on the chaotic polynomial expansion algorithm, a statistical method that uses normally distributed random inputs to describe the uncertainty of the system.
[0097] like Figure 3 As shown, step S2 further includes:
[0098] S21: Using the placement point method in the surrogate modeling cardinality of the PCE algorithm, fit the data in the water quality diffusion relationship dataset to obtain the fitted dataset;
[0099] In step S21, based on the “four-parameter” input and “three-observation-point” concentration time series output dataset obtained in step S1, the placement point method in PCE surrogate modeling technology is used to fit the dataset to establish a PCE model. The reason for choosing the placement point method is that this method has no special requirements on the probability distribution of the parameters and is widely used because of its strong applicability.
[0100] S22: Input the fitted dataset into the preset model and train the model using the chaospy software library based on the Python programming language to obtain the first PCE model.
[0101] Prepare the dataset for training the PCE model, then perform data cleaning, checking for missing values, outliers, etc., and handling them accordingly. Subsequently, normalize or standardize the data as needed to ensure all input variables have similar scales. Next, define random variables corresponding to your data in chaospy, including uniform distribution, normal distribution, etc., and then choose an appropriate multinomial basis based on the problem requirements and data characteristics. Then, use functions in the chaospy library to associate the constructed dataset with the random variables and multinomial basis. Finally, use the prepared data and the defined PCE model to train it by calling the corresponding functions or methods. During or after training, you will obtain the first PCE model described above.
[0102] S3: Optimize the approximation capability of the first PCE model based on the basic model to obtain the second PCE model.
[0103] like Figure 4 As shown, step S3 further includes:
[0104] S31: Perform MC sampling on the simulation parameters to obtain the sampling parameters;
[0105] S32: Input the sampling parameters into the basic model for simulation to obtain the sampling dataset;
[0106] S33: Input the sampling parameters into the first PCE model to obtain the first PCE dataset;
[0107] S34: Based on the sampled dataset and the first PCE dataset, perform multi-objective optimization on the first PCE model to obtain the second PCE model.
[0108] The specific process in step S3 is as follows: MC sampling is performed on the four parameters of the Delft3D model. Based on the scale of the cluster computing system, the MC "four parameters" input and "three observation points" concentration time series output dataset are obtained. Then, based on the constructed PCE model, the above MC sampling parameters are input to obtain the PCE "four parameters" input and "three observation points" concentration time series output dataset. Finally, based on the above two sets of time series datasets, the approximation ability of the PCE model is evaluated using accuracy, efficiency, and fidelity indices, respectively.
[0109] When performing multi-objective optimization on the first PCE model, the selected objectives include:
[0110] The accuracy index, namely the Nash efficiency coefficient;
[0111] The efficiency index is the time consumed in the simulation calculation;
[0112] The fidelity index is the Sobol parameter sensitivity index.
[0113] Furthermore, the Nash efficiency coefficient mentioned above is an indicator used to quantify the predictive accuracy of simulation models (such as hydrological models). It compares the degree of agreement between simulation results and observational data to evaluate the accuracy of the model. Simulation computation time refers to the total time required to complete one simulation computation. It is an important indicator for measuring the computational efficiency of the model, reflecting the model's speed and efficiency under given computing resources (such as processor speed, memory size, etc.). The Sobol parameter sensitivity index is a global sensitivity analysis tool used to quantify the contribution of input variables to the variability of model output. By analyzing the degree of influence of input variables on model output, it evaluates the model's fidelity (i.e., the degree to which the model reflects real-world phenomena).
[0114] S4: Invert the water quality diffusion relationship dataset using the second PCE model, and optimize the second PCE model based on the inversion results to obtain a third PCE model.
[0115] like Figure 5 As shown, step S4 further includes:
[0116] S41: Randomly select an inversion dataset from the sampled dataset;
[0117] S42: Use the output dataset in the inversion dataset as the observation value, input the observation value into the second PCE model, and obtain the inversion result;
[0118] S43: Using the input dataset in the inversion dataset as the true value, compare the inversion result with the true value, and optimize the second PCE model with the Nash efficiency coefficient as the objective function to obtain the third PCE model.
[0119] In step S4, similar to step S31 in step S3, firstly, MC sampling is performed on the four parameters of the Delft3D model. Based on the scale of the cluster computing system, the MC "four parameters" input and "three observation points" concentration time series output dataset are obtained. Secondly, one set of datasets is randomly selected, and its "three observation points" concentration time series dataset is used as the observed value, and its "four parameters" input dataset is used as the "true value". Using the Nash efficiency coefficient as the objective function, combined with the PCE model and multi-objective Bayesian optimization algorithm, the "four parameters" input values such as emission point, start time, end time, and emission amount are inverted. Then, the multi-objective optimization algorithm is repeated multiple times to obtain the probability distribution of the four parameters. By comparing with the "true value", the ability of the PCE model to support simulation optimization is verified, and the model is optimized based on the results.
[0120] S5: Use the third PCE model as the source tracing model to trace the source of the water pollution event to be traced.
[0121] After steps S1 to S4, a third PCE model can be obtained after multiple optimizations and reliability verification through pollution source inversion. This model can then be used to trace the source of water pollution events and obtain data such as the start time, end time, amount, and location of the emission source.
[0122] This invention provides an artificial intelligence-based method for tracing water pollution sources, shifting the traditional model-driven perspective to a data-driven perspective. It establishes a complex mechanistic model that significantly reduces computational load while maintaining accuracy, and powerfully supports uncertainty analysis and simulation optimization of complex models. The regulation and optimization of complex numerical models has long been a challenge faced by multiple disciplines and is an urgent need in practical applications. This invention integrates numerical simulation (Delft3D), big data cluster computing (Hadoop), artificial intelligence (artificial neural networks, probabilistic surrogate models), multi-objective optimization (multi-objective Bayesian optimization), and Monte Carlo simulation (MC) techniques, fully leveraging the advantages of multidisciplinary collaboration. It employs surrogate modeling methods to approximate and replace the predictive capabilities of complex water quality models and support simulation optimization, providing a new method for solving inverse pollution source inversion. With limited computational resources and a limited number of model iterations, it can perform complex pollution source inversion, replacing existing complex water quality models, improving the accuracy and efficiency of pollution inversion, and reducing computational costs.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An artificial intelligence-based water pollution tracing method, characterized in that, include: S1: Select a basic model to simulate water pollution events under various emission conditions and obtain a water quality diffusion relationship dataset; wherein, the basic model is the Delft3D model; Specifically, step S1, which involves simulating water pollution events under various emission conditions, further includes: S11: Define the simulation region, and divide the simulation region into multiple simulation meshes; S12: Set simulation parameters and simulation variables, sample the simulation parameters and select the grid of the downstream region of the simulation area as the observation point location to obtain multiple emission conditions; S13: Based on the distributed computing framework that integrates the basic model, various emission conditions are substituted to simulate water pollution events and obtain a water quality diffusion relationship dataset; S2: Based on the aforementioned water quality diffusion relationship dataset, establish the first PCE model; Step S2 further includes: S21: Using the placement point method in the surrogate modeling cardinality of the PCE algorithm, fit the data in the water quality diffusion relationship dataset to obtain the fitted dataset; S22: Input the fitted dataset into the preset model and train the model using the chaospy software library based on the Python programming language to obtain the first PCE model; S3: Optimize the first PCE model based on the basic model to obtain the second PCE model; Step S3 further includes: S31: Perform MC sampling on the simulation parameters to obtain the sampling parameters; S32: Input the sampling parameters into the basic model for simulation to obtain the sampling dataset; S33: Input the sampling parameters into the first PCE model to obtain the first PCE dataset; S34: Based on the sampled dataset and the first PCE dataset, perform multi-objective optimization on the first PCE model to obtain the second PCE model; S4: The water quality diffusion relationship dataset is inverted using the second PCE model, and the second PCE model is optimized based on the inversion results to obtain a third PCE model; Step S4 further includes: S41: Randomly select an inversion dataset from the sampled dataset; S42: Use the output dataset in the inversion dataset as the observation value, input the observation value into the second PCE model, and obtain the inversion result; S43: Using the input dataset in the inversion dataset as the true value, compare the inversion result with the true value, and optimize the second PCE model with the Nash efficiency coefficient as the objective function to obtain the third PCE model; S5: Use the third PCE model as the source tracing model to trace the source of the water pollution event to be traced.
2. The method of claim 1, wherein, The simulation parameters in step S12 include simulation duration, calculation time step, gravitational acceleration, and water density, wherein the gravitational acceleration and water density are the default values of the basic model; The simulation variables in step S12 include the start time, end time, and amount of water pollution event discharge.
3. The method of claim 1, wherein, Step S13 further includes: S131: Select Hadoop MapReduce as the distributed computing framework, integrate the basic model, and build a simulated cluster computing system. S132: Using the aforementioned simulation cluster computing system, water pollution events under various emission conditions are simulated to obtain simulation results; S133: Obtain the raw binary data of the simulation results, and parse and store the raw binary data in a columnar database; S134: Extract the record set of the aforementioned database using Spark in-memory grid technology to obtain a water quality diffusion relationship dataset.
4. The method of claim 1, wherein, The expression for the governing equations of the basic model in step S1 is: wherein, is the water level at the observation point, is the water flow in the direction, is the water flow in the direction, is the flow rate at the observation point, is the flow velocity in the direction, is the flow velocity in the direction, is the density at the observation point, is the Coriolis force coefficient, is the shear stress at the bottom of the water body in the direction, is the shear stress at the bottom of the water body in the direction, is the horizontal viscosity coefficient, is the water pressure.
5. The water pollution source tracing method based on artificial intelligence according to claim 1, characterized in that, The expression for the diffusion equation of the basic model in step S1 is as follows: in, The concentration of pollutants, for Diffusion coefficient in the direction, for Diffusion coefficient in the direction, For source and sink items, This is the reaction term.
6. The water pollution source tracing method based on artificial intelligence according to claim 1, characterized in that, When performing multi-objective optimization on the first PCE model, the selected objectives include: The accuracy index, namely the Nash efficiency coefficient; The efficiency index is the time consumed in the simulation calculation; The fidelity index is the Sobol parameter sensitivity index.