Hazardous waste landfill underground water monitoring project optimization method and system based on artificial intelligence
By classifying the data of groundwater monitoring wells numerical codes and building artificial intelligence models, the dynamic optimization of monitoring indicators in groundwater monitoring of hazardous waste landfills has been solved, cost reduction and efficiency improvement have been achieved, and the implementability of the long-term monitoring system has been improved.
Patent Information
- Application Number
- CN202510426093.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing technology lacks dynamic response adaptability adjustment to pollutant migration laws in the long-term monitoring of groundwater in hazardous waste landfills, resulting in high monitoring costs and low efficiency, making it difficult to meet the actual demands of the responsible entities for pollution, and the implementation of the long-term monitoring system is reduced.
By performing classified numerical encoding and pre-processing of groundwater monitoring well data, an artificial intelligence model is constructed for predicting groundwater monitoring wells' exceeding states, combined with the artificial intelligence model, the importance value and normalized importance value affecting groundwater monitoring wells' exceeding states are obtained, and important indicators for groundwater monitoring are determined.
The rapid and efficient optimization of groundwater monitoring indicators of hazardous waste landfills has been achieved, the monitoring costs have been reduced, the monitoring efficiency has been improved, and the implementation of the long-term monitoring system has been improved, providing support for the later management of hazardous waste landfill polluted landfills.
Smart Images

Figure CN120387536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to groundwater monitoring, and particularly to an optimization method and system for groundwater monitoring projects in hazardous waste landfills based on artificial intelligence. Background Art
[0002] Hazardous waste landfills are important means and facilities for the safe disposal of hazardous waste. Heavy metals, organic pollutants, and persistent toxic substances contained in hazardous waste can migrate to the soil and groundwater environment through leakage pathways, resulting in soil and groundwater pollution.
[0003] Once the soil and groundwater in a hazardous waste landfill site are contaminated, full-process risk control should be implemented, including investigation, risk assessment, engineering plan preparation, engineering construction, effectiveness evaluation, and post-construction environmental management. Currently, the control of contaminated sites mainly focuses on stages such as investigation, risk assessment, simulation prediction, and effectiveness evaluation, with less research on post-construction environmental management, especially on long-term management after groundwater control, and few reports on it. Due to the implementation of risk control technologies with "source control as the core and exposure pathway cut-off as the basic strategy" in contaminated sites, pollutants in groundwater usually objectively exist and have significant migration and diffusion risks. Therefore, after the construction of the control project, post-construction management of groundwater needs to be carried out to long-term monitor the migration and diffusion risks of pollutants in groundwater. As one of the core means of post-construction management of contaminated groundwater, long-term monitoring has gradually gained attention in the full-process control of contaminated sites.
[0004] However, long-term groundwater monitoring often relies on a fixed list of monitoring indicators defined during the construction period and effectiveness evaluation period, lacking an adaptive adjustment and optimization mechanism for dynamically responding to the migration laws of pollutants, resulting in high costs and low efficiency for the polluter's long-term monitoring, making it difficult to meet the actual demands of the polluter, and ultimately may lead to a reduction in the implementability of the long-term monitoring system.
[0005] Currently, many studies have been carried out by domestic and foreign scholars in the field of monitoring indicator optimization, such as monitoring indicator optimization methods constructed based on methods such as principal component analysis, water quality comprehensive index, and entropy weight water quality index. However, these methods mostly rely on static data or pollution characteristics at a single time node and are difficult to fully explore the correlation of long-term monitoring data. In addition, the dynamic changes of groundwater monitoring indicators are affected by many factors such as hydrogeological conditions and pollutant migration characteristics, and these influencing factors often have complex non-linear relationships, making it difficult for conventional statistical methods to solve the dynamic optimization problem of long-term monitoring indicators. Summary of the Invention
[0006] To solve the above problems, the present invention provides an optimization method for groundwater monitoring projects in hazardous waste landfills based on artificial intelligence. By preprocessing the data of groundwater monitoring wells through classified numerical coding, an artificial intelligence model for predicting the over-standard status of groundwater monitoring wells is constructed. Combining the artificial intelligence model, the importance value and the normalized importance value affecting the over-standard status of groundwater monitoring wells are obtained, and the important indicators for groundwater monitoring are determined. Those skilled in the art can monitor the pollution of groundwater in hazardous waste landfills based on these indicators.
[0007] The present invention provides the following technical solutions: An optimization method for groundwater monitoring projects in hazardous waste landfills based on artificial intelligence, including the following steps: (1) Collect historical monitoring data of multiple groundwater monitoring wells to construct a groundwater monitoring data set; the groundwater monitoring data set includes: the pollution over-standard status of each monitoring well, the location of each monitoring well, the monitoring items and the corresponding monitoring values, and the monitoring values are used to represent the numerical values or statuses of the monitoring items; (2) Taking the pollution over-standard status of the monitoring well as the target variable and the location and monitoring items of the monitoring well as independent variables, construct multiple decision tree models; (3) Divide the groundwater monitoring data set into a training set and a test set, use the training set and the test set to train and test the decision tree models, and screen the best decision tree model according to the model performance evaluation indicators. The model performance evaluation indicators of the decision tree model include accuracy ACC, precision PRE, recall REC, and the harmonic mean F1 of precision and recall; (4) Use the best decision tree model to determine the monitoring items that have an important impact on the pollution over-standard of groundwater monitoring wells. When and only when the four performance evaluation indicators of one of the decision tree models are all the best among all decision tree models, determine that this decision tree model is the best model. Use this decision tree model to determine the monitoring items that have an important impact on the pollution over-standard of groundwater monitoring wells.
[0008] In the embodiments of the present invention, the CART model, CHAID model, and E-CHAID model are used. As well-known common knowledge in the art, CART is a binary decision tree technology that can be used to predict classification or continuous variables. It constructs a tree structure by recursively splitting the data set until a specific stopping rule is met. For classification problems, the Gini impurity is used; for regression problems, variance minimization is adopted. CHAID is also a decision tree technology, but different from CART, it is multi-directional, which means that each split can have more than two branches. It is based on the chi-square test to determine the best split point. E-CHAID is an extended or improved version of CHAID.
[0009] Further, the pollution exceeding standard status of the monitoring well is an unordered categorical variable, which is encoded as a target variable in the range of [0, 1] according to whether it exceeds the standard and then input into the decision tree model; if the pollution of the monitoring well exceeds the standard, it is encoded as 1; if the pollution of the monitoring well does not exceed the standard, it is encoded as 0.
[0010] Further, the monitoring items in step 1 include key items and auxiliary items; the key items refer to the characteristic toxic and harmful pollutants of the hazardous waste landfill; the auxiliary items refer to the internal and external environments where the characteristic toxic and harmful pollutants are located, including air temperature, pH, COD Mn 、water temperature, turbidity, intensity level of odor and taste, presence or absence of visible substances to the naked eye, dissolved oxygen, water level, redox potential, Cl - content, Fe 3+ content, SO4 2- content, NO3 - content, etc.
[0011] Further, the location of the monitoring well is an unordered categorical variable, which is numerically encoded as [1, 2, 3] according to its relative position in the groundwater flow field. If the monitoring well is located upstream, it is encoded as 1; if the monitoring well is located in the middle reaches, it is encoded as 2; if the monitoring well is located downstream, it is encoded as 3.
[0012] Further, the presence or absence of visible substances to the naked eye is an unordered categorical variable, which is numerically encoded as [0, 1] according to its presence or absence. If there are visible substances to the naked eye, it is encoded as 1; if there are no visible substances to the naked eye, it is encoded as 0.
[0013] In some preferred embodiments of the present invention, for the monitoring items with directly measured values, the monitoring value is encoded by using the threshold judgment method. For example, for turbidity, high, medium, and low judgment thresholds are set: high (>10 NTU), medium (3 - 10 NTU), low (<3 NTU), and the high, medium, and low monitoring values are encoded as 3, 2, 1 respectively; for another example, for the intensity level of odor and taste, 6 levels are set: none, faint, weak, obvious, strong, very strong, and the monitoring values are encoded as 0, 1, 2, 3, 4, 5 respectively. It should be emphasized that the above settings of the encoding thresholds and encoding values can be set in sequence according to the actual situation, and their numerical sizes, etc., do not affect the calculation of the decision tree.
[0014] Further, the calculation methods of the accuracy ACC, precision PRE, recall REC, and the harmonic mean F1 of precision and recall are as follows: ACC=(TM0 + TM1) / (TM0 + TM1 + FM0 + FM1) PRE1=(TM1) / (TM1 + FM1), PRE0=(TM0) / (TM0 + FM0); REC1 = (TM1) / (P1), REC0 = (TM0) / (P0); F1 = (2×PRE×REC) / (PRE + REC); Where: TM0 is the number of samples correctly predicted as non - exceeding the standard in the monitoring well; TM1 is the number of samples correctly predicted as exceeding the standard in the monitoring well; FM0 is the number of samples wrongly predicted as non - exceeding the standard in the monitoring well; FM1 is the number of samples wrongly predicted as exceeding the standard in the monitoring well; PRE1 represents the percentage of samples actually exceeding the standard among the samples predicted as exceeding the standard in the monitoring well; PRE0 represents the percentage of samples actually not exceeding the standard among the samples predicted as not exceeding the standard in the monitoring well; REC1 represents the percentage of samples correctly predicted as exceeding the standard in the monitoring well among its actual samples; REC0 represents the percentage of samples correctly predicted as not exceeding the standard in the monitoring well among its actual samples; P1 is the actual number of samples exceeding the standard in the monitoring well; P0 is the actual number of samples not exceeding the standard in the monitoring well.
[0015] Furthermore, the data ratio of the training set to the test set is 7:3.
[0016] The present invention also provides a groundwater monitoring project optimization system, which at least includes: A collection module, which collects historical monitoring data of multiple groundwater monitoring wells; A first processing module, which is used to pre - process according to the groundwater monitoring well data and construct a groundwater monitoring data set; A second processing module, which is used to analyze the groundwater monitoring data set to determine the monitoring items for which the groundwater monitoring wells are polluted and exceed the standard.
[0017] The present invention also provides a computer - readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the above - mentioned steps are realized.
[0018] The beneficial effects of the present invention are as follows: Through the present invention, the optimization of groundwater monitoring indicators for hazardous waste landfills can be quickly and efficiently realized, the monitoring cost can be reduced, the monitoring efficiency can be improved, the feasibility of the long - term groundwater monitoring system can be enhanced, and support can be provided for the later management of contaminated plots of hazardous waste landfills. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figures 1 to 3 They are the confusion matrices of the performance evaluations of the CART model, CHAID model, and E - CHAID model in sequence; Figure 4 Importance analysis of input variables for the prediction ability of the CART model. DETAILED IMPLEMENTATION MANNER
[0020] In recent years, due to their powerful data mining and feature selection capabilities, machine learning algorithms such as decision trees (DTs), support vector machines (SVMs), neural networks, and random forests, as common supervised machine learning algorithms, have shown application potential in the field of soil and groundwater pollution prevention and control.
[0021] Through comparison and screening, the present invention discovers that in the technical scenario of this application, DT has advantages such as strong interpretability and high computational efficiency compared to other algorithms in judging the relevance of monitoring items.
[0022] It should be noted that constructing the above model based on the target variable and independent variables is a common method in this field, and this application will not elaborate on it.
[0023] In addition, the monitoring items described in this application are the monitoring indicators in this field, and in the embodiments of the present invention, they are described as "monitoring indicators".
[0024] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, taking a hazardous waste landfill in a certain central province of China as an object, the intelligent optimization of the groundwater monitoring indicators of this hazardous waste landfill is realized through the steps of the present invention. The specific implementation steps are as follows: (1) Collected the whole-process construction data of a hazardous waste landfill in a certain central province, including site selection, feasibility study, investigation, design, construction, and completion acceptance, etc. Analyzed in detail the pollution prevention and control situations during the operation process such as hidden danger investigation, self-monitoring, pollution investigation, risk assessment, pollution control, and its effect assessment, and mastered the pollution situation, hydrogeological conditions, and information such as the location, structure, and monitoring indicators of the monitoring wells of the hazardous waste landfill.
[0025] (2) Obtained the existing relevant data on groundwater pollution monitoring, and extracted the data variables of the groundwater monitoring wells, including the pollution exceeding standard status of 73 monitoring wells, the locations of 73 monitoring wells, the monitoring indicators, and the corresponding monitoring values. The specific monitoring indicators include 14 key items (fluoride, arsenic, cadmium, nickel, 1,2-dichloropropane, dichloromethane, cis-1,2-dichloroethylene, 1,2-dichloroethane, 1,2,4-trichlorobenzene, 1,2,3-trichlorobenzene, trichloromethane, dibromochloromethane, isophorone, petroleum hydrocarbons) and 13 auxiliary items (pH, COD Mn , water temperature, turbidity, odor and taste, visible matter to the naked eye, dissolved oxygen, water level, redox potential, Cl - , Fe 3+ , SO4 2- , NO3 -), and perform classification preprocessing on variable data of different categories. Unordered categorical variables (monitoring well location, monitoring well exceedance situation, visible substances), ordered categorical variables (turbidity, odor and taste), and continuous variables (fluoride, arsenic, cadmium, nickel, 1,2-dichloropropane, dichloromethane, cis-1,2-dichloroethylene, 1,2-dichloroethane, 1,2,4-trichlorobenzene, 1,2,3-trichlorobenzene, trichloromethane, dibromochloromethane, isophorone, petroleum hydrocarbons, pH, COD Mn 、water temperature, dissolved oxygen, water level, redox potential, Cl - 、Fe 3 + 、SO4 2- 、NO3 - ) are numerically encoded for their attribute categories according to the "Table 1 Preprocessing and Numerical Encoding Rules for Characteristic Variables" to construct a characteristic data set for optimizing the decision-making of groundwater monitoring indicators in hazardous waste landfills.
[0026] Table 1 Preprocessing and Numerical Encoding Rules for Characteristic Variables (3) Using the constructed characteristic variable data set, with the "monitoring well exceedance situation" as the target variable and the remaining characteristics as independent variables, an artificial intelligence model is built using 3 different decision tree algorithms (CHAID, E-CHAID, CART). The CART model has a significant difference from CHAID and E-CHAID in terms of the degree of growth. The maximum tree depth of CART is 5, which grows more fully compared to CHAID and E-CHAID (both with a maximum tree depth of 3), which may be beneficial to improving the model prediction performance. Since CART grows more fully, the number of internal nodes and leaf nodes is more than that of CHAID and E-CHAID. The number of internal nodes of CART, CHAID, and E-CHAID are 5, 2, and 2 respectively, and the number of leaf nodes are 7, 4, and 4 respectively. The root node of CHAID divides the data set into two relatively purer left and right subtrees, namely "≤ detection limit to standard limit" and "> detection limit to standard limit", with the splitting variable "1,2,4-trichlorobenzene" having the largest χ 2 (232.75). The right subtree has been completely split and does not need to be split again, becoming a leaf node; the left subtree is further divided into 3 leaf nodes with higher "purity" through two splitting variables, fluoride and petroleum hydrocarbons. The root node of E-CHAID-DT uses χ 2(237.34) The largest splitting variable, "1,2,4-trichlorobenzene", divides the dataset into two relatively purer left and right subtrees: "≤ detection limit to standard limit" and "> detection limit to standard limit". The right subtree has been completely split and does not need to be split again, becoming a leaf node; the left subtree is further divided into 3 leaf nodes with higher "purity" by two splitting variables, fluoride and nickel. The CART root node divides the dataset into two relatively purer left and right subtrees: "≤ detection limit to standard limit" and "> detection limit to standard limit" using the splitting variable "1,2,4-trichlorobenzene" with the fastest decrease in "impurity". The right subtree has been completely split and does not need to be split again, becoming a leaf node; the left subtree is further divided into 6 leaf nodes with higher "purity" by four splitting variables, fluoride, nickel, petroleum hydrocarbons, and chloroform. The purity of the leaf nodes affects the prediction performance of the model. The lower the purity of the leaf nodes, the lower the prediction accuracy, recall rate, and classification accuracy of the model.
[0027] To further analyze the prediction performance of the three artificial intelligence models in detail, the training set, test set, and overall model performance (the ratio of the training set to the test set is 7:3) were evaluated. The confusion matrices for the performance evaluation of the three artificial intelligence models are shown in Figures 1 to 3 . From Figures 1 to 3 it can be seen that the performance of CART in terms of accuracy, precision, recall rate, and F1 value is significantly better than that of CHAID and E-CHAID, as reflected in the following aspects: ① In terms of prediction accuracy, the accuracies of the three models for predicting "exceeding the standard" and "not exceeding the standard" in the training set and test set are relatively high, but the prediction accuracy of the CART model is better than that of the CHAID and E-CHAID models, indicating that the CART model has a stronger prediction ability for "exceeding the standard" and "not exceeding the standard" situations. The overall accuracy of the CHAID model is 95.52%, and the accuracies of the training set and test set predictions are 95.80% and 94.85% respectively; the overall accuracy of the E-CHAID model is 95.77%, and the accuracies of the training set and test set predictions are 96.28% and 94.56% respectively; the overall accuracy of the CART model is 98.13%, and the accuracies of the training set and test set predictions are 98.37% and 97.63% respectively.
[0028] ② In terms of prediction accuracy, the overall accuracy of the CART model in predicting "not exceeding the standard" is 97.69%. The accuracies of predicting "not exceeding the standard" in the training set and the test set are 97.99% and 97.04% respectively, both higher than those of the CHAID and E-CHAID models. The overall accuracy of the CHAID model in predicting "not exceeding the standard" is 94.76%, and the accuracies of predicting "not exceeding the standard" in the training set and the test set are 94.86% and 94.55% respectively; the overall accuracy of the E-CHAID model in predicting "not exceeding the standard" is 94.91%, and the accuracies of predicting "not exceeding the standard" in the training set and the test set are 95.59% and 93.26% respectively. The accuracy of the CART model in predicting "exceeding the standard" and that of the E-CHAID model are both 100%, indicating that there are no false positives when predicting "exceeding the standard" samples in the training set and the test set, while there are false positives in the test set for the CHAID model.
[0029] ③ In terms of prediction recall rate, the recall rate of the CART model in predicting "exceeding the standard" is 91.12%. In the training set and the test set, the recall rates of predicting "exceeding the standard" are 92.04% and 89.29% respectively, significantly better than those of the CHAID and E-CHAID models (the recall rate of the CHAID model in predicting "exceeding the standard" is 79.29%, and the recall rates of predicting "exceeding the standard" in the training and test sets are 81.25% and 73.17% respectively; the recall rate of the E-CHAID model in predicting "exceeding the standard" is 79.88%, and the recall rates of predicting "exceeding the standard" in the training and test sets are 80.91% and 77.97% respectively); in the training and test sets, the recall rates of the CART model in predicting "not exceeding the standard" and that of the E-CHAID model are both 100%, indicating that there are no false positives when predicting "not exceeding the standard" samples in the training and test sets, while there are false positives in the test set for the CHAID model.
[0030] ④ In addition, from the F1 index that comprehensively reflects the relationship between accuracy and recall rate, the F1 values (0.99, 0.95) of the CART model for both the non-exceeding and exceeding cases of groundwater monitoring wells are greater than those of the CHAID algorithm (0.97, 0.88) and the E-CHAID algorithm (0.97, 0.89). The larger the F1 value, the better the output result of the model.
[0031] Based on the best artificial intelligence model CART, the importance analysis of input variables on the model's prediction performance was studied, the importance indicators for groundwater monitoring in hazardous waste landfills were determined, and the optimized groundwater monitoring indicators were given. The importance analysis of input variables on the prediction ability of the CART model is shown in Figure 4 as follows.
[0032] As shown in Figure 4It can be seen that the changing trends of the importance values and the normalized importance values generally remain consistent. 1,2,4-Trichlorobenzene and nickel have a very important impact on the prediction ability of the target variable of the CART model, with the normalized importance values reaching 1 and 0.923 respectively, and the importance values reaching 0.146 and 0.135 respectively; fluoride, petroleum hydrocarbons, chloroform, dichloromethane, cadmium, and cis-1,2-dichloroethylene have a relatively important impact on the prediction ability of the CART target variable, with the normalized importance values being 0.476, 0.445, 0.429, 0.328, 0.243, and 0.195 respectively, and the importance values being 0.069, 0.065, 0.063, 0.048, 0.035, and 0.028 respectively; arsenic, 1,2-dichloropropane, 1,2-dichloroethane, 1,2,3-trichlorobenzene, iron ions, sulfate, nitrate, and dibromochloromethane also have a certain impact on the prediction ability of the CART target variable, with the normalized importance values being 0.187, 0.177, 0.173, 0.147, 0.122, 0.078, 0.023, and 0.016 respectively, and the importance values being 0.027, 0.026, 0.025, 0.021, 0.018, 0.011, 0.003, and 0.002 respectively. The importance values of 1,2,4-trichlorobenzene and nickel are significantly higher than those of other variables and have a very important impact on predicting the exceeding standard situation of groundwater monitoring wells. It is recommended to pay special attention to the monitoring of these two indicators in the subsequent long-term groundwater monitoring work; fluoride, petroleum hydrocarbons, chloroform, dichloromethane, cadmium, and cis-1,2-dichloroethylene have relative importance, and it is recommended to continue to pay attention to the monitoring of these six indicators in the subsequent long-term groundwater monitoring work.
[0033] Through the method based on the above artificial intelligence, the optimization of groundwater monitoring indicators at a hazardous waste landfill in the central region is realized, which is optimized from 27 items to 8 items.
[0034] Regarding more content such as the principle, method, and beneficial effects of the decision tree model construction in the embodiments of the present application, reference can be made to the relevant descriptions of the prior art and will not be elaborated here.
[0035] The present invention also provides a groundwater monitoring project optimization system, which at least includes: A collection module that collects historical monitoring data of multiple groundwater monitoring wells; A first processing module for preprocessing according to the groundwater monitoring well data to construct a groundwater monitoring data set; A second processing module for analyzing the groundwater monitoring data set to determine the monitoring items for the pollution exceeding standard of groundwater monitoring wells.
[0036] When the processing module in this embodiment implements the above steps, the specific control method can refer to the examples in the foregoing embodiments and will not be elaborated here.
[0037] This embodiment also provides a computer-readable storage medium, which includes a volatile or non-volatile, removable or non-removable medium implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules, or other data). The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile disc (DVD) or other optical disc storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0038] The computer-readable storage medium in this embodiment can be used to store one or more computer programs, and the one or more computer programs stored therein can be executed by a processor to implement at least one step of the above-mentioned optimization of monitoring indicators.
[0039] As mentioned above, it is only a preferred embodiment of the present invention and does not impose any formal limitations on the present invention. Although the present invention has been disclosed as above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or refinements to equivalent embodiments with equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes and refinements made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. An optimization method for the groundwater monitoring project of hazardous waste landfills based on artificial intelligence, characterized in that, It includes the following steps: (1) Collect the historical monitoring data of multiple groundwater monitoring wells to construct a groundwater monitoring data set; the groundwater monitoring data set includes: the pollution exceeding standard status of each monitoring well, the location of each monitoring well, the monitoring items and the corresponding monitoring values, and the monitoring values are used to represent the numerical values or statuses of the monitoring items; (2) Using the pollution exceeding standard status of the monitoring well as the target variable and the location and monitoring items of the monitoring well as independent variables, construct multiple decision tree models; (3) Divide the groundwater monitoring data set into a training set and a test set, use the training set and the test set to train and test the decision tree models, and screen the best decision tree model according to the model performance evaluation indexes. The performance evaluation indexes of the decision tree model include accuracy ACC, precision PRE, recall REC, and the harmonic mean F1 of precision and recall; (4) Use the best decision tree model to determine the monitoring items that have an important impact on the pollution exceeding standard of the groundwater monitoring well; when and only when the four performance evaluation indexes of one of the decision tree models are all the best among all decision tree models, determine that this decision tree model is the best model; use this decision tree model to determine the monitoring items that have an important impact on the pollution exceeding standard of the groundwater monitoring well.
2. The optimized method for the groundwater monitoring project of a hazardous waste landfill according to claim 1, wherein The pollution exceeding standard status of the monitoring well is an unordered categorical variable, and after being encoded as [0, 1] according to whether it exceeds the standard, it is input into the decision tree model as the target variable; if the monitoring well pollutes and exceeds the standard, it is encoded as 1; if the monitoring well does not pollute and exceed the standard, it is encoded as 0.
3. The optimized method for the groundwater monitoring project of a hazardous waste landfill according to claim 1, wherein The monitoring items in the above-mentioned step 1 include key items and auxiliary items; the key items refer to the characteristic toxic and harmful pollutants in the hazardous waste landfill; the auxiliary items refer to the internal and external environments where the characteristic toxic and harmful pollutants are located, including air temperature, pH, COD Mn , water temperature, turbidity, intensity level of odor and taste, presence or absence of visible substances to the naked eye, dissolved oxygen, water level, redox potential, Cl - content, Fe 3+ content, SO4 2- content, NO3 - content, etc.
4. The optimization method for the groundwater monitoring project of a hazardous waste landfill according to claim 1, wherein The location of the monitoring well is an unordered categorical variable, and it is numerically encoded as [1, 2, 3] according to its relative position in the groundwater flow field. If the monitoring well is located upstream, it is encoded as 1; if the monitoring well is located in the middle reaches, it is encoded as 2; if the monitoring well is located downstream, it is encoded as 3.
5. The optimized method for the groundwater monitoring project of the hazardous waste landfill according to claim 3, characterized in that, The presence or absence of visible substances to the naked eye is an unordered categorical variable, and it is numerically encoded as [0, 1] according to its presence or absence. If there are visible substances to the naked eye, it is encoded as 1; if there are no visible substances to the naked eye, it is encoded as 0.
6. The optimization method for the groundwater monitoring project of a hazardous waste landfill according to claim 1, wherein, The calculation methods of accuracy ACC, precision PRE, recall REC, and the harmonic mean F1 of precision and recall are as follows: ACC = (TM0 + TM1) / (TM0 + TM1 + FM0 + FM1) PRE1 = (TM1) / (TM1 + FM1), PRE0 = (TM0) / (TM0 + FM0); REC1 = (TM1) / (P1), REC0 = (TM0) / (P0); F1 = (2 × PRE × REC) / (PRE + REC); Where: TM0 is the number of samples with correct prediction of non-exceedance in the monitoring wells; TM1 is the number of samples with correct prediction of exceedance in the monitoring wells; FM0 is the number of samples with incorrect prediction of non-exceedance in the monitoring wells; FM1 is the number of samples with incorrect prediction of exceedance in the monitoring wells; PRE1 represents the percentage of samples actually exceeding the standard among the samples predicted to exceed the standard in the monitoring wells; PRE0 represents the percentage of samples actually not exceeding the standard among the samples predicted not to exceed the standard in the monitoring wells; REC1 represents the percentage of the number of samples with correct prediction of exceedance in the monitoring wells in its actual number of samples; REC0 represents the percentage of the number of samples with correct prediction of non-exceedance in the monitoring wells in its actual number of samples; P1 is the actual number of samples exceeding the standard in the monitoring wells; P0 is the actual number of samples not exceeding the standard in the monitoring wells.
7. The optimization method for the groundwater monitoring project of hazardous waste landfills according to claim 1, characterized in that, The data ratio of the training set to the test set is 7:
3.
8. An optimized system for groundwater monitoring projects according to the method described in any one of claims 1-7, characterized in that, At least including: A collection module that collects historical monitoring data of multiple groundwater monitoring wells; A first processing module for preprocessing according to the groundwater monitoring well data to construct a groundwater monitoring data set; A second processing module for analyzing the groundwater monitoring data set to determine the monitoring items with pollution exceeding the standard in the groundwater monitoring wells.
9. A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Multi-objective optimizing method of groundwater pollution monitoring network
CN107525907A
Monitoring well comprehensive investigation and evaluation method and system based on multi-source data and medium
CN114997546A
Underground water monitoring network optimization method based on physical information driven deep learning model
CN116522566A
Underground water environment monitoring well layout method based on machine learning optimization
CN118095497A
Soil heavy metal stabilization effect prediction method based on machine learning
CN118606705A