Black and odorous risk assessment index screening method based on water quality monitoring data

Through the black and odor risk assessment index screening method based on water quality monitoring data, a robust index system was built, which solved the problem of insufficient identification of key factors that cause black and odor and insufficient evaluation accuracy of existing models, and achieved efficient and accurate black and odor risk assessment.

CN120373853APending Publication Date: 2025-07-25SHANGHAI UNIVERSITY OF ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510446759.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing black and odorous water risk assessment model is insufficient in data utilization and analysis depth, and it is difficult to effectively identify key factors that cause black and odor, resulting in low prediction accuracy and difficulty in accurately assessing by mining the complex relationships contained in the data.

Method used

The black and odor risk assessment index screening method based on water quality monitoring data is adopted, including data collection and preprocessing, feature selection, initial screening of indicator systems and in-depth evaluation. Through quantitative feature contribution, integrated learning algorithms and multi-attribute decision analysis methods, a robust indicator system is built, key indicators are identified and risk assessment is carried out.

Benefits of technology

It improves the accuracy and efficiency of black and odor risk assessment, ensures reliability and robustness at different risk levels, simplifies the data processing process, and improves the generalization ability and practical application convenience of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373853A_ABST
    Figure CN120373853A_ABST
Patent Text Reader

Abstract

The invention discloses a black and odorous risk assessment index screening method based on water quality monitoring data, and the method comprises the following working steps: S1, data collection and preprocessing: collecting water quality real-time monitoring data or detection data, and carrying out the missing value processing, abnormal value elimination, data standardization and normalization of the data; s2, feature selection: carrying out importance sorting on the water quality indexes by using importance analysis; s3, index system preliminary screening: evaluating prediction model performance of different index combinations by adopting an ensemble learning algorithm, and preferably selecting an optimal index combination under each dimension; and S4, deep evaluation of the index system: carrying out deep evaluation based on the constructed water body black and odorous risk matrix to obtain an optimal index system. According to the black and odorous risk assessment index screening method based on the water quality monitoring data, the reliability of an index system on each risk level is ensured, powerful support is provided for accurate prediction of the black and odorous risk, and the accuracy of water body black and odorous prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water quality monitoring data, and particularly to a method for screening black and odorous risk assessment indicators based on water quality monitoring data. Background Art

[0002] In recent years, water quality monitoring technology has developed rapidly, accumulating a large amount of data. However, existing black and odorous water body risk assessment models still have deficiencies in data utilization rate and analysis depth. Existing models usually adopt all monitoring indicators, lacking the identification of key factors causing black and odor, resulting in low prediction accuracy and difficulty in effectively assessing the risk of black and odor occurrence. In addition, existing models often rely on simple statistical analysis and are difficult to effectively identify key factors causing black and odor by mining complex relationships contained in data, affecting the accuracy of the final prediction model. The black and odor phenomenon is a gradually deteriorating process, and each water quality indicator shows specific variation rules. For example, the ammonia nitrogen concentration rises in the initial stage, and the dissolved oxygen concentration drops in the later stage. To solve these problems, there is an urgent need for a new black and odorous water body risk assessment model that can effectively identify key factors causing black and odor and quantitatively evaluate the risk of black and odorous water bodies, providing a scientific basis and decision-making support for the prevention and control of black and odorous water bodies.

[0003] As described above, for this reason, we have designed a method for screening black and odorous risk assessment indicators based on water quality monitoring data to solve the above problems. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose a method for screening black and odorous risk assessment indicators based on water quality monitoring data.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] A method for screening black and odorous risk assessment indicators based on water quality monitoring data includes the following working steps:

[0007] Step S1: Data collection and preprocessing: Collect water quality monitoring data, including but not limited to indicators such as dissolved oxygen, chemical oxygen demand, ammonia nitrogen, total phosphorus, turbidity, etc. The data source can be a field monitoring station, sensor collection, or historical database. To ensure the accuracy of subsequent analysis, the original data is cleaned and standardized, and indicators with different dimensions are converted into a unified scale, thereby improving data quality and consistency;

[0008] Step S2: Feature selection: Evaluate the importance of water quality indicators through methods such as quantifying feature contribution, feature compression, or dimensionality reduction, and screen out key indicators that are most statistically significant for black and odorous risk prediction;

[0009] Step S3: Initial screening of the indicator system: Based on the selected key indicators, construct various indicator combinations with different numbers of dimensions and a black-odor risk prediction model driven by an ensemble learning algorithm. Use the F1 score, precision, recall rate, and accuracy to evaluate the model performance, and screen out the indicator system with the optimal water body black-odor prediction performance for different dimensions.

[0010] Step S4: In-depth evaluation of the indicator system: Use the multi-attribute decision-making analysis method to calculate the comprehensive index of the indicator system with the largest number of dimensions; construct a risk matrix of the mapping relationship between the comprehensive index and the water body black-odor probability based on the cluster analysis method; analyze the performance of the indicator system at different risk levels, and determine the robust indicator system applicable to each risk level.

[0011] Preferably, the data collection and preprocessing in the step S1 include: collecting water quality monitoring data covering multiple time periods and multiple regions; for missing values, fill them with the mean, median, or mode, or process them by combining one or more methods of the list method and pairwise deletion method; for outliers, detect them by combining one or more methods of statistical test, quantile method, or 3σ principle, and eliminate, correct, or replace them by combining one or more methods of logarithmic transformation, Box-Cox transformation, or binning processing; standardize the data by combining one or more methods of linear, non-linear, or robust standardization, and convert the original data into a high-quality data set to effectively eliminate noise.

[0012] Preferably, the feature selection in the step S2 includes: quantifying the feature contribution by one or more methods based on the Gini index, information entropy, or permutation importance of the tree model; compressing the features by one or more methods of Lasso regression, elastic net, or stepwise regression; performing feature dimensionality reduction by one or more methods of PCA, t-SNE, or factor analysis.

[0013] Preferably, the initial screening of the indicator system in the step S3 includes: based on the key indicators, construct various indicator combinations with different numbers of dimensions, and use one or more algorithms of RF, GBDT, XGBoost, AdaBoost, and CatBoost to construct a black-odor risk prediction model. By evaluating the F1 score, precision, recall rate, and accuracy, screen out the indicator system with the best performance for different numbers of dimensions.

[0014] Preferably, the in-depth evaluation of the index system in step S4 includes: calculating a comprehensive index by one or more methods among VIKOR, TOPSIS, and AHP, and using one or more methods among K-Means, K-Medoids, and DBSCAN in combination for clustering analysis to construct a risk matrix including the comprehensive index and the mapping of the probability of water body black odor, analyzing the performance of the index system at each risk level, and screening out the most robust index system. Compared with the prior art, the present invention constructs a robust index system through scientific screening and iterative verification, and has the following

[0015] Advantages:

[0016] 1. Efficient identification of key indicators

[0017] The present invention quickly evaluates and screens key indicators through the feature selection step, effectively narrowing the candidate range, eliminating irrelevant or redundant indicators, greatly improving the screening efficiency, and laying a foundation for the subsequent construction of the index system.

[0018] 2. High reliability of the index system

[0019] Through multiple rounds of iterative evaluation and risk level division, the performance changes of the index system in different black odor risk scenarios are revealed, ensuring its prediction reliability at low, medium, medium-high, and high risk levels, providing strong support for accurate black odor risk assessment, and improving the prediction accuracy under different risk conditions.

[0020] 3. Simplification and convenience of the index system

[0021] The present invention accurately locates key indicators through scientific screening and iterative optimization, eliminates redundant or highly correlated features, avoids the problem of multicollinearity, constructs a simplified and convenient index system, simplifies the data structure, and improves the generalization ability of the model and the convenience of practical application.

[0022] 4. High applicability and robustness of the index system

[0023] The selected index system shows stable performance in various prediction models, with small performance differences, indicating that it can maintain consistent prediction ability under different model configurations and data conditions. This high applicability and robustness enhance the reliability and universality of the index system in multi-scenario risk assessment, providing strong technical support for black odor risk management. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a schematic flow chart of a method for screening black odor risk assessment indicators based on water quality monitoring data proposed by the present invention;

[0025] Figure 2The box plot of the data detected by the box plot method for screening the black and odorous risk assessment indicators based on water quality monitoring data proposed by the present invention;

[0026] Figure 3 It is a comparison chart of the performance of the VIKOR comprehensive index of the present invention mapped into the constructed risk matrix. Specific implementation manners

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0028] Refer to Figures 1 - 3 , the present invention proposes a method for screening black and odorous risk assessment indicators based on water quality monitoring data. By using a large amount of monitoring data and combining technical means such as feature selection, model evaluation, and multi-index comprehensive analysis, the optimal water quality index combination under different black and odorous risk levels is screened out.

[0029] Apply the random forest algorithm to perform feature selection on water quality indicators to identify key indicators highly correlated with the black and odorous phenomenon. Based on the selected key indicators, construct indicator combinations with different dimensionality, and use the CatBoost algorithm for training and performance evaluation to screen out the optimal indicator system under each dimensionality. Use the VIKOR method to calculate the comprehensive index of the optimal indicator system under each dimensionality, and combine the risk matrix mapping the comprehensive index and the probability of water body black and odorous based on K-Means clustering analysis to determine the best indicator combination applicable to different black and odorous risk levels.

[0030] In the black and odorous risk assessment, the present invention significantly improves the prediction accuracy through systematic feature selection and multi-model evaluation. The feature selection step quickly locks in key indicators and narrows the candidate range; the model training and evaluation steps identify the optimal indicator combination under different risk levels through multiple rounds of optimization, providing a highly targeted and efficient evaluation tool to ensure the prediction effectiveness at low, medium, medium-high, and high risk levels. The selected indicator system shows stable performance under various models and environmental conditions, indicating its wide applicability, being able to adapt to different data processing methods and water quality scenarios and maintaining consistent prediction ability. In addition, the indicator system is concise and efficient, eliminating redundant features, simplifying the data processing and analysis process, improving the model operation efficiency, and making the risk assessment more intuitive and convenient.

[0031] Specifically, it includes the following working steps:

[0032] Step S1 Data collection and preprocessing:

[0033] Data source: Data collection: Collect water quality monitoring data from 2019 to 2023, a total of 124,800 records, from 10 national control section monitoring stations and 15 local monitoring stations. The monitoring indicators include 9 items: dissolved oxygen (DO, mg / L), turbidity (NTU), total phosphorus (TP, mg / L), ammonia nitrogen (NH3-N, mg / L), permanganate index (mg / L), temperature (°C), pH value, conductivity (μS / cm), total nitrogen (TN, mg / L). The data is divided into a training set (2019 - 2021, 87,360 records), a test set (2022, 24,960 records), and a validation set (2023, 12,480 records) according to the ratio of 7:2:1.

[0034] Data preprocessing: Fill in the missing values (about 0.6%) using linear interpolation. For example, the dissolved oxygen of station A03 was missing in June 2022. According to 4.8 mg / L on June 5th and 4.2 mg / L on June 7th, the interpolated value was 4.5 mg / L; for the missing total nitrogen (0.8 mg / L), the interpolated value was 0.9 mg / L based on the previous and subsequent values; use the box plot method to remove outliers, with the interquartile range threshold of 1.5 times IQR, and remove 1,872 records (such as the turbidity of a certain station suddenly increased to 180 NTU, exceeding the normal range of 10 - 50 NTU; the conductivity exceeded 5000 μS / cm, exceeding the typical range of 100 - 2000 μS / cm); scale the 9 indicators to the [0, 1] interval through Min - Max normalization. For example, DO was normalized from 2 - 8 mg / L to 0.2 - 0.8, and the temperature was normalized from 5 - 30 °C to 0.1 - 0.9. After preprocessing, the data integrity reached 99.8%, providing a high - quality basis for subsequent screening.

[0035] Step S2 Feature selection:

[0036] Build a random forest model: Use the scikit - learn library in Python to build a random forest model, set the number of trees and the maximum depth of each tree to control the complexity of the model and prevent overfitting; through model training, obtain the importance scores of each water quality indicator.

[0037] Feature importance ranking: According to the feature importance scores output by the random forest model, rank the 9 water quality indicators to identify the most critical indicators for predicting the black - odor risk. These indicators include turbidity (NTU), dissolved oxygen (mg / L), total phosphorus (mg / L), permanganate index (mg / L), ammonia nitrogen (mg / L), and these indicators are considered to have a significant predictive effect on the black - odor risk.

[0038] In the content of step S2 above, feature selection focuses on extracting the most representative features for target analysis from complex multi-dimensional data. This process extracts more possible features through the random forest algorithm, quantitatively evaluates the importance of each feature, and selects a set of optimal features according to the feature importance ranking for subsequent construction of the indicator system.

[0039] Initial screening of the indicator system in step S3:

[0040] Build CatBoost models: Use the CatBoost library in Python to build four different CatBoost models, and each model is trained based on a different number of indicators. The indicator combinations of each model are as follows:

[0041] Model 1: The smallest feature set, containing two indicators.

[0042] Model 2: The medium feature set, containing three indicators.

[0043] Model 3: The larger feature set, containing four indicators.

[0044] Model 4: The all-feature set, containing five indicators.

[0045] By training and evaluating these models, evaluate the impact of different indicator combinations on the model performance.

[0046] In step S3, the CatBoost algorithm is used to combine indicators with different dimensionality numbers to build multiple indicator systems, and through comprehensively considering model performance indicators such as accuracy, recall rate, F1 score, and precision, the indicator systems with different dimensionality numbers are preliminarily screened to determine the indicator systems with excellent performance in predicting water body black odor under each dimension for subsequent in-depth analysis.

[0047] Step S3.1 Model prediction: Use the test set data to predict the four CatBoost models to obtain the prediction results of each model under different indicator combinations. These prediction results are used to evaluate the performance of the models on actual data.

[0048] Step S3.2 Indicator system screening: Through comprehensively evaluating key performance indicators of the models such as accuracy, recall rate, F1 score, and precision, screen out the indicator systems with the best performance under each sub-model. The results are shown in Table 1:

[0049] Table 1 Performance indicators of the optimal indicator combinations under each model

[0050] model indicator combination precision accuracy recall rate F1-score Model 1 dissolved oxygen, turbidity 0.82 0.85 0.90 0.92 Model 2 dissolved oxygen, turbidity, total phosphorus 0.85 0.88 0.93 0.95 Model 3 dissolved oxygen, turbidity, permanganate index, total phosphorus 0.86 0.90 0.95 0.97 Model 4 dissolved oxygen, turbidity, permanganate index, total phosphorus, ammonia nitrogen 0.87 0.91 0.96 0.98

[0051] Deep evaluation of the indicator system in step S4:

[0052] Among them, it includes three sub-steps: comprehensive index calculation, risk matrix construction, and index system evaluation.

[0053] Step S4.1 Comprehensive index calculation

[0054] Calculate the comprehensive index through the VIKOR model. These indexes are obtained based on the optimal index system selected under different sub-models.

[0055] Step S4.2 Risk matrix construction

[0056] According to the comprehensive index calculation results, construct a risk matrix, and construct a risk matrix that maps the comprehensive index of the index system under the highest number of dimensions to the probability of water body black odor through cluster analysis.

[0057] Specifically, divide the probability of water body black odor into four risk levels: low, medium, medium-high, and high. The ranges of black odor occurrence probabilities corresponding to these risk levels are as follows:

[0058] Low risk: The probability of black odor occurrence is less than 30%

[0059] Medium risk: The probability of black odor occurrence is 30%-50%

[0060] Medium-high risk: The probability of black odor occurrence is 50%-70%

[0061] High risk: The probability of black odor occurrence is 70%-100%

[0062] On this basis, further construct a risk matrix that maps the comprehensive index to the probability of water body black odor through cluster analysis using the index system of the highest number of dimensions, as shown in Table 2. This matrix provides an intuitive reference for subsequent decision-making.

[0063] Table 2 Risk matrix of black-odored water bodies

[0064] comprehensive index risk probability risk level [0.0,0.33] 0%-30% Ⅰ (low risk) [0.33,0.47] 30%-50% Ⅱ (medium risk) [0.47,0.63] 50%-70% Ⅲ (medium-high risk) [0.63,1] 70%-100% Ⅳ (high risk)

[0065] Step S4.3 Index system evaluation

[0066] Through comprehensive evaluation of the performance of each index system at different risk levels, finally determine the optimal index system. The comprehensive evaluation includes accuracy, recall rate, F1 value, precision, robustness, and model applicability. Considering factors such as model performance, risk level, and number of indexes, it is finally determined that among the currently investigated index systems, "turbidity + dissolved oxygen" is a relatively robust and optimal-performing index combination. It shows good stability and accuracy at different risk levels, and there is no sign that adding additional characteristics such as total phosphorus and permanganate index can significantly improve the model performance. Therefore, if you pursue relatively reliable results at each risk level, the "turbidity + dissolved oxygen" combination is a better choice.

[0067] This embodiment details how to use the method of the present invention for predicting the risks of black and odorous water bodies, and demonstrates the effectiveness of the method through specific data and charts. The present invention can help the management department better understand the risk status of black and odorous water bodies, and formulate more targeted treatment plans, thereby improving the treatment efficiency and saving treatment costs.

[0068] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. Without departing from the spirit and scope of the present invention, various changes and improvements will occur to the present invention, and all these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A method for screening black and odorous risk assessment indicators based on water quality monitoring data, characterized in that, It includes the following working steps: Step S1: Data collection and preprocessing: Collect water quality monitoring data, including indicators such as dissolved oxygen, chemical oxygen demand, ammonia nitrogen, total phosphorus, turbidity, etc., and ensure data quality through data cleaning and standardization processing; Step S2: Feature selection: Evaluate the importance of water quality indicators through methods such as quantifying feature contributions, feature compression, or dimensionality reduction, and screen out the key indicators that are most statistically significant for predicting black and odorous risks; Step S3: Initial screening of the indicator system: Based on the screened key indicators, construct multiple indicator combinations with different numbers of dimensions and black and odorous risk prediction models driven by ensemble learning algorithms, and use F1 score, precision, recall rate, and accuracy to evaluate the model performance, and screen out the indicator system with the optimal performance for predicting water body black and odor at different dimensions; Step S4: In-depth evaluation of the indicator system: Calculate the comprehensive index of the indicator system with the largest number of dimensions using the multi-attribute decision-making analysis method; Construct a risk matrix that maps the comprehensive index and the probability of water body black and odor based on the cluster analysis method; Analyze the performance of the indicator system at different risk levels, and determine a robust indicator system suitable for each risk level.

2. The screening method of black and odorous risk assessment indicators based on water quality monitoring data according to claim 1, wherein The data preprocessing in Step S1 includes one or several combinations of missing value processing, outlier processing, or data standardization.

3. A screening method for black and odorous risk assessment indicators based on water quality monitoring data according to claim 1, characterized in that, The methods for quantifying feature contributions in Step S2 include one or more combinations of the Gini index of the tree model, information entropy, or permutation importance; feature compression includes one or more combinations of Lasso regression, elastic net, or stepwise regression; the dimensionality reduction method includes one or more combinations of PCA, t-SNE, or factor analysis.

4. A method for screening black and odorous risk assessment indicators based on water quality monitoring data according to claim 1, characterized in that, The ensemble learning algorithms in Step S3 include one or more combinations of random forest (RF), gradient boosting decision tree (GBTD), extreme gradient boosting (XGBoost), adaptive boosting (AdaBoost), or category boosting (CatBoost).

5. A screening method for black and odorous risk assessment indicators based on water quality monitoring data according to claim 1, characterized in that The multi-attribute decision-making analysis method described in Step S4 includes one or more combinations of the compromise ranking method (VIKOR), the technique for order preference by similarity to an ideal solution (TOPSIS), and the analytic hierarchy process (AHP); the cluster analysis method includes one or more combinations of K-Means, K-Medoids, and DBSCAN.