A water quality comprehensive evaluation method based on random forest optimization of water quality indexes

By optimizing water quality indicators using the random forest algorithm, a comprehensive river water quality evaluation system was constructed, which solved the problems of accuracy and economy in water quality evaluation and reduced evaluation costs.

CN115600810BActive Publication Date: 2026-05-01CHINA THREE GORGES CORPORATION +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA THREE GORGES CORPORATION
Filing Date
2022-10-20
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing comprehensive water quality assessment methods struggle to balance accuracy and cost-effectiveness, resulting in high assessment costs and inaccurate results.

Method used

The random forest algorithm is used to select water quality indicators. By training, predicting and selecting water quality indicators, a comprehensive evaluation system for river water quality is constructed, reducing the need to observe non-critical indicators.

Benefits of technology

It achieves a balance between accuracy and economy in water quality assessment, and significantly reduces observation and testing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115600810B_ABST
    Figure CN115600810B_ABST
Patent Text Reader

Abstract

The application provides a water quality comprehensive evaluation method based on random forest optimization of water quality indexes, comprising the following steps: determining a river section and water quality indexes, obtaining a measured data set, and calculating a water quality index data set; dividing the water quality index data set into a training set and a prediction set; constructing a training model based on the training set; constructing a prediction model based on the prediction set, predicting the water quality index, and evaluating the performance of the training model; based on the training result and the evaluation result, determining the optimized water quality indexes according to the contribution degree ranking; calculating the water quality index data set based on the optimized water quality indexes; gradually reducing the number of optimized water quality indexes, and calculating the water quality index data set; evaluating the prediction results of different numbers of optimized water quality indexes one by one, and determining the optimal water quality indexes; and calculating the water quality index of the river by using the optimal water quality indexes, so as to realize the water quality comprehensive evaluation. The method takes into account the accuracy and economy of water quality evaluation, reduces the observation of non-key indexes as much as possible, and reduces the evaluation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of water quality monitoring technology, and relates to a comprehensive water quality evaluation method, particularly a comprehensive water quality evaluation method based on random forest optimization of water quality indicators. Background Technology

[0002] Environmental problems caused by rapid social and economic development have severely constrained my country's sustainable development. Water, as the source of all life, is the most severely polluted, facing extremely serious forms of pollution. To curb the deteriorating water environment and achieve improvement, governments at all levels have introduced a series of control and management measures. Developing scientific and effective water management measures requires comprehensive water quality assessment, quantitative identification of pollutant sources and migration pathways, pollution source control and interception, and the development of water environment decision support systems. Among these, comprehensive water quality assessment is the foundation for formulating water environment management measures.

[0003] Extensive research has been conducted by scholars both domestically and internationally on methods for river water quality assessment, including single-factor evaluation, comprehensive pollution index method, hierarchical evaluation method, fuzzy evaluation method, grey evaluation method, and water quality index method. Each method has its advantages and disadvantages. For example, single-factor evaluation is simple, straightforward, and safe, directly reflecting whether water quality meets functional requirements, but it cannot comprehensively reflect the overall water quality status. The comprehensive pollution index method can determine the degree of pollution and major pollutants, and judge the trend of water quality changes, but it also cannot comprehensively reflect the overall water quality status. The fuzzy comprehensive evaluation method can well consider uncertain factors in water bodies, effectively solving fuzzy and difficult-to-quantify problems, and is suitable for indeterminate problems. However, the results are prone to distortion, invalidation, homogenization, and jumps, leading to inaccurate assessment results. The water quality index method, by converting the concentrations of multiple indicators into standard factors and assigning weights to the influence of each factor, can achieve a comprehensive evaluation of water quality, and its evaluation results are relatively accurate. However, the water quality index method usually requires a large number of water quality indicators, significantly increasing the cost of water quality assessment. Therefore, it is particularly important to reduce the number of indicators in the comprehensive evaluation process of water quality index, lower costs, and at the same time ensure the accuracy of the evaluation results using the water quality index method.

[0004] Therefore, it is evident that providing a comprehensive water quality evaluation method that balances accuracy and economy, minimizes the observation of non-critical indicators, and reduces evaluation costs has become an urgent problem for those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a comprehensive water quality evaluation method based on random forest to select water quality indicators. This comprehensive water quality evaluation method takes into account both the accuracy and economy of water quality evaluation, and minimizes the observation of non-critical indicators, thereby significantly reducing the evaluation cost.

[0006] To achieve this objective, the present invention adopts the following technical solution:

[0007] This invention provides a comprehensive water quality evaluation method based on random forest-based selection of water quality indicators. The comprehensive water quality evaluation method includes the following steps:

[0008] (1) Determine the river cross-section and water quality indicators, and obtain the measured data set of water quality indicators on the river cross-section;

[0009] (2) Based on the measured dataset obtained in step (1), the corresponding water quality index dataset is calculated;

[0010] (3) Randomly divide the water quality index dataset obtained in step (2) into a training set and a prediction set;

[0011] (4) Based on the training set obtained in step (3), a training model is constructed using the random forest algorithm;

[0012] (5) Combine the prediction set obtained in step (3) and the training model obtained in step (4) to construct a prediction model, predict the water quality index, and evaluate the performance of the training model based on the prediction results.

[0013] (6) Based on the training results of the training model obtained in step (4) and the evaluation results obtained in step (5), the preferred water quality indicators are determined according to the ranking of contribution.

[0014] (7) Based on the preferred water quality indicators obtained in step (6), the predicted water quality index dataset is calculated;

[0015] (8) Gradually reduce the number of preferred water quality indicators obtained in step (6), repeat steps (3)-(7), and calculate the predicted water quality indicator dataset.

[0016] (9) Evaluate the prediction results of different numbers of selected water quality indicators one by one, and determine the optimal water quality indicator;

[0017] (10) Calculate the river’s water quality index using the optimal water quality index obtained in step (9), thus achieving a comprehensive water quality evaluation.

[0018] This invention establishes a complete comprehensive evaluation system for river water quality. It uses a random forest algorithm to train, predict, and optimize numerous water quality indicators, and conducts a comprehensive evaluation of river water quality based on the optimal water quality indicators. This system not only considers the comprehensiveness of water quality evaluation and selects key indicators that affect water quality, ensuring the accuracy of the water quality evaluation results, but also minimizes the observation of non-key indicators, thereby significantly reducing observation and testing costs.

[0019] Preferably, the formulas involved in the calculation in step (2) include:

[0020]

[0021] In the formula, WQI is the water quality index; n is the total number of water quality indicators; C i P is the standardized value of the i-th water quality indicator; i Let be the weight of the i-th water quality indicator.

[0022] In this invention, the standardized value C i The specific method for determining it is as follows:

[0023] (A) When the corresponding water quality indicator belongs to one of the basic items in the Surface Water Environmental Quality Standard (GB 3838-2002), the standardized value C i The calculation formula is:

[0024]

[0025] In the formula, T i S represents the measured data for the i-th water quality indicator; i,k and S i,k+n Let I be the concentration of the water quality standard for the k-th and (k+n)-th categories corresponding to the i-th water quality indicator; i,k is the standardized value corresponding to the concentration of the k-th water quality standard; n is the number of water quality standard concentrations that are the same, and if there are no identical concentrations, then n = 1.

[0026] Specifically, the I i,k I can be used i,1 =20, I i,2 =40, I i,3 =60, I i,4 =80, I i,5 =100, which correspond to the standardized values ​​of Class I, Class II, Class III, Class IV and Class V in the surface water environmental quality standards, respectively.

[0027] (B) When the corresponding water quality indicator is not one of the basic items in the Surface Water Environmental Quality Standard (GB 3838-2002), the standardized value C i The determination can be made with reference to Table 1 below:

[0028] Table 1

[0029]

[0030] Preferably, the proportion of training set data in step (3) is 60-80%, for example, it can be 60%, 62%, 64%, 66%, 68%, 70%, 72%, 74%, 76%, 78% or 80%, but it is not limited to the listed values. Other unlisted values ​​within this range are also applicable.

[0031] Preferably, the training model in step (4) is constructed based on the randomForest package in the R language.

[0032] Preferably, the expression for the training model in step (4) is:

[0033] (D,θ n )=(x1,y1)……(x n ,y n (2)

[0034] In the formula, x is the independent variable; y is the dependent variable; and n is the total number of elements.

[0035] With g(D,θ) n A random forest predictor is constructed from N CART regression trees, with decision g(D,θ) as the basis. n The regression model is based on n = 1, 2, 3, ..., N, and the mean of the regression results is taken.

[0036] Preferably, the prediction model in step (5) is constructed from the training model and used to validate the training model.

[0037] This invention, based on a training set, utilizes the random forest algorithm to construct a training model. Essentially, it establishes a mapping relationship between water quality indicators and water quality indices, namely Y = f(X1, X2, X3, ... X...). n The prediction model uses the water quality indicators X1, X2, X3, ... X in the prediction set. n Substituting these values ​​into the training model (i.e., the mapping relationship) yields the predicted water quality index Y. Comparing the predicted water quality index Y with the actual water quality index Y verifies whether the performance of the training model meets the requirements.

[0038] Preferably, the proportion of the preferred water quality indicators in step (6) to the total number of water quality indicators is ≤50%, for example, it can be 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45% or 50%, but it is not limited to the listed values. Other unlisted values ​​within this range are also applicable.

[0039] Preferably, the number of water quality indicators repeated in step (8) until the preferred indicators are 2-4, for example, 2, 3 or 4.

[0040] Preferably, the parameters involved in the evaluation in step (9) include root mean square error and / or mean absolute percentage error.

[0041] In this invention, the root mean square error (RMSE) is calculated using the following formula:

[0042]

[0043] In the formula, X 实测值,i X represents the water quality index in the measured dataset; 预测值,i The water quality index is predicted based on the preferred water quality indicators; n is the total amount of data.

[0044] In this invention, the formula for calculating the Mean Absolute Percentage Error (MAPE) is as follows:

[0045]

[0046] In the formula, X 实测值,i X represents the water quality index in the measured dataset; 预测值,i The water quality index is predicted based on the preferred water quality indicators; n is the total amount of data.

[0047] Preferably, the formulas involved in the calculations in steps (7), (8) and (10) are the same as those involved in the calculations in step (2).

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] This invention establishes a complete comprehensive evaluation system for river water quality. It uses a random forest algorithm to train, predict, and optimize numerous water quality indicators, and conducts a comprehensive evaluation of river water quality based on the optimal water quality indicators. This system not only considers the comprehensiveness of water quality evaluation and selects key indicators that affect water quality, ensuring the accuracy of the water quality evaluation results, but also minimizes the observation of non-key indicators, thereby significantly reducing observation and testing costs. Attached Figure Description

[0050] Figure 1 This is a schematic diagram of the water quality comprehensive evaluation method provided by the present invention;

[0051] Figure 2 These are the prediction and evaluation results of different numbers of optimized water quality indicators in the comprehensive water quality evaluation method provided in Example 1;

[0052] Figure 3 These are the prediction and evaluation results of different numbers of optimized water quality indicators in the comprehensive water quality evaluation method provided in Example 2. Detailed Implementation

[0053] The technical solution of the present invention will be further illustrated below through specific embodiments. Those skilled in the art should understand that the embodiments described are merely illustrative of the present invention and should not be construed as limiting the invention in any way.

[0054] This invention provides a comprehensive water quality evaluation method based on random forest optimization of water quality indicators, such as... Figure 1 As shown, the comprehensive water quality evaluation method includes the following steps:

[0055] (1) Determine the river cross-section and water quality indicators, and obtain the measured data set of water quality indicators on the river cross-section;

[0056] (2) Based on the measured dataset obtained in step (1), the corresponding water quality index dataset is calculated; the formulas involved in the calculation include:

[0057]

[0058] In the formula, WQI is the water quality index; n is the total number of water quality indicators; C i P is the standardized value of the i-th water quality indicator; i Let be the weight of the i-th water quality indicator;

[0059] (3) The water quality index dataset obtained in step (2) is randomly divided into a training set and a prediction set, and the proportion of the data in the training set is 60-80%.

[0060] (4) Based on the training set obtained in step (3), a training model is constructed using the randomForest package in R language, and the expression of the training model is:

[0061] (D,θ n )=(x1,y1)……(x n ,y n (2)

[0062] In the formula, x is the independent variable; y is the dependent variable; and n is the total number of elements.

[0063] With g(D,θ) n A random forest predictor is constructed from N CART regression trees, with decision g(D,θ) as the basis. n A regressor based on n = 1, 2, 3, ..., N is used to take the mean of the regression results.

[0064] (5) Combine the prediction set obtained in step (3) and the training model obtained in step (4) to construct a prediction model, predict the water quality index, and evaluate the performance of the training model based on the prediction results.

[0065] (6) Based on the training results of the training model obtained in step (4) and the evaluation results obtained in step (5), the preferred water quality indicators are determined according to the ranking of contribution, and the proportion of the preferred water quality indicators to the total number of water quality indicators is ≤50%.

[0066] (7) Based on the preferred water quality indicators obtained in step (6), the predicted water quality index dataset is calculated;

[0067] (8) Gradually reduce the number of preferred water quality indicators obtained in step (6), repeat steps (3)-(7) until the number of preferred water quality indicators is 2-4, and calculate the predicted water quality indicator dataset.

[0068] (9) Evaluate the prediction results of different numbers of preferred water quality indicators one by one based on the root mean square error and / or mean absolute percentage error, and determine the optimal water quality indicator.

[0069] Specifically, the formula for calculating the root mean square error (RMSE) is as follows:

[0070]

[0071] Specifically, the formula for calculating the Mean Absolute Percentage Error (MAPE) is as follows:

[0072]

[0073] In the formula, X 实测值,i X represents the water quality index in the measured dataset; 预测值,i The water quality index is predicted based on the selected water quality indicators; n is the total amount of data.

[0074] (10) Calculate the river’s water quality index using the optimal water quality index obtained in step (9), thus achieving a comprehensive water quality evaluation.

[0075] The formulas involved in the calculations in steps (7), (8), and (10) are the same as those involved in the calculations in step (2).

[0076] Example 1

[0077] This embodiment provides a comprehensive water quality evaluation method based on random forest optimization of water quality indicators. The comprehensive water quality evaluation method includes the following steps:

[0078] (1) A total of 304 sets of measured data were selected from the main stream of the Yangtze River in three quarters. The measured data set included TP and NH4. + -N, TN, NO3 - -N,Mg 2+ Ca 2+ Cl - SO4 2- A total of 14 water quality indicators, including Cu, Zn, As, Se, Cd, and Pb;

[0079] (2) The corresponding water quality index is calculated using formula (1) and combined with the water quality indicators to form a water quality index dataset. Due to space limitations, this embodiment only selects a portion of the water quality index dataset as shown in Table 2 below.

[0080] Table 2

[0081]

[0082]

[0083] (3) In this embodiment, the water quality index dataset in Table 2 above is randomly divided into a training set and a prediction set, and the data volume ratio of the training set is 60% and the data volume ratio of the prediction set is 40%.

[0084] (4) Based on the training set obtained in step (3), use the randomForest package in R language to build a training model;

[0085] (5) Combine the prediction set obtained in step (3) and the training model obtained in step (4) to construct a prediction model, predict the water quality index, and evaluate the performance of the training model based on the prediction results.

[0086] (6) Based on the training results of the training model obtained in step (4) and the evaluation results obtained in step (5), according to the ranking of contribution, seven preferred water quality indicators are determined, namely Pb, TN, Cd, Zn, and NO3. - -N, AS, and TP;

[0087] (7) Based on the preferred water quality indicators obtained in step (6), the predicted water quality index dataset is calculated;

[0088] (8) Gradually reduce the number of preferred water quality indicators obtained in step (6), repeat steps (3)-(7) until the number of preferred water quality indicators is 2, and calculate the predicted water quality indicator dataset.

[0089] (9) The prediction results of different numbers of optimized water quality indicators were evaluated one by one based on the root mean square error and the mean absolute percentage error. The relevant prediction results and evaluation results are shown in [reference to relevant data]. Figure 2 Therefore, the optimal water quality indicators were determined to be TN, Pb, Cd, Zn, and NO3. - -N and As;

[0090] (10) Calculate the river’s water quality index using the optimal water quality index obtained in step (9), thus achieving a comprehensive water quality evaluation.

[0091] Example 2

[0092] This embodiment provides a comprehensive water quality evaluation method based on random forest-based optimization of water quality indicators. It selects 19 sets of water quality data from the middle and lower reaches of the Yangtze River in 2006, as recorded in the literature (Mueller, B., et al. "How polluted is the Yangtze River? Water quality downstream from the Three Gorges Dam." Science of the Total Environment 402.2-3 (2008).), and applies the seven optimized water quality indicators (TN, Pb, Cd, Zn, NO3) determined in Example 1. - -N, As and TP are used as water quality indicators to calculate the predicted water quality index dataset; gradually reduce the number of preferred water quality indicators and repeat steps (3)-(7) in Example 1 until the number of preferred water quality indicators is 2, and calculate the predicted water quality index dataset.

[0093] This embodiment evaluates the prediction results of different numbers of preferred water quality indicators one by one based on the root mean square error and the mean absolute percentage error. The relevant prediction and evaluation results are shown below. Figure 3 Therefore, the optimal water quality indicators were determined to be TN, Pb, Cd, Zn, and NO3. - -N and As are consistent with the optimal water quality indicators obtained in Example 1; the water quality index of the river is calculated using the obtained optimal water quality indicators, thus realizing the comprehensive evaluation of water quality.

[0094] Therefore, this invention establishes a complete comprehensive evaluation system for river water quality. It uses the random forest algorithm to train, predict, and optimize numerous water quality indicators, and conducts a comprehensive evaluation of river water quality based on the optimal water quality indicators. This not only takes into account the comprehensiveness of water quality evaluation and optimizes the key indicators that affect water quality, ensuring the accuracy of water quality evaluation results, but also minimizes the observation of non-key indicators, thereby significantly reducing observation and testing costs.

[0095] The applicant declares that the above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention fall within the protection and disclosure scope of the present invention.

Claims

1. A comprehensive water quality evaluation method based on random forest-based selection of water quality indicators, characterized in that, The comprehensive water quality evaluation method includes the following steps: (1) Determine the river cross-section and water quality indicators, and obtain the measured data set of water quality indicators on the river cross-section; (2) Based on the measured dataset obtained in step (1), the corresponding water quality index dataset is calculated; (3) Randomly divide the water quality index dataset obtained in step (2) into a training set and a prediction set; (4) Based on the training set obtained in step (3), a training model is constructed using the random forest algorithm; (5) Combine the prediction set obtained in step (3) and the training model obtained in step (4) to construct a prediction model, predict the water quality index, and evaluate the performance of the training model based on the prediction results. (6) Based on the training results of the training model obtained in step (4) and the evaluation results obtained in step (5), determine the preferred water quality indicators according to the ranking of contributions; (7) Based on the preferred water quality indicators obtained in step (6), the predicted water quality index dataset is calculated; (8) Gradually reduce the number of preferred water quality indicators obtained in step (6), repeat steps (3)-(7), and calculate the predicted water quality indicator dataset; (9) Evaluate the prediction results of different numbers of selected water quality indicators one by one, and determine the optimal water quality indicator; (10) Calculate the river’s water quality index using the optimal water quality index obtained in step (9), thus achieving a comprehensive water quality evaluation; The expression for the training model in step (4) is: (D,θ n )=(x1,y1),……,(x n ,y n ) (2) In the formula, x is the independent variable; y is the dependent variable; n is the total number of elements; D is the training dataset; θ n Let be the random parameter vector of the nth decision tree; With g(D, θ) n A random forest predictor is constructed from N CART regression trees, and a decision tree g(D, θ) is used. n A regressor based on n=1,2,3,…,N is used to take the mean of the regression results. The parameters involved in the evaluation in step (9) include root mean square error and / or mean absolute percentage error.

2. The water quality comprehensive evaluation method according to claim 1, characterized in that, The formulas involved in the calculation in step (2) include: (1) In the formula, WQI is the water quality index; n is the total number of water quality indicators; C i P is the standardized value of the i-th water quality indicator; i Let be the weight of the i-th water quality indicator.

3. The comprehensive water quality evaluation method according to claim 1 or 2, characterized in that, The proportion of data in the training set in step (3) is 60-80%.

4. The comprehensive water quality evaluation method according to claim 1, characterized in that, The training model described in step (4) is constructed based on the randomForest package in the R language.

5. The comprehensive water quality evaluation method according to claim 1, characterized in that, The prediction model described in step (5) is constructed from the training model and used to validate the training model.

6. The water quality comprehensive evaluation method according to claim 1, characterized in that, In step (6), the proportion of the preferred water quality indicators to the total number of water quality indicators is ≤50%.

7. The comprehensive water quality evaluation method according to claim 1, characterized in that, Step (8) involves repeating the process until 2-4 water quality indicators are selected.

8. The comprehensive water quality evaluation method according to claim 2, characterized in that, The formulas involved in the calculations in steps (7), (8) and (10) are the same as those involved in the calculations in step (2).

Citation Information

Patent Citations

  • Improvements in and relating to sticks of shaving soap and the like, and holders therefor

    GB240025A