Rapid detection method for heavy metal concentration of water body in floating rice planting area
By using the limit gradient enhancement algorithm in the water bodies in the coal mining subsidence area to construct a heavy metal concentration prediction model, the problem of difficulty in quickly monitoring the heavy metal concentration in the water bodies in the existing technology is solved, and the rapid and accurate detection of the water bodies in the floating rice planting area is achieved, ensuring the safety of the planting environment.
Patent Information
- Application Number
- CN202510116314.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to quickly and accurately monitor the concentration of heavy metals in water bodies in coal mining subsidence areas, affecting the safety of "floating rice" planting.
The extreme gradient enhancement algorithm is used to construct a heavy metal element concentration prediction model. By obtaining the physical and chemical index and nutrient index data of water bodies, a direct connection with heavy metal concentration is established to achieve rapid detection.
It realizes rapid monitoring and prediction of the heavy metal concentration of water in floating rice planting areas, improves detection efficiency and accuracy, and ensures the safety of the rice planting environment.
Smart Images

Figure CN120183528A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a water heavy metal detection technology, in particular to a rapid detection method for identifying the heavy metal concentration in the water body of a floating rice planting area by using the extreme gradient boosting algorithm. Background Art
[0002] With the rapid development of industrialization and urbanization, the problem of heavy metal pollution in water bodies has become increasingly severe and has become an important factor affecting water quality. Heavy metal pollution refers to environmental pollution caused by heavy metals or their compounds, mainly caused by human factors such as mining, waste gas emissions, sewage irrigation, and the use of products with excessive heavy metals. The accumulation of heavy metals not only affects the aquatic ecosystem but also poses a threat to human health through the food chain.
[0003] In recent years, a large amount of coal mining has provided energy power for economic development while forming a large number of subsidence areas. Planting "floating rice" in the subsidence areas makes it possible for the subsidence areas to become granaries on water. At present, as an innovative model, the "floating rice" planting technology realizes the suspended growth of rice on the water surface through advanced means such as floating beds, floating islands, and floating plates. The "floating rice" technology is particularly suitable for the ecological restoration of coal mining subsidence water areas and effectively addresses the unique challenges faced by this area.
[0004] However, the heavy metal pollution in coal mining subsidence water areas is serious, posing a potential threat to the planting safety of "floating rice". Therefore, in order to ensure the cleanliness and safety of the planting environment of "floating rice" in coal mining subsidence areas, there is an urgent need for a detection method that can quickly monitor the heavy metal concentration in the water body to determine whether the water environment in coal mining subsidence water areas is suitable for planting.
[0005] Traditional heavy metal monitoring methods rely on laboratory chemical analysis. Although they have high accuracy, they require a large amount of time to sample the water body and then bring it back to the laboratory for testing, which is time-consuming and costly, and cannot achieve rapid monitoring and prediction. As is well known, there are complex non-linear relationships among water physical and chemical indicators such as temperature (T), dissolved oxygen (DO), turbidity (TU), and pH, and nutrient indicators such as total nitrogen (TN), total phosphorus (TP), ammonia nitrogen (NH4 + -N), nitrate nitrogen (NO3--N), and orthophosphate (PO4 3 --P), as well as the heavy metal concentration in the water body such as beryllium (Be), cadmium (Cd), cobalt (Co), iron (Fe), nickel (Ni), lead (Pb), titanium (Ti), and zinc (Zn). Therefore, there is an urgent need to design a detection method to quickly and accurately measure the heavy metal concentration in the water environment. Summary of the Invention
[0006] The present invention aims to avoid the deficiencies in the above-mentioned existing technologies and provides a rapid detection method for the heavy metal concentration in the water body of a floating rice planting area, so as to achieve the rapid monitoring and prediction of the heavy metal concentration in the water body.
[0007] The present invention adopts the following technical solutions to solve the technical problems.
[0008] The present invention provides a rapid detection method for the heavy metal concentration in the water body of a floating rice planting area, including the following steps:
[0009] Step 1: Obtain the physical and chemical index data, nutrient index data, and heavy metal concentration data of the water body in the floating rice planting area;
[0010] Step 2: Preprocess the obtained physical and chemical index data, nutrient index data, and heavy metal concentration data;
[0011] Step 3: Construct a heavy metal element concentration prediction model;
[0012] Step 4: Evaluate the prediction performance of the heavy metal element concentration prediction model constructed in Step 3.
[0013] The characteristics of a rapid detection method for the heavy metal concentration in the water body of a floating rice planting area according to the present invention also lie in:
[0014] Further, in Step 1, the physical and chemical index data includes temperature T, dissolved oxygen DO, turbidity TU, and pH value;
[0015] The nutrient index data includes total nitrogen TN, total phosphorus TP, ammonia nitrogen NH4 + -N, nitrate nitrogen NO3 - -N, and orthophosphate PO4 3- -P;
[0016] The heavy metal concentration data includes beryllium Be, cadmium Cd, cobalt Co, iron Fe, nickel Ni, lead Pb, titanium Ti, and zinc Zn.
[0017] Further, in Step 2, the process of data preprocessing includes removing outliers and filling missing values.
[0018] Further, the missing values of the physical and chemical indexes Wp of the water body are filled by the method of linear interpolation; the missing values of the nutrient indexes Wn and heavy metal indexes Wm are filled by the method of interpolation combined with random forest regression.
[0019] Further, the process of data preprocessing also includes performing correlation analysis on the physical and chemical index data, nutrient index data, and heavy metal concentration data.
[0020] Further, in the step 3, the extreme gradient boosting algorithm is used to construct the heavy metal element concentration prediction model.
[0021] Further, in the step 3, the construction process of the heavy metal element concentration prediction model includes the following steps:
[0022] Step 31: Take the obtained heavy metal concentration data Wm as the output factor, and take the screened physical and chemical index data Wp and nutrient index data Wn as the input factors;
[0023] Step 32: Divide the heavy metal concentration data, physical and chemical index data, and nutrient index data into a model training data group and a model verification data group;
[0024] Step 33: Gradually construct multiple weak prediction models, gradually optimize the weak prediction models through the loss function L(φ), and take the weighted sum of all weak prediction models as the final heavy metal element concentration prediction model.
[0025] Further, in the step 33, the loss function L(φ) is the difference between the predicted value and the true value, as shown in the following formula (3).
[0026]
[0027] In formula (3), is the loss function of the i-th given sample, y i is the true value of the i-th given sample, is the predicted value of the i-th given sample, m is the total number of samples; Ω(z k ) is the complexity penalty term of the loss function L(φ), k represents the index of the splitting node in the decision tree, and is used to represent the number of penalty terms of the model complexity. Control the regularization of the loss function L(φ) to prevent overfitting by controlling the model complexity.
[0028] Further, in the step 33, the weighted sum of all weak prediction models is shown in the following formula (8);
[0029]
[0030] In formula (8), is the predicted value of the i-th given sample, t represents the t-th iteration, T is the total number of iterations, η is the learning rate used to control the contribution of the new decision tree to the overall model, x i is the input feature of the i-th sample, f t is the decision tree at the t-th iteration.
[0031] Further, the prediction ability of the heavy metal concentration prediction model of formula (8) is evaluated using the correlation coefficient R, mean absolute error MAE, root mean square error RMSE, and index of agreement IA.
[0032] Compared with the prior art, the beneficial effects of the present invention are reflected in:
[0033] The present invention discloses a rapid detection method for heavy metal concentrations in the water body of a floating rice planting area, comprising the following steps: Step 1: Obtain the physical and chemical index data, nutrient index data, and heavy metal concentration data of the water body in the floating rice planting area; Step 2: Preprocess the obtained physical and chemical index data, nutrient index data, and heavy metal concentration data; Step 3: Construct a heavy metal element concentration prediction model; Step 4: Evaluate the prediction performance of the heavy metal element concentration prediction model constructed in Step 3.
[0034] Through machine learning algorithms, especially the Extreme Gradient Boosting (XGBoost) algorithm, such non-linear problems can be effectively processed to achieve efficient prediction of heavy metal concentrations. XGBoost has the ability to process multi-dimensional data and improves the prediction accuracy by optimizing the model parameters, making it an efficient means for heavy metal monitoring and prediction.
[0035] The rapid detection method for water body heavy metal concentration proposed by the present invention shows significant advantages in predicting the heavy metal concentration in the water body of the floating rice planting area. This method not only has an efficient and fast prediction process, but also significantly improves the prediction accuracy, providing a strong technical guarantee for the water quality safety of the rice planting area. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of a rapid detection method for heavy metal concentrations in the water body of a floating rice planting area according to the present invention.
[0037] The following is a further description of the present invention through specific embodiments and in conjunction with the drawings. SPECIFIC EMBODIMENTS
[0038] See Figure 1 , the present invention provides a rapid detection method for heavy metal concentrations in the water body of a floating rice planting area, comprising the following steps:
[0039] Step 1: Obtain the physical and chemical index data, nutrient index data, and heavy metal concentration data of the water body in the floating rice planting area;
[0040] Step 2: Preprocess the obtained physical and chemical index data, nutrient index data, and heavy metal concentration data;
[0041] Step 3: Construct a heavy metal element concentration prediction model;
[0042] Specifically, input the heavy metal element concentration data and the screened water quality index data into the Extreme Gradient Boosting (XGBoost) model, correct the preliminary measurement data based on the heavy metal element concentration, and predict the heavy metal element concentration through the water quality index data and the XGBoost model.
[0043] Step 4: Evaluate the prediction performance of the heavy metal element concentration prediction model constructed in Step 3.
[0044] A rapid detection method for heavy metal concentrations in the water body of a floating rice planting area according to the present invention directly relates the heavy metal concentration to the physical and chemical indexes and nutrient indexes in the target water body, and establishes a prediction model for evaluating the heavy metal concentration in the water body through the Extreme Gradient Boosting algorithm. Different from the conventional monitoring methods or mechanism model prediction methods in the past, it has certain practical value. By coupling the Extreme Gradient Boosting algorithm with the physical and chemical indexes and nutrient indexes of the water body, rapid monitoring of the heavy metal concentration in the water body is realized.
[0045] Specifically, in Step 1, the physical and chemical index data includes temperature T, dissolved oxygen DO, turbidity TU, and pH value.
[0046] Step S100: Obtain the physical and chemical index data of the water body. In-situ measure the physical and chemical indexes of the water body at 10 cm below the water surface at the sampling point using a handheld multi-parameter water quality analyzer YSI (model EXO2, YSI Inc., USA). The physical and chemical index Wp of the water body can be represented by the vector [DO, T, TU, pH].
[0047] The nutrient index data includes total nitrogen TN, total phosphorus TP, ammonium nitrogen NH4 + -N, nitrate nitrogen NO3 - -N, and orthophosphate PO4 3- -P;
[0048] Step S200: Obtain the nutrient index data of the water body. Analyze the nutrient indexes using the standards recommended by the State Environmental Protection Administration. The nutrient index Wn of the water body can be represented by the vector [TN, NH4 + -N, NO3 - -N, PO4 3- -P, TP]. The total nitrogen (TN) in the water sample is determined using the "Alkaline Potassium Persulfate Digestion UV Spectrophotometry" (HJ 636-2012); the ammonium nitrogen (NH4 + -N) is determined using the "Flow Injection-Salicylic Acid Spectrophotometry" (HJ 666-2013); the nitrate nitrogen (NO3 - -N) is determined using the "Phenoldisulfonic Acid Spectrophotometry" (GB7480-87); the orthophosphate (PO43- -P) and total phosphorus (TP).
[0049] The heavy metal concentration data includes beryllium (Be), cadmium (Cd), cobalt (Co), iron (Fe), nickel (Ni), lead (Pb), titanium (Ti) and zinc (Zn).
[0050] Step S300: Obtain the concentration of heavy metal elements in the water body through the detection of water samples. Digest the samples according to the EPA3010 method to prepare for the analysis of heavy metal elements in the water body. The heavy metal elements Wm in the water body can be represented by the vector [Be, Cd, Co, Fe, Ni, Pb, Ti, Zn]. First, add concentrated nitric acid (5 ml) to 100 ml of water sample, and evaporate the sample to a small volume (close to dryness) below the boiling point (95 °C). Then, digest and concentrate the sample with concentrated nitric acid, hydrochloric acid and hydrogen peroxide, and evaporate it to near dryness again at 95 °C. Finally, dilute the sample to 10 ml with 2% dilute nitric acid, and measure the volume-fixed solution on the machine. Standard solutions of Rh, In, and Ce with a concentration of 10 μg / L prepared with 2% nitric acid were used to optimize and control the quality of inductively coupled plasma mass spectrometry (ICP-MS) analysis. The concentrations of Be, Cd, Co, Ni, Pb, Ti, and Zn in the extract were measured using an inductively coupled plasma mass spectrometer (ICP-MS, Elan 9000, PerkinElmer); the Fe concentration was measured by inductively coupled plasma atomic emission spectrometry (ICP-AES, Perkin Elmer, Waltham, MA, USA).
[0051] Specifically, the physical and chemical indexes include one or more of T, DO, TU and pH value, and the nutrient indexes include one or more of TN, TP, NH4 + -N, NO3 - -N, and PO4 3- -P; the heavy metals include one or more of Be, Cd, Co, Fe, Ni, Pb, Ti and Zn. Each index includes but is not limited to the parameters listed above. Specifically, the index parameters can be increased or decreased according to the actual situation. At the same time, various numerical values of the initially obtained physical and chemical index data, nutrient index data and heavy metal concentration data are preprocessed to eliminate outliers and fill in missing values.
[0052] Specifically, in step 2, the process of the data preprocessing includes eliminating outliers and filling in missing values.
[0053] Before model construction, the initial data is comprehensively preprocessed. The physical and chemical indexes Wp, nutrient indexes Wn and heavy metal indexes Wm of the water body are processed separately.
[0054] Step S400: Initial data preprocessing. First, outliers are removed. For each index group, statistical methods are used to analyze and determine outliers and remove them. Second, missing values are filled. Since the physical and chemical indexes in Wp have strong spatio-temporal continuity, linear interpolation can be used to fill them. For Wn and Wm, due to possible influence of complex environments, a method combining interpolation and random forest regression is used for filling. Finally, when screening the data, Spearman correlation analysis and redundancy analysis (RDA) are used to screen out the indexes that have a significant impact on the model.
[0055] Step S410: Use statistical methods to analyze and determine outliers in each index data and remove them. Box plots visually display the distribution of normal data and the existence of outliers in a visual form, and are suitable for detecting outliers in data such as water body physical and chemical indexes. By showing the quartile distribution of the data, outliers are judged in combination with the upper and lower limits, and the data beyond this range is regarded as an outlier and removed.
[0056] In specific implementation, the linear interpolation method is used to fill the missing values of the water body physical and chemical indexes Wp; for the nutrient index Wn and the heavy metal index Wm, a method combining interpolation and random forest regression is used to fill the missing values.
[0057] Step S420: Use the linear interpolation method to fill the physical and chemical indexes in Wp. The linear interpolation method is a method that uses the linear relationship between known data points to estimate missing values. Assume that between two known data points (x1, y1) and (x2, y2), it is necessary to estimate the value y corresponding to a missing point x. The formula for linear interpolation is shown in formula (1).
[0058]
[0059] Step S430: Use a method combining interpolation and random forest regression to fill Wn and Wm. When using the method combining interpolation and random forest regression to fill missing values, first perform preliminary interpolation (linear interpolation) on the data to generate a complete sample set, and then use the random forest regression model to perform random forest regression training on the preliminarily interpolated data. Select other variables as feature variables and the variable corresponding to the missing value as the target variable; after the training is completed, use the model to predict and fill the missing values.
[0060] In specific implementation, the process of the data preprocessing further includes performing correlation analysis on the physical and chemical index data, nutrient index data, and heavy metal concentration data.
[0061] Step S440: Perform Spearman correlation analysis on the heavy metal concentrations, nutrients, and physical and chemical index data of all samples during the sampling period. Spearman correlation analysis evaluates the monotonic relationship between two variables by comparing their ranks (rankings). When calculating, first sort the data of each variable and assign rankings, and then compare the ranking differences corresponding to the two variables to measure their correlation. The magnitude of the rank correlation coefficient reflects the strength of the monotonic relationship between variables. Different from linear correlation, it can capture non-linear but monotonic relationships. When the rank correlation coefficient Rs is close to -1 or 1, it indicates a strong negative or positive correlation; however, an Rs value close to 0 does not necessarily mean that there is no relationship between the variables being compared.
[0062] Step S450: Use the Vegan package in R language to determine the relationship between heavy metal concentrations and other physical and chemical indicators and nutrients. Using redundancy analysis (RDA), determine the relationship between water physical and chemical indicators, nutrient indicators, and water body heavy metal concentrations. The environmental explanatory variables are water physical and chemical indicators (T, DO, TU, pH) and nutrient parameters (TN, TP, NH4 + -N, NO3 - -N, PO4 3- -P), and the response variable is the water body heavy metal concentration. The angle between the arrows representing variables indicates the correlation between variables. When the angle is acute, it indicates a positive correlation between variables, and when the angle is obtuse, it indicates a negative correlation between variables. The farther the variables are apart, the weaker the correlation between variables.
[0063] In specific implementation, in step 3, the extreme gradient boosting algorithm is used to construct a heavy metal element concentration prediction model.
[0064] Select the extreme gradient boosting (XGBoost) algorithm as the extreme gradient boosting model. Use the concentration index data of the selected water quality indicators Wp and Wn as input factors, and use the concentration index data of the heavy metal element Wm in the water sample as the output factor to construct the heavy metal element concentration prediction model; the goal of the heavy metal element concentration prediction model is to learn the mapping relationship from the input Wp and Wn to the output Wm, as shown in formula (2);
[0065] W m = F(W P ,W n ; θ) (2)
[0066] In formula (2), F() is the optimized XGBoost model function, and θ is the model parameter. θ is used to describe the weights of each branch in the XGBoost model and the structure of the decision tree.
[0067] In specific implementation, in step 3, the construction process of the heavy metal element concentration prediction model includes the following steps:
[0068] Step 31: Use the obtained heavy metal concentration data Wm as the output factor, and use the screened physical and chemical index data Wp and nutrient index data Wn as the input factors;
[0069] Step 32: Divide the heavy metal concentration data, physical and chemical index data, and nutrient index data into a model training data group and a model verification data group;
[0070] Randomly divide the heavy metal concentration data, screened physical and chemical index data, and nutrient index data into two groups. The model training data group accounts for 80% of the total data, and the model verification data group is 20%;
[0071] Use SPSS data statistical analysis software to randomly select 80% of the data as training data. Establish an extreme gradient boosting model using the "input method", and the remaining 20% of the data is used for verification. The water quality data includes 4 physical and chemical indexes and 5 nutrient indexes as input factors, and the concentrations of 8 heavy metal elements are used as output factors respectively.
[0072] Step 33: Gradually construct multiple weak prediction models, and gradually optimize the weak prediction models through the loss function L(φ). Take the weighted sum of all weak prediction models as the final heavy metal element concentration prediction model.
[0073] In the extreme gradient boosting algorithm, the weak prediction model is constructed through a preset extreme gradient boosting function. In the present invention, the physical and chemical index data and nutrient index data of the target water body are used as the input factors of the extreme gradient boosting function, and a one-class factor (nutrient or physical and chemical index) model and a two-class factor (nutrient and physical and chemical index) model are constructed using RStudio. During the extreme gradient boosting process, gradually construct multiple decision trees as weak prediction models, minimize the difference between the prediction result of each new decision tree and the prediction result of the previous decision tree, and optimize the weak prediction model step by step, thereby improving the prediction performance of the overall model. By gradually constructing multiple decision trees, minimize the error through the continuously optimized objective function of the decision tree.
[0074] In specific implementation, in step 33, the loss function L(φ) is the difference between the predicted value and the true value, as shown in formula (3).
[0075]
[0076] In formula (3), is the loss function of the i-th given sample, y i is the true value of the i-th given sample, is the predicted value of the i-th given sample, and m is the total number of samples; Ω(z k ) is the complexity penalty term of the loss function L(φ). k represents the index of the splitting node in the decision tree, which is used to represent the number of penalty terms for the model complexity. The regularization of the loss function L(φ) is controlled to prevent overfitting by controlling the model complexity.
[0077] The initial weak prediction model is set to a constant value, which is taken from the average value of the model training data set. In each iteration, the residual r i (i.e., the gradient) between the current predicted value of the weak prediction model and the true value is calculated. The formula for the residual r i is shown in Formula (4).
[0078]
[0079] In Formula (4), t represents the t-th iteration; y i is the true value of the i-th given sample, is the predicted value of the i-th given sample. is the loss function of the i-th given sample, y i is the true value of the i-th given sample, is the predicted value of the i-th given sample. is the loss function L(φ).
[0080] Based on the residual r i calculated by the above formula, a new decision tree f t (which is the decision tree for the t-th iteration) is generated. XGBoost (Extreme Gradient Boosting) fits the residual at the leaf nodes of each decision tree to minimize the error as much as possible.
[0081] The new decision tree f t is used to correct the previous predicted value, and the correction formula is shown in Formula (5) as follows.
[0082]
[0083] In Formula (5), η is the learning rate used to control the contribution of the new decision tree to the overall model, is the predicted value of the i-th given sample at the t-th iteration, is the predicted value of the i-th given sample at the (t - 1)-th iteration; f t is the decision tree at the t-th iteration; x i is the input feature of the i-th sample.
[0084] After generating the new decision tree, the optimization objective function is achieved through the second-order Taylor expansion approximation of the loss function L(φ). The loss function L at the t-th iteration(t) See the following formula (6).
[0085]
[0086] In formula (5), is the gradient (i.e., the residual) of the loss function L(φ) at the (t - 1)-th iteration, is the second derivative of the loss function L(φ) at the (t - 1)-th iteration; Ω(f t ) is the complexity penalty term of the loss function L(φ) at the t-th iteration; y i is the true value of the i-th given sample, is the predicted value of the i-th given sample, and m is the total number of samples.
[0087] For the decision tree at each iteration, the best split point is selected by maximizing the gain Gain. The calculation formula for maximizing the gain Gain is shown in the following formula (7).
[0088]
[0089] In formula (7), G L and G R are the gradients of the left and right nodes of the best split point respectively, and H L and H R are the second-order gradients of the left and right nodes of the best split point respectively. λ and γ are regularization parameters used to control the complexity of the model.
[0090] After repeating the above steps multiple times, a new decision tree is constructed until the early stopping condition is met. The final model is the weighted sum of the prediction results of all the trees.
[0091] In specific implementation, in step 33, the weighted sum of all weak prediction models is shown in the following formula (8);
[0092]
[0093] In formula (8), is the predicted value of the i-th given sample, t represents the t-th iteration, T is the total number of iterations, η is the learning rate used to control the contribution of the new decision tree to the overall model, x i is the input feature of the i-th sample, and f t is the decision tree at the t-th iteration;
[0094] After obtaining the final heavy metal element concentration prediction model, the prediction performance of the heavy metal element concentration prediction model is evaluated using the model verification data set to further verify the generalization ability of the heavy metal concentration prediction model and verify the accuracy of the model.
[0095] In specific implementation, the correlation coefficient R, mean absolute error MAE, root mean square error RMSE, and index of agreement IA are used to evaluate the prediction ability of the heavy metal concentration prediction model of formula (8).
[0096] In the present invention, four evaluation indexes, namely the correlation coefficient R, mean absolute error MAE, root mean square error RMSE, and index of agreement IA, are used to evaluate the prediction ability of the model. Among them, R is used to characterize the goodness of fit between the predicted value and the actual value of the model, MAE and RMSE are used to characterize the magnitude of the difference between the predicted value and the actual value of the model, and IA is used to test the similarity between the measured value and the simulated predicted value of the heavy metal element. Generally, the closer the R value and the IA index are to 1, and the closer the MAE value and the RMSE value are to 0, the more reliable the model is and the better the model simulation effect is. The calculation formulas of the correlation coefficient R, mean absolute error MAE, root mean square error RMSE, and index of agreement IA are shown in the following formulas (9) to (12).
[0097]
[0098] In formulas (9) to (12), n represents the number of data, and y i respectively represent the predicted value and the true value of the heavy metal element X concentration in the i-th given sample, represents the average value of the predicted values of the heavy metal element.
[0099] Table 1
[0100]
[0101] The above Table 1 is a list of the prediction performance evaluation of the extreme gradient boosting model of the present invention for the concentrations of different heavy metals (Be, Cd, Co, Fe, Ni, Pb, Ti, Zn) on the dataset of water quality indicators. Model I is a type of factor model, and the input factor is Wn.
[0102] The above Table 2 is a list of the prediction performance evaluation of the extreme gradient boosting model of the present invention for the concentrations of different heavy metals (Be, Cd, Co, Fe, Ni, Pb, Ti, Zn) on the dataset of water quality indicators. Model II is a type of factor model, and the input factor is Wp.
[0103] Table 2
[0104]
[0105] Generally speaking, the training R values of Cd and Co elements in Model I are between 0.7 and 0.8, and the validation R values are between 0.6 and 0.8. Among them, the element with the best simulation effect is Cd, with a training R value of 0.793 and a validation R value of 0.615. The training and validation IAs are 0.884 and 0.736 respectively. Model I has a good simulation effect on Cd and Co elements, among which Cd has the best simulation effect, and the simulation effect on Ni, Ti and Zn elements is the worst. In Model II, the training R values of Be, Co, Fe and Ti elements are all greater than 0.9, and the validation R values are also greater than 0.8. Among them, the element with the best simulation effect is Be, with a training R value reaching 0.932 and a validation R value of 0.906. The training and validation IAs are 0.961 and 0.906 respectively. The training R values of Ni, Pb and Zn elements are between 0.8 and 0.9, and the validation R values are between 0.7 and 0.9. Among all elements, the worst simulation result is Cd, but its training R value is 0.755, the validation R value is 0.723, and the training and validation IAs are 0.812 and 0.794 respectively. Model II has a good simulation effect on each heavy metal. The elements with the best simulation effect are Be, Co, Fe and Ti, and among them, Be element is the best. The simulation effect on Cd is relatively poor.
[0106] Table 3
[0107]
[0108] The above Table 3 is a list of the prediction performance evaluation of the extreme gradient boosting model of the present invention on the dataset of water quality indicators for the concentrations of different kinds of heavy metals (Be, Cd, Co, Fe, Ni, Pb, Ti, Zn). Model III is a two-factor model, and the input factors are Wn and Wp. The training R values of Be, Co, Fe and Ti elements are all greater than 0.9, and the validation R values are also all greater than 0.8. Except that the validation IA of Fe is 0.894, the training and validation IAs of other elements are all greater than 0.9. The training R values of Cd and Zn elements are between 0.8 and 0.9, and the validation R values are between 0.7 and 0.8. The relatively poor simulation results are for Ni and Pb elements, with training R values of 0.754 and 0.746 respectively, validation R values of 0.755 and 0.651 respectively, and training and validation IAs between 0.7 and 0.8. Generally speaking, Model 9 has the best simulation effect on Be, Co, Fe and Ti elements, and the highest training R value can reach 0.937. The simulation effect on Cd and Zn elements is good. The simulation effect on Ni and Pb elements is relatively the worst.
[0109] In summary, in Model I, the training R values of each heavy metal are between 0.559 and 0.793, and the average training R value is 0.652. The validation R values are between 0.572 and 0.781, and the average validation R value is 0.644. The simulation prediction results of Ni are the worst, with training and validation R values of 0.581 and 0.572 respectively. The training and validation R values of Be, Cd, and Co are all above 0.6. In Model II, the training R values of each heavy metal are between 0.755 and 0.932, and the average training R value is 0.877. The validation R values are between 0.723 and 0.926, and the average validation R value is 0.834. In Model III, the training R values of each heavy metal are between 0.746 and 0.939, and the average training R value is 0.856. The validation R values are between 0.651 and 0.877, and the average validation R value is 0.797. Therefore, when the input factors are Wn and Wp, the concentration prediction of Wm is more accurate.
[0110] Table 4
[0111]
[0112] Inputting the water quality index data information into the established extreme gradient boosting model can predict the heavy metal element concentrations. Table 4 is a list of the prediction accuracy results of the extreme gradient model of the present invention for the concentrations of different types of heavy metals (Be, Cd, Co, Fe, Ni, Pb, Ti, Zn) on the water quality index dataset.
[0113] A rapid detection method for the heavy metal concentrations in the water body of a floating rice planting area according to the present invention couples the physical and chemical indexes and nutrient salt indexes of the water body, and establishes an extreme gradient boosting model to achieve efficient prediction of the heavy metal concentrations, so as to meet the real-time detection requirements of heavy metals in the water body. The specific steps include: first, obtaining the physical and chemical indexes of the target water body such as temperature (T), dissolved oxygen (DO), turbidity (TU), pH, and nutrient salt indexes such as total nitrogen (TN), total phosphorus (TP), ammonia nitrogen (NH4+-N), nitrate nitrogen (NO3--N), orthophosphate (PO4 3--P), etc. and concentration data of target heavy metal elements such as beryllium (Be), cadmium (Cd), cobalt (Co), iron (Fe), nickel (Ni), lead (Pb), titanium (Ti), zinc (Zn), etc. Then, the extreme gradient boosting (XGBoost) algorithm is used to construct a machine learning prediction model. The physical and chemical indexes and nutrient indexes are used as inputs, and the heavy metal concentration is used as the output. The extreme gradient boosting algorithm is used to construct the prediction model to gradually optimize the loss function and reduce the prediction error. Finally, the heavy metal concentration is predicted based on the constructed model. In the model training stage, the training set and the validation set (80% training set, 20% validation set) are divided, and the RStudio software and SPSS tool are used to realize the construction and parameter optimization of the model. Finally, the model performance is verified by evaluation indexes such as the correlation coefficient, root mean square error, mean absolute error, and index of agreement. The rapid detection method provided by the present invention is efficient, low-cost, and real-time compared with the traditional laboratory chemical analysis, and can be used for the rapid detection of heavy metal concentrations in the water body of the "floating rice" planting area.
[0114] A rapid detection method for heavy metal concentrations in the water body of a floating rice planting area. After the heavy metal concentration prediction model is trained and mature, a direct connection is established between the heavy metal element concentration and the physical and chemical indexes and nutrient indexes in the water body. Through the extreme gradient boosting algorithm, a prediction model for evaluating the heavy metal element concentration in the water body is established, which is different from the conventional monitoring methods or mechanism model prediction methods in the past, and has certain practical value. By coupling the extreme gradient boosting model with the physical and chemical indexes and nutrient indexes of the water body, the rapid monitoring of the heavy metal element concentration in the water body is realized.
[0115] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0116] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A rapid detection method for heavy metal concentration in water bodies of floating rice planting areas, characterized in that: The steps include: Step 1: Obtain the physical and chemical index data, nutrient index data and heavy metal concentration data of the water body in the floating rice planting area; Step 2: Preprocess the obtained physical and chemical index data, nutrient index data and heavy metal concentration data; Step 3: Construct a heavy metal element concentration prediction model; Step 4: Evaluate the prediction performance of the heavy metal element concentration prediction model constructed in step 3.
2. The rapid detection method for heavy metal concentration in water bodies in floating rice planting areas according to claim 1 is characterized by: In step 1, the physical and chemical index data include temperature T, dissolved oxygen DO, turbidity TU and pH value; The nutrient index data include total nitrogen TN, total phosphorus TP, ammonia nitrogen NH4 + -N, nitrate nitrogen NO3 - -N and orthophosphate PO4 3- -P; The heavy metal concentration data include beryllium Be, cadmium Cd, cobalt Co, iron Fe, nickel Ni, lead Pb, titanium Ti and zinc Zn.
3. The rapid detection method for heavy metal concentration in water bodies of floating rice planting areas according to claim 1 is characterized by: In step 2, the data preprocessing process includes removing outliers and filling missing values.
4. The rapid detection method for heavy metal concentration in water bodies of floating rice planting areas according to claim 3 is characterized by: Linear interpolation was used to fill missing values for water body physical and chemical indicators Wp; interpolation combined with random forest regression was used to fill missing values for nutrient indicators Wn and heavy metal indicators Wm.
5. The rapid detection method for heavy metal concentration in water bodies in floating rice planting areas according to claim 3 is characterized by: The data preprocessing process also includes correlation analysis of physical and chemical index data, nutrient index data and heavy metal concentration data.
6. The rapid detection method for heavy metal concentration in water bodies in floating rice planting areas according to claim 1 is characterized by: In step 3, an extreme gradient boosting algorithm is used to construct a heavy metal element concentration prediction model.
7. The rapid detection method for heavy metal concentration in water bodies in floating rice planting areas according to claim 6 is characterized by: In step 3, the construction process of the heavy metal element concentration prediction model includes the following steps: Step 31: using the acquired heavy metal concentration data Wm as output factors, and using the screened physical and chemical index data Wp and nutrient index data Wn as input factors; Step 32: Divide the heavy metal concentration data, physical and chemical index data, and nutrient index data into a model training data group and a model verification data group; Step 33: gradually construct multiple weak prediction models, gradually optimize the weak prediction models through the loss function L(φ), and take the weighted sum of all weak prediction models as the final heavy metal element concentration prediction model.
8. The rapid detection method for heavy metal concentration in water bodies in floating rice planting areas according to claim 7 is characterized by: In step 33, the loss function L(φ) is the difference between the predicted value and the true value, as shown in the following formula (3). In formula (3), is the loss function of the i-th given sample, y i is the true value of the i-th given sample, is the predicted value of the i-th given sample, m is the total number of samples; Ω(z k ) is the complexity penalty term of the loss function L(φ), k represents the index of the split node in the decision tree, and is used to represent the number of penalty terms for model complexity.
9. The rapid detection method for heavy metal concentration in water bodies of floating rice planting areas according to claim 7 is characterized by: In step 33, the weighted sum of all weak prediction models is shown in the following formula (8): In formula (8), is the predicted value of the i-th given sample, t represents the t-th iteration, T is the total number of iterations, η is the learning rate used to control the contribution of the new decision tree to the overall model, and x i The input features of the i-th sample, f t is the decision tree at the tth iteration.
10. The method for rapid detection of heavy metal concentration in water bodies in floating rice planting areas according to claim 9, characterized in that: The correlation coefficient R, mean absolute error MAE, root mean square error RMSE and consistency index IA were used to evaluate the prediction ability of the heavy metal concentration prediction model of formula (8).