A total nitrogen and total phosphorus inversion method based on sample migration and lasso regression
By employing a total nitrogen and total phosphorus inversion method based on sample migration and lasso regression, and utilizing a water quality parameter-driven TrAdaboost transfer learning model, the problems of low detection efficiency and low accuracy in existing technologies are solved, achieving rapid and accurate monitoring of total nitrogen and total phosphorus while avoiding chemical pollution and high costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
- Filing Date
- 2022-12-21
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies have low detection efficiency and low accuracy in total nitrogen and total phosphorus. Manual chemical methods are cumbersome and prone to errors, while automated instrument methods are costly and prone to contamination. Optical remote sensing methods have poor model robustness and limited application scope.
A total nitrogen and total phosphorus inversion method based on sample transfer and lasso regression is adopted. Driven by water quality parameters, a TrAdaboost transfer learning model based on weak learner lasso algorithm is constructed, and the total nitrogen and total phosphorus are predicted by sample transfer learning algorithm, bypassing chemical processes and complex hardware structures.
It achieves rapid and accurate monitoring of total nitrogen and total phosphorus, with a speed increase of several thousand times and an accuracy of less than 10% relative error. It can provide early warning of eutrophication and algal blooms, avoiding chemical pollution and high costs.
Smart Images

Figure CN116092597B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of quantitative analysis of water quality parameters, specifically relating to a method for inverting total nitrogen and total phosphorus based on sample migration and lasso regression. Background Technology
[0002] Total nitrogen (TNI) generally refers to the sum of all organic and inorganic nitrogen in water, including inorganic nitrogen such as nitrate, nitrite, and ammonium, and organic nitrogen such as proteins, amino acids, and organic amines. Total phosphorus (TP) is a general term for orthophosphate, condensed sulfate, pyrophosphate, metaphosphate, and organically bound phosphate. Simply put, TNI and TTP are key indicators of eutrophication, and the primary control indicators for preventing eutrophication are TNI and TTP. Currently, methods for determining TNI and TTP are limited, mainly consisting of three types: artificial chemical methods, automated instrument methods, and remote sensing spectroscopy. Firstly, in artificial chemical methods, TNI is determined using the alkaline potassium persulfate digestion ultraviolet spectrophotometric method (GB11894-89); TTP is determined using the ammonium molybdate spectrophotometric method (GB11893-89). Secondly, automated instrument methods for TNI and TTP are also based on the national standard methods used in artificial chemical methods. Finally, with the development of optical remote sensing and machine learning technologies, remote sensing methods such as multispectral satellites and drones equipped with hyperspectral cameras have provided new ideas for the determination of total nitrogen and total phosphorus.
[0003] However, each of the three methods mentioned above has its drawbacks. The manual chemical method is cumbersome, highly specialized, prone to human error, and susceptible to secondary contamination. While automated instrument methods save labor and time, they cannot avoid secondary chemical contamination and still require an average analysis time of 1 hour; the complex chemical processes easily lead to cross-contamination between reagents, introducing significant uncertainty into the determination of total nitrogen and total phosphorus; and the sophisticated and complex structure of the instruments naturally results in high costs. The optical remote sensing method uses machine learning models to find complex coupling relationships in water spectra to estimate total nitrogen and total phosphorus. However, this method is cumbersome, suffers from poor spatiotemporal resolution of source data, has poor model robustness and transferability, and has limited application scope.
[0004] In summary, there is an urgent need to study a method for determining total nitrogen and total phosphorus with high detection efficiency and high accuracy. Summary of the Invention
[0005] The purpose of this invention is to provide a total nitrogen and total phosphorus inversion method based on sample migration and lasso regression, so as to solve the technical problems of low detection efficiency and low accuracy of measurement results in the prior art.
[0006] To achieve the above objectives, the present invention employs the following technical solution:
[0007] A method for retrieving total nitrogen and total phosphorus based on sample migration and lasso regression specifically includes the following steps:
[0008] Step 1: Collect water quality data from multiple stations and sort the water quality data from each station by time.
[0009] Step 2: Preprocess the water quality data obtained in Step 1;
[0010] Step 3: Divide the dataset: Divide the water quality data of each station obtained in Step 2 into a target domain training set and a target domain test set; at the same time, water quality data from stations with the same structure as the water quality data collected in Step 1 but different from those collected in Step 1 are used as the source domain auxiliary training set for that station.
[0011] Step 4: Construct a TrAdaboost transfer learning model based on the weak learner lasso algorithm. Use the combined source domain auxiliary training set and target domain training set obtained in Step 3 to train the constructed sample transfer learning model for the site, and obtain the trained sample transfer learning model for the site. Input the target domain test set for the site in Step 3 into the trained sample transfer learning model for the site to obtain the model output of the total nitrogen and total phosphorus results for the site.
[0012] Furthermore, in step 1, the water quality data includes water temperature, pH value, dissolved oxygen, conductivity, turbidity, permanganate index, ammonia nitrogen, total phosphorus, and total nitrogen.
[0013] Furthermore, step 2 is performed as follows: First, iterate through all the water quality data obtained in step 1, and determine whether there is a null value in the current data. If there is, delete the data; otherwise, keep the data. Then, use the SG smoothing algorithm to process the data.
[0014] Furthermore, in step 3, the portion of the water quality data obtained in step 2 that has more than 60% of the data from earlier times is used as the target domain training set, and the remaining portion is used as the target domain test set.
[0015] Furthermore, in step 4, when constructing the TrAdaboost transfer learning model based on the weak learner lasso algorithm, the weak learner is used as a hyperparameter in the TrAdaboost model.
[0016] Furthermore, the weak learner employs a lasso algorithm.
[0017] Compared with the prior art, the beneficial effects of the method of the present invention are as follows:
[0018] 1. It bypasses the complex chemical processes of the national standard method, saving reagents and avoiding chemical pollution. Driven by data purity, it eliminates the need for chemical reagents found in traditional national standard methods.
[0019] 2. This method provides a new approach for the inversion of total nitrogen and total phosphorus parameters, solving problems such as secondary chemical pollution, high cost, and cumbersome processes in existing technologies. It only requires several commercially available sensor probes monitoring seven water quality parameters (water temperature, pH, dissolved oxygen, conductivity, turbidity, permanganate index, and ammonia nitrogen) and a machine learning algorithm to complete the inversion and prediction of total nitrogen and total phosphorus parameters. It is purely data-driven, eliminating the need for complex hardware structures in automated instruments.
[0020] 3. Driven by water quality parameters, it employs non-spectral techniques, avoiding the tens or hundreds of dimensions found in spectral methods, thus preventing data redundancy, resulting in low feature dimensionality and high algorithm efficiency.
[0021] 4. The total nitrogen and total phosphorus detection speed is fast, reaching the second level. Compared with commercially available automated total nitrogen and total phosphorus analyzers, the speed is thousands of times faster. The method of this invention can complete the monitoring within a few seconds, while commercially available equipment for total nitrogen and total phosphorus takes an average of more than 1 hour.
[0022] 5. Excellent inversion accuracy; the vast majority of prediction results can guarantee a relative error of less than 10%. Based on the highly similar trends in total nitrogen and total phosphorus over time, obvious peaks can be observed, which can be used to warn of potential eutrophication and algal blooms in water bodies. The results from the examples demonstrate that the method of this invention possesses high performance.
[0023] 6. Using the lasso regression algorithm as a weak learner as the hyperparameter of the sample transfer learning model, the prediction accuracy was improved by utilizing the sample transfer learning algorithm TrAdaboost. Attached Figure Description
[0024] Figure 1 This is a flowchart of the method of the present invention.
[0025] Figure 2 This is a comparison chart of the predicted total nitrogen concentration at Bengbu Gate from April to June 2022.
[0026] Figure 3 This is a comparison chart of the predicted total phosphorus concentration of the Tuojiang Bridge from April to June 2022.
[0027] Figure 4 This is a comparison chart of the predicted total phosphorus concentration in Wangjiaba from April to May 2022.
[0028] Figure 5 This is a comparison chart of the predicted total phosphorus concentration in Zongguan from April to June 2022.
[0029] Figure 6 This is a comparison chart of the predicted total phosphorus concentrations in Zhutuo from May to June 2022.
[0030] The present invention will be further explained and described below with reference to the accompanying drawings and specific embodiments. Detailed Implementation
[0031] Various water quality parameters in water bodies exhibit highly complex coupling relationships, which can be automatically explored using machine learning. The National Surface Water Quality Automatic Monitoring Real-Time Data Release System displays real-time water quality data for key sections of various water bodies across China, including stable 4-hour interval data for water temperature (°C), pH (dimensionless), dissolved oxygen (mg / L), conductivity (μS / cm), turbidity (NTU), permanganate index (mg / L), ammonia nitrogen (mg / L), total phosphorus (mg / L), and total nitrogen (mg / L). Total nitrogen and total phosphorus can be used as labels, with the other seven water quality parameters as variables, to study suitable inversion algorithms for total nitrogen and total phosphorus. Currently, the aforementioned seven water quality probes are commercially available and readily integrated and deployed in locations requiring total nitrogen and total phosphorus monitoring.
[0032] The total nitrogen and total phosphorus inversion method based on sample migration and lasso regression provided in this invention specifically includes the following steps:
[0033] Step 1: Collect water quality data from multiple stations and sort the water quality data from each station by time.
[0034] Specifically, the water quality data includes water temperature (°C), pH value (dimensionless), dissolved oxygen (mg / L), conductivity (μS / cm), turbidity (NTU), permanganate index (mg / L), ammonia nitrogen (mg / L), total phosphorus (mg / L), and total nitrogen (mg / L).
[0035] Step 2: Preprocess the water quality data obtained in Step 1.
[0036] Specifically, the process is as follows: First, iterate through all the water quality data obtained in step 1 and determine whether there is a null value in the current data. If there is, delete the data; otherwise, keep the data. Then, use the SG smoothing algorithm to process the data, thereby replacing any possible outliers in the water quality data with surrounding values.
[0037] Step 3: Dataset Splitting: For each station, the portion of the water quality data obtained in Step 2 that has more than 60% of the data from earlier times is used as the target domain training set, and the remaining portion is used as the target domain test set. Additionally, water quality data from stations different from those collected in Step 1 but with the same data structure as those collected in Step 1 (including water temperature (°C), pH (dimensionless), dissolved oxygen (mg / L), conductivity (μS / cm), turbidity (NTU), permanganate index (mg / L), ammonia nitrogen (mg / L), total phosphorus (mg / L), and total nitrogen (mg / L)) are all used as the source domain auxiliary training set for that station.
[0038] Step 4: Construct a TrAdaboost transfer learning model based on the weak learner lasso algorithm. For each site, the constructed sample transfer learning model is trained using the combined source domain auxiliary training set and target domain training set obtained in Step 3, resulting in a trained sample transfer learning model for that site. The target domain test set from Step 3 is then input into the trained sample transfer learning model to obtain the model output for the total nitrogen and total phosphorus results for that site. This concludes the total nitrogen and total phosphorus inversion method.
[0039] Specifically, in this invention, when constructing the TrAdaboost transfer learning model based on the weak learner lasso algorithm, the existing TrAdaboost model used is the TrAdaboost model in Pardoe's paper "Boosting for Regression Transfer" (ICML 2010). In this TrAdaboost model, the weak learner (specifically the decision tree algorithm) is used as a hyperparameter. In this embodiment, the lasso algorithm is selected as this hyperparameter.
[0040] The reason for using the lasso algorithm here is that the coupling relationships of water quality parameters among different types of water bodies are very complex. Furthermore, the potential for malfunctions in water quality monitoring probes can negatively impact the total nitrogen and total phosphorus retrieval model due to the potential for certain features to malfunction. This invention makes the lasso regression algorithm a suitable base learner for the total nitrogen and total phosphorus retrieval model because it uses L1 regularization to dominate sparsity, thus compressing the coefficients of unimportant variables to zero. This achieves both relatively accurate parameter estimation and feature selection, i.e., dimensionality reduction.
[0041] Example:
[0042] Step 1: Since the webpage of the National Surface Water Quality Automatic Monitoring Real-time Data Release System only displays data for the current day and lacks historical storage functionality, this embodiment utilizes a web crawler script written in Python to crawl and store the data locally. Water quality data from five stations were collected: Bengbu Sluice in Anhui Province, located in the Huai River Basin (1210 data entries from February 28, 2021 to December 31, 2021, and 325 data entries from April 18, 2022 to June 14, 2022); Tuojiang Bridge in Sichuan Province, located in the Yangtze River Basin (1450 data entries from February 27, 2021 to December 31, 2021, and 333 data entries from April 18, 2022 to June 14, 2022); and Wangjiaba in Anhui Province, located in the Huai River Basin (data entries from February 27, 2021 to December 31, 2021). 1944 data points, 150 data points from April 18, 2022 to May 13, 2022; Zongguan in Hubei Province, located in the Yangtze River Basin, has 1624 data points from February 27, 2021 to December 31, 2021, and 328 data points from April 18, 2022 to June 14, 2022; Zhutuo in Chongqing Municipality, located in the Yangtze River Basin, has 785 data points from February 28, 2021 to December 31, 2021, and 142 data points from May 5, 2022 to June 14, 2022; 56873 water quality data points from 24 key sections in the Yangtze River Basin in 2021.
[0043] Step 2: Iterate through all the water quality data obtained in Step 1, and determine whether there is a null value in the current data. If there is, delete the data; otherwise, keep the data. Use the SG smoothing algorithm to process the data.
[0044] Step 3: Divide the water quality data of each station in Step 1 into a target domain training set, a target domain test set, and a source domain auxiliary training set for each station. Specifically, for each station, the water quality data of 2021 is used as the target domain training set, the water quality data of 2022 is used as the target domain test set, and the water quality data of 24 key sections of the Yangtze River Basin in 2021 is used as the source domain auxiliary training set.
[0045] Step 4: Construct a TrAdaboost transfer learning model based on the weak learner lasso algorithm. For each site, the constructed sample transfer learning model is trained using the combined source domain auxiliary training set and target domain training set obtained in Step 3, resulting in a trained sample transfer learning model for each site. The target domain test set from Step 3 is then input into the trained sample transfer learning model for that site to obtain the model output of the total nitrogen and total phosphorus results for that site. This concludes the total nitrogen and total phosphorus inversion method.
[0046] When constructing the TrAdaboost transfer learning model based on the weak learner lasso algorithm, the existing TrAdaboost model used is the TrAdaboost model in Pardoe's paper "Boosting for Regression Transfer" (ICML 2010). In this TrAdaboost model, the type of weak learner (specifically, the decision tree algorithm) is used as a hyperparameter. In this embodiment, the lasso algorithm is selected as this hyperparameter.
[0047] To demonstrate the feasibility and effectiveness of the method of this invention, a comparative evaluation experiment was added in step 4: the common Adaboost model was selected as the comparison model, that is, while implementing step 4, the benchmark model Adaboost corresponding to TrAdaboost was simultaneously constructed (the implementations of Adaboost and the lasso algorithm directly call the sklearn library). The hyperparameters of these two models are consistent, and their hyperparameter n_estimators are both selected as 50. The weak learner is the lasso algorithm, and the hyperparameter alpha of the lasso algorithm is both selected as 0.001. In particular, the other hyperparameters steps and fold of TrAdaboost are selected as 10 and 5, respectively. The model is trained using the target domain training set obtained in step 3 to obtain the trained Adaboost model. The target domain test set from step 3 is input into the model to obtain the model output of the total nitrogen and total phosphorus results. The mean squared error (MSE) and coefficient of determination (R²) are selected. 2 Using accuracy (Acc) as the evaluation metric, the actual total nitrogen and total phosphorus concentrations in the target domain test set and the predicted total nitrogen and total phosphorus concentrations output by the two models are calculated and substituted into the following formula. The performance differences before and after model migration are compared and visualized graphically. Wherein:
[0048]
[0049]
[0050]
[0051]
[0052] y i These are measured values. It is a predicted value. It measures the average value, RE is the relative error, and N is the sample size. a It is the number of samples with a relative error of less than 10%.
[0053] The following five tables compare the performance of five websites using the Adaboost model and the TrAdaboost model.
[0054] Table 1 Performance Comparison of Bengbu Gate Station
[0055] Bengbu Gate / TN MSE <![CDATA[R 2 ]]> Acc Adaboost 0.194 0.66 0.93 TrAdaboost 0.179 0.69 0.95
[0056] Table 2 Performance Comparison of Tuojiang Bridge Stations
[0057] Tuojiang Bridge / TP MSE <![CDATA[R 2 ]]> Acc Adaboost 0.00024 0.57 0.59 TrAdaboost 0.00011 0.8 0.83
[0058] Table 3 Performance Comparison of Wangjiaba Station
[0059] Wangjiaba / TP MSE <![CDATA[R 2 ]]> Acc Adaboost 0.0001259 0.867 0.813 TrAdaboost 0.0001116 0.882 0.832
[0060] Table 4 Performance Comparison of Zhutuo Station
[0061] Zhu Tuo / TP MSE <![CDATA[R 2 ]]> Acc Adaboost 0.000274 0.48 0.58 TrAdaboost 0.000144 0.72 0.8
[0062] Table 5 Performance Comparison of Zongguan Sites
[0063] Zongguan / TP MSE <![CDATA[R 2 ]]> Acc Adaboost 0.000496 0.41 0.73 TrAdaboost 7.871e-5 0.87 0.96
[0064] From Table 1-5 and Figure 2-6 As can be seen, the TrAdaboost model, which utilizes water quality data from 24 key locations in 2021, shows significant performance improvements, with accuracies exceeding 0.8 and more similar curve trends. The TrAdaboost model maximizes the use of water quality data with similar distributions to specific sites, significantly improving the prediction accuracy of total nitrogen and total phosphorus.
Claims
1. A method for retrieving total nitrogen and total phosphorus based on sample migration and lasso regression, characterized in that, Specifically, the steps include the following: Step 1: Collect water quality data from multiple stations and sort the water quality data from each station according to time; the water quality data includes water temperature, pH value, dissolved oxygen, conductivity, turbidity, permanganate index, ammonia nitrogen, total phosphorus, and total nitrogen; Step 2: Preprocess the water quality data obtained in Step 1; the specific operation is as follows: First, iterate through all the water quality data obtained in Step 1, and determine whether there is a null value in the current data. If there is, delete the data; otherwise, keep the data. Use the SG smoothing algorithm to process the data. Step 3: Divide the dataset: Divide the water quality data of each station obtained in Step 2 into a target domain training set and a target domain test set; at the same time, water quality data from stations with the same structure as the water quality data collected in Step 1 but different from those collected in Step 1 are used as the source domain auxiliary training set for that station; among them, the portion of the water quality data obtained in Step 2 with more than 60% of the data from earlier times is used as the target domain training set, and the remaining portion is used as the target domain test set. Step 4: Construct a TrAdaboost transfer learning model based on the weak learner lasso algorithm. For each site, the constructed sample transfer learning model is trained using the combined source domain auxiliary training set and target domain training set obtained in Step 3, resulting in a trained sample transfer learning model for that site. The target domain test set obtained in Step 3 is then input into the trained sample transfer learning model for that site to obtain the model output of the total nitrogen and total phosphorus results for that site. When constructing the TrAdaboost transfer learning model based on the weak learner lasso algorithm, the weak learner is used as a hyperparameter in the TrAdaboost model; the weak learner employs the lasso algorithm.