Monitoring method for high-frequency real-time monitoring of TP pollution condition

Through the combined use of multiple sensors and random forest model analysis, high-frequency real-time monitoring of TP pollution is achieved, which solves the problem of the inability to detect water pollution in a timely manner in existing technologies and improves the accuracy and timeliness of water quality monitoring.

CN120685873APending Publication Date: 2025-09-23CHINESE RES ACAD OF ENVIRONMENTAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510664067.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-23

Smart Images

  • Figure CN120685873A_ABST
    Figure CN120685873A_ABST
Patent Text Reader

Abstract

The invention discloses water quality monitoring, and relates to the technical field of monitoring methods for high-frequency real-time monitoring of TP pollution conditions. The method comprises the following steps: step a, arranging a plurality of sensors and monitoring cameras at a plurality of monitoring point positions of a river channel water body monitoring point, and transmitting water body data shot and collected by the sensors and the monitoring cameras to an analysis terminal; by optimizing the data transmission and processing process, the monitoring period can be shortened to 1 min per time, before data is used, water quality data is subjected to data cleaning, abnormal values generated by machine faults are eliminated, abnormal data generated by rainfall and other conditions are reserved, the model precision is enhanced, after the data are uploaded to the cloud platform, the cloud platform automatically organizes the data, and the data are stored in the cloud platform. And a data calibration training model can be obtained through regular manual sampling, the model precision is continuously improved, and high-frequency real-time monitoring of the river TP concentration is realized through a machine learning algorithm in combination with an intelligent sensor technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water quality monitoring, and in particular to a monitoring method for high-frequency and real-time monitoring of TP pollution. Background Art

[0002] Compared with more frequent measurements of hydrological and meteorological variables, traditional water quality monitoring is usually conducted at a coarser temporal resolution. Due to the high cost of water quality testing, monitoring is usually carried out on a monthly or biweekly basis. Although the effectiveness of traditional monitoring programs is unquestionable, short-term, extreme and unpredictable events with time scales shorter than the sampling frequency of traditional monitoring programs are often not detected. This limitation makes it difficult to study the impact of extreme events on water quality because the sampling dates often do not coincide with the dates of occurrence of such events. Many studies in this field either focus on specific seasons or have limited sample sizes, which makes it difficult to study the impact of phosphorus pollutants on river water quality during heavy rainfall periods during the migration of phosphorus pollutants into river channels.

[0003] Current monitoring of water quality is often too limited in time and space to detect and address factors that influence the development of harmful events, such as harmful algal blooms, oxygen depletion, and fish kills.

[0004] To this end, we provide a high-frequency real-time monitoring method for TP pollution to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a monitoring method for TP pollution in real time at a high frequency. By using a combination of multiple sensors to collect water body data in real time and combining it with historical pollution data of river water bodies, the present invention solves the problem that water quality monitoring in the existing technology is usually too limited in time and space scales and cannot detect and predict water pollution in a timely manner.

[0006] To solve the above technical problems, the present invention is achieved through the following technical solutions:

[0007] The present invention provides a high-frequency, real-time monitoring method for TP pollution, comprising the following steps: Step a: installing multiple sensors and monitoring cameras at multiple monitoring points in a river water monitoring point, transmitting the water data captured and collected by the sensors and monitoring cameras to an analysis terminal, establishing a database based on historical data, grading the water pollution situation, and setting an early warning value for each level; Step b: analyzing the historical water quality index monitoring data of the monitoring points to determine appropriate water quality indicators. The analysis method is based on the eigenvalue importance analysis of the random forest model, and the formula is as follows: Among them, the out-of-bag data (OOB) error is calculated for each decision tree in the random forest model using the corresponding out-of-bag data (errOOB1); noise interference is randomly added to all sample features of the out-of-bag data OOB, and the out-of-bag data error is calculated again, which is errOOB2; the number of random forest model trees is N; step c: automatic water quality monitoring equipment is set at the monitoring point to obtain the determined conventional water quality index data; step d: the collected water quality data is cleaned to solve missing values ​​and outliers; step e: the dissolved oxygen content, algae and aquatic grass density, water color and water light intensity of the water body are monitored in real time, and the collected data are summarized, cleaned and missing values ​​and outliers are solved. The analysis terminal analyzes the data and compares the new data with the historical data. The old data is used to generate a water pollution timeline, and pollution nodes are preset on the timeline. The pollution situation of the water body in each time period is predicted based on the new data; step f: the water quality data is processed using a machine learning algorithm based on the random forest model to invert TP concentration data. The random forest model formula is: Where T is the number of trees, and ht(x) is the prediction result of the t-th tree for the input x. Step g: Generate a chart of water oxygen content, algae and aquatic plant density, and water turbidity based on the summarized data, so that staff can judge the water pollution situation based on the chart trend. The analysis terminal compares the chart data with the historical data trend to automatically judge the current water pollution situation. Step h: Invert the TP concentration data and upload it to the cloud platform and user terminal. TP concentration data is obtained through regular manual sampling to calibrate the training model and continuously improve the model prediction accuracy.

[0008] The present invention is further configured such that the sensors in step a are respectively arranged at three sections of the river channel A, B, and C, and the sensors include a dissolved oxygen content sensor, a high-definition camera, a turbidity sensor, a TP sensor, a temperature sensor, and a pH sensor.

[0009] The present invention is further configured such that the multiple sensors and high-definition cameras in step a transmit data to the analysis terminal in real time via a wireless communication module, and statistically integrate the historical pollution data of the river water body to form a database.

[0010] The present invention is further configured such that the data cleaning method in step d is based on Spearman correlation analysis and Z-score method. After screening out abnormal values ​​of indicators, whether to retain them is determined based on the changes in other indicators with strong correlation with them. The relevant formula is as follows:

[0011]

[0012] Where ρ is the Spearman correlation coefficient, which is used to measure the monotonic relationship between two variables; di is the absolute value of the difference between the ranks (or orders) of the two variables in the i-th pair of observations; and n is the total number of observations.

[0013]

[0014] Among them, x is the specific data, μ is the mean of the data set, and σ is the standard deviation.

[0015] The present invention is further configured such that in step e, the various data collected in real time are summarized and aggregated to form a comprehensive data set, and the analysis terminal performs in-depth analysis on the comprehensive data set, including data cleaning, outlier processing, etc., compares the newly collected data with the data in the historical database, analyzes their changing trends and abnormal conditions, and predicts the water pollution situation based on the pollution nodes preset on the timeline.

[0016] The present invention is further configured such that statistical information such as the average value, maximum value and minimum value of each pollution level in the historical data, as well as the comparison between the current data point and these statistical information, will be marked on the water parameter chart in step g. When the data point of a monitoring indicator approaches or exceeds the preset warning value, the analysis terminal will automatically highlight the data point on the chart and issue a corresponding warning signal so that the staff can take timely countermeasures.

[0017] The present invention is further configured such that the eigenvalue importance analysis based on the random forest model in step b must ensure that the random forest model used is accurate and fully trained so as to accurately evaluate the importance of each water quality indicator and the representativeness of the data. The historical data used to train the random forest model should be representative and able to cover various water quality conditions.

[0018] The present invention is further configured as follows: in the training and verification of the model in step f, when using the random forest model to invert TP concentration data, it is necessary to ensure that the model has been fully trained and verified to improve the prediction accuracy; and for data preprocessing, before inputting the data into the model, the water quality data must be properly preprocessed, such as normalization and denoising, to improve the prediction performance of the model.

[0019] The present invention is further configured to ensure clarity of the chart in step g, that the generated chart should be clear and easy to read, and be able to intuitively reflect the changing trends and abnormal conditions of water quality parameters, and the accuracy of the early warning signal. When the data point of a certain monitoring indicator approaches or exceeds the preset early warning value, the analysis terminal should accurately issue an early warning signal so that the staff can take timely countermeasures.

[0020] The present invention is further configured such that manual sampling is performed regularly in step h to obtain accurate TP concentration data for calibrating the training model and continuously improving the model based on newly acquired data and actual conditions to improve prediction accuracy and reliability.

[0021] The present invention has the following beneficial effects:

[0022] 1. The present invention can shorten the monitoring cycle to 1 minute per time by optimizing the data transmission and processing process. It cleans the water quality data before using it, eliminates abnormal values ​​caused by machine failures, and retains abnormal data caused by conditions such as precipitation, thereby enhancing the model accuracy. After uploading the data to the cloud platform, the cloud platform automatically organizes the data and can obtain data through regular manual sampling to calibrate the training model, continuously improving the model accuracy. Through machine learning algorithms and combined with intelligent sensor technology, high-frequency real-time monitoring of river TP concentrations can be achieved, which can provide sufficient data for river TP pollution flux accounting and change research.

[0023] 2. The present invention uses a variety of different types of water quality monitoring sensors to monitor water pollution in real time, collect a variety of water data in real time, and improve the prediction accuracy of water pollution. Once the oxygen content in the water body continues to decrease, the turbidity of the water body continues to increase, and algae and aquatic plants reproduce in large numbers, it indicates that the severity of phosphorus pollution in the water body is getting higher and higher. Compared with monitoring the water body with a single TP sensor, the present invention can integrate multiple data and improve the accuracy of TP pollution monitoring of the water body. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present invention will be described below. The described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0026] Example 1

[0027] A monitoring method for high-frequency real-time monitoring of TP pollution includes the following steps: step a: setting multiple sensors and monitoring cameras at multiple monitoring points of river water monitoring points, transmitting water data captured and collected by the sensors and monitoring cameras to an analysis terminal, establishing a database based on historical data, grading water pollution conditions, and setting warning values ​​for each level. Sensors are arranged in three sections A, B, and C of the river, respectively. The sensors include a dissolved oxygen content sensor, a high-definition camera, a turbidity sensor, a TP sensor, a temperature sensor, and a pH sensor. The multiple sensors and high-definition cameras transmit data to the analysis terminal in real time via a wireless communication module, and statistically integrate historical river water pollution data to form a database; step b: analyzing historical water quality index monitoring data of the monitoring points to determine appropriate water quality indicators. The analysis method is based on eigenvalue importance analysis of a random forest model, and the formula is as follows: Among them, the out-of-bag data (OOB) error is calculated for each decision tree in the random forest model using the corresponding out-of-bag data (errOOB1); noise interference is randomly added to all sample features of the out-of-bag data OOB, and the out-of-bag data error is calculated again, which is errOOB2; the number of random forest model trees is N, and the eigenvalue importance analysis based on the random forest model is to ensure that the random forest model used is accurate and fully trained so that the importance of each water quality indicator can be accurately evaluated. The representativeness of the data, the historical data used to train the random forest model should be representative and can cover various water quality conditions; step c: set up automatic water quality monitoring equipment at the monitoring point to obtain the determined routine water quality indicator data; step d: clean the collected water quality data to solve missing values ​​and outliers. The data cleaning method is based on Spearman correlation analysis and Z-score method. After screening out the outliers of the indicators, it is determined whether they need to be retained based on the changes in other indicators with strong correlation. The relevant formula is as follows: Where ρ is the Spearman correlation coefficient, which is used to measure the monotonic relationship between two variables; di is the absolute value of the difference between the rankings (or orders) of the two variables in the i-th pair of observations; n is the total number of observations, Where x is the specific data, μ is the mean of the data set, and σ is the standard deviation; Step e: Real-time monitoring of the dissolved oxygen content, algae and aquatic plant density, water color, and water light intensity of the water body. After the collected data are summarized and cleaned and missing values ​​and outliers are resolved, the analysis terminal analyzes the data and compares the new data with the historical data. The old data is used to generate a water pollution timeline, and pollution nodes are preset on the timeline. The pollution status of the water body in each time period is predicted based on the new data. The various data collected in real time are summarized and aggregated to form a comprehensive data set. The analysis terminal conducts in-depth analysis of the comprehensive data set, including data cleaning, outlier processing, etc. The newly collected data is compared with the data in the historical database, and its change trend and anomalies are analyzed. The water pollution status is predicted based on the pollution nodes preset on the timeline; Step f: The machine learning algorithm based on the random forest model is used to process the water quality data to invert TP concentration data. The formula of the random forest model is: Where T is the number of trees, ht(x) is the prediction result of the t-th tree for the input x, model training and verification, when using the random forest model to invert TP concentration data, it is necessary to ensure that the model has been fully trained and verified to improve the prediction accuracy, data preprocessing, before inputting the model, the water quality data must be properly preprocessed, such as normalization, denoising, etc., to improve the prediction performance of the model; step g: Generate charts for water oxygen content, algae and aquatic plant density and water turbidity based on the summarized data, so that staff can judge the water pollution situation according to the chart trend. The analysis terminal compares the chart data with the historical data trend, automatically judges the current water pollution situation, and the pollution levels in the historical data will be marked on the water parameter chart. Statistical information such as the average, maximum and minimum values, as well as the comparison between the current data point and these statistical information. When the data point of a monitoring indicator approaches or exceeds the preset warning value, the analysis terminal will automatically highlight the data point on the chart and issue a corresponding warning signal so that the staff can take timely response measures; Step h: After inverting the TP concentration data, upload it to the cloud platform and user terminal, obtain TP concentration data through regular manual sampling, calibrate the training model, and continuously improve the model prediction accuracy. Regular manual sampling should be carried out regularly to obtain accurate TP concentration data for calibrating the training model. Continuous improvement of the model should be carried out according to the newly acquired data and actual conditions to continuously improve and optimize the model to improve prediction accuracy and reliability.

[0028] Example 2

[0029] Implementation background: The river water in a certain area is threatened by TP (total phosphorus) pollution. In order to monitor and control the pollution in a timely manner, it was decided to adopt the above-mentioned patented high-frequency real-time monitoring method of TP pollution.

[0030] Implementation steps:

[0031] Sensor layout and data collection: Dissolved oxygen content sensors, high-definition cameras, turbidity sensors, TP sensors, temperature sensors and pH sensors are respectively arranged at the three sections A, B and C of the river. These sensors transmit data to the analysis terminal in real time through wireless communication modules to form a database.

[0032] Determination of water quality indicators: Based on the importance analysis of the eigenvalues ​​of the random forest model, appropriate water quality indicators are screened from the historical database to ensure that the historical data used to train the random forest model is representative and can cover various water quality conditions.

[0033] Automatic water quality monitoring: Automatic water quality monitoring equipment is set up at monitoring points to obtain routine water quality index data.

[0034] Data cleaning and processing: The collected water quality data were cleaned to resolve missing values ​​and outliers. Spearman correlation analysis and Z-score method were used to screen outliers, and the need to retain them was determined based on the changes in other indicators with strong correlations.

[0035] Real-time monitoring and early warning: The system monitors indicators such as dissolved oxygen content, algae and aquatic plant density, water color, and light intensity in real time. Newly collected data is compared with historical data, and a water pollution timeline is generated using old data. Pollution nodes are preset on the timeline. When a data point for a monitoring indicator approaches or exceeds the preset warning value, the analysis terminal automatically highlights the data point on the chart and issues an early warning signal.

[0036] TP concentration data inversion: Water quality data is processed using a machine learning algorithm based on a random forest model to invert TP concentration data. Before inputting the data into the model, the water quality data is preprocessed by normalization and denoising to improve the model's prediction performance.

[0037] Chart generation and response: Generate water parameter charts based on the summarized data, marking statistical information such as the average, maximum and minimum values ​​of each pollution level in the historical data. Staff will judge the water pollution situation based on the chart trends and take corresponding response measures.

[0038] Model calibration and training: Manual sampling is performed regularly to obtain accurate TP concentration data for calibration and training of the model. Based on the newly acquired data and actual conditions, the model is continuously improved and optimized to improve prediction accuracy and reliability.

[0039] Implementation effect: Through the implementation of this method, the TP pollution in the river water bodies in the region has been timely and effectively monitored and controlled, and the water quality has been significantly improved.

[0040] Example 3

[0041] Implementation Background: A certain city's river is affected by industrial wastewater discharge and is seriously polluted by TP. In order to strengthen water quality monitoring and improve pollution control efficiency, it was decided to adopt the above-mentioned patented high-frequency real-time monitoring method for TP pollution.

[0042] Implementation steps:

[0043] Monitoring point layout: Multiple monitoring points are set up at key locations of the river, each equipped with multiple sensors and high-definition cameras to collect water data in real time.

[0044] Data integration and analysis: The data collected by sensors and cameras are transmitted to the analysis terminal, and the historical pollution data of the river water body are statistically integrated to form a database. The importance of water quality indicators is analyzed based on the random forest model to determine the key indicators.

[0045] Establishment of a real-time monitoring and early warning system: Establish a real-time monitoring and early warning system to monitor indicators such as dissolved oxygen, algae density, and water color in real time, compare and analyze new data with historical data, and predict water pollution trends.

[0046] TP concentration inversion and verification: The random forest model is used to invert TP concentration data, and accurate data is obtained through regular manual sampling to verify and adjust the model.

[0047] Chart generation and decision support: Generate charts based on the summarized data to show the changing trends and abnormal conditions of water quality parameters, providing intuitive data support for decision makers so that they can take timely pollution control measures.

[0048] Continuous model optimization: Based on newly acquired data and actual conditions, the random forest model is continuously optimized to improve prediction accuracy and reliability.

[0049] Implementation effect: Through the implementation of this method, TP pollution in the city's rivers has been effectively controlled, and the water quality has gradually improved. At the same time, this method provides strong data support for decision makers and improves the efficiency and effectiveness of pollution control.

[0050] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can better understand and utilize the present invention.

Claims

1. A method for high-frequency real-time monitoring of TP pollution, characterized by: The following steps are involved: Step a: Install multiple sensors and monitoring cameras at multiple monitoring points in the river water monitoring area. Transmit the water data captured and collected by the sensors and monitoring cameras to the analysis terminal. Build a database based on historical data, classify the water pollution situation, and set warning values ​​for each level. Step b: Analyze the historical water quality index monitoring data of the monitoring point to determine the appropriate water quality index. The analysis method is based on the importance analysis of the eigenvalues ​​of the random forest model. The formula is as follows: Among them, the out-of-bag data (OOB) error is calculated for each decision tree in the random forest model using the corresponding out-of-bag data, which is errOOB1; noise interference is randomly added to all sample features of the out-of-bag data OOB, and the out-of-bag data error is calculated again, which is errOOB2; the number of random forest model trees is N; Step c: setting up automatic water quality monitoring equipment at the monitoring point to obtain the determined conventional water quality index data; Step d: Clean the collected water quality data to resolve missing values ​​and outliers; Step e: Real-time monitoring of the dissolved oxygen content, algae and aquatic plant density, water color, and light intensity of the water body is performed. The collected data is summarized and aggregated, cleaned, and missing and outliers are resolved. The analysis terminal then analyzes the data and compares the new data with historical data. A water pollution timeline is generated using the old data, and pollution nodes are preset on the timeline. The pollution status of the water body in each time period is predicted based on the new data. Step f: Use a machine learning algorithm based on a random forest model to process water quality data and invert TP concentration data. The random forest model formula is: Where T is the number of trees, ht(x) is the prediction result of the t-th tree for input x; Step g: Generate a chart of water oxygen content, algae and aquatic plant density, and water turbidity based on the aggregated data, so that staff can easily determine the water pollution situation based on the chart trends. The analysis terminal compares the chart data with historical data trends to automatically determine the current water pollution situation. Step h: After inverting the TP concentration data, upload it to the cloud platform and the user end. Obtain TP concentration data through regular manual sampling, calibrate the training model, and continuously improve the model prediction accuracy.

2. The method for high-frequency real-time monitoring of TP pollution according to claim 1, characterized in that: In step a, the sensors are respectively arranged at three sections of the river channel A, B, and C. The sensors include a dissolved oxygen content sensor, a high-definition camera, a turbidity sensor, a TP sensor, a temperature sensor, and a pH sensor.

3. The method for high-frequency real-time monitoring of TP pollution according to claim 1, characterized in that: The multiple sensors and high-definition cameras in step a transmit data to the analysis terminal in real time through the wireless communication module, and the historical pollution data of the river water body are statistically integrated to form a database.

4. The method for high-frequency real-time monitoring of TP pollution according to claim 1, characterized in that: The data cleaning method in step d is based on Spearman correlation analysis and Z-score method. After screening out abnormal values ​​of indicators, whether to retain them is determined based on the changes in other indicators with strong correlation with them. The relevant formula is as follows: Where ρ is the Spearman correlation coefficient, which is used to measure the monotonic relationship between two variables; di is the absolute value of the difference between the rankings (or orders) of the two variables in the i-th pair of observations; n is the total number of observations, Among them, x is the specific data, μ is the mean of the data set, and σ is the standard deviation.

5. The method for high-frequency real-time monitoring of TP pollution according to claim 1, characterized in that: In step e, the various data collected in real time are summarized and aggregated to form a comprehensive data set. The analysis terminal conducts in-depth analysis of the comprehensive data set, including data cleaning and outlier processing. The newly collected data is compared with the data in the historical database to analyze its changing trends and anomalies, and the water pollution situation is predicted based on the pollution nodes preset on the timeline.

6. The method for high-frequency real-time monitoring of TP pollution according to claim 1, characterized in that: In step g, the water parameter chart will be marked with statistical information such as the average, maximum and minimum values ​​of each pollution level in the historical data, as well as the comparison between the current data point and these statistical information. When the data point of a monitoring indicator approaches or exceeds the preset warning value, the analysis terminal will automatically highlight the data point on the chart and issue a corresponding warning signal so that the staff can take timely response measures.

7. The method for high-frequency real-time monitoring of TP pollution according to claim 1, characterized in that: The eigenvalue importance analysis based on the random forest model in step b must ensure that the random forest model used is accurate and fully trained so that the importance of each water quality indicator can be accurately evaluated. The representativeness of the data, and the historical data used to train the random forest model should be representative and able to cover various water quality conditions.

8. The method for high-frequency real-time monitoring of TP pollution according to claim 1, characterized in that: The training and validation of the model in step f, when using the random forest model to invert TP concentration data, must ensure that the model has been fully trained and validated to improve the prediction accuracy. Before inputting the data into the model, the water quality data must be properly preprocessed, such as normalization and denoising, to improve the prediction performance of the model.

9. The method for high-frequency real-time monitoring of TP pollution according to claim 1, characterized in that: The clarity of the chart in step g should be clear and easy to read, and can intuitively reflect the changing trends and abnormal conditions of water quality parameters. The accuracy of the early warning signal should be such that when the data point of a monitoring indicator approaches or exceeds the preset early warning value, the analysis terminal should accurately issue an early warning signal so that the staff can take timely response measures.

10. The method for high-frequency real-time monitoring of TP pollution according to claim 1, characterized in that: In step h, manual sampling is performed regularly to obtain accurate TP concentration data for calibration and training of the model. The model is continuously improved and optimized based on newly acquired data and actual conditions to improve prediction accuracy and reliability.