Control method for quick response to overflow pollution

By quickly evaluating water quality parameters through a linear regression model, the problems of rough setting of the overflow device's interception multiple and time-consuming water quality parameter measurement were solved, rapid response and control of overflow pollution were achieved, and the timeliness of water quality monitoring and environmental supervision capabilities were improved.

CN120611129APending Publication Date: 2025-09-09CHONGQING THREE GORGES ECO-ENVIRONMENTAL TECH INNOVATION CENT CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510692691.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In the existing technology, the interception multiple of the overflow device is roughly set, resulting in a large amount of low-concentration sewage at the front end of the pipeline occupying the capacity or high-concentration mixed sewage being directly discharged into the river, making it difficult to achieve precise overflow control, and the water quality parameter measurement process is time-consuming, affecting the rapid response to water quality emergencies and the timeliness of water resources management.

Method used

By establishing a rapid assessment and control method, using a linear regression model combined with correlation analysis of water quality parameters, water quality variables can be quickly predicted to achieve rapid response control of overflow pollution, including data collection, preprocessing, feature selection, model establishment and parameter updating, and interception control is carried out in combination with the river environment capacity and downstream pipeline carrying capacity.

Benefits of technology

It achieves rapid prediction of water quality parameters in a short period of time, improves the timeliness of water quality monitoring, reduces the potential risks caused by detection delays, effectively controls water pollution incidents, and improves environmental supervision capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611129A_ABST
    Figure CN120611129A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of municipal and environmental engineering, and relates to an overflow pollution quick response control method, which specifically comprises the steps of initial data collection, data preprocessing, feature selection, model establishment, data prediction, parameter updating, closure control and the like. According to the method, water quality parameters which are relatively slow in measurement process can be predicted in a short time, so that the monitoring speed of the water quality change is accelerated. The quick response capability can effectively reduce potential risks caused by detection delay. According to the method, the timeliness is improved by rapidly predicting the water quality parameters, the water quality monitoring data can be rapidly obtained, the water pollution problem can be timely identified and treated, and the rapid monitoring and response mechanism is beneficial to enhancing environment supervision and effectively controlling and reducing the occurrence of water pollution events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of municipal engineering and environmental engineering, and in particular relates to a rapid response control method for overflow pollution. Background Art

[0002] During urbanization, new pipeline networks often adopt separate rainwater and sewage systems, but most older urban areas still retain combined sewer systems. During rainfall, the flow rate within the combined sewer increases. When the flow rate exceeds the intercepting main's capacity, the excess mixed rainwater and sewage is discharged into the receiving water body through overflow wells. Furthermore, wastewater from road washouts such as night markets, farmers' markets, and urban roads currently primarily enters urban stormwater pipes, where it easily accumulates. Rainfall carries these deposited pollutants into rivers, and control over initial rainfall interception remains relatively crude, leading to large amounts of low-concentration rainwater entering sewage pipes.

[0003] Currently, commonly used overflow devices include weir overflow, gate-controlled overflow, and a combination of weir and gate. The interception ratio is controlled by the height of the overflow outlet. In practice, however, the interception ratio is crudely set, resulting in a large amount of low-concentration wastewater at the front end of the pipeline, which overloads the pipeline capacity, or high-concentration mixed wastewater being directly discharged into the river, making it difficult to improve the quality and efficiency of sewage treatment plants. The interception ratio should be determined comprehensively based on the concentration of combined wastewater, the environmental capacity of the river, and the load of downstream pipelines. Water quality parameters such as chemical oxygen demand (COD), ammonia nitrogen, total phosphorus, and total nitrogen are key indicators for assessing water quality. Accurate measurement of these parameters is crucial for the rational use of water resources, pollution source control, and water environment management. Although existing water quality monitoring technologies are quite mature, they still have limitations. In particular, the measurement process for some water quality parameters is time-consuming. For example, analysis of COD often requires complex pretreatment and instrumental analysis steps, resulting in a long time to produce results. This delay not only hinders the rapid response to water quality emergencies but also limits the timeliness and dynamic control capabilities of water resource management. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention provides a control method for rapid response to overflow pollution. In order to solve the problem that online water quality detection is difficult to achieve accurate overflow and abandonment control, a rapid assessment and control method is established to accelerate the judgment speed of overall water quality parameters and improve the control speed and efficiency of overflow pollution.

[0005] The technical purpose of the present invention is achieved through the following technical solution: a method for controlling overflow pollution with rapid response, comprising the following steps: Step 1: Initial data collection: First, collect an initial complete water quality data set for subsequent new model building; Step 2: Data preprocessing: 1. Handle missing values ​​and outliers; 2. When the dimensions of different dimensional features of the original data are inconsistent, standardize or normalize the data; otherwise, do not perform standardization or normalization; Step 3: Feature selection: Conduct correlation analysis on different water quality variables and select characteristic fast indicator variables that are significantly correlated with the target slow indicator variables from the collected data for subsequent model building; Step 4: Model building: 1. Build a linear regression model based on the selected feature variables; 2. Randomly select some data from the initial data set to determine the parameters of the linear regression model equation and evaluate it using the F test; 3. Use the remaining data to test the established model; Step 5: Data prediction: A new round of water samples is collected and tested. After preprocessing the fast indicator variable data, the model equation established above is used to calculate the predicted value of the slow indicator variable, thereby quickly obtaining complete water quality variable data for guiding diversion; Step 6: Parameter update: After the measured values ​​of the slow indicator variables of the new water sample in step 5 are obtained, a new data set is formed with the fast indicator variable data. After the data set is preprocessed, the parameters of the model equation established in step 4 are updated. The updated equation can be used to guide the next round of water quality prediction; Step 7: Interception Control: Based on the river's environmental capacity and the downstream pipeline's carrying capacity, and using the water quality variable data quickly acquired in Step 5, the water quality is determined and the interception height is controlled to quickly control overflow pollution.

[0006] Preferably, the step three is to perform correlation analysis between different variables and calculate the Pearson correlation coefficient r between pollutants: ; in 、 as well as 、 It is a sample 、 The mean and standard deviation of .

[0007] Preferably, the method for establishing the model in step 4 is linear regression, and the matrix composed of independent variable data is recorded as , the matrix composed of dependent variable data is , then the established linear model equation is: ; The parameters to be estimated are , is the residual.

[0008] Preferably, the F test is to test the significance of the regression equation from the regression effect. If it is significant, it means that the linear relationship of the regression equation exists; if it is not significant, it means that the linear relationship of the regression equation does not exist. The specific steps are: First, make the assumption: H0: The linear relationship is not significant, that is, the parameters β are all 0; H1: At least one of the parameters β is not 0; Then calculate the test statistic F and the corresponding p-value: ; Where MSR stands for the mean square of the regression sum of squares, which measures the variance explained by the model due to the independent variables; MSE is the mean square error, which is the average of the sum of the squares of the differences between the true value and the predicted value; Among them, R² can be used to measure the effect of this linear model fitting, and the calculation formula is: ; in, is the residual sum of squares, is the total sum of squares, When the remaining data, i.e., test data, is used to test the established model, the detection results of the test data on the model can be measured by MSE mean square error and R².

[0009] Preferably, the triggering condition for the parameter update in step 6 is any one of the following: The absolute error between the measured value and the predicted value of the slow indicator of the new water sample exceeds the preset threshold (e.g., 10%); The cumulative newly added dataset reaches 20% of the initial dataset size; Update periodically, with the update frequency being every 24 hours or every 5 sets of new data received.

[0010] Preferably, in step five, the fast indicator variables include at least three of pH value, conductivity, dissolved oxygen, and turbidity, and the slow indicator variables include at least two of chemical oxygen demand (COD), total nitrogen (TN), total phosphorus (TP), and heavy metal concentration.

[0011] Preferably, the method also includes: Step 8, risk warning and visualization: displaying the predicted value, measured value and diversion status of the river water quality in real time through the GIS map. When the predicted value exceeds the water quality safety threshold, an audible and visual alarm is triggered and the warning information is pushed to the management terminal.

[0012] Preferably, in the data preprocessing in step 2, the missing values ​​are processed by interpolation or deletion, and the outliers are processed by Z-score standardization combined with threshold judgment method, specifically: when the absolute value of the Z-score of a data point exceeds 3, it is determined to be an outlier and eliminated; the standardization process uses the Z-score formula: ; Among them, μ is the mean of the characteristic variable and σ is the standard deviation.

[0013] Preferably, if the p-value is less than the significance level determined in advance, the null hypothesis is rejected and it is considered that the linear relationship of the regression equation exists; otherwise, the null hypothesis cannot be rejected, that is, there is no linear relationship in the regression equation.

[0014] Preferably, after using the F test to evaluate the model, if the linear relationship of the regression equation is not significant, the stepwise regression method is used to screen the characteristic variables, adding or deleting one variable from the model in each iteration, and recalculating the F test statistic and p value until the regression equation reaches a significant level and the model structure is optimized.

[0015] Compared with the existing methods, the present invention has the following beneficial effects: The present invention can quickly predict water quality parameters that are slower to measure, thereby accelerating the monitoring of water quality changes. This rapid response capability is crucial for handling water quality emergencies and can effectively reduce potential risks caused by delayed detection.

[0016] This invention improves timeliness by rapidly predicting water quality parameters, enabling faster acquisition of water quality monitoring data and timely identification and resolution of water pollution issues. This rapid monitoring and response mechanism helps strengthen environmental regulation and effectively control and reduce the occurrence of water pollution incidents.

[0017] The method of the present invention is easy to integrate with other water quality monitoring systems and equipment, and has compatibility and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Flowchart of the present invention.

[0019] Figure 2 is the residual of the chemical oxygen demand under the prediction model in the embodiment of the present invention.

[0020] Figure 3 is the residual of total phosphorus under the prediction model in the embodiment of the present invention.

[0021] Figure 4 is the residual of total nitrogen under the prediction model in the embodiment of the present invention.

[0022] Figure 5The figure is a schematic diagram of the implementation structure of an embodiment of the present invention.

[0023] In the figure: 1. Water inlet pipe; 2. Overflow control panel; 3. Overflow pipe. DETAILED DESCRIPTION

[0024] The present invention is described in detail below with reference to specific embodiments.

[0025] This embodiment provides a method for determining the rationality of the construction of a rural sewage treatment station and its supporting pipe network, including the following steps: Step 1: Initial data collection: First, collect an initial complete water quality data set for subsequent new model building; Step 2: Data preprocessing: 1. Handle missing values ​​and outliers; 2. When the dimensions of different dimensional features of the original data are inconsistent, standardize or normalize the data; otherwise, do not perform standardization or normalization; Step 3: Feature selection: Conduct correlation analysis on different water quality variables and select characteristic fast indicator variables that are significantly correlated with the target slow indicator variables from the collected data for subsequent model building; Step 4: Model building: 1. Build a linear regression model based on the selected feature variables; 2. Randomly select some data from the initial data set to determine the parameters of the linear regression model equation and evaluate it using the F test; 3. Use the remaining data to test the established model; Step 5: Data prediction: A new round of water samples is collected and tested. After preprocessing the fast indicator variable data, the model equation established above is used to calculate the predicted value of the slow indicator variable, thereby quickly obtaining complete water quality variable data for guiding diversion; Step 6: Parameter update: After the measured values ​​of the slow indicator variables of the new water sample in step 5 are obtained, a new data set is formed with the fast indicator variable data. After the data set is preprocessed, the parameters of the model equation established in step 4 are updated. The updated equation can be used to guide the next round of water quality prediction; Step 7. Interception control: Based on the river environment capacity and the downstream pipeline carrying capacity, the water quality variable data quickly obtained in step 5 is used to judge the water quality and control the interception height to achieve rapid control of overflow pollution.

[0026] In the above embodiment, the commonly used parameters and indicators in the water quality dataset established in step 1 include ammonia nitrogen, suspended solids, chemical oxygen demand, total phosphorus, and total nitrogen. This method proposes that these parameters and indicators can be obtained by an online monitor. Requirements for the online water quality testing equipment used in this method are as follows: ammonia nitrogen is measured using the ammonia gas-sensitive electrode method, meeting the "Technical Requirements and Testing Methods for Online Automatic Ammonia Nitrogen Water Quality Monitors" (HJ 101-2019); suspended solids are measured using an optical sensor method, which calculates the concentration of suspended solids in water by measuring the intensity of light signals scattered or transmitted by suspended solids after exposure to visible or near-infrared light; and chemical oxygen demand is measured using the dichromate method, meeting the "Technical Requirements and Testing Methods for Online Automatic Chemical Oxygen Demand (CODCr) Water Quality Monitors" (HJ377-2019). Total phosphorus is measured using ammonium molybdate spectrophotometry, meeting the "Technical Requirements for Automatic Water Quality Analyzers for Total Phosphorus" (HJ / T 103-2003). Total nitrogen is measured using alkaline potassium persulfate digestion-UV spectrophotometry, meeting the "Technical Requirements for Automatic Water Quality Analyzers for Total Nitrogen" (HJ / T 102-2003). The detection time for chemical oxygen demand, total phosphorus, and total nitrogen is approximately 20-40 minutes, considered slow indicator variables. The detection time for ammonia nitrogen and suspended solids is less than 2 minutes, considered fast indicator variables.

[0027] In addition, the data preprocessing involved in step 2 requires standardization or normalization of the data if the scales or dimensions of the different dimensional features of the original data are inconsistent. Normalization refers to transforming the data into a fixed interval. For example, a linear transformation can be used to map the data values ​​to between [0, 1]. ; Data standardization is to transform the data into a distribution with a mean of 0 and a standard deviation of 1, while retaining the original data distribution. It can be expressed as: .

[0028] In some embodiments, step three performs correlation analysis between different variables to calculate the Pearson correlation coefficient r between pollutants: ; in 、 as well as 、 It is a sample 、 The mean and standard deviation of .

[0029] In other embodiments, based on the above embodiment, the method of establishing the model in step 4 is linear regression, and the matrix composed of independent variable data is recorded as , the matrix composed of dependent variable data is , then the established linear model equation is: ; The parameters to be estimated are , is the residual.

[0030] There are two approaches to solving the estimated values ​​of the above parameters. One is to start from the direction of making the residual as small as possible, that is, the least squares method; the other is to start from the probability distribution of the residual, that is, the maximum likelihood estimation method.

[0031] When using the least squares method to solve the estimated values ​​of the above parameters, it is necessary to find the parameters Appropriate estimate , so that the residual sum of squares is minimized, that is: ; in ; Depend on ; It can be solved .

[0032] When the least squares method is used to solve the estimated values ​​of the above parameters, the maximum likelihood estimation method: For linear regression, , , then the maximum likelihood function can be written as: ; After taking the logarithm, we have ; Depend on , which can be solved as .

[0033] In step 4, the F test is used to test the significance of the regression equation based on the regression effect. If it is significant, it means that the linear relationship of the regression equation exists; if it is not significant, it means that the linear relationship of the regression equation does not exist.

[0034] The specific steps are as follows: First, make the hypothesis: H0: The linear relationship is not significant, that is, the parameters β are all 0; H1: At least one of the parameters β is not 0; Then calculate the test statistic F and the corresponding p-value.

[0035] ; Where MSR stands for the mean square of the regression squares, which measures the variance explained by the model due to the independent variables; MSE is the mean squared error, which is the average of the sum of the squares of the differences between the true values ​​and the predicted values.

[0036] If the p-value is less than the significance level we determined in advance, we reject the null hypothesis and believe that the linear relationship of the regression equation exists. Otherwise, we cannot reject the null hypothesis, that is, the regression equation does not have a linear relationship.

[0037] R² can be used to measure the effect of this linear model fitting. The calculation formula is ; in, is the residual sum of squares, is the total sum of squares.

[0038] The remaining data, i.e., the test data, is used to test the established model. The test results of the model on the test data can be measured by MSE mean square error and R².

[0039] The diversion control involved in step seven needs to combine the river environment capacity and the downstream pipeline carrying capacity to set the threshold of the water quality index. The water quality variable data quickly obtained in step five above is compared with the set threshold to judge the water quality situation, control the diversion height, and achieve rapid control of overflow pollution.

[0040] In the above embodiment, the triggering condition for the parameter update in step 6 is any one of the following: The absolute error between the measured value and the predicted value of the slow indicator of the new water sample exceeds the preset threshold (e.g., 10%); The cumulative newly added dataset reaches 20% of the initial dataset size; Update periodically, with the update frequency being every 24 hours or every 5 sets of new data received.

[0041] Among them, in step five, the fast indicator variables include at least three of pH value, conductivity, dissolved oxygen, and turbidity, and the slow indicator variables include at least two of chemical oxygen demand (COD), total nitrogen (TN), total phosphorus (TP), and heavy metal concentration.

[0042] In a preferred embodiment, the method also includes: Step 8, risk warning and visualization: the predicted value, measured value and diversion status of the river water quality are displayed in real time through the GIS map. When the predicted value exceeds the water quality safety threshold, an audible and visual alarm is triggered and the warning information is pushed to the management terminal.

[0043] In the data preprocessing in step 2, missing values ​​are processed by interpolation or deletion, and outliers are processed by Z-score standardization combined with threshold judgment. Specifically, when the absolute value of the Z-score of a data point exceeds 3, it is determined to be an outlier and eliminated; the standardization process uses the Z-score formula: ; Among them, μ is the mean of the characteristic variable and σ is the standard deviation.

[0044] In a feasible preferred embodiment, after using the F test to evaluate the model, if the linear relationship of the regression equation is not significant, the stepwise regression method is used to screen the characteristic variables, adding or deleting one variable from the model in each iteration, and recalculating the F test statistic and p value until the regression equation reaches a significant level, thereby optimizing the model structure.

[0045] Next, the present invention will be further analyzed and explained with reference to the following examples.

[0046] In a specific embodiment, after the collected data are preprocessed, there are 70 groups of data as shown in Table 1, 56 groups are randomly selected as training data for establishing model equations, and the remaining 14 groups are test data for verifying the established model equations.

[0047] Table 1 Water quality data after pretreatment

[0048] Next, we conduct correlation analysis on these five different variables and calculate the Pearson correlation coefficient between pollutants: ; in 、 as well as 、 It is a sample 、 Then, according to the correlation analysis results, appropriate characteristic variables are selected for subsequent linear model establishment.

[0049] Table 2 Correlation coefficient table

[0050] Note: ***, **, and * represent significance levels of 1%, 5%, and 10%, respectively. From the results presented in the above table, it can be seen that there is a significant correlation between chemical oxygen demand, total phosphorus and total nitrogen and ammonia nitrogen and suspended solids. Therefore, it is possible to consider establishing a linear relationship between chemical oxygen demand, total phosphorus and ammonia nitrogen and ammonia nitrogen and suspended solids.

[0051] 1. Determine the relationship between chemical oxygen demand, ammonia nitrogen and suspended solids.

[0052] The model equation established by 56 randomly selected groups is: chemical oxygen demand = 8.168 + 2.152*ammonia nitrogen + 0.338*suspended solids.

[0053] Table 3 Linear regression analysis results between chemical oxygen demand, ammonia nitrogen and suspended solids

[0054] Note: ***, **, and * represent significance levels of 1%, 5%, and 10%, respectively. Evaluation: The R² value is used to determine the goodness of fit of the regression line to the linear model, and the result is 0.865. The F-test is used to determine whether there is a significant linear relationship. A linear regression model requires that the overall regression coefficient is not zero, that is, there is a regression relationship between the variables. The model is tested based on the F-test results. The F-test results show that the level is significant, rejecting the null hypothesis that the regression coefficient is 0, indicating that the model basically meets the requirements. The VIF value represents the severity of multicollinearity and is used to test whether the model exhibits collinearity. All VIF values ​​are less than 10, indicating that the model does not have multicollinearity issues and is well constructed.

[0055] 2. Determine the relationship between total phosphorus, ammonia nitrogen and suspended solids.

[0056] The model equation established by the training data is: Total phosphorus = 0.08 + 0.079 * ammonia nitrogen - 0.0002 * suspended solids Table 4 Linear regression analysis results between total phosphorus, ammonia nitrogen and suspended solids

[0057] Note: ***, **, and * represent significance levels of 1%, 5%, and 10%, respectively. Evaluation: The R² result is 0.981. The F-test results indicate that the regression coefficient is significant at the level of significance, rejecting the null hypothesis of a regression coefficient of 0. Therefore, the model generally meets the requirements. All VIF values ​​are less than 10, indicating that the model has no multicollinearity issues and is well constructed.

[0058] 3. Determine the relationship between total nitrogen, ammonia nitrogen and suspended solids.

[0059] The model equation established by the training data is: Total nitrogen = 1.096 + 0.979 * ammonia nitrogen + 0.024 * suspended solids Table 5 Linear regression analysis results between total nitrogen, ammonia nitrogen and suspended solids

[0060] Note: ***, **, and * represent significance levels of 1%, 5%, and 10%, respectively. Evaluation: The R² result is 0.988. The F-test results indicate that the regression coefficient is significant at the level of significance, rejecting the null hypothesis of a regression coefficient of 0. Therefore, the model generally meets the requirements. All VIF values ​​are less than 10, indicating that the model has no multicollinearity issues and is well constructed.

[0061] Figure 2 Figures 3 and 4 show the residuals of chemical oxygen demand, total phosphorus, and total nitrogen under the prediction model. From the R2 and mean square error (MSE) of the training data and the prediction, it can be seen that the model has established a good prediction.

[0062] Table 5 Prediction results of the linear regression model for training data and test data

[0063] After establishing the above model, it was used to make predictions and guide diversion. After a new water sample was collected, the measured ammonia nitrogen and suspended solids values ​​were 10.2 mg / L and 27 mg / L. Substituting these values ​​into the established model equations, the predicted chemical oxygen demand, total phosphorus, and total nitrogen values ​​were 39 mg / L, 0.88 mg / L, and 11.7 mg / L, respectively. These five data points were compared with the set thresholds to control diversion.

[0064] After the actual measured values ​​of chemical oxygen demand, total phosphorus and total nitrogen of the new water sample are 39 mg / L, 0.88 mg / L and 11.7 mg / L, a new data set is formed with the fast indicator variable data of the first measured results. After being added to the original 70 sets of data sets, a new data set is formed, and then the model equation parameters are updated to guide subsequent interception control.

[0065] The method of the present invention can be used in Figure 5 In the facility shown, water enters the facility through the water inlet pipe 1, and water samples are taken to obtain water quality data. The system then uses this data to perform predictive analysis based on an established mathematical model. The system then compares the predicted results with preset thresholds to determine the interception control strategy. For water with high concentrations exceeding the standard, the system raises the overflow control plate 2 to prevent it from overflowing through the overflow pipe 3, ensuring that the water is transported to the treatment station for necessary purification. Conversely, for water with low concentrations that do not exceed the threshold, the system lowers the overflow control plate 2, allowing it to be discharged directly from the overflow pipe 3. The entire process can be managed automatically, reducing manual intervention and improving processing efficiency and response speed.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for quickly responding to overflow pollution, comprising the following steps: Step 1: Initial data collection: First, collect an initial complete water quality data set for subsequent new model building; Step 2: Data preprocessing:

1. Handle missing values ​​and outliers; 2. When the dimensions of different dimensional features of the original data are inconsistent, standardize or normalize the data; otherwise, do not perform standardization or normalization; Step 3: Feature selection: Conduct correlation analysis on different water quality variables and select characteristic fast indicator variables that are significantly correlated with the target slow indicator variables from the collected data for subsequent model building; Step 4: Model building:

1. Build a linear regression model based on the selected feature variables; 2. Randomly select some data from the initial data set to determine the parameters of the linear regression model equation and evaluate it using the F test; 3. Use the remaining data to test the established model; Step 5: Data prediction: A new round of water samples is collected and tested. After preprocessing the fast indicator variable data, the model equation established above is used to calculate the predicted value of the slow indicator variable, thereby quickly obtaining complete water quality variable data for guiding diversion; Step 6: Parameter update: After the measured values ​​of the slow indicator variables of the new water sample in step 5 are obtained, a new data set is formed with the fast indicator variable data. After the data set is preprocessed, the parameters of the model equation established in step 4 are updated. The updated equation can be used to guide the next round of water quality prediction; Step 7. Interception control: Based on the river environment capacity and the downstream pipeline carrying capacity, the water quality variable data quickly obtained in step 5 is used to judge the water quality and control the interception height to achieve rapid control of overflow pollution.

2. The overflow pollution rapid response control method according to claim 1 is characterized in that: The third step is to perform correlation analysis between different variables and calculate the Pearson correlation coefficient r between pollutants: ; in 、 as well as 、 It is a sample 、 The mean and standard deviation of .

3. The overflow pollution rapid response control method according to claim 1, characterized in that: The method of establishing the model in step 4 is linear regression, and the matrix composed of independent variable data is recorded as , the matrix composed of dependent variable data is , then the established linear model equation is: ; The parameters to be estimated are , is the residual.

4. The overflow pollution rapid response control method according to claim 1, characterized in that: The F test is to test the significance of the regression equation from the regression effect. If it is significant, it means that the linear relationship of the regression equation exists. If it is not significant, it means that the linear relationship of the regression equation does not exist. The specific steps are: First, make the assumption: H0: The linear relationship is not significant, that is, the parameters β are all 0; H1: At least one of the parameters β is not 0; Then calculate the test statistic F and the corresponding p-value: ; Where MSR stands for the mean square of the regression sum of squares, which measures the variance explained by the model due to the independent variables; MSE is the mean square error, which is the average of the sum of the squares of the differences between the true value and the predicted value; Among them, R² can be used to measure the effect of this linear model fitting, and the calculation formula is: ; in, is the residual sum of squares, is the total sum of squares, When the remaining data, i.e., test data, is used to test the established model, the test results of the model on the test data are measured by MSE mean square error and R².

5. The overflow pollution rapid response control method according to claim 1 is characterized in that: The triggering condition for the parameter update in step 6 is any one of the following: The absolute error between the measured value and the predicted value of the slow indicator of the new water sample exceeds the preset threshold; The cumulative newly added dataset reaches 20% of the initial dataset size; Update periodically, with the update frequency being every 24 hours or every 5 sets of new data received.

6. The overflow pollution rapid response control method according to claim 1, characterized in that: In step 5, the fast indicator variables include at least three of pH value, conductivity, dissolved oxygen, and turbidity, and the slow indicator variables include at least two of chemical oxygen demand (COD), total nitrogen (TN), total phosphorus (TP), and heavy metal concentration.

7. The overflow pollution rapid response control method according to claim 1, characterized in that: The method also includes: Step 8, risk warning and visualization: real-time display of river water quality prediction value, measured value and diversion status through GIS map; when the predicted value exceeds the water quality safety threshold, an audible and visual alarm is triggered and the warning information is pushed to the management terminal.

8. The overflow pollution rapid response control method according to claim 1, characterized in that: In the data preprocessing in step 2, missing values ​​are processed by interpolation or deletion, and outliers are processed by Z-score standardization combined with threshold judgment. Specifically, when the absolute value of the Z-score of a data point exceeds 3, it is determined to be an outlier and eliminated; the standardization process uses the Z-score formula: ; Among them, μ is the mean of the characteristic variable and σ is the standard deviation.

9. The overflow pollution rapid response control method according to claim 4, characterized in that: If the p-value is less than the significance level determined in advance, the null hypothesis is rejected and it is considered that the linear relationship of the regression equation exists; otherwise, the null hypothesis cannot be rejected, that is, there is no linear relationship in the regression equation.

10. The overflow pollution rapid response control method according to claim 9, characterized in that: After using the F test to evaluate the model, if the linear relationship of the regression equation is not significant, the stepwise regression method is used to screen the characteristic variables. Each iteration, one variable is added or deleted from the model, and the F test statistic and p value are recalculated until the regression equation reaches a significant level and the model structure is optimized.

Citation Information

Patent Citations

  • Double-gate mixed flow rainwater interception and storage device, system and method based on water quality monitoring

    CN113502896A

  • Rapid prediction method based on combined system overflow sewage pollutant removal rate

    CN113962493A

  • Confluence system pipe network pollution control device based on water quality and liquid level real-time monitoring

    CN114892782A

  • Water quality monitoring index prediction method based on multiple regression and random forest

    CN116108941A