Combined system overflow pollution real-time control method based on rainfall data
By using multi-source water and rainfall data and the K-means clustering algorithm to construct an overflow category set, and combining it with the current water and rainfall identification categories, a differentiated control strategy was formulated, which solved the pollution control problem of the combined sewer system during heavy rainfall and achieved precise and automated overflow control.
Patent Information
- Application Number
- CN202510809753.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
The existing combined sewer system is unable to effectively utilize historical data during heavy rainfall, resulting in the direct discharge of untreated pollutants into environmental water bodies. The lack of multi-scenario modeling and classification mechanisms makes it difficult to achieve refined and real-time overflow control.
Multi-source water and rainfall data and the K-means clustering algorithm are used to construct an overflow category set. By mapping the cluster centers with historical water quality events, differentiated real-time control strategies are established. The corresponding strategies are called in combination with the current water and rainfall identification categories to achieve precise pollution reduction.
It improves the response speed and decision-making accuracy of combined sewer overflow control, simplifies the calculation process, adapts to variable rainfall scenarios, and reduces the risk of urban non-point source pollution.
Smart Images

Figure CN120705657A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of overflow control, and in particular relates to a real-time control method for combined sewer overflow pollution based on rainfall data. Background Art
[0002] Combined sewer systems are characterized by rainwater and sewage sharing the same pipe network. During light to moderate rainfall, initial rainwater carries surface pollutants into the pipes, passing through interception facilities to the sewage treatment plant. However, during periods of heavy rainfall, when the pipe network flow exceeds the interception capacity or the sewage treatment plant's processing capacity, mixed rainwater and sewage are discharged directly into the receiving water body through the overflow, forming a combined sewer overflow. While this overflow behavior has a flood prevention effect, it also directly discharges untreated pollutants (such as chemical oxygen demand (COD), ammonia nitrogen (NH3-N), and total suspended solids (SS)) into the ambient water body, polluting the urban water environment.
[0003] Combined sewer overflow pollution control is a core challenge in urban water environment management. Existing methods fail to effectively utilize historical data resources. While a vast amount of historical rainfall, water level, and water quality monitoring data is available, most systems lack in-depth analysis and modeling, failing to extract typical scenarios to support real-time control strategies. Combined sewer system operation is influenced by numerous factors, including rainfall duration, water distribution, and node capacity. Actual operational scenarios are complex and highly variable, making fixed rules inadequate to address diverse events. Furthermore, there is a lack of multi-scenario modeling and classification mechanisms.
[0004] There is an urgent need for a scenario recognition and classification method based on historical data that can fully combine rainfall data, water level fullness and water quality monitoring information to construct typical combined sewer overflow categories and establish a corresponding set of pollution control strategies based on clustering results to achieve more refined and efficient real-time overflow control and reduce the risk of urban non-point source pollution. Summary of the Invention
[0005] The present invention aims to utilize multi-source water and rainfall data and cluster analysis technology to identify the overflow category of the current rainfall event, and implement differentiated real-time control strategies accordingly to achieve precise pollution reduction and drainage scheduling.
[0006] To achieve the above objectives, the present invention provides the following technical solutions:
[0007] A real-time combined sewer overflow pollution control method based on rainfall data comprises the following steps:
[0008] S1. Current water and rainfall data collection:
[0009] Obtain current rainfall observation data, future rainfall forecast data, and real-time water level fullness observation values of multiple key nodes in the pipe network to form the current water and rainfall situation set;
[0010] S2. Overflow category construction and cluster analysis:
[0011] Construct a set of historical water and rainfall events and a set of historical water quality events, use a clustering algorithm to perform cluster analysis on the set of historical water and rainfall events, and generate an overflow category set;
[0012] By mapping the cluster centers to historical event points, we establish a corresponding relationship between each overflow category and its corresponding historical water quality event set, and then extract water quality characteristics and construct a real-time combined sewer overflow control strategy set.
[0013] S3. Identification of current water and rainfall conditions and retrieval of control strategies:
[0014] The current water and rainfall conditions are obtained in real time, the distance to each cluster center is calculated, the overflow category to which it belongs is identified, and the corresponding control strategy set is extracted to guide the real-time control of combined sewer facilities.
[0015] Step S1 forms a water and rainfall regime collection that reflects the characteristics of the current rainfall process, ensuring that the control logic is based on the latest state. Step S2 uses a clustering algorithm to classify historical water and rainfall events, forming a set of overflow categories. This is then correlated with historical water quality events to extract the pollution characteristics and control requirements corresponding to different categories. Step S3 calculates the distance between the current water and rainfall conditions and each cluster center, determines its category, and invokes the corresponding control strategy, thus implementing a full-process control system from "observation-identification-decision-making", significantly improving response speed and decision-making accuracy.
[0016] Furthermore, the current water and rainfall condition set includes at least one of the following data types:
[0017] a) Hourly rainfall at the current time and in the previous hours;
[0018] b) Rainfall forecast data for the next several hours;
[0019] c) Water level fullness data of multiple key nodes at each moment.
[0020] Multi-source data fusion improves the description accuracy of current water and rainfall conditions and provides more reliable data support for category identification.
[0021] Furthermore, the historical water quality event set includes at least one of the following water quality parameters: chemical oxygen demand, ammonia nitrogen, and dissolved oxygen.
[0022] Ensure that pollution intensity and water quality characteristics can be extracted under different overflow categories, so as to formulate differentiated and feasible pollution control strategies.
[0023] Furthermore, the clustering algorithm is a K-means clustering algorithm, which includes the following sub-steps:
[0024] a) Randomly select k historical water and rainfall event samples as initial cluster centers;
[0025] b) Divide historical events into corresponding cluster centers according to the minimum Euclidean distance;
[0026] c) Update the cluster center based on the mean of each type of data sample;
[0027] d) Repeat steps b and c until the cluster center change is less than the threshold or the maximum number of iterations is reached.
[0028] Ensure that the overflow classification results are stable, representative and physically meaningful, and enhance the correlation between categories and pollution characteristics.
[0029] Furthermore, the real-time combined sewer overflow control strategy includes at least one of the following control operations:
[0030] a) Open or close a specific overflow port;
[0031] b) Start or stop the operation of the storage tank;
[0032] c) Adjust the pump station operation strategy.
[0033] Provides a variety of executable control actions to adapt to response requirements under different categories and achieve precise and automated regulation.
[0034] Furthermore, the cluster category to which the current water and rainfall conditions belong is determined by calculating the Euclidean distance between the current water and rainfall conditions and each cluster center, and selecting the category corresponding to the minimum distance.
[0035] It uses efficient classification criteria to quickly locate categories, ensuring real-time performance while taking into account clustering stability.
[0036] Furthermore, the control strategy set has differentiated response rules for different categories, including:
[0037] a) In light rain and high pollution scenarios, close all overflow outlets to reduce pollution load;
[0038] b) When a single node is about to be flooded due to moderate rain, the overflow port of the corresponding node is partially opened;
[0039] c) In the case of moderate rain with many nodes where the water level is not full but the pollution concentration is high, close the overflow outlet;
[0040] d) In the event of heavy rain, high water level and low pollution, all overflow outlets should be opened immediately to give priority to drainage.
[0041] Through situational awareness, differentiated strategy configuration can be achieved to enhance the overall pollution reduction capacity and operational resilience of the urban drainage system.
[0042] The present invention has the following beneficial effects:
[0043] Traditional combined sewer overflow control relies on static thresholds or empirical scheduling, and cannot respond to complex and changeable rainfall scenarios in real time, and it is difficult to balance drainage safety and water environment protection goals. The present invention makes full use of the highly correlated relationship between rainfall data and combined sewer overflow pollution control. The proposed technical solution constructs a dynamic water-rainfall dataset containing observation and forecast information; extracts typical overflow categories from historical data and associates pollution characteristics; quickly identifies the category and calls the matching pollution control strategy during the current rainfall process; and realizes a multi-scenario, multi-strategy intelligent response control system. Through the method of the present invention, only the water-rainfall situation is combined with the clustering algorithm to simplify the calculation process; and real-time control of multiple combined sewer network drainage control facilities is obtained over a period of time in the future, without the need for too frequent data updates. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 Schematic diagram of the process of this method;
[0045] Figure 2 It is a sequence diagram of the preceding rainfall observation, future rainfall forecast and real-time water level fullness of multiple key nodes of the pipe network in the embodiment.
[0046] Figure 3 4 is a historical water and rainfall sequence diagram of an embodiment.
[0047] Figure 4 4 is a historical water quality event sequence diagram of an embodiment.
[0048] Figure 5 The rainfall and water level filling degree series diagrams of multiple key nodes of the pipe network are shown for the first type of clustering characteristics after clustering.
[0049] Figure 6 The rainfall and water quality series plots showing the first cluster characteristics after clustering.
[0050] Figure 7 The rainfall and water level filling degree series diagrams of multiple key nodes of the pipe network are displayed for the third clustering feature after clustering.
[0051] Figure 8 Rainfall and water quality series plots showing the third cluster characteristics after clustering. DETAILED DESCRIPTION
[0052] A real-time control method for combined sewer overflow pollution based on rainfall data. The overall implementation process is as follows: Figure 1 , comprising the following steps S1 to S3:
[0053] S1. Current water and rainfall data collection
[0054] Obtain the current rainfall observation data, future rainfall forecast data, and real-time water level fullness observation values of multiple key nodes in the pipeline network to form the current water and rainfall situation set.
[0055] in:
[0056] Rainfall observation data: collected by on-site rainfall monitoring equipment and obtained through IoT transmission;
[0057] Rainfall forecast data: can be obtained from meteorological departments, commercial meteorological companies, or through simulation of basic / non-basic models;
[0058] Water level fullness data: collected in real time by water level monitoring devices installed at key nodes through the Internet of Things.
[0059] The current water and rainfall conditions are expressed as follows:
[0060]
[0061] in, is the current water and rainfall condition set; n represents different types of data dimensions.
[0062] Taking a large-scale urban combined sewer overflow pollution real-time control project as an example, the current water and rainfall conditions are as follows: Figure 2 As shown in the figure, it includes historical rainfall observations at time t+0 and hours t-1 to t-5, future rainfall forecasts from hours t+1 to t+6, and water level fullness of three key pipe network nodes from hours t-5 to t+0, including a total of 12 types of data features.
[0063] S2. Overflow category construction and cluster analysis
[0064] In this step, the historical water and rainfall events and water quality event data are analyzed through clustering algorithms to construct multiple representative overflow categories to support the formulation of subsequent control strategies.
[0065] S21. Historical data collection and construction
[0066] The following two types of historical data are collected:
[0067] Historical water and rainfall event collection X:
[0068] X={(x1,x2,…,x n )1,(x1,x2,…,x n )2,…,(x1,x2,…,x n ) m1}
[0069] Among them, X is the set of historical water and rainfall events; m1 represents the number of historical water and rainfall events; n is the data dimension contained in each event, which is consistent with step S1.
[0070] Historical water quality event set Y:
[0071] Y={(y1,y2,…,y p )1,(y1,y2,…,y p )2,…,(y1,y2,…,y p ) m2}
[0072] Among them, Y is the set of historical water quality events of water quality observation data of multiple key nodes in the pipeline network; m2 represents the number of historical water quality events; and p represents the dimension of water quality parameters (such as COD, ammonia nitrogen, dissolved oxygen, etc.).
[0073] Historical water and rainfall conditions, historical water quality events such as Figure 3 、 4 As shown in the figure, the data includes the m′th event at a certain historical time point (t+0), rainfall observations from one hour before (t-1) to five hours before (t-5), and from one hour after (t+1) to six hours after (t+6), water level fills at multiple key pipe network nodes, and real-time water quality data (COD, dissolved oxygen, and ammonia nitrogen). The m′th event data in the X set includes rainfall data from t-5 to t+0 and real-time water level fills at three key pipe network nodes from t-5 to t+6. The m′th event data in the Y set includes three types of water quality data from t-5 to t+6.
[0074] S22. Cluster analysis to establish overflow category set C
[0075] The K-means clustering algorithm is used to analyze the historical water and rainfall set X. The specific process includes:
[0076] Step S211, initialize cluster centers
[0077] At the beginning of the K-means model algorithm, the number of clusters k is set, and k event samples are randomly selected as initial cluster centers to form the initial cluster center set Q0:
[0078]
[0079] Q0 is the initial cluster center set; Represents the initial cluster center value on the nth water and rainfall characteristic dimension; Represents the initial center of the kth cluster category; the initial cluster center set is used for step S212.
[0080] Step S212: assign the event to the nearest cluster category
[0081] For each historical event point x∈X in the historical water and rainfall event set X, calculate its Euclidean distance to all cluster centers and assign it to the nearest cluster category C k :
[0082]
[0083] Step S213: Update cluster center
[0084] For each cluster category C k , calculate the new cluster center based on all the data points inside it, and use the mean as the updated cluster center Q t :
[0085]
[0086]
[0087] Q t is the set of cluster centers obtained in the tth iteration; is the new cluster center of the nth type of water and rainfall data at the tth iteration; n Is the data value representing the nth feature dimension; |C k | is the number of events in the k-th cluster category.
[0088] Step S214, iterate until termination
[0089] Repeat steps S212 to S213 to repeatedly assign cluster categories and update cluster centers until a termination condition is met, which includes no significant changes in cluster centers and reaching the maximum number of iterations, thereby obtaining a final set of cluster centers Q.
[0090] Q={(q1,q2,…,q n )1,(q1,q2,…,q n )2,…,(q1,q2,…,q n ) k}
[0091] And according to the distance between the event point and the cluster center, the cluster category to which each historical water and rainfall event belongs can be obtained to form an event classification result set:
[0092] C={c1,c2,…,c m}
[0093] Among them, C is the cluster category set; c m ∈{1,2,...,k} represents the category to which the mth historical event belongs; multiple events can belong to the same category or be distributed in different categories.
[0094] In this embodiment, the number of clusters is set to 4, and the clustering results are as follows: Figure 5 、 6 , 7, and 8, Figure 5 、 6 Figures 7 and 8 only show the clustering results for two of the categories, with the distribution characteristics of each category displayed in bar graphs and solid curves. Each cluster category contains the following time series data dimensions:
[0095] Rainfall data: hourly rainfall from t-5 to t+0 hours;
[0096] Pipeline network status: changes in water level at three key nodes from t-5 to t+6 hours;
[0097] Water quality data: dissolved oxygen, COD, ammonia nitrogen and other parameters;
[0098] Future trend: The rainfall forecast trend for each cluster category from t+1 to t+6 hours.
[0099] The cluster categories may be 2 to k categories, so the categories in each sub-set within the cluster category set may be the same or different. The cluster analysis results are used to support the construction of the overflow pollution risk assessment model and the formulation of the early warning strategy in the subsequent step S3.
[0100] S23. Construct real-time overflow control strategy set Z
[0101] Since there is a one-to-one correspondence or dual mapping relationship between historical water and rainfall events X and historical water quality events Y, the cluster category C can be simultaneously mapped to the water quality event set Y, and then the typical water quality characteristics of various overflow scenarios can be analyzed to provide a basis for control strategies.
[0102] Based on the above mapping relationship, the water quality data corresponding to the events contained in each cluster category can be extracted, and a characteristic profile of each cluster category can be constructed. Then, through scenario-based simulation methods (such as establishing a simulation model to simulate rainfall response) or empirical rule-based judgment, a targeted combined sewer overflow control strategy can be formulated for each category, thus forming a real-time control strategy set:
[0103] Z={(z1,z2,…,z r )1,(z1,z2,…,z r )2,…,(z1,z2,…,z r ) k}
[0104] Where Z is the set of real-time combined sewer overflow control strategies; (z1, z2,…, z r ) k represents the control strategy group containing r strategies established for the k-th cluster category, z rIt represents a specific combined sewer system control strategy item, such as opening or closing a specific overflow port, starting and stopping a storage tank, and scheduling pump station operations.
[0105] For example, Figure 5 In the four cluster categories shown, the control strategy for each category can be formulated based on the following judgment principles:
[0106] Category 1: Light rain scenario. In the future, the water levels at the network nodes will not reach the full water threshold, but the pollutant concentration will be high. The proposed strategy is to close all overflow outlets to reduce the risk of pollutant discharge.
[0107] Category 2: Moderate rain scenario, where only the first node will be full in the next hour and the pollution concentration is low. The proposed strategy is to open the overflow port of the first node at t+1 to relieve local hydraulic pressure.
[0108] Category 3: Moderate rain scenario, all nodes will not be full of water in the future, but the pollutant concentration is high, and the proposed strategy is to close the overflow outlet;
[0109] Category 4: Heavy rain scenario. All nodes are currently full of water and pollutant concentrations are low. The proposed strategy is to open all overflow outlets immediately (t+0) to prioritize drainage safety.
[0110] S3. Identification of current water and rainfall conditions and retrieval of control strategies
[0111] Get current water and rainfall conditions in real time By calculating the distance from the cluster center Q, the overflow category k' to which it belongs is identified, and the corresponding control strategy Z is extracted. k ′:
[0112] Z k′ =(z1,z2,…,z r ) k′
[0113] Z k′ It is a real-time combined sewer overflow control strategy set that can perform real-time control on multiple combined sewer network drainage control facilities.
[0114] like Figure 3 Current water and rainfall conditions shown and Figure 5 The second category is closest, so it is identified as a medium rain scenario. It is predicted that the first node is about to overflow and the pollution concentration is low. The system will automatically open the overflow device of the first node at t+1.
Claims
1. A real-time control method for combined sewer overflow pollution based on rainfall data, characterized in that: The steps include: S1. Current water and rainfall data collection: Obtain current rainfall observation data, future rainfall forecast data, and real-time water level fullness observation values of multiple key nodes in the pipe network to form the current water and rainfall situation set; S2. Overflow category construction and cluster analysis: Construct a set of historical water and rainfall events and a set of historical water quality events, use a clustering algorithm to perform cluster analysis on the set of historical water and rainfall events, and generate an overflow category set; By mapping the cluster centers to historical event points, we establish a corresponding relationship between each overflow category and its corresponding historical water quality event set, and then extract water quality characteristics and construct a real-time combined sewer overflow control strategy set. S3. Identification of current water and rainfall conditions and retrieval of control strategies: The current water and rainfall conditions are obtained in real time, the distance to each cluster center is calculated, the overflow category to which it belongs is identified, and the corresponding control strategy set is extracted to guide the real-time control of combined sewer facilities.
2. The real-time control method for combined sewer overflow pollution based on rainfall data according to claim 1 is characterized in that: The current water and rainfall condition set includes at least one of the following data types: a) Hourly rainfall at the current time and in the previous hours; b) Rainfall forecast data for the next several hours; c) Water level fullness data of multiple key nodes at each moment.
3. The real-time control method for combined sewer overflow pollution based on rainfall data according to claim 1 is characterized in that: The historical water quality event set includes at least one of the following water quality parameters: chemical oxygen demand, ammonia nitrogen, and dissolved oxygen.
4. The real-time control method for combined sewer overflow pollution based on rainfall data according to claim 1 is characterized in that: The clustering algorithm is a K-means clustering algorithm, which includes the following sub-steps: a) Randomly select k historical water and rainfall event samples as initial cluster centers; b) Divide historical events into corresponding cluster centers according to the minimum Euclidean distance; c) Update the cluster center based on the mean of each type of data sample; d) Repeat steps b and c until the cluster center change is less than the threshold or the maximum number of iterations is reached.
5. The real-time control method for combined sewer overflow pollution based on rainfall data according to claim 1 is characterized in that: The real-time combined sewer overflow control strategy includes at least one of the following control operations: a) Open or close a specific overflow port; b) Start or stop the operation of the storage tank; c) Adjust the pump station operation strategy.
6. The real-time control method for combined sewer overflow pollution based on rainfall data according to claim 1 is characterized in that: The cluster category to which the current water and rainfall conditions belong is determined by calculating the Euclidean distance between the current water and rainfall conditions and each cluster center, and selecting the category corresponding to the minimum distance.
7. The real-time control method for combined sewer overflow pollution based on rainfall data according to claim 1 is characterized in that: The control strategy set has differentiated response rules for different categories, including: a) In light rain and high pollution scenarios, close all overflow outlets to reduce pollution load; b) When a single node is about to be flooded due to moderate rain, the overflow port of the corresponding node is partially opened; c) In the case of moderate rain with many nodes where the water level is not full but the pollution concentration is high, close the overflow outlet; d) In the event of heavy rain, high water level and low pollution, all overflow outlets should be opened immediately to give priority to drainage.
Citation Information
Cited By
Intelligent catch basin control method and system based on water quality on-line monitoring
CN121478004A