Water environment monitoring data processing method and device, and electronic equipment
By performing collinearity, time series, and cluster analysis on water quality parameters and flow data of drainage pipe network monitoring nodes, abnormal time periods were screened out, solving the problems of unrepresentative statistical distribution of monitoring data and difficulty in identifying abnormal states in drainage pipe networks, and realizing the accuracy of water component analysis and the effectiveness of Monte Carlo simulation.
Patent Information
- Application Number
- CN202511609682.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-11-05
Smart Images

Figure CN121071852B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of water environment monitoring technology, and in particular to a method and apparatus for processing water environment monitoring data, and electronic equipment. Background Technology
[0002] Since the changes in water quality parameters (or "water quality characteristic factors") in actual drainage pipe networks over time generally conform to a certain statistical distribution, the statistical distribution of water quality parameters in drainage pipe networks can be inferred through representative sample collection, thereby solving the proportion of different types of incoming water in the pipe network.
[0003] In practical engineering scenarios, due to interference from factors such as cost, sampling point distribution, sampling time, and testing errors, there is a problem of unrepresentative statistical distribution of sampled data. That is, the statistical distribution of the samples deviates too much from the actual distribution. A wide distribution of actual monitoring data can lead to problems such as non-convergence of calculation results or abnormally high coefficients of variation when using Monte Carlo simulation methods. At the same time, abnormal operating states in drainage pipe networks, such as sudden large-scale intrusion of external water or discharge of high-concentration sewage, are shorter in duration and harder to identify than normal operating states. This problem directly makes it difficult to obtain water quality parameter data under abnormal operating states through monitoring methods or to identify them as outliers (i.e., abnormal data) for exclusion. Summary of the Invention
[0004] In view of this, this disclosure proposes a method, apparatus, and electronic device for processing water environment monitoring data, which can classify the monitoring data at the monitoring nodes, so that Monte Carlo simulation method can be used to analyze the water composition of the drainage network based on the classification results, and obtain reasonable results.
[0005] According to one aspect of this disclosure, a method for processing water environment monitoring data is provided, comprising: performing a first screening based on the acquired time series of monitoring values of multiple water quality parameters for each monitoring node in a target drainage network during a first time period to determine the target water quality parameters for each monitoring node during the first time period, wherein each monitoring node is determined based on the distribution information of the target drainage network, and the types of target water quality parameters include at least two; and determining at least one first target time period for each monitoring node during the first time period based on the results of time series analysis of the monitoring values of the target water quality parameters and the monitoring values of the target flow rate for each monitoring node during the first time period, wherein the time series analysis refers to using a time series model to perform... Prediction of water quality parameters and flow rates; based on the results of a second screening of the time series of target water quality parameters and target flow rates of each monitoring node in each corresponding first target time period, at least one second target time period of each monitoring node is determined from all first target time periods of each monitoring node; based on the time series of target water quality parameters and target flow rates of each monitoring node in each corresponding second target time period, cluster analysis is performed to obtain the classification results corresponding to each monitoring node, the classification results indicating the pipeline operation problems existing at the corresponding monitoring node, and the classification results are used as the basis for water composition analysis of the target drainage pipeline network for each monitoring node.
[0006] In one possible implementation, the first screening based on the acquired time series of monitoring values of multiple water quality parameters for each monitoring node in the target drainage network during a first time period to determine the target water quality parameters for each monitoring node during the first time period includes: performing collinearity analysis on the time series of monitoring values of at least some water quality parameters for each monitoring node during the first time period to obtain multiple optional parameters with collinearity for each monitoring node; and taking the optional parameters matched for each water sample type at each monitoring node during the first time period as the target water quality parameters for the corresponding monitoring node during the first time period.
[0007] In one possible implementation, the first time period includes a first sub-time period and a second sub-time period following the first sub-time period, each of which includes multiple monitoring time points. Specifically, determining at least one first target time period for each monitoring node within the first time period, based on the time series analysis of the monitoring values of the target water quality parameters and the target flow rate monitoring values for each monitoring node within the first time period, includes: inputting the monitoring values of the target water quality parameters and the target flow rate monitoring values for each monitoring node within the first sub-time period into a time series model for calculation to obtain prediction results for each monitoring node at each monitoring time point within the second sub-time period, the prediction results including predicted water quality parameter values and predicted flow rates; and determining at least one first target time period for each monitoring node within the second sub-time period based on the analysis results of the predicted water quality parameter values and the target water quality parameter monitoring values, and the analysis results of the predicted flow rates and the target flow rates monitoring values for each monitoring node at each monitoring time point within the second sub-time period.
[0008] In one possible implementation, the prediction result further includes a water quality parameter confidence interval and a flow confidence interval; wherein, based on the analysis results of the predicted water quality parameter values and target water quality parameter monitoring values at each monitoring time point in the second sub-time period for each monitoring node, and the analysis results of the predicted flow value and target flow monitoring value, at least one first target time period for each monitoring node is determined from the second sub-time period for each monitoring node, including: for each monitoring node, taking the time period formed by the monitoring time points corresponding to the target water quality parameter monitoring values outside the water quality parameter confidence interval as the first target time period for the monitoring node, and taking the time period formed by the monitoring time points corresponding to the target flow monitoring values outside the flow confidence interval as the first target time period for the monitoring node.
[0009] In one possible implementation, determining at least one second target time period for a monitoring node from all first target time periods based on the results of a second screening of the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each corresponding first target time period includes: performing equipment anomaly detection on the monitoring equipment at each monitoring node based on the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each first target time period, obtaining equipment detection results for each monitoring node, wherein the equipment detection results include monitoring equipment anomaly or monitoring equipment normal operation; if there is a monitoring node whose equipment detection result is monitoring equipment anomaly, then determining the equipment anomaly time period corresponding to the monitoring equipment anomaly, and removing the equipment anomaly time period from the first target time period of that monitoring node to obtain at least one second target time period for that monitoring node.
[0010] In one possible implementation, the step of determining at least one second target time period for each monitoring node from all first target time periods based on the results of a second screening of the time series of monitoring values of target water quality parameters and the time series of monitoring values of target flow for each of the monitoring nodes in each corresponding first target time period further includes: performing data anomaly detection on the time series of monitoring values of target water quality parameters and the time series of monitoring values of target flow for each of the monitoring nodes in each first target time period; determining abnormal data of target water quality parameters from the time series of monitoring values of target water quality parameters and abnormal data of flow from the time series of monitoring values of target flow; wherein the data anomaly detection includes at least one of abnormal noise detection, data jitter detection, baseline anomaly detection, and abnormal jump detection; removing abnormal data that matches the abnormal data of target water quality parameters and / or abnormal data of flow in the first target time periods of each of the monitoring nodes to obtain at least one second target time period for each of the monitoring nodes.
[0011] In one possible implementation, the method further includes: before performing the clustering analysis, for each monitoring node, removing second target time periods that do not meet the selection conditions according to preset selection criteria.
[0012] In one possible implementation, the method further includes: performing Monte Carlo calculations based on the classification results corresponding to each monitoring node, and determining the water composition analysis results of the target drainage network corresponding to each monitoring node, wherein the composition analysis results represent the basis for determining the flow rate of each water sample type at the corresponding monitoring node.
[0013] According to another aspect of this disclosure, a water environment monitoring data processing device is provided, comprising: a first screening module, configured to perform a first screening based on the acquired time series of monitoring values of multiple water quality parameters for each monitoring node in a target drainage network during a first time period, to determine the target water quality parameters for each monitoring node during the first time period, wherein each monitoring node is determined based on the distribution information of the target drainage network, and the types of target water quality parameters include at least two; and a time series analysis module, configured to determine at least one first target time period for each monitoring node during the first time period based on the results of time series analysis of the monitoring values of the target water quality parameters and the monitoring values of the target flow rate for each monitoring node during the first time period, wherein the time series analysis refers to using a time series model for... The system includes: a prediction module for water quality parameters and flow rates; a second screening module for determining at least one second target time period for each monitoring node from all first target time periods based on the second screening results of the time series of the monitoring values of the target water quality parameters and the time series of the monitoring values of the target flow rates for each monitoring node in the corresponding first target time periods; and a cluster analysis module for performing cluster analysis based on the time series of the monitoring values of the target water quality parameters and the time series of the monitoring values of the target flow rates for each monitoring node in the corresponding second target time periods to obtain the classification results corresponding to each monitoring node. The classification results indicate the pipeline operation problems existing at the corresponding monitoring node, and the classification results are used as the basis for conducting water composition analysis for the target drainage pipeline network at each monitoring node.
[0014] In one possible implementation, the first screening based on the acquired time series of monitoring values of multiple water quality parameters for each monitoring node in the target drainage network during a first time period to determine the target water quality parameters for each monitoring node during the first time period includes: performing collinearity analysis on the time series of monitoring values of at least some water quality parameters for each monitoring node during the first time period to obtain multiple optional parameters with collinearity for each monitoring node; and taking the optional parameters matched for each water sample type at each monitoring node during the first time period as the target water quality parameters for the corresponding monitoring node during the first time period.
[0015] In one possible implementation, the first time period includes a first sub-time period and a second sub-time period following the first sub-time period, each of which includes multiple monitoring time points. Specifically, determining at least one first target time period for each monitoring node within the first time period, based on the time series analysis of the monitoring values of the target water quality parameters and the target flow rate monitoring values for each monitoring node within the first time period, includes: inputting the monitoring values of the target water quality parameters and the target flow rate monitoring values for each monitoring node within the first sub-time period into a time series model for calculation to obtain prediction results for each monitoring node at each monitoring time point within the second sub-time period, the prediction results including predicted water quality parameter values and predicted flow rates; and determining at least one first target time period for each monitoring node within the second sub-time period based on the analysis results of the predicted water quality parameter values and the target water quality parameter monitoring values, and the analysis results of the predicted flow rates and the target flow rates monitoring values for each monitoring node at each monitoring time point within the second sub-time period.
[0016] In one possible implementation, the prediction result further includes a water quality parameter confidence interval and a flow confidence interval; wherein, based on the analysis results of the predicted water quality parameter values and target water quality parameter monitoring values at each monitoring time point in the second sub-time period for each monitoring node, and the analysis results of the predicted flow value and target flow monitoring value, at least one first target time period for each monitoring node is determined from the second sub-time period for each monitoring node, including: for each monitoring node, taking the time period formed by the monitoring time points corresponding to the target water quality parameter monitoring values outside the water quality parameter confidence interval as the first target time period for the monitoring node, and taking the time period formed by the monitoring time points corresponding to the target flow monitoring values outside the flow confidence interval as the first target time period for the monitoring node.
[0017] In one possible implementation, determining at least one second target time period for a monitoring node from all first target time periods based on the results of a second screening of the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each corresponding first target time period includes: performing equipment anomaly detection on the monitoring equipment at each monitoring node based on the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each first target time period, obtaining equipment detection results for each monitoring node, wherein the equipment detection results include monitoring equipment anomaly or monitoring equipment normal operation; if there is a monitoring node whose equipment detection result is monitoring equipment anomaly, then determining the equipment anomaly time period corresponding to the monitoring equipment anomaly, and removing the equipment anomaly time period from the first target time period of that monitoring node to obtain at least one second target time period for that monitoring node.
[0018] In one possible implementation, the step of determining at least one second target time period for each monitoring node from all first target time periods based on the results of a second screening of the time series of monitoring values of target water quality parameters and the time series of monitoring values of target flow for each of the monitoring nodes in each corresponding first target time period further includes: performing data anomaly detection on the time series of monitoring values of target water quality parameters and the time series of monitoring values of target flow for each of the monitoring nodes in each first target time period; determining abnormal data of target water quality parameters from the time series of monitoring values of target water quality parameters and abnormal data of flow from the time series of monitoring values of target flow; wherein the data anomaly detection includes at least one of abnormal noise detection, data jitter detection, baseline anomaly detection, and abnormal jump detection; removing abnormal data that matches the abnormal data of target water quality parameters and / or abnormal data of flow in the first target time periods of each of the monitoring nodes to obtain at least one second target time period for each of the monitoring nodes.
[0019] In one possible implementation, the device further includes a selection module for: before performing the clustering analysis, removing second target time periods from the second target time periods of each monitoring node that do not meet the selection conditions, based on preset selection conditions.
[0020] In one possible implementation, the device further includes a calculation module for: performing Monte Carlo calculations based on the classification results corresponding to each of the monitoring nodes, and determining the water composition analysis results of the target drainage network corresponding to each of the monitoring nodes, wherein the composition analysis results represent the basis for determining the flow rate of each water sample type at the corresponding monitoring node.
[0021] According to another aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-described method when executing instructions stored in the memory.
[0022] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided that stores computer program instructions thereon, wherein the computer program instructions, when executed by a processor, implement the above-described method.
[0023] According to another aspect of this disclosure, a computer program product is provided, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.
[0024] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0025] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.
[0026] Figure 1 This diagram illustrates a method for processing water environment monitoring data according to an embodiment of the present disclosure.
[0027] Figure 2 This diagram illustrates the predicted water quality parameter values and confidence intervals provided in the embodiments of this disclosure.
[0028] Figure 3 This diagram illustrates a method for processing water environment monitoring data according to an embodiment of the present disclosure.
[0029] Figure 4 A block diagram of a water environment monitoring data processing apparatus provided in an embodiment of this disclosure is shown. Detailed Implementation
[0030] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0031] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0032] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0033] To facilitate understanding of the technical solutions provided by the embodiments of this disclosure by those skilled in the art, the technical environment for implementing the technical solutions will be described below.
[0034] Damage and misconnections in municipal pipe networks can cause numerous problems. For example, they can lead to sewage being discharged into river and lake systems via stormwater systems, increase the operating costs of sewage treatment plants and pumping stations, reduce the treatment efficiency of sewage treatment plants, and cause combined sewer and stormwater pipe capacity to be occupied by external water, reducing flood drainage capacity. Furthermore, when reclaimed water plants are impacted by upstream water flow, the complexity of upstream pipelines and the effects of dilution or degradation from branch pipelines increase the difficulty of accurately tracing the source of anomalies within the overall complex system.
[0035] Related technologies provide quantitative assessment systems, such as quantitative analysis methods for external water inflow and infiltration based on the chemical mass balance equation of stormwater and sewage pipe networks. However, since the changes in water quality parameters over time in actual drainage systems generally conform to a certain statistical distribution, the statistical distribution of water quality parameters in actual drainage systems can be inferred through representative sample collection, so as to solve the proportion of different types of incoming water in stormwater and sewage pipe networks using Monte Carlo simulation methods.
[0036] However, in actual monitoring scenarios, the statistical distribution of the monitored data is quite wide, which can lead to problems such as non-convergence of calculation results or abnormally high coefficients of variation when using Monte Carlo simulation methods. At the same time, abnormal operating states in drainage pipe networks (such as a sudden large influx of external water or discharge of high-concentration sewage) are shorter in duration and harder to identify than normal operating states. This directly makes it difficult to obtain or identify water quality parameters under abnormal operating states as abnormal data, thus hindering the monitoring of the water environment.
[0037] To address the aforementioned technical problems, this disclosure provides a method for processing water environment monitoring data. Now, in conjunction with... Figures 1 to 3 The method for processing water environment monitoring data provided in the embodiments of this disclosure is illustrated.
[0038] like Figure 1 As shown, the processing method may include the following steps S101 to S104.
[0039] Step S101: Based on the obtained time series of monitoring values of various water quality parameters of each monitoring node in the target drainage network in the first time period, perform the first screening to determine the target water quality parameters of each monitoring node in the first time period.
[0040] Each monitoring node This is determined based on the distribution information of the target drainage network, where i can be a positive integer. For example, the distribution information of the target drainage network of interest is determined based on the topology map of the municipal pipeline network. Based on this distribution information, monitoring areas are first divided, and each monitoring area corresponds to a monitoring node, thereby establishing an online water quality and flow monitoring system to acquire real-time monitoring data including at least various water quality parameters and the target flow rate. The monitoring nodes can be determined by the land use and wastewater type of the monitoring area. Water sample type of the corresponding drainage network Water sample types may include, but are not limited to, black water, grey water, industrial water, groundwater, and river water. Since each monitoring node corresponds to a different monitoring area, and the land use and wastewater types may differ in different monitoring areas, the water sample types corresponding to each monitoring node may also differ.
[0041] Water quality parameters may include, but are not limited to, chemical oxygen demand (COD), conductivity, temperature, ammonia nitrogen, total hardness, five-day biochemical oxygen demand (BOD5), total phosphorus, total nitrogen, suspended solids, total dissolved solids, petroleum hydrocarbons, pH, anionic surfactants, cyanide, sulfide, fluoride, chloride, organophosphorus compounds, sulfate, mercury, chromium, cadmium, arsenic, lead, nickel, beryllium, silver, selenium, copper, zinc, manganese, iron, volatile phenols, benzene series compounds, aniline compounds, and nitrobenzene. The specific types of various water quality parameters obtained are set according to actual needs, and this disclosure does not limit this aspect.
[0042] Water quality parameters can be obtained in situ and in real time using a spectral sensor. For example, the sampling frequency can be 3-60 minutes / time, preferably 5-30 minutes / time, particularly preferably 8-20 minutes / time, and most preferably 10-15 minutes / time. Compared with the method of sampling water quality and then conducting laboratory tests to measure water quality concentration, this treatment method can obtain water quality concentration at a higher frequency, thereby effectively capturing the water quality characteristics of the drainage network.
[0043] A monitoring value time series refers to a sequence of values for the same indicator arranged chronologically, where the indicator refers to water quality parameters or flow rate. At each monitoring node, target water quality parameters can be selected from multiple water quality parameters acquired in the first time period. The target water quality parameters can be of at least two types. Step S101 may include: performing collinearity analysis on the monitoring value time series of at least some water quality parameters at each monitoring node in the first time period to obtain multiple optional parameters with collinearity for each monitoring node; and using the optional parameters matching each water sample type corresponding to each monitoring node in the first time period as the target water quality parameters for that monitoring node in the first time period. In this way, this processing method can identify collinear water quality parameters at each monitoring node through collinearity analysis, thereby avoiding the participation of all water quality parameters in subsequent calculations, saving computational resources and time, improving subsequent computational efficiency, and by selecting water quality parameters matching the water sample type, it can more accurately monitor water quality issues related to specific water sample types.
[0044] Before conducting collinearity analysis, all water quality parameters can be initially screened. Specifically, this can be done by selecting water quality parameters that are more relevant to the land use and water type corresponding to the monitoring node (such as industrial water, residential water, and river water). Then, collinearity analysis can be performed based on these selected water quality parameters. This avoids all water quality parameters from participating in collinearity analysis and subsequent calculations, saving computational resources and analysis time.
[0045] Alternatively, collinearity analysis can be performed directly on all acquired water quality parameters. For example, this can be done from monitoring nodes. The water conditions at the corresponding drainage network were used to determine the monitoring nodes. The drainage network, under normal operating conditions, contains a mixture of groundwater and industrial water samples. Normal operating conditions refer to the network's operation without equipment malfunctions, water quality fluctuations, or other abnormalities. This is observed at monitoring nodes. The monitoring time series obtained included the monitoring values of various water quality parameters for the first time period at 7 monitoring time points: Chemical Oxygen Demand (COD) {A1, A2, A3, A4, A5, A6, A7}, Temperature {B1, B2, B3, B4, B5, B6, B7}, Ammonia Nitrogen {E1, E2, E3, E4, E5, E6, E7}, and Total Hardness {D1, D2, D3, D4, D5, D6, D7}. The monitoring value A1 represents the monitoring node. The chemical oxygen demand (COD) monitoring value at the first monitoring time point is used, and the rest are analyzed similarly. Based on the collinearity analysis of the time series of all water quality parameters, including COD, temperature, and ammonia nitrogen, the monitoring nodes can be determined. The present invention discloses a number of optional parameters with strong collinearity. Collinearity analysis can be achieved by correlation coefficient method, eigenvalue method, etc., and the present invention does not limit the specific parameters.
[0046] Taking the correlation coefficient method as an example, and referring to the "monitoring value time series" as "time series" for short, this explains the process of collinearity analysis. The first correlation coefficient between the time series of any two water quality parameters can be determined, resulting in multiple first correlation coefficients. The time series of two water quality parameters whose first correlation coefficients satisfy a preset first screening condition are identified as the target water quality parameter time series. This first screening condition can be that the first correlation coefficient is greater than a preset first value (e.g., 0.8). For example, the first correlation coefficient between the time series of water quality parameters M and N is 0.9, which is greater than 0.8, indicating strong collinearity between the time series of M and N. The water quality parameters with strong collinearity are then selected as optional parameters. In this example, the optional parameters include oxygen demand, temperature, and total hardness. Due to the monitoring nodes... The mixture contains groundwater and industrial water. Based on chemical oxygen demand (COD), temperature, and total hardness, the optimal parameters for matching groundwater are total hardness and for matching industrial water are COD. Therefore, both total hardness and COD can be used as monitoring nodes for the first time period. The target water quality parameters for the first time period. The same principles apply to the remaining monitoring nodes. For the sake of brevity, this will not be elaborated further. Since the types of water samples at each monitoring node may differ, the target water quality parameters at each monitoring node may also differ in the first time period. For example, in this case, the monitoring node... The target water quality parameters for the first time period were total hardness and chemical oxygen demand (COD), with other monitoring nodes... The target water quality parameters for the first time period may be petroleum hydrocarbons, chemical oxygen demand (COD), and ammonia nitrogen.
[0047] Step S102: Based on the time series analysis of the monitoring values of the target water quality parameters and the monitoring values of the target flow rate of each monitoring node in the first time period, at least one first target time period for each monitoring node is determined from the first time period.
[0048] The first time period includes a first sub-time period and a second sub-time period following the first sub-time period. The first and second sub-time periods each include multiple monitoring time points. Time series analysis refers to the prediction of water quality parameter values and flow rates using time series models. The time series model used here refers to a model that predicts water quality parameter values and flow rates based on the time series of monitored values, such as the Periodic Autoregressive Integrated Moving Average Model (PARIMA), the Autoregressive Integrated Moving Average Model (ARIMA), the Generalized Autoregressive Conditional Heteroskedasticity Model (GARCH), and the Long Short-Term Memory (LSTM) network, etc. This disclosure does not limit the specific models used.
[0049] The time series model provided by this processing method can predict future data based on actual acquired monitoring data and predicted data. Specifically, for future data within a given time period to be predicted, the time series model can predict the future data at the starting time point within the given time period based on the actual acquired monitoring data prior to the given time period. The time series model can also predict the future data at non-starting time points within the given time period based on the actual acquired monitoring data prior to the given time period and the predicted data at the previous time point, where the previous time point is relative to the time point corresponding to the future data. For example, taking time points t1, t2, t3, and t4 as examples, assuming we want to predict the future data from time points t1 to t4, and the future data for time point t1 is the first data to be predicted after the time series model runs, then the future data for t1 is predicted based on the measured data from a period prior to t1. The future data for t2 is obtained from the measured data before t2 and the predicted data t1 at the time point prior to the predicted data t2. The prediction methods for the future data of t3 and t4 are similar to those for the future data of t2, and will not be elaborated further.
[0050] Step S102 may include: inputting the time series of the target water quality parameter monitoring value and the time series of the target flow monitoring value of each monitoring node in the first sub-time period into the time series model for calculation, to obtain the prediction results of each monitoring node at each monitoring time point in the second sub-time period, the prediction results including the predicted water quality parameter value and the predicted flow value; based on the analysis results of the predicted water quality parameter value and the target water quality parameter monitoring value, and the analysis results of the predicted flow value and the target flow monitoring value at each monitoring time point in the second sub-time period, determining at least one first target time period for each monitoring node from the second sub-time period, wherein the target water quality parameter monitoring value comes from the time series of the target water quality parameter monitoring value, and the target flow monitoring value comes from the time series of the target flow monitoring value. In this way, this processing method predicts future water quality monitoring data based on the actual acquired monitoring data using a time series model, which can adapt to different water quality and flow change trends, maintain high prediction accuracy, and, based on the data prediction value and the actual acquired monitoring data at the same monitoring time point as the data prediction value, preliminarily screen out the first target time period where the actual acquired monitoring data deviates from the prediction result.
[0051] Taking chemical oxygen demand (COD) as an example, one of the target water quality parameters, assuming that the actual monitoring nodes are... The time series of chemical oxygen demand (COD) monitoring values and target flow rate monitoring values are calculated for the period from 00:00 on March 15th to 00:00 on March 18th (i.e., the first time period). The COD monitoring value time series includes COD values from multiple monitoring time points (i.e., one of the target water quality parameters), and the target flow rate monitoring value time series includes target flow rate values from multiple monitoring time points. The period from 00:00 on March 15th to 00:00 on March 17th can be designated as the first sub-time period (T1), and the period from 00:00 on March 17th to 00:00 on March 18th as the second sub-time period (T2). The COD monitoring value at 00:00 on March 17th is included in the COD monitoring value time series of the first sub-time period, and the target flow rate monitoring value at 00:00 on March 17th is also included in the target flow rate monitoring value time series of the first sub-time period. However, the specific division of sub-time periods can be adjusted according to actual conditions; this is only an example. Monitoring nodes... The monitoring time series of chemical oxygen demand (COD) values at T1 is input into a time series model for calculation (or "prediction") to obtain the monitoring nodes. The predicted concentration values of chemical oxygen demand (i.e., predicted water quality parameter values) at each monitoring time point in T2, and the monitoring nodes The target flow monitoring value time series of T1 is input into the time series model for calculation (or "prediction") to obtain the monitoring nodes. The predicted flow rates at each monitoring time point in T2. Then, based on the monitoring nodes... The analysis results of the predicted chemical oxygen demand (COD) concentration and the monitored COD value at each monitoring time point in T2, as well as the analysis results of the predicted flow rate and the monitored target flow rate, determine the monitoring nodes from T2. At least one first target time period.
[0052] In this example, the analysis results can be determined by judging whether the difference between the actual monitored chemical oxygen demand (COD) value and the predicted COD concentration value predicted by the time series model at each monitoring time point in T2 falls within the corresponding preset interval. The first target time period is determined based on the monitoring time points where the difference does not fall within the preset interval. The analysis and judgment process for the target flow rate monitoring value and the predicted flow rate value is similar to that for the COD monitoring value and the predicted COD concentration value, and will not be repeated here. The first target time periods determined based on the analysis and judgment process for the COD monitoring value and the predicted COD concentration value, and the first target time periods determined based on the analysis and judgment process for the target flow rate monitoring value and the predicted flow rate value, are collectively used as monitoring nodes. At least one first target time period is defined, and each first target time period is different. The determination of the first target time periods for the remaining monitoring nodes is similar to that for the monitoring nodes. For the sake of brevity, this will not be elaborated upon further. Additionally, the division of sub-time periods for other monitoring nodes can be correlated with the monitoring nodes themselves. The sub-time periods can be the same (see T1 and T2 above), or different for different monitoring nodes, depending on actual needs. The time series model can be trained using monitoring data from the drainage network under normal operating conditions. Normal operating conditions refer to the operating state of the drainage network when there are no abnormalities such as equipment malfunctions or water quality fluctuations. The initial model can be, but is not limited to, PARIMA, Autoregressive Integrated Moving Average (ARIMA), Long Short-Term Memory (LSTM), or Generalized Autoregressive Conditional Heteroskedasticity (GARCH) models. Both water quality parameters and flow rates can be predicted using the same time series model, or two separate models can be used. Using the same time series model is preferred because different time series models will produce errors, and using the same model can save computational resources to some extent. Since water quality parameter monitoring data for drainage pipe networks typically exhibit obvious periodicity, such as daily or seasonal cycles, the PARIMA model, by introducing periodic components (such as seasonal differences and periodic autoregressive terms), can effectively capture these periodic changes and improve prediction accuracy compared to other models. Therefore, the PARIMA model is preferred as the time series model. However, in reality, the specific initial model and training method can be set according to the actual situation, and this disclosure does not limit this aspect.
[0053] The prediction results may also include confidence intervals for water quality parameters and confidence intervals for flow rates. The confidence intervals for water quality parameters at each monitoring time point in the second sub-time period for each monitoring node are determined based on the predicted water quality parameter values at the corresponding monitoring time points. Similarly, the confidence intervals for flow rates at each monitoring time point in the second sub-time period for each monitoring node are determined based on the predicted flow rates at the corresponding monitoring time points. Step S102, based on the analysis results of the predicted water quality parameter values and target water quality parameter monitoring values, and the analysis results of the predicted flow rates and target flow rates at each monitoring time point in the second sub-time period for each monitoring node, determines at least one first target time period for that monitoring node. This may include: for each monitoring node, using the time period formed by the monitoring time points corresponding to the target water quality parameter monitoring values outside the water quality parameter confidence interval as the first target time period for that monitoring node, and using the time period formed by the monitoring time points corresponding to the target flow rate monitoring values outside the flow rate confidence interval as the first target time period for that monitoring node. Thus, by introducing confidence intervals for water quality parameters and flow rates, this processing method can quantify the uncertainty of the prediction results, thereby more accurately determining the first target time period when the drainage network's operating status is suspected of being abnormal.
[0054] The monitoring nodes mentioned above are still in use. For example, Figure 2 As shown, the predicted water quality parameter values (i.e., predicted chemical oxygen demand concentration values) at each monitoring time point in T2 can be calculated using the time series model, as well as the chemical oxygen demand concentration range (denoted as S1) when the drainage network is under normal operating conditions. Figure 2 The confidence interval for normal pipeline network conditions and the chemical oxygen demand (COD) concentration interval for abnormal pipeline network conditions (denoted as S2, i.e., S2) are both considered normal. Figure 2 The confidence interval for characterizing pipeline network anomalies and the chemical oxygen demand concentration interval for severe pipeline network anomalies (denoted as S3, i.e.) Figure 2 The confidence interval for characterizing severe anomalies in the pipeline network is defined in the data. S1, S2, and S3 at each monitoring time point can be determined based on the predicted water quality parameter values at the corresponding monitoring time point. S1 (or S2 or S3, preferably S1) at each monitoring time point can be used as the monitoring node. The confidence intervals for water quality parameters are determined. In this example, S1 at each monitoring time point is ±1 standard deviation of the predicted water quality parameter value, S2 is ±2 standard deviation of the predicted water quality parameter value, and S3 is ±3 standard deviation of the predicted water quality parameter value. Monitoring nodes are then determined. The process of determining the flow confidence intervals at each monitoring time point in T2 is similar to determining the monitoring nodes. The process of defining the confidence intervals for water quality parameters at each monitoring time point in T2 will not be elaborated further. At each monitoring time point in T2, it is determined whether the actual monitored chemical oxygen demand (COD) value (i.e., one of the target water quality parameters) at the corresponding monitoring time point is within the confidence interval for the water quality parameter at that monitoring time point. If the COD value is within the confidence interval, it is considered as monitoring data indicating that the drainage network is in normal operation. If the COD value is outside the confidence interval, a corresponding time period is formed or divided based on the monitoring time point corresponding to the COD value. The specific method of forming or dividing the time period can be flexibly set according to the actual situation, and this embodiment does not limit this. The process of determining whether the target flow rate is within the flow rate confidence interval is similar to the process of determining whether the target water quality parameter is within the water quality parameter confidence interval, and will not be elaborated further. The time periods determined based on the judgment results of the target water quality parameter monitoring values and the water quality parameter confidence interval, and the time periods determined based on the judgment results of the target flow monitoring values and the flow confidence interval, are together used as monitoring nodes. At least one first target time period, representing a period during which the drainage network's operational status is suspected of being abnormal. The same applies to the remaining monitoring nodes. For the sake of brevity, this article will not elaborate further.
[0055] After identifying the first target time period where the drainage network's operational status is suspected to be abnormal, anomaly detection can be used to further filter out a second target time period where the drainage network's operational status is abnormal. The main difference between the first and second target time periods is that the anomaly corresponding to the first target time period may be caused by network operation problems or non-network operation problems, while the anomaly corresponding to the second target time period may be caused by network operation problems such as inflow / seepage, cross-connections, sudden large-scale intrusion of external water, or discharge of high-concentration sewage.
[0056] Step S103: Based on the results of the second screening of the time series of the monitoring values of the target water quality parameters and the time series of the monitoring values of the target flow rate of each monitoring node in the corresponding first target time period, at least one second target time period of the monitoring node is determined from all the first target time periods of each monitoring node.
[0057] The second screening process provided by this method specifically involves equipment anomaly detection and data anomaly detection. Generally, equipment anomaly detection is performed first, followed by data anomaly detection. This allows for a faster and more accurate determination of the detection results, as well as whether the cause of the detected results is equipment anomaly or environmental anomaly.
[0058] Equipment anomalies in equipment anomaly detection refer to deviations of equipment operating parameters or states from the normal range. Specifically, this can manifest as performance degradation (such as freezing), equipment malfunction, or unexpected output. Many factors can cause equipment anomalies. For example, if the equipment's operating temperature does not meet the monitoring requirements (e.g., the equipment needs to operate between -10°C and 30°C), but the actual operating temperature does not meet this requirement, the equipment may malfunction or freeze. Another example is if the equipment fails to contact the water body it should be testing, or if it does not, it will also cause equipment anomalies.
[0059] Data anomalies in data anomaly detection are usually caused by environmental anomalies, such as the device's lens being entangled in aquatic plants, causing fluctuations in the collected data.
[0060] Step S103 may include: performing equipment anomaly detection on the monitoring equipment at each monitoring node based on the time series of the monitoring values of the target water quality parameters and the time series of the monitoring values of the target flow rate at each monitoring node in each first target time period, and obtaining the equipment detection results of each monitoring node. The equipment detection results indicate whether the monitoring equipment is abnormal or normal. For example, for a certain monitoring node, the change in the time series of the target flow rate monitoring value can be used to determine whether the flow monitoring equipment at that monitoring node is abnormal. For example, if the flow rate monitoring value in the sequence after a certain time point is higher than the range or lower than the detection limit, it indicates that the equipment detection result of that monitoring node is abnormal. The equipment detection result can be obtained by jointly judging the time series of the monitoring values of the target water quality parameters and the time series of the target flow rate monitoring value, or only one of the sequences can be selected. If there is a monitoring node whose equipment detection result is abnormal, the equipment abnormal time period corresponding to the equipment abnormality is determined, and the equipment abnormal time period in the first target time period of that monitoring node is removed to obtain at least one second target time period of that monitoring node. In this way, this processing method ensures the accuracy and reliability of the actual collected monitoring data by detecting equipment anomalies. The equipment detection results help identify whether the monitoring equipment is malfunctioning, thereby avoiding errors caused by equipment failure or equipment being out of water. Furthermore, if the equipment detection result indicates an anomaly, an alert message can be issued to remind staff to address the anomaly promptly, reducing the loss or error of monitoring data due to equipment failure and minimizing the impact of equipment malfunction on water quality monitoring in the drainage network.
[0061] Step S103 may further include: performing data anomaly detection on the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each first target time period; identifying abnormal data of target water quality parameters from the time series of monitoring values of target water quality parameters and abnormal data of flow from the time series of monitoring values of target flow. The data anomaly detection includes at least one of abnormal noise detection, data jitter detection, baseline anomaly detection, and abnormal jump detection. The step S103 also includes removing abnormal time periods from the first target time period of each monitoring node that match abnormal data of target water quality parameters and / or abnormal data of flow, thereby obtaining at least one second target time period for each monitoring node. The abnormal data can be either outlier values or abnormal time series. For abnormal noise detection, the abnormal data is a single value; for data jitter detection, baseline anomaly detection, and abnormal jump detection, the abnormal data is a time series. In this way, by using data anomaly detection methods such as abnormal noise detection, data jitter detection, baseline anomaly detection, and abnormal jump detection, this processing method can eliminate data anomalies caused by non-pipeline operation problems and obtain the second target time period corresponding to water quality anomalies caused by pipeline operation problems. This helps to reduce misjudgments of water quality anomalies and improve the reliability of water environment monitoring.
[0062] Abnormal noise detection is used to detect abnormal noise in monitoring data that shows a large abrupt change at a certain moment (i.e., abnormal values of target water quality parameters or abnormal flow rates), but quickly returns to normal.
[0063] Data jitter typically manifests as instability in data values, causing significant discrepancies between data obtained within the same time period and the actual situation. Data jitter detection is used to identify abnormal jitter data (i.e., abnormal time series of target water quality parameters, abnormal time series of flow rates) from monitoring data. For example, the data to be detected can be obtained by first removing the data from the collected monitoring data. The data from the air refers to the data values measured by the monitoring equipment when the water is removed. Then, the relative difference value of the data to be detected can be calculated using ABS (absolute value) differencing. ABS differencing is the process of performing a difference operation on two or more ABS values. Differential analysis refers to the operation of subtracting adjacent values from a set of values. Taking chemical oxygen demand (COD) as a target water quality parameter as an example, abs difference calculation can be performed based on the time series of COD monitoring values to identify outliers of the target water quality parameter. For example, the relative difference value of COD at three adjacent monitoring time points (denoted as Q1, Q2, and Q3) can be calculated as abs(Q3-Q1) / Q2, where abs represents the absolute value, and the data in the monitoring time series is not repeated. If the value of abs(Q3-Q1) / Q2 is greater than 0.1, then Q1~Q3 is considered an abnormal time series of COD, thereby removing the abnormal time periods in the first target time period that match Q1~Q3 (i.e., the time periods formed based on y1, y2, and y3). In fact, the specific methods for implementing data jitter detection and other anomaly detection can be selected according to the actual situation, and this disclosure does not limit them.
[0064] Baseline anomaly, also known as baseline drift, refers to the linear change in the difference between monitored data and actual data over time, which may be caused by zero-point drift, temperature drift, etc. Baseline anomaly detection is used to detect abnormal drift data (i.e., abnormal time series of target water quality parameters, abnormal time series of flow rates) from the monitoring data. Baseline anomaly detection can be achieved using existing technologies, such as linear regression, etc., and this application does not impose specific limitations on this.
[0065] Anomaly jump detection is used to detect abnormal data (i.e., abnormal time series of target water quality parameters, abnormal flow time series) that are fixed multiples of the true values from monitoring data. The specific identification method for anomaly jump detection is explained below. Similar to data jitter detection, the water-free data in the collected monitoring data can be removed first to obtain the data to be detected. Then, the following calculations are performed on the data to be detected. Taking chemical oxygen demand (COD) as a target water quality parameter, based on the COD monitoring value time series (denoted as H) of the first target time period, a first moment (denoted as t0) is selected. The values of COD (denoted as Y0) for multiple preset time periods (e.g., per hour) before the first moment are determined. The increase ratio of COD (denoted as Y1) at the first moment to the multiple Y0s is calculated. The increase ratio can be calculated by [(Y1-Y0) / Y0]*100%. If each increase ratio exceeds the preset ratio threshold, and the monitoring value of COD in each preset time period after the first moment meets the aforementioned increase ratio threshold compared with the multiple Y0s, then it is determined that there is an abnormal jump in H of the first target time period. In the case that the monitoring value time series of COD has an abnormal jump, the abnormal time period can be determined based on the first moment and the preset time period after the first moment.
[0066] This processing method may further include: before performing cluster analysis, based on preset selection criteria, removing second target time periods from the second target time periods of each monitoring node that do not meet the selection criteria. For example, suppose the monitoring node... The target water quality parameters include two types. The time series of the monitoring values of these two target water quality parameters are denoted as C1 and C2, respectively, and the monitoring nodes are... The time series of the target flow monitoring value is denoted as Q. Thus, by monitoring the nodes... Anomaly detection of C1, C2, and Q can identify abnormal values for the target water quality parameters in C1 and C2, and abnormal values for the flow rate in Q. Based on the abnormal values of the target water quality parameters and the abnormal flow rate, the corresponding abnormal time periods can be determined, thus obtaining... Figure 3 The anomaly detection results are shown. For example... Figure 3 As shown, the anomaly detection results for C1 include 5 second target time periods, the anomaly detection results for C2 include 4 second target time periods, and the anomaly detection results for Q include 3 second target time periods. From... Figure 3 It can be seen that the second target time periods corresponding to these anomalies caused by actual pipeline operation problems overlap to some extent. Therefore, time periods in which multiple anomalies occur simultaneously can be selected (e.g., Figure 3In this method, T' is used as the second target time period for subsequent cluster analysis. The preset selection criterion can be that the outlier score is greater than or equal to a preset threshold. The outlier score for each second target time period is calculated, and it is determined whether the outlier score meets the selection criterion. The second target time periods corresponding to outlier scores that do not meet the selection criterion are then removed, resulting in the second target time periods for subsequent cluster analysis. Outlier scores are calculated using an additive method. For example, for a given second target time period, if a C1 outlier appears within that time period, the outlier score is increased by 1. If a Q outlier also appears, the outlier score is increased by another 1. The final outlier score for that second target time period is 2. Figure 3 For the second target time period (denoted as T'), C1, C2, and Q are all identified as monitoring data under abnormal operating conditions at T', so the abnormality score for T' is 3. The other second target time periods are similar to T', and will not be elaborated further. Thus, we can obtain... Figure 3 The figure shows the anomaly scores for each second target time period. In this example, all second target time periods with anomaly scores less than 2 can be removed, thus allowing second target time periods with anomaly scores equal to or greater than 2 to be used as monitoring nodes. The time periods corresponding to anomalies caused by actual pipeline operation problems are identified so that subsequent cluster analysis can be performed based on these time periods. The same principle applies to other monitoring nodes. For the sake of brevity, this will not be elaborated further. It should be noted that the preset selection conditions can also be to select only the time period corresponding to the abnormal values of the target water quality parameters, which helps to focus on key information and improve the targeting of water quality monitoring. In fact, the selection conditions can be set according to actual needs, and this embodiment does not limit this.
[0067] Step S104: Perform cluster analysis based on the time series of the monitoring values of the target water quality parameters and the time series of the monitoring values of the target flow rate at each monitoring node in the corresponding second target time period to obtain the classification results corresponding to each monitoring node.
[0068] Each classification result indicates the operational status of the pipeline network at the corresponding monitoring node. The classification results can be used as a basis for analyzing the water composition of the target drainage pipeline network at each monitoring node. For example, suppose the monitoring node... There are 18 corresponding second target time periods. The monitoring data for each of these 18 second target time periods includes time series of monitoring values for three target water quality parameters and time series of monitoring values for target flow. For each second target time period, the average values of the time series of monitoring values for the three target water quality parameters and the target flow were calculated, resulting in 72 average values. Based on these 72 average values, k-means clustering analysis was performed, yielding classification results including four major anomaly categories. Each element in each major anomaly category represents an anomalous event, which has a corresponding second target time period and the time series of monitoring values for the three target water quality parameters and the target flow within that second target time period. Therefore, each of the three target water quality parameters and the target flow has a corresponding anomaly category label. Monitoring Nodes The classification results include the labels of the three target water quality parameters and target flow rates corresponding to each of the 18 second target time periods, and the anomaly categories to which they belong. The classification for other monitoring nodes is similar. For the sake of brevity, this article will not elaborate further. In this way, this processing method refines the pipeline network operation problem by splitting monitoring data through abnormal events, which helps to break down the water composition (proportion) of the drainage network under abnormal events, thereby making it easier to determine the cause of the abnormal event, such as pipeline rupture and river backflow, illegal discharge by industrial enterprises, etc. In addition, this processing method uses a large amount of online monitoring data, which can increase the representativeness of the statistical distribution.
[0069] This processing method may further include: performing Monte Carlo calculations based on the classification results corresponding to each monitoring node, and determining the water composition analysis results of the drainage network at each monitoring node. The composition analysis results represent the basis for determining the flow rate of each water sample type at the corresponding monitoring node. The monitoring nodes mentioned above... For example, based on monitoring nodes The corresponding classification results and monitoring nodes Normal monitoring data (i.e., time series of target water quality parameters and target flow rate monitoring values under normal pipeline network operation conditions; in this example, the monitoring data of the remaining time periods after removing all first target time periods from the first time period) can generate statistical distribution results corresponding to each major category of anomalies and normal categories. This yields the maximum, minimum, mean, and standard deviation of each target water quality parameter and target flow rate. The purpose is to perform Monte Carlo calculations and chemical mass balance model calculations to obtain the monitoring nodes. The composition of the water in the corresponding pipe network, i.e. the proportion of each water sample type, is used to carry out routine operation and maintenance and supervision of the drainage pipe network.
[0070] With monitoring nodes Taking one of the major categories of anomalies as an example (monitoring nodes) The determination process for other major categories of anomalies and the components of incoming water (the proportion of each water sample type) under the normal category is similar to that of this major category of anomalies (and will not be repeated here). It is assumed that this major category of anomalies includes 10 elements, each element representing an anomaly event. The 10 anomaly events corresponding to these 10 elements have corresponding second target time periods and time series of monitoring values of three target water quality parameters and target flow monitoring values for the second target time period. The three target water quality parameters are ammonia nitrogen, chemical oxygen demand, and total hardness. Based on the time series of monitoring values of the three target water quality parameters and the target flow rate corresponding to this anomaly category, the statistical distribution results (maximum, minimum, mean, and standard deviation) of this anomaly category are obtained. Then, Monte Carlo calculations are performed based on these statistical distribution results to generate random values. These random values include multiple random values for each target water quality parameter and multiple random values for flow rate, with the number of random values for water quality parameters being the same as the number of random values for flow rate. One set of input data (including one random value for each target water quality parameter and one random value for flow rate) is randomly selected from these random values and input into the chemical mass balance model to calculate the component proportions, obtaining the component proportions of the incoming water from the target drainage network (i.e., the component analysis results). Based on these component proportions, it is determined whether the sum of the component proportions equals 1. If not, another set of input data is randomly selected from these random values for chemical mass balance model calculation until the sum of the component proportions equals 1. The final calculated component analysis result that satisfies the condition of a total component proportion equal to 1 is used as the monitoring node. The percentage of each water sample type is shown below, and the calculation process for the percentage of each component can be described as follows:
[0071] Since one monitoring node is set up for each monitoring area, let's assume that in this example, the monitoring node... The monitoring point contains three types of water samples: black water, grey water, and groundwater. This monitoring node is the outflow point (usually at the end of the pipeline network in the monitoring area). In this case, the calculation process for the proportion and flow rate of different water sample types in different anomaly categories is as follows. Taking ammonia nitrogen as a target water quality parameter as an example, the mass balance equation for ammonia nitrogen is as follows:
[0072] C 黑 Q 黑 +C 灰 Q 灰 +C 地下 Q 地下 =C 监测点 Q 监测点 Formula 1
[0073] In Equation 1, C 黑 Q represents the ammonia nitrogen content in black water. 黑 C represents the flow rate of black water. 灰 This indicates the ammonia nitrogen content in the greywater. C represents the flow rate of grey water. 地下 Q represents the ammonia nitrogen in groundwater. 地下 C represents the flow rate of groundwater. 监测点 The ammonia nitrogen at the monitoring node is determined by the time series of monitoring values of ammonia nitrogen, a target water quality parameter, and the specific method is not limited. This represents the flow rate at the monitoring node (determined through a time series of target flow rate monitoring values, the specific method is not limited). It can be understood that C in Equation 1... 黑 C 灰 C 地下 This is just an example; there might also be C. 工业 C 清水 The specific number depends on the type of water sample at each monitoring node, with at least two types of water samples at each monitoring node. The parameters in Equation 1 can be various monitoring values at the same time point, or values obtained by averaging monitoring values over a period of time.
[0074] Since the three target water quality parameters in this example are ammonia nitrogen, chemical oxygen demand (COD), and total hardness, similarly, the mass balance equation for COD is derived from the mass balance equation for ammonia nitrogen. The mass balance equation for total hardness is: The interpretations of the parameters in these two equations are similar to those in Equation 1, and for the sake of brevity, they will not be repeated here.
[0075] C in Equation 1 黑 C 灰 C 地下 These ammonia nitrogen data, and the mass balance equation corresponding to chemical oxygen demand, Isochemical oxygen demand data, and Y in the mass balance equation corresponding to total hardness 黑 The total hardness data, and other three types of data, can be obtained as follows: Data on drainage users within the monitoring area are compiled; representative drainage users of different types (such as residential communities, commercial plazas, industrial enterprises, etc.) are selected at their access manholes before connecting to municipal pipelines (where monitoring equipment is installed); and representative external clean water bodies (such as groundwater, river water, construction precipitation, etc.) are monitored online. Water quality characteristic data for different clean water bodies are collected, thus providing the required ammonia nitrogen, chemical oxygen demand, and total hardness data. It is clear that the data collection method is consistent with that at each monitoring point, using a spectral sensor for water quality monitoring, thereby collecting the corresponding time series of monitoring values. Now, it is necessary to solve for Q. 黑 Q 灰 Q 地下These three unknowns. Since this treatment method has already determined the main water quality parameters corresponding to black water, grey water, and groundwater sample types, such as ammonia nitrogen, COD, and total hardness, we can obtain the three unknowns in the mass balance equations for ammonia nitrogen, chemical oxygen demand, and total hardness mentioned above. These three equations can then be solved to obtain... However, the solution may not satisfy the relationship shown in Equation 2:
[0076] Q 黑 / Q 监测点 +Q 灰 / Q 监测点 +Q 地下 / Q 监测点 =1 Formula 2
[0077] The parameters in Equation 2 are explained above, and for the sake of brevity, they will not be repeated here. Therefore, C 监测点 O 监测点 , and Q 监测点 Based on the time series of monitoring values of the three target water quality parameters and the target flow rate corresponding to this anomaly category in this example, the statistical distribution results (maximum, minimum, mean, and standard deviation) of this anomaly category are obtained. Then, based on these statistical distribution results, random values are generated using the Monte Carlo method to determine the C values corresponding to the three water sample types: black water, grey water, and groundwater. 黑 C 灰 C 地下 O 黑 O 灰 O 地下 Y 黑 Y 灰 Y 地下 Methods and , , and The method is the same, also obtained through the Monte Carlo method, and will not be elaborated further. Thus, data usable for Equation 1 is obtained through the Monte Carlo calculation method. This data is then substituted into Equation 1 to calculate the components, and it is determined whether it conforms to Equation 2, i.e., whether the sum of the component proportions equals 1, until the final Q is obtained. 黑 Q 灰 Q 地下 The relationship shown in Formula 2 is satisfied, that is, the sum of the component proportions equals 1, thus obtaining the proportion of each water sample type.
[0078] This disclosure also provides a water environment monitoring data processing device, comprising: a first screening module, used to perform a first screening based on the acquired time series of monitoring values of multiple water quality parameters for each monitoring node in a target drainage network in a first time period, to determine the target water quality parameters of each monitoring node in the first time period, wherein each monitoring node is determined based on the distribution information of the target drainage network, and the types of target water quality parameters include at least two; and a time series analysis module, used to determine at least one first target time period for each monitoring node in the first time period based on the time series of monitoring values of target water quality parameters and the time series of monitoring values of target flow in the first time period, wherein the time series analysis refers to using a time series model to analyze water quality parameters. The system includes: a prediction of numerical values and flow rates; a second screening module, used to determine at least one second target time period for each monitoring node from all first target time periods based on the second screening results of the monitoring time series of target water quality parameters and target flow rates for each monitoring node in the corresponding first target time periods; and a cluster analysis module, used to perform cluster analysis based on the monitoring time series of target water quality parameters and target flow rates for each monitoring node in the corresponding second target time periods to obtain the corresponding classification results for each monitoring node. The classification results indicate the pipeline operation problems existing at the corresponding monitoring node, and the classification results are used as the basis for analyzing the water composition of the target drainage pipeline network for each monitoring node.
[0079] In one possible implementation, the first screening based on the acquired time series of monitoring values of multiple water quality parameters for each monitoring node in the target drainage network during a first time period to determine the target water quality parameters for each monitoring node during the first time period includes: performing collinearity analysis on the time series of monitoring values of at least some water quality parameters for each monitoring node during the first time period to obtain multiple optional parameters with collinearity for each monitoring node; and taking the optional parameters matched for each water sample type at each monitoring node during the first time period as the target water quality parameters for the corresponding monitoring node during the first time period.
[0080] In one possible implementation, the first time period includes a first sub-time period and a second sub-time period following the first sub-time period, each of which includes multiple monitoring time points. Specifically, determining at least one first target time period for each monitoring node within the first time period, based on the time series analysis of the monitoring values of the target water quality parameters and the target flow rate monitoring values for each monitoring node within the first time period, includes: inputting the monitoring values of the target water quality parameters and the target flow rate monitoring values for each monitoring node within the first sub-time period into a time series model for calculation to obtain prediction results for each monitoring node at each monitoring time point within the second sub-time period, the prediction results including predicted water quality parameter values and predicted flow rates; and determining at least one first target time period for each monitoring node within the second sub-time period based on the analysis results of the predicted water quality parameter values and the target water quality parameter monitoring values, and the analysis results of the predicted flow rates and the target flow rates monitoring values for each monitoring node at each monitoring time point within the second sub-time period.
[0081] In one possible implementation, the prediction result further includes a water quality parameter confidence interval and a flow confidence interval; wherein, based on the analysis results of the predicted water quality parameter values and target water quality parameter monitoring values at each monitoring time point in the second sub-time period for each monitoring node, and the analysis results of the predicted flow value and target flow monitoring value, at least one first target time period for each monitoring node is determined from the second sub-time period for each monitoring node, including: for each monitoring node, taking the time period formed by the monitoring time points corresponding to the target water quality parameter monitoring values outside the water quality parameter confidence interval as the first target time period for the monitoring node, and taking the time period formed by the monitoring time points corresponding to the target flow monitoring values outside the flow confidence interval as the first target time period for the monitoring node.
[0082] In one possible implementation, determining at least one second target time period for a monitoring node from all first target time periods based on the results of a second screening of the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each corresponding first target time period includes: performing equipment anomaly detection on the monitoring equipment at each monitoring node based on the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each first target time period, obtaining equipment detection results for each monitoring node, wherein the equipment detection results include monitoring equipment anomaly or monitoring equipment normal operation; if there is a monitoring node whose equipment detection result is monitoring equipment anomaly, then determining the equipment anomaly time period corresponding to the monitoring equipment anomaly, and removing the equipment anomaly time period from the first target time period of that monitoring node to obtain at least one second target time period for that monitoring node.
[0083] In one possible implementation, the step of determining at least one second target time period for each monitoring node from all first target time periods based on the results of a second screening of the time series of monitoring values of target water quality parameters and the time series of monitoring values of target flow for each of the monitoring nodes in each corresponding first target time period further includes: performing data anomaly detection on the time series of monitoring values of target water quality parameters and the time series of monitoring values of target flow for each of the monitoring nodes in each first target time period; determining abnormal data of target water quality parameters from the time series of monitoring values of target water quality parameters and abnormal data of flow from the time series of monitoring values of target flow; wherein the data anomaly detection includes at least one of abnormal noise detection, data jitter detection, baseline anomaly detection, and abnormal jump detection; removing abnormal data that matches the abnormal data of target water quality parameters and / or abnormal data of flow in the first target time periods of each of the monitoring nodes to obtain at least one second target time period for each of the monitoring nodes.
[0084] In one possible implementation, the device further includes a selection module for: before performing the clustering analysis, removing second target time periods from the second target time periods of each monitoring node that do not meet the selection conditions, based on preset selection conditions.
[0085] In one possible implementation, the device further includes a calculation module for: performing Monte Carlo calculations based on the classification results corresponding to each of the monitoring nodes, and determining the water composition analysis results of the target drainage network corresponding to each of the monitoring nodes, wherein the composition analysis results represent the basis for determining the flow rate of each water sample type at the corresponding monitoring node.
[0086] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0087] This disclosure also provides a water environment monitoring system, which performs water environment monitoring based on the above-described water environment monitoring data processing method. In some embodiments, the functions or modules included in the water environment monitoring system provided by this disclosure can be used to execute the methods described in the above method embodiments. Specific implementations can be referred to the description of the above processing method embodiments, and for brevity, will not be repeated here.
[0088] This disclosure also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the above-described method. The computer-readable storage medium can be volatile or non-volatile.
[0089] This disclosure also proposes an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0090] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.
[0091] Figure 4 A block diagram of a water environment monitoring data processing apparatus provided in an embodiment of this disclosure is shown. For example, apparatus 1900 may be provided as a server or terminal device. (Refer to...) Figure 4 The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0092] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932.TM macOS X TM Unix TM Linux TM FreeBSD TM Or similar.
[0093] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of the device 1900 to perform the above-described method.
[0094] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.
[0095] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0096] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0097] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0098] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0099] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0100] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0102] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for processing water environment monitoring data, characterized in that, include: Based on the obtained time series of monitoring values of various water quality parameters for each monitoring node in the target drainage network in the first time period, a first screening is performed to determine the target water quality parameters of each monitoring node in the first time period. The monitoring node is determined based on the distribution information of the target drainage network, and the types of target water quality parameters include at least two. Based on the time series analysis of the monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in the first time period, at least one first target time period is determined for each monitoring node in the first time period. The time series analysis refers to using a time series model to predict water quality parameter values and flow values. The actual monitoring data obtained in the first target time period deviates from the prediction results of the time series model. Based on the results of the second screening of the time series of the monitoring values of the target water quality parameters and the time series of the monitoring values of the target flow rate of each monitoring node in the corresponding first target time period, at least one second target time period of the monitoring node is determined from all the first target time periods of the monitoring node. Cluster analysis is performed on the time series of the target water quality parameters and the time series of the target flow rate monitoring values of each monitoring node in the corresponding second target time period to obtain the classification results corresponding to each monitoring node. The classification results indicate the pipeline operation problems existing at the corresponding monitoring node. The classification results are used as the basis for water composition analysis of each monitoring node for the target drainage pipeline network. The method further includes: performing Monte Carlo calculations based on the classification results corresponding to each monitoring node, and determining the water composition analysis results of the target drainage network corresponding to each monitoring node, wherein the composition analysis results represent the basis for determining the flow rate of each water sample type at the corresponding monitoring node; The step of determining at least one second target time period for each monitoring node from all first target time periods based on the results of a second screening of the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each corresponding first target time period includes: performing equipment anomaly detection on the monitoring equipment at each monitoring node based on the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each first target time period, and obtaining the equipment detection results of each monitoring node, wherein the equipment detection results include monitoring equipment anomaly or monitoring equipment normal operation; if there is a monitoring node whose equipment detection results are monitoring equipment anomaly, then the equipment anomaly time period corresponding to the monitoring equipment anomaly is determined, and the equipment anomaly time period in the first target time period of the monitoring node is removed to obtain at least one second target time period for the monitoring node.
2. The method according to claim 1, characterized in that, The first screening, based on the acquired time series of monitoring values of various water quality parameters for each monitoring node in the target drainage network during a first time period, determines the target water quality parameters for each monitoring node during the first time period, including: Collinearity analysis is performed on the time series of at least some water quality parameters monitored at each monitoring node during the first time period to obtain multiple optional parameters with collinearity for each monitoring node. The optional parameters matched for each type of water sample at each monitoring node during the first time period are used as the target water quality parameters for each monitoring node during the first time period.
3. The method according to claim 1, characterized in that, The first time period includes a first sub-time period and a second sub-time period following the first sub-time period. The first sub-time period and the second sub-time period each include multiple monitoring time points. Based on the time series analysis of the monitoring values of target water quality parameters and target flow rates of each monitoring node within the first time period, at least one first target time period is determined for each monitoring node within the first time period, including: The time series of the target water quality parameters and the time series of the target flow rate monitoring values of each monitoring node in the first sub-time period are respectively input into the time series model for calculation to obtain the prediction results of each monitoring node at each monitoring time point in the second sub-time period. The prediction results include the predicted water quality parameter values and the predicted flow rate values. Based on the analysis results of the predicted water quality parameter values and target water quality parameter monitoring values, and the analysis results of the predicted flow rate values and target flow rate monitoring values at each monitoring time point in the second sub-time period of each monitoring node, at least one first target time period of the monitoring node is determined from the second sub-time period of each monitoring node.
4. The method according to claim 3, characterized in that, The prediction results also include confidence intervals for water quality parameters and confidence intervals for flow rates; wherein, based on the analysis results of the predicted water quality parameter values and target water quality parameter monitoring values, and the analysis results of the predicted flow rate values and target flow rate monitoring values at each monitoring time point in the second sub-time period for each monitoring node, at least one first target time period for that monitoring node is determined from the second sub-time period for each monitoring node, including: for each monitoring node, the time period formed by the monitoring time points corresponding to the target water quality parameter monitoring values outside the confidence intervals for water quality parameters is taken as the first target time period for that monitoring node, and the time period formed by the monitoring time points corresponding to the target flow rate monitoring values outside the confidence intervals for flow rates is taken as the first target time period for that monitoring node.
5. The method according to claim 1, characterized in that, The step of determining at least one second target time period for each monitoring node from all first target time periods based on the results of a second screening of the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each corresponding first target time period further includes: Data anomaly detection is performed on the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each first target time period. Abnormal data of target water quality parameters are determined from the time series of monitoring values of target water quality parameters and abnormal data of flow are determined from the time series of monitoring values of target flow. The data anomaly detection includes at least one of abnormal noise detection, data jitter detection, baseline anomaly detection, and abnormal jump detection. Abnormal data that match the target water quality parameters and / or abnormal flow data in the first target time period of each monitoring node are removed to obtain at least one second target time period for each monitoring node.
6. The method according to claim 1 or 5, characterized in that, The method further includes: Before performing the clustering analysis, according to preset selection conditions, for each monitoring node, the second target time period that does not meet the selection conditions is removed from the second target time period of that monitoring node.
7. An electronic device, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to implement the method of any one of claims 1 to 6 when executing the instructions stored in the memory.
8. A device for processing water environment monitoring data, characterized in that, include: The first screening module is used to perform a first screening based on the acquired time series of monitoring values of multiple water quality parameters for each monitoring node in the target drainage network in a first time period, to determine the target water quality parameters of each monitoring node in the first time period, wherein each monitoring node is determined based on the distribution information of the target drainage network, and the types of target water quality parameters include at least two. The time series analysis module is used to determine at least one first target time period for each monitoring node based on the time series of the monitoring values of the target water quality parameters and the time series of the monitoring values of the target flow rate of each monitoring node in the first time period. The time series analysis refers to using a time series model to predict the water quality parameter values and flow rate values. The actual monitoring data obtained in the first target time period deviates from the prediction results of the time series model. The second screening module is used to determine at least one second target time period of each monitoring node from all the first target time periods of each monitoring node based on the results of the second screening of the time series of the monitoring values of the target water quality parameters and the time series of the monitoring values of the target flow rate of each monitoring node in the corresponding first target time periods. The cluster analysis module is used to perform cluster analysis based on the time series of the monitoring values of the target water quality parameters and the time series of the monitoring values of the target flow rate of each monitoring node in the corresponding second target time period, so as to obtain the classification results corresponding to each monitoring node. The classification results indicate the pipeline operation problems existing at the corresponding monitoring node. The classification results are used as the basis for the water composition analysis of each monitoring node for the target drainage pipeline network. The device further includes a calculation module for: performing Monte Carlo calculations based on the classification results corresponding to each monitoring node, and determining the water composition analysis results of the target drainage network corresponding to each monitoring node, wherein the composition analysis results represent the basis for determining the flow rate of each water sample type at the corresponding monitoring node; The step of determining at least one second target time period for each monitoring node from all first target time periods based on the results of a second screening of the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each corresponding first target time period includes: performing equipment anomaly detection on the monitoring equipment at each monitoring node based on the time series of monitoring values of target water quality parameters and target flow monitoring values of each monitoring node in each first target time period, and obtaining the equipment detection results of each monitoring node, wherein the equipment detection results include monitoring equipment anomaly or monitoring equipment normal operation; if there is a monitoring node whose equipment detection results are monitoring equipment anomaly, then the equipment anomaly time period corresponding to the monitoring equipment anomaly is determined, and the equipment anomaly time period in the first target time period of the monitoring node is removed to obtain at least one second target time period for the monitoring node.
Citation Information
Patent Citations
Power system abnormal data identifying and correcting method based on time series analysis
CN104766175A
Anti-interference method and device for industrial wireless network and computer equipment
CN117858090A
Drainage pipe network fault determination method, device and equipment
CN118378137A