Inflow infiltration position identification method and device, electronic equipment and storage medium
By combining time series data of water quality indicators with iterative processing of global optimization algorithms, the problem of accuracy in identifying inflow and infiltration locations in drainage pipe networks was solved, achieving higher identification accuracy and reducing errors, thus improving the identification capability of drainage pipe networks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CORE VISION (BEIJING) TECH CO LTD
- Filing Date
- 2025-11-05
- Publication Date
- 2026-07-24
Smart Images

Figure CN121479388B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of drainage network monitoring, and in particular to a method and apparatus for identifying inflow and infiltration locations, electronic equipment, and storage medium. Background Technology
[0002] Many cities suffer from combined sewer systems and infiltration of external water into their drainage networks. Untreated sewage mixed with rainwater and discharged directly into receiving water bodies exacerbates surface water pollution and seriously harms the urban water environment. Groundwater infiltration not only reduces the drainage capacity of rainwater pipes but can also wash soil particles from the surrounding area into the sewage pipes, creating erosion pits that significantly accelerate pipe instability and even lead to pipe bursts. When rainwater mixes with sewage or when large amounts of clean water infiltrate, it results in low influent concentrations at sewage treatment plants, a significant indicator of the inefficiency of urban sewage collection systems.
[0003] Therefore, accurately determining the inflow and seepage locations of drainage pipe networks is crucial and provides scientific guidance for subsequent mixed-connection renovation projects. However, the accuracy of related technologies in identifying the inflow and seepage locations of pipe networks is not satisfactory. Summary of the Invention
[0004] In view of this, this disclosure proposes a scheme for identifying the location of inflow and infiltration.
[0005] According to one aspect of this disclosure, a method for identifying inflow and infiltration locations is provided, comprising: obtaining multiple inflow water quality index time series based on water quality indicators of water flowing into a drainage unit in a drainage network from multiple inflow sources, and obtaining multiple outflow water quality index time series based on water quality indicators of water flowing out of the drainage unit; performing a first iterative process on the multiple inflow water quality index time series and the multiple outflow water quality index time series until the sum of the flow proportions of each inflow source at each time point meets a first preset condition, stopping the first iterative process, and obtaining a target flow proportion time series for each inflow source; determining inflow and infiltration units based on the target flow proportion time series of each drainage unit; the first iterative process comprising: determining the flow proportion time series of each inflow source based on the multiple inflow water quality index time series and the multiple outflow water quality index time series; if the sum of the flow proportions of each inflow source at at least one time point does not meet the first preset condition, performing a second iterative process on the multiple inflow water quality index time series and the multiple outflow water quality index time series using a global optimization algorithm until a second preset condition is met, stopping the second iterative process.
[0006] In one possible implementation, the method further includes: determining the flow time series of each inflow source of a single inflow infiltration unit based on the first flow time series of a single inflow infiltration unit and the target flow percentage time series; and determining the processing priority corresponding to each inflow infiltration unit according to the flow time series of each inflow infiltration unit, wherein the processing priority is used to formulate an inflow infiltration treatment plan.
[0007] In one possible implementation, the method further includes: acquiring the initial inflow water quality index time series and the initial outflow water quality index time series of each drainage unit; identifying and removing abnormal operating condition data from the initial inflow water quality index time series and the initial outflow water quality index time series to obtain an operating condition screening time series; performing anomaly detection on the operating condition screening time series, removing abnormal data and adding normal data to obtain an anomaly screening time series, wherein the anomaly screening time series includes: a first anomaly screening time series and a second anomaly screening time series; aligning the first anomaly screening time series and the second anomaly screening time series within the same drainage unit in the time dimension to obtain the plurality of inflow water quality index time series and the plurality of outflow water quality index time series.
[0008] In one possible implementation, the anomaly detection includes: data jitter anomaly detection; the anomaly detection of the working condition screening time series includes: sliding a first time window on the working condition screening time series to determine multiple first indicators, the first indicators representing the differences between data within the first time window; if the first indicator is greater than a first difference threshold, the data for which the first indicator is calculated is determined as the abnormal data.
[0009] In one possible implementation, the anomaly detection includes: anomaly baseline detection. The anomaly detection of the operating condition screening time series includes: determining first selected data from the operating condition screening time series; determining a first sub-time series located before the corresponding time of the first selected data in the operating condition screening time series based on a second time window; determining a second sub-time series in the operating condition screening time series starting with the first selected data and its corresponding time, based on a third time window; determining a second indicator characterizing the numerical change of the second sub-time series compared to the first sub-time series; determining a third indicator characterizing the trend of change of the first sub-time series, and a fourth indicator characterizing the trend of change of the second sub-time series; and determining the second sub-time series as the anomalous data when the second indicator falls within a first numerical range, the third indicator falls within a second numerical range, and the fourth indicator is greater than a first threshold.
[0010] In one possible implementation, determining the second indicator characterizing the numerical change of the second sub-time series compared to the first sub-time series includes: determining first target data based on a first quantile of the first sub-time series, determining second target data based on a second quantile of the second sub-time series, and determining the second indicator based on the first target data and the second target data.
[0011] In one possible implementation, the anomaly detection includes: anomaly jump detection. The anomaly detection of the working condition screening time series includes: identifying multiple first data points prior to a first acquisition time in the working condition screening time series; grouping the multiple first data points according to a first duration to obtain multiple first data groups; determining the median of each first data group; if the growth ratio of the second data points acquired at the first acquisition time relative to each of the medians is greater than a ratio threshold, acquiring each third data point acquired after the first acquisition time and within an adjacent second duration; if the growth ratio of each of the third data points relative to each of the medians is greater than the ratio threshold, identifying the second data points and each of the third data points as the anomaly data.
[0012] According to another aspect of this disclosure, an inflow / infiltration location identification device is provided, comprising:
[0013] The water quality index time series acquisition unit is used to obtain multiple inflow water quality index time series based on the water quality index of water flowing into the drainage unit in the drainage network from multiple inflow sources, and to obtain multiple outflow water quality index time series based on the water quality index of water flowing out of the drainage unit.
[0014] The target flow percentage time series acquisition unit is used to perform a first iterative process on the multiple inflow water quality index time series and the multiple outflow water quality index time series until the sum of the flow percentages of each inflow source at each time point meets a first preset condition, and then stop the first iterative process to obtain the target flow percentage time series of each inflow source.
[0015] An inflow and infiltration determination unit is used to determine the inflow and infiltration unit based on the time series of the target flow rate ratio of each drainage unit;
[0016] The first iterative process includes:
[0017] Based on the time series of the multiple inflow water quality indicators and the multiple outflow water quality indicators, determine the flow rate proportion time series of each inflow source;
[0018] If the sum of the flow proportions of each of the inflow sources at at least one moment does not meet the first preset condition, a global optimization algorithm is used to perform a second iteration on the time series of the multiple inflow water quality indicators and the time series of the multiple outflow water quality indicators until the second preset condition is met, and then the second iteration is stopped.
[0019] In one possible implementation, the device further includes:
[0020] A flow time series determination unit is used to determine the flow time series of each inflow source of a single inflow infiltration unit based on the first flow time series of a single inflow infiltration unit and the target flow proportion time series.
[0021] The processing priority determination unit is used to determine the processing priority corresponding to each of the inflow infiltration units based on the flow time series of each of the inflow infiltration units. The processing priority is used to formulate an inflow infiltration treatment plan.
[0022] In one possible implementation, the device further includes:
[0023] The initial water quality index time series acquisition unit is used to acquire the initial inflow water quality index time series and the initial outflow water quality index time series of each drainage unit.
[0024] The operating condition screening time series determination unit is used to identify and remove abnormal operating condition data from the initial inflow water quality index time series and the initial outflow water quality index time series to obtain the operating condition screening time series.
[0025] An anomaly screening time series determination unit is used to perform anomaly detection on the working condition screening time series, remove abnormal data from the working condition screening time series and fill in normal data to obtain an anomaly screening time series, wherein the anomaly screening time series includes: a first anomaly screening time series and a second anomaly screening time series.
[0026] The inflow and outflow water quality index time series determination unit is used to align the first anomaly screening time series and the second anomaly screening time series within the same drainage unit in the time dimension to obtain the multiple inflow water quality index time series and the multiple outflow water quality index time series.
[0027] In one possible implementation, the anomaly detection includes: data jitter anomaly detection, and the anomaly filtering time series determination unit is further configured to:
[0028] Slide a first time window on the working condition screening time series to determine multiple first indicators, where the first indicators characterize the differences between data within the first time window;
[0029] If the first indicator is greater than the first difference threshold, the data used to calculate the first indicator will be identified as the abnormal data.
[0030] In one possible implementation, the anomaly detection includes: anomaly baseline detection, and the anomaly screening time series determination unit is further configured to:
[0031] The first selected data is determined from the time series of the operating conditions.
[0032] Based on the second time window, a first sub-time series located before the time corresponding to the first selected data is determined in the working condition filtering time series;
[0033] Based on the third time window, in the working condition screening time series, a second sub-time series is determined with the first selected data and the corresponding time as the starting point;
[0034] Determine a second index characterizing the numerical change of the second sub-time series compared to the first sub-time series;
[0035] A third indicator characterizing the trend of change in the first sub-time series and a fourth indicator characterizing the trend of change in the second sub-time series are determined.
[0036] If the second indicator falls within the first value range, the third indicator falls within the second value range, and the fourth indicator is greater than the first threshold, the second sub-time series is identified as the abnormal data.
[0037] In one possible implementation, the anomaly screening time series determination unit is further configured to:
[0038] The first target data is determined based on the first quantile of the first sub-time series, and the second target data is determined based on the second quantile of the second sub-time series.
[0039] The second indicator is determined based on the first target data and the second target data.
[0040] In one possible implementation, the anomaly detection includes: anomaly jump detection, and the anomaly screening time series determination unit is further configured to:
[0041] Determine multiple first data points prior to the first acquisition time in the working condition screening time series;
[0042] The plurality of first data are grouped according to a first duration to obtain a plurality of first data groups;
[0043] Determine the median of each of the first data groups;
[0044] If the growth rate of the second data collected at the first acquisition time relative to each of the medians is greater than the ratio threshold, then the third data collected within the second time period after the first acquisition time and adjacent to it are acquired.
[0045] If the growth rate of each of the third data points relative to each of the medians is greater than the ratio threshold, then the second data point and each of the third data points are identified as the abnormal data.
[0046] According to another aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-described method when executing instructions stored in the memory.
[0047] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided that stores computer program instructions thereon, wherein the computer program instructions, when executed by a processor, implement the above-described method.
[0048] According to another aspect of this disclosure, a computer program product is provided, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.
[0049] In this embodiment, the flow rate proportion time series of each inflow source is determined using inflow and outflow water quality index time series. This time series is then combined with a global optimization algorithm, and through a first iteration, the target flow rate proportion time series is finally obtained. This combination reduces errors in the sampling and measurement process of the monitoring equipment and in the calculation of the flow rate proportion of each inflow source, and minimizes the negative impact on identifying inflow and infiltration units. Therefore, using the target flow rate proportion of each inflow source of the drainage unit determined by the method of this disclosure to identify inflow and infiltration units can improve the accuracy of identifying inflow and infiltration units.
[0050] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0051] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.
[0052] Figure 1 This is a flowchart illustrating the inflow and infiltration location identification method provided in an embodiment of the present disclosure.
[0053] Figure 2This is a schematic diagram of the structure of the inflow and infiltration location identification device provided in an embodiment of this disclosure.
[0054] Figure 3 This is a schematic diagram of the structure of an electronic device for identifying the location of inflow / infiltration, provided in an embodiment of this disclosure. Detailed Implementation
[0055] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0056] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0057] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0058] The core of the characteristic factor method for quantitatively determining the source of mixed water in drainage networks is the chemical mass balance model, which directly substitutes the mass concentration data of water quality characteristic factors of a certain type of mixed connection into the model in the form of mean values. However, the water quality at multiple mixed connection points of the same type of mixed source inevitably varies spatially. Errors in the mass concentration of water quality characteristic factors due to chemical sampling and measurement and the use of mean values, as well as the asynchronous nature of mass concentration measurement of mixed sources and mass concentration observation at the end discharge outlet, will inevitably lead to deviations in the calculation results.
[0059] In view of this, this disclosure proposes a method for identifying the inflow and infiltration location, which can reduce the above-mentioned deviations and improve the accuracy of determining the inflow and infiltration location.
[0060] Figure 1 This is a flowchart illustrating the inflow / infiltration location identification method provided in an embodiment of this disclosure. Figure 1 As shown, the method includes:
[0061] S11. Based on the water quality indicators of the water flowing into the drainage unit of the drainage network from multiple inflow sources, multiple inflow water quality indicator time series are obtained, and based on the water quality indicators of the water flowing out of the drainage unit, multiple outflow water quality indicator time series are obtained.
[0062] In this embodiment, the drainage network can be divided into multiple drainage units. For example, the drainage network can be divided into multiple drainage units according to its topological relationship. A drainage unit can be an area composed of multiple pipe segments in the drainage network, or it can be a main pipe in the drainage network. The sum of the flow rates of each drainage unit can be equal to the flow rate of the entire drainage network. Water quality indicators can include a variety of the following: chemical oxygen demand, conductivity, ammonia nitrogen, total hardness, temperature, pH value, turbidity, total organic carbon, five-day biochemical oxygen demand, total phosphorus, total nitrogen, suspended solids, total dissolved solids, petroleum hydrocarbons, anionic surfactants, cyanide, sulfides, fluorides, organophosphorus compounds, sulfates, mercury, chromium, cadmium, arsenic, lead, nickel, beryllium, silver, selenium, copper, zinc, manganese, iron, volatile phenols, benzene series compounds, aniline compounds, and nitrobenzene, etc.
[0063] The water quality indicators of the water flowing into the drainage unit of the drainage network are referred to as inflow water quality indicators in this application, and the water quality indicators of the water flowing out of the drainage unit are referred to as outflow water quality indicators in this application. It should be noted that the types of inflow water quality indicators and outflow water quality indicators are the same.
[0064] For a single drainage unit, water quality monitoring equipment can be used to collect monitoring values of multiple water quality indicators, thereby forming time series data of water quality indicators. For example, this can be done using water quality monitoring equipment that includes a spectral sensor. Monitoring equipment with a spectral sensor can achieve online, in-situ, high-frequency, real-time acquisition of water quality indicator monitoring values. For example, the acquisition frequency can be increased from once a day to once every 3-60 minutes, preferably 5-30 minutes, particularly preferably 8-20 minutes, and most preferably 10-15 minutes, which is far higher than the traditional testing method of sampling water and then conducting laboratory analysis. Therefore, water quality data can be obtained at a higher frequency to effectively capture the water quality characteristics of the drainage network.
[0065] In this embodiment of the disclosure, inflow water quality indicators can be selected for different inflow sources corresponding to the drainage pipe network, and a time series of inflow water quality indicators can be formed based on the high-frequency acquisition of the above-mentioned spectral sensor.
[0066] Collinearity analysis can be performed on at least some water quality index time series to obtain multiple selectable water quality indices with collinearity. Each selectable water quality index corresponding to each inflow source is then treated as multiple inflow water quality indices. In this way, this processing method can identify collinear water quality indices through collinearity analysis, thereby avoiding the participation of all water quality indices in subsequent calculations, saving computational resources and time, improving subsequent computational efficiency, and by selecting water quality indices that match each inflow source, it can more accurately monitor water quality problems related to specific inflow sources and improve the accuracy of inflow and infiltration location identification.
[0067] For example, time-series data of the 38 water quality indicators mentioned above can be collected using a spectral sensor and subjected to collinearity analysis to obtain three major categories of water quality indicators. Each category includes multiple selectable water quality indicators. Then, based on the type of each inflow source, one corresponding water quality indicator can be selected, ultimately forming multiple inflow source water quality indicators. For instance, if the inflow source types include groundwater, sewage, and river water, total hardness can be selected as the inflow water quality indicator for groundwater, ammonia nitrogen for sewage, and dissolved oxygen for river water.
[0068] The selected water quality indicators should possess two characteristics: specificity and stability. Specificity refers to water quality indicators that show significant differences among various inflow sources within the drainage network. Stability refers to water quality indicators that are relatively stable in water and undergo minimal physical, chemical, and biological reactions.
[0069] Inflow sources can be classified according to the type of water flowing into the drainage unit. For example, inflow sources can include: clean water inflow sources and sewage inflow sources.
[0070] For drainage units with simple topological relationships, such as those mainly consisting of main pipes, water quality monitoring can be performed at each inflow source corresponding to that single drainage unit. For each inflow source, the monitored values of various water quality indicators obtained at that single inflow source are arranged in time series to obtain multiple inflow water quality indicator time series. Furthermore, since the drainage unit has an outflow point, water quality monitoring can be performed at this outflow point. The monitored values of water quality indicators consistent with those at each inflow source are arranged in time series to obtain an outflow water quality time series. The types of water quality indicators monitored at each inflow source are the same as those monitored at the outflow point of the drainage unit.
[0071] For drainage units with complex topologies, such as those comprising main and branch pipes, for each inflow source, the monitored values of various water quality indicators corresponding to each inflow source are arranged in time series, resulting in multiple inflow water quality indicator time series. A single drainage unit typically includes multiple inflow points, with at least one outflow point downstream of each inflow point. Each inflow source flows into and out of the drainage unit. Therefore, water quality and flow rate can be monitored at both the inflow and outflow points. The types of water quality indicators monitored at the inflow and outflow points are the same as those at each inflow source. Then, the outflow water quality indicator time series is calculated based on the water quality monitoring time series data at the inflow and outflow points and the flow rate time series data obtained from flow rate monitoring. The specific calculation method is detailed below regarding the calculation of the ai value, and will not be repeated here.
[0072] S12, perform a first iterative process on the multiple inflow water quality index time series and the multiple outflow water quality index time series until the sum of the flow proportions of each inflow source at each time point meets a first preset condition, then stop the first iterative process to obtain the target flow proportion time series of each inflow source; the first iterative process includes: determining the flow proportion time series of each inflow source based on the multiple inflow water quality index time series and the multiple outflow water quality index time series; if the sum of the flow proportions of each inflow source at at least one time point does not meet the first preset condition, use a global optimization algorithm to perform a second iterative process on the multiple inflow water quality index time series and the multiple outflow water quality index time series until the second preset condition is met, then stop the second iterative process.
[0073] The inflow and outflow of water in drainage units should follow the principle of chemical mass balance. However, most existing technologies use chemical sampling methods for water quality data collection, which are inherently prone to errors. Because sampling is manual, it is easy to cause asynchrony in water quality concentration observations at multiple sampling points from the same inflow source, leading to errors in the monitored water quality indicators. Therefore, directly using a chemical mass balance model would reduce the accuracy of identifying inflow and infiltration units. Thus, we use water quality monitoring equipment to acquire water quality time-series data (i.e., inflow and outflow water quality indicator time series) in situ, online, and at high frequency. However, even with in-situ and high-frequency data acquisition, it is impossible to completely avoid spatial differences and asynchronous monitoring among inflow sources, all of which will cause errors in the collected water quality indicator monitoring values. Furthermore, the inflow and outflow water quality indicator time series contain monitoring values of multiple water quality indicators, which can easily amplify the negative effects of errors in the monitoring values, reducing the accuracy of identifying inflow and infiltration units. Moreover, some errors accumulate over time, which further reduces the accuracy of inflow infiltration unit judgment.
[0074] Therefore, combining the principle of chemical mass balance with the global optimization algorithm can improve the accuracy of determining the flow ratio of each inflow source, thereby improving the accuracy of identifying inflow into the infiltration unit.
[0075] Furthermore, in related technologies, the Monte Carlo method is commonly used to determine the water volume proportion of each inflow source. The Monte Carlo method requires generating random numbers based on the distribution of data in the time series, thus exhibiting strong randomness. Moreover, the water quality in this disclosure is susceptible to environmental influences (e.g., sudden discharge of high-concentration wastewater), which may cause sudden data fluctuations, resulting in a large span of distribution in the monitoring values of water quality indicators over time. Using the Monte Carlo method would further accentuate this randomness. Consequently, the accuracy and reliability of the determined flow proportion are poor. Compared to using the Monte Carlo method, this disclosure combines the chemical mass balance principle with a global optimization algorithm, which can improve the accuracy and reliability of determining the water volume proportion of each inflow source.
[0076] In this embodiment, the first iterative process includes two steps. Step one: Based on the principle of chemical mass balance, process multiple inflow water quality index time series and at least one outflow water quality index time series corresponding to each drainage unit to obtain the flow rate proportion time series of each inflow source for the drainage unit. Step two: Determine whether the flow rate proportion time series meets a first preset condition. If not, it indicates that the inflow and outflow water quality index time series are not accurate enough, and a global optimization algorithm is needed for processing until the flow rate proportion time series determined in step one meets the first preset condition, at which point the first iterative process stops. The determined flow rate proportion time series can be used as the target flow rate proportion time series.
[0077] To facilitate understanding of the process of determining the flow rate proportion time series using the principle of chemical mass balance, the following example is provided:
[0078] The time series of each inflow water quality index and each outflow water quality index of a single drainage unit can be input into the mass balance model to obtain the flow rate proportion of each inflow source of the drainage unit at each collection time. For ease of understanding, formula (1) is used to demonstrate a chemical mass balance model.
[0079] (1)
[0080] in, Characterizes the monitored value of water quality index i (e.g., ammonia nitrogen, hardness, etc.) in the j-th inflow source at a single moment; This represents the proportion of the flow rate of the j-th inflow source to the total flow rate of the drainage unit at a single moment, i.e., the flow rate percentage; n indicates that there are n inflow sources in the drainage unit. This represents the total value of water quality index i in the water flow exiting the drainage unit at a single moment. This total value can be either a monitored value or a calculated value of water quality index i.
[0081] For drainage units with simple topological relationships, such as those mainly consisting of main pipes, It can be the monitoring value of water quality indicators obtained at the outflow point at a single moment.
[0082] For drainage units with complex topological relationships, such as those comprising main pipes and branch pipes, The calculated value of the water quality index can be obtained by using the monitoring values of water quality indicators and flow rate obtained at each outflow point and at each inflow point at a single moment.
[0083] For ease of understanding, formula (2) is used to represent the process of determining the calculated values of water quality indicators.
[0084] (2)
[0085] in, This represents the flow rate at the h-th outflow point at a single time. This represents the flow rate at the c-th inflow point at a single time. This represents the monitored value of water quality index i at the h-th outflow point at a single time point. The value of water quality index i at the c-th inflow point is represented at a single moment, f represents the total number of outflow points in the drainage unit, and m represents the total number of inflow points in the drainage unit.
[0086] For each water quality index i, an equation can be determined using formula (1). Thus, for multiple water quality indexes i, multiple equations can be obtained. Then, by solving the equations simultaneously, the proportion of each inflow source to the total flow of the drainage unit at a single moment can be obtained, that is, the flow proportion of each inflow source. The flow proportion of a single inflow source in a single drainage unit is arranged in chronological order to obtain a flow proportion time series corresponding to that inflow source.
[0087] Because errors may occur during the collection and use of water quality indicator monitoring values, the flow rate percentage time series may not be accurate. To verify whether the flow rate percentage time series is accurate, it can be determined whether the sum of the flow rates of each flow rate percentage time series at the same time satisfies the first preset condition. For ease of understanding, formula (3) is used to represent the first preset condition.
[0088] (3)
[0089] If the conditions are met, it indicates that the accuracy of the flow percentage time series has reached the required level and can be used to identify inflow / infiltration units, and the identification accuracy meets the usage requirements. If the conditions are not met, it indicates that the flow percentage time series is not accurate enough. Here, the first preset condition can be that the sum of the flow percentages of each flow percentage time series at the same time is equal to 1, or falls within the first value range.
[0090] If the sum of the flow proportions of all inflow sources at at least one time point does not meet the first preset condition, a global optimization algorithm is used to perform a second iteration on the inflow and outflow water quality index time series. The input to the global optimization algorithm is the inflow and outflow water quality index time series. After each iteration, new inflow and outflow water quality index time series are obtained and used as the output of the current iteration. The output of the current iteration is used as the input of the next iteration. The second iteration stops when the second preset condition is met. The second preset condition can be that the difference between the output of the current iteration and the output of the previous iteration is less than a result difference threshold. Here, the difference can be the difference between the outputs of two iterations, or the growth / decrease ratio. Alternatively, the second preset condition can be that the current iteration number is greater than an iteration number threshold. In this way, after the second iteration, more accurate inflow and outflow water quality index time series can be obtained.
[0091] Global optimization algorithms are more objective and stable than existing Monte Carlo methods. Suitable global optimization algorithms include, but are not limited to, differential evolution, genetic algorithms, simulated annealing, gradient descent, and particle swarm optimization. Differential evolution is preferred because its superior performance stems from its effective utilization of parent individuals. By leveraging the differences between parent individuals, the algorithm can generate diverse offspring and grandchildren populations. Compared to Monte Carlo methods, this process not only ensures the objectivity of the algorithm during optimization but also makes it more targeted and stable in solving problems, thus demonstrating significant advantages in continuous optimization.
[0092] Specifically, differential evolution is a method that iteratively evolves individuals in a population, using mutation, crossover, and selection operations to find the global optimum. Compared to other methods, it has the advantages of fewer input parameters, faster convergence, and better robustness.
[0093] To facilitate understanding of the process of using the differential evolution algorithm in this disclosure, the following example is provided:
[0094] At each time point and To initialize the population, a second iteration is performed. A single round of the second iteration includes:
[0095] A. Initialization: Generate the initial population. Each individual in the population represents a candidate solution in the problem space, which is denoted by X for ease of description.
[0096] B. Mutation: The current individual undergoes a mutation operation. This step typically includes the following sub-steps:
[0097] a. Select 4 individuals from the population and perform a second vector difference to generate a difference vector.
[0098] b. Select the optimal individual and sum it with the two difference vectors to generate the mutated individual. For ease of understanding, the process of b is represented by formula (4).
[0099] (4)
[0100] Where F represents the scaling factor, and F can take any value between [0,2]. , , , This represents the current individual in the g-th iteration. Four different random individuals; Let be the globally optimal individual in the g-th iteration, which is also the basis vector in the g-th iteration; there are two difference vectors, namely . and .
[0101] Two difference vectors can make the perturbation stronger, the randomness better, and the global search capability better. , The difference is calculated, then scaled using F, and then... , The difference between the two vectors is calculated, and then scaled using F. The difference between the scaled vectors is added to the basis vectors to generate the mutated individual obtained after the g-th iteration. .
[0102] C. Crossover: In the g-th iteration, for the current individual ( ) and corresponding variant individuals ( Perform a crossover operation to generate new offspring individuals. For ease of understanding, formula (5) is used to represent the crossing process.
[0103] (5)
[0104] in, Denotes the offspring individuals after the crossover in the g-th iteration. The value in the q-th dimension. rand is a random number between 0 and 1, usually following a uniform distribution; CR is the crossover probability; qrand is a randomly generated integer between 1 and D, where D is the individual gene dimension, i.e., the vector dimension, which means that the value in the q-th dimension must come from V; Denotes the mutated individual in the g-th iteration. The value in the qth dimension; This represents the current individual in the g-th iteration. The value in the q-th dimension.
[0105] D. Selection: In the current individual ( ) and offspring individuals ( The selection process involves choosing between the current individual and its offspring. Both the current individual and its offspring are substituted into the objective function to obtain their respective objective function values. If the objective function value of the offspring is less than that of the current individual, the offspring is considered a candidate solution for the next generation. If the objective function value of the offspring is not less than that of the current individual, the current individual is considered a candidate solution for the next generation. For ease of understanding, formula (6) is used to represent the selection process.
[0106] (6)
[0107] in, This represents the individual obtained after the (g-1)th iteration.
[0108] E. Iteration: The candidate solution includes optimizations at each time step. and optimization .
[0109] The process of determining whether the second preset condition is met has already been described and will not be repeated here. The second iteration continues until the second preset condition is met.
[0110] S13, Based on the time series of the target flow rate ratio of each drainage unit, determine the inflow infiltration unit.
[0111] Drainage networks are typically categorized by their drainage system. Common types include separate sewage systems, separate stormwater systems, and combined sewer systems. Each type of network has one primary water flow. If the flow rate of other inflow sources is excessively high, exceeding a certain threshold, it indicates potential infiltration. For example, in a separate sewage system, sewage is the primary flow source, while clean water is a secondary source. The flow rate of clean water inflow sources within each drainage unit should be considered to ensure it does not exceed the threshold. Similarly, in a separate stormwater system, rainwater is the primary flow source, while clean water and sewage are secondary sources. The flow rate of clean water and sewage inflow sources within each drainage unit should be considered to ensure they do not exceed the threshold. Likewise, in a combined sewer system, sewage and rainwater are the primary flow sources, while clean water is a secondary source. The flow rate of clean water inflow sources within each drainage unit should be considered to ensure it does not exceed the threshold.
[0112] Therefore, inflow and infiltration units can be identified based on the time series of the target flow rate ratio of the drainage unit.
[0113] For example, the median or average of the time series of the target flow proportions of each drainage unit can be extracted as the flow proportion of each inflow source of the drainage unit. Then, based on the type of drainage network, it is determined whether the flow proportion of non-primary transport objects is greater than the flow proportion threshold. If it is greater, the drainage unit is identified as an inflow infiltration unit.
[0114] For example, the primary transport targets can be determined based on the type of drainage network. Then, based on the type of drainage network, the flow weight of each inflow source is determined. The time series of the target flow proportion of non-primary transport targets is weighted using these flow weights to obtain a weighted flow proportion time series. The median or mean of each weighted flow proportion time series can be extracted as the target flow proportion of non-primary transport targets for the drainage unit. It is then determined whether the target flow proportion of non-primary transport targets is greater than a flow threshold. If it is, the drainage unit is identified as an inflow infiltration unit.
[0115] The above are merely examples. The specific method for determining the inflow infiltration unit is not the focus of this disclosure. The embodiments of this disclosure do not limit the method for determining the inflow infiltration unit.
[0116] In this embodiment, the flow rate proportion time series of each inflow source is determined using inflow and outflow water quality index time series. This time series is then combined with a global optimization algorithm, and through a first iteration, the target flow rate proportion time series is finally obtained. This combination reduces errors in the sampling and measurement process of the monitoring equipment and in the calculation of the flow rate proportion of each inflow source, and minimizes the negative impact on identifying inflow and infiltration units. Therefore, using the target flow rate proportion of each inflow source of the drainage unit determined by the method of this disclosure to identify inflow and infiltration units can improve the accuracy of identifying inflow and infiltration units.
[0117] In one possible implementation, based on the first flow time series of a single inflow infiltration unit and the target flow percentage time series, the flow time series of each inflow source of a single inflow infiltration unit is determined; according to the flow time series of each inflow infiltration unit, the processing priority corresponding to each inflow infiltration unit is determined, and the processing priority is used to formulate an inflow infiltration treatment plan.
[0118] In this embodiment of the disclosure, flow measurement devices can be installed in the drainage unit to obtain a flow measurement time series. The flow measurement time series can characterize the flow rate of the drainage unit at each time point. For ease of description, the flow measurement time series corresponding to the inflow into the infiltration unit is named the first flow time series.
[0119] By multiplying the values in the first flow rate time series at the same time point by the values in the time series of each target flow rate percentage, the flow rate time series of each inflow source in the infiltration unit can be obtained. A single value in the flow rate time series can represent the flow rate of a single type of inflow source in the infiltration unit at a single time point.
[0120] In this embodiment of the disclosure, non-primary inflow sources can be identified based on the type of drainage network. An anomaly index characterizing the severity of inflow and infiltration is determined from the flow time series corresponding to the non-primary inflow sources. The treatment priority of the inflow and infiltration unit is determined based on the anomaly index. The anomaly index and priority can be positively correlated.
[0121] For example, a target time period subsequence can be extracted from the traffic time series corresponding to a non-primary transport object. The target time period can be a peak traffic period or several sampled periods. The average of the target time period subsequences is used as an indicator of the anomaly level corresponding to the non-primary transport object.
[0122] For example, the median of the traffic time series corresponding to a non-primary transport object can be used as an indicator of the degree of anomaly corresponding to that non-primary transport object.
[0123] The above are merely examples, and the methods for determining the degree of abnormality indicators in this disclosure are not limited.
[0124] In this embodiment, the inflow and outflow infiltration flow rate in the inflow and outflow infiltration unit can be determined, intuitively quantifying the severity of inflow and outflow infiltration and providing a reliable basis for formulating inflow and outflow treatment plans. For example, inflow and outflow infiltration units can be ranked according to anomaly indexes to determine the treatment order. As another example, the inflow and outflow water quality index time series of each drainage unit can be dynamically acquired over time to dynamically identify inflow and outflow infiltration units. Then, the treatment order is dynamically determined, and inflow and outflow infiltration units ranked higher than a preset position in the multiple determined treatment orders are given priority for treatment.
[0125] In one possible implementation, the method further includes: acquiring the initial inflow water quality index time series and the initial outflow water quality index time series of each drainage unit; identifying and removing abnormal operating condition data from the initial inflow water quality index time series and the initial outflow water quality index time series to obtain an operating condition screening time series; performing anomaly detection on the operating condition screening time series, removing abnormal data and adding normal data to obtain an anomaly screening time series, wherein the anomaly screening time series includes: a first anomaly screening time series and a second anomaly screening time series; aligning the first anomaly screening time series and the second anomaly screening time series within the same drainage unit in the time dimension to obtain the plurality of inflow water quality index time series and the plurality of outflow water quality index time series.
[0126] The initial inflow water quality index time series and the initial outflow water quality index time series can be obtained directly from the monitoring values of water quality indicators collected by water quality monitoring equipment arranged in time.
[0127] Abnormal operating conditions can include: equipment being removed from water, abnormal water quality data, equipment malfunction, abnormal operation, and environmental changes. Abnormal water quality data can be caused by water quality indicators exceeding the design range (generally, exceeding the design value by ±5%-10% is considered abnormal). Equipment malfunction can be caused by equipment failure, jamming, or other mechanical or control system abnormalities. Abnormal operation can be caused by human error such as misoperation. Environmental changes can be caused by external interference such as power fluctuations or extreme weather.
[0128] Abnormal operating condition data can be the monitored values of water quality indicators collected by water quality monitoring equipment under abnormal operating conditions. Abnormal operating condition data cannot reflect the true condition of the monitored water body. Therefore, it is necessary to detect and remove abnormal operating condition data. Then, correct data is added to obtain the operating condition screening time series. Specifically, after removing abnormal operating condition data from the initial inflow water quality indicator time series and adding correct data, a first operating condition screening time series can be obtained; after removing abnormal operating condition data from the initial outflow water quality indicator time series and adding correct data, a second operating condition screening time series can be obtained. In this embodiment of the disclosure, the operating condition screening time series includes: a first operating condition screening time series and / or a second operating condition screening time series.
[0129] Taking the abnormal operating condition of equipment leaving water as an example, the conductivity of the water quality monitoring equipment can be monitored according to the data acquisition frequency. If the monitored conductivity value of the water quality monitoring equipment is less than the conductivity threshold, it indicates that the water quality monitoring equipment has experienced water leaving, and the data collected during the period when water leaving occurred should be removed. Different conductivity thresholds can be set according to different weather conditions.
[0130] In another example, the liquid level of the water environment being tested can be monitored according to the data acquisition frequency. The installation elevation of the water quality monitoring equipment can be obtained in advance. The installation elevation is the distance between the installation location of the water quality monitoring equipment and a reference object (e.g., the bottom of a pipe or the ground) under the condition that the water quality monitoring equipment is in contact with the water body and can accurately complete water quality monitoring. If the liquid level is higher than the installation elevation, it indicates that the equipment has not left the water. If the liquid level is not higher than the installation elevation, it indicates that the equipment has left the water.
[0131] If the first iteration processes a water quality index time series with abnormal operating conditions, the accuracy of the results will be reduced, and it may be impossible to stop the first iteration, wasting computational resources. Therefore, removing abnormal operating condition data before the first iteration can increase the probability of obtaining the correct results and improve the accuracy of the results.
[0132] Anomaly detection can include: data jitter detection, baseline drift detection, abnormal jump detection, and abnormal noise detection. Abnormal data may increase the number of iterations in the first iteration process, prolonging computation time. Therefore, removing abnormal data from the operating condition screening time series and adding normal data can reduce the number of iterations in the first iteration process and improve computational efficiency. After adding normal data, the anomaly screening time series can be obtained. Specifically, removing abnormal data from the first operating condition screening time series and adding normal data yields the first anomaly screening time series; removing abnormal data from the second operating condition screening index time series and adding correct data yields the second anomaly screening time series.
[0133] After abnormal operating condition data and / or abnormal data identified through anomaly detection are removed, gaps appear at their corresponding time points. New values can be generated using estimation or simulation methods. For example, estimated values for water quality indicators can be calculated to fill these gaps using methods such as ARIMA prediction models, smoothing algorithms, and averaging data before and after the abnormal operating condition data or the abnormal data identified through anomaly detection.
[0134] Because spatial differences exist between inflow sources and between inflow sources and outflows within the same drainage unit, time differences arise between the first anomaly screening time series and between the first and second anomaly screening time series within the same drainage unit. This time difference weakens the correlation between water flows entering and leaving the same drainage unit, reducing the accuracy of identifying inflow / infiltration units. Therefore, aligning the first and second anomaly screening time series within the same drainage unit along the time dimension preserves the aforementioned correlation and improves the reliability of inflow / infiltration unit identification. A time alignment algorithm can be used to align the first and second anomaly screening time series within the same drainage unit along the time dimension. This time alignment algorithm includes, but is not limited to, one of the following: sine / cosine signal method, complex number representation method, or trigonometric function relationship method. The sine signal method is preferred. The specific method is as follows:
[0135] For example, the phase difference between the first and second anomaly screening time series in the same drainage unit can be determined. To eliminate the phase difference, each of the first and second anomaly screening time series is adjusted; the adjusted first anomaly screening time series is named the influent water quality index time series, and the adjusted second anomaly screening time series is named the effluent water quality index time series.
[0136] For ease of understanding, the process of determining the phase difference is illustrated using formula (7).
[0137] (7)
[0138] in, This represents the phase difference between two anomaly screening time series (which can be a first anomaly screening time series and a second anomaly screening time series). This represents the correlation coefficient between the two anomaly screening time series. , These represent the autocorrelation coefficients of the two anomaly screening time series, respectively.
[0139] For example, the phase difference can be determined by observation.
[0140] In this embodiment of the disclosure, the time difference between two anomaly screening time series (which can be a first anomaly screening time series and a second anomaly screening time series) can be determined based on the phase difference. The anomaly screening time series are then aligned using the time difference.
[0141] For ease of understanding, Equation (8) is used to demonstrate the process of determining the time difference using the phase difference.
[0142] (8)
[0143] in, Indicates time difference, This indicates the periodic frequency of change in the time series used for anomaly screening. For example, T can be equal to 24 hours.
[0144] In this embodiment, by identifying and replacing abnormal operating condition data and abnormal data identified through anomaly detection, the accuracy of the input data for the first iterative processing is improved, thereby increasing the accuracy of identifying inflow and infiltration units. Furthermore, the number of iterations in the first iterative processing is reduced, improving computational efficiency. Aligning the first and second anomaly screening time series within the same drainage unit reduces spatial asynchrony in the measurement data, preserves the correlation between measurement data, and further improves the accuracy of identifying inflow and infiltration units.
[0145] In one possible implementation, the anomaly detection includes: data jitter anomaly detection; the anomaly detection of the working condition screening time series includes: sliding a first time window on the working condition screening time series to determine multiple first indicators, the first indicators representing the differences between data within the first time window; if the first indicator is greater than a first difference threshold, the data for calculating the first indicator is determined as the abnormal data.
[0146] Data jitter anomaly detection can detect anomalous data that shows sudden increases or decreases in value. Data jitter anomalies typically refer to short-term, non-trend-based, rapid, and random fluctuations in data. "Jitter" can be a layer of rapid, irregular "noise" or "spikes" superimposed on an overall stable or regular data trend. These jitters are usually caused by measurement errors, random noise, or factors that cannot be explained in the short term.
[0147] The sliding step size of the first time window can be no greater than the length of the first time window, which can improve the precision of data jitter detection. For example, the sliding step size can be 1. This embodiment does not limit the length of the first time window. The first index can characterize the degree of difference in data size within the first time window. The first index can be the range, standard deviation, or first relative difference of the data within the first time window. Here, the first relative difference can be the ratio of the difference between other data points to a baseline value, using one data point in the first time window as the baseline. The first relative difference can more precisely and accurately reflect the degree of difference between the data points in the first window, improving the accuracy of data jitter detection.
[0148] The first time window slides multiple times on the working condition screening time series according to the time sequence, and each slide can determine at least one first indicator. For example, the length of the first time window is a multiple of 3, and the sliding step size of the first time window is equal to the length of the first time window. For each slide of the first time window, the data in the first time window are grouped into sets of three according to the time sequence. For each set, a first relative difference can be determined. For ease of understanding, formula (9) is used to describe the method for determining the first relative difference in this example. Taking a set as an example, the set contains three data Y1, Y2, and Y3. Among them, the time corresponding to Y1 is earlier than the time corresponding to Y2, the time corresponding to Y2 is earlier than the time corresponding to Y3, and Y2 is used as the baseline data.
[0149] (9)
[0150] Here, A represents the first relative difference.
[0151] Typically, the differences in the monitored values of the same water quality indicator in the monitored water environment at several adjacent times are not significant, for example, not exceeding a preset difference threshold. Therefore, the first indicator can be compared with the first difference threshold. If the first indicator is not greater than the first difference threshold, it indicates that the monitored value of the water quality indicator used to determine the first indicator has not experienced abnormal data fluctuation. If the first indicator is greater than the first difference threshold, and the monitored value of the water quality indicator used to determine the first indicator has experienced abnormal data fluctuation, then all monitored values of the water quality indicator used to determine the first indicator will be determined as abnormal data.
[0152] In this embodiment of the disclosure, a first time window can be slid across the working condition screening time series, and a first indicator can be determined using the data in the first time window. Based on the relationship between the first indicator and a first difference threshold, data jitter anomaly detection is performed, thereby identifying and removing abnormal data from the data in the working condition screening time series.
[0153] In one possible implementation, the anomaly detection includes: anomaly baseline detection. The anomaly detection of the operating condition screening time series includes: determining first selected data from the operating condition screening time series; determining a first sub-time series located before the corresponding time of the first selected data in the operating condition screening time series based on a second time window; determining a second sub-time series in the operating condition screening time series starting with the first selected data and its corresponding time based on a third time window; determining a second indicator characterizing the numerical change of the second sub-time series compared to the first sub-time series; determining a third indicator characterizing the trend of change of the first sub-time series, and a fourth indicator characterizing the trend of change of the second sub-time series; and determining the second sub-time series as the anomalous data when the second indicator falls within a first numerical range, the third indicator falls within a second numerical range, and the fourth indicator is greater than a first threshold. In another possible implementation, the anomaly detection includes: anomaly baseline detection. The anomaly detection of the operating condition screening time series includes: determining first selected data from the operating condition screening time series; determining a first sub-time series located before the corresponding time of the first selected data in the operating condition screening time series based on a second time window; determining a second sub-time series in the operating condition screening time series starting with the first selected data and its corresponding time based on a third time window; determining a second indicator characterizing the numerical change of the second sub-time series compared to the first sub-time series; determining a third indicator characterizing the changing trend of the first sub-time series and a fourth indicator characterizing the changing trend of the second sub-time series; and determining the first sub-time series and the second sub-time series as the anomalous data when the second indicator falls within a first numerical range, the third indicator is greater than a first threshold, and the fourth indicator is greater than a second threshold.
[0154] An abnormal baseline, also known as baseline drift or baseline shift, refers to a persistent, non-periodic, slow trend change in the overall baseline level of a time series dataset. Normally, water quality indicator time series data remain stable around a baseline or fluctuate within a certain range. If some data in a water quality indicator time series no longer remain stable around the original baseline, and the entire baseline level exhibits a persistent, non-periodic, slow trend change, then this data is considered abnormal, and abnormal baseline detection can detect these abnormal data.
[0155] In this embodiment, the first selected data can be any data in the working condition filtering time series. For example, the first selected data can be data located at the midpoint or data located at a third point. The lengths of the second and third time windows can be the same or different. This embodiment does not limit the lengths of the second and third time windows. Data in the second time window can be used as the first sub-time series, and data in the third time window can be used as the second sub-time series. The first and second sub-time series are adjacent.
[0156] The third and fourth indicators can be of the same type. The third and fourth indicators can be the average slope of adjacent data in the corresponding time series, or the range of the slope of adjacent data in the corresponding time series, or the median slope of the corresponding time series, or the third indicator can be the Tilson slope of the first sub-time series and the fourth indicator can be the Tilson slope of the second time series. The above are just examples and the embodiments disclosed herein do not limit this.
[0157] The second indicator can be calculated from specific data in the first and second sub-time series. For ease of description, the specific data in the first sub-time series is named the first target data, and the specific data in the second sub-time series is named the second target data. The first target data and the second target data can be the median of the first and second sub-time series, or data at their respective third quartiles. This disclosure does not limit this.
[0158] The second indicator can be the difference between the first target data and the second target data, or a ratio, etc. This disclosure does not limit it.
[0159] The second indicator represents the magnitude of the increase in the second sub-time series relative to the first sub-time series. If the second indicator falls within the first numerical range, it indicates a significant increase in the second sub-time series compared to the first sub-time series. The lower limit of the first numerical range can be selected from (0, 0.2), and the upper limit can be selected from (0.2, 1). The second numerical range can be (0, 0.1).
[0160] A third indicator exceeding the first threshold indicates an upward trend in the first sub-time series; a fourth indicator exceeding the second threshold indicates an upward trend in the second sub-time series. This suggests that baseline anomalies have occurred in the data of both the first and second sub-time series. The first and second thresholds can be the same. For example, the first and second thresholds can be set to values of 0.2, 0.3, or 0.5, etc.
[0161] Furthermore, the second numerical range can be (0, 0.2). If the third indicator falls within this range, it indicates that the first sub-time series is relatively stable around the baseline and shows no significant trend. If the fourth indicator is greater than the second threshold, it indicates that the second sub-time series is showing an upward trend and is no longer relatively stable around the baseline. This suggests a baseline anomaly has occurred in the data from the second sub-time series. Although the second indicator can characterize the numerical relationship between the two sub-time series, it cannot indicate the internal trends of the first and second sub-time series themselves. The second indicator alone is insufficient to determine whether the data in the working condition screening time series is changing around the baseline. Therefore, it is necessary to determine the third and fourth indicators, and to assess the relationship between the third indicator and the first threshold, and the fourth indicator and the second threshold; or to assess whether the third indicator falls within the second numerical range, and the relationship between the fourth indicator and the first threshold. This improves the accuracy of the overall working condition screening time series' trend and enhances the accuracy of identifying anomalous data.
[0162] For example, each acquisition time in the operating condition screening time series can be used as a potential drift time. The data corresponding to each potential drift time is used as the first selected data.
[0163] The system iterates through the potential drift moments in chronological order to determine the corresponding second, third, and fourth indicators. If a single potential drift moment's corresponding second indicator falls within a first value range, the second indicator falls within a second value range, and the third indicator is greater than a first threshold, then that single potential drift moment is designated as the start moment. Potential drift moments chronologically later than this single potential drift moment are designated as remaining potential drift moments. The system then iterates through the remaining potential drift moments in chronological order to determine the corresponding second, third, and fourth indicators. If a single remaining potential drift moment's corresponding second indicator falls within a first value range, the third indicator is greater than the first threshold, and the fourth indicator is less than the first threshold, then that single remaining potential drift moment is designated as the end moment. Data collected at the start moment, data collected at the end moment, and data collected between the start and end moments are designated as abnormal data. This refines the determination of the range of abnormal data.
[0164] In one possible implementation, determining the second indicator characterizing the numerical change of the second sub-time series compared to the first sub-time series includes: determining a first target data at a first quantile from the first sub-time series, determining a second target data at a second quantile from the second sub-time series, and determining the second indicator based on the first target data and the second target data.
[0165] In this embodiment of the disclosure, a second relative difference (the ratio of the difference between the second target data and the first target data to the benchmark value) can be determined by using the first target data as a benchmark value or by using the average of the first target data and the second target data as a benchmark value. This second relative difference is then used as a second indicator.
[0166] In this way, the second indicator can be determined using only two data points, requiring less data acquisition and computation, and resulting in high computational efficiency.
[0167] The first and second quantiles are preset quantiles. For example, the first quantile is greater than the second quantile. For instance, the first quantile is the 90th quantile, and the second quantile is the 10th quantile.
[0168] In this embodiment of the disclosure, the data in the sub-time series can be arranged in order of size. The larger the quantile, the closer the value of the data corresponding to the quantile is to the maximum value of the sub-time series, and it belongs to the larger value of the sub-time series. Conversely, the larger the quantile, the closer the value of the data corresponding to the quantile is to the minimum value of the sub-time series, and it belongs to the smaller value of the sub-time series.
[0169] Therefore, if the first quantile is greater than the second quantile, and the second indicator falls within the first threshold range, it means that comparing the smaller numerical data in the second sub-time series with the larger numerical data in the first sub-time series both meet the criteria for identifying outlier data. Similarly, using other data from the first and second sub-time series to determine the second indicator would also meet the criteria for identifying outlier data. Therefore, the second indicator calculated from the selected first and second target data is more convincing, improving the accuracy of the determined numerical changes in the second sub-time series compared to the first sub-time series.
[0170] In this embodiment of the disclosure, the first quantile and the second quantile can indicate the numerical ranking of the data in the first sub-time series and the second sub-time series, making the selected first target data and second target data more targeted, which helps to improve the persuasiveness of the second indicator and the accuracy of the numerical changes of the two sub-series represented by the second indicator.
[0171] In one possible implementation, the anomaly detection includes: anomaly jump detection. The anomaly detection of the working condition screening time series includes: identifying multiple first data points prior to a first acquisition time in the working condition screening time series; grouping the multiple first data points according to a first duration to obtain multiple first data groups; determining the median of each first data group; if the growth ratio of the second data points acquired at the first acquisition time relative to each of the medians is greater than a ratio threshold, acquiring each third data point acquired after the first acquisition time and within an adjacent second duration; if the growth ratio of each of the third data points relative to each of the medians is greater than the ratio threshold, identifying the second data points and each of the third data points as the anomaly data.
[0172] Anomaly detection can identify anomalous data that differs significantly from normal data in value and is adjacent to normal data. These anomalies can be multiple data points adjacent to normal data, including those whose numerical differences from normal data exceed a threshold.
[0173] The first data collection time can be the moment when abnormal data needs to be detected. For example, during the process of collecting water quality indicator monitoring values, abnormal data can be detected every three preset time intervals. That is, every three time intervals marks a first data collection time. The first data collection time can be the current data collection time.
[0174] The first data point can be the monitoring value of the water quality indicator collected before the current collection time. For example, the first data point can be the monitoring value of the water quality indicator collected within one or more third time periods closest to the current collection time and before the current collection time.
[0175] The first data point can correspond to a specific moment in time; therefore, the multiple first data points in this embodiment are sequential. These multiple first data points can be arranged sequentially to obtain a first data time series. Starting from the moment corresponding to the first first data point in the first data time series, a first time period can be obtained after each first duration. A single first data group can be the first data point corresponding to the moments contained within a single first time period. For each first data group, the median of the first data group can be calculated.
[0176] At the current data collection moment, water quality indicator monitoring values can be collected. For ease of description, these monitoring values are named the second data. The growth rate of the second data relative to each median is calculated. A growth rate threshold can be preset. If the growth rate of each second data point relative to each median exceeds the threshold, it indicates that the second data point may have experienced an abnormal jump, requiring further verification. Therefore, the third data point can be collected. Here, the difference between the second data point and a single median can be determined, and the ratio of this difference to the second data point is taken as a growth rate.
[0177] The third data point consists of the water quality indicator monitoring values collected within the second time interval following the current collection time. Then, the growth rate of each third data point relative to each median can be calculated. If the growth rate of each third data point relative to each median within each second time interval exceeds a growth rate threshold, it indicates an abnormal jump in the water quality indicator monitoring values (i.e., each third data point) within the second time interval. Here, the difference between the third data point and a single median can be determined, and the ratio of this difference to the third data point is taken as a growth rate.
[0178] Therefore, the monitoring values of the second data, the third data, and the water quality indicators between the second and third data were identified as abnormal data.
[0179] The method used in the embodiments of this disclosure for abnormal transition detection requires little computation, is easy to implement, and has high computational efficiency.
[0180] In one possible implementation, the anomaly detection includes: anomaly noise detection, wherein the anomaly detection of the working condition screening time series includes: determining the difference between each data in the working condition screening time series and its adjacent data; and determining the data corresponding to the difference greater than a second difference threshold as the anomaly noise.
[0181] In this embodiment of the disclosure, the difference between each data point in the working condition screening time series and one or two adjacent data points can be determined. When a single data point has one adjacent data point, the difference between that single data point and the adjacent data point is determined, resulting in one difference. If this difference is greater than a second difference threshold, the single data point is identified as an anomalous data point. When a single data point has two adjacent data points, the differences between that single data point and each of the two adjacent data points are determined separately, resulting in two differences. If both of these differences are greater than the second difference threshold, the single data point is identified as an anomalous data point.
[0182] In the embodiments disclosed herein, identifying abnormal noise points is simple, easy, and computationally efficient.
[0183] Figure 2A schematic diagram of the inflow / infiltration location identification device provided in an embodiment of this disclosure. The device 20 includes:
[0184] The water quality index time series acquisition unit 21 is used to obtain multiple inflow water quality index time series based on the water quality index of water flowing into the drainage unit in the drainage network from multiple inflow sources, and to obtain multiple outflow water quality index time series based on the water quality index of water flowing out of the drainage unit.
[0185] The target flow percentage time series acquisition unit 22 is used to perform a first iterative process on the multiple inflow water quality index time series and the multiple outflow water quality index time series until the sum of the flow percentages of each inflow source at each time point meets a first preset condition, and then stop the first iterative process to obtain the target flow percentage time series of each inflow source.
[0186] Inflow and infiltration determination unit 23 is used to determine the inflow and infiltration unit based on the target flow rate ratio time series of each drainage unit;
[0187] The first iterative process includes:
[0188] Based on the time series of the multiple inflow water quality indicators and the multiple outflow water quality indicators, determine the flow rate proportion time series of each inflow source;
[0189] If the sum of the flow proportions of each of the inflow sources at at least one moment does not meet the first preset condition, a global optimization algorithm is used to perform a second iteration on the time series of the multiple inflow water quality indicators and the time series of the multiple outflow water quality indicators until the second preset condition is met, and then the second iteration is stopped.
[0190] In one possible implementation, the device 20 further includes:
[0191] A flow time series determination unit is used to determine the flow time series of each inflow source of a single inflow infiltration unit based on the first flow time series of a single inflow infiltration unit and the target flow proportion time series.
[0192] The processing priority determination unit is used to determine the processing priority corresponding to each of the inflow infiltration units based on the flow time series of each of the inflow infiltration units. The processing priority is used to formulate an inflow infiltration treatment plan.
[0193] In one possible implementation, the device 20 further includes:
[0194] The initial water quality index time series acquisition unit is used to acquire the initial inflow water quality index time series and the initial outflow water quality index time series of each drainage unit.
[0195] The operating condition screening time series determination unit is used to identify and remove abnormal operating condition data from the initial inflow water quality index time series and the initial outflow water quality index time series to obtain the operating condition screening time series.
[0196] An anomaly screening time series determination unit is used to perform anomaly detection on the working condition screening time series, remove abnormal data from the working condition screening time series and fill in normal data to obtain an anomaly screening time series, wherein the anomaly screening time series includes: a first anomaly screening time series and a second anomaly screening time series.
[0197] The inflow and outflow water quality index time series determination unit is used to align the first anomaly screening time series and the second anomaly screening time series within the same drainage unit in the time dimension to obtain the multiple inflow water quality index time series and the multiple outflow water quality index time series.
[0198] In one possible implementation, the anomaly detection includes: data jitter anomaly detection, and the anomaly filtering time series determination unit is further configured to:
[0199] Slide a first time window on the working condition screening time series to determine multiple first indicators, where the first indicators characterize the differences between data within the first time window;
[0200] If the first indicator is greater than the first difference threshold, the data used to calculate the first indicator will be identified as the abnormal data.
[0201] In one possible implementation, the anomaly detection includes: anomaly baseline detection, and the anomaly screening time series determination unit is further configured to:
[0202] The first selected data is determined from the time series of the operating conditions.
[0203] Based on the second time window, a first sub-time series located before the time corresponding to the first selected data is determined in the working condition filtering time series;
[0204] Based on the third time window, in the working condition screening time series, a second sub-time series is determined with the first selected data and the corresponding time as the starting point;
[0205] Determine a second index characterizing the numerical change of the second sub-time series compared to the first sub-time series;
[0206] A third indicator characterizing the trend of change in the first sub-time series and a fourth indicator characterizing the trend of change in the second sub-time series are determined.
[0207] If the second indicator falls within the first value range, the third indicator falls within the second value range, and the fourth indicator is greater than the first threshold, the second sub-time series is identified as the abnormal data.
[0208] In one possible implementation, the anomaly screening time series determination unit is further configured to:
[0209] The first target data is determined based on the first quantile of the first sub-time series, and the second target data is determined based on the second quantile of the second sub-time series.
[0210] The second indicator is determined based on the first target data and the second target data.
[0211] In one possible implementation, the anomaly detection includes: anomaly jump detection, and the anomaly screening time series determination unit is further configured to:
[0212] Determine multiple first data points prior to the first acquisition time in the working condition screening time series;
[0213] The plurality of first data are grouped according to a first duration to obtain a plurality of first data groups;
[0214] Determine the median of each of the first data groups;
[0215] If the growth rate of the second data collected at the first acquisition time relative to each of the medians is greater than the ratio threshold, then the third data collected within the second time period after the first acquisition time and adjacent to it are acquired.
[0216] If the growth rate of each of the third data points relative to each of the medians is greater than the ratio threshold, then the second data point and each of the third data points are identified as the abnormal data.
[0217] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0218] This disclosure also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the above-described method. The computer-readable storage medium can be volatile or non-volatile.
[0219] This disclosure also proposes an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0220] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.
[0221] Figure 3 This is a schematic diagram of an electronic device for identifying inflow / infiltration locations, provided as an embodiment of this disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. (Refer to...) Figure 3 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0222] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). Electronic device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM Mac OS X TM Unix TM Linux TM FreeBSD TM Or similar.
[0223] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.
[0224] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.
[0225] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0226] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0227] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0228] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0229] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0230] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0231] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0232] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for identifying the location of inflow / infiltration, characterized in that, include: Based on the water quality indicators of water flowing into the drainage unit of the drainage network from multiple inflow sources, multiple inflow water quality indicator time series are obtained; and based on the water quality indicators of water flowing out of the drainage unit, multiple outflow water quality indicator time series are obtained. The first iteration process is performed on the multiple inflow water quality index time series and the multiple outflow water quality index time series until the sum of the flow proportions of each inflow source at each time point meets the first preset condition, and the first iteration process is stopped to obtain the target flow proportion time series of each inflow source. Based on the time series of the target flow rate proportion of each drainage unit, the inflow infiltration unit is determined; The first iterative process includes: Based on the time series of the multiple inflow water quality indicators and the multiple outflow water quality indicators, determine the flow rate proportion time series of each inflow source; If the sum of the flow proportions of each of the inflow sources at at least one moment does not meet the first preset condition, a global optimization algorithm is used to perform a second iteration on the time series of the multiple inflow water quality indicators and the time series of the multiple outflow water quality indicators until the second preset condition is met, and then the second iteration is stopped.
2. The method according to claim 1, characterized in that, The method further includes: Based on the first flow time series of a single inflow infiltration unit and the target flow percentage time series, the flow time series of each inflow source of a single inflow infiltration unit is determined; Based on the flow time series of each inflow infiltration unit, the processing priority corresponding to each inflow infiltration unit is determined, and the processing priority is used to formulate an inflow infiltration treatment plan.
3. The method according to claim 1, characterized in that, The method further includes: Obtain the initial inflow water quality index time series and the initial outflow water quality index time series for each drainage unit; Identify and remove abnormal operating condition data from the initial inflow water quality index time series and the initial outflow water quality index time series to obtain the operating condition screening time series; Anomaly detection is performed on the working condition screening time series, abnormal data is removed from the working condition screening time series and normal data is added to obtain an anomaly screening time series, which includes: a first anomaly screening time series and a second anomaly screening time series. Align the first anomaly screening time series and the second anomaly screening time series within the same drainage unit in the time dimension to obtain the multiple inflow water quality index time series and the multiple outflow water quality index time series.
4. The method according to claim 3, characterized in that, The anomaly detection includes: data jitter anomaly detection; the anomaly detection of the working condition screening time series includes: Slide a first time window on the working condition screening time series to determine multiple first indicators, where the first indicators characterize the differences between data within the first time window; If the first indicator is greater than the first difference threshold, the data used to calculate the first indicator will be identified as the abnormal data.
5. The method according to claim 3, characterized in that, The anomaly detection includes: anomaly baseline detection; the anomaly detection of the operating condition screening time series includes: The first selected data is determined from the time series of the operating conditions. Based on the second time window, a first sub-time series located before the time corresponding to the first selected data is determined in the working condition filtering time series; Based on the third time window, in the working condition screening time series, a second sub-time series is determined with the first selected data and the corresponding time as the starting point; Determine a second index characterizing the numerical change of the second sub-time series compared to the first sub-time series; A third indicator characterizing the trend of change in the first sub-time series and a fourth indicator characterizing the trend of change in the second sub-time series are determined. If the second indicator falls within the first value range, the third indicator falls within the second value range, and the fourth indicator is greater than the first threshold, the second sub-time series is identified as the abnormal data.
6. The method according to claim 5, characterized in that, The determination of the second indicator characterizing the numerical change of the second sub-time series compared to the first sub-time series includes: The first target data is determined based on the first quantile of the first sub-time series, and the second target data is determined based on the second quantile of the second sub-time series. The second indicator is determined based on the first target data and the second target data.
7. The method according to claim 3, characterized in that, The anomaly detection includes: anomaly jump detection; the anomaly detection of the time series of the selected operating conditions includes: Determine multiple first data points prior to the first acquisition time in the working condition screening time series; The plurality of first data are grouped according to a first duration to obtain a plurality of first data groups; Determine the median of each of the first data groups; If the growth rate of the second data acquired at the first acquisition time relative to each of the medians is greater than the ratio threshold, then the third data acquired within the second time period after the first acquisition time and adjacent to it are obtained. If the growth rate of each of the third data points relative to each of the medians is greater than the ratio threshold, then the second data point and each of the third data points are identified as the abnormal data.
8. An inflow / infiltration location identification device, characterized in that, include: The water quality index time series acquisition unit is used to obtain multiple inflow water quality index time series based on the water quality index of water flowing into the drainage unit in the drainage network from multiple inflow sources, and to obtain multiple outflow water quality index time series based on the water quality index of water flowing out of the drainage unit. The target flow percentage time series acquisition unit is used to perform a first iterative process on the multiple inflow water quality index time series and the multiple outflow water quality index time series until the sum of the flow percentages of each inflow source at each time point meets a first preset condition, and then stop the first iterative process to obtain the target flow percentage time series of each inflow source. An inflow and infiltration determination unit is used to determine the inflow and infiltration unit based on the time series of the target flow rate ratio of each drainage unit; The first iterative process includes: Based on the time series of the multiple inflow water quality indicators and the multiple outflow water quality indicators, determine the flow rate proportion time series of each inflow source; If the sum of the flow proportions of each of the inflow sources at at least one moment does not meet the first preset condition, a global optimization algorithm is used to perform a second iteration on the time series of the multiple inflow water quality indicators and the time series of the multiple outflow water quality indicators until the second preset condition is met, and then the second iteration is stopped.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the method of any one of claims 1 to 7 when executing instructions stored in the memory.
10. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.