Method for processing measurements acquired in water installations, and associated computer program and processing system

Deep learning models effectively address the challenge of missing data in water treatment time series by analyzing and predicting missing data points, enhancing precision in water treatment plant operations.

WO2026017714A1PCT designated stage Publication Date: 2026-01-22SUEZ INTERNATIONAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/070276
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-16
Filing Date
2025-07-16
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing methods for handling missing data in time series from water installations, such as water and wastewater treatment plants, lack accuracy and are not suitable for all cases, particularly when data is missing or likely to skew results.

Method used

A method utilizing deep learning models, specifically neural networks, to analyze time series data, identify outliers, and predict missing data points by training on existing data, excluding outliers, and completing the series with predicted values.

Benefits of technology

Achieves exceptional accuracy in determining missing data points, improving prediction precision compared to traditional interpolation and extrapolation methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025070276_22012026_PF_FP_ABST
    Figure EP2025070276_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for processing measurements acquired in water installations, the measurements forming at least one timeseries, the method comprising the following steps: - analysing (120) the timeseries in order to determine abnormal data and / or missing-datum instants corresponding to instants at which at least one datum is missing; - when at least one abnormal datum is determined, excluding (130) this datum from the timeseries, the corresponding instant then being considered to be a missing-datum instant; - training (140) a deep learning model on at least some of the data of the timeseries; - determining (160) the data missing at the missing-datum instants using the deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TITLE: Process for processing measurements acquired in water installations, associated computer program and processing system

[0002] The present invention relates to a method for processing measurements acquired in water installations. The present invention also relates to a computer program and a processing system associated with this processing method.

[0003] The present invention relates more particularly to the field of water treatment implemented by various water installations such as a water treatment plant, a water distribution plant, a wastewater treatment plant, etc.

[0004] In order to size these various installations and monitor their operation, it is often necessary to have data characterizing the water within these installations, such as water consumption, flow rate at a given point in the system, chemical composition, etc. This data can then be used to determine, for example, water quality indicators to meet current standards, to alert consumers, or to carry out maintenance work when necessary.

[0005] The corresponding data typically comes from sensors operating at predetermined measurement frequencies and located in known locations. Furthermore, each data point is usually associated with a known time, which is also relevant for data analysis. For example, it is often necessary not only to know the magnitude of a given parameter but also the time at which that parameter was equal to that magnitude.

[0006] The data generated by the sensors and associated with known times thus form time series. Here, each time can correspond to any time reference and can, for example, be defined by the exact time the measurement was taken or by a period such as an hour, a day, a week, a month, etc. The known times can, for example, be determined according to a known frequency (for example, every hour).

[0007] In some cases, time series data may be missing. A data point may be missing for a given time due to various events that can occur in the data acquisition and transmission chain. For example, such an event could be a sensor, transmission, or storage device failure. Data may also be considered missing when the determined measurement frequency is insufficient to use the corresponding measurements for a desired purpose. To address the problem of missing data in time series, several methods are already known in the state of the art.

[0008] Most of these methods rely on the use of business formulas that allow for the deduction of certain data types from other data types by leveraging known relationships between them. However, these relationships are often derived empirically and cannot be applied to all cases with the desired accuracy.

[0009] Methods involving interpolation and / or extrapolation of known data are also known, but these methods may also lack the desired accuracy.

[0010] Furthermore, state-of-the-art methods are not always suitable for dealing with data considered not missing but likely to skew the results.

[0011] The present invention aims to remedy these drawbacks and therefore to propose a method and a processing system enabling the processing of time series relating to measurements acquired in water installations in a particularly precise manner.

[0012] To this end, the invention relates to a method for processing measurements acquired in water installations, said measurements forming at least one time series, the time series presenting a set of instants and a set of data associated with these instants, the method comprising the following steps:

[0013] - analysis of the time series to determine outliers and / or missing data points corresponding to times in which at least one data point is missing;

[0014] - when at least one outlier is determined, exclusion of that data from the time series, the corresponding instant then being considered as an instant of missing data;

[0015] - training a deep learning model from at least some of the time series data;

[0016] - Determining missing data at missing data times using the deep learning model.

[0017] Thanks to these features, the invention enables the determination of missing data with exceptional accuracy. Indeed, the deep learning model allows for the identification of specific data behaviors and / or correlations between these data points. This, in turn, enables the prediction of missing data with greater precision. The invention is therefore particularly advantageous compared to state-of-the-art methods that rely solely on business rules and / or data interpolation / extrapolation.

[0018] According to other advantageous aspects of the invention, the method comprises one or more of the following features, taken individually or in all technically possible combinations:

[0019] - the determination of each outlier is carried out by applying business rules;

[0020] - the determination of each outlier includes comparing the corresponding data with at least one threshold and / or calculating a Mahalanobis distance and / or verifying at least one relationship and / or implementing an unsupervised machine learning model;

[0021] - each missing data point is due to a sensor failure or a transmission fault;

[0022] - training the deep learning model includes determining the behavior of the data over time and / or determining correlations between the data;

[0023] - the training of the deep learning model is carried out using a portion of the time series data;

[0024] - the process further comprising a step of testing the deep learning model using time series data not used for its training;

[0025] - each measurement corresponds to a parameter of the water passing through the water facilities;

[0026] - the process further comprising a step of completing the time series with the missing data thus determined and / or a step of transmitting the completed time series to any interested system.

[0027] The invention also relates to a computer program comprising software instructions which, when executed by a computer, implement the process as defined above.

[0028] The invention also relates to a system for processing measurements acquired in water installations, comprising technical means configured to implement the process as defined above.

[0029] The invention will become clearer upon reading the following description, given solely by way of non-limiting example, and made with reference to the drawings in which: - [Fig. 1] Figure 1 is a schematic view of a processing system according to the invention; and

[0030] - [Fig. 2] Figure 2 is a flowchart of a treatment process according to the invention, the treatment process being implemented by the treatment system of Figure 1.

[0031] Figure 1 illustrates a processing system 10 for processing measurements acquired in water installations 12.

[0032] Each water installation 12 includes, for example, a water treatment plant, a water distribution plant, a wastewater treatment plant, etc. Such a plant is specifically designed to treat or distribute water using methods known per se. Throughout this text, "water" refers to clean water intended for use by consumers (e.g., drinking water) or wastewater.

[0033] The 12 water facilities can be located in different geographical locations.

[0034] Each water installation 12 includes one or more sensors 14 capable of acquiring measurements relating to at least one parameter characterizing the water in that installation. Such a parameter may relate to the chemical composition of the water or to the manner in which it flows in the corresponding water installation 12, such as its flow rate, pressure, etc. For example, a sensor 14 may measure the ammonium (NH4) content or suspended solids (SS) in wastewater or the pH in clean water.

[0035] The 14 sensors can be of the same nature (i.e., capable of measuring the same parameters) or of different natures (i.e., capable of measuring different parameters).

[0036] Each sensor 14 is capable of acquiring the corresponding measurements and generating data, for example, numerical data, corresponding to these measurements. Each sensor 14 is also capable of transmitting each generated data to the processing system 10 via any known means of communication, such as a wired or wireless computer network. In addition, or alternatively, each sensor 14 is also capable of transmitting each generated data to a centralized storage system, such as a SCADA (Supervisory Control and Data Acquisition) system.

[0037] Advantageously, each data point generated by a sensor 14 is associated by that same sensor with a specific time. Such a time corresponds, for example, to a precise time and date when the measurement was taken, or to a period during which the measurement was taken, such as an hour, a day, a week, a month, etc. Alternatively, each data point generated by a sensor 14 is associated with a time by any other means, for example, by any other means of the installation 12 or by the processing system 10.

[0038] In some cases, each sensor 14 is capable of acquiring the corresponding measurements at a predetermined frequency. This frequency can then be used to associate times with the corresponding data.

[0039] The data generated by the 14 sensors, along with the time points associated with those data, form one or more time series. In particular, each time series presents a set of time points and a set of data associated with those time points.

[0040] For example, several data points can be associated with the same time point. In this case, it is called vector data. Furthermore, a time point may not be associated with any data. This is then called missing data. Such a time point is subsequently referred to as a missing data point.

[0041] The processing system 10 is capable of processing the data generated by the sensors 14, as will be explained in more detail later. The processing system 10 is, for example, located away from any water installation 12 or is part of at least one of these water installations 12.

[0042] With reference to Figure 1, the processing system 10 comprises an input module 21, a processing module 22 and an output module 23.

[0043] The input module 21 is connected directly or indirectly to each of the sensors 14 (possibly via a centralized storage system) using a communication means as defined above and capable of acquiring the data generated by these sensors 14. Advantageously, this data is acquired in the form of one or more time series, i.e., with the corresponding instants. Alternatively, the input module 21 is capable of associating the corresponding instant with each acquired data point to form one or more time series. Such an instant corresponds, for example, to the instant of reception of the corresponding data point.

[0044] In some embodiments, the input module 21 includes a storage unit for storing data for a specified period.

[0045] The processing module 22 is capable of processing the data acquired by the input module 21. In particular, the processing module 22 is capable of processing this data in the form of at least one time series in order to impute the missing data to it, as will be explained later.

[0046] The output module 23 is capable of transmitting the data processed by the processing module 22 to any interested external system. Such an external system allows, for example, the implementation of processing other than that performed by system 10, for example, processing for a specific task, such as determining a water quality indicator, the total water consumption for the installations 12, etc.

[0047] Each of the modules 21 to 23 is presented, for example, in the form of one or more software programs, that is, in the form of a computer program, also called a computer program product. The software program(s) is implemented by one or more processors and random access memory (RAM), which are also part of the processing system 10. The software program(s) is, for example, stored on a computer-readable medium (not shown). A computer-readable medium is, for example, a medium capable of storing electronic instructions and being connected to a bus of a computer system. For example, a readable medium is an optical disc, a magneto-optical disc, ROM, RAM, any type of non-volatile memory (e.g., FLASH or NVRAM), or a magnetic card. A computer program comprising software instructions is then stored on the readable medium.

[0048] Alternatively or in addition, at least one of the modules 21 to 23 is presented at least partially in the form of a programmable logic circuit such as an FPGA (Field Programmable Gate Array), or an integrated circuit, such as an ASIC (Application Specific Integrated Circuit).

[0049] The processing system 10 is capable of implementing a processing method according to the invention, which will now be explained with reference to Figure 2, which shows a flowchart of its steps.

[0050] Initially, it is assumed that the sensors 14 acquire measurements and generate data corresponding to these measurements which are then transmitted to the processing system 10, possibly with the corresponding times.

[0051] In an initial step 110, the input module 21 of the processing system 10 acquires this data along with the corresponding time points or associates time points with the acquired data, as explained previously. In other words, the input module 21 acquires or forms at least one time series during this step 110.

[0052] In some embodiments, the time series is acquired or generated by the input module 21 after a predetermined time period. In other words, during this time period, the input module 21 can receive data from different sensors 14 at different frequencies and generate the corresponding time series only after this period has elapsed. To store the intermediate data, the input module 21 can use its storage unit. This time period can correspond to an hour, a day, a week, a month, a year, etc.

[0053] The total number of data in the formed time series can be greater than 1000, advantageously greater than 10,000. At the end of step 110, the input module 21 transmits the corresponding time series to the processing module 22.

[0054] In the next step 120, the processing module 22 analyzes the received time series to determine outliers and / or missing data moments.

[0055] In particular, as explained previously, no data corresponds to a missing data point. This could correspond to a failure in a sensor 14 or in a communication means.

[0056] Furthermore, to identify outliers, the processing module 22 uses business rules. An outlier can also correspond to a failure in a sensor 14 or in a communication device. It is therefore likely to distort the processing results.

[0057] Business rules, for example, impose thresholds for each type of data and / or relationships between data of different types or between data corresponding to different times.

[0058] For example, business rules may include the following relationship for wastewater linking the quantity of ammonium NH4, suspended solids (SS), and the Kjeldahl total nitrogen (TTN) indicator:

[0059] NTK = alpha * MES + beta * NH4 where alpha and beta correspond to coefficients that can be determined empirically, using for example multilinear regression.

[0060] Thus, the determination of an outlier includes comparing each data point in the time series with at least one threshold and / or calculating a Mahalanobis distance and / or verifying at least one relationship.

[0061] Alternatively or in addition, the determination of an outlier may involve the implementation of an unsupervised machine learning model, for example of the IsolationForest type.

[0062] In the following step 130, when at least one outlier has been identified in step 120, the processing module 22 excludes that outlier from the time series. The time corresponding to this outlier is then considered a missing data point.

[0063] In the next step 140, the processing module 22 uses at least some of the time series data to train a deep learning model.

[0064] This section could, for example, include between 70% and 90% of all non-missing data. This data could be selected randomly or according to a predetermined rule. In addition, other data, such as data relating to past measurements verified by other means, could also be used.

[0065] The deep learning model is trained on relevant data to determine the behavior of the data over time and / or correlations between the data, using methods known per se. In particular, in some embodiments, a method enabling bidirectional learning (i.e., from the past to the future and from the future to the past) can be used.

[0066] The deep learning module includes, in particular, a neural network, advantageously a recurrent neural network (RNN), comprising an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.

[0067] More specifically, each layer comprises neurons taking their inputs from the outputs of the neurons in the previous layer, or from the input variables for the first layer.

[0068] Alternatively, more complex neural network structures can be considered with a layer that can be linked to a layer further away than the immediately preceding layer.

[0069] Each neuron is also associated with an operation, that is, a type of processing, to be carried out by said neuron within the corresponding processing layer.

[0070] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a link between two neurons.

[0071] Each neuron performs a weighted summation of the value(s) received from the neurons in the preceding layer. Each value is then multiplied by the respective synaptic weight of each synapse, or connection, between that neuron and the neurons in the preceding layer. Next, an activation function, typically a non-linear function, is applied to this weighted summation. The resulting value is then delivered to the neuron's output, particularly to the neurons in the next layer connected to it. The activation function introduces non-linearity into the processing performed by each neuron. The sigmoid function, the hyperbolic tangent function, and the Heaviside function are examples of activation functions.

[0072] As an optional complement, each neuron is also capable of applying, in addition, a multiplicative factor, also called bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the bias value and the value from the activation function.

[0073] In the next step, step 150, the processing module 22 performs a test of the deep learning model using the unused time series data for its training. In other words, for this purpose, the processing module 22 can use, for example, 10% to 30% of the unused data to train the model.

[0074] To test the model, the processing module 22 determines, for example, test data corresponding to the unused data for training the model—that is, data at the time points associated with this unused data. Then, the processing module 22 compares this test data with the unused data.

[0075] When the test data differs considerably from the unused data, the processing module 22 can readjust the model, or even train a new model, by returning to step 140. In particular, during this step 140, other data can be chosen to train the model.

[0076] Otherwise, during step 160, the processing module 22 uses the model to determine the missing data and completes the time series with this data.

[0077] Then, in the following step 170, the output module 23 transmits the completed time series to any interested system. Such an interested system may be located near or far from the processing system 10 and allows, for example, the implementation of a treatment other than the treatment performed by system 10, for example, a treatment for a particular task, such as determining a water quality indicator, the total water consumption for the installations 12, etc.

[0078] It is therefore understandable that the present invention has a number of advantages.

[0079] First, as explained previously, the invention makes it possible to determine missing data in a time series with exceptional accuracy. The invention also makes it possible to exclude outliers that could skew the results. Finally, the invention makes it possible to test the performance of the deep learning model using data not used for its training.

[0080] The inventors have also conducted numerous tests demonstrating that the method according to the invention provides more accurate results compared to those obtained using prior art methods, particularly those employing industry-standard formulas. Indeed, the accuracy of the results obtained using the method according to the invention is, on average, at least twice as high as that obtained using prior art methods.

Claims

DEMANDS 1. A method for processing measurements acquired in water installations (12), said measurements forming at least one time series, the time series presenting a set of instants and a set of data associated with these instants, the method comprising the following steps: - analysis (120) of the time series to determine outliers and / or missing data points corresponding to times in which at least one data point is missing; - when at least one outlier is determined, exclusion (130) of this data from the time series, the corresponding instant then being considered as an instant of missing data; - training (140) of a deep learning model from at least some of the time series data; - determination (160) of missing data at missing data times using the deep learning model.

2. Processing method according to claim 1, further comprising a step of transmitting (170) the completed time series to any interested system.

3. Processing method according to claim 1 or 2, wherein the determination of each outlier is carried out by applying business rules.

4. Processing method according to any one of the preceding claims, wherein the determination of each outlier includes comparing the corresponding data with at least one threshold and / or calculating a Mahalanobis distance and / or verifying at least one relationship and / or implementing an unsupervised machine learning model.

5. Processing method according to any one of the preceding claims, wherein each missing data is due to a failure of a sensor (14) or to a transmission fault.

6. A processing method according to any one of the preceding claims, wherein the training of the deep learning model includes determining the behavior of data over time and / or determining correlations between data.

7. Processing method according to any one of the preceding claims, wherein the training of the deep learning model is carried out using a portion of the time series data.

8. Processing method according to claim 7, further comprising a testing step (150) of the deep learning model using the time series data not used for its training.

9. A treatment method according to any one of the preceding claims, wherein each measurement corresponds to a parameter of the water passing through the water installations (12).

10. Computer program comprising software instructions which, when executed by a computer, implement the processing method according to any one of the preceding claims.

11. Processing system (10) for measurements acquired in water installations (12), comprising technical means (21, 22, 23) configured to implement the processing method according to any one of claims 1 to 9.