Method for processing measurements acquired in water installations, associated computer program and processing system

The method addresses the challenge of precise data handling in water installations by using deep learning models to identify and predict missing data, significantly improving accuracy over existing methods.

FR3164819A1Pending Publication Date: 2026-01-23SUEZ INTERNATIONAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2024007791
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing methods for handling missing data in time series from water installations lack precision and accuracy, particularly when using empirical relationships and interpolation/extrapolation, and are not suitable for all cases.

Method used

A method involving time series analysis to identify outliers and missing data, followed by training a deep learning model to determine missing data points using neural networks, specifically recurrent neural networks, to predict missing data with greater accuracy.

Benefits of technology

The method achieves precise determination of missing data, excluding outliers and providing accurate predictions, with results being at least twice as accurate as prior art methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for processing measurements acquired in water installations, associated computer program and processing system. The present invention relates to a method for processing measurements acquired in water installations, said measurements forming at least one time series, the method comprising the following steps: - analysis (120) of the time series to determine outliers and / or missing data times corresponding to times in which at least one data point is missing; - when at least one outlier is determined, exclusion (130) of this data point from the time series, the corresponding time point then being considered as a missing data point; - training (140) of a deep learning model from at least some of the data points of the time series; - determination (160) of the missing data at the missing data points using the deep learning model.Figure for the abbreviation: Figure 2.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method for processing measurements acquired in water installations, associated computer program and processing system

[0001] The present invention relates to a method for processing measurements acquired in water installations. The present invention also relates to a computer program and a processing system associated with this processing method.

[0002] The present invention relates more particularly to the field of water treatment implemented by various water installations such as a water treatment plant, a water distribution plant, a wastewater treatment plant, etc.

[0003] In order to size these various installations and monitor their operation, it is often necessary to have data characterizing the water in these installations, corresponding, for example, to water consumption, its flow rate at a given point in the system, its chemical composition, etc. This data can then be used to determine, for example, water quality indicators in order to meet current standards or to alert consumers or carry out maintenance work, when necessary.

[0004] The corresponding data generally come from sensors operating at predetermined measurement frequencies and located in known locations. Furthermore, each data point is generally associated with a known time, which is also relevant when analyzing that data. For example, it is often necessary not only to know the magnitude of a given parameter but also the time at which that parameter had that value.

[0005] The data generated by the sensors and associated with known times thus form time series. Here, each time can correspond to any time reference and can, for example, be defined by the exact time the measurement was taken or by a period such as an hour, a day, a week, a month, etc. The known times can, for example, be determined according to a known frequency (for example, every hour).

[0006] In some cases, time series data may be missing. Data may be missing for a given instant due to various events that can occur in the data acquisition and transmission chain. For example, such an event could be a failure of a sensor, a transmission device, or a storage device. Data may also be considered missing when the determined measurement frequency does not allow the corresponding measurements to be used for a desired purpose.

[0007] To overcome the problem of missing data in time series, various methods are already known in the state of the art.

[0008] Most of these methods are based on the use of business formulas that allow, in particular, the deduction of certain types of data from other types of data using known relationships between these data types. However, these relationships are often obtained empirically and cannot be applied to all cases with the desired accuracy.

[0009] Methods involving interpolation and / or extrapolation of known data are also known, but these methods may also lack the desired precision.

[0010] Furthermore, state-of-the-art methods are not always suitable for dealing with data considered not missing but likely to distort the results.

[0011] The present invention aims to remedy these drawbacks and therefore to propose a method and a processing system enabling the processing of time series relating to measurements acquired in water installations in a particularly precise manner.

[0012] To this end, the invention relates to a method for processing measurements acquired in water installations, said measurements forming at least one time series, the time series presenting a set of instants and a set of data associated with these instants, the method comprising the following steps:

[0013] - time series analysis to determine outliers and / or missing data moments corresponding to moments in which at least one piece of data is missing;

[0014] - when at least one outlier is identified, exclusion of that outlier from the time series, the corresponding instant then being considered as an instant of missing data;

[0015] - training a deep learning model from at least some of the data from the time series;

[0016] - determination of missing data at missing data times in using the deep learning model.

[0017] Thanks to these features, the invention makes it possible to determine missing data with particular precision. Indeed, the deep learning model makes it possible to determine a particular behavior of the data and / or correlations between these data. This then makes it possible to predict missing data with greater accuracy. The invention is therefore particularly advantageous compared to state-of-the-art methods that use, for example, only business rules and / or data interpolation / extrapolation.

[0018] According to other advantageous aspects of the invention, the method comprises one or more of the following features, taken individually or in all technically possible combinations:

[0019] - the determination of each outlier is carried out by applying rules job ;

[0020] - the determination of each outlier includes comparing the corresponding data with at least one threshold and / or the calculation of a Mahalanobis distance and / or the verification of at least one relationship and / or the implementation of an unsupervised machine learning model;

[0021] - each missing data point is due to a sensor failure or a defect in transmission;

[0022] - training the deep learning model includes determining of the behavior of the data over time and / or the determination of correlations between the data;

[0023] - the training of the deep learning model is carried out from a part of the time series data;

[0024] - the method further comprising a step of testing the learning model in depth by using time series data not used for its training;

[0025] - each measurement corresponds to a parameter of the water passing through the installations of water.

[0026] The invention also relates to a computer program comprising software instructions which, when executed by a computer, implement the process as defined above.

[0027] The invention also relates to a system for processing measurements acquired in water installations, comprising technical means configured to implement the process as defined above.

[0028] The invention will become clearer upon reading the following description, given solely by way of non-limiting example, and made with reference to the drawings in which:

[0029] - [Fig. 1] [Fig. 1] is a schematic view of a processing system according to the invention; and

[0030] - [Fig.2] [Fig.2] is a flowchart of a treatment process according to the invention, the treatment process being implemented by the treatment system of [Fig.1].

[0031] Fig. 1 illustrates a processing system 10 for processing measurements acquired in water installations 12.

[0032] Each water installation 12 includes, for example, a water treatment plant, a water distribution plant, a wastewater treatment plant, etc. Such a plant is particularly capable of treating or distributing water using methods known per se. In all that follows, "water" means clean water intended for use by consumers (e.g., drinking water) or wastewater.

[0033] The water installations 12 can be arranged in different geographical locations.

[0034] Each water installation 12 includes one or more sensors 14 capable of acquiring measurements relating to at least one parameter characterizing the water in that installation. Such a parameter may relate to the chemical composition of the water or to the manner in which it flows in the corresponding water installation 12, such as its flow rate, pressure, etc. For example, a sensor 14 may measure the ammonium NH4 content or suspended solids (SS) in wastewater or the pH in clean water.

[0035] The sensors 14 can be of the same nature (i.e., capable of measuring the same parameters) or of different natures (i.e., capable of measuring different parameters).

[0036] Each sensor 14 is capable of acquiring the corresponding measurements and generating data, for example, numerical data, corresponding to these measurements. Each sensor 14 is also capable of transmitting each generated data to the processing system 10 via any means of communication known per se, such as a wired or wireless computer network. In addition, or alternatively, each sensor 14 is also capable of transmitting each generated data to a centralized storage system, such as a SCADA (Supervisory Control and Data Acquisition) system.

[0037] Advantageously, each data point generated by a sensor 14 is associated by that same sensor with a specific time. Such a time corresponds, for example, to a precise time and date of measurement or to a period in which the measurement was taken, such as an hour, a day, a week, a month, etc.

[0038] Alternatively, each data generated by a sensor 14 is associated with a time by any other means, for example by any other means of the installation 12 or by the processing system 10.

[0039] In certain cases, each sensor 14 is capable of acquiring the corresponding measurements at a predetermined frequency. This frequency can then be used to associate times with the corresponding data.

[0040] The data generated by the sensors 14, together with the times associated with this data, form one or more time series. In particular, each time series presents a set of times and a set of data associated with these times.

[0041] Several data points can, for example, be associated with the same instant. In this case, it is a vector data point. Furthermore, an instant may not be associated with any data. This is then a missing piece of data. Such a moment is subsequently called a missing data moment.

[0042] The processing system 10 is capable of processing the data generated by the sensors 14, as will be explained in more detail later. The processing system 10 is, for example, remote from any water installation 12 or is part of at least one of these water installations 12.

[0043] With reference to [Fig.1], the processing system 10 comprises an input module 21, a processing module 22 and an output module 23.

[0044] The input module 21 is connected directly or indirectly to each of the sensors 14 (optionally via a centralized storage system) using a communication means as defined above and capable of acquiring the data generated by these sensors 14. Advantageously, this data is acquired in the form of one or more time series, i.e., with the corresponding instants. Alternatively, the input module 21 is capable of associating the instant corresponding to each acquired data point to form one or more time series. Such an instant corresponds, for example, to the instant of reception of the corresponding data point.

[0045] In some embodiments, the input module 21 includes a storage unit for storing data for a specified period.

[0046] The processing module 22 is capable of processing the data acquired by the input module 21. In particular, the processing module 22 is capable of processing this data in the form of at least one time series in order in particular to impute the missing data to it, as will be explained later.

[0047] The output module 23 is capable of transmitting the data processed by the processing module 22 to any interested external system. Such an external system allows, for example, the implementation of processing other than that performed by the system 10, for example, processing for a particular task, such as determining a water quality indicator, the total water consumption for the installations 12, etc.

[0048] Each of the modules 21 to 23 is, for example, in the form of one or more software programs, that is, in the form of a computer program, also called a computer program product. The software program(s) is implemented by one or more processors and random access memory (RAM), which are also part of the processing system 10. The software program(s) is, for example, stored on a computer-readable medium (not shown). The computer-readable medium is, for example, a medium capable of storing electronic instructions and being connected to a bus of a computer system. By way of example, the readable medium is an optical disc, a magneto-optical disc, ROM, RAM, any type of Non-volatile memory (e.g., FLASH or NVRAM) or a magnetic card. A computer program containing software instructions is then stored on the readable medium.

[0049] Alternatively or in addition, at least one of the modules 21 to 23 is at least partially in the form of a programmable logic circuit such as an FPGA (Field Programmable Gate Array) or an integrated circuit, such as an ASIC (Application Specified Integrated Circuit).

[0050] The processing system 10 is capable of implementing a processing method according to the invention which will henceforth be explained with reference to [Fig.2] showing a flowchart of its steps.

[0051] It is initially considered that the sensors 14 acquire measurements and generate data corresponding to these measurements which are then transmitted to the processing system 10, possibly with the corresponding times.

[0052] During an initial step 110, the input module 21 of the processing system 10 acquires this data with the corresponding time points or associates time points with the acquired data as explained previously. In other words, the input module 21 acquires or forms at least one time series during this step 110.

[0053] According to certain embodiments, the time series is acquired or generated by the input module 21 after a predetermined time period. In other words, during this time period, the input module 21 can receive data from different sensors 14 at different frequencies and generate the corresponding time series only after this period. To store the intermediate data, the input module 21 can use its storage unit. This time period can correspond to an hour, a day, a week, a month, a year, etc.

[0054] The total number of data in the time series formed can be greater than 1000, advantageously greater than 10,000.

[0055] At the end of step 110, the input module 21 transmits the corresponding time series to the processing module 22.

[0056] During the next step 120, the processing module 22 analyzes the received time series to determine outliers and / or missing data moments.

[0057] In particular, as explained previously, no data corresponds to a missing data point. This may correspond to a failure occurring in a sensor 14 or in a communication means.

[0058] Furthermore, to identify outliers, the processing module 22 uses business rules. An outlier can also correspond to a failure in a sensor 14 or in a communication device. It is therefore likely to distort the processing results.

[0059] Business rules impose, for example, thresholds for each type of data and / or relationships between data of different types or between data corresponding to different times.

[0060] For example, business rules may include for wastewater the following relationship linking the quantity of ammonium NH4, suspended solids TSS and the NTK indicator (total nitrogen Kjeldahl):

[0061] NTK = alpha * MES + beta * NH4

[0062] where

[0063] alpha and beta correspond to coefficients that can be determined empirically, for example using multilinear regression.

[0064] Thus, the determination of an outlier includes comparing each data point in the time series with at least one threshold and / or calculating a Mahalanobis distance and / or verifying at least one relationship.

[0065] Alternatively or in addition, the determination of an outlier may include the implementation of an unsupervised machine learning model, for example of the IsolationForest type.

[0066] In the following step 130, when at least one outlier has been identified in step 120, the processing module 22 excludes that outlier from the time series. The time corresponding to that outlier is then considered a missing data point.

[0067] In the next step 140, the processing module 22 uses at least part of the time series data to train a deep learning model.

[0068] This part may, for example, comprise between 70% and 90% of all non-missing data. This data may, for example, be chosen randomly or according to a predetermined rule. In addition, certain other data, for example, data relating to measurements carried out in the past and verified by other means, may also be used.

[0069] The deep learning model is trained with the corresponding data to determine the behavior of the data over time and / or correlations between these data, according to methods known per se. In particular, in certain embodiments, a method allowing bidirectional learning (i.e., from the past to the future and from the future to the past) may be used.

[0070] The deep learning module includes in particular a neural network, advantageously a recurrent neural network (RNN), comprising an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.

[0071] More precisely, each layer comprises neurons taking their inputs from the outputs of the neurons of the previous layer, or from the input variables for the first layer.

[0072] Alternatively, more complex neural network structures can be envisaged with a layer that can be linked to a layer further away than the immediately preceding layer.

[0073] Each neuron is also associated with an operation, that is to say a type of processing, to be carried out by said neuron within the corresponding processing layer.

[0074] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a link between two neurons.

[0075] Each neuron is designed to perform a weighted sum of the value(s) received from the neurons of the preceding layer, each value being multiplied by the respective synaptic weight of each synapse, or connection, between said neuron and the neurons of the preceding layer, and then to apply an activation function, typically a non-linear function, to said weighted sum, and to deliver at the output of said neuron, in particular to the neurons of the next layer connected to it, the value resulting from the application of the activation function. The activation function introduces non-linearity into the processing performed by each neuron. The sigmoid function, the hyperbolic tangent function, and the Heaviside function are examples of activation functions.

[0076] As an optional complement, each neuron is also capable of applying, in addition, a multiplicative factor, also called bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the bias value and the value from the activation function.

[0077] In the next step 150, the processing module 22 performs a test of the deep learning model using the unused time series data for its training. In other words, for this purpose, the processing module 22 can use, for example, 10% to 30% of the unused data to train the model.

[0078] To test the model, the processing module 22 determines, for example, using the model, test data corresponding to the unused data for training the model, that is, data at the times associated with this unused data. Then, the processing module 22 compares this test data with the unused data.

[0079] When the test data differs considerably from the unused data, the processing module 22 can readjust the model, or even train a new model, by returning to step 140. In particular, during this step 140, other data can be chosen to train the model.

[0080] Otherwise, during step 160 the processing module 22 uses the model to determine the missing data and completes the time series with this data.

[0081] Then, in the next step 170, the output module 23 transmits the completed time series to any interested system.

[0082] It is therefore understood that the present invention has a number of advantages.

[0083] First, as explained previously, the invention makes it possible to determine missing data in a time series with particular precision. The invention also makes it possible to exclude outliers that could skew the results. Finally, the invention makes it possible to test the operation of the deep learning model using data not used for its training.

[0084] The inventors have also carried out numerous tests demonstrating that the method according to the invention provides more accurate results compared to those obtained using prior art methods, particularly those using industry-standard formulas. Indeed, the accuracy of the results obtained using the method according to the invention is on average at least twice as high as that obtained using prior art methods.

Claims

Demands

1. A method for processing measurements acquired in water installations (12), said measurements forming at least one time series, the time series presenting a set of instants and a set of data associated with these instants, the method comprising the following steps: - analysis (120) of the time series to determine outliers and / or missing data instants corresponding to instants in which at least one data point is missing; - when at least one outlier is determined, exclusion (130) of this data point from the time series, the corresponding instant then being considered as a missing data instant; - training (140) of a deep learning model from at least some of the data from the time series; - determination (160) of the missing data at the missing data instants using the deep learning model.

2. Processing method according to claim 1, wherein the determination of each outlier is carried out by applying business rules.

3. Processing method according to claim 1 or 2, wherein the determination of each outlier includes comparing the corresponding data with at least one threshold and / or calculating a Mahalanobis distance and / or verifying at least one relationship and / or implementing an unsupervised machine learning model.

4. A processing method according to any one of the preceding claims, wherein each missing data is due to a sensor failure (14) or a transmission fault.

5. Processing method according to any one of the preceding claims, wherein the training of the deep learning model includes determining the behavior of the data over time and / or determining correlations between the data.

6. A processing method according to any one of the preceding claims, wherein the training of the learning model in-depth analysis is performed using a portion of the time series data.

7. Processing method according to claim 6, further comprising a testing step (150) of the deep learning model using the time series data not used for its training.

8. A treatment method according to any one of the preceding claims, wherein each measurement corresponds to a parameter of the water passing through the water installations (12).

9. A computer program comprising software instructions which, when executed by a computer, implement the processing method according to any one of the preceding claims.

10. System for processing (10) measurements acquired in water installations (12), comprising technical means (21, 22, 23) configured to implement the processing method according to any one of claims 1 to 8.