A method of modelling a production system

EP4673872A1Pending Publication Date: 2026-01-07SOLUTION SEEKER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024764260
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-02
Filing Date
2024-02-29
Publication Date
2026-01-07

AI Technical Summary

Technical Problem

Current production system modeling relies heavily on mechanistic approaches, which are costly and time-consuming, limiting the adoption of new technologies and innovation, especially in industries with limited resources, and faces challenges with paucity, inaccuracies, and unavailability of production data.

Method used

A parametric generative model is used to generate synthetic production data, simulating the full probability distribution of production data, allowing for unconditional data generation and improved applicability across various tasks, reducing reliance on recorded data and enhancing data efficiency.

Benefits of technology

The generative model provides accurate and efficient modeling of production systems, enabling better monitoring, prediction, and optimization of production performance, even with scarce data, and supports real-time analytics and soft sensing applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure NO2024050053_06092024_PF_FP
    Figure NO2024050053_06092024_PF_FP
Patent Text Reader

Abstract

A method of modelling a production system comprises using a parametric generative model to generate synthetic production data relating to the production system, wherein the parametric generative model is adapted to simulate how production data associated with the production system is generated.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]A METHOD OF MODELLING A PRODUCTION SYSTEM TECHNICAL FIELD The invention relates to methods of modelling a production system and to methods of training a model for a production system. The invention further extends to a corresponding model. The invention further relates to methods of analysing production performance for a production system, methods of determining predicted production performance for a production system, and methods of improving or optimising production performance for a production system. A computer programme product comprising instructions that permit execution of one or more methods of the invention is also provided by the invention. Similarly, a computer system configured to carry out one or more methods of the invention is also provided by the invention. BACKGROUND A production system, as used herein, may refer to a production unit that is a single (sub-)system that produces a particular resource or product or that is involved in the production of that particular resource or product. For example, a production unit may constitute a hydrocarbon production well, a wind turbine or a solar panel, and may equally constitute a sub-system of these devices, e.g. a pump, valve, separator, generator. Equally, a production system as used herein may refer to a production asset, which typically consists of a collection of production units and / or equipment and apparatus provided in support of the production unit. For example, a production asset may be a wind farm comprising a plurality of wind turbines or a hydrocarbon production system / facility comprising a plurality of production wells. Production systems also commonly come in the form of processing units or processing facilities, and hence these terms are commonly be used interchangeably in field of the invention. In the context of a processing unit / facility, the resource or product that is deemed to be produced is the output of the process unit / facility. For example, a milk processing facility may have raw milk as an input and output, as a product, pasteurised and homogenised milk. Production systems come in many different forms. For example, a production system may be an energy production system or a physical resource production system, e.g. a natural resource production system. Examples of production systems, and more specifically production assets, include power plants: such as coal-fired, nuclear power, fuel cell based, hydrogen fuel cell based, wind, hydroelectric, natural gas and solar panel based power plants. Further examples, include farms, such as arable or pastoral (e.g. fish) farms. Hydrocarbon (oil and gas) production facilities and production reserves are also types of production asset. Chemical production facilities and plants also constitute production assets. Manufacturing plants and facilities, along with mining plats and facilities, further constitute types of production assets. It will be appreciated that each of the exemplary assets specified above comprise a combination of different production units that each serve to produce a product or resource and which, collectively, constitute the relevant production asset. Desired and safe operation of a production system requires knowledge about the operating performance of the production system, along with knowledge relating to operating and other parameters associated with the production system. Parameters associated with actuators / control devices that control operation of the production system, parameters associated with the rate of production of the resource or product of interest, parameters relating to the physical properties of the production system are all exemplary parameters that may be desired or essential to know for proper, safe and optimised operation of and production from the production system. For example, with respect to the example of a production well, desired operation of the production well may require knowledge relating to production flow rate, produced fluid composition, etc. It may also be necessary to have knowledge of parameters such as pressures, temperatures, choke valve position, pipe diameters, choke valve dimensions etc. to ensure desired and safe operation of the production well. Production data relating to the operating performance and other parameters of the production system may be measured (i.e. recorded) to allow for desired and safe operation of the production system to be determined and maintained. For example, and with further reference to the example a production well, Figure 1 shows a schematic of a typical production well 1 for which production data is being recorded. The production well 1 is connected to a hydrocarbon reservoir (not shown). Hydrocarbons from the reservoir are produced via the well 1 and transferred to a separator or production manifold (not shown) via conduit 3. A choke valve 5 is positioned within the conduit 3 and can be opened and closed to permit and prevent the flow of hydrocarbons from the production well 1. The choke valve 5 thus acts as an actuator that impacts on the production of fluid from production well 1. Situated in proximity to the production well 1 (i.e. “downhole”) is a pressure sensor 7 configured to record the pressure of the hydrocarbons as they emanate from the production well 1. In proximity to the choke valve 5 and situated upstream thereof are a pressure sensor 9 and temperature sensor 11. The pressure sensor 9 and temperature sensor 11 are configured to record the pressure and temperature of the hydrocarbons, respectively, immediately upstream of the choke valve 5. Also in proximity to the choke valve 5 and situated downstream thereof are a pressure sensor 13 and temperature sensor 15. The pressure sensor 13 and temperature sensor 15 are configured to record the pressure and temperature of the hydrocarbons, respectively, immediately downstream of the choke valve 5. The measurements at the pressure sensor 13 and temperature sensor 15 can be compared to the measurements at the pressure sensor 9 and temperature sensor 11 to determine how the choke valve 5 and its degree of opening impacts on the hydrocarbons passing therethrough. These measurements and the determination made therefrom are forms of production data that are recorded in relation to the production well. Attached to the conduit 3 at a location downstream of the pressure sensor 13 and temperature sensor 15 is a bypass conduit 3A which leads to a test separator 17. The bypass conduit 3A and test separator 17 permit well tests to be carried out such that production performance of the production well can be determined (i.e. another form of recorded production data). When it is desired to carry out a well test, e.g. to record the flow rate of the fluid produced, the produced hydrocarbons from the well 1 can be diverted from the separator or production manifold and to the test separator 17 via bypass conduit 3A. The test separator 17 separates out the produced fluid out into multiple single phases (e.g. water, oil and gas) and then the flow rate of each of these phases is measured. A multiphase flow meter (MPFM) 19 is also situated proximate and upstream to the choke valve 5. The MPFM 19 hosts multiple different sensors which provide measurements that are fused in a model in real time to estimate flow rates of the fluid produced from the well 1 in real time, which is indicative of production performance (i.e. a form of production data). The production data recorded using the MPFM 19, test separator 17, the pressure sensors 7, 9, 13, and the temperature sensors 11, 15 allow, to a degree, operational performance of the production well to be monitored and to thereby allow efficient, safe and optimised production from the production well to be managed. However, limits are imposed on how well production performance and operation of the production well can be monitored by virtue of the recorded production data from the MPFM 19, test separator 17, the pressure sensors 7, 9, 13, and the temperature sensors 11, 15. This is in large part due to a paucity in the data that is recorded, inaccuracies in the data that is recorded, and the relevant data not being recorded and thus being unavailable. For example, there is typically a paucity in the data recorded at the test separator 17. Well tests using the test separator 17 are performed on an intermittent basis such that the flow rates for the well 1 can be recorded. However, a well test may take several hours (or even days) to complete, and the test separator 17 will typically be used for a large number of wells (i.e. several wells other than production well 1). Routing of produced fluid to the test separator 17 can also result in an impedance of production performance of production well 1. Thus, there are limitations imposed on the frequency with which well tests can be performed. A typical frequency for well testing of a given well is once a month and as such the flow rate data as measured by a well test are relatively scarce. Moreover, whilst flow rate measurements for the production well 1 can be made far more frequently with the MPFM 19 as compared to those determined from a well test using test separator 17, MPFM are rarely provided / used since their expense is not always commercially warranted. Hence data from MPFM 19 will, for a majority of production wells, be unavailable. MPFM measurements also tend to be inaccurate and they have limited applicability since they are typically only designed to determine flow rates for a given type of flow regime from the production well. Therefore, if the flow regime slightly differs from that for which the MPFM was designed then the measurements of flow rate may be associated with significant inaccuracies. Pressure sensors and temperature sensors such as pressure sensors 7, 9, 13, and temperature sensors 11, 15 are commonplace (i.e. are found in association with a vast majority of production wells) and, moreover, the data these sensors record is typically highly accurate. In addition, the data recorded with pressure sensors and temperature sensors of the type shown in Figure 1 is typically highly frequently recorded. Accordingly, it might be assumed that the data recorded using pressure sensors and temperature sensors of the type shown in Figure 1 is not afflicted with the same deficiencies that impact the data recorded using the MPFM 19 or the test separator 17. However, pressure sensors, and temperature sensors of the type depicted in Figure 1 are prone to dropout, drift or damage which prevents data from being recorded accurately (or at all). Moreover, given their position within the production system, it is not straightforward for these sensors to be repaired / replaced or brought back online. Accordingly, there can be significant periods of time during which pressure sensors and temperature sensors of the type depicted in Figure 1 are unable to record data. More generally, there are also limitations in the recorded production data for a given production well (e.g. production well 1) in that the data is only indicative of the past and present performance of that production well under the very specific operational constraints which said production well has experienced or is experiencing. Accordingly, the recorded data gives limited insight into future performance and / or performance of the production well under different operating constraints / operating states. Hence, sole reliance on the data that is directly recorded (i.e. measured) in connection with a given production well to allow for monitoring and determinations to be made in relation to that production well can result in inaccurate or unreliable determinations related to production and operating performance given the paucity of the data available. Moreover, limited insight can be garnered in relation to the future performance of the production well given the inherent limitations in the data that has been recorded with that production well. Production systems more generally (i.e. not just limited to the specific example of a production well as discussed above) are similarly impacted by the paucity in production data that is directly recorded. As above, this may result due to an absence of the relevant sensors required to record the relevant production data, dropout / failure / drift of the relevant sensors, an inability to measure the desired production data, inaccuracies or infrequency of production data measurements etc. The production data recorded with production systems, generally, also shares similar inherent limitations to those discussed above in connection with the production data recorded in respect of production wells. That is, the recorded data is indicative only of past or present performance under a specific set of operating constraints / operating states and hence said data may not provide any particular insight into future performance of the production system. To overcome such shortcomings associated, modelling of production systems to, amongst other things, garner insights into production system performance has become more prevalent in recent years. Modern production and processing industries that make use of production systems rely heavily on information processing technologies to safely and efficiently operate said systems. Technologies important in operations include real-time analytics, soft sensing, process surveillance, condition performance monitoring, process control, and real-time optimization. These technologies are enabled by process / production models, which can be queried or used to perform application- specific inference tasks. For example, in soft sensing, a process model operates on process measurements to infer key performance indicators of a production system. The prevailing approach to modelling of industrial production processes is mechanistic (physics-based) modelling, where increasingly complex models are derived from first principles. Mechanistic modelling of this type is a ‘bottom-up’ approach comprising: 1. Using the scientific method, a model of a single production unit / process unit operation (henceforth called a production unit model) is derived from first principles and refined through experimentation 2. A library of production unit models is built 3. A model of an asset is built by assembling / connecting production unit models in a configuration that reflects the overall process 4. During operation, production unit models are calibrated to new experiment data (while attempting to honour prior beliefs based on the physical interpretation of adjusted model parameters) State of the art production modelling of this type is mostly facilitated by advanced computer simulation software, which can be generic or specialised to some production processes. Simulation software thus plays a crucial role in the operation of many industrial production systems. The rigorous mechanistic modelling approach often results in high model accuracy, but it may be tedious and costly. The high costs related to mechanistic model development and maintenance is slowing the implementation, and in some cases preventing the adoption, of new technologies. Thus, it may be unsustainable for many production industries to rely solely on mechanistic modelling in their future digital transition. Even for well-funded industries, it may be a hindrance to technological innovation. In the more recent past, production system modelling reliant on data driven modelling and machine learning techniques has been utilised. For example, and again with reference to the specific example of a production well, data driven and machine learning modelling techniques have been applied in order to model well parameters, flow parameters and information on control points associated with production wells (e.g. choke valves). Exemplary data driven modelling techniques are disclosed, for example, in the applicant’s own patent publication: WO 2019 / 110851 A1. These data driven modelling techniques differ from the mechanistic / physics based modelling techniques discussed above in that the models absorb and learn the relevant physics and behaviours of the production system from the data recorded in relation to the production system rather than having the relevant physics imputed therein. Thus, step 1 outlined above in relation to mechanistic models is absent from data driven modelling techniques since the data driven models learn and absorb the necessary physics and behaviours from the production system directly from the data upon which the model is trained. Hence, the data driven modelling techniques are significantly less laborious and costly in terms of the model that needs to be developed. Mechanistic models require significant investments in model development to ensure they accurately reflect the production system, whereas data driven modelling techniques do not require this same investment since the behaviours and properties of the production system are ‘learnt’ by virtue of the modelling undertaken. Data modelling techniques are valuable since they allow a production system variable or parameter of interest to be modelled without necessitating a ‘complete' set of measurements for that given variable / parameter to have been actually recorded. Thus, drawbacks associated with infrequency, unavailability and inaccuracies in recorded production data are not shared to the same extent by known data modelling techniques for production systems and hence drawbacks in performance determination and monitoring based on recorded production data can be moderated by using data-driven modelling in parallel with these measurement techniques. That is, the resultant model from the data driven modelling techniques can be used in support of recorded production data and can be used to provide predictions / estimates in the intervals where no recorded production data exists or as a back-up if a measurement (e.g. MPFM) device fails. Still further improvements in data driven modelling techniques for modelling production wells, and for modelling production systems more generally, are desired. SUMMARY OF THE INVENTION According to a first aspect of the invention, there is provided a method of modelling a production system, the method comprising: using a parametric generative model to generate synthetic production data relating to the production system, wherein the generative model is adapted to simulate how production data associated with the production system is generated. The invention of the first aspect uniquely makes use of a generative model to model the production system. A generative model, in the context of the present invention, captures the full probability distribution of the production data associated with the production system. That is, the generative model simulates how production data associated with the production system is generated in the real world. Accordingly, the generative model can be seen to be a ‘digital twin’ of the production system, with the synthetic data that it outputs being a simulation of data that would actually, in the real world, be associated with the production system. Known data driven models from the prior art (e.g. those disclosed in WO 2019 / 110851 A1) for modelling a given parameter of a production system are discriminative models. Discriminative models aim to learn an output for a variable of interest given known observations. For example, a discriminative model for a production well may aim to learn to infer / predict / estimate production flow rate based on values of pressure, temperature, choke position etc. A discriminative model effectively models the conditional probability distribution of a response / predicted / target variable (i.e. the parameter of interest) given a set of known or measured or recorded parameters / observations. It can therefore be understood how the known prior art data driven models for production wells can model and make predictions, estimations for a parameter having a paucity of measurements based on measurements of other parameters associated with the production well and which are more prevalently available as discussed above. A generative model as in the first aspect of the invention is in direct contrast to a discriminative model as discussed above. A discriminative model describes the probability distribution of a parameter of interest (y) given known (i.e. observed or measured) parameters (x) and as such is a conditional model (i.e. P(y|x)). As such, discriminative models solve a single learning task – i.e. knowledge of the parameter (y) of interest given known parameters (x) – and are typically trained by supervised training using “labelled” data pairs (x,y). By contrast, generative models capture the full probably distribution of the data (e.g. P(x, y)), by approximating both the conditional distribution P(y|x) and the marginal distribution P(x). In some situations they can be trained by semi- supervised learning using a combination of “labelled” data (x,y) and “unlabelled” measurements of x. They need not be limited to a single learning task, but may be more broadly applicable to a greater number of learning tasks. A generative model is thus not conditionally limited in same way that a discriminative model is and hence may be used to generate unconditional synthetic production data. To say this another way, since a generative model, in embodiments of the first aspect, can model the full distribution of the production data, the generative model may provide information on how likely a given example of the production data is. In contrast, a discriminative model cannot unconditionally indicate the likelihood of a given example of production data. A discriminative model can only provide information on the likelihood of a single parameter of interest (i.e. one for which the discriminative model has been designed to model) based (i.e. conditional) on other measured parameters (which the model is conditioned on). It will thus be appreciated that the modelling of the first aspect of the invention provides for greater applicability than known prior art modelling techniques for production systems, such as the discriminative modelling techniques discussed above. This is because the generative model gives a simulation of the total behaviour of the production system, and not just a specific behaviour based on specific known (recorded or observed) variables. As a result, the generative model used in the first aspect of the invention is less directly reliant on data that has been recorded in association with the production system. The generative model is also able to produce unconditional synthetic production data in relation to the production system, which is contrast to the discriminative models discussed above which require knowledge of at least some variables (i.e. recorded data) to model other variables. The generative model may also be adapted to a wide variety of different tasks / modelling problems, which discriminative models generally cannot. The parametric generative model of the first aspect produces synthetic production data in relation to the production system. This data is ‘synthetic’ as it has not actually been recorded or measured, but is instead synthetically generated by the model and reflects data that has been / will be / would have been / could have been recorded. Accordingly, the synthetic production data is to be seen as distinct from recorded or measured production data, which is data that has actually (i.e. in the ‘real world’) been recorded / measured in relation to the production system. The parametric generative model may be a trained model. It may have been trained once, or it may be trained repeatedly over time, at regular or irregular intervals (e.g. in response to receiving more recorded production data). It may be trained on a combination of labelled and unlabelled production data. Training the model partly on unlabelled data can allow the model to learn from instances of unlabelled measurement data that are available in large quantities (e.g. frequent temperature measurements), rather than only from a smaller set of labelled data (e.g. temperature measurements that are paired a parameter of interest such as flow rate, which may be measured relatively infrequently). This can enable the model to learn from a much greater amount of data than if only relatively-sparse labelled data were used for the training, e.g. as when training a discriminative model. The production system may be associated with one or more sensors, meters and / or measurement devices arranged to measure / record production data associated with the production system. For example, one or more pressure sensor(s), temperature sensors(s), flow meter(s), amp meter(s), power meter(s), voltage meter(s), speed sensor(s) / meter(s), time recording device(s), anemometer(s), device(s) for measuring mass(es), device(s) for measuring weight(s), device(s) for measuring length(s), device(s) for measuring molar amounts of a substance, device(s) for measuring concentration(s), device(s) for measuring forces, device(s) for measuring electrical resistance(s), device(s) for measuring luminous flux and / or intensity, stress and / or strain gauge(s), device(s) for detecting and / or measuring the presence / amount of a given substance, radiation detector(s) and / or any other sensor(s) / measurement device(s) / meter(s) typically used to record data associated with production systems. The exact sensor(s) / measurement device(s) / meter(s) present will depend upon the specific nature of the production system. The method of the first aspect may comprise recording / measuring / detecting / sensing production data associated with the production system, optionally with the one or more sensor(s), meter(s) and / or measurement device(s). The parametric generative model may be configured to transform recorded production data into a latent space. This latent space may be of lower dimension than an observable space of the production data. The generative model may have been trained to learn a joint probability distribution over the latent space and the observable space. The parametric generative model may be used for unconditional generation of synthetic production data. However, in some embodiments, the step of using a parametric generative model to generate synthetic production data relating to the production system may comprise generating conditional synthetic production data. The method may hence comprise, prior to or as part of the step of using a parametric generative model to generate synthetic production data, inputting production data (optionally recorded production data) into the parametric generative model. In such a scenario, the generative model may comprise an encoder or inference network and a generative part of the model (e.g. decoder). Inputting the production data into the generative model may comprise inputting the production data into the encoder or inference network. Input of the production data into the encoder / inference network may produce a set of latent variables in a latent space that are conditioned on the production data such that some latent variables are more probable than others. Generating the synthetic production data relating to the production system may comprise mapping the conditioned latent variables through the generative part of the model to generate the synthetic production data in the initial (e.g. ‘real-world’) space. By inputting (e.g. recorded) production data into the generative model prior to using the parametric generative model to generate synthetic production data the generative model is made to generate synthetic production data conditioned on the input production data. For example, in the specific example of a production well, recorded or hypothetical temperature and / or pressure data may be input into the generative model prior to using the generative model to generate synthetic flow rate / composition data in order to determine a flow rate / composition. Input of production data prior to generation of the synthetic production data can be seen as a ‘calibration’ of the generative model, such that the generative model has a greater likelihood of modelling likely / accurate production data. The production data that is optionally input into the generative model may be recorded production data, e.g. of the type discussed previously. By using the generative model to produce conditioned synthetic production data, the model might be thought of as sharing some similarities with prior art discriminative models as discussed previously. However, the skilled person would recognise that a generative model conditioned in this way remains technically distinct from a discriminative model. In particular, and as discussed above, a given discriminative model is designed for only a single learning task (e.g. modelling flow rates given know temperature measurements). In contrast, the generative model of the invention of the first aspect is not limited in this way since the generative model describes the entire probability distribution of production data, including modelling the input parameters (e.g. the temperature measurements). Accordingly, the generative model may be conditioned on any type of production data (i.e. not limited to a single type or types of data for which the model has been designed as is the case for the discriminative model) and can be used to model any production data of interest (i.e. is not limited to modelling a single or specified set of parameters). Whilst the generative model of the first aspect of the invention may be used conditionally as discussed above, it is not necessary for the generative model to be used conditionally. Rather the model can, in some embodiments, be used unconditionally. That is to say, the step of using a parametric generative model to generate synthetic production data relating to the production system may comprise generating unconditional synthetic production data. Inputting further production data in order to enable the trained parametric generative model to generate synthetic production data is hence not necessary. Since the model is generative rather than discriminative it does not need an input of data to provide an output and may be used to generate unconditional synthetic production data. This is something that a discriminative model is simply not capable of given its nature. The recorded production data that is optionally input into the generative model when the model is used conditionally may optionally be cleaned and / or compacted prior to input into the generative model, and hence the method of the first aspect may comprise cleaning and / or compacting the recorded production data. The cleaning and / or compacting of the recorded production data may comprise (1) gathering historical recorded production data and / or live recorded production data; (2) identifying time intervals in the recorded production data during which production data is in a steady state; and (3) extracting statistical data representative of some or all steady state intervals identified in step (2) to thereby represent the original data from step (1) in a compact form. Alternatively, the cleaning and / or compacting of the recorded production data may comprise: (1) gathering recorded production data covering a period of time; (2) identifying multiple time intervals in the data during which the recorded production data can be designated as being in a category selected from multiple categories relating to different types of stable production and multiple categories relating to different types of transient events, wherein the data hence includes multiple datasets each framed by one of the multiple time intervals; (3) assigning a selected category of the multiple categories to each one of the multiple datasets that are framed by the multiple time intervals; and (4) extracting statistical data representative of some or all of the datasets identified in step (2) to thereby represent the original recorded production data from step (1) in a compact form including details of the category assigned to each time interval in step (3). The production system may be associated with one or more actuators arranged to alter operation / performance of the production system. The actuator may be a control point. The actuator / control point may be a means / mechanism capable of applying a controlled adjustment to the operation or performance of the production system. Valves, pumps, compressors, switches, clutches, hydraulic actuators etc. etc. are all forms of actuators that may be associated with the production system. As will be apparent from the above discussion, the generative model may be a machine learning model—optionally, a deep learning model. The generative model may comprise or be based on one or more neural networks, optionally one or more deep neural networks. The model may comprise a neural network, trained to represent a conditional probability distribution relating to the production data. The generative model may be a generative adversarial network or an autoregressive model. The generative model may be, or be comprised as part, of a variational autoencoder (VAE). A VAE is a probabilistic generative model. The VAE may comprise an encoder (also termed an ‘inference model’ or a ‘recognition model’) and a decoder. The decoder in itself is a form of generative model. Accordingly, in scenarios where the generative model of the first aspect is comprised as part of the VAE, the generative model may be the decoder. Alternatively, the VAE itself (comprising both the encoder and the decoder) may be the generative model of the first aspect. The encoder and the decoder may be parameterized models. The encoder and decoder may or may not be independently parameterized. The encoder may be configured to transpose production data from an initial space (i.e. in decoded form) into a latent space whilst reducing its dimensionality, optionally without loss of information. The decoder may be configured to produce synthetic production by transposing data from the latent space back into the initial space (i.e. back in decoded form). It will be appreciated from the above discussion in connection with conditional synthetic production data, that generation of conditional synthetic production data may comprise use of a VAE. More specifically, the parametric generative model of the method of the first aspect may be a VAE comprising an encoder and a decoder, and before or as part of the step of using the VAE to generate synthetic production data relating to the production system the method of the first aspect may comprise inputting production data (e.g. recorded production data) into the encoder to produce a set of latent variables in the latent space that are conditioned on the production data such that some latent variable are more probable than others. The step of generating the synthetic production data relating to the production system may then comprise mapping the conditioned latent variables through the decoder to generate the synthetic production data. Equally, it will be appreciated from the above discussion in connection with unconditional synthetic production data that generation of unconditional synthetic production data may comprise use of a VAE. More specifically, the parametric generative model of the method of the first aspect may be a decoder of a VAE. The step of generating unconditional synthetic production data relating to the production system may comprise mapping latent variables through the decoder to generate the unconditional synthetic production data. The method of modelling in the first aspect of the invention may be a method of estimating production data. Accordingly, the generated synthetic production data may comprise estimated production data. The method may hence be considered as a method of ‘soft’ or ‘virtual’ sensing, particularly in a scenario where no actual recorded data correspondent to the estimated production data exists. The estimated production data may comprise conditional estimated production data. As such, prior to or as part of the step of using the parametric generative model to generate estimated production data, the method may comprise inputting production data (e.g. recorded production data) into the (trained) parametric generative model as discussed above. The estimated production data may comprise unconditional estimated production data. Estimations are useful as they allow for determinations to be made regarding production data for the production system. These determinations can then be used to make inferences and assessments in connection with the production system and its performance – i.e. they allow past, present or future performance of the production system to be analysed. The method of estimating production data may be a method of estimating production data from a period of time in the past. Hence, the estimated production data may comprise estimated production data for the period of time in the past, optionally a period of time in the past where no correspondent recorded production data exists. Thus, the method may be used to backfill production data (i.e. fill in for missing recorded production data from the past). This may be useful in scenarios where that data is otherwise unavailable, e.g. because no appropriate sensor / measurement device is available for recording said production data or because no production data was recorded (e.g. due to sensor dropout, sensor freeze, or because no recording was taken at the given time). The period of time in the past may be an instantaneous point in time from the past, or may be a period of time from the past, e.g.1 minute, 2 minutes, 5 minutes, 10 minutes, 30 minutes, 1 hour, 2 hours, 5 hours, 10 hours, 12 hours, 1 day, 2 days, 5 days, 1 week, 2 weeks, 3 weeks, 1 month and / or any period of time within the range of 1 minute to 1 month. The method of estimating production data may be a method of estimating present (i.e. real-time) or future production data. Where the method comprises estimating present (i.e. real-time) production data the estimated production data may comprise estimated real-time production data. Some or all production data recorded in real-time may be unavailable or unavailable in real-time. Where the method comprises estimating future production data the estimated production data may comprise estimated future production data. In a second aspect of the invention, the method provides a method of analysing production performance of a production system, the method comprising estimating production data in accordance with the optional forms of the first aspect of the invention described above; and analysing production performance of the production system based on the estimated production data. In a third aspect of the invention, there is provided a method of identifying an outlier or incorrect data (e.g. data resulting from a frozen or stray measurement) in recorded production data associated with a production system, the method comprising: estimating production data associated with the production system in accordance with the relevant optional forms of the first aspect of the invention described above; comparing the estimated synthetic production data with production data recorded for the same period of time; identifying any deviation between the recorded production data and the estimated synthetic production data; and determining that the recorded production data comprises an outlier or incorrect data when the deviation exceeds a threshold quantity. In a fourth aspect of the invention, there is provided a method of cleaning recorded production data, the method comprising identifying an outlier or incorrect data in recorded production data associated with a production system in accordance with the third aspect of the invention; and removing that production data from the recorded production data. The method of the first aspect of the invention may be a method of predicting potential future production data. As such, the generated synthetic production data may comprise potential future production data. Predicted potential future production data is a form of conditional synthetic production data. This is because, in order to make a prediction of production data, it is necessary to input production data into the (trained) parametric generative model, with the input production data being reflective of an operating state (for example, a hypothetical state) of the production system for which the prediction is being made. Accordingly, where the method of the first of the invention is a method of predicting potential future production data, the method may comprise, prior to or as part of the step of using a parametric generative model to generate predicted potential future production data, inputting production data (optionally recorded production data) into the parametric generative model, wherein the input production data is optionally reflective of an operating state of the production system for which the prediction is being made. The input of production data used in this optional form of the first aspect of the invention may be in accordance with the input of data as already discussed above. Predictions can give an indication of how a hypothetical change (i.e. a proposed or theoretical change) in operation of the production system impacts on the performance of the production. Thus, proposed or theoretical predictions and / or developments in the performance of the production system can be determined. These predictions can then be used to determine an improved or optimised state for a production system (i.e. one in which production is optimised) since it can be determined what state of the production system provides improved production performance. In a fifth aspect of the invention, there is provided a method of determining predicted production performance of a production system, comprising: (i) providing a proposed change in an operating state of the production system; (ii) predicting, based on the proposed change, potential future production data in accordance with the relevant optional form of the first aspect of the invention as discussed above; and (iii) determining production performance of the production system based on the potential future production data. Optionally, steps (i) to (iii) are repeated, with each repetition of the steps (i) to (iii) being based on a different proposed change in operating state of the production system, until a desired improvement in production performance is determined. Further optionally, steps (i) to (iii) are repeated until optimised production performance is determined. In a sixth aspect of the invention, there is provided method of improving or optimising production performance of a production system, the method comprising: determining improved or optimised production performance for the production system in accordance with the optional form of the fifth aspect of the invention as discussed above; and changing the operation of the production system in accordance with the proposed change in operation of the production system that gives rise to the improved or optimised production performance. The production system may have a production output (e.g. total volumetric flow rate) that can be, or is, measured. Measurements of this output may contribute to recorded production data used to train the model. Thus thee output may form part of training data to train the generative model. The production data (recorded and / or synthetic) may comprise a quantity or rate of production—e.g. of an output the production system. The type of the production system referred to in the above aspects may be a physical resource production system. Accordingly, the production system may be a system that produces or is involved in the production of a (physical) product. The physical resource may be a natural resource, such as hydrocarbons (e.g. oil and gas), geological materials (e.g. metals, ores, precious stones), crops, livestock etc. Alternatively, the physical resource may be a man-made (synthetic) resource. Synthetically produced chemicals, e.g. bioethanol, may be the physical resource. The production system may be a hydrocarbon production well, a hydrocarbon production facility, a hydrocarbon production network, or a producing hydrocarbon reserve. It may produce one or more types of hydrocarbon. The production system may be an arable farm, a pastoral farm or an arable and pastoral farm. It may produce one or more crops. The production system may be a fish farm. It may produce fish. The production system may be a bioethanol production facility. It may produce bioethanol. The production system may be an energy production system. That is, the production system may produce (i.e. output) energy (or power)—e.g. electricity. For example, the production system may be or comprise a power plant. The power plant may be a hydrocarbon power plant, nuclear power plant, fuel cell based power plant, hydrogen fuel cell based power plant, wind based power plant, hydroelectric based power plant, solar panel based power plant, geothermal based power plant, waste heat power plant or biomass-fueled power plant. The production system may be a wind turbine or a wind farm. The production system may be a mine or mining facility. It may produce one or more minerals. The productions system may be a forestation system, wherein the resultant product is a forest or biomass. For example, the production system may be a eucalyptus forestation system. It may produce wood or trees. The production system may be a carbon sequestration facility, which may be a natural carbon sequestration facility (e.g. a forest) wherein the product is sequestered carbon dioxide. It may produce negative carbon. A production output may be the rate of output of what the production system produces. The type of the production system may be a production unit that is a single sub-system that produces a particular resource or product or that is involved in the production of that particular resource or product. For example, a production unit may constitute a hydrocarbon production well, a wind turbine or a solar panel, and may equally constitute a sub-system of these devices, e.g. a pump, valve, separator, generator, fuel cell, multiphase pump, compressor, turbine, distillation column, separator, motor, electrical motor, internal combustion engine, pipe, solar panel, fuel burner, wind turbine, rotor, water jet, heat exchanger, wing, conveyor belt, CNC machine, lathe, robot, robot arm, robot joint, furnace, wind tunnel, reservoir, joint, wheel, wheel bearings, tank, fish tank, scrubber, hydrocyclone, dryer, washing machine, washer, storage unit or storage facility, farm area, forest area, etc. However, in some embodiments, a production system may comprise a collection of production units, and optionally equipment and apparatus provided in support of the production units—e.g. collectively forming a production asset. For example, the production system may be a wind farm comprising a plurality of wind turbines or a hydrocarbon production system / facility comprising a plurality of production wells. A production system may be a processing unit or processing facility. In the context of a processing unit / facility, the resource or product that is deemed to be produced is the output of the process unit / facility. For example, a milk processing facility may have raw milk as an input and output, as a product, pasteurised and homogenised milk. The production system may be or comprise a chemical production facility or plants. The production system may be or comprise a manufacturing plant or facility, or a mining plat or facility. The production system may be or comprise a hydrocarbon production well, wherein the production well is associated with at least one control point, wherein the generative model is adapted to simulate one or more flow parameters, one or more well parameters and / or a status of the at least one control point associated with the production well, and wherein the synthetic production data comprises synthetic data relating to one or more flow parameters, one or more well parameters and / or an associated status of the status of the at least one control point associated with the production well. In the scenario where the production system is or comprises a production well, the at least one control point may comprise at least one of: a flow control valve; a pump; a compressor; a gas lift injector; an expansion devices; a choke control valve; gas lift valve settings or rates on wells or riser pipelines; ESP (Electric submersible pump) settings, effect, speed or pressure lift; down hole branch valve settings, down hole inflow control valve settings; or topside and subsea control settings on one or more: separators, compressors, pumps, scrubbers, condensers / coolers, heaters, stripper columns, mixers, splitters, chillers. In the scenario where the production system is or comprises a production well, the flow parameters may include one or more of: pressures; a flow rate (by volume, mass or flow speed / velocity), a gas flow rate (by volume, mass or flow speed / velocity), an oil flow rate (by volume, mass or flow speed / velocity), a water flow rate (by volume, mass or flow speed / velocity), a liquid flow rate (by volume, mass or flow speed / velocity), a hydrocarbon flow rate (by volume, mass or flow speed / velocity), a hydrogen sulphide fluid flow rate (by volume, mass or flow speed / velocity), a multiphase fluid flow rate (by volume, mass or flow speed / velocity), a flow rate that is the sum of one or more of any of the previous rates (by volume, mass or flow speed / velocity); an oil fraction, a gas fraction, a carbon dioxide fraction, a multiphase fluid fraction, a hydrogen sulphide fraction, a multiphase fluid fraction, temperatures, a ratio of gas to liquid, densities, viscosities, molar weights, pH, water cut (WC), productivity index (PI), Gas Oil Ratio (GOR), BHP and wellhead pressures, rates after topside separation, separator pressure, other line pressures, flow velocities, a well health indicator, a liquid loading risk indicator, a slug severity, or sand production. In the scenario where the production system is or comprises a production well, the well parameters may include one or more of: depth, length, number and type of joints, inclination, cross-sectional area (e.g. diameter or radius) within / of a production well, wellbore, well branch, pipe, pipeline or sections thereof; choke valve Cv-curve; choke valve discharge hole cross-sectional area; heat transfer coefficient (U-value); coefficients of friction; material types; isolation types; skin factors; and external temperature profiles. In the scenario where the production system is or comprises a production well, the optional forms of the method of the first aspect of the invention that comprise estimating or predicting production data as described above may comprise estimating or predicting a production flow rate. Alternatively, they may comprise estimating or predicting a flow rate (by volume, mass or flow speed / velocity), a gas flow rate (by volume, mass or flow speed / velocity), an oil flow rate (by volume, mass or flow speed / velocity), a water flow rate (by volume, mass or flow speed / velocity), a liquid flow rate (by volume, mass or flow speed / velocity), a hydrocarbon flow rate (by volume, mass or flow speed / velocity), a hydrogen sulphide fluid flow rate (by volume, mass or flow speed / velocity), a multiphase fluid flow rate (by volume, mass or flow speed / velocity), a flow rate that is the sum of one or more of any of the previous rates (by volume, mass or flow speed / velocity); an oil fraction, a gas fraction, a carbon dioxide fraction, a multiphase fluid fraction, a hydrogen sulphide fraction, a multiphase fluid fraction, temperatures, a ratio of gas to liquid, densities, viscosities, molar weights, pH, water cut (WC), productivity index (PI), Gas Oil Ratio (GOR), BHP and wellhead pressures, rates after topside separation, separator pressure, other line pressures, flow velocities, a well health indicator, a liquid loading risk indicator, a slug severity, or sand production, depth, length, number and type of joints, inclination, cross-sectional area (e.g. diameter or radius) within / of a production well, wellbore, well branch, pipe, pipeline or sections thereof; choke valve Cv-curve; choke valve discharge hole cross-sectional area; heat transfer coefficient (U-value); coefficients of friction; material types; isolation types; skin factors; and external temperature profiles. In the scenario where the production system is or comprises a production well, in the fifth aspect of the invention step (i) may comprise providing a proposed change in one or more flow parameters, one or more well parameters and / or a status of the at least one control point, step (ii) may comprise predicting, based on the proposed change, potential future production data relating to one or more flow parameters, one or more well parameters and / or an associated status of the at least one control point associated with the production well, and step (iii) may comprise determining predicted production performance of the production well based on the potential future production data relating one or more flow parameters, one or more well parameters and / or an associated status of the at least one control point. In some embodiments, the method of modelling the production system further comprises training the parametric generative model. The training may use recorded production data from the production system and / or from one or more further production systems (e.g. a plurality of production systems of the same type as the production system being modelled). The training may comprise minimising a loss function for the generative model. It may comprise updating parameters of the generative model based on the minimised loss function. In a seventh aspect of the invention, there is provided a method of training a parametric generative model that is adapted to simulate how production data associated with a production system is generated, the method comprising: (i) modelling in accordance with any of the preceding aspects of the invention, and optionally in the optional forms thereof; (ii) minimising a loss function for the generative model; and (iii) updating parameters of the generative model based on the minimised loss function. Step (ii) of minimising a loss function for the generative model may be based on the generated synthetic production data. In a scenario where the parametric generative model is (or comprises part of) a variational autoencoder comprising an encoder and a decoder, the encoder and the decoder being parameterized models, minimising a loss function may comprise minimising a first loss function for the encoder and minimising a second loss function for the decoder, and updating parameters of the generative model may comprise updating parameters of the encoder and the decoder based on the first and second loss functions respectively. During training, recorded production data associated with the production system may be input into the generative model prior to or as part of generation of the synthetic production data. In such embodiments, minimising a loss function may comprise minimising the loss function based on the recorded production data. Some or all of the recorded production data in this optional form of the invention may be unlabeled production data. Accordingly, minimising a loss function may comprise, at least in part, minimising an unsupervised loss function based on the unlabeled recorded production data. Alternatively, or additionally, some or all of the recorded production data may be labelled production data. Accordingly, minimising a loss function may comprise, at least in part, minimising a supervised loss function based on the labelled recorded production data. In some embodiments, the model is trained using semi-supervised learning. Unlabelled production data as used herein is used to refer to instances of data where some variables / parameters are available in the recorded production data (e.g. ‘x’ variable), but the variables / parameters of interest (e.g. ‘y’ variables) are not available. In the exemplary case of a production well, unlabeled production may arise for example when temperature and / or pressure data are available but correspondent flow rate data is not. In contrast, labelled production data is data that has been recorded and for which the variables / parameters of interest are available, in addition to other variables / parameters. Taking the exemplary case of a production well again, data may be considered labelled where data for flow rate, temperatures and pressures are all available. In some embodiments, the parametric generative model may be trained with recorded production data from a single production system. The synthetic production data may relate to an output (e.g. rate of production) of a single production system. However, in other embodiments, the parametric generative model is trained with recorded production data from a plurality of production systems. The parametric generative model may model all of the plurality of production systems. It may be adapted to simulate how production data associated with any or all of a plurality of production systems is generated. Training the model on multiple, similar production systems can advantageously allow for improved modelling performance and / or greater data efficiency. Some embodiments use transfer learning (e.g. multi- task learning) to model the plurality of productions systems jointly (rather than developing a separate model for each production system). Each production system of the plurality may be considered a different respective task in a multi-task learning process. Integrating semi-supervised generative modelling with transfer learning has been found to be especially powerful for efficiently modelling production systems. In some embodiments, the parametric generative model may have been trained on recorded production data from each of the plurality of production systems that is obtained from different combinations of sensor types for at least two of the production system. The generative model may be used to impute production data that is absent from the production data obtained for at least one of the plurality of production systems. The parametric generative model of any of the above aspects may comprise a set of first parameters representative of properties common to a plurality of production systems of a same type as the production system being modelled. For example, the generative model may comprise a first set of parameters representative of properties common to a plurality of production wells where the production system to be modelled is a production well. The first set of parameters representative of properties common to all of the plurality of production systems may permit the generative model to, fairly robustly and without requiring a significant retraining, model any of the production systems within the plurality of production systems, or any further production system of a same type, relatively accurately and successfully. The parametric generative model as discussed above may comprise a set of second parameters that are representative of properties that are specific to the production system that is being modelled. The second set of parameters may be reflective of behaviours, characteristics and traits unique to the production system to which it relates, and hence the generative model can model behaviours that are specific and idiosyncratic to an individual production system. In some embodiments, modelling the production system may comprise providing recorded production data from the production system as input to the parametric generative model, and using the provided data for determining the second set parameters. The generative model may be trained (or have been trained) using variational inference. The parametric generative model may, in some embodiments, be conditioned on a respective context of each of a plurality of production systems. The generative model may comprise a set of latent variables trained to represent one or more differences between respective production systems of a plurality of production systems. The parametric generative model may be configured to transform recorded production data from any production system of the plurality of production systems into a first latent space representative of properties common to the plurality of production systems (e.g. corresponding to the first set of parameters), and additionally into a second latent space representative of properties specific to the specific production system (e.g. corresponding to a respective second set of parameters). In a further aspect of the invention, there is provided a computer system configured to carry out the method of any of the above aspects, optionally in accordance with any optional form thereof. In yet a further aspect of the invention, there is provided a computer program product (e.g. a non-transitory computer-readable storage medium) comprising instructions for execution on a computer system, wherein the instructions, when executed, will configure the computer system to carry out a method as defined those method aspects of the invention and the optional forms thereof. BRIEF DESCRIPTION OF THE DRAWINGS Certain preferred embodiments of the invention will now be described, by way of example only, and with reference to the accompanying drawings, in which: Figure 1 is a schematic of a conventional production well; Figure 2 is a schematic representation of a generative hierarchical common cause (HCC) model; Figure 3 is a schematic representation of an inference model of the HCC model of Figure 2; Figure 4 is a schematic of a discriminative multi-task learning model; and Figures 5a & 5b depict respective graphs showing the few shot learning performance of the HCC model in a real and an artificial data test case respectively. DETAILED DESCRIPTION OF EMBODIMENTS Described below is an example of a parametric generative model that is for modelling a production system. The motivation and background to the model is first discussed, followed by a discussion of the application of the model in test cases that demonstrate its suitability for modelling production systems, more specifically production wells. The parametric generative model may be implemented on a computer system comprising one or more processors and memory storing software which, when executed by the one or more processors, causes the computer system to implement any one more of the methods described herein. The computer system may be configured to access recorded data from a set of one or more production systems (e.g. oil and gas wells). The data may be stored in a memory and / or received over a network. The model may be trained on data recorded from the production systems. It may be updated / retrained as more data becomes available over time. The computer system may use the model to generate synthetic data for one or more of the production systems, e.g. for soft sensing or “what-if” analysis. 1. Overview In this example it is demonstrated how the two learning paradigms of multi- task learning (MTL) and semi-supervised learning (SSL) are leveraged to obtain a data-efficient method for data-driven ‘soft sensing’ for production systems using a deep latent variable model (DLVM). Soft sensing comprises use of a model, in this case the DVLM, to operate on process measurements (x) to make timely inferences about key process variables (y). DVLM is a type of generative model for which conditional probability distributions are represented by deep neural networks. The generic model has a hierarchical structure and relies on the principle of common causes to explain the observations (x, y). Building on the framework of the variational auto-encoder, it is shown how to make inferences for latent variables and estimate the parameters of the model from data. The generative model has additional uses relevant / similar to soft sensing, which are not admitted by prior art discriminative models. For example, the synthetic production data generated by virtue of the generative DLVM model as discussed in this example can be used in what-if scenario analyses and can be used to impute missing production data. The data-driven soft sensing modelling as discussed in this example is associated with the following advantages: • Data-driven. The method has a low modelling cost. • Data efficient. The method exploits unlabelled data via semi-supervised learning (SSL) and information shared by multiple production systems via multi-task learning (MTL). • Parameter efficient. The nonlinear model is parameterized by neural networks for amortized inference. • Few-shot learning. The model can be calibrated to a new production system using few observations. • Data imputation. The method allows for inference of production data (including input variables) with missing values (i.e. values that were not recorded). The method and modelling regime underpinning this first example is based on a variational autoencoder (VAE). Variants of semi-supervised VAE have been disclosed in the prior art which are specialized to semi-supervised classification. Combination of SSL and MTL for both classification and regression problems has been disclosed in the prior art; however none of the prior art disclose the use a of a generative model for modelling a production system and which utilises a combination of SSL and MTL. 2. Soft-sensing problem to be solved Consider a set of M > 1 distinct, but related production systems indexed by i ∈ {1, ... ,M}. Each production system may be single unit such as a pump, valve, or solar panel, or a complex system (i.e. production asset), such as one or more distillation columns, production wells. The production systems are assumed to be related so that it is reasonable to expect, a priori to seeing any data, that transfer learning between systems is possible. It is assumed that the same explanatory variables x ∈ RDxand target variables y ∈ RDycan be considered for all production systems. In soft sensor applications, x typically represents cheap and abundant measurements from which we wish to infer a target variable y which is expensive or difficult to measure. For each system i, labelled data unlabelled data are available. For brevity it can be denoted that . Each dataset Di consists of independent and identically distributed observations drawn from a probability distribution pi over ^^^×^^ . The aim of the generative model is to efficiently learn the behaviour of the production systems, as captured by p1,...,pM, from the data collection D = {D1,...,DM}. In particular, to exploit that some structures or patterns that are common between production systems. Furthermore, the aim is to have the generative model fully utilize the information in D such that it can learn from unlabelled data. Invariably, distinct production systems differ in various ways, e.g., they may be of different design or construction, and they may be in different condition or operate under different conditions. The properties or conditions that make systems distinct are referred to as the context of a system. An observed context may be included in the dataset as an explanatory variable. An unobserved context must be treated as a latent variable to be inferred from data, and it is only possible to make such inferences by studying data from multiple, related systems. We assign the letter c to latent context variables. It is assumed there is some underlying process p at population level, such that pi = p(·|ci), where ci ∈ RKis a latent variable drawn from p(c). The variable ci represents unobserved properties (or an unobserved context) of system i. We can now restate our goal as learning from D a generative model ^^ which approximates p. 3. Notation Sets of variables are compactly represented by bold symbols, e.g. c = {ci}i. Double-indexed variables can be written as x = {xij}ijand xi= {xij}j, which permits x = {xi}ito also be written. When the indices are irrelevant to the discussion, x instead of xijmay be written. For some positive integer K, the K-vector of zeros is denoted by 0Kand the K × K identity matrix by IK. For a K-vector σ = , log σ denotes the element- wise logarithm, and diag(σ2) denotes the diagonal (K ×K)-matrix with diagonal elements ( The normal distribution is denoted by N(µ,Σ), where µ and Σ is the mean and covariance matrix, respectively. For a random variable X ∼ p(x), the expectation operator is denoted by E [X]. Sometimes EX[X] or Ep[X] are used to indicate that the expectation is of a specific random variable or probability distribution. 4. The generative model used The generative model of example 1 and that is used to simulate how production data relating to a production system is generated is a hierarchical common cause (HCC) model. An HCC model is a probabilistic model that is configured for generating observations (i.e. synthetic production data) for a sample of production systems. Observations are split semantically so that x labels cheap and frequent observations, and y labels expensive and occasional observations. The model has three levels of latent variables that explain the observations. At observation level, z represents a latent state which is (partly) observed by x and y. At system level, the latent variable c represents the context or specificity of a system. At the universal top-level, θ represents properties which are shared by all systems and observations. The hierarchical structure allows the model to capture variations at the level of systems and observations. A schematic representation of the HCC model of this example can be seen in Figure 2. Random variables in the model are encircled in Figure 2. A fully shaded circle indicates that the variable is an observed (i.e. recorded or measured variable). A non-shaded circle indicates that the variable is latent. A part shaded circle indicates that the variable comprises either latent or observed variables. The nested plates (rectangles) in Figure 2 group variables at different levels. The variables depicted in Figure 2 are discussed in greater detail below. The HCC model assumes the following generative process for a collection D of observations from M systems (where each system is considered a respective task for the multi-task learning). For each system i = ci∼ p(c) = N(0K,IK), zij ∼ p(z) = N(0D,ID), j = 1,...,Ni, (1) xij|zij,ci∼ pθ(x|zij,ci), j = 1,...,Ni, yij|zij,ci∼ pθ(y |zij,ci), j = 1,...,Ni. At the system level, latent context variables ci ∈ RKare generated from a common prior distribution p(c) = N(0K,IK). At the level of observation, latent variables zij ∈ RDare generated from a common prior p(z) = N(0D,ID). The standard normal priors on the latent variables allow for easy sampling and reparameterization, and is commonly used in variational autoencoders. Observations (xij,yij) are generated conditionally on the latent variables zijand ci. The conditional distributions, pθ(x|z,c) and pθ(y |z,c), are parameterized by the universal parameters θ, as discussed below. Several simplifying assumptions are made in the above model: 1) the dimensions of the latent variables, K and D, are fixed and considered to be choices of design; 2) θ is treated as a parameter vector to be estimated, not as a latent variable with a prior p(θ); 3) no hyper-priors are put on the parameters of p(z) and p(c). In a fuller Bayesian approach, it would be natural to remove assumption 2 and 3, and include a hyper-prior also for p(θ). The conditional distributions in equation (1), pθ(x|z,c) and pθ(y |z,c), are modelled by deep neural networks. For the continuous variable y, an appropriate likelihood function may be the multivariate normal pθ(y |z,c) = N(µy,Σy) where (µy,Σy) = fθ(z,c), and fθ is a neural network with parameters θ. The neural network provides an expressive family of conditional distributions for modelling y |x,c. The likelihood pθ(x|z,c) can be modelled similarly for continuous x. Note that the observations x and y are conditionally independent given the latent variables c and z. We can thus write pθ(x,y |z,c) = pθ(x|z,c)pθ(y |z,c). Our notation indicates that some elements of θ may be shared by the two distributions; for example, the two distributions can be modelled by a multiheaded neural network. The resulting model is a type of deep latent variable model. The joint distribution of all variables in equation (1) above factorises as follows pθ(x,y,z,c) =∏^ ^^^ ^(^^) ∏^ ^^^^^(^^^ ∨ ^^^, ^^)^^(^^^ ∨ ^^^, ^^)^(^^^)(2) In equation (2), bold symbols are used for sets of variables, e.g. x = {xij}ij = {xi}i. As indicated in Figure 2 by the partly shaded circle for y, both labelled and unlabelled data may be available for a production system. For system i, observed variables are grouped in {xli,yil,xui } and latent variables With this notation, xi = xl i ∪ xu i , where xl i and xu i are disjoint sets (likewise for yi). Observations for system i are denoted by Di = Dil ∪Diu, where Dil= {xil,yil} is labelled data and Diu= xui is unlabelled data. When considering all systems, the subscript i is dropped. For the model in equation (1), the marginal likelihood of a data collection D = {D1,...,DM} is obtained by marginalizing out all latent variables in equation (2) as shown in equation (3) below: The equation in (3) is intractable due to the integration over complex conditional distributions (parameterized by deep neural networks). Maximum likelihood estimation of θ is thus not possible. However, an approximate inference procedure can be used to enable the estimation of θ. Application of Bayes’ theorem to the model in equation (2) results in an alternative expression (shown in equation (4)) for the marginal likelihood to that shown in equation (3). Equation (4) includes the posterior distribution of the latent variables, pθ(yu,z,c|D). Since the model in equation (2) is tractable by design, it follows from the intractability of the marginal distribution in (3) that the posterior must be intractable. Exact inference of the latent variables is thus unachievable. This is a general issue with DLVMs and motivates the use of approximate inference methods such as variational inference. With variational inference, the posterior is approximated by a simpler distribution, sometimes called an inference model, which we denote by qϕ(yu,z,c|D). The variational parameters ϕ will be described shortly. The approximate posterior makes it possible to derive the lower bound on the marginal log-likelihood as shown in equation (5): In equation (5), the expectation is over values of latent variables {yu,z,c} drawn from qϕ. This bound is referred to as the evidence lower bound (ELBO). Maximization of the ELBO with respect to the parameters (θ,ϕ) achieves two goals: i) it maximizes the marginal likelihood, and ii) it minimizes the KL divergence between the approximation qϕ(yu,z,c|D) and the posterior pϕ(yu,z,c|D). Another advantage with this strategy is that it permits the use of stochastic gradient descent methods which can scale the learning process to large datasets. The specific choice of inference model removes some of the complicating dependencies in the posterior. To motivate the simplifications, we consider the following factorization of the posterior distribution shown in equation (6): pθ(yu,z,c|D) = pθ(z |D,yu,c)pθ(yu|D,c)pθ(c|D) (6) The first factor in equation (6) can be written as shown in equation (7): pθ(z |D,yu,c) = ∏ ^ ^^^∏ ^ ^^^ ^^(^^^ ∨ ^^^, ^^^, ^^). (7) From this expression, we see that the latent variable zij is inferred from xij, yij and ci. Evidently, inference of zij requires access to yij, which may be observed (yij ∈ yl) or inferred (yij∈ yu). The posterior factors in equation (7) are approximated by a variational distribution qϕz(z |x,y,c) with parameters ϕz. The second factor in equation (6) can be written as shown in equation (8): A variational distribution qϕy(y |x,c) with parameters ϕyis used to approximate the posterior factors in equation (8). This variational distribution resembles a discriminative model, where y is conditioned on an input x and context c. The third factor in equation (6) can be written as shown in equation (9): pθ(c|D) =∏^ ^^^^^(^^ ∨ ^^)(9) The latent variable ci is inferred from all observations in Di and can be thought of as a vector of summary statistics of the dataset. Conditioning the latent variable on the full dataset can be challenging for large datasets and it requires a model that can handle different dataset cardinalities. The inference is simplified by allowing an unconditional approximation of the form qϕci(ci), where ϕci are parameters related to variable ci. An argument for this simplification is that the latent variable is likely to change little when new datapoints are added to the dataset; i.e., as evidence is accumulated for the context variable ci, the variance of pθ(ci |Di) is expected to shrink. To summarize, the above discussion motivates how the inference model in equation (10) approximates the model of equation (6): where the variational parameters have been denoted by ϕ = {ϕz,ϕy,ϕc} and ϕc = {ϕc1,...,ϕcM}. The inference model of equation (10) is schematically shown in Figure 3. Again, random variables in the model are encircled in Figure 3. A fully shaded circle indicates that the variable is an observed (i.e. recorded or measured variable). A non-shaded circle indicates that the variable is latent. A part shaded circle indicates that the variable comprises either latent or observed variables. The nested plates (rectangles) in Figure 3 group variables at different levels. The variational distributions in equation (10) are chosen to be mean-field normal as shown in equations (11)-(13): where the symbols m and s have been used for the mean and standard deviation, respectively. The conditional distributions of y and z are computed as (my,logsy) = gϕy(x,c) and (mz,logsz) = gϕz(x,y,c), where gϕyand gϕzare neural networks. The inference networks amortize the cost of inference by utilizing a set of global variational parameters (ϕy and ϕz), instead of computing variational parameters per datapoint. In variational autoencoders, an inference network is usually called an encoder. The unconditional variational distribution for ciis simply computed as (mci,logsci) = ϕci. That is, for each task i, the distribution of the context variable ci∈ RKhas 2K parameters. With the inference model fully described as in equation (10), a leaner notation is afforded and subsequently variational distributions are referred to by qϕ(·), where the arguments indicate included factors and parameters. For example, it is possible to write qϕ(y,z |x,c) = qϕz(z |x,y,c)qϕy(y |x,c). With the inference model represented in equation (10), the ELBO cost function in equation (5) becomes: ^ ^ ^(^, ^)=∑^^^^^^(^, ^)(14) where the terms related to system I are collected in Notice that the expectations in ^^^(^, ^)are over the context variable ci of system i. The final term is the Kullback-Leibler divergence of the variational distribution qϕ(ci), which can be computed analytically when both qϕ(ci) and pθ(ci) are normal distributions. Optimization of the ELBO in equation (14) using a gradient-based method, requires computation of equation (15): Where the operator produces the gradient with respect to the parameters θ and ϕ. The gradient computation in equation (15) is problematic since the expectation on the right-hand side depends on ϕ. To circumvent the problem it is possible to reparameterise the latent variables, yu, z, and c. Using such a reparameterisation, a one-sample Monte-Carlo approximation of the ELBO is derived as shown in equation (16): Where~^^^(^, ^)is represented in equation (17): In equation (17), the context variable is sampled once (per system) by computing where εc ∼ N(0K,IK). The latent variables z and yuare sampled per data point using the inference networks: ^^∼ ^(0^, ^^), ~ ^ = ^^+ ^^⊙ ^^, Note that in the second summation in equation (17) both x and y are observed, and z is sampled by first evaluating (^^, ^^^^^) = The gradient estimator in equation (16) is known as the Stochastic Gradient Variational Bayes (SGVB) estimator. It is an unbiased estimator of the ELBO gradient ∇θ,ϕLθ,ϕ(D), and can be computed efficiently using back-propagation (due to the reparameterization). 5. Results of using the generative model To investigate the abilities of the HCC model discussed above, various experimental case studies were undertaken and the findings of these studies are summarised below. 5.1. Artificial dataset - Single-phase well flow To investigate the abilities of the HCC model of improving label prediction by exploiting grouped data and unlabelled data in a soft sensing setting, a case study on an artificial dataset was performed. The artificial data was generated using equations describing single-phase flow through a long pipe with a choke valve at the end, a hypothetical production system use case. The objective in this case is to predict the mass flow rate through the pipe given measurements of pressure and the percent-wise valve opening. Data were generated for 50 different systems. For each of the systems, 5 labelled data points and 100 additional unlabelled data points were generated. A test set consisting of 5000 data points, distributed equally among the systems, was also generated. Multi-task learning (MTL) and HCC model architectures were compared. The MTL model used is depicted in Figure 4 and is given as^^(^|^, ^)^(^).Random variables in the model are encircled in Figure 4. A fully shaded circle indicates that the variable is an observed (i.e. recorded or measured variable). A non-shaded circle indicates that the variable is latent. The nested plates (rectangles) in Figure 4 group variables at different levels. This MTL model is a discriminative model since^(^) is not modelled. The MTL model was trained using only labelled data, while the HCC models were trained on labelled data and increasing amounts of unlabelled data. The resulting test set root-mean-square-error (RMSE) for the different models are shown in Table 1. The average, the 10th, the 50th and the 90th percentile of the RMSE is reported. The lowest value of each error measure is set in boldface. The HCC model has lower RMSE than the MTL model, and the RMSE decreases as the number of available unlabelled samples increases. Table 1: Model Unlabeled Average P10 P50 P90 ratio MTL 0 0.091 0.078 0.086 0.111 HCC 1 0.083 0.066 0.077 0.104 HCC 2 0.082 0.067 0.077 0.102 HCC 4 0.080 0.065 0.075 0.099 HCC 10 0.076 0.063 0.074 0.089 HCC 20 0.073 0.061 0.071 0.086 5.2. Data-driven virtual flow metering To further investigate the abilities of the HCC model in a soft sensing setting, a case study on a real dataset from a production well was performed. The soft sensing case study used a dataset consisting of real measurements from 70 petroleum production wells (i.e. production systems) was used to investigate the abilities of the HCC model of learning from grouped and partly unlabelled data in a real-life setting. The objective was to predict the total volumetric flow rate given measurements of pressure, temperatures, mass fractions and choke valve opening. In practice, these measurements are often abundant while measurements of mass flow rate can be scarce, motivating the use of model structures (such as the HCC model) which can exploit unlabelled data, and learn simultaneously from data from multiple systems. The behaviour of wells can be expected to be qualitatively similar, but different when it comes to physical parameters, which motivates the use of MTL and HCC models where a well is considered to be a system. Just as in the artificial data case study, the models considered are the MTL model structure, as well as HCC models having access to increasing amounts of unlabelled data. The resulting models were evaluated on a test set consisting of the most recent measurements available for each production well. This method for selecting the test set was chosen to mimic how a mass flow prediction model (often called a virtual flow meter) would be used in practice Test set RMSE is shown in Table 2. The average, the 10th, the 50th and the 90th percentile of the RMSE is reported in Table 2. The lowest value of each error measure is in boldface. The differences between the MTL and HCC models as shown in Figure 2 are small. Table 2: Model Unlabeled Average P10 P50 P90 ratio MTL 0 0.168 0.041 0.108 0.368 HCC 1 0.173 0.054 0.110 0.444 HCC 2 0.188 0.049 0.126 0.384 HCC 3 0.183 0.047 0.105 0.386 HCC 4 0.178 0.045 0.124 0.397 HCC 5 0.178 0.053 0.106 0.404 5.3. Few-shot learning The few-shot learning capabilities of the HCC model on the artificial and real data case was investigated. The experiment involved: • For each case, five new production wells which were not included in the original training dataset were considered. • For each well variational inference of a new context variable c⋆ ∼ qϕc⋆(c⋆) by estimating ϕc⋆∈R2Kwas estimated while keeping the θ, ϕy, and ϕzparameters fixed. • The number of observations were varied and the effect of labelled and unlabelled data on the inference were considered. • For labelled data, {1,...,10} observations were considered. For unlabelled data, {1,5,10,20,50,100,200,500,1000} observations for the artificial data case were considered and {1,2,3,5,10,15,20,30,40} for the real data case were considered. • Some statistical uncertainty is present in the mean values due to the small sample of five wells / systems. The results of the experiment are shown in Figures 5a & 5b. The plot in Figure 5a shows the few shot learning performance on labelled data whilst the plot in Figure 5b shows the few shot learning performance on the artificial data. Each plot shows the mean absolute percentage error (MAPE) of five new wells / systems on the y-axis potted against the number of data points (on a logarithmic scale) on the x-axis are logarithmic for values above one. It can be seen from Figures 5a & 5b that for both the real and artificial data case, labelled observations are more valuable than unlabelled observations in terms of lowering the prediction error. This was expected. As is evident in the figure, the error decreases quickly and obtains a much lower error even with only 1- 2 observations. For the artificial data case, the context learned from unlabelled data improves the prediction performance considerably. However, as more unlabelled data becomes available the performance asymptotes to a higher error than the error obtained when labelled observations are available, even with fewer labelled observations 5.4. Data imputation To investigate the imputation capabilities of the HCC model, another case study on the petroleum dataset was performed. The situation considered was the case of a single sensor breaking, specifically the pressure sensor at the wellhead (see sensor 7 in Figure 1) and also in an alternative study the pressure sensor downstream of the choke valve (see Figure 1, sensor 13). This meant that all values in the test dataset being removed for the relevant sensor in the respective study. The objective is then to impute a good estimate of this value. The HCC model models the distribution p(x). When x is only partly observed (e.g. when a sensor breaks), which can be denoted by partitioning it into an observed part xoand an unobserved part xu, the model implicitly defines the conditional distribution p(xu|xo). This conditional distribution can be used to impute missing data. To perform imputation with the HCC model, the Metropolis-within- Gibbs (MwG) method was used. This method leverages the exact likelihood evaluation which VAE-based models provide, by using it in an MCMC method which in the limit gives an exact evaluation of conditional probabilities on the form p(xu|xo). The model trained with unlabelled ratio 1 seemed to be the most reliable, and was used for the study. The algorithm used for modelling was initialized with the training set mean for each well, and run for 400 iterations. The last half of these iterations were used to form an empirical estimate of p(xu|x0). Based on these last half of iterations, the population mean was used for imputation as a comparator to the MwG method as shown in table 3. To evaluate imputation performance, all of the measurements from a single sensor were removed (either the wellhead pressure sensor (PWH) or the pressure sensor downstream of the choke valve) and subsequently imputed by the MwG method. The test set performance is shown in Table 3, which shows the imputation results. The average, the 10th, the 50th and the 90th percentile of the mean MAPE is reported. There are significant differences between the measurements between the MwG method and the mean imputation. For most measurements however, the iterative MwG method significantly outperforms the mean imputation method. Table 3: In this way, it can be possible to use generative modelling as described herein to perform model inference by transfer learning between production systems of a common type (e.g. similar hydrocarbon wells) but having a different combination of sensors for recording production data, or where a sensor has failed on one of the production systems.

Claims

Claims 1. A method of modelling a production system, the method comprising: using a parametric generative model to generate synthetic production data relating to the production system, wherein the parametric generative model is adapted to simulate how production data associated with the production system is generated.

2. The method of claim 1, wherein: the parametric generative model is a trained parametric generative model; the method comprises inputting production data, recorded from the production system, into the trained parametric generative model; and the generated synthetic production data is conditional synthetic production data which depends on the recorded production data that is input to the trained parametric generative model.

3. The method of claim 2, wherein the trained parametric generative model is a variational autoencoder comprising an encoder and a decoder, the encoder and the decoder both being trained parametric models, and wherein the method comprises inputting the recorded production data to the encoder and producing the conditional synthetic production data from the decoder.

4. The method of any preceding claim, wherein the parametric generative model has been trained with recorded production data from each of a plurality of production systems of a same type as the production system being modelled.

5. The method of any preceding claim, wherein the parametric generative model comprises a set of first parameters representative of properties common to a plurality of production systems of a same type as the production system being modelled, and further comprises a set of second parameters representative of properties that are specific to the production system that is being modelled.

6. The method of claim 4 or 5, wherein the parametric generative model is adapted to simulate how production data associated with each of the plurality of production systems is generated.

7. The method of any of claims 4 to 6, wherein the parametric generative model is configured to transform recorded production data from the plurality of production systems into a first latent space representative of properties common to the plurality of production systems, and to transform the recorded production data into a second latent space representative of properties specific to a respective production system of the plurality of production systems.

8. The method of any preceding claim, wherein the production system is a physical resource production system.

9. The method of claim 8, wherein the production system is a hydrocarbon production well, a hydrocarbon production facility, a hydrocarbon production network, or a producing hydrocarbon reserve.

10. The method of claim 9, wherein the production system is a hydrocarbon production well, wherein the production well is associated with at least one control point, wherein the parametric generative model is adapted to simulate one or more flow parameters, or one or more well parameters, or a status of the at least one control point associated with the production well, and wherein the synthetic production data comprises synthetic data relating to one or more flow parameters, or one or more well parameters, or a status of the at least one control point associated with the production well.

11. The method of any preceding claim, wherein the generative model is a variational autoencoder comprising an encoder and a decoder, the encoder and the decoder both being parametric models; or wherein the generative model is a decoder of a variational autoencoder, the variational autoencoder comprising the decoder and an encoder, the encoder and the decoder both being parametric models.

12. The method of claim 1 or 2, wherein the generative model is a generative adversarial network or an autoregressive model.

13. The method of any preceding claim, wherein the method of modelling is a method of estimating production data and the generated synthetic production data comprises estimated production data.

14. The method of claim 13, wherein the method of estimating production data is a method of estimating production data from a period of time in the past and the estimated production data comprises estimated production data for the period of time in the past.

15. The method of claim 14, wherein some or all recorded production data associated with the production system from the period of time in the past is unavailable.

16. The method of claim 13, wherein the method of estimating production data is a method of estimating production data in real-time and the estimated production data comprises estimated real-time production data.

17. The method of claim 16, wherein some or all production data recorded in real-time is unavailable or is unavailable in real-time.

18. A method of analysing production performance of a production system, the method comprising estimating production data in accordance with any of claims 13 to 17; and analysing production performance of the production system based on the estimated production data.

19. The method of claim 1, wherein the synthetic production data relating to the production system is unconditional synthetic production data.

20. The method of any of claims 1 to 12, wherein the method of modelling is a method of predicting potential future production data and the generated synthetic production data comprises potential future production data.

21. The method of claim 20, comprising, prior to or as part of the step of using a parametric generative model to generate predicted potential future production data, inputting production data into the parametric generative model, wherein the inputproduction data is reflective of an operating state of the production system for which the prediction is being made.

22. A method of determining predicted production performance of a production system, comprising: (i) providing a proposed change in an operating state of the production system; (ii) predicting, based on the proposed change, potential future production data in accordance with claim 20 or 21; and (iii) determining production performance of the production system based on the potential future production data.

23. The method of claim 22, wherein steps (i) to (iii) are repeated, with each repetition of the steps (i) to (iii) being based on a different proposed change in operating state of the production system, until a desired improvement in or optimisation of production performance is determined.

24. A method of improving or optimising production performance of a production system, the method comprising: determining improved or optimised production performance for the production system in accordance with claim 23; and changing the operation of the production system in accordance with the proposed change in operating state of the production system that gives rise to the improved or optimised production performance.

25. A method of identifying an outlier or incorrect data in recorded production data associated with a production system, the method comprising: estimating production data associated with the production system in accordance with any of claims 13 to 17; comparing the estimated synthetic production data with production data recorded for a same period of time; identifying any deviation between the recorded production data and the estimated synthetic production data; and determining that the recorded production data comprises an outlier or incorrect data when the deviation exceeds a threshold quantity.

26. A method of cleaning recorded production data, the method comprising identifying an outlier or incorrect data in recorded data associated with a productionsystem in accordance with claim 25; and removing that production data from the recorded production data.

27. The method of any preceding claim, wherein the parametric generative model is configured to transform recorded production data into a latent space of lower dimension than an observable space of the recorded production data, and has been trained to learn a joint probability distribution over the latent space and the observable space.

28. The method of any preceding claim, wherein the parametric generative model is a deep generative model comprising one or more deep neural networks.

29. The method of any preceding claim, further comprising training the parametric generative model.

30. The method of claim 29, wherein the training uses recorded production data from the production system, or from one or more of a plurality of production systems of a same type as the modelled production system.

31. The method of claim 29 or 30, wherein the training comprises minimising a loss function for the generative model, and updating parameters of the generative model based on the minimised loss function.

32. The method of any of claims 29 to 31, comprising repeatedly training the parametric generative model at intervals in response to receiving more recorded production data.

33. The method of any of claims 29 to 32, wherein the training uses variational inference.

34. The method of any of claims 29 to 33, comprising training the parametric generative model on a combination of labelled and unlabelled production data.

35. The method of any of claims 29 to 34, comprising training the parametric generative model using semi-supervised learning.

36. A computer system configured to carry out the method of any preceding claim.

37. A computer program product comprising instructions for execution on a computer system, wherein the instructions, when executed, instruct the computer system to carry out the method of any of claims 1 to 35.