Method, computer program and associated prediction system for predicting clogging of a distillation column in a refinery
The prediction model constructed through machine learning and using sensor data preprocessing and synchronization technology, the error problem of the general prediction of the distillation tower liquid in the refinery is solved, and the prediction accuracy and production efficiency are improved.
Patent Information
- Application Number
- CN202180022149.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-18
- Filing Date
- 2021-03-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-03-17
AI Technical Summary
When the prior art predicts that the distillation tower liquid in the refinery is still significant, resulting in low production efficiency and unnecessary shutdown. The existing methods fail to effectively identify the operating parameters that trigger the liquid liquid.
The prediction model is constructed using machine learning methods, and by collecting, preprocessing, synchronizing and calculating derivatives from sensor data, forming a transformed data set, using a random forest model to predict the liquid state, and combining multiple standard methods to reduce errors.
Improve the detection efficiency of liquid pan prediction by 10% to 50%, realize earlier risk diagnosis and prevention, and reduce unnecessary downtime and production losses.
Smart Images

Figure CN115427905B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting flooding in at least one distillation column of a refinery.
[0002] The present invention also relates to a computer program comprising software instructions which, when executed by a computer, implement such a prediction method.
[0003] The present invention also relates to a system for predicting flooding in at least one distillation column of a refinery, the system being implemented by machine learning. Background Art
[0004] The documents US Patent No. 2018 / 275690A1 and US Patent No. 2002 / 116079A1 both relate to monitoring the general operation of a refinery; in particular, US Patent No. 2018 / 275690A1 discloses comparing the performance of a refinery or a refinery unit with the performance predicted by one or more process models in order to identify differences or defects in the operation.
[0005] US Patent No. 2008 / 082265A1 relates to systems and methods for facilitating monitoring and diagnosis, aiming to prevent any abnormalities in the coking heating device of a coking unit during the product refining process.
[0006] However, none of these documents disclose a solution specifically for monitoring the state of a distillation column in a refinery and particularly relating to predicting flooding in one or more distillation columns.
[0007] The work of Khairiyah et al. in the article titled "Development of neural Networks Models for a Crude Oil Distillation Column" (Jurnal Teknologi, January 20, 2012) discusses the development of artificial neural network models for optimizing the operation of a distillation column, but neither discloses nor suggests a specific solution for predicting flooding. Summary of the Invention
[0008] The present invention specifically relates to monitoring the state of a distillation column in a refinery and particularly relates to predicting flooding in the column.
[0009] Distillation is a process for separating or purifying different liquid substances from a mixture. When the vapor flow rate inside a distillation column exceeds a predetermined flow rate threshold, flooding of the distillation column occurs, and thus the liquid no longer flows through the distillation column. Such flooding is frequent, for example, for an atmospheric distillation column, it occurs 8 to 9 times per month on average, due to various reasons, such as excessive vapor flow rate and / or heating, etc.
[0010] Typically, due to an increase in the pressure difference greater than a predetermined threshold and a decrease in production performance, as well as a decrease in separation quality, such a flooding event is unfortunately detected after its occurrence. After flooding, the stabilization of the distillation column is slow, for example about 8 hours for an atmospheric distillation column, and is disadvantageous in terms of profitability.
[0011] To remedy this, it is necessary to predict such flooding in order to anticipate and, if possible, prevent it. For example, to prevent flooding in an atmospheric distillation column, it is necessary to predict the occurrence of such an event at least 20 minutes in advance.
[0012] So far, predictors have been developed for a first capacity predictor based on theoretical equations (such as equations developed from the work of H.Z. Kister et al. in the article "Predict Entrainment Flooding on Sieve and valve trays" (Chemical Engineering Progress, 1990, Vol. 86, No. 9, pp. 63 - 69)) and / or for a second predictor based on temperature and differential pressure analysis using conventional methods. However, such first and second predictors are not currently very effective: the number of false predictions (i.e., false positives predicting flooding when the operation remains nominal) or the number of non - predictions (i.e., false negatives where flooding occurs and has not been predicted in advance) is still too significant, and furthermore, false predictions trigger an expensive and useless slowdown of the production rate to avoid a flooding that is unlikely to occur, while the occurrence of undetected flooding slows down, or even simply stops the operation of the column during the time to re - stabilize it, which is also expensive.
[0013] The object of the present invention is to propose a method and a prediction system that, compared to current predictors, make it possible to reduce the prediction error corresponding to the sum of false predictions (i.e., false positives) and non - predictions (i.e., false negatives) and better identify the (one or more) operating parameters or (one or more) characteristics that may trigger flooding.
[0014] To this end, the subject of the present invention relates to a prediction method of the above - mentioned type, i.e., a method for predicting flooding in at least one distillation column of an oil refinery, the method being implemented by machine learning and comprising:
[0015] - a construction and training phase of a machine - learning model for predicting flooding, the machine - learning model being obtained from a data set pre - collected during a predetermined previous period and at least from a set of sensors of the oil refinery, each collected data being associated with sensor time data,
[0016] - The operating phase for predicting flooding, comprising the following steps:
[0017] - Collect a current data stream from the set of sensors of the refinery until a data buffer of a predetermined size is filled, and each data of the current data stream is further associated with sensor time data.
[0018] - Preprocess the data from the data buffer through a predetermined cleaning and classification, thereby delivering a current clean and classified data set.
[0019] - Synchronize the sensor time data associated with the data of the current collected data stream of the current clean and classified data set, thereby delivering a current clean and classified data set.
[0020] - Determine at least one value of a current variable representing at least one current performance of the at least one distillation column from the current clean, classified, and synchronized data set, and add the at least one value of the variable to the current clean, classified, and synchronized data set so as to form a current data set to be processed.
[0021] - Form a current transformed data set by calculating a predetermined derivative of the current data set to be processed.
[0022] - Predict the current state of the at least one distillation column by applying the learning model to the current transformed data set, the current state corresponding to a binary value representing the presence or absence of current pre-flooding.
[0023] Then, the prediction method according to the present invention is suitable for effectively improving the prediction of flooding in a distillation column in real time, no longer based on a single predictor, but on a multi-criteria method, which is processed via a series of steps executed according to an order specific to the present invention by machine learning so as to construct in real time a relevant data set for processing by a learning model, the learning model being pre-constructed and trained according to a previously collected data set.
[0024] In particular, the data synchronization step followed by the step of determining and adding at least one value of a current variable representing at least one current performance of the at least one column is used synergistically and according to this specific sequence of steps to enrich the data set collected in real time and refine the flooding prediction implemented by the learning model.
[0025] The position of the synchronization step in the specific sequence of steps according to the invention is particularly important as it allows any delay between a cause and its effect to be eliminated while following the kinetics of the distillation process implemented in the column under consideration. In the absence of such positioning of the synchronization step in the specific sequence of steps according to the invention, the data representing the cause is not correctly correlated with its consequences and the resulting calculated characteristics are insignificant and irrelevant.
[0026] In fact, according to the inventors' assessment, this method of predicting flooding in at least one distillation column of a refinery makes it possible to obtain a detection efficiency performance improvement of 10% to 50% compared to the above-mentioned first and second current predictors. Thus, the invention makes it possible to better diagnose in advance and specifically the risk of flooding in at least one distillation column of a refinery, which is not the purpose of the above-mentioned documents US Patent No. 2018 / 275690A1, US Patent No. 2002 / 116079A1, US Patent No. 2008 / 082265A1 and the work of Khairiyah et al. in the article entitled "Development of Neural Networks Models for a Crude Oil Distillation Column" (Jurnal Teknologi, January 20, 2012), none of which explicitly disclose the prediction of flooding in a distillation column nor the specific sequence and all steps of the process for predicting flooding in at least one distillation column of a refinery according to the invention.
[0027] According to other advantageous aspects of the invention, the prediction method comprises one or more of the following features taken individually or according to all technically possible combinations:
[0028] - The construction and training phase of the flooding prediction machine learning model comprises the following steps:
[0029] - Preprocessing the pre-collected data set by a predetermined cleaning and classification, which delivers preliminary clean and classified data,
[0030] - Synchronizing the sensor time data associated with the collected data of the preliminary clean and classified data set, thus delivering a preliminary clean, classified and synchronized data set,
[0031] - Determining at least one value of a variable representing at least one previous performance of the at least one distillation column from the preliminary clean, classified and synchronized data set and adding the at least one value of the variable to the preliminary clean, classified and synchronized data set in order to form a preliminary data set to be processed,
[0032] - Performing regression on the learning model by calculating and filtering a predetermined derivative of the preliminary data set to be processed, thereby forming two classes generated by the learning model, the learning model being associated with the normal operation of the at least one distillation column and the pre-flooding of the at least one distillation column respectively,
[0033] - Resampling the two classes generated by the learning model at a predetermined sampling rate,
[0034] - Using all the samples from the resampling step to determine, train, and validate the learning model;
[0035] - The synchronization implemented within the construction and training phase of the machine learning model for flood prediction and / or implemented within the operational phase of the flood prediction includes the application of a time lag determined according to the position of each sensor in the set of sensors;
[0036] - The learning model is a random forest model including a predetermined number of estimators and a maximum depth, the maximum depth being configured to expand each node of the random forest until all leaves are pure or until all leaves contain fewer than two samples;
[0037] - The current variable representing at least one current performance and / or the variable representing at least one previous performance is of a type belonging to a group including at least the following:
[0038] - The change in the total flow rate within the at least one distillation column,
[0039] - A flooding characteristic corresponding to the difference between a predetermined reflux flow rate set point and the reflux flow rate measured and collected during the data collection step associated with the operational phase or the construction phase,
[0040] - The upper recirculation index of the at least one distillation column corresponding to the ratio of the liquid-gas ratio to the reflux ratio of the extraction tray, the liquid-gas ratio and the reflux ratio being measured and collected during the data collection step associated with the operational phase or the construction phase,
[0041] - A risk index determined at least from the temperature and pressure data of the at least one distillation column measured and collected during the data collection step associated with the operational phase or the construction phase,
[0042] - A flooding index obtained from a predetermined theoretical equation and an associated binary index,
[0043] - A set of predetermined temperature deviations and ratios obtained from at least two sensors in the set of sensors located at different positions relative to the position of the at least one distillation column,
[0044] - Material balance,
[0045] - Enthalpy;
[0046] - The prediction step associates a probability with the binary value, and wherein after predicting the current state of the at least one distillation column, the method further includes generating an alert and returning the alert to at least one operator located within the refinery in the case of obtaining a binary value representing the presence of a current pre-flooding with an associated probability value greater than a predetermined probability threshold during the prediction step;
[0047] - The prediction method further includes the steps of: storing the data of the current data stream within a previously collected data set for subsequent iterations of the construction and training phases of the prediction flooding machine learning model, and updating the machine learning model for subsequent iterations of the prediction operation phase;
[0048] - The prediction method includes a compression step implemented at a predetermined compression during the collection of the current data stream, and a step of verifying the maintenance of the compression ratio at each subsequent collection step.
[0049] The present invention also relates to a computer program including software instructions which, when executed by a computer, implement the method as defined above for monitoring the execution of an application on an electronic computer.
[0050] The present invention also relates to a prediction system for at least one distillation column of a refinery implemented by machine learning, the system including at least one database, and the prediction system further includes:
[0051] - A unit for initially constructing and training a machine learning model for predicting flooding, the machine learning model obtained from a data set previously collected and stored in the database during a predetermined previous period and at least from a set of sensors of the refinery, each collected data being associated with sensor time data,
[0052] - A unit for predicting flooding, the unit for predicting flooding including:
[0053] - A data buffer of a predetermined size and a collection module configured to collect a current data stream from the set of sensors of the refinery until the data buffer of the predetermined size is filled, each data in the current data stream being further associated with sensor time data,
[0054] - A preprocessing module configured to preprocess the data from the data buffer by predetermined cleaning and classification, thereby delivering a current clean and classified data set,
[0055] - A synchronization module configured to synchronize the sensor time data associated with the data of the current collection data stream of the current clean and classified data set, thereby delivering the current clean and classified data set.
[0056] - A determination module configured to determine at least one current value of a variable representing at least one current performance of the at least one distillation column from the current clean, classified, and synchronized data set, and add the at least one value of the variable to the current clean, classified, and synchronized data set to form a current data set to be processed.
[0057] - A formation module configured to form a current transformed data set by calculating a predetermined derivative of the current data set to be processed.
[0058] - A prediction module configured to predict the current state of the at least one distillation column by applying the training model to the current transformed data set, the current state corresponding to a binary value representing the presence or absence of a current pre-flooding.
[0059] According to another advantageous aspect of the present invention, the prediction system is such that
[0060] - The collection module is located within the refinery itself, and
[0061] - The unit for initially constructing and training the general prediction machine learning model, the data buffer of the flooding prediction unit, the preprocessing module, the synchronization module, the determination module, the formation module, and the prediction module are outside the refinery and organized in a cloud computing manner.
[0062] The collection module is also adapted to directly load the data buffer, and the prediction system further includes a receiving module adapted to receive the prediction representing the current state of the at least one distillation column and return the information to at least one operator existing within the refinery via the return device of the prediction system. Description of the Drawings
[0063] The features and advantages of the present invention will become more apparent when reading the following description given only as a non-limiting example and referring to the drawings, in which:
[0064] Figure 1 is a schematic diagram of the material elements of the prediction system according to the present invention;
[0065] Figure 2 is a flowchart of the prediction method according to the present invention, which is implemented by the Figure 1 prediction system shown. Detailed implementation mode
[0066] Figure 1 An example of the architecture of the prediction system 10 according to the present invention is shown. According to such an architecture, the prediction system 10 according to the present invention is distributed over two different parts, namely within the refinery R itself and remotely, for example within a cloud computing CL system.
[0067] More precisely, the system 10 for predicting flooding in at least one distillation column of a refinery R implemented by machine learning includes at least one database BD or a set of databases BD organized in a cloud computing manner, a unit 12 for initially constructing and training a machine learning model for predicting flooding (which is obtained from a previously collected data set stored in the database BD), and a unit 14 for predicting flooding.
[0068] It should be noted that the unit 12 for initially constructing and training the machine learning model for predicting flooding is usually furthest from the refinery. For example, in a manner not shown, it is integrated into the personal computer of the refinery operator, who is mobile if appropriate and able to move inside and outside the refinery, and such a unit 12 is not necessarily directly integrated into the cloud computing, but can only communicate with the said cloud computing, for example.
[0069] As Figure 1 shown, the flooding prediction unit 14 consists of two parts (part 14 within the refinery R itself A and part 14 distributed over a set of servers organized in a cloud computing CL manner B ).
[0070] More precisely, the database BD has been pre-constructed by storing a set of data previously collected and obtained by means of a set of sensors C1 to C N (where N is an integer greater than or equal to 1) distributed within the refinery R, in particular about a thousand sensors for each distillation column of the refinery R, where, for example, a temperature sensor is in contact with the wall of the considered distillation column and a pressure sensor is in contact with the fluid flowing through the pipe of the same column under consideration. For example, sensors C1 to C N measure various types of data within the considered distillation column, namely temperature, pressure, flow rate, valve opening, etc. These data have been archived in the database for many years.
[0071] According to a specific aspect, the unit 12 for initially constructing and training the machine learning model for predicting flooding is adapted to extract from the database a set of data previously collected during a predetermined previous period and sampled at a predetermined sampling rate of, for example, one minute (i.e., each data of the same data type is associated with sensor time data and is spaced one minute from the previous and the next data).
[0072] Thereafter, reference will be made toFigure 2 Describe the detailed operation of such a construction and training unit 12.
[0073] The flooding prediction unit composed of two parts 14 A and 14 B more precisely includes a collection module 16 in part 14 located within the refinery R itself. The collection module 16 is configured to collect the current data stream from a set of sensors C1 to C of the refinery A until a data buffer (not shown) of a predetermined size is filled, and each data in the current data stream is further associated with sensor time data. N According to a first variant, the data buffer is filled within the refinery R itself and then, once the data buffer is filled, is sent via the transceiver module 18 of the refinery R to a receiver module (not shown) of a set of servers organized in the manner of cloud computing CL.
[0074] As an alternative, the buffer is directly located on one of the servers organized in the manner of cloud computing CL, and the collection module 16 is configured to fill this server in real time via the transceiver module 18 of the refinery R.
[0075] According to a particular aspect, the transceiver module 18 of the refinery R is dedicated only to the prediction system 10 according to the present invention. In this case, the prediction system 10 according to the present invention includes such a dedicated transceiver module 18, and according to one variant (not shown), such a transceiver module 18 is also directly integrated within part 14 of the flooding prediction unit
[0076] within. A
[0077] According to a particular aspect, such a collection module is also adapted to pre-feed the database BD.
[0078] The flooding prediction unit also includes a preprocessing module 20 in part 14 located within the cloud computing CL B which is configured to preprocess the data of the data buffer by a predetermined cleaning and classification. The preprocessing module 20 delivers the current clean and classified data set. In other words, the preprocessing module 20 has an input connected to the data buffer.
[0079] The flooding prediction unit also includes a synchronization module 22 in part 14 located within the cloud computing CL B which is configured to synchronize the sensor time data associated with the current data stream of the current clean and classified data set. The synchronization module 22 delivers a common clean, classified and synchronized data set. In other words, the synchronization module 22 has an input connected to the output of the preprocessing module 20.
[0080] The flooding prediction unit in part 14 located within the cloud computing CL B also includes a determination module 24, which is configured to determine at least one current value of a variable representing at least one current performance of the at least one distillation column from a current clean, classified, and synchronized data set, and to add the at least one value of the variable to the current clean, classified, and synchronized data set to form a current data set to be processed. The determination module 24 has an input connected to the output of the synchronization module 22.
[0081] The flooding prediction unit in part 14 located within the cloud computing CL B also includes a formation module 26, which is configured to form a current transformed data set by calculating a predetermined derivative of the data set to be processed. In other words, the formation module 26 has an input connected to the output of the determination module 24.
[0082] The flooding prediction unit in part 14 located within the cloud computing CL B also includes a prediction module 28, which is configured to predict the current state of the at least one distillation column by applying a learning model to the current transformed data set, where the current state corresponds to a binary value indicating the presence or absence of current pre-flooding. In other words, the prediction module 28 has an input connected to the output of the formation module 26.
[0083] As an optional supplement, the prediction module 28 includes a calculation tool 30 and a module 32. The calculation tool 30 is configured to calculate the probability of flooding (i.e., the confidence index) and compare the probability of flooding (i.e., the confidence index) with a predetermined probability threshold to obtain a binary value indicating the presence or absence of current pre-flooding. The module 32 is used to generate an alarm during the prediction step when a binary value indicating the presence of current pre-flooding is obtained and the associated probability value is greater than the predetermined probability threshold.
[0084] According to another optional specific aspect, the prediction system 10 further includes a receiving module, such as Figure 1 the transceiver module 18 shown, which is adapted to receive information representing the prediction of the current state of the at least one distillation column and is adapted to be returned to at least one operator present within the refinery via the return device E of the prediction system 10.
[0085] In particular, such representative information directly corresponds, for example, to the alarm generated by the optional alarm generation module 32.
[0086] According to a specific aspect, part 14 located within the cloud computing CL BIt is also configured to store the data of the current data stream within a previously collected data set of the database BD for subsequent iterations of the construction and training phase of a machine learning model for predicting flooding, as implemented by the construction and training unit 12, and to update the machine learning model for subsequent iterations of the operational prediction phase.
[0087] In Figure 1 the example shown, the portion 14 located within the cloud computing CL of the flooding prediction unit B includes one or more information processing units 34, which are formed, for example, by a memory 36 associated with a processor 38 such as a CPU (Central Processing Unit).
[0088] In Figure 1 the example shown, the preprocessing module 20, the synchronization module 22, the determination module 24, the formation module 26, the prediction module 28, and optionally the calculation module 30 and the generation module 32 are all implemented in the form of software executable by the processor 38.
[0089] Then, the memory 36 of the information processing unit 34 is adapted to store preprocessing software that is configured to preprocess the data of a data buffer (sent by the transceiver module 18 of the refinery R or, according to a variant (not shown), the buffer is directly stored within the memory 36) by means of a predefined cleaning and classification, so as to deliver a current clean and classified data set. The memory 36 of the information processing unit 34 is also adapted to store synchronization software, determination software, formation software, prediction elements, the synchronization software being configured to synchronize sensor time data associated with the data of the current collected data stream of the current clean and classified data set, so as to deliver a current clean and classified data set, the determination software being configured to determine at least one current value of a variable representing at least one current performance of the at least one distillation column from the current clean, classified and synchronized data set, and to add the at least one value of the variable to the current clean, classified and synchronized data set in order to form a current data set to be processed, the formation software being configured to form a current transformed data set by calculating a predefined derivative of the current data set to be processed, the prediction elements being configured to predict the current state of the at least one distillation column by applying the learning model to the current transformed data set, the current state corresponding to a binary value representing the presence or absence of current pre-flooding. Optionally, the memory 36 of the information processing unit also includes software for calculating a probability (i.e., a confidence index) associated with the binary value representing the presence or absence of current pre-flooding and for generating an alarm in the case of obtaining a binary value representing the presence of current pre-flooding during the prediction step, where the associated probability value is greater than a predefined probability threshold.
[0090] Then, the processor 38 is adapted to run in series preprocessing software, synchronization software, determination software, formation software, prediction software, and optionally calculation software and alert generation software.
[0091] In a variant (not shown), the preprocessing module 20, the synchronization module 22, the determination module 24, the formation module 26, the prediction module 28, and optionally the calculation module 30 and the generation module 32 are all implemented in the form of programmable logic components (such as FPGAs (Field Programmable Gate Arrays)) or in the form of application specific integrated circuits (such as ASICs (Application Specific Integrated Circuits)).
[0092] When at least a part of the prediction system 10 is implemented in the form of one or more software programs (i.e., in the form of computer programs), at least a part of the prediction system 10 is also adapted to be recorded on a computer-readable medium (not shown). The computer-readable medium is, for example, a medium that is adapted to store electronic instructions and is coupled to the bus of a computer system. By way of example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (e.g., EPROM, EEPROM, flash memory, NVRAM). Then a computer program including software instructions is stored on the readable medium.
[0093] Now, the operation of the prediction system 10 will be explained with the aid of Figure 2 explaining the operation of the prediction system 10, Figure 2 a flowchart of a process 38 for predicting flooding in at least one distillation column of a refinery, implemented by machine learning, is shown.
[0094] Such a method 38 includes two different phases 40 and 42, namely, a phase A of constructing and training a machine learning model M for predicting flooding and an operational phase B of predicting flooding, the machine learning model M being obtained from a data set pre-collected during a predetermined previous period and at least from a set of sensors of a refinery, each collected data being associated with sensor time data, phase A being implemented by a unit 12 for initially constructing and training a machine learning model for predicting flooding, and the operational phase B being implemented in real time by a flooding prediction unit 14.
[0095] The construction and training phase 40 of the machine learning model M must therefore be implemented before the flooding prediction phase B, since phase A feeds the learning model M into the flooding phase implemented in real time.
[0096] More precisely, the construction and training phase 40 of the machine learning model M for predicting flooding includes a series of steps in a specific order according to the invention, and includes a first step 44 of preprocessing PREP-P-P the pre-collected and particularly stored data set in the database BD by a predetermined cleaning and classification so as to provide a preliminary clean and classified data set.
[0097] Specifically, such preprocessing specifically includes:
[0098] - Using the pandas library to convert the format of the previously collected data, such as from an Excel spreadsheet format, into a data frame according to a format suitable for the Python language, and then
[0099] - In this converted form, filtering the collected values by identifying and then removing outliers, as outliers are duplicate or redundant or even constant and thus irrelevant to determining flooding, which reduces the amount of data to be processed by 30% to 40%, and then
[0100] - Cleaning and classifying the filtered data by successively implementing the following sub-steps:
[0101] - When the learning model expects numerical values, first replace strings with the item NaN (not a number), and then
[0102] - Classifying the data from the first replacement into five categories, namely:
[0103] - The data category associated with the non-operating state,
[0104] - The data category associated with the pre-flooding state, which groups all data collected, for example, during a 60-minute period before flooding,
[0105] - The data category associated with the post-flooding state, which groups all data collected, for example, within eight hours after flooding,
[0106] - The data category associated with the flooding state, which groups all data collected during flooding,
[0107] - The category including all remaining data that does not belong to any of the four previous categories and thus represents the normal operation of the tower,
[0108] - In each of the five categories, perform a second replacement of NaN items by forward filling, where the missing numerical value NaN is filled from the corresponding value in the previous row, and
[0109] - Reset columns that contain only zero values,
[0110] - Memorize the cleaned and classified data frame using a suitable storage format that reduces the size of the collected data, for example, by using the Pickle tool in Python, which is suitable for implementing the binary protocol for serializing and deserializing Python object structures.
[0111] More precisely, for example, in the presence of collected data indicating that the distillation column under consideration is in working order (not shut down) and based on measurements of the controlled flow rate level of the three-phase separator of the column condenser, a flooding state is detected. For example, the flooding state is detected in the presence of three conditions, namely:
[0112] - The distillation column under consideration is in working order (i.e., effective operation), and
[0113] - The difference between the measured value collected from the flow rate level and the setpoint value of the flow rate level is greater than 10, and
[0114] - The output value of the flow rate level controller of the column under consideration is greater than 70.
[0115] When the three conditions are met, flooding is detected, and the sensor time data associated with the measured value of the flow rate level is used for classification:
[0116] - All collected data associated with the same sensor time data is grouped into a data category associated with the flooding state,
[0117] - All collected data associated with sensor time data, for example, up to 60 minutes before, is grouped together into a data category associated with the pre-flooding state, where the sensor time data is associated with the measured value of the flow rate level,
[0118] - All collected data associated with sensor time data, for example, up to eight hours after, is grouped together into a data category associated with the post-flooding state, where the sensor time data is associated with the measured value of the flow rate level.
[0119] According to a specific practical aspect, the classification includes assigning a value representing one of the above five categories (i.e., the categories "not working (i.e., shut down)", "pre-flooding", "flooding", "post-flooding", "not relevant") to a variable (e.g., called FI) representing the category of each collected data.
[0120] The construction and training phase 40 of the machine learning model M for predicting flooding includes a second synchronization step SYNC 46 of the sensor time data associated with the collected data in the preliminary clean and classified data set, which delivers a preliminary clean, classified, and synchronized data set.
[0121] In particular, the synchronization 46 implemented within the construction and training phase 40 of the machine learning model M for predicting flooding includes the application 47 of (one or more) time lags TL determined according to the position of each sensor in the set of sensors C1 to C N in the set.
[0122] More precisely, such a time lag corresponds to the delay in the response of the distillation process to changing conditions, such as the delay caused by a change in the feedstock state at a specific location in the distillation column. For example, the time lag to be applied to the data collected by sensor C1 depends on the distance between sensor C1 and the location of the flooding point in the distillation column, which is known and constant for a given distillation column and application. Such a (one or more) time lag is automatically determined based on the knowledge of the distillation process implemented within the considered distillation column and is confirmed by a mutual information method based on the work of O. Ludwig et al. in the article titled "Applications of information theory, genetic algorithms, and neural models to predict oil flow" (CNSNS 14 (2009) 2870 - 2885).
[0123] Such synchronization 46 specifically consists in retrieving the automatically determined time lag value TL and then applying 47 said time lag value TL to the sensor time data associated with the collected data.
[0124] According to the previous clean, classified, and synchronized data set, the construction and training phase 40 of the flooding prediction machine learning model M includes a third step 48 for determining at least one value of a variable DET - EF - P representing at least one preliminary performance of the at least one distillation column and adding at least one value of said variable to the preliminary clean, classified, and synchronized data set in order to form a preliminary data set to be processed.
[0125] More precisely, the variable representing at least one previous performance is of a type belonging to a group including at least the following:
[0126] - The variation in the total flow rate within the at least one distillation column,
[0127] - The flooding characteristic corresponding to the difference between the measured value of the flow rate level collected during the data collection step associated with the operation phase or the construction phase and the setpoint value of the controlled flow rate level of the condenser column three - phase separator,
[0128] - The upper recirculation index of the at least one distillation column corresponding to the ratio of the liquid - to - gas ratio to the reflux ratio of the extraction tray, where the liquid - to - gas ratio and the reflux ratio are measured and collected during the data collection step associated with the operation phase or the construction phase.
[0129] - A risk metric determined at least from the temperature and pressure data of the at least one distillation column measured and collected during the data collection step associated with the operation phase or the construction phase,
[0130] - A flooding metric obtained from a predetermined theoretical equation (such as the above-mentioned first capacity predictor) and an associated binary metric,
[0131] - A set of predetermined temperature differences and ratios obtained from at least two sensors in the set of sensors located at different positions relative to the position of the at least one distillation column,
[0132] - Material balance,
[0133] - Enthalpy.
[0134] In other words, such variables representing at least one previous performance or preferably all of the above variables are calculated for each sensor time data and based on the data collected at the moment associated with the considered sensor time data using a predetermined customized engineering equation, which is specific to each type of distillation column and relevant in the industrial field of determining flooding in refinery distillation columns. Thus, each category of the preliminary clean, classified, and synchronized dataset is enriched with variables representing performance to facilitate the modeling of flooding.
[0135] It should be noted that, according to the present invention, such an enrichment step is specifically implemented after synchronization, which enables the use of the data collected with the same sensor time data after synchronization in each engineering equation to obtain one of the above variables and avoid bias when calculating the enriched variables representing at least one previous performance of the distillation column.
[0136] The construction and training phase 40 of the machine learning model M for predicting flooding further includes a fourth regression step REG50 of the learning model M, which forms two categories generated by the learning model by calculating and filtering predetermined derivatives of the preliminary dataset to be processed, and the two categories are respectively associated with the normal operation of the at least one distillation column and the pre-flooding of the at least one distillation column.
[0137] More precisely, such a calculation and filtering of a predetermined derivative of a preliminary data set to be processed according to the present invention includes determining a gradient calculated using exact central differences of second order applied to an internal subset of the preliminary data set to be processed (in other words, applied to predetermined internal points of the set) and exact first or second order one-sided differences (backward or forward) of data of the preliminary data set to be processed located outside the internal subset, such that the resulting gradient has a shape similar to the shape of the preliminary data set to be processed and used as input. It should be noted that such a calculation is not applied to binary metrics or variables representing classes associated with the first capacity predictor (e.g., called FI).
[0138] In other words, according to step 50, the five classes obtained and enriched up to the previous step 48 are reduced to two unique result classes Fi, each class being associated with a different binary value, i.e., for example, for the class associated with normal operation, Fi = 0, and for the class associated with pre-flooding, Fi = 1.
[0139] The construction and training phase 40 of the machine learning model M for predicting flooding also includes a fifth step 52 of resampling RS the two result classes FI respectively associated with the normal operation (FI = 0) and pre-flooding (FI = 1) of the learning model M at a predetermined sampling rate.
[0140] In fact, the two resulting classes are unbalanced in terms of size, the size of the class associated with normal operation being much larger than the size of the class associated with pre-flooding, since the frequency of flooding is, for example, eight times per month on average. For this purpose, for example, during step 52, multiple sampling ratios between the two resulting classes are tested in order to provide optimal results, such as ratios 10:1, 5:1, 5:5, and 5:10 in the case where the class associated with normal operation has ten times the samples of the class associated with pre-flooding. Preferably, during the resampling step 52, according to the present invention, the ratio 5:5 or the class associated with normal operation has as many samples as the class associated with pre-flooding.
[0141] The construction and training phase 40 of the machine learning model M for predicting flooding also includes a sixth step T54 of determining, training, and validating the learning model M using all the samples from the resampling step.
[0142] According to a particular aspect, during step 54, a cross-validation method is used to perform the determination of the learning model by specifically dividing all the samples from the resampling step into two non-overlapping subsets, one dedicated to training and the other dedicated to validating the learning model M. The subset dedicated to training is further subdivided into a predetermined number of non-overlapping cross-validation subsets and processed, for example, by means of a sliding window technique for time data series, specifically using previous sample steps to predict subsequent sample steps by time lag.
[0143] According to the present invention, the determination of the most effective learning model M for predicting past flooding associated with a pre-collected data set is performed within a list of predetermined model types having an interpretability degree greater than a desired and predetermined interpretability threshold. Such a list includes, for example, the following types of models: logistic regression, decision tree, random forest, artificial neural network, and support vector machine, among others.
[0144] The performance of each model in the list is measured using the area under the receiver operating characteristic (ROC) curve representing the performance of the classification model for all classification thresholds, and the true positive (actual flooding) rate is plotted against the false positive (false flooding) rate.
[0145] For example, for an atmospheric distillation column at the Donges refinery in France, for example, the best-performing learning model M is a random forest model having a predetermined number of estimators and a maximum depth configured to extend each node of the random forest until all leaves are pure or until all leaves contain fewer than two samples.
[0146] Thus, through all the above steps 44 to 54, the construction and training phase of the machine learning model makes it possible to determine and train the most effective learning model M for real-time prediction of the "pre-flooding" situation, which is the originality of the present invention, which predicts the "pre-flooding" phenomenon approximately 60 minutes before flooding, rather than the actual flooding that no longer allows the operator to take measures to reverse the process and prevent flooding.
[0147] At Figure 2 this point, the operation phase B42 of predicting flooding, as implemented in real time by the flooding prediction unit 14, then includes the following steps implemented in real time.
[0148] According to the first step 56, the collection module 16 of the flooding prediction unit 14 collects data from the current data stream DC of COLLECT_DC until a data buffer of a predetermined size is filled, and each data of the current data stream is also associated with sensor time data.
[0149] In particular, according to an optional supplementary aspect, such a collection COLLECT_DC includes a compression step 58COMP with a predetermined compression ratio of, for example, one data per minute with a priority memory, and a step 60 for verifying that the compression ratio is maintained at each subsequent collection step 56. Such compression makes it possible to maintain the quality of the collected data required for the subsequent effective training of the learning model M.
[0150] In particular, according to another optional supplementary aspect, after the collection 56, there is a step 62 which stores S the data of the current data stream DC within a previously collected data set in the database BD for subsequent iterations of the construction and training phase 40 of the machine learning model M for predicting flooding, and updates the machine learning model for subsequent iterations of the prediction operation phase 42. The construction and training phase 40 of the machine learning model M for predicting flooding is iterated, for example, after a predetermined number of actual flooding events for updating.
[0151] Then, according to step 64, the operation phase B42 of predicting flooding, as implemented in real time by the flooding prediction unit 14, includes preprocessing PREP-P-C the data from the data buffer by a predetermined cleaning and classification, thereby delivering a current clean and classified data set similar to that implemented during the construction and training phase 40 with the pre-collected data with the machine learning model M.
[0152] The synchronization step 66SYNC is also carried out after the preprocessing step 64 implemented during the construction and training phase 40 of the machine learning model M with the pre-collected data, but this time by applying such synchronization to the sensor time data associated with the data of the current collection data stream of the current clean and classified data set, which delivers a current clean, classified and synchronized data set.
[0153] In the same way as carried out during the construction and training phase 40, such a synchronization 66 includes the application 68 of a time lag TL determined according to the position of each sensor in the set of sensors C1 to C N in the sensor.
[0154] Then, according to step 70, the determination DET-EF-C of at least one current value of a variable representing at least one current performance of at least one distillation column is carried out according to the current clean, classified and synchronized data set, and is added to the current clean, classified and synchronized data set in order to form a current data set to be processed.
[0155] According to step 72, the formation of the current transformed data set is carried out by calculating a predetermined derivative DERIV of the current data set to be processed.
[0156] Finally, according to step 74, the current state of the at least one distillation column is predicted by applying the learning model M to the current transformed data set, the current state corresponding to a binary value representing the presence or absence of a current pre-flooding.
[0157] In particular, according to an optional supplementary aspect, the prediction step 74 determines a probability PROB and compares the probability PROB with a predetermined probability threshold during step 76 in order to obtain a binary value representing the presence or absence of a current pre-flooding.
[0158] After predicting the current state of the at least one distillation column in 74, the operating phase B42 of predicting flooding, as implemented in real time by the flooding prediction unit 14, further includes step 78 of generating and in particular returning an alert ALERT to at least one operator located within the refinery R via a screen E in the case of obtaining, during the prediction step 74, a binary value representing the presence of a current pre-flooding with an associated probability value greater than the predetermined probability threshold.
[0159] According to the practical aspect of real-time processing, once a prediction is obtained, the oldest data collected in the buffer is deleted so that data can be collected after the most recent data that can be collected in the buffer, and then steps 64 to 68 are repeated, and so on.
[0160] In other words, during the operating phase B42 of predicting flooding as implemented in real time by the flooding prediction unit 14, the data of the current data stream is processed in a manner similar to the manner in which the pre-collected data was utilized during the construction and training phase 40 of the machine learning model M, such that the learning model M is equally effective by using the current data set to be processed corresponding to the current clean, classified, synchronized and rich data set.
[0161] Therefore, it should be understood that the method for predicting flooding in at least one distillation column of a refinery according to the present invention is particularly useful for the real-time and pre-diagnosis of the risk of flooding in the distillation columns of a refinery. Using this method, an early alert can be sent to the operator of the distillation column of the refinery because the pre-flooding of the distillation column of the refinery is detected in this way rather than the flooding, which enables the operator to react before the flooding occurs and is a source of leakage or at least efficiency loss of the distillation column, and enables the downtime of the distillation column that may be subject to flooding to be reduced and the safety of the associated distillation process to be improved.
[0162] Compared with U.S. Patent No. 2018 / 275690A1, which particularly discloses comparing the performance of a refinery or a refinery unit with the performance predicted by one or more process models in order to identify differences or defects in operation, the present invention proposes a solution that is prior in time to the presence of a fault corresponding to flooding in a distillation column of a refinery.
[0163] Therefore, the solution according to the present invention makes it possible to avoid a loss of performance which, according to the document US Patent No. 2018 / 275690A1, is necessary for detecting an overall failure of an oil refinery in the absence of both precisely locating the cause of such a failure and locally detecting a flooding in a distillation column of the oil refinery and even less pre-flooding in the distillation column of the oil refinery.
Claims
1. A method (38) for predicting flooding in at least one distillation column of a refinery implemented by machine learning, the method comprising: - A construction and training phase (40) of a machine learning model for predicting flooding, the machine learning model obtained from a dataset pre-collected during a predetermined previous period and at least from a set of sensors of the refinery, each collected data being associated with sensor time data, - An operation phase (42) for predicting flooding, including the steps of: - Collecting (56) a current data stream from the set of sensors of the refinery until a data buffer of a predetermined size is filled, each data of the current data stream being further associated with sensor time data, - Preprocessing (64) the data from the data buffer by a predetermined cleaning and classification, delivering a current clean and classified dataset, - Synchronizing (66) the sensor time data associated with the data of the current collected data stream of the current clean and classified dataset, delivering a current clean, classified and synchronized dataset, - Determining (70) at least one current value of a variable representing at least one current performance of the at least one distillation column from the current clean, classified and synchronized dataset, and adding the at least one value of the variable to the current clean, classified and synchronized dataset to form a current dataset to be processed, - Forming (72) a current transformed dataset by calculating a predetermined derivative of the current dataset to be processed, - Predicting (74) the current state of the at least one distillation column by applying the learning model to the current transformed dataset, the current state corresponding to a binary value indicating the presence or absence of current pre-flooding.
2. The method (38) according to claim 1, wherein the construction and training phase (40) of constructing and training the machine learning model for predicting flooding includes the steps of: - Preprocessing (44) the pre-collected dataset by a predetermined cleaning and classification, the preprocessing (44) delivering preliminary clean and classified data, - Synchronizing (46) the sensor time data associated with the collected data of the preliminary clean and classified dataset, delivering a preliminary clean, classified and synchronized dataset, - Determining (48) at least one value of a variable representing at least one preliminary performance of the at least one distillation column from the preliminary clean, classified and synchronized dataset, and adding the at least one value of the variable of the at least one preliminary performance to the preliminary clean, classified and synchronized dataset to form a preliminary dataset to be processed, - Performing regression (50) on the learning model by calculating and filtering a predetermined derivative of the preliminary dataset to be processed, forming two classes generated by the learning model, the learning model being respectively associated with the normal operation of the at least one distillation column and the pre-flooding of the at least one distillation column, - Resampling (52) the two classes generated by the learning model at a predetermined sampling rate, - Use all the samples from the resampling (52) to determine, train, and validate (54) the learning model.
3. The method (38) according to claim 2, wherein the synchronization (46, 66) implemented during the construction and training phase (40) of the machine learning model for predicting flooding and / or during the operating phase (42) for predicting flooding comprises: Apply (47, 68) the time lag determined based on the position of each sensor in the set of sensors.
4. The method (38) according to any one of the preceding claims 1 - 3, wherein the learning model is a random forest model including a predetermined number of estimators and a maximum depth, the maximum depth being configured to expand each node of the random forest until all leaves are pure or until all leaves contain fewer than two samples.
5. The method (38) according to claim 2 or 3, wherein the current variable representing at least one current performance and / or the variable representing at least one previous performance are of types belonging to a group including at least the following: - A change in the total flow rate within the at least one distillation column, - A flooding characteristic corresponding to the difference between a predetermined reflux flow rate set point and the reflux flow rate measured and collected during a data collection step associated with the operation phase (42) or the construction and training phase (40), - An upper recycle metric of the at least one distillation column corresponding to the ratio of the liquid - to - gas ratio to the reflux ratio of the draw tray, the liquid - to - gas ratio and the reflux ratio being measured and collected during a data collection step associated with the operation phase (42) or the construction and training phase (40), - A risk metric determined at least from the temperature and pressure data of the at least one distillation column measured and collected during a data collection step associated with the operation phase (42) or the construction and training phase (40), - A flooding metric obtained from a predetermined theoretical equation and an associated binary metric, - A set of predetermined temperature differences and ratios obtained from at least two sensors in the set of sensors located at different positions relative to the position of the at least one distillation column, - A material balance, - An enthalpy.
6. The method (38) according to any one of the preceding claims 1 - 3, wherein the prediction (74) associates a probability with the binary value, and wherein after predicting the current state of the at least one distillation column, the method further comprises the steps of: During the prediction (74), in the case of obtaining a binary value representing the presence of a current pre - flooding with an associated probability value greater than a predetermined probability threshold, generate an alarm (78) and return the alarm to at least one operator located within the refinery.
7. The method (38) according to any one of the preceding claims 1 to 3, further comprising the following steps: Store (62) the data of the current data stream within the pre - collected data set for subsequent iterations of the construction and training phase (40) of the machine learning model for predicting flooding, and update the machine learning model for subsequent iterations of the operation phase (42) for predicting flooding.
8. The method (38) according to any one of the preceding claims 1-3, comprising: A compression step (58) implemented at a predetermined compression ratio during the collection of the current data stream, and a step of verifying the maintenance of the predetermined compression ratio at each subsequent collection step.
9. A computer program including software instructions which, when executed by a computer, implement the method according to any one of the preceding claims 1 - 8.
10. A system (10) for predicting flooding in at least one distillation column of a refinery implemented by machine learning, the system (10) including at least one database, characterized in that, The system (10) further includes: - A unit (12) for initially constructing and training a machine learning model for predicting flooding, the machine learning model obtained from a data set pre-collected and stored in the database during a predetermined previous period and at least from a set of sensors of the refinery, each data collected being associated with sensor time data, - Flooding prediction unit (14 A , 14 B ), the flooding prediction unit (14 A , 14 B ) includes: - A data buffer of a predetermined size and a collection module (16), the collection module (16) being configured to collect a current data stream from the set of sensors of the refinery until the data buffer of the predetermined size is filled, each data of the current data stream being further associated with sensor time data, - A preprocessing module (20), the preprocessing module (20) being configured to preprocess the data from the data buffer by a predetermined cleaning and classification, delivering a current clean and classified data set, - A synchronization module (22), the synchronization module (22) being configured to synchronize the sensor time data associated with the data of the current collected data stream of the current clean and classified data set, the synchronization module (22) delivering a current clean, classified and synchronized data set, - A determination module (24), the determination module (24) being configured to determine at least one current value of a variable representing at least one current performance of the at least one distillation column from the current clean, classified and synchronized data set, and adding the at least one value of the variable to the current clean, classified and synchronized data set so as to form a current data set to be processed, - A formation module (26), the formation module (26) being configured to form a current transformed data set by calculating a predetermined derivative of the current data set to be processed, - A prediction module (28), the prediction module (28) being configured to predict the current state of the at least one distillation column by applying the learning model to the current transformed data set, the current state corresponding to a binary value representing the presence or absence of current pre-flooding.
11. The system (10) according to claim 10, wherein: - The collection module (16) is located within the refinery itself, and - The unit (12) for initially constructing and training a machine learning model for predicting flooding, the data buffer of the flooding prediction unit, the preprocessing module, the synchronization module, the determination module, the formation module and the prediction module are outside the refinery and organized in a cloud computing manner, The collection module (16) is further adapted to directly load the data buffer, the system (10) further includes a receiving module, the receiving module being configured to receive information predicting the current state of the at least one distillation column and adapted to return via a return device of the system (10) to at least one operator present within the refinery.
Citation Information
Patent Citations
Process unit monitoring program
US20020116079A1
Accurate positioning system for a vehicle and its positioning method
US20080082265A1
Operating slide valves in petrochemical plants or refineries
US20180275690A1