Determining the appropriate sequence of actions to be taken during operational states of an industrial plant
The method employs a state-action network to determine the correct sequence of actions for addressing abnormal conditions in industrial plants, enhancing operator response and plant safety.
Patent Information
- Application Number
- JP2024530471
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-23
- Filing Date
- 2022-10-28
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2042-10-28
AI Technical Summary
Industrial plant operators face challenges in automatically determining the correct sequence of actions to remedy abnormal operating conditions, as existing systems primarily focus on alerting operators rather than providing the necessary actions.
A computer-implemented method using a state-action network that encodes multiple state variables into a lower-dimensional representation, maps this representation to a sequence of actions, and decodes it into a specific action sequence to be taken by the operator or automatically executed.
This method enables operators to take appropriate actions in response to abnormal conditions more effectively, reducing the risk of exacerbating the situation and improving plant safety and efficiency.
Smart Images

Figure 0007676668000001 
Figure 0007676668000002 
Figure 0007676668000003
Abstract
Description
[Technical field]
[0001] The present invention relates to monitoring industrial plants and determining actions to take in response to particular operating conditions, particularly abnormal operating conditions. [Background technology]
[0002] The intended operation of many industrial plants is controlled by a distributed control system (DCS), which adjusts the set points of low-level controllers to optimize one or more key performance indicators (KPIs). The plant is also continuously monitored for any abnormal conditions, such as equipment failure or malfunction. This monitoring may be at least partially integrated with the control by the DCS.
[0003] When an abnormal condition is detected, it is not always possible to remedy it automatically by the DCS. In every plant, there are abnormal situations that need to be remedied by an operator by performing a certain sequence of actions. WO2019 / 104296A1 discloses an alarm management system that assists an operator in identifying high priority alarms. However, the output of an alarm to an operator is only the first step to remedy an abnormal condition. It is also necessary for the operator to perform the correct sequence of actions. US 2020 / 026976 A1 discloses a method for matching an industrial machine with an intelligent industrial assistant having a set of predefined commands. The states of the machine's own I / O interfaces are mapped to the predefined commands. Based on this mapping, a customized interface for the machine is generated. US 2021 / 097401 A1 discloses a neural net system for processing input data values, where the input data values are encoded by an encoder network, the encoded products are aggregated by an aggregator, and the combination of the aggregated output and a target input value is decoded by a decoder to obtain a final output. US 2014 / 237487 A1 discloses an event processing system and method of operation, where the occurrence of an event is detected based on observations obtained from measurements, and based on the event, a state transition to be performed and an action to be performed are determined.
[0004] Object of the invention Therefore, the object of the present invention is to assist plant operators in remedying abnormal situations by calculating a sequence of actions to be taken from information obtained during monitoring of an industrial plant or part thereof.
[0005] This object is achieved by a first method for determining a suitable sequence of events according to the first independent claim and by a second method for training a machine learning network used in the first method according to the second independent claim. Further advantageous embodiments are specified in the respective dependent claims. Summary of the Invention
[0006] The present invention is defined by the following claims. Non-claimed embodiments and examples are presented to illustrate and facilitate understanding of the claimed invention. The present invention provides a computer-implemented method for determining an appropriate sequence of actions to take during operation of an industrial plant or part thereof.
[0007] In the course of the method, values of a number of state variables are obtained that characterize the operating state of the plant or part thereof. For example, these values may be combined into a vector or tensor. The specific set of state variables required to characterize the state of the plant or part thereof is plant specific. Examples of such state variables include pressure, temperature, mass flow rate, voltage, current, fill level, and concentration of a substance in a mixture of substances.
[0008] A plurality of state variables (e.g., a vector or tensor containing values of the state variables) is encoded into a representation of an operating state of the plant or a portion thereof by at least one trained state encoder network. In particular, such a representation may have a much lower dimensionality than the original plurality of state variables. That is, the representation may depend on a much smaller number of variables than the original vector having values of the state variables.
[0009] Examples of state variables include, but are not limited to, the following: · Variables that at least indicate a process variable such as pressure, mass flow, voltage or current; Events that indicate discrete changes in the plant, such as a motor switching on or off, a signal exceeding an alarm limit, or a valve opening or closing; · Combining process variables with events.
[0010] When the state variables include both process variables and events, the state encoder network may be a combined state encoder network that takes both process variables and events as inputs simultaneously, but a combination of two state encoding networks, one that encodes process variables and the other that encodes events, may also be used.
[0011] A trained state-action network maps a representation of an operational state into a representation of a sequence of actions to be taken in response to the operational state. In particular, just as a small number of variables can encode a complex operational state of a plant that depends on a large number of state variables, a small number of variables can also encode a complex sequence of actions with many different actions.
[0012] The representation of the sequence of actions is decoded by the trained action decoder network into a desired sequence of actions to be taken. The sequence of actions so determined may be used in any possible way to remedy the abnormal situation. For example, insofar as actions in the sequence may be performed automatically, automatic execution of these actions may be triggered. For actions that may only be performed by the operator, the operator may be notified in any suitable way to perform them. Also, any suitable action may be taken to assist the operator in performing such actions. For example, if the operator needs to operate some particular control element in the human machine interface of the DCS, this may be highlighted. If the action requires the operation of a physical control (such as a button, switch, or knob) and this is controlled and guarded by a cover to prevent accidental or inadvertent operation, the cover may be unlocked and / or opened. Similarly, if the action requires physical access to a particular piece of equipment, an access door to the equipment location may be unlocked and / or the equipment location may be made safe for the operator to enter, such as by stopping equipment that may cause harm to the operator. Hybrid solutions are also possible in the sense that the actions available in the HMI of the DCS are executed automatically without human intervention or optionally at least with human supervision (i.e., a human observes as the sequence of actions proposed by the system is executed automatically).
[0013] Examples of actions in a sequence of actions include: · Enabling or disabling equipment of the plant or any part thereof; · Opening and closing valves of the plant or parts thereof; · Changing the setpoint of at least one low-level controller in the plant or part thereof.
[0014] The primary evaluation of the plant's operational state occurs when this representation of the operational state is mapped to a representation of a sequence of actions. Since the representations usually have a much lower dimension than the operational states, respectively the sequences of actions, this means that the evaluation simply involves a mapping between the two low-dimensional spaces. This makes state-action networks much easier to train than networks that directly take high-dimensional operational states and directly output high-dimensional action sequences. In particular, the ease of training state-action networks comprises that training requires a smaller amount of training samples.
[0015] Furthermore, separate tasks in the overall processing chain, i.e. Going from a high-dimensional input to a low-dimensional representation, Processing this representation into another representation in another lower dimensional space; and Going from a second low-dimensional representation space to a second high-dimensional output space; are assigned to separate networks that can specialize for each job. This improves the overall accuracy of the final determined output compared to using a single network that must do all the tasks at once and may have to make tradeoffs between competing goals. This is somewhat similar to the philosophy of Unix tools such as sed, awk, and grep. Each tool is built to do exactly one single, simple job, and excels at that one single, simple job. For more complex jobs, the output of one tool is piped as input to the other tool.
[0016] In a particularly advantageous embodiment, the state encoder network is selected to be the encoder part of an encoder-decoder configuration that first maps a number of state variables to a representation of the operating state of the plant or part thereof, and then reconstructs the number of state variables from this representation. Such an encoder-decoder configuration may be trained in a self-supervised manner, i.e., training samples may be input to the configuration and the results of the configuration may be evaluated for how well they match the original inputs. For this training to work, the training samples do not need to be labeled with a "ground truth". In machine learning applications, obtaining the "ground truth" is the most expensive part of the training.
[0017] Moreover, training for reconstruction in this way trains the encoder-decoder configuration to force the parts of the input that are most salient for reconstruction to pass through the "information bottleneck" of the lower-dimensional representation, which filters out less important parts of the input, such as noise. Thus, the encoder also gains some denoising capability by training.
[0018] Alternatively or in combination with this, in another particularly advantageous embodiment, the action decoder network is chosen to be the decoder part of an encoder-decoder arrangement which first maps the sequence of actions and / or the processing results derived from this sequence of actions into a representation and then reconstructs the sequence of actions and / or the processing results derived therefrom from this representation, with the same advantages.
[0019] Reconstructing a "processing result" means, for example, that an encoder-decoder arrangement may be trained to predict the next action in a sequence of actions based on a portion of this sequence of actions.
[0020] Examples of networks that may be used as encoder and / or decoder networks for behavioral states and / or sequences of actions include recurrent neural networks, RNNs, and transformer networks. In recurrent neural networks, the output is fed back as input and the network is run for a predefined number of iterations. Transformer neural networks consist of a sequence of encoding units, each with an attention head that calculates the correlation between different parts of the input. Both architectures are particularly useful for processing sequences as input.
[0021] The state-action network may comprise, for example, a convolutional neural network. In a particularly advantageous embodiment, the state-action network comprises a fully connected neural network. This architecture offers maximum flexibility for training, at the expense of including many parameters per input and output size. As mentioned above, the dimensionality of the representation of the action states, as well as the dimensionality of the representation of the sequence of actions, is fairly low. Thus, many more parameters than a fully connected neural network would have can be accommodated.
[0022] Exemplary plants in which the method is particularly advantageous include continuous or process plants configured to emit alarm and event data, particularly for remediating abnormal situations. For example, the method may be used to remediate abnormal situations such as: · Waste incineration plants; · Hydrocarbon separation plants; · Reinjection systems for injecting water into hydrocarbon wells; Hydrocarbon utilization facilities; and / or · Glycol removal regeneration plant.
[0023] These plants have in common that abnormal situations often cause safety problems. In the case of a high priority alarm, if an incorrect sequence of actions is initiated or if an error is made when executing the correct sequence (such as omitting a step or swapping the order of two steps), this can exacerbate the abnormal situation and possibly miss the last chance to get the plant under control again. Also, safety-critical abnormal situations are fortunately rare, so there is little training data available for these situations. The method therefore has the advantage that it can work with less training data, since the main inference is made between two spaces of rather low dimensionality, as mentioned before.
[0024] The present invention also provides a method for training the configuration of a network for use in the above-described method.
[0025] In the course of the method, a first pre-trained encoder-decoder configuration of an action encoder network and an action decoder network is obtained, and a second pre-trained encoder-decoder configuration of a state encoder network and a state decoder network is obtained.
[0026] Samples of training data are obtained. Each such sample includes values of a number of state variables that characterize an operating state of the plant or a portion thereof. These state variables are input data for the configuration to be trained. Each sample also comprises a sequence of actions to be taken in response to this operating state. This sequence is a "ground truth" label attached to the operating state of the sample.
[0027] The values of the state variables in each sample are encoded into a representation of the respective action state by a pre-trained state encoder network. The resulting representation of the action state is then mapped to a representation of a sequence of actions. The sequence of actions encoded in this representation is the sequence of actions that the network configuration would propose given the action state characterized by the state variables.
[0028] The correspondence between this sequence of actions and the “ground truth” associated with the sample is measured by a cost function, which can be achieved in one or a combination of two ways:
[0029] According to the first method, the loss function measures how well the representation of a sequence of actions matches a representation obtained by encoding the sequence of actions in a training sample by a pre-trained action encoder network.
[0030] According to the second method, the loss function measures how well a sequence of actions obtained by decoding a representation of the sequence of actions by a pre-trained action decoder network matches the sequences of actions in the training samples.
[0031] During the training process, the parameters characterizing the behavior of the state-action network to be trained are optimized such that its rating by the loss function is likely to improve as further training samples are processed.
[0032] The state variables may be obtained during the actual execution of the process on the plant, or after such execution from a plant historian, or from a simulation run with a process simulator that generates the same state variables as the process on the plant. The use of a process simulator is particularly beneficial for a newly commissioned plant with little historical data, when the model is initially trained on simulated data capturing the general process dynamics and later on a limited amount of data from the actual process execution. Similarly, the sequence of actions may be monitored during the execution of the process, or obtained after such execution from the action log, or for initial training, from simulation experiments with actual plant operators or predefined action sequences. Both the state variables in the plant historian and the actions in the action log are time-stamped, so that they can be correlated with each other. Thus, training may be understood as the plant operator "mining" a workflow that reacts to a particular situation and teaching the network configuration to suggest this workflow when this situation or a substantially similar situation occurs again. In this way, even knowledge that exists in the operator's mind but is difficult to put into words or communicate to other operators can be used.
[0033] For example, if operators have learned to "open the valve slowly if the flame turns bluish", each operator may perceive the moment when the flame turns bluish slightly differently. Also, different operators may have different concepts of opening a valve "slowly". The present training method allows knowledge to be captured in an automated way that leaves no room for interpretation.
[0034] In a particularly advantageous embodiment, obtaining a first pre-trained encoder-decoder configuration of the action encoder network and the action decoder network comprises: Obtaining training samples of sequences of actions; providing the sequence of actions in each training sample and / or the processing results derived therefrom to an action encoder network to be trained, thereby obtaining a representation; providing this representation to the action decoder network to be trained, thereby obtaining a sequence of actions and / or a processing result derived therefrom; measuring how well this sequence of actions and / or this processing result matches the sequences of actions and / or processing results in training samples using a predefined loss function; Optimizing parameters characterizing the behavior of the action encoder network to be trained and the action decoder network to be trained such that the rating by the loss function is likely to improve as further training samples are processed; It is equipped with:
[0035] The training samples used for this pre-training may have training samples in common with the main training described above, but may be performed on a set of training samples separate from those used for the main training. For example, the pre-training may be performed in a generic way once for a particular type of plant. For each instance of the plant that is subsequently installed, the pre-trained encoder-decoder configuration may be used in training the state-sequence network. Optionally, when moving from generic training to a specific instance of the plant, the pre-training of the encoder-decoder configuration may be refined using further training samples taken from this instance of the plant.
[0036] Pre-training can be performed with a variety of tasks for which "ground truth labels" can be easily generated from available process state data and action sequences. Examples of such tasks are reconstructing an input (plant state variables or action sequences), predicting the next n elements of a sequence (plant state variables or action sequences), identifying the correct next sequence segment among several presented sequence segments, identifying the correct previous sequence segment among several presented sequence segments, identifying whether presented sequences are adjacent in the entire sequence, etc. Such tasks may also be combined in parallel or in sequence, which is beneficial as it further increases the amount of training data for pre-training and also prevents overfitting the pre-trained model to a single task.
[0037] The same advantage applies in a similar manner to a further particularly advantageous embodiment, where obtaining a second pre-trained encoder-decoder configuration of the state encoder network and the state decoder network comprises: obtaining training samples comprising values of a plurality of state variables characterizing an operating state of a plant or a portion thereof; providing values of state variables at each training sample to the state encoder network to be trained, thereby obtaining a representation; providing this representation to the state decoder network to be trained, thereby obtaining values of the state variables; Measuring how well these values match those of the training samples by a predefined loss function; Optimizing parameters characterizing the behavior of the state encoder network to be trained and the state decoder network to be trained such that the rating by the loss function is likely to improve as further training samples are processed; Equipped with.
[0038] In a further particularly advantageous embodiment, the action encoder network to be trained and the state encoder network to be trained are combined into one single network architecture, which may depend on fewer parameters than the combination of the two individual architectures, resulting in easier training. Also, since the tasks performed by both networks have something in common, the two trainings can benefit to some extent from each other by "sharing" knowledge in one single network architecture.
[0039] In a further particularly advantageous embodiment, obtaining training samples for training the first encoder-decoder configuration, and / or the second encoder-decoder configuration, and / or the state-action network comprises aggregating training samples obtained in multiple industrial plants. This improves the overall variability in the set of training samples, resulting in a better performance of the final configuration of the network in terms of accuracy. As mentioned above, abnormal situations that pose safety risks tend to occur very rarely in any given plant. Due to safety risks, it is usually not practical to induce such situations just for the purpose of obtaining training samples. However, in larger plants, there will be enough instances of abnormal situations occurring on their own that a reasonable amount of training samples can be collected.
[0040] As mentioned above, the method is computer-implemented. Thus, the present invention also relates to one or more computer programs having machine-readable instructions that, when executed on one or more computers and / or computing instances, cause the one or more computers to perform the above-mentioned methods. In this context, virtualization platforms, hardware controllers, network infrastructure devices (such as switches, bridges, routers, or wireless access points), as well as end devices in a network (such as sensors, actuators, or other industrial field devices) capable of executing machine-readable instructions shall also be considered as computers.
[0041] The present invention therefore also relates to a non-transitory storage medium and / or a download product having one or more computer programs. A download product is a product that may be sold in an online shop for immediate fulfillment by download. The present invention also provides one or more computers and / or computing instances having one or more computer programs and / or one or more non-transitory machine-readable storage media and / or a download product.
[0042] In the following, the present invention is illustrated by means of drawings, which are not intended to limit the scope of the present invention. [Brief description of the drawings]
[0043] [Figure 1] FIG. 1 is an exemplary embodiment of a method 100 for determining an appropriate sequence of actions to take during operation of an industrial plant. [Diagram 2] FIG. 2 is an exemplary embodiment of a method 200 for training a configuration of a network for use in method 100 . [Diagram 3] FIG. 3 illustrates two ways in which the state-action network 4 can be trained.
[0044] FIG. 1 is a schematic flow chart of one embodiment of a method 100 for determining an appropriate sequence of actions to take during operation of an industrial plant 1.
[0045] In step 110, values of a number of state variables 2 characterizing the operating state of the plant 1 or part thereof are obtained.
[0046] In step 120, a number of state variables 2 are encoded by at least one trained state encoder network 3 into a representation 2a of the operating state of the plant 1 or part thereof.
[0047] According to block 121, the state encoder network 3 may be selected to be the encoder part of an encoder-decoder configuration which first maps a number of state variables 2 to a representation 2a of the operating state of the plant 1 or part thereof and then reconstructs the number of state variables 2 from this representation 2a.
[0048] In step 130, the representation 2a of the action state is mapped by the trained state-action network 4 to a representation 6a of a sequence 6 of actions to be taken in response to the action state.
[0049] In step 140 , the representation 6 a of the sequence of actions 6 is decoded by the trained action decoder network 5 into the desired sequence 6 of actions to be taken.
[0050] According to block 141, the action decoder network 5 may be selected to be the decoder part of an encoder-decoder arrangement which first maps a sequence of actions 6 and / or processing results derived from this sequence of actions 6 into a representation 6a and then reconstructs the sequence of actions 6 and / or processing results derived therefrom from this representation 6a.
[0051] FIG. 2 is a schematic flow chart of one embodiment of a method 200 for training configurations of networks 3, 4, 5 for use in the method 100 described above.
[0052] In step 210, a pre-trained first encoder-decoder configuration of the action encoding network 5# and the action decoder network 5 is obtained. In the example shown in FIG. 2, this obtaining comprises the following additional steps: - obtaining training samples for a sequence of actions 6 according to block 211; providing, according to block 212, the sequence of actions 6 in each training sample and / or the processing results derived therefrom to an action encoder network 5# to be trained, thereby obtaining a representation 6a; providing this representation 6a to an action decoder network 5 to be trained, according to block 213, thereby obtaining a sequence of actions 6' and / or a processing result derived therefrom; measuring, according to a block 214, how well this sequence of actions 6' and / or this processing result matches the sequence of actions 6 and / or the processing result in a training sample by means of a predefined loss function 9; Optimizing the parameters 5a#, 5a characterizing the behavior of the action encoder network 5# to be trained and the action decoder network 5 to be trained such that the rating 9a according to the loss function 9 is likely to improve when further training samples are processed, according to block 215. The final optimized states of the parameters 5a#, 5a are labelled with the reference characters 5a#*, 5a*.
[0053] In step 220, a pre-trained second encoder-decoder configuration of the state encoder network 3 and the state decoder network 3# is obtained. In the example shown in FIG. 2, this includes the following additional steps: - obtaining training samples comprising values of a number of state variables 2 characterizing an operating state of the plant 1 or part thereof, according to a block 221; providing the values of state variable 2 at each training sample to the state encoder network to be trained, thereby obtaining a representation 2a, according to block 222; providing this representation 2a to a state decoder network 3# to be trained, thereby obtaining the values 2' of the state variables, according to block 223; measuring, according to block 224, how well these values 2' match the values 2 of the training samples by means of a predetermined loss function 10; Optimizing the parameters 3a, 3a# characterizing the behavior of the state encoder network 3 to be trained and the state decoder network 3# to be trained such that the rating 10a according to the loss function 10 is likely to improve when further training samples are processed, according to block 225. The final optimized states of parameters 3a, 3a# are labelled with reference characters 3a*, 3a#*.
[0054] In step 230, samples 7 of training data are obtained. Each sample 7 comprises values of a number of state variables 2 characterizing an operating state of the plant 1 or part of it and a sequence 6* of actions to be taken in response to this operating state.
[0055] According to block 231, obtaining 230 training samples 7 for training the first encoder-decoder configuration and / or the second encoder-decoder configuration and / or the state-action network 4 comprises aggregating training samples 7 obtained in multiple industrial plants 1.
[0056] In step 240, the values of the state variables 2 in each sample 7 are encoded into a representation 2a of the respective motion state by a pre-trained state encoder network 3.
[0057] In step 250, the representation 2a of the action state is mapped to a representation 6a of the sequence 6 of actions.
[0058] In step 260, the degree of the loss is calculated by a predetermined loss function 8. the representation 6a of the sequence of actions corresponds to a representation 6a* obtained by encoding the sequence of actions 6 in the training sample 7 by means of a pre-trained action encoder network 5#, and / or whether the sequence of actions 6* obtained by decoding the representation 6a of the sequence of actions 6 by the pre-trained action decoder network 5 matches the sequence of actions 6* in the training samples 7; is measured.
[0059] In step 270, parameters 4a characterizing the behavior of the state-action network 4 to be trained are optimized such that the rating 8a by the loss function 8 is likely to improve as further training samples 7 are processed. The final optimized state of the parameters 4a is labelled with the reference sign 4a*.
[0060] FIG. 3 illustrates two ways in which the state-action network 4 may be trained.
[0061] The values of one or more state variables 2 from the training samples 7 are encoded into a representation 2a of an operational state of the plant 1. This representation 2a is mapped by the state-action network 4 into a representation 6a of a sequence of actions 6. This representation 6a needs to be compared to a "ground truth" sequence of actions 6* in the training samples 7 to measure how correct the representation 6a output by the state-action network 4 is.
[0062] In the first method, a "ground truth" sequence of actions 6* is encoded into a "ground truth" representation 6a* by a pre-trained action encoder network 5#. A loss function 8 measures how well the representation 6a output by the state-action network 4 matches the "ground truth" representation 6a*.
[0063] In the second method, the representation 6a output by the state-action network 4 is decoded by a pre-trained action decoder network 5 to obtain a sequence of actions 6'. A loss function 8 measures how well this sequence of actions 6 matches a "ground truth" sequence of actions 6*. The invention as originally claimed in the present application is set forth below. [1] A computer-implemented method (100) for determining an appropriate sequence of actions (6) to be taken during operation of an industrial plant (1) or part thereof, comprising: - obtaining (110) values of a number of state variables (2) characterizing an operating state of the plant (1) or part thereof; - encoding (120) said plurality of state variables (2) into a representation (2a) of said operating state of said plant (1) or part thereof by at least one trained state encoder network (3); a step (130) of mapping (2a) said operational states to a representation (6a) of a sequence (6) of actions to be taken in response to said operational states by a trained state-action network (4); and decoding (140) the representation (6a) of the sequence of actions (6) into a desired sequence of actions to be taken (6) by a trained action decoder network (5). [2] The method (100) according to [1], wherein the state encoder network (3) is selected to be the encoder part of an encoder-decoder configuration (121) that first maps a plurality of state variables (2) to a representation (2a) of the operating state of the plant (1) or part thereof, and then reconstructs the plurality of state variables (2) from this representation (2a). [3] The method (100) according to any one of [1] or [2], wherein the action decoder network (5) is selected (141) to be the decoder part of an encoder-decoder arrangement which first maps a sequence of actions (6) and / or processing results derived from this sequence of actions (6) into a representation (6a) and then reconstructs the sequence of actions (6) and / or processing results derived therefrom from this representation (6a). [4] The method (100) of any one of [1] to [3], wherein the state encoder network (3) and / or the action decoder network (5) comprise a recurrent neural network, an RNN, and / or a transformer neural network. [5] The method (100) of any one of [1] to [4], wherein the state-action network (4) comprises a convolutional neural network and / or a fully connected neural network. [6] The method (100) according to any one of [1] to [5], wherein the state variables (2) comprise one or more of pressure, temperature, mass flow rate, voltage, current, fill level, and / or concentration of a substance in a mixture of substances. [7] The action is: - enabling or disabling the equipment of said plant (1) or part thereof; opening and closing valves of said plant (1) or of parts thereof; - changing the setpoints of at least one low-level controller in the plant (1) or part thereof. [8] The method (100) according to any one of [1] to [7], wherein the plant (1) or part thereof is a continuous plant or a process plant configured to emit alarm and event data. [9] The plant (1) or part thereof waste incineration plants; Hydrocarbon separation plants; A reinjection system for injecting water into hydrocarbon wells; Hydrocarbon utilization facilities; and / or a deglycolization regeneration plant.
[10] A computer-implemented method (200) for training a configuration of a network (3, 4, 5) for use in the method (100) according to any one of [1] to [9], comprising: obtaining (210) a pre-trained first encoder-decoder configuration of an action encoder network (5#) and an action decoder network (5); Obtaining (220) a pre-trained second encoder-decoder configuration of the state encoder network (3) and the state decoder network (3#); a step (230) of obtaining samples (7) of training data, each sample (7) comprising values of a number of state variables (2) characterizing said operating state of said plant (1) or of a part thereof and a sequence (6*) of actions to be taken in response to said operating state, encoding (240) the values of the state variables (2) in each sample (7) into a representation (2a) of a respective motion state by the pre-trained state encoder network (3); a step (250) of mapping (2a) said representations of operational states (2a) to representations (6a) of sequences of actions (6) by said state-action network (4) to be trained; Using a predefined loss function (8), to what extent the representation (6a) of the sequence of actions corresponds to a representation (6a*) obtained by encoding the sequence (6*) of actions in a training sample (7) by the pre-trained action encoder network (5#), and / or whether the sequence of actions (6) obtained by decoding the representation (6a) of the sequence of actions (6) by the pre-trained action decoder network (5) matches the sequence of actions (6*) in the training sample (7); measuring (260); optimizing (270) parameters (4a) characterizing the behavior of the state-action network (4) to be trained such that its rating (8a) by the loss function (8) is likely to improve when further training samples (7) are processed.
[11] Obtaining (210) a first pre-trained encoder-decoder configuration of an action encoder network (5#) and an action decoder network (5) includes: Obtaining (211) training samples of a sequence of actions (6); providing (212) said sequence of actions (6) for each training sample and / or a processing result derived therefrom to said action encoder network (5#) to be trained, thereby obtaining a representation (6a); providing this representation (6a) to the action decoder network (5) to be trained, thereby obtaining (213) a sequence of actions (6') and / or a processing result derived therefrom; measuring (214) how well this sequence of actions (6') and / or this processing result matches the sequence of actions (6) and / or the processing result in the training sample using a predefined loss function (9); optimizing (215) parameters (5a#, 5a) characterizing the behavior of the action encoder network (5#) to be trained and the action decoder network (5) to be trained such that the rating (9a) by the loss function (9) is likely to improve when further training samples are processed.
[12] Obtaining (220) a second pre-trained encoder-decoder configuration of the state encoder network (3) and the state decoder network (3#) includes: obtaining (221) training samples comprising values of a plurality of state variables (2) characterizing an operating state of the plant (1) or a part thereof; providing (222) values of the state variables (2) for each training sample to the state encoder network (3) to be trained, thereby obtaining a representation (2a); providing (223) this representation (2a) to the state decoder network (3#) to be trained, thereby obtaining the values (2') of the state variables; measuring (224) how well these values (2') match the values (2) of the training samples using a predefined loss function (10); optimizing (225) parameters (3a, 3a#) characterizing the behavior of the state encoder network (3) to be trained and the state decoder network (3#) to be trained such that the rating (10a) by the loss function (10) is likely to improve when further training samples are processed.
[13] The method (200) described in
[12] citing
[11] , in which the action encoder network to be trained (5#) and the state encoder network to be trained (3) are combined into one single network architecture.
[14] The method (200) according to any one of
[10] to
[13] , wherein obtaining (230) training samples (7) for the training of the first encoder-decoder configuration and / or the second encoder-decoder configuration and / or the state-action network (4) comprises aggregating (231) training samples (7) obtained in a plurality of industrial plants (1).
[15] One or more computer programs comprising machine-readable instructions that, when executed on one or more computers, cause the one or more computers to perform the method (100, 200) described in any one of [1] to
[14] .
[16] A non-transitory storage medium and / or download product comprising a computer program as described in
[15] .
[17] One or more computers having a computer program as described in
[15] and / or a non-transitory storage medium and / or a download product as described in
[16] .
[0064] List of References 1. Industrial plants 2 State variables characterizing the operating state of plant 1 2' State variables to be decoded during encoder-decoder training 2. Display of the operating status of Plant 1 3-State Encoder Network 3a Parameters characterizing the behavior of network 3 3a* The final optimized state of parameter 3a 3# State decoder network 3a# Parameters characterizing the behavior of network 3# 3a#* The final optimized state of parameter 3a# 4. State-Action Networks 4a Parameters characterizing the behavior of network 4 4a* The final optimized state of parameter 4a 5. Action Decoder Network 5a Parameters characterizing the behavior of network 5 5a* The final optimized state of parameter 5a 5# Action Encoder Network 5a# Parameters characterizing the behavior of network 5# 5a#* The final optimized state of parameter 5a# 6 Sequence of Actions 6' Sequence decoded during encoder-decoder training 6* Sequence of actions in training sample 7 6a Representation of the sequence of actions 6 6a* Encoded representation from sequence 6* 7 Training samples for training the state-action network 4 8 Loss Functions for Training State-Action Networks4 8a Rating by loss function 9 Action Encoder-Decoder Configuration 5#, Loss Function for 5 9a Rating by loss function Loss function for 10-state encoder-decoder configuration 3, 3# 10a Rating by loss function 100 Methods for determining the appropriate sequence of actions6 110 Get state variable 2 120 Encode state variables into representation 120 121 Select state encoder 3 from the encoder-decoder configuration 130 Mapping Representation 2a to Sequence Representation 6a 140 Decode expression 6a into the desired sequence 6b 141 Select action decoder 5 from the encoder-decoder configuration 200 How to train configurations of networks 3, 4 and 5 210 Get the first pre-trained encoder-decoder configuration 5#,5 211 Get training sample for sequence 6 212 Provide training sequence 6 to action encoder network 5# 213 Representation 6a is provided to the action decoder network 5 214 Rate the decoded sequence 6' with loss function 9 215 Optimize parameters 5a# and 5b# of networks 5a and 5b 220 Get the pre-trained second encoder-decoder configuration 3, 3# 221 Get training samples for state variable 2 222 Provide state variable 2 to state encoder network 3 223 Representation 2a is provided to state decoder network 3# 224 Rating value 2' with loss function 10# 225 Optimize parameters 3a and 3a# of networks 3 and 3# Take sample 7 of 230 training data Aggregate training samples 7 across 231 plants 240 Encode training state variable 2 into representation 2a 250 Mapping state representation 2a to sequence representation 6a 260 Rating sequence representation 6a / sequence 6 with loss function 8 270 Optimize parameter 4a of state-action network 4
Claims
1. A computer-implemented method (100) for determining an appropriate sequence of actions (6) to be taken during operation of an industrial plant (1) or part thereof, comprising: - obtaining (110) values of a number of state variables (2) characterizing an operating state of said plant (1) or part thereof; a step (120) of encoding said plurality of state variables (2) into a representation (2a) of said operating states of said plant (1) or part thereof, by at least one trained state encoder network (3), wherein said representation (2a) of said operating states depends on fewer variables than said operating states of said plant (1); - mapping (130) by a trained state-action network (4) said representation (2a) of said operational state to a representation (6a) of a sequence of actions (6) to be taken in response to said operational state, where this representation (6a) of said sequence of actions (6) depends on fewer variables than a complex sequence of actions (6) comprising many different actions; - decoding (140) said representation (6a) of said sequence of actions (6) into a sequence of actions to be taken (6) by a trained action decoder network (5).
2. 2. The method (100) of claim 1, wherein the state encoder network (3) is selected (121) to be the encoder part of an encoder-decoder configuration that first maps a plurality of state variables (2) to a representation (2a) of the operating state of the plant (1) or part thereof and then reconstructs the plurality of state variables (2) from this representation (2a).
3. The method (100) according to any one of claims 1 or 2, wherein the action decoder network (5) is selected (141) to be the decoder part of an encoder-decoder arrangement which first maps a sequence of actions (6) and / or processing results derived from this sequence of actions (6) into a representation (6a) and then reconstructs the sequence of actions (6) and / or processing results derived therefrom from this representation (6a).
4. The method (100) of claim 1 or 2, wherein the state encoder network (3) and / or the action decoder network (5) comprise a recurrent neural network, an RNN, and / or a transformer neural network.
5. The method (100) of claim 1 or 2, wherein the state-action network (4) comprises a convolutional neural network and / or a fully connected neural network.
6. The method (100) of claim 1 or 2, wherein the state variables (2) comprise one or more of the following: pressure, temperature, mass flow rate, voltage, current, fill level, and / or concentration of a substance in a mixture of substances.
7. The action is: - Enabling or disabling equipment of the plant (1) or parts thereof; - opening and closing valves of said plant (1) or parts thereof; - changing the setpoints of at least one low-level controller in the plant (1) or part thereof.
8. The method (100) of claim 1 or 2, wherein the plant (1) or part thereof is a continuous or process plant configured to emit alarm and event data.
9. The plant (1) or a part thereof - a waste incineration plant; - a hydrocarbon separation plant; - a reinjection system for injecting water into a hydrocarbon well; Hydrocarbon utilization facilities, and / or The method (100) of claim 1 or 2, comprising one or more of the following: a deglycolization regeneration plant.
10. A computer-implemented method (200) for training a configuration of a network (3, 4, 5) for use in the method (100) of claim 1, comprising the steps of: Obtaining (210) a pre-trained first encoder-decoder configuration of the action encoder network (5#) and the action decoder network (5); - obtaining (220) a pre-trained second encoder-decoder configuration of a state encoder network (3) and a state decoder network (3#); - obtaining (230) samples (7) of training data, each sample (7) comprising values of a number of state variables (2) characterizing said operating state of said plant (1) or part thereof and a sequence (6*) of actions to be taken in response to said operating state, encoding (240) by the pre-trained state encoder network (3) the values of the state variables (2) in each sample (7) into a representation (2a) of a respective motion state, where this representation (2a) of the motion state depends on fewer variables than the respective motion state; a step (250) of mapping (2a) said representation (2a) of said action states to a representation (6a) of a sequence of actions (6) by said state-action network (4) to be trained, where this representation (6a) of said sequence of actions (6) depends on fewer variables than said sequence of actions (6); Using a predefined loss function (8), to what extent the representation (6a) of the sequence of actions corresponds to a representation (6a*) obtained by encoding the sequence of actions (6*) in a training sample (7) by the pre-trained action encoder network (5#), and / or whether the sequence of actions (6) obtained by decoding the representation (6a) of the sequence of actions (6) by the pre-trained action decoder network (5) matches the sequence of actions (6*) in the training sample (7); (260) measuring the score, thereby obtaining a rating (8a); - optimizing (270) parameters (4a) characterizing the behavior of the state-action network (4) to be trained such that the rating (8a) by the loss function (8) is likely to improve when further training samples (7) are processed.
11. Obtaining (210) a first pre-trained encoder-decoder configuration of an action encoder network (5#) and an action decoder network (5) includes: Obtaining (211) training samples of a sequence of actions (6); - providing (212) the sequence of actions (6) for each training sample and / or the processing results derived therefrom to the action encoder network (5#) to be trained, thereby obtaining a representation (6a); - providing this representation (6a) to the action decoder network (5) to be trained, thereby obtaining (213) the sequence of actions (6') and / or the processing results derived therefrom; measuring (214) how well this sequence of actions (6') and / or this processing result matches the sequence of actions (6) and / or the processing result in the training sample using a predefined loss function (9); - optimizing (215) parameters (5a#, 5a) characterizing the behavior of the action encoder network (5#) to be trained and the action decoder network (5) to be trained such that the rating (9a) by the loss function (9) is likely to improve when further training samples are processed.
12. Obtaining (220) a second pre-trained encoder-decoder configuration of the state encoder network (3) and the state decoder network (3#) includes: - obtaining (221) training samples comprising values of a plurality of state variables (2) characterizing an operating state of said plant (1) or a part thereof; providing (222) the values of the state variables (2) at each training sample to the state encoder network (3) to be trained, thereby obtaining a representation (2a); providing (223) this representation (2a) to the state decoder network (3#) to be trained, thereby obtaining the values (2') of the state variables; measuring (224) how well these values (2') match the values (2) of the training samples using a predefined loss function (10); Optimizing (225) parameters (3a, 3a#) characterizing the behavior of the state encoder network (3) to be trained and the state decoder network (3#) to be trained such that the rating (10a) by the loss function (10) is likely to improve when further training samples are processed.
13. Obtaining (220) a second pre-trained encoder-decoder configuration of the state encoder network (3) and the state decoder network (3#) includes: - obtaining (221) training samples comprising values of a plurality of state variables (2) characterizing an operating state of said plant (1) or a part thereof; providing (222) the values of the state variables (2) at each training sample to the state encoder network (3) to be trained, thereby obtaining a representation (2a); providing (223) this representation (2a) to the state decoder network (3#) to be trained, thereby obtaining the values (2') of the state variables; measuring (224) how well these values (2') match the values (2) of the training samples using a predefined loss function (10); optimizing (225) parameters (3a, 3a#) characterizing the behavior of the state encoder network (3) to be trained and the state decoder network (3#) to be trained such that the rating (10a) by the loss function (10) is likely to improve when further training samples are processed; The method (200) of claim 11, wherein the action encoder network (5#) to be trained and the state encoder network (3) to be trained are combined into one single network architecture.
14. The method (200) of claim 10, wherein obtaining (230) training samples (7) for the training of the first encoder-decoder configuration and / or the second encoder-decoder configuration and / or the state-action network (4) comprises aggregating (231) training samples (7) obtained in a plurality of industrial plants (1).
15. One or more computer programs comprising machine-readable instructions that, when executed on one or more computers, cause the one or more computers to perform the method (100) of claim 1.
16. One or more computer programs comprising machine-readable instructions that, when executed on one or more computers, cause the one or more computers to perform the method (200) of claim 10.
17. A non-transitory storage medium having the computer program according to claim 15 recorded thereon.
18. A non-transitory storage medium having the computer program according to claim 16 recorded thereon.
19. One or more computers configured to carry out the method (100) of claim 1 by a computer program according to claim 15 and / or by a non-transitory storage medium and / or a download product according to claim 17.
20. One or more computers configured to carry out the method (200) of claim 10 by a computer program according to claim 16 and / or by a non-transitory storage medium and / or a download product according to claim 18.
Citation Information
Patent Citations
Symbolizing device and process controller and control supporting device using the symbolizing device
JP1991166601A
Failure restoration device, failure restoration method, and program
JP2021174348A