CONTROL DEVICE FOR CONTROLLING A TECHNICAL SYSTEM AND METHOD FOR CONFIGURING THE CONTROL DEVICE
Patent Information
- Application Number
- DE502021008111
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-01-29
- Publication Date
- 2025-08-14
- Estimated Expiration
- 2041-01-29
AI Technical Summary
Existing machine learning methods for controlling complex technical systems face convergence problems and repeatability issues due to incomplete consideration of the system's state space, noisy sensor data, and time delays in control actions, which impair learning success.
A control device and method utilizing three machine learning modules to reproduce system behavior without and with control actions, followed by optimizing control action performance based on deviations and specific behavioral signals, allowing efficient training without implicitly learning system behavior.
This approach enhances convergence behavior, improves training repeatability, reduces data and computing requirements, and stabilizes training against data variations, resulting in more robust and efficient control system configurations.
Description
[0001] Machine learning methods are increasingly being used to control complex technical systems, such as gas turbines, wind turbines, internal combustion engines, robots, manufacturing plants, or power grids. Using such learning methods, a machine learning model of a control device can be trained on the basis of training data to determine, based on current operating signals from a technical system, those control actions for controlling the technical system that specifically bring about a desired or optimized behavior of the technical system and thus optimize its performance. Such a machine learning model for controlling a technical system is often also referred to as a policy or control model. A variety of well-known training methods, such as reinforcement learning methods, are available for training such a policy.
[0002] However, in control optimizations in industrial environments, many known training methods experience convergence problems and / or problems with the repeatability of learning processes. This can be due, for example, to the fact that only a small part of the technical system's state space is considered, that sensor data from the technical system is noisy, and / or that control actions generally have a time delay, with different control actions often resulting in different time delays. The above symptoms frequently occur in complex real-world systems and can significantly impair learning success.
[0003] EP 3 588 211 A1 discloses a computer-implemented method for configuring a control device for a technical system.
[0004] It is an object of the present invention to provide a control device for controlling a technical system and a method for configuring the control device, which allow more efficient training.
[0005] This object is achieved by a method having the features of patent claim 1, by a control device having the features of patent claim 11, by a computer program product having the features of patent claim 12 and by a computer-readable storage medium having the features of patent claim 13.
[0006] To configure a control device for a technical system, an operating signal of the technical system is fed into a first machine learning module, which is trained to use an operating signal of the technical system to reproduce a behavioral signal of the technical system that specifically arises without the current application of a control action and to output the reproduced behavioral signal as a first output signal. The first output signal is fed into a second machine learning module, which is trained to use a control action signal to reproduce a resulting behavioral signal of the technical system and to output the reproduced behavioral signal as a second output signal. Furthermore, an operating signal of the technical system is fed into a third machine learning module, and a third output signal of the third machine learning module is fed into the trained, second machine learning module.Based on the second output signal, a control action performance is determined. This trains the third machine learning module to optimize the control action performance based on an operating signal from the technical system. Finally, based on the third machine learning module, the control device is configured to control the technical system using a third output signal from the third machine learning module.
[0007] To carry out the method according to the invention, a control device, a computer program product and a preferably non-volatile computer-readable storage medium are provided.
[0008] The method according to the invention and the control device according to the invention can be carried out or implemented, for example, by means of one or more computers, processors, application-specific integrated circuits (ASICs), digital signal processors (DSPs) and / or so-called "field programmable gate arrays" (FPGAs).
[0009] The invention allows a control device to be configured or trained considerably more efficiently. If the trained second machine learning module is used to train the third machine learning module, essential components of a system's behavior generally no longer need to be learned or represented implicitly when training the third machine learning module. This often leads to significantly improved convergence behavior and / or better repeatability of training results. Furthermore, training is often more stable and / or robust against variations in the training data. Furthermore, in many cases, less training data, computing time, and / or computing resources are required.
[0010] Advantageous embodiments and further developments of the invention are specified in the dependent claims.
[0011] According to an advantageous embodiment of the invention, the third machine learning module can be trained using the first output signal. This often allows the third machine learning module to be trained particularly effectively, since the third machine learning module has access to specific information about a system behavior without the current application of a control action.
[0012] According to a particularly advantageous embodiment of the invention, the control action performance can be determined for a particular point in time based on a single time step of a behavioral signal. A complex determination or estimation of future effects on performance is often unnecessary. Thus, dynamic effects occurring on different time scales can also be efficiently considered. Furthermore, the time step can vary in length depending on a control action and / or a behavioral signal and can also represent effects of control actions further in the future.
[0013] Advantageously, first and / or second parts of an operating signal of the technical system can be specifically selected based on whether or not they comprise a control action. Thus, first parts of the operating signal that do not comprise a control action can be specifically used to train the first machine learning module, and / or second parts of the operating signal that comprise a control action can be specifically used to train the second machine learning module. Through a specific selection of training data aligned with a respective training objective, the first and / or second machine learning module can be trained particularly effectively.
[0014] According to a further advantageous embodiment of the invention, a behavior signal setpoint can be read in, and the second output signal can be compared with the behavior signal setpoint. This allows the control action performance to be determined based on the comparison result. In particular, a deviation between the second output signal and the behavior signal setpoint can be determined, e.g., in the form of a difference or difference square. The control action performance can then be determined based on the deviation, with a larger deviation generally leading to lower control action performance.
[0015] The behavioral signal setpoint can then be fed into the third machine learning module. This allows the third machine learning module to be trained to optimize control action performance based on the behavioral signal setpoint.
[0016] According to a further advantageous embodiment of the invention, the control action performance can be determined based on the first output signal. In particular, a deviation between the first output signal and the second output signal can be determined, e.g., in the form of a difference or difference square. Alternatively or additionally, a deviation of a sum of the first and second output signals from a behavior signal target value can be determined. The control action performance can then be determined depending on the deviation thus determined. In this case, the deviation can be used, in particular, to evaluate how system behavior with application of a control action differs from system behavior without application of this control action. It has been shown that the determination of control action performance can be significantly improved in many cases through this distinction.
[0017] According to an advantageous development of the invention, the first and / or second machine learning module can be trained to separately reproduce multiple behavioral signals of different processes occurring in the technical system. The control action performance can then be determined depending on the reproduced behavioral signals. For this purpose, the first and / or second machine learning module can, in particular, comprise a set of machine learning models or submodels, each of which models a specific process occurring in the technical system in a process-specific manner. Such separate training proves to be more efficient in many cases than combined training, since the underlying individual dynamics generally exhibit a simpler response behavior than a combined system dynamics.
[0018] To the extent that the invention allows the control action performance to be determined at a given point in time based on a single, possibly adjustable, time step of a behavioral signal, fewer synchronization problems generally arise between processes with different execution speeds, especially during training of the third machine learning module. In many cases, it has been shown that a relatively precise and robust evaluation of the control action performance can be performed in a single step for various process-specific machine learning models.
[0019] Furthermore, a specific behavioral signal setpoint can be read in for each behavioral signal. The control action performance can then be determined by comparing the reproduced behavioral signals with the specific behavioral signal setpoints.
[0020] In particular, the third machine learning module can be trained to optimize the control action performance based on the specific behavioral signal setpoints.
[0021] An embodiment of the invention is explained in more detail below with reference to the drawings, each of which shows a schematic representation: Figure 1 shows a gas turbine with a control device according to the invention, Figure 2 shows a control device according to the invention in a first training phase, Figure 3 shows the control device in a second training phase and Figure 4 shows the control device in a third training phase.
[0022] Figure 1illustrates, by way of example, a gas turbine as a technical system TS with a control device CTL. Alternatively or additionally, the technical system TS can also include a wind turbine, an internal combustion engine, a manufacturing plant, a chemical, metallurgical, or pharmaceutical manufacturing process, a robot, a motor vehicle, an energy transmission network, a 3D printer, or another machine, device, or system.
[0023] The gas turbine TS is coupled to the control device CTL, which can be implemented as part of the gas turbine TS or entirely or partially external to the gas turbine TS. Figure 1 For reasons of clarity, the control device CTL is shown externally to the technical system TS.
[0024] The control unit CTL is used to control the technical system TS and is trained for this purpose using a machine learning process. Controlling the technical system TS also includes regulating the technical system TS as well as outputting and using control-relevant data or signals, i.e., data or signals that contribute to controlling the technical system TS.
[0025] Such control-relevant data or signals may include, in particular, control action signals, forecast data, monitoring signals, status data and / or classification data, which may be used, in particular, for operational optimization, monitoring or maintenance of the technical system TS and / or for wear or damage detection.
[0026] The gas turbine TS is equipped with sensors S that continuously measure one or more operating parameters of the technical system TS and output them as measured values. The measured values of the sensors S, as well as any other operating parameters of the technical system TS, are transmitted as operating signals BS from the technical system TS to the control unit CTL.
[0027] The operating signals BS can, in particular, include physical, chemical, control-related, effect-related, and / or design-related operating variables, property data, performance data, effect data, status signals, behavior signals, system data, default values, control data, control action signals, sensor data, measured values, environmental data, monitoring data, forecast data, analysis data, and / or other data arising from the operation of the technical system TS and / or describing an operating state or a control action of the technical system TS. This can, for example, be data on temperature, pressure, emissions, vibrations, oscillation states, or resource consumption of the technical system TS. Specifically in the case of a gas turbine, the operating signals BS can relate to turbine power, rotational speed, vibration frequencies, vibration amplitudes, combustion dynamics, combustion alternating pressure amplitudes, or nitrogen oxide concentrations.
[0028] Based on the operating signals BS, the trained control unit CTL determines control actions that optimize the performance of the technical system TS. The performance to be optimized can, in particular, relate to power, yield, speed, runtime, precision, error rate, resource requirements, efficiency, pollutant emissions, stability, wear, service life, and / or other target parameters of the technical system TS.
[0029] The determined, performance-optimizing control actions are initiated by the control unit CTL by transmitting corresponding control action signals AS to the technical system TS. These control actions can, for example, adjust a gas supply, gas distribution, or air supply in a gas turbine.
[0030] Figure 2shows a schematic representation of a learning-based control device CTL according to the invention in a first training phase. The control device CTL is to be configured to control a technical system TS. Where the same or corresponding reference symbols are used in the figures, these reference symbols denote the same or corresponding entities.
[0031] In the present exemplary embodiment, the control device CTL is coupled to the technical system TS. The control device CTL comprises one or more processors PROC for executing the method according to the invention and one or more memories MEM for storing method data.
[0032] The control device CTL receives operating signals BS from the technical system TS as training data. The operating signals contain, in particular, time series, i.e., temporal sequences of values of operating parameters of the technical system TS. In the present exemplary embodiment, the operating signals BS contain state signals SS specifying states of the technical system TS over time, control action signals AS specifying or initiating control actions of the technical system TS, and behavior signals VS specifying a system behavior of the technical system TS. The latter can, for example, specify changes in combustion pressure change amplitudes, emissions, a speed, or a temperature of a gas turbine. State signals of the technical system TS that are relevant, in particular, for the performance of the technical system TS can be recorded as behavior signals VS.
[0033] The operating signals BS can also be received or originate, at least in part, from a technical system similar to the technical system TS, from a database with stored operating signals of the technical system TS or a technical system similar thereto and / or from a simulation of the technical system TS or a technical system similar thereto.
[0034] The control device CTL further comprises a first machine learning module NN1, a second machine learning module NN2, and a third machine learning module NN3. A respective machine learning module NN1, NN2, or NN3 can be configured, in particular, as an artificial neural network or as a set of neural subnetworks. The first machine learning module NN1 can, in particular, be configured as a submodule of the third machine learning module NN3.
[0035] Preferably, the machine learning modules NN1, NN2, and / or NN3 can use or implement a supervised learning method, a reinforcement learning method, a recurrent neural network, a convolutional neural network, a Bayesian neural network, an autoencoder, a deep learning architecture, a support vector machine, a data-driven trainable regression model, a k-nearest neighbor classifier, a physical model, a decision tree, and / or a random forest. A variety of efficient implementations are available for the specified variants and their training.
[0036] Training is generally understood to mean the optimization of a mapping from input signals to output signals. This mapping is optimized during a training phase according to predefined, learned and / or yet-to-be-learned criteria. These criteria can be, for example, a prediction error for prediction models, a classification error for classification models, or the success of a control action for control models. Through training, for example, the network structures of neurons in the neural network and / or the weights of connections between the neurons can be adjusted or optimized so that the predefined criteria are met as closely as possible. Training can therefore be viewed as an optimization problem. A variety of efficient optimization methods are available for such optimization problems in the field of machine learning.In particular, gradient descent methods, particle swarm optimization and / or genetic optimization methods can be used.
[0037] In the Figure 2 In the first training phase illustrated, the first machine learning module NN1 is trained. This module is to be trained to predict or reproduce, based on an operating signal of the technical system TS, a behavior of the technical system TS that would develop without the current application of a control action.
[0038] To improve training success, the training data BS is filtered by a filter F1 coupled to the first machine learning module NN1 to preferentially obtain training data without control action or without the effects of a control action. For this purpose, the operating signals BS are fed into the filter F1. The filter F1 includes a control action detector ASD for detecting control actions in the operating signals BS based on the control action signals AS contained therein.
[0039] Depending on the detection of control actions by the control action detector ASD, initial portions of the operating signals BS are selected by the filter F1 and extracted from the operating signals BS. Priority is given to initial portions of the operating signals BS that do not contain a control action and / or do not contain any effects of a control action. For example, the initial portions of the operating signals BS can be extracted from a time window after a detected current control action, with the time window selected such that this control action cannot yet affect system behavior.
[0040] The filtered first parts of the operating signals BS comprise first parts SS1 of the state signals SS and first parts VS1 of the behavior signals VS. The first parts SS1 and VS1 are output by the filter F1 and used to train the first machine learning module NN1.
[0041] The first parts SS1 of the state signals SS are fed into the first machine learning module NN1 as an input signal for training. The aim of the training is for the first machine learning module to reproduce, based on an operating signal of the technical system TS, a behavioral signal of the technical system that arises without the current application of a control action as closely as possible. This means that an output signal VSR1 of the first machine learning module NN1, hereinafter referred to as the first output signal, matches the actual behavioral signal of the technical system TS as closely as possible. For this purpose, a deviation D1 is determined between the first output signal VSR1 and the corresponding first parts VS1 of the behavioral signals VS. The deviation D1 represents a reproduction or prediction error of the first machine learning module NN1.The deviation D1 can be calculated in particular as the square or amount of a difference, in particular a vector difference, according to D1 = (VS1-VSR1) 2< or D1 = |VS1-VSR1|.
[0042] The deviation D1 is calculated as in Figure 2indicated by a dashed arrow, is fed back to the first machine learning module NN1. Based on the fed-back deviation D1, the first machine learning module NN1 is trained to minimize this deviation D1 and thus the reproduction error. As already indicated above, a variety of optimization methods are available to minimize the deviation D1, such as gradient descent methods, particle swarm optimizations, or genetic optimization methods. In this way, the first machine learning module NN1 is trained using a supervised learning process. The trained first machine learning module NN1 uses the first output signal VSR1 to reproduce a behavioral signal of the technical system TS as it would occur without the current application of a control action.
[0043] By using the filtered operating signals SS1 and VS1 for training, the first machine learning module NN1 is trained particularly effectively for this training goal. It should also be noted that the first machine learning module NN1 can also be trained outside of the control unit CTL.
[0044] The above training method can be used particularly advantageously to separately reproduce multiple behavioral signals from different processes occurring in the technical system TS. For this purpose, the first machine learning module NN1 can comprise multiple process-specific neural subnetworks, each of which is trained separately or individually with process-specific behavioral signals, as described above. Such separate training often proves to be more efficient than combined training, since the underlying individual dynamics generally exhibit a simpler and / or more uniform response behavior.
[0045] Figure 3illustrates the control device CTL in a second training phase. In the second training phase, the second machine learning module NN2 is to be trained to predict or reproduce a behavior of the technical system TS induced by a respective control action based on an operating signal BS of the technical system TS, and in particular based on a control action signal AS contained therein.
[0046] To train the second machine learning module NN2, the control device CTL receives operating signals BS from the technical system TS as training data. As already mentioned above, the operating signals BS contain, in particular, time series of state signals SS, control action signals AS, and behavior signals VS. The trained first machine learning module NN1 is also used to train the second machine learning module NN2. In the present exemplary embodiment, the training of the first machine learning module NN1 has already been completed during the training of the second machine learning module NN2.
[0047] To improve training success, the training data BS are filtered by a filter F2 coupled to the second machine learning module NN2 in order to preferentially obtain training data that contains control actions or effects of control actions.
[0048] For this purpose, the operating data BS is fed into filter F2. Filter F2 contains a control action detector ASD for specifically detecting control actions in the operating signals BS based on the control action signals AS contained therein. Depending on the detection of control actions by the control action detector ASD, second parts of the operating signals BS are selected by filter F2 and extracted from the operating signals BS. Preferably, second parts of the operating signals BS are selected that include control actions and / or effects of control actions. For example, the second parts of the operating signals BS can be extracted from a time window around a respectively detected control action and / or from a time window in which an effect of the respective control action is to be expected.The filtered second parts of the operating signals BS include, in particular, second parts VS2 of the behavior signals VS and second parts AS2 enriched with control action signals. The second parts AS2 and VS2 of the operating signals BS are output by the filter F2 and used to train the second machine learning module NN2.
[0049] For training, the second parts AS2 of the operating signals BS are fed into the second machine learning module NN2 as an input signal. Furthermore, the operating signals BS are fed into the already trained first machine learning module NN1, which derives a behavior signal VSR1 from them and outputs it as the first output signal. As described above, the behavior signal VSR1 reproduces a behavior signal of the technical system TS as it would occur without the current application of a control action. The behavior signal VSR1 is fed into the second machine learning module NN2 as a further input signal.
[0050] The aim of the training is for the second machine learning module NN2 to reproduce, as accurately as possible, a behavioral signal of the technical system TS induced by the control actions based on an operating signal (here AS2) containing control actions, as well as a behavioral signal (here VSR1) generated without the current application of a control action. This means that an output signal (VSR2) of the second machine learning module NN2, hereinafter referred to as the second output signal, matches the actual behavioral signal of the technical system TS under the influence of control actions as closely as possible.
[0051] During training, a deviation D2 is determined between the second output signal VSR2 and the corresponding second parts VS2 of the behavioral signals VS. The deviation D2 represents a reproduction or prediction error of the second machine learning module NN2. The deviation D2 can be calculated, for example, as the square or absolute value of a difference, in particular a vector difference, according to D2 = (VS2-VSR2) 2< or D2 = |VS2-VSR2|.
[0052] The deviation D2 is calculated as in Figure 3 indicated by a dashed arrow, is fed back to the second machine learning module NN2. Using the fed-back deviation D2, the second machine learning module NN2 is trained to minimize this deviation D2 and thus the reproduction error. As already mentioned above, a variety of well-known optimization methods, particularly supervised learning, can be used to minimize the deviation D2.
[0053] The trained second machine learning module NN2 uses the second output signal VSR2 to reproduce a behavioral signal of the technical system TS that is induced by the current application of a control action.
[0054] By using the filtered operating signals AS2 and VS2 for training, the second machine learning module NN2 is trained particularly effectively for this training goal. Furthermore, by feeding the behavioral signal VSR1 into the second machine learning module NN2, the training success can be significantly increased in many cases, as the second machine learning module NN2 has specific information about the difference between a control-action-induced system behavior and a system behavior without control action. It should also be noted that the second machine learning module NN2 can also be trained outside of the control unit CTL.
[0055] The above training method can be used particularly advantageously to separately reproduce multiple behavioral signals of different processes occurring in the technical system TS. For this purpose, the second machine learning module NN2, like the first machine learning module NN1, can comprise multiple process-specific neural subnetworks, each of which is trained separately or individually with process-specific behavioral signals, as described above.
[0056] Figure 4illustrates the control unit CTL in a third training phase. In the third training phase, the third machine learning module NN3 is to be trained to generate a performance-optimizing control action signal for controlling the technical system based on an operating signal from the technical system TS. Optimization is also understood to mean approaching an optimum. By training the third machine learning module NN3, the control unit CTL is configured to control the technical system TS.
[0057] To train the third machine learning module NN3, the control device CTL receives operating signals BS from the technical system TS as training data. For this training, the first machine learning module NN1 and the second machine learning module NN2, which were trained as described above, are used in particular. In the present exemplary embodiment, the training of the machine learning modules NN1 and NN2 has already been completed during the training of the third machine learning module NN3.
[0058] In addition to the components described above, the control device CTL includes a performance evaluator EV coupled to the machine learning modules NN1, NN2, and NN3. Furthermore, the first machine learning module NN1 is coupled to the machine learning modules NN2 and NN3, and the second machine learning module NN2 is coupled to the third machine learning module NN3.
[0059] The performance evaluator EV is used to determine the performance of the behavior of the technical system TS triggered by a particular control action. For this purpose, a reward function Q is evaluated. The reward function Q determines and quantifies a reward, in this case the performance of a current system behavior already mentioned several times. Such a reward function is often also referred to as a cost function, loss function, objective function, reward function, or value function. The reward function Q can, for example, be implemented as a function of an operating state, a control action, and one or more target values OB for a system behavior.
[0060] If several behavioral signals are evaluated by the machine learning modules NN1, NN2, NN3 and / or by the performance evaluator EV, several behavioral signal target values OB can be specified specifically for each behavioral signal.
[0061] To train the third machine learning module NN3, the operating signals BS are fed as input signals into the trained machine learning modules NN1 and NN2 as well as into the third machine learning module NN3.
[0062] Based on the operating signals BS, the trained first machine learning module NN1 reproduces a behavior signal VSR1 of the technical system TS, as it would occur without the current application of a control action. The reproduced behavior signal VSR1 is fed from the first machine learning module NN1 to the second machine learning module NN2, the third machine learning module NN3, and the performance evaluator EV. In addition, one or more behavior signal target values OB are fed into the third machine learning module NN3 and the performance evaluator EV.
[0063] An output signal AS of the third machine learning module NN3 resulting from the operating signals BS, the reproduced behavior signals VSR1, and one or more behavior signal target values OB, hereinafter referred to as the third output signal, is further fed into the trained second machine learning module NN2 as an input signal. Based on the third output signal AS, the reproduced behavior signal VSR1, and the operating signals BS, the trained second machine learning module NN2 reproduces a control action-induced behavior signal VSR2 of the technical system TS, which is fed into the performance evaluator EV by the trained second machine learning module NN2.
[0064] The performance evaluator EV quantifies the current performance of the technical system TS based on the reproduced behavior signal VSR2, taking into account the reproduced first behavior signal VSR1 and the one or more behavior signal target values OB. In particular, the performance evaluator EV determines a first deviation of the control-action-induced behavior signal VSR2 from the one or more behavior signal target values OB. As the deviation increases, a reduced control action performance is generally determined. In addition, a second deviation between the control-action-induced behavior signal VSR2 and the behavior signal VSR1 is determined. Based on the second deviation, the performance evaluator EV can evaluate how system behavior with application of a control action differs from system behavior without application of this control action.It turns out that in many cases the performance evaluation can be significantly improved by this distinction.
[0065] The control action performance determined by the reward function Q is, as in Figure 4 indicated by a dashed arrow, is fed back to the third machine learning module NN3. Based on the fed-back control action performance, the third machine learning module NN3 is trained to maximize control action performance. As mentioned several times above, a variety of well-known optimization methods can be used to maximize control action performance.
[0066] To the extent that the second machine learning module NN2 specifically expects a control action signal as an input signal, the third machine learning module NN3 is implicitly trained to output such a control action signal, here AS. By optimizing the control action performance, the third machine learning module NN3 is thus trained to output a performance-optimizing control action signal AS.
[0067] By using the reproduced behavioral signal VSR1 in addition to the operating signal BS to train the third machine learning module NN3, the latter can be trained particularly effectively, since the third machine learning module NN3 has specific information about a control-action-free system behavior at its disposal.
[0068] A particular advantage of the invention is that when training the third machine learning module NN3, it is often sufficient for the performance evaluator EV to evaluate only a single, possibly adjustable, time step of the behavioral signals for a given point in time. A complex determination or estimation of future rewards is often unnecessary. Thus, effects occurring on different time scales can also be efficiently considered.
[0069] Furthermore, a given data set of operating signals can be used multiple times for training the third machine learning module NN3 by specifying varying behavior signal setpoints OB for the behavior signal VSR2. This makes it possible to learn different, setpoint-specific control action signals from the same operating signals, thus achieving better coverage of a control action space.
[0070] By training the third machine learning module NN3, the control device CTL is configured to control the technical system TS in a performance-optimizing manner using the control action signal AS of the trained third machine learning module NN3.
Claims
1. Computer-implemented method for configuring a control device (CTL) for a technical system (TS), wherein a) an operating signal (BS) of the technical system is fed into a first machine learning module (NN1) that is trained, on the basis of the operating signal (BS) of the technical system, to reproduce a behaviour signal of the technical system that arises specifically without current application of a control action and to output the reproduced behaviour signal (VSR1) as first output signal, b) the first output signal (VSR1) is fed into a second machine learning module (NN2) that is trained, on the basis of a control action signal (AS) that specifies or triggers a control action of the technical system, to reproduce a resultant behaviour signal of the technical system and to output the reproduced behaviour signal (VSR2) as second output signal, c) the operating signal (BS) of the technical system is fed into a third machine learning module (NN3), d) a third output signal (AS) of the third machine learning module (NN3) is fed into the trained second machine learning module (NN2), e) a control action performance (Q) is determined on the basis of the second output signal (VSR2), f) the third machine learning module (NN3) is trained to optimize the control action performance (Q) on the basis of the operating signal (BS) of the technical system, and g) the control device (CTL) is configured, on the basis of the third machine learning module (NN3), to control the technical system by way of an output signal (AS) of the trained third machine learning module (NN3).
2. Method according to Claim 1, characterized in that the third machine learning module (NN3) is also trained on the basis of the first output signal (VSR1).
3. Method according to either of the preceding claims, characterized in that first (SS1, VS1) and second (AS2, VS2) parts of an operating signal (BS) of the technical system are selected specifically according to whether or not they comprise a control action, and in that the first parts (SS1, VS1) of the operating signal (BS) that do not comprise a control action are used specifically to train the first machine learning module (NN1) and the second parts (AS2, VS2) of the operating signal (BS) that comprise a control action are used specifically to train the second machine learning module (NN2).
4. Method according to one of the preceding claims, characterized in that a behaviour signal setpoint value (OB) is read in, in that the second output signal (VSR2) is compared with the behaviour signal setpoint value (OB), and in that the control action performance (Q) is determined depending on the comparison result.
5. Method according to Claim 4, characterized in that the behaviour signal setpoint value (OB) is fed into the third machine learning module (NN3), and in that the third machine learning module (NN3) is trained also to optimize the control action performance (Q) on the basis of the behaviour signal setpoint value (OB).
6. Method according to one of the preceding claims, characterized in that the control action performance (Q) is also determined on the basis of the first output signal (VSR1).
7. Method according to Claim 6, characterized in that a deviation between the first output signal (VSR1) and the second output signal (VSR2) is determined, and in that the control action performance (Q) is determined depending on the deviation.
8. Method according to one of the preceding claims, characterized in that the first (NN1) and / or the second (NN2) machine learning module is trained to separately reproduce multiple behaviour signals of different processes running in the technical system, and in that the control action performance (Q) is determined depending on the reproduced behaviour signals.
9. Method according to Claim 8, characterized in that a specific behaviour signal setpoint value (OB) is read in for a respective behaviour signal (VS), and in that the control action performance (Q) is determined on the basis of a comparison between the reproduced behaviour signals and the specific behaviour signal setpoint values.
10. Method according to Claim 9, characterized in that the third machine learning module (NN3) is trained to optimize the control action performance (Q) on the basis of the specific behaviour signal setpoint values (OB).
11. Control device (CTL) for controlling a technical system (TS), comprising means for carrying out a method according to one of the preceding claims.
12. Computer program product, comprising commands that cause the control device (CTL) according to Claim 11 to carry out the method steps according to one of Claims 1 to 10.
13. Computer-readable storage medium containing a computer program product according to Claim 12.