Controlling a machine using a learning-based control device
The integration of a dynamic weighting mechanism in learning-based control systems balances performance and predefined rules, addressing reliability and safety issues in machine control by adapting to untested conditions, ensuring efficient and safe operation.
Patent Information
- Application Number
- PCT/EP2025/050751
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2025-01-14
- Publication Date
- 2025-08-14
AI Technical Summary
Existing learning-based control systems for machines are unreliable and risky due to untested output signals and difficulty in assessing their consequences, especially under conditions not covered by training data, leading to potential unsafe operation.
A method and device that integrates a learning-based control system with a weighting mechanism to balance performance optimization and adherence to predefined action selection rules, using a trained control agent to determine operating control signals based on machine state signals and error variables, adjusting the weighting dynamically to ensure safe and efficient operation.
The system provides safe and efficient control by automatically adjusting the balance between performance and adherence to predefined rules, minimizing risks in unforeseen conditions, thus enhancing the reliability and safety of machine operation.
Smart Images

Figure EP2025050751_14082025_PF_FP_ABST
Abstract
Description
[0001] Description
[0002] Controlling a machine with a learning-based control device
[0003] The invention relates to a computer-implemented method for controlling a machine using a learning-based control device. Operating status signals of the machine are provided to the previously trained control device, and operating control signals are then fed from the control device to the machine.
[0004] Data-driven machine learning methods, particularly reinforcement learning methods, are increasingly being used to control complex technical systems such as robots, engines, manufacturing plants, power generation facilities, gas turbines, wind turbines, steam turbines, milling machines, or other machines. Control agents, particularly artificial neural networks, are trained using a large amount of training data to generate a control signal optimized for a given state of the technical system with respect to a given target function. Control devices for controlling technical systems are therefore used in the form of trained control agents, also known as policies.
[0005] The objective function used during training can be used to evaluate, in particular, a performance of the technical system to be optimized, e.g., power, efficiency, resource consumption, yield, pollutant emissions, product quality, and / or other operating parameters of the technical system. Such an objective function is often also referred to as a reward function, cost function, or loss function. Using such a performance-optimizing control agent, the performance of a technical system controlled by it can, in many cases, be significantly increased. However, the productive use of such a control agent also entails risks because the output operating status signals are often untested and thus potentially unreliable, or their consequences for the state of the technical system are difficult to assess.
[0006] This is especially true under operating conditions that are not adequately covered by the training data. Therefore, it is being considered to orient the control system at least partially on predefined action selection rules. This would involve rules specifying which operating control signal should be issued when a specific operating state signal is present. This could, for example, be a control sequence previously performed by a user.
[0007] The invention is based on the object of demonstrating a method for controlling by means of a learning-based control device that enables efficient and safe control of a machine.
[0008] This object is achieved by a method having the features of claim 1. Furthermore, the invention relates to a corresponding control device, a corresponding computer program, a corresponding computer-readable, preferably non-volatile, storage medium, and a corresponding transmission signal. Advantageous embodiments and further developments are the subject of subclaims.
[0009] In the inventive computer-implemented method for controlling a machine by means of a trained learning-based control device, the control device has previously been trained using a target function, wherein this target function includes a weighting of the machine's performance against a deviation from a specific action selection rule. In the method for controlling the machine, the previously trained control device determines and outputs operating control signals based on the machine's operating state signals and weighting values provided to it. These weighting values are determined by the control device by calculating an error variable relating to a prediction of state signals by the control device compared to the machine's operating state signals.
[0010] The control device knows the state of the machine insofar as it is described by the most recently received operating state signal. Based on this and a weighting value, the control device determines an action to be performed on or by the machine, described by the operating control signal output by the control device. This determination of a suitable operating control signal is possible because the control device learned during training how to determine operating control signals that optimize the objective function. The objective function used during training includes two variables: the performance of the machine and a deviation from a specific action selection rule. The latter is a specification of which operating control signal should be issued when a certain operating state signal is present.The goal here is to maximize performance and minimize deviations from the specific action selection rule. Since these two goals are not equally important in every situation, they are weighted so that, depending on the weighting value used, one of the two goals can be emphasized more than the other, or even one of the two goals can be ignored, and only the other goal is considered.
[0011] The control unit also uses a weighting value to control the machine. Based on this value, the machine is controlled based on the two objectives described above. The weighting value used for this purpose can be changed over time. To do this, it is necessary to specify the weighting value to be used. This determination is made by the control unit, i.e., the control unit itself determines which weighting of the two objectives is to be used. However, this does not preclude a user from temporarily intervening and specifying the weighting.
[0012] To determine the weighting value, the control unit calculates an error value. This error value is used to compare a prediction of state signals by the control unit with the machine's operating state signals. This is based on the fact that the control system is more effective the better the control unit can predict the machine's future state when, based on a state corresponding to a specific operating state signal, an action corresponding to a specific operating control signal is performed. In addition to the prediction of the state signals, the prediction of the machine's performance can also be included in the error value.
[0013] The controlled machine can be, for example, a robot, a motor, a manufacturing plant, a factory, a power plant, a gas turbine, a wind turbine, a steam turbine, or a milling machine. The control device may be integrated into the machine.
[0014] Preferably, the error magnitude is used when determining the weighting values in such a way that an increasing value of the error magnitude corresponds to an increasing significance of the deviation from the specific action selection rule in the objective function. As already explained, a large value for the error magnitude means that the control device's predictions of the operating state signals are not good. It can therefore be assumed that no or insufficient training has taken place for these machine states. Accordingly, in such a situation, it is advantageous to implement the control in such a way that the goal of a small deviation from the specific action selection rule is pursued predominantly or even exclusively. This is because the action selection rule can correspond to a predetermined control sequence which, in the respective situation, demonstrates safe and low-risk control of the machine.Conversely, a decreasing error magnitude can correspond to a decreasing significance of the deviation from the specific action selection rule in the objective function, meaning that the machine's performance can be increasingly weighted. This corresponds to the principle that if the control system can accurately predict the machine's state, it can be trusted to act autonomously, oriented toward achieving the highest possible machine performance.
[0015] To calculate the error magnitude, consecutive past control steps can be considered, i.e., the control sequence that actually occurred - preferably in the recent past. For each control step, there is an operating state signal and a subsequent operating control signal output by the control device. For a specific control step, a difference is determined between the operating state signal of the control step following the specific control step and a state signal predicted by the control device based on the operating state signal and the operating control signal of the specific control step. This error magnitude can be calculated for several consecutive control steps, and the several calculated error magnitudes can then be combined to form a total error magnitude.In this way, a number N of recent control steps can be considered, and based on this history of control steps, a decision can be made as to which weighting of the two objectives should be applied. It is advantageous if, when summarizing, the multiple calculated error variables are weighted in such a way that control steps further in the past receive a lower weighting. This allows for the fact that circumstances have changed since the beginning of the history under consideration, which is why the control steps further in the past should only play a minor role in the decision regarding the design of the values currently to be used for the weighting of the two objectives. The weighting can be determined using a parameter between 0 and 1.Furthermore, the objective function can be a weighted linear combination of the machine's performance and the deviation from the action selection rule, where the weighting factors of the weighted linear combination are the parameter and the number 1 minus the parameter. This means that with a parameter value of 0 or 1, only one or the other objective is included in the objective function, while with a parameter value of 0.5, both objectives are considered equally. This parameter is then the value for the weighting, which is determined and applied by the control device.
[0016] When calculating the linear combination, the performance can be defined as an overall performance calculated over several control steps and the deviation can be defined as an overall deviation from the action selection rule calculated over the several control steps.
[0017] Two components of the control device can be used to train the machine, namely: a performance evaluator that determines the performance of the machine based on an operating state signal and an operating control signal, and an action evaluator that determines a deviation from the specific action selection rule based on the operating state signal and the operating control signal.
[0018] These two components can be learning-based agents that have been previously trained with the same training data that is also used to train the controller.
[0019] The method according to the invention and / or one or more functions, features and / or steps of the method according to the invention and / or one of its embodiments can be carried out with computer support. It can be carried out or implemented, for example, using one or more computers, processors, application-specific integrated circuits (ASICs), digital signal processors (DSPs) and / or so-called "field programmable gate arrays" (FPGAs). It can also be carried out at least partially in a cloud and / or in an edge computing environment. One or more interacting computer programs are used for the computer-supported process. If multiple programs are used, they can be stored and executed together on one computer, or on different computers at different locations.Since this is functionally equivalent, “the computer program” and “the computer” are formulated in the singular.
[0020] The invention is explained in more detail below using an exemplary embodiment. In the following:
[0021] Figure 1: a control device when controlling a machine,
[0022] Figure 2: a training of a control agent of the control device,
[0023] Figure 3: a flowchart.
[0024] Figure 1 illustrates a control device CTL controlling a machine M, e.g. a robot, a motor, a production plant, a factory, an energy supply facility, a gas turbine, a wind turbine, a steam turbine, a milling machine or another device or another plant. In particular, a component or subsystem of a machine or plant can also be understood as a machine M. The machine M has sensors SK for the preferably continuous detection and / or measurement of system states or subsystem states of the machine M and, if applicable, also of its environment. The control device CTL is shown in Figure 1 external to the machine M and coupled to it. Alternatively, the control device CTL can also be fully or partially integrated into the machine M.
[0025] The control device CTL has one or more processors PROC for executing method steps of the control device CTL as well as one or more memories MEM coupled to the processor PROC for storing data, e.g. information to be processed by the control device CTL, information received by the control device CTL, information output by the control device CTL. Furthermore, the control device CTL has a learning-based control agent POL, which is trained or trainable using reinforcement learning methods. Such a control agent is often also referred to as a policy or, for short, as an agent. In the present exemplary embodiment, the control agent POL is implemented as an artificial neural network. Furthermore, the control device CTL has a learning-based transition model NN and a learning-based action evaluator VAE.The control agent POL, the transition model NN, and the action evaluator VAE, and thus the control unit CTL, are trained in advance using predefined training data and thus configured for optimized control of the machine M. The transition model NN and the action evaluator VAE are required only for the purpose of training the control agent POL; the machine M is then controlled by the trained control agent POL, without the transition model NN and the action evaluator VAE being used for this purpose. An exception to this is the calculation explained below by the component CALC W, which accesses the transition model NN.
[0026] The training of the control agent POL is aimed in particular at two control criteria. On the one hand, the machine M controlled by the control agent POL should achieve the highest possible performance, while on the other hand, at least under certain circumstances, the control should not deviate too much from a reference policy. The performance of the machine M can, for example, relate to power, efficiency, resource consumption, yield, pollutant emissions, product quality and / or other operating parameters of the machine M. The reference policy is a specific action selection rule, i.e. the specification of control signals that should be issued when certain state signals are present. A reference policy can originate from a past control of the machine M, a similar machine, or a simulation of the machine M while it was controlled by a reference control agent.A verified, validated, and / or rule-based control agent can preferably be used as the reference control agent. The reference policy is included in the training data, as explained in Figure 2.
[0027] For optimized control of the machine M by the trained control agent POL, operating state signals SO, i.e. state signals determined during ongoing operation of the machine M, are transmitted to the control device CTL. The operating state signals SO each specify a state, in particular an operating state of the machine M and, if applicable, its environment, and are preferably each represented by a numerical state value or vector or a time series of state values or vectors. The operating state signals SO can represent measurement data, sensor data, environmental data, or other data arising during the operation of the machine M or influencing its operation. Examples include data on a temperature, a pressure, a setting, an actuator position, a valve position, a pollutant emission, a utilization, a resource consumption, and / or a performance of the machine M or its components.In a production plant, the operating status signals SO can also relate to product quality or another product property. The operating status signals SO can be measured at least partially by the sensor system SK and / or determined by simulation using a simulator of the machine M.
[0028] Furthermore, the control device CTL uses a weight value W. The weight value W weights two variables against each other, namely the performance of the machine M on the one hand and a deviation of the control by the control agent POL from the reference policy on the other hand.
[0029] The operating state signals SO transmitted to the control device CTL, together with the weight value W, are fed into the trained control agent POL as input signals. Based on a respective supplied operating state signal SO and weight value W, the trained control agent POL generates a respective optimized operating control signal AO. The latter specifies one or more control actions that can be performed on the machine M. The generated operating control signals AO are transmitted from the control agent POL or from the control device CTL to the machine M. The transmitted operating control signals AO control the machine M in an optimized manner by executing the control actions specified by the operating control signals AO on the machine M.
[0030] By means of the weight value W, the trained control agent POL can be adjusted during the ongoing operation of the machine M to both weight proximity to the reference policy higher than the performance of the machine M when determining optimized operating control signals AO, and conversely to weight proximity to the reference policy lower than the performance of the machine M. How the weight value W is determined for this purpose is explained further below.
[0031] First, the control agent POL is trained using the training data TD from the database DB. This training data contains a large number of data points, where each data point is a tuple consisting of a state signal, a control signal, a reward for executing the control action corresponding to the control signal based on the respective state signal—this is equal to the machine's performance after executing the corresponding control action—and a subsequent state resulting from the application of the respective control action corresponding to the control signal. The same training data TD was previously used to train the transition model NN and the action evaluator VAE. The training data contains the reference policy, with each data point containing a state signal and a corresponding control signal. It is possible that the entire reference policy is more extensive than the portion contained in the training data.This is because the training data may not cover the entire state / action space, so that there are pairs of state signals and control signals produced by the generator of the reference policy that are not part of the training data.
[0032] The control agent POL is trained in such a way that a relative weighting of the two optimization criteria, performance of the machine M and proximity to the reference policy, can be changed by the weight value W during the ongoing operation of the trained control agent POL. A sequence of this training is explained in more detail with reference to Figure 2. This figure illustrates a training of the control agent POL using the trained transition model NN and the trained action evaluator VAE. To illustrate successive work steps, several instances of the control agent POL, the trained transition model NN, and the trained action evaluator VAE are shown in Figure 2. The different instances can, in particular, correspond to different calls or evaluations of routines by means of which the control agent POL, the trained transition model NN, or the trained action evaluator VAE are implemented.
[0033] The trained transition model NN, which can be implemented in particular as an artificial neural network, was trained to predict, based on a respective state signal and a respective control signal, a subsequent state of the machine M resulting from the application of the corresponding control action, as well as a resulting performance value of the machine M, the reward for the application of the corresponding control action, as accurately as possible. The training of such a transition model NN is described, for example, in publications EP 3940596 A1 and EP 4235317 A1. Instead of a transition model NN, other models can also be used to evaluate the reward, e.g., a value function.The trained action evaluator VAE, which can in particular be a variational autoencoder implemented as a feed-forward neural network, evaluates arbitrary pairs of a state signal and a control signal by determining a reproduction error D for the pair. The reproduction error D is a measure of how well a state signal-control signal pair is covered by the training data TD, or how frequently or how likely it occurs there. Since the training data TD contains the reference policy, as explained above, the reproduction error indicates how strongly the state-control signal pair deviates from the reference policy. The training of such an action evaluator VAE is described, for example, in the publications EP 3940596 A1 and EP 4235317 A1, where the reference policy is referred to as the predetermined control sequence.
[0034] As already mentioned above, the control agent POL should be trained to output a control signal A for a respective state signal S of the machine M, which is optimized, on the one hand, with regard to the resulting performance of the machine M and, on the other hand, with regard to a proximity to or deviation from the reference policy. Obviously, proximity to the reference policy can also be represented by a negatively weighted deviation from the reference policy, and vice versa. The optimization aims at higher performance and greater proximity to or smaller deviation from the reference policy. The weighting of the two optimization criteria, performance and proximity or deviation, can be adjusted using the weight value W.
[0035] The weight value W can be set between 0 and 1. With a weight value of W=1, the trained control agent POL should output a control signal A that exclusively optimizes performance, whereas with a weight value of W=0, a control signal A that exclusively minimizes deviations from the reference policy should be output. With weight values between 0 and 1, the two optimization criteria should be weighted proportionally accordingly. To train the control agent POL for different weight values, a generator GEN of the control device CTL generates a plurality of weight values W, preferably randomly, which lie in the interval from 0 to 1. The weight values W generated by the generator GEN are fed, as can be seen in Figure 2, both into the control agent POL to be trained and into an objective function TF to be optimized through training.To train the control agent POL, a large number of state signals S from the training data TD are fed into the control agent POL as input signals. In parallel, these state signals S are also fed into the trained transition model NN and the trained action evaluator VAE as input signals. Based on the respective state signal S, the trained transition model NN predicts the subsequent states of the machine M resulting from the application of a control signal A. Furthermore, the pair formed by the respective state signal S and the corresponding control signal A is evaluated by the trained action evaluator VAE.
[0036] The control agent POL derives an output signal from the respective state signal S and outputs it as control signal A. The control signal A, together with the respective state signal S, is then fed into the trained transition model NN, which predicts a subsequent state from it and outputs a subsequent state signal S1 specifying this, as well as an associated performance value R1. Furthermore, the control signal A, along with the respective state signal S, is fed into the trained action evaluator VAE, which determines and outputs a reproduction error DO for the pair of state signal S and control signal A. As explained above, the reproduction error can be viewed as a measure of a deviation of control signals from the reference policy.
[0037] The subsequent state signal S1 is in turn fed to the control agent POL, which derives a next control signal A1 for this subsequent state. The control signal A1, together with the subsequent state signal S1, is fed into the trained transition model NN, which predicts a further subsequent state and outputs a subsequent state signal S2 specifying this, along with a corresponding performance value R2. Furthermore, the control signal A1, together with the subsequent state signal S1, is fed into the trained action evaluator VAE, which determines and outputs a reproduction error D1 for the pair of state signal S1 and control signal A1.
[0038] The above steps can be repeated iteratively, determining performance values and reproduction errors for subsequent states. The iteration can be terminated if a termination condition is met, e.g., if a specified number of iterations is exceeded. In this way, a control trajectory spanning several time steps, progressing from subsequent state to subsequent state, and extrapolated into the future can be determined, with associated performance values R1, R2, ... and reproduction errors DO, D1, ... Such an extrapolation is often referred to as a rollout or virtual rollout.
[0039] From the performance values R1, R2, ... of a respective control trajectory, a cumulative total performance RET of this control trajectory over several time steps is determined. Such cumulative total performance is referred to as return in reinforcement learning. The total performance RET is preferably assigned to the respective state signal S at the beginning of the respective control trajectory and thus evaluates the ability of the control agent POL to determine a control signal A for the respective state signal S, which initiates a control sequence with high performance over several time steps. To determine the total performance RET, the performance values R1, R2, ... determined for future time steps are discounted, i.e., assigned weights that decrease for each time step. As a concrete example, the total performance RET can be expressed as the weighted sum of the performance values R1, R2, ...whose weights correspond to a discount factor < 1 that decreases with each journal into the future. In this way, the overall performance RET can be determined, for example, according to RET = R1 + R2*G + R3*G. 2 + R4*G 3 + .... For example, a value of 0.99, 0.9, 0.8, or 0.5 can be used for G. Alternatively, the overall performance can be calculated without discounting or using a discount factor of 1.
[0040] The transition model NN and the above discounting method together form the performance evaluator PEV, which uses control signals A, A1, ... and state signals S, S1, ... to determine the overall performance RET of the machine M resulting from the application of the control signals. Alternatively or additionally, the performance evaluator PEV can also be implemented using a Q-learning method and trained to determine the overall performance RET accumulated over a future period.
[0041] Furthermore, a total reproduction error D accumulated over several time steps is determined from the reproduction errors DO, D1, ... of the respective control trajectory. The latter serves in the further process as a measure of the deviation of this control trajectory from the reference policy. In the present embodiment, the total reproduction error D is determined as the sum of the individual reproduction errors DO, D1, ... according to D = D0 + D1 + .... The respectively determined total performance RET and the respectively determined total reproduction error D are both fed into the objective function TF to be optimized. The objective function TF weights the total performance RET and the total reproduction error D with the respective weight value W, which was also fed into the control agent POL together with the respective status signal S.By means of a training that optimizes the objective function TF, the control agent POL can be trained to output a control signal A when a state signal S and a weight value W are input, which optimizes the objective function TF according to the input weight value W, at least on average.
[0042] In the present embodiment, the objective function TF determines a target value TV of the objective function TF from the overall performance RET, the overall reproduction error D, and the weight value W according to TV = W*RET - (1-W)*D. Since the goal is to ensure that the control trajectory is as close as possible to the reference policy, the overall reproduction error D is included in the target value TV with a negative sign. If necessary, the overall performance RET and / or the overall reproduction error D can be provided with a constant normalization factor before calculating the target value TV.
[0043] The determined target value TV is assigned to the respective state signal S at the beginning of the respective control trajectory. The respective target value TV thus evaluates the current ability of the control agent POL to initiate a control sequence for the respective state signal S according to the weight value W that is both performance-optimizing and close to the reference policy. To train the control agent POL, i.e., to optimize its processing parameters, the determined target values TV are fed back to the control agent POL—as indicated by a dashed arrow in Figure 2. The processing parameters of the control agent POL are set or configured such that the target values TV are maximized, at least on average.To carry out the training, a variety of efficient standard methods can be used, in particular stochastic gradient descent, population-based optimization methods, or other methods of supervised learning.
[0044] After successful training of the control agent POL, it can be used for the optimized control of the machine M, as described in Figure 1. For this purpose, the trained control agent POL requires an operating state signal and a weight value W and outputs an operating control signal based on these. The weighting of the optimization criteria performance and proximity to the reference policy can be set and changed during operation. In this way, the control of the machine M can be adjusted depending on whether performance or reliability should be given greater weight. A suitable setting of this weighting during runtime, i.e. after training and deployment of the control device CTL, is described below. For this purpose, the component CALC W of the control device CTL is used, which calculates the weight value W and makes it available to the control agent POL for controlling the machine M.
[0045] As described, the objective function TF used during training is a linear combination of the overall performance RET, which represents the return of the control agent POL, and the overall reproduction error D, which represents a penalty term. The factor that sets the balance between these two components is the weight value W, which thus represents a hyperparameter controlling the balance. Since no interaction with the machine M takes place before the start of the actual use of the control agent POL, the deployment, it is difficult to estimate which weight value W is the most suitable. In particular, a user will have difficulty specifying a weight value W to be used by the control agent POL.This typically leads to either an overly conservative approach, in which control is strongly oriented towards the reference policy; this means that no improvement in control occurs compared to the known process. Otherwise, an overly risky approach would be taken, in which orientation towards the reference policy is neglected and instead only the machine's performance is focused, which can be dangerous for state signals that are not strongly represented in the training data. It is precisely such state signals that occur rarely in training that can correspond to rather rarely occurring safety-critical operating phases of the machine M. In these situations, it is sensible to follow the reference policy, which shows a way out of the safety-critical operating phase.
[0046] Therefore, the weight value W should be automatically adjusted to achieve the best possible balance between the two goals of high performance and adherence to the reference policy. This adjustment of the weight value W is better than choosing a constant value because, depending on the state of the machine M and its environment, a different control approach proves advantageous: if the current situation corresponds to a state that was frequently presented to the control unit CTL during training, the control agent POL can take a riskier approach and achieve a higher return, since it can be assumed that the control unit CTL knows how to proceed based on the training it has undergone.If, however, the current situation corresponds to a state that is not or hardly known to the control device CTL from the training, a conservative approach should be taken and the machine M should be controlled as strictly as possible according to the reference policy.
[0047] To enable automatic adjustment of the weight value W, information is used that becomes available during runtime. This includes, on the one hand, the operating status signals SO of Figure 1, indicated in the following formulas with i, where i indicates the time. Furthermore, it includes the operating control signals AO of Figure 1, indicated in the following formulas with a t where i indicates the time. Finally, the reward resulting from a control signal AO is used, denoted by r in the following formulas. twhere i indicates time. The reward in the training explained in Figure 2 corresponds to the values R1, R2, ...; however, in contrast to the training, the reward is not predicted by the transition model NN, but results from the measured state of the machine.
[0048] With h t is the history of interactions during runtime over a certain number t of journals: h t =< >. How long the history under consideration is, ie how large t is, depends on the inertia of the system.
[0049] For example, t can be approximately 10 when effects occur quickly, and approximately 1000 when they occur very slowly. The length of a time step, i.e. the time interval between one triple s^a^rt and the following triple s i+1 , a i+1 , r i+1 can be seconds to minutes, for example.
[0050] Part of the control unit CTL is the trained transition model NN. To determine the weight value W, the extent to which the actually observed transitions between the state signals match the predictions of the transition model NN is checked. This is done by comparing the actual state signals with those determined by the transition model NN based on the previous state signal and the control signal output by the control agent POL:
[0051] Here, t stands for the trained transition model NN and e t gives the total squared error related to the history h t t (s^ai ) is the state signal which the transition model NN generates from a state signal-control signal pair predicts. Sj +1is the next state signal actually determined at the machine, following the state signal Sj. Using the bracket expression, the predicted state signal is compared with the state signal corresponding to the actually occurred state. N is the number of complete triples s^a^rt in h t .
[0052] This total squared error e t can be used as a basis for further action. Alternatively, in addition to checking the predictions of the transition model NN with respect to the state signals, it is also possible to check the predictions of the transition model NN with respect to the reward r t This can be realized mathematically either by considering the reward instead of the state signals in the above formula, ie by using a pair consisting of state signal and reward, or by defining a separate squared total error e for the reward.t according to the calculation of the total error e t for the state signal and these two total errors are then combined.
[0053] If the calculated total error is low, the NN transition model is able to accurately estimate the machine's behavior during the period of the observed history. However, if the calculated total error is high, this is not the case. This total error should now be used to set a suitable weight value W. Since the error is bounded by zero at the bottom but has no upper limit, a mapping of the total error from the range (0, oo), within which the value range can lie based on the above formula, to the range (0, 1) of the weight value W is calculated.
[0054] Such a mapping of the areas can be done via A = e t) = 1 / (1 + <? t ). Small values for the error e tlead to values for A close to 1, while high error values bring A close to 0. Thus, the quantity A can be used as the weight value W. Because with small values for the error e t The transition model works reliably, which is why attempts should be made to enable free optimization in the control system, ie the highest possible performance of the machine is sought and the orientation towards the reference policy is of little importance. For large values for the error e t the transition model does not work reliably, which is why a control of the machine that is strongly focused on performance would be risky; in this case, the weight value W causes the control to be regularized by proceeding strongly according to the reference policy.
[0055] It is possible that in history h tThere are status signals that no longer accurately describe the current situation. This is even more true the older the values are, i.e., the further they are removed from the current point in time. According to the notation of history h t =< s Q , a Q , r Q , S a , r 1; ... , s t-1 , a t-1 , r t-1 , s t > This affects more status signals the closer they are to the beginning of the history. If expert knowledge exists about the number of status signals relevant to the present, the history can be limited to these most recent points in time. Since this is usually not the case, individual errors can be weighted using a decreasing exponential function:
[0056] The further the time of the status signal is at the beginning of the history, the smaller the value of i - t and thus the exponential function, and the less the respective individual error contributes to the total error e t T is a constant parameter that determines the strength of the exponential function's decline. The above statements regarding the total error related to the inclusion of the reward r t apply accordingly.
[0057] As described, the CALC W component of the control device CTL calculates the weight value W and makes it available to the control agent POL. The calculation of the weight value can be done continuously by using the history as a sliding window, so that at each new point in time a new T ripel s t , a t , r tis placed at the end of the history, and the triple previously at the beginning is deleted. Thus, a new value for the weight value W can be determined at each new point in time and made available to the control agent POL. Alternatively, it is possible to determine the weight value W less frequently, e.g., by waiting t points in time until a completely new history is available and performing a calculation based on this new history.
[0058] The described procedure enables automatic and independent calculation of the weight value W in the CTL control unit, without requiring user interaction. However, user intervention by specifying a weight value W is possible. Therefore, if a user considers this type of control to be disadvantageous, they can take control by specifying a weight value W to be used via a user interface III of the CTL control unit.
[0059] Because the user is not required to specify the weight value W for machine control, this saves the precious time of an expert who would have to laboriously and extensively attempt to determine and enter a weight value adapted to the respective situation. Instead, the weight value is determined automatically and repeated continuously or from time to time based on the updated history. The user can enter the frequency of recalculations of the weight value via the user interface U1. In this way, the user's observation that conditions adapt at a changing rate can be taken into account, for example.
[0060] In summary, Figure 3 shows a flowchart of the described procedure. In a training step TRAIN 1, which precedes the use of the control unit CTL to control the machine, the transition model NN and the action evaluator VAE are first trained using training data. In a further upstream training step TRAIN 2, the control agent POL is then trained using the same training data and with the help of the trained transition model NN and the trained action evaluator VAE. After deployment, in the BEGIN step, the control unit begins controlling the machine with the three trained units: the transition model NN, the action evaluator VAE, and the control agent POL. After an initial period, the data of the history h tBased on this, the weight value W is calculated in the CALCULATE step, which is then used in the regular control mode OPERATE for control by the control agent POL, which outputs control signals AO to the machine. This provides an updated history h t which in turn can be used in the CALCULATE step to calculate an updated weight value W, etc.
[0061] It can be advantageous for the initial period in which no history has yet been t To define a behavior, you could, for example, initially use W=0 to be on the safe side, ie, to strictly follow the reference policy, or a value like W=0.5 if security is not a major concern.
[0062] The invention has been described above using an exemplary embodiment. It is understood that numerous changes and modifications are possible without departing from the scope of the invention.
Claims
Patent claims 1. Computer-implemented method for controlling a machine (M) by means of a trained learning-based control device (CTL), wherein the training of the control device (CTL) was carried out using an objective function which includes a weighting of a performance of the machine (M) against a deviation from a specific action selection rule, in which the previously trained control device (CTL) for controlling the machine (M) determines and outputs operating control signals (AO) based on operating state signals (SO) of the machine (M) made available to it and values for the weighting, wherein these values for the weighting are determined by the control device (CTL) by calculating an error variable relating to a prediction of state signals by the control device (CTL) compared to operating state signals (SO) of the machine (M).
2. The method according to claim 1, wherein the error magnitude is used in determining the values for the weighting such that an increasing value of the error magnitude corresponds to an increasing significance of the deviation from the determined action selection rule in the objective function.
3. Method according to claim 1 or 2, in which successive past control steps are considered to calculate the error size, wherein for each control step there is an operating state signal (SO) and an operating control signal (AO) subsequently output by the control device (CTL), for a specific control step a difference is determined between the operating state signal (SO) of the control step following the specific control step and a state signal predicted by the control device (CTL) based on the operating state signal (SO) and the operating control signal (AO) of the specific control step.
4. Method according to claim 3, wherein the error magnitude is calculated for several consecutive control steps, and the several calculated error sizes are combined into a total error size.
5. Method according to claim 4, wherein, when combining, the plurality of calculated error variables is weighted such that control steps further in the past receive a lower weighting.
6. Method according to one of claims 1 to 5, in which the weighting is carried out by means of a parameter between 0 and 1, and the objective function is a weighted linear combination of the performance of the machine (M) and the deviation from the action selection rule, the weighting factors of the weighted linear combination being the parameter and the number 1 minus the parameter.
7. Method according to one of claims 1 to 6, wherein, when calculating the linear combination, a total performance calculated over several control steps is used as the performance and a total deviation from the action selection rule calculated over the several control steps is used as the deviation.
8. Method according to one of claims 1 to 7, in which a performance evaluator (PEV) is used for training, which determines a performance of the machine (M) on the basis of an operating state signal (SO) and an operating control signal (AO), and an action evaluator (VAE) is used, which determines a deviation from the determined action selection rule on the basis of the operating state signal (SO) and the operating control signal (AO).
9. Method according to one of the preceding claims, characterized in that the machine (M) is a robot, a motor, a manufacturing plant, a factory, a power plant, a gas turbine, a wind turbine, a steam turbine, a milling machine or another device or installation.
10. Control device (CTL) for controlling a machine (M), arranged to carry out a method according to one of claims 1 to 9.
11. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the steps of the method according to any one of claims 1 to 9.
12. A computer-readable storage medium (MEM) with a computer program according to claim 11.
13. A transmission signal transmitting the computer program according to claim 11.
Citation Information
Patent Citations
Method for configuring a control agent for a technical system and control device
EP3940596A1
Method for controlling a machine by means of a learning based control agent and control device
EP4235317A1
Method for training a control policy for controlling a technical system
US20240037393A1