Method and apparatus for training a machine learning system

Through machine learning methods, the optimization problems in model prediction adjustment are approximately solved, and the relaxation auxiliary conditions of the slack variable are used to solve the problem that it is difficult to solve the optimization problems in the existing technology in real time, realizing the safety and efficiency of the system.

CN120068982APending Publication Date: 2025-05-30ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411735251.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When using model prediction and adjustment, it is difficult to solve optimization problems accurately in real time, resulting in suboptimal solutions of manipulated signals, affecting the safety and efficiency of the system.

Method used

Through machine learning methods, the optimization problems in model prediction adjustment are approximately solved, and the relaxation auxiliary conditions of the slack variable are used to improve the stability of the system state and the accuracy of the manipulation signal.

Benefits of technology

Real-time accurate solution determination in model prediction and regulation is realized, ensuring the safety and efficiency of the system, and improving the robustness of the system performance under the low requirements of resources and computing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068982A_ABST
    Figure CN120068982A_ABST
Patent Text Reader

Abstract

The invention relates to a method and apparatus for training a machine learning system. A computer-implemented method (300) for training a machine learning system (60) for use in model predictive conditioning of a technical system, the machine learning system (60) being designed to determine a value with respect to an operating state and / or an environmental state of the technical system to be conditioned, the value characterizing a quality of the state, wherein the method comprises the following steps: determining a plurality of operating states and / or environmental states of the technical system to be adjusted; determining a quality value (Vi) of a state (xi) of the plurality of operating states and / or environmental states, wherein the quality value (Vi) characterizes a quality of the state (xi) with respect to the model predictive adjustment; the machine learning system (60) is trained by supervised training of the machine learning system (60), wherein the state (xi) is used as an input to the machine learning system (60) and the determined quality value (Vi) is used as a desired output to the machine learning system (60).
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The present invention relates to a method for training a machine learning system, a method for performing model predictive regulation of a technical system, a training system, a regulation system, a computer program, and a machine-readable storage medium. Background Art

[0002] "Soft Constrained Model Predictive Control With Robust Stability Guarantees" by Zeilinger et al., 2014, https: / / ieeexplore.ieee.org / abstract / document / 6730917 / discloses a method for performing model predictive regulation of a technical system.

[0003] Advantages of the Invention

[0004] Model predictive control (MPC) is a key method for regulating technical systems. In the case of known schemes for performing model predictive regulation, the optimal regulation problem is formulated and solved at each sampling time point. In the case of this optimization problem, starting from the current state measurement of the system and using the model of the system, the future behavior of the system over a finite time horizon is predicted. The advantages of model predictive control over other advanced regulation methods are that model predictive control can optimally handle systems with multiple inputs and outputs and can systematically incorporate system constraints. These two advantages are used, for example, in the motion control of a vehicle. For example, a vehicle must comply with road restrictions, and steering, acceleration, and braking must be coordinated to ensure safe and efficient operation. Model predictive control has proven its high efficiency in many studies and at the same time provides strict safety guarantees.

[0005] When using model predictive control, typically an optimal regulation problem is defined at the current time point and state, and the optimal regulation problem determines an optimal sequence of control signals for the system for a finite number of future states by means of the model of the system to be regulated. For this purpose, typically the optimization problem is solved at each time point and thus for each state, where the optimization is performed by (or with respect to) the control signal. The main challenge results from the following requirement, namely that in order to regulate a technical system, the optimization problem must be solved in real time in order to be able to ensure sufficiently precise regulation.

[0006] Known solutions for solving optimization problems work iteratively. A disadvantage of these solutions is that the number of iterations available is typically not sufficient to determine the exact solution of the optimization problem and at the same time maintain real-time capabilities. The effect of the application of a manipulation signal determined based on such a suboptimal solution of the optimization problem usually does not allow an estimation of how the system to be regulated will behave, especially with regard to the safety-critical behavior of the system.

[0007] The inventors may surprisingly find that the solution of the optimization problem for model predictive control can be approximated extremely precisely by machine learning methods. This enables the exact determination of the solution at runtime of the model predictive control. The inventors are furthermore able to determine that, by approximating the adapted optimization problem under slack variables, the stability of the manipulation signal can be ensured for the states of the technical system, and the compliance with the auxiliary conditions defined in the original optimization problem of the model predictive control can be ensured. Summary of the Invention

[0008] In a first aspect, the invention relates to a computer-implemented method for training a machine learning system for use in model predictive control of a technical system, wherein the machine learning system is configured to determine values characterizing the quality of a state with respect to the operating state and / or the environmental state of the technical system to be regulated, and wherein the method comprises the following steps:

[0009] · Determining a plurality of operating states and / or environmental states of the technical system to be regulated;

[0010] · Determining quality values of the states in the plurality of operating states and / or environmental states, wherein the quality values characterize the quality of the states with respect to model predictive control;

[0011] · Training the machine learning system by supervised training of the machine learning system, wherein the states are used as inputs to the machine learning system and the determined quality values are used as the desired outputs of the machine learning system.

[0012] The operating state and / or the system state are also referred to as the state hereinafter. In the sense of regulation technology, the state of a technical system can in particular be understood as being fully measurable or at least observable. For example, it may be possible that at least part of the state is measured by means of sensors of the system to be regulated or is derived based on sensor recordings (also referred to as indirect measurement). For example, it may be possible that a mobile robot should be controlled, which records images of its environment by means of a camera. Based on these images, it can be determined as at least part of the state of the robot how far the robot deviates from a defined travel path in its environment.

[0013] After training, in particular, a machine learning system can be used within a policy to evaluate the quality of future states. Based on this evaluation, a control signal can then be selected for the technical system, the control signal having the best quality with respect to a plurality of possible control signals.

[0014] The method for training can be understood as first collecting training data and then using the training data to train a machine learning system in a supervised manner. A data set of tuples is typically used for supervised training, where the tuples each include an input signal and a desired output signal. During supervised training, the machine learning system is adapted such that when the machine learning system receives an input signal as input, the machine learning system outputs an output signal that is as close as possible to the desired output signal. For the method for training, a state can be understood as the input signal used for training. Such an input signal can in particular be determined by recording the state of the system during operation of the system. Alternatively, it is also possible to generate the state synthetically, for example based on a scan of physically meaningful values of the state.

[0015] For training, a quality measure is then determined for at least one state. The quality measure can be understood as an annotation or a desired output signal in the sense of supervised training. The quality measure is determined in this case with respect to the model predictive control.

[0016] In particular, a quality value can be determined based on a value function of the model predictive control, where the value function determines the quality of the state with respect to the state and the value function furthermore includes auxiliary conditions, where the auxiliary conditions characterize physical constraints of the state and / or the auxiliary conditions characterize constraints of the control signal, which the model predictive control can select.

[0017] For example, the state can characterize the acceleration of the system, where, determined by the electric motor used in the system, the acceleration cannot exceed a certain value. In this case, the auxiliary condition can characterize this value.

[0018] The states used for this method can especially be determined by the system, and the system should be adjusted in a model-prediction manner based on a machine learning system. For example, the system can be placed in a test run where another adjustment is used, and typical states are determined by means of the test run. The system can also be placed in unusual operating conditions in order to thereby determine states that simulate the marginal regions of typical adjustments. Alternatively or additionally, it is also possible to determine the possible states of the system in a simulation manner, for example, by means of a computer simulation of the system. The possibilities for determining states mentioned here can equally be applied to systems with the same or similar structures. For example, it is possible to use a prototype or a system with the same or similar structure as the system to be adjusted in order to thereby determine states that may approximately be like this or similarly also occur in this system.

[0019] The state can especially be understood as a set of values that describe the physical parameters of the system and / or the environment of the system.

[0020] Preferably, the state of the system can be characterized by a vector of real numbers. Thus, the expression "determine multiple operating states and / or environmental states of the technical system to be adjusted" can also equivalently be replaced by "determine multiple states of the technical system to be adjusted". Each individual state among the multiple states can be understood as a vector of values, where each value characterizes the operating state or the environmental state of the system.

[0021] The value function can especially be the solution of an optimization convex function, where this function characterizes the cost of the regulation technical problem, and thus this function can be understood as a cost function.

[0022] Advantageously, the optimized solution can be determined before using the model predictive control of the system (English: offline). Thereby, it is possible to determine the quality value very precisely because there is no time requirement for the real-time ability of the optimization.

[0023] Subsequently, it is advantageously to train a machine learning method for predicting the quality value with respect to the state. This prediction can be understood as an approximate solution of the optimization problem. The inventors were able to determine that even when using a machine learning system with low requirements for resources and computing power, the prediction of the quality value is still precise enough for model predictive control.

[0024] The optimization can especially be minimization, where the quality value determined by minimization can be understood as: a lower value has higher quality. Thus, the quality value can be understood as a cost value, where a lower cost has higher quality, in other words, a state with a lower cost is better than a second state with a higher cost.

[0025] In various embodiments of the present invention, the value function may include an auxiliary condition that permits a numerical deviation of the state from the physical constraints by means of a slack variable.

[0026] By means of the slack variable (Schlupfvariablen), the auxiliary condition of the optimization problem is relaxed in the sense of mathematical optimization. Thereby, even for demanding regulation tasks, a sufficiently precise solution to the optimization problem can be advantageously determined. The inventors were able to determine that the performance of the machine learning system is improved thereby, since the slack variable enables a higher generalization ability of the machine learning system through the relaxation. In particular, it is also possible that there are multiple auxiliary conditions with slack variables.

[0027] In particular, it is advantageous if the cost function includes a term that represents the sum of the values of the slack variables. Advantageously, the solution to the optimization problem can thus take into account that the slack variables cannot become arbitrarily large, and thus solutions far outside the permitted constraints of the auxiliary conditions are determined. In particular, this term can be multiplied by a factor that can be understood as a hyperparameter of the optimization problem and controls the influence of the slack variables.

[0028] In the case of a vector state, the constraints and the slack variables can also exist in vector form in particular. In these cases, the terms of the cost function can in particular include the sum of the lengths or squared lengths of the slack variables.

[0029] In particular, it is possible that the auxiliary condition of the optimization problem is additionally a factor that strengthens the physical constraints of the state.

[0030] The inventors were able to advantageously determine that the strengthening of the constraints results in the later model predictive regulation for which the machine learning system is used becoming more robust with respect to the errors that occur through the approximation of the optimization problem by the machine learning system.

[0031] In particular, if instead of the exact solution of the optimization problem a machine learning system is used to determine the quality value in the model predictive regulation, it can be shown by the factor used for the slack variables and the factor used for the physical constraints that the auxiliary conditions are adhered to even under the errors that occur when the quality value is approximated by the machine learning system.

[0032] The optimization problem is preferably characterized by the following formula

[0033]

[0034] u.d.Nx 0|k =x(k),

[0035] x N|k ∈aX f

[0036]

[0037] 0 ≤ α ≤ 1,

[0038]

[0039] 0 ≤ ξ i|k ,

[0040] x i+1|k = Ax i|k + Bu i|k ,

[0041] H u u i|k ≤ h u ,

[0042] H x x i|k ≤ h x (1 - η) + ξ i|k + ξ N|k

[0043] Starting from the current state x, the optimization variables are: the sequence of the control signals of the system determined by model predictive control the sequence of the subsequent states predicted from the sequence of the control signals and optionally but preferably the sequence of the slack variables The auxiliary conditions of the optimization problem can be relaxed by using the slack variables. The cost function J is preferably characterized by the following formula

[0044]

[0045] l(x, u) = x T Qx + u T Ru

[0046] The matrix H u 、H x 、Q and R are the hyperparameters of the optimization problem. The index |k indicates: related to the time points starting from the time point k. Thus, for example, x i|k is the predicted state for the time point i predicted at the time point k. The auxiliary condition H u u i|k ≤ h u characterizes the constraint of the control signal at each time point i starting from the time point k. The auxiliary condition

[0047] H x x i|k ≤ h x (1 - η) + ξ i|k + ξ N|k

[0048] denotes a preferred auxiliary condition, which further strengthens the state x through the factor η i|k constraint h x . The factors η and ρ can be understood as interacting hyperparameters that trade off the degree of relaxation of the optimization problem (through ρ) against the robustness with respect to the approximation error (through η). Equation x i+1|k = Ax i|k + Bu i|k characterizes the linear model of the system to be regulated, which is based on the current state x i|k and the control signal u i|k to predict the state transition from time point i to the subsequent time point i + 1, i.e., to forecast the subsequent state. The variable N characterizes the prediction horizon of the model predictive control, in other words how far into the future the model predictive control looks.

[0049] In this method, it can be stipulated that the first quality value corresponds to the result of the value function at the position of the state. In other words, it can be stipulated that the machine learning system predicts the result of the value function in a single step (end-to-end).

[0050] However, preferably, it can be stipulated that the machine learning system predicts two parts of the value function and thus divides the prediction.

[0051] In particular, it can thus be stipulated that the first quality value corresponds to the first part of the value function at the position of the state, where the first part does not contain slack variables, and in addition, the machine learning system determines a second quality value, where the second quality value corresponds to the second part of the value function at the position of the state, the second part containing slack variables, and in the training step, the second quality value is used as another desired output of the machine learning system.

[0052] The inventors were able to determine that the value function typically behaves such that from a certain state value (which is related to the system to be regulated), the value function has a greatly increased value, because the slack variables become "activated" for this part of the state space, i.e., the optimization has to utilize the slack variables in order to solve the optimization. Thereby, a greatly increased local Lipschitz constant and curvature of the function result, which makes it difficult to learn the value function by the machine learning system.

[0053] Advantageously, dividing the value function into a part with slack variables and a part without slack variables allows two separate parts of the value function to be learned, which together describe the value function, where each part on its own has a locally small Lipschitz constant and curvature. This significantly simplifies the learning problem, and the machine learning system can better predict the value of the value function for specific states.

[0054] For training, it is preferably first necessary to solve the optimization problem and only perform the partitioning of the value function using the values obtained in this way. In other words, it is preferably possible to determine the training data in such a way that state values are selected separately, the optimization problem is solved for this state value and subsequently the value function is partitioned with respect to the values obtained into a part containing slack variables and a part not containing slack variables.

[0055] Accordingly, the first quality value can preferably be determined according to the following formula:

[0056]

[0057] where the value x represents in this case the value that has been determined by optimizing V(x).

[0058] The second quality value can preferably be described by the following formula

[0059]

[0060] where x again represents the value that has been determined by optimizing V(x).

[0061] Furthermore, it is preferably possible to provide that the machine learning system comprises two machine learning models, in particular two neural networks, where the first model of the two models is set up to predict the first quality value based on the state as input to the first model, and the second model of the two models is set up to predict the second quality value based on the state as input to the second model.

[0062] In other words, there are preferably corresponding models in order to be able to predict the first or second quality value. As set out above, the inventors have been able to determine that the two models can be trained more easily separately for predicting the values of the value function than a single model.

[0063] Here, the two models do not necessarily have to be present in the same computer, but can also be present in a distributed environment on different computers. Accordingly, the machine learning system is not limited to a single computer, but can also be provided by a distributed system.

[0064] Alternatively, it is also conceivable that the machine learning system comprises one machine learning model, in particular a neural network, where the model is set up to predict the first quality value and the second quality value.

[0065] If the model is constructed as a neural network, the model can in particular comprise two so-called heads which each predict the first or second quality value.

[0066] A machine learning system that is set up to predict two quality values can in particular be trained based on a first loss function and a second loss function, where the first loss function characterizes the difference between the prediction of the first quality value and the first quality value and the second loss function characterizes the difference between the prediction of the second quality value and the second quality value.

[0067] In other words, the machine learning system can be trained in a supervised manner in view of the first and second quality values, where there is a loss function for each quality value separately.

[0068] For training, in particular, a dataset of respective states and pairs of quality values determined for that state by optimization can be used. For an embodiment with two quality values, after optimization, the quality values for the part of the value function with slack variables and the quality values for the part of the value function without slack variables can be determined separately. In these cases, the dataset preferably consists of tuples of three elements: a state value, a first quality value, and a second quality value.

[0069] Alternatively, the dataset can also be divided into a dataset of tuples consisting of a state value and a first quality value respectively and another dataset of tuples consisting of a state value and a second quality value respectively.

[0070] In particular, the second loss function, i.e., the loss function with respect to the quality value (including the slack variable component), can be characterized by the following formula:

[0071]

[0072] where (x, y) ∈ D slack,+ respectively characterizes all tuples including the state value x and the positive second quality value y, and (x, y) ∈ D slack,0 respectively characterizes all tuples having the state value x and the second quality value y equal to zero. In other words, the subset D slack,0 includes the states for which the slack variable is not activated.

[0073] Advantageously, thus if the output of the second quality value is negative, it is not penalized at the neural network. The inventors were able to determine that the model for the second quality value to be learned by the machine learning system can thereby be further simplified, and thus the second quality value can be predicted more precisely.

[0074] In various embodiments of the method for training, it is possible that, in the training step, the machine learning system is trained at least until an approximation quality criterion is met, where the approximation quality criterion characterizes the numerical value (Betrag) of the difference between the value determined by the value function for a state and the value determined by the machine learning system for that state.

[0075] The approximate quality criterion can be understood as evaluating a machine learning system in terms of the extent to which it can determine the same value for a state as has already been determined by means of a value function. As long as the approximation is not good enough, the machine learning system can be further trained. If the quality criterion should not have been met after a predefined number of iterations, it is additionally possible, for example, to collect additional training data for the method or to increase the machine learning system in terms of the number of its degrees of freedom, for example by increasing the number of parameters of the machine learning system. If a neural network is used as the machine learning system, the architecture of the neural network can in particular be changed such that the neural network becomes deeper and / or the layers of the neural network have more parameters.

[0076] Advantageously, the inventors were able to determine that when the approximate quality criterion is met, it can be ensured that the auxiliary conditions of the value function are complied with despite the approximation error when using the machine learning system. This results in the fact that the found approximation can enable a provably safe regulation of the technical system with the aid of the machine learning system.

[0077] On the other hand, the invention relates to a computer-implemented method for performing model predictive control of a technical system, comprising the following steps:

[0078] · Determining the operating state and / or environmental state of the system;

[0079] · Determining a control signal for the system, wherein the control signal is determined based on a control law, wherein the control law comprises the term where characterizes the output of the machine learning system when it receives the input f(x, u), where f is a model of the system, preferably a linear model, and predicts the next operating state and / or system state of the system based on the determined operating state and / or environmental state x and the control signal u, wherein the machine learning system has moreover been trained according to a method for training;

[0080] · Actuating the system with the determined control signal.

[0081] The method for performing model predictive control described above describes the case where the machine learning system directly, i.e., end-to-end, predicts the result of the value function.

[0082] The concept of a control law is generally better known by its English term "policy". Basically, a policy can be understood as a function that determines a control signal for the state of the system to be controlled, by means of which the system should be controlled.

[0083] In a method for performing model predictive control, the strategy can in particular be formulated as another optimization problem. As the manipulation signal, the one can be selected which, when applied to the current state, results in a state with the highest possible quality.

[0084] To determine the next state resulting from the manipulation signal based on the current state, a linear model of the following form can in particular be used

[0085] x(k + 1) = f(x(k), u(k)) = Ax(k) + Bu(k),

[0086] where x(k) is the state at time point k, u(k) is the manipulation signal at time point k, and A and B are matrices.

[0087] Advantageously, the execution of the method can be significantly accelerated by using a machine learning system, since it is no longer necessary to perform the optimization under auxiliary conditions to determine the quality of the state, but rather the quality can be determined with the aid of a machine learning system.

[0088] The optimization problem in the method for performing model predictive control can furthermore include a term characterizing the amplitude of the manipulation signal.

[0089] Advantageously, if both terms occur in the optimization problem, it is thus possible to weigh the quality of the state to be achieved against the influence exerted by the control on the system. This prevents the system from possibly oscillating due to excessive manipulation signals, but can also lead to a softer control, for example when controlling an electric motor, which contributes to less load and wear on the electric motor and thus on the system to be controlled.

[0090] On the other hand, the invention relates to a computer-implemented method for performing model predictive control of a technical system, comprising the following steps:

[0091] · Determining the operating state and / or environmental state of the system;

[0092] · Determining a manipulation signal for the system, where the manipulation signal is determined based on a control law, where the control law includes a term where is determined according to the following formula

[0093]

[0094] where characterizes a first quality value output by a machine learning system for the input f(x, u) and Characterize a second quality value output by a machine learning system for the same input and output, where f is the model of the system, preferably a linear model, and predicts the next operating state and / or system state of the system based on the determined operating state and / or environmental state x and the control signal u, where the machine learning system has in addition been trained according to the method described above for training a machine learning system to output a first quality value and a second quality value.

[0095] This method is basically equivalent to the previous method for predictive regulation, except that it includes the specific structure of the machine learning system for outputting the first quality value and the second quality value.

[0096] In the above method, the machine learning system can in particular be a neural network or include a neural network.

[0097] The inventors were able to determine that the use of a neural network leads to the best approximation of the quality of the state. However, it is also possible to use other machine learning methods. Especially in the case of simpler regulation tasks with few auxiliary conditions or a small state space or when used on embedded hardware, more resource-saving methods such as linear regression or support vector machines can still be used. Description of the Drawings

[0098] The embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. In the drawings:

[0099] Figure 1 Schematically shows a method for training a machine learning system;

[0100] Figure 2 Schematically shows a training system;

[0101] Figure 3 Schematically shows the structure of a regulation system for regulating a technical system;

[0102] Figure 4 Schematically shows an embodiment for regulating at least a partially autonomous robot;

[0103] Figure 5 Schematically shows an embodiment for regulating a production system. Detailed Description of the Embodiments

[0104] Figure 1 Schematically shows the flow of a method (300) for training a machine learning system, where the machine learning system can be used in model predictive regulation of a technical system after training.

[0105] In a first step (301), a plurality of states of a technical system to be adjusted are determined. Preferably, the states can be determined by operating the technical system, wherein at different points in time during operation, at least one state of the system is preferably determined at regular intervals. The state can in particular be understood as the operating state and / or the environmental state of the system. Preferably, a plurality of states are determined with the aid of sensors, so that the states are preferably each the result of a measurement by a sensor. Alternatively, the states can be determined based on indirect measurements. An indirect measurement can be understood as a measurement that is determined from a physical measurement by a sensor or a previous indirect measurement and cannot be directly identified from the measured values in the original measurement. For example, the environmental state of the system can characterize the position of an object in the environment of the system. For example, an object can be recorded as an image by means of a camera, and the position of the object can be determined based on object detection in the image, internal and external camera parameters. In this case, the image can be understood as a direct measurement, the value of which is the pixel value of the image, while the position of the object is understood as an indirect measurement derived from the (directly measured) pixel values.

[0106] In particular, the state can also characterize a plurality of physical properties. For example, the state can characterize the position of various objects in the environment of the system and / or the position of the system.

[0107] Additionally or alternatively, a plurality of states can also be determined based on simulations. For example, a physical model of the system can be provided, which then simulates the course of change of the state of the system. Alternatively or additionally, it is also possible to scan the state space by rastering and use the states determined in this way among the plurality of states.

[0108] Alternatively or additionally, it is also possible to determine the states among the plurality of states based on a prototype of the system, a system with the same or similar structure.

[0109] In a second step (302), a quality value is determined for each of the determined states, preferably for all of the determined states, wherein the quality value characterizes the quality with respect to the state adjusted by model prediction. Preferably, the quality value is determined based on an optimization problem, wherein the optimization problem characterizes optimal model predictive control. The optimization problem is preferably characterized by the following formula

[0110]

[0111] u.d.Nx 0|k = x(k),

[0112] x N|k ∈aX f

[0113]

[0114] 0 ≤ α ≤ 1,

[0115]

[0116] 0 ≤ ξ i|k ,

[0117] x i+1|k = Ax i|k + Bu i|k ,

[0118] H u u i|k ≤ h u ,

[0119] H x x i|k ≤ h x (1 - η) + ξ i|k + ξ N|k

[0120] Starting from the current state x, the optimization variables are: the sequence of the system's manipulation signals determined by model predictive control the sequence of the subsequent states predicted from the sequence of the manipulation signals and optionally but preferably the sequence of the slack variables Using the slack variables, the auxiliary conditions of the optimization problem can be relaxed. The cost function J is preferably characterized by the following formula

[0121]

[0122] l(x, u) = x T Qx + u T Ru

[0123] The matrices H u 、H x 、Q and R are the hyperparameters of the optimization problem. The index |k indicates: related to the time points starting from the time point k. Thus, for example, x i|k is the predicted state for the time point i predicted at the time point k. The auxiliary condition H u u i|k ≤ h u characterizes the constraint of the manipulation signal at each time point i starting from the time point k. The auxiliary condition

[0124] H x x i|k ≤ h x (1 - η) + ξ i|k + ξ N|k

[0125] represents the preferred auxiliary condition, and the preferred auxiliary condition further strengthens the constraint h of the state x by the factor η i|k of hx The factors η and ρ can be understood as interacting hyperparameters that trade off the degree of relaxation of the optimization problem (through ρ) against the robustness with respect to the approximation error (through η). Equation x i+1|k = Ax i|k + Bu i|k characterizes a linear model of the system to be regulated, which is based on the current state x i|k and the control signal u i|k to predict the state transition from time point i to the subsequent time point i + 1, i.e., to predict the subsequent state. The variable N characterizes the prediction horizon of the model predictive control, in other words how far into the future the model predictive control looks ahead.

[0126] In this embodiment, the state is normalized such that the state can respectively characterize the deviation from the desired state. In other words, the aim of the optimization can in particular be: to reach a state corresponding to the zero vector. It is also possible to regulate to other states. For example, the desired state can be transformed using an appropriate transformation such that the transformed state is regulated to the zero vector, and thus the regulation point is determined overall by the transformation.

[0127] For the determined states, a quality value can be determined respectively as the result of the optimization with respect to the state. Preferably, quality values are thus determined for all the determined states, thereby determining a data set that respectively includes pairs of a state and the associated quality value.

[0128] This data set can be used in a third step (303) for the supervised training of a machine learning system.

[0129] Optionally, it is also possible to train the machine learning system (60) batch by batch by an iterative method (English: batched training). In this case, subsets of states can also be determined respectively from the set of states, and a quality value can be determined for each state before the thus determined subsets (this optional embodiment is indicated by a dashed line in Figure 1 ).

[0130] Preferably, the machine learning system (60) is trained at least until an approximation quality criterion is met, where the approximation quality criterion characterizes the numerical difference between the value determined for a state by the value function and the value determined for the state by the machine learning system. In particular, the approximation quality criterion can be characterized by the following equation

[0131]

[0132] V max = N(||Q||·||X|| + ||R||·||U||) + ||Q||·||X||,

[0133]

[0134] wherein is the prediction of the machine learning system (60) for the state x. Preferably, the quality criterion can also be simplified, and a subset of the states of the state space can be selected instead of considering all states (represented by the expression ). For example, the subset can be determined by scanning the state space.

[0135] Alternatively, it is also possible that the machine learning system (60) is respectively set up to determine a first quality value and a second quality value. In this example, the data set can in particular be partitioned such that after the above optimization for a specific state, the first part of the value function is provided as the first quality value, where the first part includes the terms of the value function that do not include slack variables respectively, and the second part of the value function is provided as the second quality value, where the second part includes the terms of the value function that include slack variables respectively. Regarding the preferred optimization problem and the function J, the quality values can preferably be determined according to the following formula

[0136]

[0137] where the value x represents the value that has been determined by optimizing V(x) in this case.

[0138] The second quality value can preferably be described by the following formula

[0139]

[0140] where x again represents the value that has been determined by optimizing V(x).

[0141] Thus, the data set can include, according to the tuples of the state, the first and the second quality values:

[0142]

[0143] where n characterizes the number of data points in the data set. Alternatively, it is also possible to partition the data set into two data sets:

[0144]

[0145] Then the machine learning system can be trained based on two loss functions

[0146]

[0147] where is the prediction of the machine learning system (60) regarding the first quality value, is a prediction of the machine learning system (60) regarding a second quality value, c is a positive hyperparameter of the training method, D slack,+ ={(x, y) ∈ D slack | y > 0} and D slack,+ ={(x, y) ∈ D slack | y = 0}.

[0148] The use of two loss functions completely and advantageously avoids the problem of learning rapidly changing values when transitioning from a slack variable equal to zero to a slack variable greater than zero. This is achieved in such a way that the machine learning system (60) can also predict negative values for the second quality value. These negative values can be very simply filtered out during subsequent adjustment by approximating the complete result of the value function by the following formula

[0149]

[0150] Figure 2 shows an embodiment of a training system (140) which is set up to carry out the third step (303) of the method (300) shown in Figure 1 . The machine learning system (60) is trained with the aid of a training data set (T) which respectively comprises tuples of a state (x i ) and an associated quality value (V i ).

[0151] For training, the training data unit (150) accesses a computer-implemented database (St 2 ), where the database (St 2 ) provides the training data set (T). The training data unit (150) preferably randomly determines at least one state (x i ) and the quality value (V i ) corresponding to the state (x i ) from the training data set (T), and transmits the state (x i ) to the machine learning system (60). The machine learning system (60) determines a prediction of the quality value (V i ) based on the state (x i )

[0152] The quality value (V i ) and the prediction are transmitted to the change unit (180).

[0153] Based on the quality value (V i ) and the prediction ​Then, the change unit (180) determines new parameters (Φ') for the machine learning system (60). To this end, the change unit (180) compares the quality value (V i ) and the prediction using a loss function. The loss function determines a first loss value that characterizes how much the quality value (V i ) deviates from the prediction . In this embodiment, a function that represents the squared distance between the quality value (V i ) and the prediction is selected as the loss function. In alternative embodiments, other loss functions are conceivable.

[0154] The change unit (180) determines the new parameters (Φ') based on the first loss value. In this embodiment, this occurs by means of the gradient descent method, preferably stochastic gradient descent, Adam, or AdamW. In other embodiments, the training can also be based on an evolutionary algorithm or second-order optimization.

[0155] The determined new parameters (Φ') are stored in the model parameter memory (St 1 ). Preferably, the determined new parameters (Φ') are provided to the machine learning system (60) as parameters (Φ).

[0156] In other preferred embodiments, the described training is repeated iteratively for a predefined number of iteration steps or iteratively until the first loss value does not exceed a predefined threshold. Alternatively or additionally, it is also conceivable to end the training when the average first loss value with respect to a test or validation data set does not exceed a predefined threshold. In at least one iteration of the iteration, the new parameters (Φ') determined in the previous iteration are used as the parameters (Φ) of the machine learning system (60).

[0157] Furthermore, the training system (140) can include at least one processor (145) and at least one machine-readable storage medium (146) that contains instructions that, when executed by the processor (145), cause the training system (140) to perform a training method according to one of the aspects of the present invention.

[0158] Figure 3The regulating system (40) is shown, which determines the control signal (u) of the system with the aid of a machine learning system (60). In this embodiment, the control signal (u) represents the control signal of the actuator (10) or the display device (10a) of the technical system. At preferably regular time intervals, at least one operating state and / or environmental state of the technical system is detected in the sensor (30), which can also be given by a plurality of sensors. The sensor signal (S) of the sensor (30) - or in the case of a plurality of sensors, each one sensor signal (S) - is transmitted to the regulating system (40). The regulating system (40) thus receives a sequence of sensor signals (S). The regulating system (40) determines the control signal (u) therefrom, which is transmitted to the actuator (10).

[0159] The regulating system (40) receives the sequence of sensor signals (S) of the sensor (30) in an optional receiving unit (50), which converts the sequence of sensor signals (S) into a sequence of states (x) (alternatively, the sensor signal (S) can also be directly used as the state (x) at any time). The state (x) can be, for example, a segment or further processing of the sensor signal (S), or the state (x) can be determined on the basis of the sensor signal (S) by means of an indirect measurement. In other words, the state (x) is determined according to the sensor signal (S). The state (x) is transmitted to the regulation law module (π).

[0160] The regulation law module (π) determines the control signal (u) on the basis of the state (x) and the machine learning system (60). The regulation module preferably determines the control signal (u) on the basis of a regulation law, which is preferably characterized by the following formula

[0161]

[0162] term prediction of the quality value (V) characterizing the state reached when the control signal u is executed in the state x In other words, the quality of the state reached when the control signal u is executed in the state x.

[0163] If the machine learning system (60) should be used for regulation and the machine learning system predicts a first quality value and a second quality value as described above, it can preferably be determined by the following formula

[0164]

[0165] preferably by the same model

[0166] x(k + 1) = f(x(k), u(k)) = Ax(k) + Bx(k)

[0167] to determine the reached state, the model having been used to determine a quality value (V i ). Prediction is performed by a machine learning system (60) In other words, by means of the machine learning system (60), a quality value is determined for the subsequent state predicted by the linear model.

[0168] The regulation law includes a preferred term u T Ru, the preferred term additionally causing a trade-off between reaching the subsequent state with the highest quality value and the manipulation signal (u) to be applied. The matrix R represents a hyperparameter which can preferably be deleted or can be set to the identity matrix. The optimization problem can be solved, for example, using derivative-free and / or parallelizable methods, such as sampling or scanning of the search space U of the possible manipulation signals (u). The solution of the optimization problem is then provided as the manipulation signal (u) to the actuator (10) and / or the display device (10a).

[0169] The machine learning system (60) is preferably parameterized by parameters (Φ) stored in a parameter memory (P) and provided by the parameter memory (P).

[0170] The actuator (10) receives the manipulation signal (A), is accordingly manipulated and performs the corresponding action. The actuator (10) can in this case include (not necessarily structurally integrated) manipulation logic which determines a second manipulation signal from the manipulation signal (A) and then uses the second manipulation signal to manipulate the actuator (10).

[0171] In other embodiments, the regulation system (40) includes a sensor (30). In still other embodiments, the regulation system (40) alternatively or additionally further includes an actuator (10).

[0172] In other preferred embodiments, the regulation system (40) includes at least one processor (45) and at least one machine-readable storage medium (46) on which commands are stored which, if executed on the at least one processor (45), cause the regulation system (40) to perform the method according to the invention.

[0173] Figure 4 It is shown how the regulation system (40) can be used to control at least a partially autonomous robot, here at least a partially autonomous motor vehicle (100).

[0174] The sensor (30) can be, for example, a video sensor preferably arranged in a motor vehicle (100). The conversion unit (50) can in particular be set up such that it determines the drivable area, such as the lanes of a road or a multi-lane road. The state (x) can then in particular characterize the deviation of the robot (100) from the desired position on the drivable area, such as the deviation of the midpoint of the robot from the midpoint of the drivable area. The state (x) can additionally include the longitudinal and lateral accelerations of the robot. The state can additionally characterize the slip (Schlupf) of the propulsion means of the robot, such as the tires or chains of the robot. Furthermore, it is possible that the state includes the positions of other objects in the environment of the robot (100), such as traffic participants. The optimization problem that has been solved in order to train the machine learning system (60) can then in particular include a secondary condition characterizing the minimum distance of the robot from other objects.

[0175] The actuator (10) preferably arranged in the robot (100) can be, for example, a brake, a drive or a steering device of the robot (100). The control signal (A) can then be determined such that the one actuator or the plurality of actuators (10) is controlled such that the robot (100) moves, for example, along the center line of the drivable area. Additionally, it is possible that the one or more actuators (10) are controlled such that a collision with other objects in the environment of the robot (100) is avoided.

[0176] Alternatively or additionally, the display unit (10a) can be controlled by means of the control signal (A) and, for example, represent the identified objects and / or display the planned travel trajectory of the robot (100).

[0177] Alternatively, the at least partially autonomous robot can also be other mobile robots (not shown), such as robots that move forward by flying, swimming, diving or walking. The mobile robot can also be, for example, an at least partially autonomous lawn mower or an at least partially autonomous cleaning robot. Even in these cases, the control signal (A) can be determined such that the drive means and / or the steering means of the mobile robot are controlled such that the at least partially autonomous robot, for example, prevents a collision with the objects identified by the machine learning system (60).

[0178] Figure 5 An embodiment is shown in which an adjustment system (40) is used to adjust a production machine (11) of a production system (200) by adjusting the actuator (10) that controls the production machine (11). The production machine (11) can be, for example, a machine for stamping, sawing, drilling and / or cutting. Furthermore, it is conceivable that the production machine (11) is configured to grip production products (12a, 12b) by means of grippers.

[0179] The sensor (30) can then be, for example, a video sensor that detects, for example, the conveying surface of the conveyor belt (13) on which the manufactured products (12a, 12b) can be located.

[0180] The state (x) can in particular include the position of the manufactured products (12a, 12b) and the position of the processing tool of the production machine (11), such as the position of a gripper. The actuator (10) of the production machine (11) can then be controlled based on the determined position of the manufactured products (12a, 12b). For example, the actuator (10) can be controlled such that the actuator punches, saws, drills, and / or cuts the manufactured products (12a, 12b) at a predetermined position of the manufactured products (12a, 12b).

[0181] Furthermore, it is conceivable that the machine learning system (60) is configured to determine other characteristics of the manufactured products (12a, 12b) instead of or in addition to the position determination. In particular, it is imaginable that the machine learning system (60) determines whether the manufactured products (12a, 12b) are defective and / or damaged. In this case, the actuator (10) can be controlled such that the production machine (11) picks out the defective and / or damaged manufactured products (12a, 12b).

[0182] The term "computer" includes any device for performing pre-given computing procedures. These computing procedures can exist in the form of software, or can exist in the form of hardware or also in a hybrid form consisting of software and hardware.

[0183] Generally, a plurality can be understood as indexed, i.e., preferably by assigning successive integers to the elements included in the plurality, a unique index is assigned to each element in the plurality. Preferably, when the plurality includes N elements, where N is the number of elements in the plurality, integers from 1 to N are assigned to the elements.

Claims

1. A computer-implemented method (300) for training a machine learning system (60) for use in a model-predictive control of a technical system (100, 200), wherein the machine learning system (60) is designed to determine a first quality value for an operating state and / or an environmental state of the technical system (100, 200) to be controlled, the first quality value characterizing the quality of the state, wherein the method comprises the following steps: Determining (301) a plurality of operating states and / or environmental states of the technical system to be regulated; Determine (302) a state (x) from the plurality of operating states and / or environmental states i )'s first quality value The first mass value Represents the state (x) with respect to the model prediction regulation i )’s quality; Training (303) the machine learning system (60) by supervised training of the machine learning system (60), wherein the state (x i ) is used as input to the machine learning system (60) and the first quality value determined is used as the desired output of the machine learning system (60).

2. The method (300) according to claim 1, wherein the first quality value is determined based on a cost function adjusted by the model prediction The value function is related to the state (x i )Determine the state (x i ) and the cost function further comprises auxiliary conditions, wherein the auxiliary conditions characterize physical constraints of the state and / or the auxiliary conditions characterize constraints of a control signal, and the model predictive regulation can select the control signal.

3. The method (300) according to claim 2, wherein the value function characterizes the result of optimizing the cost function under the auxiliary conditions.

4. The method (300) according to claim 2 or 3, wherein the cost function comprises the admission of states (x i )Auxiliary conditions for numerical deviations from physical constraints.

5. The method (300) according to claim 4, wherein the auxiliary conditions further provide factors which enforce the physical constraints of the state.

6. The method according to any one of claims 1 to 5, wherein the cost function is described by the following formula: u.d.N x 0|k =x(k), 0≤α≤1, 0≤ξ i|k , x i+1|k =Ax i|k +Bu i|k , H u u i|k ≤h u , H i x i|k ≤h x (1-h)+ξ i|k +ξ N|k , The function J is described by the following formula:

7. The method according to any one of claims 2 to 6, wherein the first quality value (V i ) corresponds to the state (x i ) is the result of the value function at the position.

8. The method according to any one of claims 4 or 6, wherein the first quality value corresponding to a first portion of the value function at the location of the state (xi), wherein the first portion does not include a slack variable, and furthermore a second quality value is determined by the machine learning system The second mass value (V i slack ) corresponds to the state (x i ), the second part of the value function at the position of ), the second part includes a slack variable, and in the training step, the second quality value Serves as another desired output of the machine learning system (60).

9. The method according to claim 7 when dependent on claim 6, wherein the first quality value is described by the following formula: And the second quality value is described by the following formula 10. The method according to any one of claims 8 to 9, wherein the machine learning system (60) comprises two machine learning models, in particular two neural networks, wherein a first model of the two models is set up to generate a prediction based on the state (x i ) as the input of the first model to predict the first quality value And the second model of the two models is set up to be based on the state (x i ) as the input of the second model to predict the second quality value 11. The method according to any one of claims 6 to 8, wherein the machine learning system (60) comprises a machine learning model, in particular a neural network, wherein the model is set up to predict the first quality value and the second quality value 12. The method according to claim 10 or 11, wherein the machine learning system is trained based on a first loss function and a second loss function, wherein the first loss function characterizes the prediction of the first quality value and the first quality value The second loss function characterizes the prediction of the second quality value and the second quality value difference.

13. The method according to claim 12, wherein the second loss function is described by the following formula:

14. A method (300) according to any one of claims 2 to 13, wherein in the step of training (303), the machine learning system (60) is trained at least until an approximation quality criterion is met, wherein the approximation quality criterion characterizes the difference between a value of a first value and / or a second value determined by the value function for a state and at least one value determined by the machine learning system for the state.

15. A computer-implemented method for performing model predictive regulation of a technical system (100, 200), comprising the following steps: Determining the operating state and / or environmental state of the system (100, 200); Determining a control signal (A) for the system (100, 200), wherein the control signal (A) is determined based on a regulation law (π), wherein the regulation law (π) includes the term in characterizing an output of a machine learning system (60) when it receives an input f(x, u), wherein f is a model, preferably a linear model, of the system (100, 200), and predicting a next operating state and / or system state of the system (100, 200) based on the determined operating state and / or environmental state x and the control signal u, wherein the machine learning system (60) has also been trained according to a method according to any one of claims 1 to 7; The system (100, 200) is controlled by means of the determined control signal.

16. A computer-implemented method for performing model predictive regulation of a technical system (100, 200), comprising the following steps: Determining the operating state and / or environmental state of the system (100, 200); Determining a control signal (A) for the system (100, 200), wherein the control signal (A) is determined based on a regulation law (π), wherein the regulation law (π) includes the term in Determine according to the following formula in characterizes a first quality value output by the machine learning system (60) for an input f(x,u) and Characterizing a second quality value output by a machine learning system (60) for the same input, wherein f is a model of the system (100, 200), preferably a linear model, and predicting the next operating state and / or system state of the system (100, 200) based on the determined operating state and / or environmental state x and the control signal u, wherein the machine learning system (60) has also been trained according to the method according to any one of claims 8 to 14. 17 . The method as claimed in claim 15 , wherein the control law further comprises a term which characterizes the amplitude of the control signal.

18. A training device (140) configured to carry out the method according to any one of claims 1 to 13.

19. A regulating system (40) configured to carry out the method according to claim 15.

20. A computer program, which is designed to carry out the method according to any one of claims 1 to 17 when the computer program is executed by a processor (45, 145).

21. A machine-readable storage medium (46, 146) having stored thereon a computer program according to claim 20.