Device and method for configuring a planner of a technical system
The contract-based planner and motion control strategy addresses model inconsistency in hierarchical control architectures by synchronizing planners and controllers using a feasibility function, enhancing trajectory feasibility and system performance.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2025-11-11
- Publication Date
- 2026-05-21
AI Technical Summary
Hierarchical control architectures in technical systems, particularly in robots, often face model inconsistency between different layers, leading to significant deviations between planned and actual behavior, resulting in suboptimal performance and safety issues.
A contract-based planner and motion control strategy is implemented, where a feasibility function is provided to the planner to determine feasible trajectories, allowing for combining different planner and controller models, and using a soft-constrained model predictive control optimization to synchronize the planner and controller without human intervention.
This approach ensures that the planner and controller automatically adapt to each other, reducing computational complexity and enhancing the feasibility of trajectories, thereby improving system performance and safety.
Smart Images

Figure EP2025082664_21052026_PF_FP_ABST
Abstract
Description
[0001] R. 414010
[0002] - 1 -
[0003] Description
[0004] Title
[0005] Device and method for
[0006]
[0007] a planner of a technical
[0008]
[0009] Technical field
[0010] The invention concerns a method for configuring a planner, a method for determining a control signal, a method for controlling a technical system, a technical system, a training system, a control system, a computer program, and a machine-readable storage device.
[0011] Prior art
[0012] Di Cairano et al. “Vehicle tracking control on piecewise-clothoidal trajectories by MPC with guaranteed error bounds”, 2016, IEEE 55th Conference on Decision and Control (CDC), pages 709-714 discloses a vehicle steering controller with limited preview ensuring that the vehicle constraints are satisfied, and that any piecewise clothoidal trajectory, that is possibly generated by a path planner or supervisory algorithm and satisfies constraints on the desired yaw rate and the change of desired yaw rate, is tracked within a preassigned lateral error bound.
[0013] Gupta et al. “A framework for vehicle lateral motion control with guaranteed tracking and performance”, 2019, IEEE Intelligent Transportation Systems Conference (ITSC), pages 3607-3612 discloses a framework to design a controller for the vehicle lateral dynamics, which guarantees to meet desired safety and performance requirement.
[0014] Technical background R. 414010
[0015] - 2 -
[0016] When controlling technical systems, a commonly used paradigm is that of sense-plan-act. This may be understood as first sensing certain characteristics of the system to be controlled, then planning how control should be conducted and then actually controlling the device physically in an act step. Especially in robots and in particular mobile robots, this paradigm may be used to determine trajectories of motions of the robot and then controlling the robot to actually follow the desired trajectory.
[0017] In a typical software architecture for such robotic applications, the path planning and execution may be structured hierarchically, with a higher-level path planning or trajectory generation (path planner or simply planner in the following), and a lower-level motion control unit that executes the desired plan using the robot’s actuators. The path planning layer may generate a trajectory for the robot’s motion control module to follow. It may especially consider the robot’s current position, desired destination, and a model of the environment the robot operates in. Typically, a planner only considers simple models like point mass or kinematic models since longer prediction horizons are desired, but available computing resources may not allow for more complex models. The motion control layer may take the generated path and convert it into low-level control commands for the robot’s actuators, such as steering, throttle, and brakes. The motion control layer may be understood as ensuring that the robot accurately follows the planned path while maintaining comfort and safety. The control layer typically uses more detailed dynamical models of the vehicle, such as the dynamic bicycle model, to generate the control commands.
[0018] While such hierarchical control architectures may be used effectively for controlling technical systems, especially robots, it also poses the problem of model inconsistency: Typically, different models are used at different layers of the control architecture. Trajectories planned using the simple models in the planning layer, such as the point mass or kinematic models, may not be feasible for dynamics models used in the control layer or even not be executable altogether by the technical system. This can lead to a significant deviation between planned and actual behavior which typically result in suboptimal performance and safety issues of the overall technical system. R. 414010
[0019] - 3 -
[0020] Advantageously, the method with features of claim 1 proposes to solve this problem using a contract-based planner and motion control strategy. By providing information from the controller to the planner about feasibility of trajectories determined by the planner, the combined system of planner and controller is able to efficiently determine feasible trajectories of a technical system, irrespective of the models used in the planner or the controller. This allows for combining different planner models and controller models, wherein planner and controller then perform an automated “handshake” to account for differences in their models.
[0021] Disclosure of the invention
[0022] In a first aspect, the invention concerns a computer-implemented method for configuring a planner, wherein the planner is configured for providing trajectory signals for controlling a technical system, the method comprising the steps of: • Providing a feasibility function to the planner, wherein the feasibility function is configured to determine a feasibility value for a trajectory signal, wherein the feasibility value indicates a minimal cost of a soft-constrained model predictive control optimization used by a controller of the technical system for controlling the technical system, wherein the controller is configured for controlling the technical system based on trajectory signals provided by the planner;
[0023] • Including the feasibility function as a component of an optimization problem of the planner thereby configuring the planner, wherein the planner provides a trajectory signal based on the optimization problem.
[0024] The method may, in general, be understood as allowing for adapting an existing planner with respect to a controller, i.e., providing a connection between controller and planner. In terms of software development, this may be understood as a contract-based planner-controller combination, wherein the feasibility function provided to the planner serves as contract between the planner and the controller.
[0025] The term "planner" may be understood as a software component, algorithm, or module that generates a plan or sequence of actions for a technical system to R. 414010
[0026] - 4 -
[0027] achieve a desired goal. In particular, the planner may provide a trajectory signal that is indicative of a trajectory to be executed by the technical system. The planner may operate on a digital computer or a network of digital computers, receiving input data about the system's current state, its environment, and / or the desired objective. The planner makes use of optimization-based planning, e.g., optimizing a cost function, for determining the trajectory signal.
[0028] Including the feasibility function as a component of the optimization problem of the planner may be understood as including the feasibility function as a potentially weighted term in a cost function of the optimization problem. The feasibility function may, however, also be included elsewhere in the optimization to influence the generation of trajectories. For example, the feasibility function could be included as an additional constraint of the optimization problem with an optional weighting factor.
[0029] In some embodiments, the planner may be integrated with the controller or other components of the system. In other embodiments, the planner may be a separate module or process that communicates with the controller and other components of the system. The planner may be generally configured to use a simplified model of the technical system for improved computational efficiency, e.g., as proposed by Di Cairano et al. or Gupta et al.
[0030] The planner may further be part of an autonomous control system for an at least partially automated vehicle. The term "planner" may also be referred to as a "trajectory generator" or "path planner".
[0031] The initial planner, i.e., the planner used at the outset of the method before including the feasibility function, may especially provide its trajectory signals based on a discrete-time model. In particular, a discrete-time model according to the formulae:
[0032] xp(k + 1) = fp(xp(fc), Up(fc)},
[0033] yp( / c) = gp^xP(Ji),up(Ji)^,
[0034]
[0035] where xp( / c) is a state of the system at discrete time k, up(Ji) is a control input, and yp( / c) is an output of the system. An example for a model used in a planner R. 414010
[0036] for a mobile robot can be expressed according to the formulae (see also Di Cairano et al. and Gupta et al.):
[0037] -J- " I” COS
[0038] sy,p(Jt + sy,p( / c) + rsvx p( / c) sin (ipp( / c)) ^x,p (fc =
[0039] ^x,p " I” Ts&x '
[0040] + 1)
[0041] ipp(k) + Tsipp(Ji)
[0042] %(k) + Tsi[>p(Jc)
[0043] yp(Ji) = ipp(k),
[0044] with input up(Ji) = [i / ip( / c), ax( / c)]Tand state xp( / c) =
[0045]
[0046] [sx,p(fc\ sy p(k), vX; P(k),ipp(k),tpp(k)]Twhere sxp( / c),syp( / c) are desired x- and y-coordinates respectively of the robot in a global coordinate frame at time t,p(k) is a desired rotation of the robot at time t,p(k) is a desired yaw rate at time t, [>p(k) is a desired yaw acceleration at time t, vxp(k) is a velocity of the robot (preferably in a coordinate frame of the robot) at time t, ax(Ji) is a desired acceleration of the robot at time t and Tsis a sample time.
[0047] The trajectory signal may, in general, comprise or consist of a sequence of states xpand / or a sequence of outputs ypcorresponding to states xp. The trajectory signal may also include an initial state x0. For example, the trajectory signal may comprise or consist of the output ypas a trajectory signal provided by the planner. The controller may then determine its control signals based on the provided trajectory signal, in particular based on the sequence of states xpand optionally the initial state x0. Alternatively, the controller may also take into account information regarding outputs ypor base its outputs solely on outputs yp. For example, as shown above, the controller may receive only yp=p(k), i.e., a desired yaw rate for each timestep. The output ypmay, however, be at least parts of the state xp. In other words, the function gpmay be an identity function in its first argument or provide a subset of components of its first argument as output.
[0048] Other examples for models applicable in the planner are a point mass model or a kinematic bicycle model.
[0049] The planner determines the trajectory signal based on the cost function. This may especially be understood as the planner optimizing the cost function, wherein a result of the optimization comprises or consists of the trajectory signal. R. 414010
[0050] - 6 -
[0051] The planner may especially solve a finite-horizon optimal control problem according to the formulae:
[0052] min Jp(%p,-\k>
[0053] up, |fc,xp, |fc,yp, |fc
[0054] S
[0055]
[0056] . t.xp,0\k Xp(Jt)
[0057] forZ = 0,..., N — 1:
[0058] xp,l+l\k fpxp,l\k> p,l\k)
[0059] yp,l\k 9p. p,l\k> p,l\k)
[0060] c(yP,i\k o
[0061] where ]pis a cost function that may consider constraints for the planned trajectory such as driving comfort measures, energy consumption, etc. The timediscrete model from above may especially be used to predict the system state sequence xp.]k= {xp,i|fe}J=0and the output sequence yp.]k= {yp,i|fe}J=0for a given input sequence u
[0062]
[0063] Pi.\k= |up,(|k}(-0- The model may be initialized with a state xp( / c), which may be measured or reconstructed using a state estimator or an observer. The function c(yp,(|k) is a constraint function, i.e., a function that assesses whether certain constraints of a generated output yp>i\ksatisfies certain requirements posed to the technical system. For example, in the context of mobile robots, the constraint function may assign a positive value to an output yP,i\k if the output exceeds a certain allowed rotation speed. In further embodiments the constraint function c may use the state of input or yp>i\kmay be at least a subset of the state xp>i\k. In these cases, the constraint may also comprise information about a position, speed, acceleration, and / or jerk of the system to be controlled. The function c may then provide positive values if, for example, a position of a state is outside of a predefined driving area (e.g., a road), a speed exceeds a desired speed limit or the state characterizes an undesired jerk in the driving motion.
[0064] The discrete-time model from above may be used in the planner to predict a system state sequence xp.\k= and an output sequence yp.\k=
[0065]
[0066] {y^lk}^1for an input sequence up / |k=
[0067]
[0068] The model is initialized with the state xp( / c), which is measured or reconstructed using an estimator or observer. The constraints may consider for instance road boundaries, obstacles, etc., Z e N is a prediction horizon. The optimization variables are the predicted input sequence up.\k, the predicted state sequence xp.\k, and the predicted output sequence yp.\k. Solving the optimization problem, one may obtain optimal R. 414010
[0069] - 7 -
[0070] sequences x*^k, y^, and u* for the predicted optimal states, outputs, and inputs respectively. All or some of these optimal sequences may be provided in the trajectory signal. In particular, the optimal output y*.|kmay be provided in or as the trajectory signal to the controller of the technical system.
[0071] In general, the term trajectory signal may be understood as a sequence of desired future states and / or outputs of the technical system. The trajectory signal may serve as input to the controller, i.e., serve as interface between the planner and the controller.
[0072] The controller may especially be configured for determining control signals for controlling at least one actuator of the technical system. The controller controls the technical system, in particular actions of the technical system, by means of a soft-constrained model predictive control optimization. The controller may especially use a model of the state of the technical system according to the formula:
[0073] x(k + 1) = f ^x( / c),u( / c), w( / c),yp( / c)^,
[0074]
[0075] where x( / c) is a state of the technical system, u( / c) is a control input, w( / c) is an external disturbance acting on the technical system or a model uncertainty, and yp( / c) is the desired trajectory signal. Concrete implementations of this model include a dynamic bicycle model. In case the technical system is a mobile robot, e.g., an least partially automated vehicle, a formulation of the dynamic bicycle model for the mobile robot may be given by system states x =
[0076] [
[0077]
[0078] e(at, ev, vy,r]Tand inputs u = [5, T]T, where e(atis a lateral error, i.e., a shortest distance from the robot’s center of gravity to the path,
[0079]
[0080] is an error in the heading angle, vxand vyare velocities in the x- and y-direction, r is a yaw rate, 8 is a steering angle and T is a drivetrain and braking command. For track relative coordinates, the model may require path information yp( / c), e.g., a curvature of the path or a desired path orientation. The dynamic models may consider factors like tire-road friction, vehicle dynamics, and actuator limitations.
[0081] The model used by the controller may in particular be different from the model used by the planner. R. 414010
[0082] - 8 -
[0083] From the viewpoint of the controller, the states and inputs of the technical system may be subject to constraints, i.e., x( / c) e X:= {x|cx(x) < 0} and u( / c) e U ■= {Ulcu(u) °} for all e N where cx(x) and cy(u) may be linear or nonlinear functions that define the constraints.
[0084] The input constraints U may be defined by the technical systems physical control limits. For a mobile robot such as an at least partially automated vehicle, the physical limits may comprise a maximum steering angle, a maximum acceleration, or other physical limitations. The state constraints X may be given by safety and / or comfortability specifications. For a mobile robot such as an at least partially automated vehicle, the state constraints may comprise a maximum deviation from a desired path, a maximal velocity, a maximal acceleration, a maximal jerk, or other safety and / or comfortability related features.
[0085] The soft-constrained model predictive control optimization used by the controller may be defined according to the formulae:
[0086] min JMPC(.x.\k, u.\k~) + wf / f(^.|k)
[0087] X\k, U\kX\k
[0088] s. t. x0|k= x( / c)
[0089] ξN|k ≥ 0
[0090] cf(.xN\k) %N\k
[0091] for I e {0, — 1}:
[0092] %i+l|fc = f(xl\k’Ul\k’0,ypi\k]
[0093] cX,l(xllk) — cU,l(ul|k) ≤ 0
[0094]
[0095] ξl,k ≥ 0
[0096] The cost function ]MPCmay especially comprise or consists of two terms, namely a first term characterizing the model predictive control cost and a second term characterizing a feasibility cost term. Alternatively or additionally, x*^kmay be used in f.
[0097] The const function for a common soft-constrained model predictive control (MPC) may be defined according to the formula:
[0098] N — l
[0099] JMPc[x-\k>u-\k)=l(xl\k>ul\k) + Vf(.xN\k)
[0100]
[0101] 1 = 0 R. 414010
[0102] - 9 -
[0103] where I is a stage cost and Vj is a terminal cost The feasibility cost function h&\k) = I jvifc || + Sf o1l i|fc II penalizes the slack variables
[0104]
[0105] > 0. The slack variables allow for softening of the terminal and state constraints in the optimization problem. By selecting a weight factor
[0106]
[0107] large enough, the optimal input sequence of the soft-constrained MPC problem is identical to the optimal input sequence of the hard-constrained MPC problem for initial states where the hard-constrained MPC is feasible and
[0108]
[0109] = 0 for all
[0110]
[0111] I e {0, Moreover, the soft-constrained MPC problem is also feasible for states where the hard-constrained MPC problem is infeasible through the softening of the state constraints. Since the disturbance w( / c) may be unknown to the controller, a nominal prediction model may be used and the constraints may be tightened using common robustifaction techniques to compensate the deviation between actual and nominal model. Given a state x( / c) and a trajectory signal, e.g., yp( / c)* and / or xp( / c)*, the soft-constrained MPC problem may be solved online to obtain an optimal control input sequence u^kand corresponding slack variables ^
[0112]
[0113] k.
[0114] As part of a method, a feasibility function is provided to the planner based on information from the controller. The planner hence adapts its optimization in order to account for the feasibility function. The adapted optimization problem of the planner may be defined according to the formulae:
[0115] min / p(xp / |fc,up / |k) + w^x^k^yp.^ up,·|k, xp,·|k, yp,·|k
[0116] s. t. xp,0|k = xp(k)
[0117] for I = 0,..., N — 1:
[0118] xp,l+1|k = fp(xp,l|k, up,l|k) yp,l|k = gp(xp,l|k, up,l|k)
[0119]
[0120] c(yp,l|k) ≤ 0
[0121] Where whis an optional and predefined weight for weighting a cost term h, wherein h is a measure quantifying feasibility of the MPC optimization problem used by the controller.
[0122] In some embodiments, the cost term h may be selected as a solution to the minimal feasibility problem according to the formulae:
[0123] h(x(k), yp,·|k) = min Jξ(ξ·|k)
[0124] s. t. x0|fc = x(Ji)
[0125]
[0126] ξN|k ≥ 0 R. 414010
[0127] - 10 -
[0128] cf {.xN\k) %N\k
[0129] for / e {0,..., N - 1}:
[0130] %i + l|fc = f(xl\k>ul\k> ^> yp,l\k)
[0131] cX,l(xl|k) ≤ ξl|k
[0132] cU,l(ul|k) ≤ 0
[0133]
[0134] ξl,k ≥ 0
[0135] This may be understood as the feasibility problem checking whether for a given output yp / |kand state x( / c) there exists an input sequence for the controller which complies with the constraints. As above, the trajectory signal may comprise yp.\kand xp.\k. In case constraints are violated, the cost term h is positive. If the planner cannot find a feasible trajectory, it can minimize the amount of constraint violation by choosing the trajectory signal with low cost h. The weighting factor whcan be used to trade-off constraint violation and cost minimization. If h > 0, this can be understood as a violation of the original constraints in the soft-constrained MPC problem. Advantageously, the soft-constraint MPC can bring back the system to a feasible state in a later point in time.
[0136] The objective functions J may all be defined according to known MPC objective functions.
[0137] While the formulae above can be understood to characterize an exact measure of feasibility, the solution requires solving a nested optimization problem. Hence, in preferred embodiments the feasibility function is configured for approximating a result of the soft-constrained model predictive control optimization for a given trajectory signal.
[0138] Advantageously, this leads to reduced computational complexity when determining the trajectory signal.
[0139] In preferred embodiments, the approximation may be achieved by means of a machine learning system, in particular a neural network. The neural network may be trained offline, i.e., before using the planner for determining a trajectory signal for the technical system, using initial states and a dataset of feasible and infeasible trajectory signals. The trajectory signals may especially be generated by the planner before adding the feasibility function, i.e., with a not-yet adapted R. 414010
[0140] - 11 -
[0141] version of the planner. When the planner is executed online, i.e., for determining the trajectory signal of the technical system,
[0142] the neural network can be used to evaluate the feasibility of the generated trajectory signal. The inventors found that, advantageously, the neural network is capable of approximating the true feasibility value with high accuracy thus allowing for a suitable trade-off between correctness of the output and computational complexity. The neural network may be a multi-layer perceptron. The multi-layer perceptron advantageously allows for fast training times, high accuracy of predicting the feasibility value, as well as small computational footprint when used online.
[0143] In preferred embodiments, providing the feasibility function to the planner may comprise the steps of:
[0144] • Providing at least one trajectory signal as training input
[0145] • Providing for the at least one trajectory signal a feasibility value as desired training output, wherein the feasibility value is determined by providing the trajectory signal to the soft-constrained model predictive optimization;
[0146] • Performing supervised training of a machine learning system using the training input and desired training output as training data;
[0147] • Providing the machine learning system as the feasibility function.
[0148] This may be understood as the planner providing trajectory signals to form a training dataset for the machine learning system. The machine learning system may then be trained in a supervised fashion, i.e., using the trajectory signals as input and the determined feasibility value as desired output of the machine learning system. The trained machine learning system may then be provided back to the planner.
[0149] In preferred embodiments, the at least one trajectory signal is used for training the machine learning system is provided by the planner.
[0150] This may be understood as the planner working preliminarily without the additional cost term represented by the feasibility function. The planner could provide trajectories „stand alone", i.e., without taking into account specifics of the R. 414010
[0151] - 12 -
[0152] controller. In other words, the controller and planner may be understood as operating in a „set up phase", where the planner provides trajectory signals the planner „thinks“ are feasible wherein the controller provides feedback in terms of the feasibility function in terms of what the controller is capable of achieving given the soft-constraints used in the MPC of the controller.
[0153] Advantageously, the entire process allows for synchronizing the planner and the controller automatically without the need for human intervention. This has the further advantage of configuring different planners and controllers for one another without the need for human intervention. For example, a first controller may use a different model in its MPC for determining its control signal with respect to a second controller. Depending on which controller shall be used for controlling the technical system, the respective controller may automatically perform a “handshake” with the planner using the method above.
[0154] In preferred embodiments, the steps of providing the feasibility function are conducted by the controller.
[0155] This may be understood as modularizing the planner and controller and having the controller “introduce” itself to the planner.
[0156] In another aspect, the invention concerns a computer-implemented method for determining a control signal of a technical system comprising the steps of:
[0157] • Configuring a planner of the technical system according to any one of the preceding claims;
[0158] • Determining a trajectory signal using the adapted cost function of the planner;
[0159] • Providing the trajectory signal to a controller of the technical system; • Determining, by the controller, an output signal characterizing an optimal control input to of the actuator, wherein the controller determines the output signal based on the trajectory signal using a soft-constrained model predictive control optimization;
[0160] • Determining the control signal based on the output signal. R. 414010
[0161] - 13 -
[0162] This may be understood as a method for combining the “offline” with the “online” phase of the planner and / or the controller.
[0163] In the embodiments provided above, examples typically revolved around the technical system being a robot, in particular a mobile robot, in particular an at least partially automated vehicle. However, the method may be used for general control tasks, where some form of “trajectory” needs to be followed by the technical system. The trajectory does not necessarily need to be positions in 2D or 3D space. For example, an opening of a valve over time may also be understood as a trajectory signal. In general, any physical position or state over time may be comprises by a trajectory signal or the trajectory signal may consist thereof. The output ypfor each state may also be used in the trajectory signal, e.g., as a proxy of the states xp.
[0164] In other aspect, the invention concerns a technical system comprising:
[0165] • A planner configured for providing trajectory signals for controlling a technical system;
[0166] • A controller configured for determining control signals for controlling at least one actuator of the technical system, preferably a subset or all actuators of the technical system or all actuators of the technical system, wherein the controller determines the control signal using a soft-constrained model predictive control optimization,
[0167] wherein the planner uses a cost function for determining trajectory signals, wherein the cost function comprises a potentially weighted cost term, wherein the cost term is a function characterizing a minimal cost of the soft- constrained model predictive control optimization for a given trajectory signal.
[0168] The planner and the controller may especially be considered as being part of the technical system. For example, the technical system may comprise or consist of a computer that is configured to use the planner and / or the controller as described above.
[0169] Alternatively, the technical system comprising the planner and / or the controller may also be understood as the technical system having a network connection to R. 414010
[0170] - 14 -
[0171] a remote computer, wherein the remote computer runs the planner and / or controller. The interaction between planner and controller may then be conducted through the network connection. Alternatively, the planner and controller may both be run by the remote computer, wherein the technical system is provided a control signal obtained from the controller.
[0172] The system inherits all the advantages of the combination of planner and controller as described above, in particular an automatic configuration or “handshake” between the planner and the controller. This allows for providing planners and / or controllers using different models for their respective optimization procedures to automatically “introduce” themselves to one another and automatically configuring themselves for cooperating in order to control the technical system.
[0173] Embodiments of the invention will be discussed with reference to the following figures in more detail. The figures show:
[0174] Figure 1 a contract-based planner and controller component;
[0175] Figure 2 a training system for training a feasibility function of the contractbased planner controller component;
[0176] Figure 3 a control system comprising the contract-based planner controller component for controlling an actuator in its environment;
[0177] Figure 4 the control system controlling an at least partially autonomous robot;
[0178] Figure 5 the control system controlling a manufacturing machine.
[0179] Description of the embodiments
[0180] Figure 1 shows an embodiment of a contract-based planner and controller component (63), wherein the contract-based planner and controller combination is configured for providing a control output (c) for controlling a technical system. The contract-based planner and controller component (63) comprises a planner R. 414010
[0181] - 15 -
[0182] (61) and a controller (62). The planner (61) is configured for determining a trajectory signal (t) characterizing a desired trajectory of the technical system (100). The trajectory signal may, for example, comprise desired positions in 2D or 3D space of the technical system. The planner (61) is configured to determine the trajectory signal (t) based on a cost function. The planner (61) may especially solve a finite-horizon optimal control problem according to the formulae:
[0183] min Jp(%p,-\k>
[0184] up,·|k,xp,·|k,yp,·|k
[0185] s. t. xp,0|k = xp(k)
[0186] for I = 0,..., N — 1:
[0187] xp,l+1|k = fp(xp,l|k, up,l|k)
[0188] yp,l|k = gp(xp,l|k, up,l|k)
[0189]
[0190] c(yp,l|k) ≤ 0
[0191] The definition of the symbols in the formulae is the same as described above. Solving the optimization problem allows for generating a trajectory signal (t) comprising x*p,·|k and / or y*p,·|k. In other words, the trajectory signal (t) may comprise, for example, a sequence of desired future states xp, such as planned positions, velocities, and headings of the vehicle over a given time horizon, and / or a sequence of desired outputs ypcorresponding to the states xp. The trajectory signal (t) may be understood as an interface between the planner (61) and the controller (62).
[0192] The controller (62) is configured to determine an output signal (c) based on the trajectory signal (t), wherein the output signals (c) allows for controlling an actuator of the technical system such that the desired trajectory represented by the trajectory signal (t) is achieved. The output signal (c) may comprise or consist of an optimal input u*kas determined by the controller (62). The controller (62) determines the output signal (c) based on an optimization soft-constrained model predictive control problem. The soft-constrained model predictive control optimization used by the controller may be defined according to the formulae:
[0193] min JMPC(x·|k, u·|k, ξ·|k) + wξJξ(ξ·|k)
[0194] x·|k, u·|k, ξ·|k
[0195] s. t. x0|k = x(k)
[0196] ξN|k ≥ 0
[0197] cf(xN|k) ≤ ξN|k
[0198] for l ∈ {0,..., N - 1}:
[0199]
[0200] %i+l|fc = f{.xl\k’ul\k>^>yp,l\k) R. 414010
[0201] - 16 -
[0202] cX,l(xl|k) ≤ ξl|k
[0203] cU,l(ul|k) ≤ 0
[0204]
[0205] ξl,k ≥ 0
[0206] The definition of the symbols in the formulae is the same as described above.
[0207] Before controlling the technical system, the planner (61) is configured by providing a feasibility function (fi.) to the planner (61). The feasibility function may be defined according to the formulae:
[0208] h(x(k), yp,·|k) = min Jξ(ξ·|k)
[0209] s. t. x0|k = x(k)
[0210] ξN|k ≥ 0
[0211] cf(xN|k) ≤ ξN|k for l ∈ {0,..., N - 1}:
[0212] xl+1|k = f(xl|k, ul|k, 0, yp,l|k)
[0213] cX,l(xl|k) ≤ ξl|k
[0214] cU,l(ul|k) ≤ 0
[0215]
[0216] ξl,k ≥ 0
[0217] The definition of the symbols in the formulae is the same as described above. As such a definition may lead to a nested optimization problem, the feasibility function (fi.) may preferably be given by a function approximating the exact feasibility, in particular a machine learning system, specifically a neural network. The machine learning system may be trained with initial state and trajectory signals (t) originating from the initial states and generated by the planner as input and feasibility scores as desired output, wherein the feasibility scores are obtained using the initial states and trajectory signals (t) in the optimization above. This specific embodiment is visualized by dashed lines in the figure: The planner (61) may provide trajectory signals (t) comprising optimal states x*p,·|k and / or optimal outputs y*p,·|k based on initial states x( / c) to the controller (62) in a “handshake” phase, i.e., before controlling the technical system (100, 200), wherein the controller (62) may then determine feasibility values for the respective trajectories characterized by the trajectory signal (t by optimizing the cost function from above given a respective trajectory signal (t. The controller (62) may then train a machine learning system, in particular a neural network, with this data and provide the machine learning system as feasibility function (fi.) back to the planner (61). The planner (61) may then integrate the feasibility R. 414010
[0218] - 17 -
[0219] function (fi.) into its cost function thereby incorporating information about the controller’s abilities into its process of generating trajectories.
[0220] When using a neural network as feasibility function (fi.) the parametrized layers of the neural network may all be fully connected, i.e., all parametrized layers may be fully connected layers. In particular, the neural network may be a multi-layer perceptron. The architecture of the neural network may especially be defined such that each fully connected layer comprises 64 neurons. As activation functions, common activation functions such as ReLU, ELU, GELU or the like may be used.
[0221] Figure 2 shows an embodiment of a training system (140) for training the machine learning system (60) as a feasibility function (h) to be used in the contract-based planner and controller combination (63). The machine learning system (60) is trained by means of a training data set (T). The training data set (T) comprises a plurality of training inputs (x which are used for training the machine learning system (60), wherein the training data set (T) further comprises, for each training input (x, a desired training output (o which corresponds to the training input (x. In the embodiment, each training input (x comprises an initial state x( / c) and a trajectory signal (t) characterizing a trajectory starting from the initial state. The trajectory signal (t) may especially be an optimal sequence of states x*kand / or outputs y*kas generated by the planner (61) starting from the initial state x( / c). In other words, the planner (61) may be used to generate trajectory signals (t for training. For this, the planner (61) may be provided with initial states x( / c) samples preferably at random. The planner may then determine the trajectory signals (t for training, which are then stored in the training dataset (T). The desired training outputs (o may especially represent feasibility values as obtained by the controller. In particular, some or all of the trajectory signals (t generated by the planner (61) may be provided to the controller (62). The controller may then determine a feasibility value for each of the provided trajectory signals (t, which may then be stored as desired training outputs (o in the training data set (T). In the embodiment, the planner (61) may especially generate at least one million trajectory signals (t to be included in the training data set (T), in particular at least 6.9 million trajectory signals (t. Each of R. 414010
[0222] - 18 -
[0223] these trajectory signals (t may be assigned a feasibility value as described above.
[0224] For training the machine learning system (60), a training data unit (150) accesses a computer-implemented database (St2), the database (St2) providing the training data set (T). The training data unit (150) determines from the training data set (T) preferably randomly at least one training input (x and the desired output (t corresponding to the training input (x and transmits the training input (x to the machine learning system (60). The machine learning system (60) determines a training output (y based on the training input (x.
[0225] The desired output (o and the determined training output (y are transmitted to a modification unit (180).
[0226] Based on the desired output (o and the determined training output (y, the modification unit (180) then determines new parameters (O') for the machine learning system (60). For this purpose, the modification unit (180) compares the desired output signal (o and the determined output signal (y using a loss function. The loss function determines a first loss value that characterizes how far the determined output signal (y) deviates from the desired output signal (o). In the given embodiment, a negative log-likelihood function is used as the loss function, in particular a mean squared error loss function or an L2-loss function. Other loss functions are also conceivable in alternative embodiments.
[0227] The modification unit (180) determines the new parameters (O') based on the first loss value. In the given embodiment, this is done using a gradient descent method, preferably stochastic gradient descent, Adam, or AdamW. In further embodiments, training may also be based on an evolutionary algorithm or a second-order method for training neural networks.
[0228] In other preferred embodiments, the described training is repeated iteratively for a predefined number of iteration steps or repeated iteratively until the first loss value falls below a predefined threshold value. Alternatively or additionally, it is also conceivable that the training is terminated when an average first loss value with respect to a test or validation data set comprising trajectory signals and R. 414010
[0229] - 19 -
[0230] feasibility values falls below a predefined threshold value. In at least one of the iterations the new parameters (O') determined in a previous iteration are used as parameters ( ) of the machine learning system (60).
[0231] Furthermore, the training system (140) may comprise at least one processor (145) and at least one machine-readable storage medium (146) containing instructions which, when executed by the processor (145), cause the training system (140) to execute a training method according to one of the aspects of the invention.
[0232] The training system (140) may especially be comprised by the controller (62), i.e., the controller (62) may comprise the training system (140) as a subcomponent of the controller (62).
[0233] After training, the machine learning system (60) may be provided as feasibility function (fi.) to the planner (61).
[0234] Figure 3 shows an embodiment of an actuator (10) of a technical system (100, 200) in its environment (20). The actuator (10) interacts with a control system (40) comprising the contract-based planner and controller component (63). The actuator (10) and its environment (20) will be jointly called actuator system. At preferably evenly spaced points in time, a sensor (30) senses a condition of the actuator system. The sensor (30) may comprise several sensors. Preferably, the sensor (30) is an optical sensor that takes images of the environment (20). An output signal (S) of the sensor (30) (or, in case the sensor (30) comprises a plurality of sensors, an output signal (S) for each of the sensors) which encodes the sensed condition is transmitted to the control system (40). The output signals (S) may hence be understood as indicative of the environment (20) of the technical system (100, 200)
[0235] Thereby, the control system (40) receives a stream of sensor signals (S). It then computes a series of control signals (A) depending on the stream of sensor signals (S), which are then transmitted to the actuator (10). R. 414010
[0236] - 20 -
[0237] The control system (40) receives the stream of sensor signals (S) of the sensor (30) in a receiving unit (50). The receiving unit (50) determines a current state (x( / c)) of the technical system (100, 200) based on the sensor signals (S). The receiving unit (50) transmits the current state (x( / c)) and a desired behavior (v) to the contract-based planner and control component (63). The desired behavior (v) may especially characterize a desired motion of the technical system (100, 200), e.g., desired velocities of the technical system (100, 200) and or a desired path or route of the technical system (100, 200).
[0238] The planner (61) receives the state (x( / c)) and the desired behavior (v) as input and determines a trajectory signal (t) based on this input. In the embodiment, the planner (61) uses an adapted optimization problem for determining the trajectory signal (t). The adapted optimization problem of the planner (61) may be defined according to the formulae:
[0239] min Jp(xp,·|k, up,·|k) + wh·ĥ(x(k), yp,·|k)
[0240] up,·|k, xp,·|k, yp,·|k
[0241] s. t. xp,0|k = xp(k)
[0242] for I = 0,..., N — 1:
[0243] xp,l+1|k = fp(xp,l|k, up,l|k) yp,l|k = gp(xp,l|k, up,l|k)
[0244]
[0245] c(yp,l|k) ≤ 0
[0246] where wh is an optional and predefined weight for weighting a cost term ĥ, wherein ĥ is the feasibility function (h) provided to the planner (61) as described above. The planner (61) determines a trajectory signal (t), which is provided as input to the controller (62) alongside the state (x( / c)). Based on these inputs, the controller (62) determines an output signal (c). The output signal (c) may comprise or consist of an optimal input u*
[0247]
[0248] as determined by the controller (62).
[0249] The output signal (c) is transmitted to an optional conversion unit (80), which converts the output signal (c) into the control signals (A). The control signals (A) are then transmitted to the actuator (10) for controlling the actuator (10) accordingly. Alternatively, the output signal (c) may directly be taken as control signal (A).
[0250] The actuator (10) receives control signals (A), is controlled accordingly and carries out an action corresponding to the control signal (A). The actuator (10) R. 414010
[0251] - 21 -
[0252] may comprise a control logic which transforms the control signal (A) into a further control signal, which is then used to control actuator (10).
[0253] In further embodiments, the control system (40) may comprise the sensor (30). In even further embodiments, the control system (40) alternatively or additionally may comprise the actuator (10).
[0254] In still further embodiments, it can be envisioned that the control system (40) controls a display (10a) instead of or in addition to the actuator (10).
[0255] Furthermore, the control system (40) may comprise at least one processor (45) and at least one machine-readable storage medium (46) on which instructions are stored which, if carried out, cause the control system (40) to carry out a method according to an aspect of the invention.
[0256] Figure 4 shows an embodiment in which the control system (40) is used to control robot (100), e.g., an at least partially autonomous vehicle (100).
[0257] The sensor (30) may comprise one or more video sensors and / or one or more radar sensors and / or one or more ultrasonic sensors and / or one or more LiDAR sensors. Some or all of these sensors are preferably but not necessarily integrated in the robot (100).
[0258] The constraints used in the optimizations of planner (62) and / or the controller (62) may comprise an information, which characterizes where objects are located in the vicinity robot (100) and what distance to keep from these objects Alternatively or additionally, the constraints may characterize a deviation from a drivable surface area such as a road. The control signal (A) may then be determined in accordance with this information, for example to avoid collisions with the objects and / or to stay on the drivable surface area.
[0259] The actuator (10), which is preferably integrated in the robot (100), may be given by a brake, a propulsion system, an engine, a drivetrain, or a steering of the vehicle (100). The control signal (A) may be determined such that the actuator (10) is controlled such that robot (100) avoids collisions with the detected objects. R. 414010
[0260] - 22 -
[0261] Alternatively or additionally, the control signal (A) may also be used to control the display (10a), e.g., for displaying the trajectory signal (t) and / or output signal (c). It can also be imagined that the control signal (A) may control the display (10a) such that it produces a warning signal if feasibility value determined for a trajectory signal (t) is equal to or above a predefined threshold. The warning signal may be a warning sound and / or a haptic signal, e.g., a vibration of a steering wheel of the vehicle.
[0262] In further embodiments, the at least partially autonomous robot may be given by another mobile robot (not shown), which may, for example, move by flying, swimming, diving or stepping. The mobile robot may, inter alia, be an at least partially autonomous lawn mower, or an at least partially autonomous cleaning robot. In all of the above embodiments, the control signal (A) may be determined such that propulsion unit and / or steering and / or brake of the mobile robot are controlled such that the mobile robot may avoid collisions with identified objects.
[0263] In a further embodiment, the at least partially autonomous robot may be given by a gardening robot (not shown), which uses the sensor (30), preferably an optical sensor, to determine a state of plants in the environment (20). The actuator (10) may control a nozzle for spraying liquids and / or a cutting device, e.g., a blade. Depending on an identified species and / or an identified state of the plants, an control signal (A) may be determined to cause the actuator (10) to spray the plants with a suitable quantity of suitable liquids and / or cut the plants.
[0264] In even further embodiments, the at least partially autonomous robot may be given by a domestic appliance (not shown), like e.g. a washing machine, a stove, an oven, a microwave, or a dishwasher. The sensor (30), e.g., an optical sensor, may detect a state of an object which is to undergo processing by the household appliance. For example, in the case of the domestic appliance being a washing machine, the sensor (30) may detect a state of the laundry inside the washing machine. The control signal (A) may then be determined depending on a detected material of the laundry.
[0265] Figure 5 shows an embodiment in which the control system (40) is used to control a manufacturing machine (11), e.g., a punch cutter, a cutter, a gun drill or R. 414010
[0266] - 23 -
[0267] a gripper, of a manufacturing system (200), e.g., as part of a production line. The manufacturing machine may comprise a transportation device, e.g., a conveyer belt or an assembly line, which moves a manufactured product (12). The control system (40) controls an actuator (10), which in turn controls the manufacturing machine (11).
[0268] The sensor (30) may be given by an optical sensor which captures properties of, e.g., a manufactured product (12). The manufacturing machine (11) may be controlled by the contract-based planner and controller component to move manufacturing machine (11) along a trajectory, especially in order to interact with the manufactured product (12). The planner (61) may determine a trajectory signal (t) comprising or consisting of joint positions and / or joint speeds and / or joint accelerations for joints of the manufacturing machine (11), especially for actuators (10) of the joints.
[0269] The machine learning system (60) may especially be trained to predict a feasibility value for the trajectory of joint positions. The actuators (10) may then be controlled depending on the determined feasibility value. For example, the actuators (10) may be controlled to cut the manufactured product at a specific location of the manufactured product itself.
[0270] The term "computer" may be understood as covering any devices for the processing of pre-defined calculation rules. These calculation rules can be in the form of software, hardware or a mixture of software and hardware.
[0271] In general, a plurality can be understood to be indexed, that is, each element of the plurality is assigned a unique index, preferably by assigning consecutive integers to the elements contained in the plurality. Preferably, if a plurality comprises N elements, wherein N is the number of elements in the plurality, the elements are assigned the integers from 1 to N. It may also be understood that elements of the plurality can be accessed by their index.
[0272] If an element is defined according to a formula or formulae, this may be understood as the element being similar to the provided formulae. The exact formulae defining the element may be different from the provided formulae if an R. 414010
[0273] - 24 -
[0274] overall functionality is kept. For example, when using constraint optimization, the constraints may each be weighted by a weight or the cost function may comprise a term that is simply a linear offset, i.e., an addition of a number.
Claims
R. 414010- 25 -Claims1. Computer-implemented method for configuring a planner (61), wherein the planner (61) is configured for providing trajectory signals (t) for controlling a technical system (100, 200), the method comprising the steps of:• Providing a feasibility function (fi.) to the planner (61), wherein the feasibility function (fi.) is configured to determine a feasibility value for a trajectory signal (t), wherein the feasibility value indicates a minimal cost of a soft-constrained model predictive control optimization used by a controller (62) of the technical system (100, 200) for controlling the technical system (100, 200), wherein the controller (62) is configured for controlling the technical system (100, 200) based on trajectory signals provided by the planner (61);• Including the feasibility function (fi.) as a component of an optimization problem of the planner (61) thereby configuring the planner (61), wherein the planner (61) provides a trajectory signal (t) based on the optimization problem.
2. Method according to claim 1, wherein the feasibility function (ĥ) is configured for approximating a result of the soft-constrained model predictive control optimization for a given trajectory signal (t, tᵢ).
3. Method according to claim 2, wherein providing the feasibility function (ĥ) comprises the steps of:• Providing at least one trajectory signal (tᵢ) as training input (xᵢ)• Providing for the at least one trajectory signal (tᵢ) a feasibility value as desired training output (oᵢ), wherein the feasibility value is determined by providing the trajectory signal (tᵢ) to the soft-constrained model predictive optimization;• Performing supervised training of a machine learning system (60) using the training input (xᵢ) and desired training output (oᵢ) as training data;R. 414010- 26 -• Providing the machine learning system (60) as the feasibility function (ĥ).
4. Method according to claim 3, wherein the at least one trajectory signal (t is provided by the planner (61).
5. Method according to any one of the claims 2 to 4, wherein the steps of providing the feasibility function (ĥ) are conducted by the controller (62).
6. Method according to any one of the claims 2 to 5, wherein the machine learning system (60) is a neural network.
7. Computer-implemented method for determining a control signal (A) for controlling a technical system (100, 200) comprising the steps of:• Configuring a planner (61) of the technical system (100, 200) according to any one of the preceding claims;• Determining a trajectory signal (t) using the adapted cost function of the planner (61);• Providing the trajectory signal (t) to a controller (62) of the technical system (100, 200);• Determining, by the controller (62), an output signal (c) characterizing an optimal control input to of the actuator (10), wherein the controller (62) determines the output signal (c) based on the trajectory signal (t) using a soft-constrained model predictive control optimization;• Determining the control signal (A) based on the output signal (c).
8. Computer-implemented method for controlling a technical system (100, 200) using a planner (61) and a controller (62), the method comprising the steps of:• Providing a planner (61) and a controller (62) to the technical system (100, 200);• Performing a handshake operation between the planner (61) and the controller (62), wherein the handshake operation comprises configuring the planner (61) using a feasibility function (ĥ) according to any one of the claims 1 to 6;R. 414010- 27 -• Controlling the technical system (100, 200) comprising:• Determining a trajectory signal (t) using the planner (61);• Providing the trajectory signal (t) as input to the controller (62); • Determining, by the controller (62), a control output (c);• Determining a control signal (A) based on the trajectory signal control output (c);• Controlling the technical system (100, 200) using the control signal (A).
9. Method according to any one of the preceding claims, wherein the technical system is a robot (100), in particular a mobile robot, in particular an at least partially automated vehicle, or a manufacturing machine (200), or a valve.
10. Technical system (100, 200) comprising:• A planner (61) configured for providing trajectory signals (t) for controlling the technical system (100, 200);• A controller (62) configured for determining control output (c), wherein a control signals (A) is determined based on the control output (c), wherein the control signal (A) is configured for controlling at least one actuator (10) of the technical system (100), preferably a subset or all actuators (10) of the technical system (100, 200) or all actuators (10) of the technical system (100, 200), wherein the controller (62) determines the control output (c) using a soft-constrained model predictive control optimization,wherein the planner (61) uses a cost function for determining trajectory signals (t), wherein the cost function comprises a potentially weighted cost term, wherein the cost term is a feasibility function (ĥ) characterizing a minimal cost of the soft-constrained model predictive control optimization for a given trajectory signal (t).
11. Training system (140), which is configured to carry out the training method according to claim 3.R. 414010- 28 -12. Control system (40), which is configured to carry out the method according to claim 7.
13. Computer program that is configured to cause a computer to carry out the method according to any one of the claims 1 to 9 with all of its steps if the computer program is carried out by a processor (45, 145).
14. Machine-readable storage medium (46, 146) on which the computer program according to claim 13 is stored.