Real-time trajectory generation for a movable robot

WO2026175527A1PCT designated stage Publication Date: 2026-08-27ABB (SCHWEIZ) AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/054900
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2026-08-27

Smart Images

  • Figure EP2025054900_27082026_PF_FP_ABST
    Figure EP2025054900_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method (100) for computing a trajectory (3) for a movable robot (1), said trajectory (3) comprising at least time series of quantities qk that characterize a motion of the movable robot (1) with respect to available degrees of freedom k = 1,..., K, wherein this motion moves at least one reference point (2) of the movable robot (1) from a given initial position qt to a given final position qf, the method (100) comprising the steps of: • providing (110) at least one physical relationship (4a) between the quantities qk, and / or at least one constraint (4b) for one or more of the quantities qk; • determining (120), by at least one first machine learning model (5), based at least in part on the initial position qt and the final position qf, an output (5a) that is indicative of one or more of the quantities qk; and • computing (140) one or more quantities qk from the output (5a), and / or rating (150) the usability (5b) of the output (5a) for indicating the one or more of the quantities qk, based at least in part on the physical relationship (4a) and / or constraint (4b).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ABB Schweiz AG 24.02.2025

[0002] A 19488 WO

[0003] REAL-TIME TRAJECTORY GENERATION FOR A MOVABLE ROBOT

[0004] FIELD OF THE INVENTION

[0005] The invention relates to the field of robotics, and in particular to the trajectory planning for moving at least one reference point of a movable robot, such as a tool center point, TCP, from a given initial position to a given final position.

[0006] BACKGROUND

[0007] A movable robot usually has a fixed number of degrees of freedom along which it can move. For example, members of the robot may be rotated about a number of joints, or they may be extended and retracted linearly. Therefore, if a given reference point of the movable robot, such as a tool center point, TCP, is to be moved from a given initial position qtto a given final position qf, this desired motion will have to be translated into motions along said degrees of freedom. Also, there may be obstacles between the initial position qtand the final position qf. Therefore, it may not be possible to move the reference point from qtto qfin a straight line. Moreover, for a practical application, what is sought is not just any trajectory in space of time that is mechanically possible. Rather, a trajectory that is optimal with respect to at least one optimization goal is sought.

[0008] The usual way to obtain a trajectory that leads from qtto qfis to model the movable robot by means of an explicit model and solve an optimization problem based on this model. However, such an optimization poses a major challenge in terms of computational effort, especially for integrated and path trajectory optimization.

[0009] Integrated path and trajectory optimization are therefore not feasible in real-time.

[0010] Rather, generation of the path geometry is usually separated from trajectory generation that, given the path, boils down to a one-dimensional optimization problem.

[0011] Therefore, trajectory planners based on machine learning models have been proposed. The basic idea is to train neural networks or other machine learning models in a one-time effort to predict a trajectory given the initial position qtand the final position qf. However, such training requires an amount of training examples that is usually not available in most practical applications.

[0012] OBJECTIVE OF THE INVENTION

[0013] It is therefore an objective of the invention to allow a fast computation of a trajectory for a movable robot by means of machine learning models in a manner that is trainable with a lesser amount of training examples.

[0014] This objective is achieved by a method for computing a trajectory according to a first independent claim, a method for training an arrangement of machine learning models according to a second independent claim, and an apparatus according to a third independent claim. Further advantageous embodiments are detailed in the respective dependent claims.

[0015] DISCLOSURE OF THE INVENTION

[0016] The invention provides a computer-implemented method for computing a trajectory for a movable robot. The trajectory comprising at least time series of quantities qkthat characterize a motion of the movable robot with respect to available degrees of freedom k =

[0017]

[0018] This motion moves at least one reference point of the movable robot from a given initial position qt to a given final position qf.

[0019] In the course of the method, at least one physical relationship between the quantities qk, and / or at least one constraint for one or more of the quantities qk, is provided. At least one first machine learning model determines, based at least in part on the initial position qtand the final position qf, an output that is indicative of one or more of the quantities qk. Herein, “indicative” comprises that the output may already directly contain the one or more of the quantities qk, but also that the output may be usable for calculating the one or more of the quantities qktogether with further information, such as the output from a second machine learning model.

[0020] The one or more quantities qkmay be computed from the output based at least in part on the physical relationship and / or constraint. In particular, this may mean that an algorithm or other process for computing the quantities qkincorporates knowledge about the physical relationship and / or constraint. Alternatively or in combination to this,the usability of the output of the first machine learning model for indicating the one or more of the quantities qkmay be rated based at least in part on the physical relationship and / or constraint. Again, this may mean that the rating employs some kind of knowledge about the physical relationship and / or constraint. The rating may, for example, be done during training of the first machine learning model. That is, a loss function used for the training may incorporate a term that penalizes non-compliance with the physical relationship and / or constraint. But the rating may also be done during inference after the training. For example, the output of the machine learning model, and / or any values of quantities qkdetermined from it, may be given more weight if the rating is better.

[0021] The inventors have found that the consideration of the physical relationship and / or constraint relieves the first machine learning model from learning knowledge that is already available from another source. In particular, if such knowledge is used in an explicit computation of one or more of the sought quantities qkfrom the output of the first machine learning model, some of the total work required for obtaining the qkmay be shifted from the first machine learning model into the explicit computation. This means that the first machine learning model may be smaller, so that its behaviour is characterized by a lesser number of trainable parameters. This in turn means that the training may be performed using a lesser number of training examples. But even if no explicit computation is used, and the first machine learning model directly predicts the sought quantities qkfrom the initial position qtand the final position qf, the consideration of the knowledge about the physical relationship and / or constraint in the loss function will guide the training to convergence based on a much smaller amount of training examples. What is already known (and is available as an explicit first-principles model) does not need to be learned (from data).

[0022] This is based on the insight that standard fully connected neural networks, such as multi-layer perceptrons, MLP, are general function approximators that can approximate then behaviour of any given function, but are inefficient in terms of a requirement for training examples because they do not exploit a known structure of the solution.

[0023] Moreover, the explicit computation allows to strictly enforce known physical relationship or constraints. If such a physical relationship or constraint is merely considered in the loss function during the training of the first machine learning model, then the trainingmay converge on a solution where a small violation of said physical relationship or constraint is accepted for the sake of improving some other aspect and achieving an overall better rating by the loss function. But such violation, however small, may in fact make this solution inaccessible in reality because some fundamental constraints, such as the conservation laws for energy or momentum and the fundamental laws of thermodynamics, are inviolable. Figuratively speaking, the first machine learning model might just as well predict that a person’s IQ will double as soon as the body temperature reaches 45°C. Even if this is true, the practical use is zero because this temperature is invariably fatal.

[0024] In a simple analogy, the result of the trajectory planning may be understood as a musical score for an orchestra. Each degree of freedom corresponds to one instrument in this orchestra. The overall goal of the musical score is to move the at least one reference point of the movable robot from qtto qf.

[0025] Apart from facilitating the computation, the consideration of the physical relationship and / or constraint, in particular by means of explicit computation, also improves the overall explainability. When using machine learning, it is always an issue that the processing is “black box” to some degree and it is difficult to understand the reasons for a decision of the machine learning model. By explicitly modelling hard-and-fast knowledge and constraints, consideration of this is guaranteed.

[0026] In a particularly advantageous embodiment, the physical relationship comprises that at least one of the quantities qkis a derivative or integral of another quantity qk. This is a very frequent use case when completely describing motion, as the state of motion is characterized not only by a momentary position, but also by a velocity and an acceleration. Therefore, in a further particularly advantageous embodiment, the quantities qkcomprise a position q, a velocity q and an acceleration q with respect to at least one degree of freedom.

[0027] In particular, the acceleration q may be determined by the first machine learning model. The velocity q may then be determined as an integral of the acceleration q. The position q may then be determined as an integral of the velocity q. In this manner, only one single quantity q is predicted by the first machine learning model, and everything else is computed explicitly. The final solution will strictly adhere to the fundamentalphysical laws that velocity is the integral of acceleration, and position is the integral of velocity.

[0028] Despite “baking” the hard constraints into the model structure, the outputs for the position q and the velocity q may still be used as feedback in the loss function of the training process for the first and / or second machine learning model. Thus, all dimensions of the available training data may be leveraged, i.e. , position q, velocity q, acceleration q and motion time AT, instead of “only” formulating a loss function that penalizes deviations of the predicted acceleration q from the acceleration time series in the data.

[0029] In a further particularly advantageous embodiment, a second machine learning model computes a further output based at least in part on the initial position qt and the final position qf. One or more quantities qkare computed based on the outputs of the first machine learning model and the second machine learning model. In this manner, it may be fixed explicitly how the final quantities qkare computed from their main constituents. But any unknown influences on these constituents that are hard (or even impossible) to model explicitly may be learned by the first and second machine learning models.

[0030] If such a further output is computed, then this may go into the computation of the one or more quantities qkon top of the output of the first machine learning model. Also, a usability of the further output of the second machine learning model for indicating, together with the output of the first machine learning model, may be rated based at least in part on the physical relationship and / or constraint. To this end, the further output may be rated directly, or indirectly by rating the computed quantities qkinto which the further output has gone.

[0031] In one example, the further output may comprise a total motion time T, and / or a duration of discretized time segments AT for a given number of discretization points. This may be learned from training examples where actual motion is recorded as a function of the time. This output T, AT is not only usable for being combined with the output of the first machine learning model into the sought quantity qkor time series thereof. Rather, it is also good on its own, e.g., to understand the achievable performance of the movable robot in terms of cycle time.In particular, the first and / or second machine learning models may comprise a neural network that is specifically adapted for processing time series, such as a recurrent neural network, RNN, or a long short-term memory, LSTM. However, this is not required. Any other neural network architecture, such as a fully connected multi-layer perceptron, MLP, may be used just as well.

[0032] In a further particularly advantageous embodiment, usability of a quantity qkthat has been computed from the output of the first machine learning model based at least in part on the physical relationship and / or constraint is rated based at least in part on the physical relationship and / or constraint. For example, the result of the computation may be rated by a loss function of the first machine learning model based on whether this result is compliant with the physical relationship and / or constraint. In this manner, the first machine learning model may be trained to produce outputs whose processing into quantities qkwill yield compliant results. For example, if an acceleration q as output of the first machine learning model is numerically integrated to form a velocity q, and this velocity q is integrated further to form a position q, even small amounts of noise on the acceleration q may accumulate to considerable deviations of the velocity q and the position q. Measuring these deviations in the loss function for the training of the first machine learning model helps to suppress this noise at its source.

[0033] In a further particularly advantageous embodiment, the at least one physical constraint comprises

[0034] • a maximum force and / or torque that may be exerted in at least one degree of freedom of the movable robot; and / or

[0035] • a maximum force and / or moment that may be borne by at least one member and / or bearing of the movable robot.

[0036] These are the most important constraints of a given movable robot. If the planned trajectory demands more force and / or torque that the movable robot is able to provide, then the planned trajectory is not feasible. At the very least, the movable robot will take too much time to move the designated reference point to designated positions.

[0037] Likewise, if a member and / or bearing of the movable robot is subjected to an overly high force and / or moment, the member and / or bearing may be damaged permanently. Because these constraints are dynamic constraints, they are the main complications that are responsible for the complexity of trajectory planning. It is a goal of thetrajectory planning to operate the movable robot close to the border of what is mechanically possible, but without ever crossing this border. Furthermore, the maximum force and / or torque that may be exerted, i.e. , the maximum actuator effort, may itself be a function of the velocity q. Typically, available torque decreases with the velocity q, which further complicates the optimization. The machine learning based approach may implicitly take this further constraint into account.

[0038] Said constraints are differentiable, meaning that backpropagation of any errors regarding these constraints onto parameters of the first and / or second machine learning model is possible.

[0039] In a further particularly advantageous embodiment, at least one quantity qkrepresents the linear progress of a movement along a path of a given shape. In this manner, if the shape of the sought trajectory is already known, the problem at hand can be reduced to a one-dimensional problem that is a lot easier.

[0040] In a further particularly advantageous embodiment, at least one quantity qk, and / or at least one time series thereof, is further optimized towards a given goal by means of an iterative optimization algorithm. By providing a faster way to obtain the quantity qk, it is possible to obtain a first candidate for qkand / or the corresponding time series that already goes some way towards the optimal qkor time series, and use the optimization algorithm only for covering the “last mile” towards the optimum. It is also possible to obtain multiple candidates for qkand / or the corresponding time series, and test each candidate for whether it is optimal with respect to the given goal. If the obtaining of one single candidate already took too long, there simply would not be any time available for performing a further optimization, or even testing any further candidates.

[0041] In a further particularly advantageous embodiment, the iterative optimization algorithm is chosen to be too slow for a direct optimization of the quantity qk, and / or the time series thereof, from the given initial position qt and the given final position qf. That is, because the machine learning based determination of quantity qk, and / or the time series thereof, already provides a good starting point for the iterative optimization algorithm, fewer iterations are needed than would be required for a direct optimization from qtand qf. Every single of these iterations may then take more time.In a further particularly advantageous embodiment, the first machine learning model, and / or the second machine learning model, further gets a time derivative qkof at least one of the quantities qkas input. This facilitates the computation of trajectories in a situation where the direction has to be changed suddenly, at the price that the first machine learning model becomes more complex and takes more computation time. But because the consideration of the physical relationship and / or constraint has saved so much time, part of this savings may be given back in order to accommodate changes in direction. How quick such a change in direction can be made chiefly depends on the current velocity. A small change with respect to the current direction can be made quickly, whereas it is most difficult to bring the movable robot to a complete standstill and then get it moving in the opposite of the current direction of motion. That is, considering also the time derivative qkfacilitates reactive motion planning.

[0042] The invention also provides a computer-implemented method for training an arrangement of a first machine learning model and a second machine learning model for use in the embodiment of the method described above where the output of the second machine learning model comprises a total motion time T, and / or a duration of discretized time segments AT for a given number of discretization points.

[0043] In the course of this method, a set of training examples is provided. Each training example comprises a given initial position qi ta given final position qf, and at least one time series of at least one quantity qkthat characterizes a motion of the movable robot with respect to available degrees of freedom k = 1,

[0044]

[0045] such as position q, velocity q, and / or acceleration q. The motion described by the respective training example moves at least one reference point of the movable robot from the given initial position qt to the given final position qf.

[0046] The second machine learning model is trained towards the goal that the outputted time segment durations AT match corresponding time segment durations AT in the time series of the respective training example.

[0047] The first machine learning model is trained towards the goal that a time series of the at least one quantity qkobtained based on the outputs of the first and second machine models matches the time series qkin the respective training example.In this manner, the two machine learning models may be at least pre-trained independently of one another, which saves a lot of effort during this pre-training phase. In particular, if the second machine learning model is trained first, and has trouble reproducing the correct time segments AT, one need not bother starting with any optimization of the first machine learning model before this problem is sorted out.

[0048] In a particularly advantageous embodiment, for computing qkduring training of the first machine learning model, time segments AT from the respective training examples are used in addition to the output of the first machine learning model. In this manner, the first machine learning model may be trained in parallel with the second machine learning model. That is, the training of the first machine learning model does not have to wait for the second machine learning model to gain sufficient proficiency at predicting time segments AT.

[0049] The invention also provides an apparatus for computing a trajectory for a movable robot, said trajectory comprising at least time series of quantities qkthat characterize a motion of the movable robot with respect to available degrees of freedom k = 1,

[0050]

[0051] This motion moves at least one reference point of the movable robot from a given initial position qtto a given final position qf.

[0052] This apparatus comprises a first machine learning model that is configured to predict, based at least in part on the initial position qtand the final position qf, an output that is indicative of one or more of the quantities qk. As discussed before, “indicative” comprises that that the output may already directly contain the one or more of the quantities qk, but also that the output may be usable for calculating the one or more of the quantities qktogether with further information, such as the output from a second machine learning model.

[0053] The apparatus comprises such a second machine learning model. This second machine learning model is configured to predict, based at least in part on the initial position qtand the final position qf, a duration of discretized time segments AT for a given number of discretization points in the time series. In this manner, an output of the first machine learning model may be brought together with time information.For this purpose, the apparatus also comprises an explicit computation module. This explicit computation module computes, from the outputs of the first machine learning model and the second machine learning model, quantities qkand / or time derivatives thereof. In particular, as discussed before, the first machine learning model may predict an acceleration q, and from this and the time segments AT, the explicit computation module may compute velocity q and position q.

[0054] As discussed before, this architecture has multiple advantageous effects. First, a part of the total work is offloaded to the explicit computation module, so less work remains to be performed by the first and second machine learning models. Second, the remaining work to be performed by machine learning is split into two sub-tasks. Each of these sub-tasks may be performed by a relatively small, specialized machine learning model. As discussed before, the first and second machine learning models may be trained independently from one another. The total amount of training, and thus also the total amount of training examples, that is required is considerably less than what would be required to train one monolithic machine learning model for directly predicting all of acceleration q, velocity q and position q from the initial position qtand the final position f-

[0055] Because they may be fully or at least partially computer-implemented, the present methods may be embodied in the form of a software. The invention therefore also relates to a computer program with machine-readable instructions that, when executed by one or more computers and / or compute instances, cause the one or more computers and / or compute instances to perform one of the methods described above. Examples for compute instances include virtual machines, containers or serverless execution environments in a cloud. The invention also relates to a machine-readable data carrier and / or a download product with the computer program. A download product is a digital product with the computer program that may, e.g., be sold in an online shop for immediate fulfilment and download to one or more computers. The invention also relates to one or more compute instances with the computer program, and / or with the machine-readable data carrier and / or download product.

[0056] DESCRIPTION OF THE FIGURES

[0057] In the following, the invention is described using Figures without any intention to limit the scope of the invention. The Figures show:Figure 1: Exemplary embodiment of the method 100 for computing a trajectory 3 for a movable robot 1;

[0058] Figure 2: Exemplary embodiment of the method 200 for training an arrangement of a first machine learning model 5 and a second machine learning model 6;

[0059] Figure 3: Exemplary embodiment of an apparatus 10 for computing a trajectory 3 for a movable robot 1;

[0060] Figure 4: Comparison of the predictions of position q, velocity q and acceleration q obtained using a standard neural network (Figure 4a) and using the proposed method 100 (Figure 4b).

[0061] Figure 1 is a schematic flow chart of an exemplary embodiment of the method 100 for computing a trajectory 3 for a movable robot 1. The sought trajectory 3 comprises at least time series of quantities qkthat characterize a motion of the movable robot 1 with respect to available degrees of freedom k = 1,

[0062]

[0063] This motion moves at least one reference point 2 of the movable robot 1 from a given initial position qt to a given final position qf. That is, the method 100 starts from a situation where the movable robot 1, the reference point 2, the initial position qtand the final position qfare given.

[0064] In step 110, at least one physical relationship 4a between the quantities qk, and / or at least one constraint 4b for one or more of the quantities qk, is provided.

[0065] According to block 111 , the physical relationship 4a may comprise that at least one of the quantities qkis a derivative or integral of another quantity qk.

[0066] According to block 112, the quantities qkmay comprise a position q, a velocity q and an acceleration q with respect to at least one degree of freedom.

[0067] According to block 113, the at least one physical constraint 4b may comprise

[0068] • a maximum force and / or torque that may be exerted in at least one degree of freedom of the movable robot (1); and / ora maximum force and / or moment that may be borne by at least one member and / or bearing of the movable robot (1).

[0069] In step 120, based at least in part on the initial position qtand the final position qf, at least one first machine learning model 5 determines an output 5a that is indicative of one or more of the quantities qk.

[0070] According to block 121, an acceleration q may be determined by the first machine learning model 5. The velocity q may then be determined, according to block 122, as an integral of the acceleration q. The position q may then be determined, according to block 123, as an integral of the velocity q. In the example shown in Figure 1 , the integrals are computed by an explicit computation module 8.

[0071] According to block 124, the first machine learning model 5 may further get a time derivative qkof at least one of the quantities qkas input.

[0072] In the example shown in Figure 1, in step 130, a second machine learning model 6 computes a further output 6a based at least in part on the initial position qtand the final position qf.

[0073] According to block 131, this further output 6a may comprise a total motion time T, and / or a duration of discretized time segments AT for a given number of discretization points.

[0074] According to block 132, the second machine learning model 6 may further get a time derivative qkof at least one of the quantities qkas input.

[0075] In step 140, one or more quantities qkmay be computed from the output 5a.

[0076] According to block 141, if there is an output 6a from a second machine learning model 6, then one or more quantities qkmay be computed based at least in part on the output 5a of the first machine learning model 5 and the output 6a of the second machine learning model 6.According to block 142, at least one quantity qkmay represent the progress of a movement along a path of a given shape.

[0077] Alternatively or in combination to computing one or more quantities qkfrom the output 5a, in step 150, the usability 5b of the output 5a for indicating the one or more of the quantities qkmay be rated based at least in part on the physical relationship 4a and / or constraint 4b.

[0078] According to block 151, the usability 5b of a quantity qkthat has been computed from the output 5a of the first machine learning model 5 based at least in part on the physical relationship 4a and / or constraint 4b may be rated based at least in part on the physical relationship 4a and / or constraint 4b. That is, it may be determined how well the final result is compliant with the physical relationship 4a and / or constraint 4b.

[0079] According to block 152, if there is a further output 6a from a second machine learning model 6, a usability 6b of this further output 6a for indicating, together with the output 5a from the first machine learning model 5, the one or more of the quantities qk, may be rated based at least in part on the physical relationship 4a and / or constraint 4b.

[0080] The quantities qk, and / or their time series, computed in step 140 constitute the sought trajectory 3. In the example shown in Figure 1, in step 160, at least one such quantity qkis further optimized towards a given goal by means of an iterative optimization algorithm. This yields an optimized quantity qk’*, and thus also an optimized trajectory 3*. This optimized trajectory 3* may be applied to the movable robot 1.

[0081] According to block 161, the iterative optimization algorithm may be chosen to be too slow for a direct optimization of the quantity qk, and / or the time series thereof, from the given initial position qt and the given final position qf. As discussed before, such an algorithm becomes usable in the present method by virtue of part of the work being offloaded to the first machine learning model 5 (and optionally second machine learning model 6).

[0082] Figure 2 is a schematic flow chart of an embodiment of the method 200 for training an arrangement of a first machine learning model 5 and a second machine learning model 6 for use in the method 100 discussed above.In step 210, a set of training examples 7 is provided. Each training example 7 comprises a given initial position qi ta given final position qf, and at least one time series of at least one quantity qkthat characterizes a motion of the movable robot with respect to available degrees of freedom k = 1,

[0083]

[0084] As discussed before, this motion moves at least one reference point 2 of the movable robot 1 from the given initial position qtto the given final position qf.

[0085] In step 220, the second machine learning model 6 is trained towards the goal that the outputted time segments AT match corresponding time segments AT in the time series of the respective training example 7. That is, the training is a supervised training.

[0086] In step 230, the first machine learning model 5 is trained towards the goal that a time series of the at least one quantity qkobtained based on the outputs of the first and second machine models 5, 6 matches the time series qkin the respective training example 7.

[0087] The first machine learning model 5 can also be trained before the second machine learning model 6, or concurrently with the second machine learning model 6.

[0088] According to block 231, for computing qkduring training of the first machine learning model 5, time segments AT from the respective training examples 7 may be used in addition to the output of the first machine learning model 5. As discussed before, this facilitates concurrent training of both the first machine learning model 5 and the second machine learning model 6 because the training of the first machine learning model 5 does not have to wait for the training of the second machine learning model 6 to make a notable progress.

[0089] Figure 3 is a schematic block diagram of an apparatus 10 for computing a trajectory 3 for a movable robot 1. Again, the trajectory 3 comprises at least time series of quantities qkthat characterize a motion of the movable robot with respect to available degrees of freedom k = 1,

[0090]

[0091] This motion moves at least one reference point (2) of the movable robot (1) from a given initial position qt to a given final position qf.The apparatus 10 comprises a first machine learning model 5. This first machine learning model 5 is configured to predict, based at least in part on the initial position qtand the final position qf, an output 5a that is indicative of one or more of the quantities qk. In the example shown in Figure 3, the output 5a comprises accelerations qk. These accelerations qkalready belong to the sought quantities qk, but it is also indicative of the velocities qkand the positions qkthat are determinable from the accelerations qkby integration.

[0092] The apparatus 10 also comprises a second machine learning model 6. This second machine learning model 6 is configured to predict, based at least in part on the initial position qtand the final position qf, a duration of discretized time segments AT for a given number of discretization points in the time series. These time segments AT are the output 6a of the second machine learning model 6.

[0093] The apparatus 10 further comprises an explicit computation module 8. From the output 5a of the first machine learning model 5 and the output 6a of the second machine learning model 6, this explicit computation module 8 computes quantities qkand / or time derivatives thereof that characterize the sought trajectory 3. In the example shown in Figure 3, the explicit computation module 8 computes velocities qkand positions qkby integration from the accelerations qk.

[0094] The accelerations qk, the velocities qkand the positions qkare rated by a loss function that measures their usability 5b in terms of compliance with the known physical relationship that speed is the integral of acceleration and position is the integral of speed. Gradients of this loss function may be calculated by any suitable algorithmic differentiation framework abstractly labelled V in Figure 3, such as CasADi. Thereby, it is possible to integrate penalty terms based on such dynamic constraints in machine learning frameworks such as PyTorch or Tensorflow. By means of the gradients, parameters that characterize the behaviour of the first machine learning model 5, and / or of the second machine learning model 6, may be changed towards the goal of minimizing the value of the loss function.

[0095] Figure 4 illustrates, on a concrete use case, the advantages of determining the position q, velocity q and acceleration q according to the method 100 described above overusing a normal (“vanilla”) neural network that does not use any physical prior knowledge.

[0096] Figure 4a shows the results obtained with a standard neural network. The acceleration q was only predicted using this neural network (P). For the velocity q, results obtained by direct prediction with the neural network on the one hand (P) and by calculation from the predicted acceleration q on the other hand (C) are shown. Likewise, for the position q, results obtained by direct prediction with the neural network on the one hand (P) and by calculation from the predicted acceleration q and velocity q on the other hand (C) are shown. All results are shown together with the respective ground truth (GT).

[0097] It is evident from Figure 4a that the computing of the velocity q and the position q from the predicted acceleration q results in velocities q and positions q that are physically not compatible. Also, at the end of the motion, the position q obtained by calculation (C) has a remaining offset from the ground truth (GT). That is, the motion does not reach its intended target position.

[0098] Figure 4b shows the results obtained using the method 100 described above. The acceleration q was predicted using the neural network, the velocity q and position q were obtained by integration from the acceleration q, and compliance of the complete ensemble of acceleration q, velocity q and position q with the physical rules was rated by the loss function that was used for training the neural network. This causes the acceleration q, velocity q and position q to better obey the physical rules, and also improves the generalizability and explainability of the obtained results to improve. In particular, at the end of the motion, there is no more remaining offset between the calculated position q (C) and the ground truth (GT). This means that the desired target position is reached. Moreover, because the physical rules are explicitly “baked in” instead of learning them during training, a lesser amount of training data is required.List of reference signs:

[0099] 1 movable robot

[0100] 2 reference point of movable robot

[0101] 3 trajectory of movable robot

[0102] 3* optimized trajectory 3

[0103] 4a physical relationship between quantities qk

[0104] 4b constraint for at least one quantity qk

[0105] 5 first machine learning model

[0106] 5a output of first machine learning model 5

[0107] 5b usability of output 5a

[0108] 6 second machine learning model

[0109] 6a output of second machine learning model 6

[0110] 6b usability of output 6a

[0111] 7 training examples

[0112] 8 explicit computation module in apparatus 10

[0113] 10 apparatus for computing trajectory 3

[0114] 100 method for computing trajectory 3

[0115] 110 providing physical relationship 4a, constraint 4b

[0116] 111 choosing differential / integral physical relationship 4a

[0117] 112 choosing position q, velocity q and acceleration q as quantities qk113 choosing maximum forces, torques, moments as constraints 4b 120 determining output 5a of first machine learning model 5

[0118] 121 determining acceleration q by first machine learning model 5 122 determining velocity q by integration of acceleration q

[0119] 123 determining position q by integration of velocity q

[0120] 124 supplying time derivative qkto first machine learning model 5 130 determining output 6a of second machine learning model 6 131 choosing motion time T, and / or time segments AT, as outputs 6a 132 supplying time derivative qkto second machine learning model 6 140 computing quantities qkfrom output 5a based on information 4a / 4b 141 computing quantities qkfrom both outputs 5a and 6a

[0121] 142 choosing quantity qkthat represents linear movement progress 150 rating usability 5b based on information 4a / 4b151 rating usability 5b based at least in part on information 4a / 4b 152 rating usability 6b based at least in part on information 4a / 4b 160 further optimizing trajectory 3

[0122] 161 choosing iterative algorithm that is too slow for full direct optimization qksought quantities (positions) that characterize trajectory 3 qk’* optimized quantities qk

[0123] qkvelocities of quantities qk

[0124] qkaccelerations of quantities qk

[0125] qffinal position

[0126] qtinitial position

[0127] C computation result for quantities qk

[0128] GT ground truth for quantities qk

[0129] P prediction for quantities qk

[0130] T motion time

[0131] AT time segments

Claims

Claims:

1. A computer-implemented method (100) for computing a trajectory (3) for a movable robot (1), said trajectory (3) comprising at least time series of quantities qkthat characterize a motion of the movable robot (1) with respect to available degrees of freedom k = 1,wherein this motion moves at least one reference point (2) of the movable robot (1) from a given initial position qt to a given final position qf, the method (100) comprising the steps of:• providing (110) at least one physical relationship (4a) between the quantities qk, and / or at least one constraint (4b) for one or more of the quantities qk;• determining (120), by at least one first machine learning model (5), based at least in part on the initial position qt and the final position qf, an output (5a) that is indicative of one or more of the quantities qk; and• computing (140) one or more quantities qkfrom the output (5a), and / or rating (150) the usability (5b) of the output (5a) for indicating the one or more of the quantities qk, based at least in part on the physical relationship (4a) and / or constraint (4b).

2. The method (100) of claim 1, wherein the physical relationship (4a) comprises (111) that at least one of the quantities qkis a derivative or integral of another quantity qk.

3. The method (100) of any one of claims 1 to 2, wherein the quantities qkcomprise (112) a position q, a velocity q and an acceleration q with respect to at least one degree of freedom.

4. The method (100) of claims 1, 2 and 3, wherein• the acceleration q is determined (121) by the first machine learning model (5),• the velocity q is determined (122) as an integral of the acceleration q, and • the position velocity q is determined (123) as an integral of the velocity q.

5. The method (100) of any one of claims 1 to 4, wherein• a second machine learning model (6) computes (130) a further output (6a) based at least in part on the initial position qtand the final position qf, and • one or more quantities qkare computed (141) based on the outputs (5a, 6a) of the first machine learning model (5) and the second machine learning model (6).

6. The method (100) of claim 5, wherein the further output (6a) comprises (131) a total motion time T, and / or a duration of discretized time segments AT for a given number of discretization points.

7. The method (100) of any one of claims 5 to 6, wherein a usability (6b) of the further output (6a) for indicating, together with the output (5a) from the first machine learning model (5), the one or more of the quantities qk, is rated (152) based at least in part on the physical relationship (4a) and / or constraint (4b).

8. The method (100) of any one of claims 1 to 6, wherein the usability (5b) of a quantity qkthat has been computed from the output (5a) of the first machine learning model (5) based at least in part on the physical relationship (4a) and / or constraint (4b) is rated (151) based at least in part on the physical relationship (4a) and / or constraint (4b).

9. The method (100) of any one of claims 1 to 7, wherein the at least one physical constraint (4b) comprises (113)• a maximum force and / or torque that may be exerted in at least one degree of freedom of the movable robot (1); and / or• a maximum force and / or moment that may be borne by at least one member and / or bearing of the movable robot (1).

10. The method (100) of any one of claims 1 to 8, wherein at least one quantity qkrepresents (142) the progress of a movement along a path of a given shape.

11. The method (100) of any one of claims 1 to 9, wherein at least one quantity qk, and / or at least one time series thereof, is further optimized (160) towards a given goal by means of an iterative optimization algorithm.

12. The method (100) of claim 10, wherein the iterative optimization algorithm is chosen (161) to be too slow for a direct optimization of the quantity qk, and / or the time series thereof, from the given initial position qtand the given final position qf.

13. The method (100) of any one of claims 1 to 11 , wherein the first machine learning model (5), and / or the second machine learning model (6), further gets (124, 132) a time derivative qkof at least one of the quantities qkas input.

14. A computer-implemented method (200) for training an arrangement of a first machine learning model (5) and a second machine learning model (6) for use in the method (100) according to claim 6, and optionally further any one of claims 2 to 4 or 7 to 12, comprising the steps of:• providing (210) a set of training examples (7), each training example (7) comprising a given initial position qi ta given final position qf, and at least one time series of at least one quantity qkthat characterizes a motion of the movable robot with respect to available degrees of freedom k = 1,wherein this motion moves at least one reference point (2) of the movable robot (1) from the given initial position qt to the given final position qf, • training (220) the second machine learning model (6) towards the goal that the outputted time segments AT match corresponding time segments AT in the time series of the respective training example (7); and• training (230) the first machine learning model (5) towards the goal that a time series of the at least one quantity qkobtained based on the outputs of the first and second machine models (5, 6) matches the time series qkin the respective training example (7).

15. The method (200) of claim 13, wherein, for computing qkduring training of the first machine learning model (5), time segment durations AT from the respective training examples (7) are used (231) in addition to the output of the first machine learning model (5).

16. An apparatus (10) for computing a trajectory (3) for a movable robot (1), said trajectory (3) comprising at least time series of quantities qkthat characterize a motion of the movable robot with respect to available degrees of freedom k = 1,wherein this motion moves at least one reference point (2) of the movable robot (1) from a given initial position qtto a given final position qf, said apparatus (10) comprising:• a first machine learning model (5) that is configured to predict, based at least in part on the initial position qtand the final position qf, an output (5a) that is indicative of one or more of the quantities qk• a second machine learning model (6) that is configured to predict, based at least in part on the initial position qtand the final position qf, a duration of discretized time segments AT for a given number of discretization points in the time series; and• an explicit computation module (8) that computes, from the outputs (5a, 6a) of the first machine learning model (5) and the second machine learning model (6), quantities qkand / or time derivatives thereof.

17. A computer program, comprising machine-readable instructions that, when executed by one or more computers and / or compute instances, causes the oneor more computers to perform the method (100, 200) according to any one of claims 1 to 14.

18. A non-transitory machine-readable data carrier, and / or a download product, with the computer program of claim 16.

19. One or more computers and / or compute instances with the computer program of claim 16, and / or with the machine-readable data carrier and / or download product of claim 17.