Method for evaluating a control of a robot device
By training a machine learning model with quantile regression loss to simulate object behaviors in control scenarios, the method addresses the challenge of evaluating robot device controls against diverse and worst-case scenarios, ensuring safety and robustness.
Patent Information
- Application Number
- DE102023200231
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-01-12
- Publication Date
- 2025-05-22
- Estimated Expiration
- 2043-01-12
AI Technical Summary
Existing methods for evaluating controls of robot devices, particularly in autonomous driving, struggle to comprehensively test control actions against a wide range of realistic scenarios, including worst-case behaviors, which are crucial for ensuring safety.
A method involving the training of a machine learning model using quantile regression loss to simulate the behavior of objects in control situations, allowing for the simulation of realistic worst-case scenarios and thorough evaluation of control methods.
This approach enables a comprehensive and realistic simulation of various control scenarios, including worst-case behaviors, thereby ensuring the safety and robustness of control methods employed by robot devices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
State of the art
[0001] The present disclosure relates to methods for evaluating a control of a robot device, in particular for testing a control method for a robot device or for selecting a control (ie, a control action) in the operation of a robot device.
[0002] In recent years, autonomous driving has become a topic of great interest both in research and among the public. Autonomous vehicles have enormous potential, not only economically, but also for improving mobility options and potentially reducing carbon emissions. Like any control, autonomous driving involves making decisions in a given control situation to select a control action. These control actions should be safe, meaning they should not lead to dangerous situations. To test their reliability and safety, control systems for autonomous driving must be extensively tested. Since this would be too complex or even too dangerous with real-world tests, this is done using simulations.To this end, other road users will be simulated in such a way that the broadest possible spectrum of traffic situations that could occur in reality is covered, so that the control system under test can be comprehensively tested. In particular, worst-case scenarios (those caused by particularly unfavorable behavior by another road user) will also be covered.
[0003] Thus, methods for evaluating control systems for vehicles, or generally for robotic devices (such as robot arms, walking robots, etc.), are desirable that comprehensively cover possible sequences of control scenarios (i.e., various sequences that are plausible to occur in reality). The problem is solved as claimed.
[0004] The paper “Autoregressive Quantile Flows for Predictive Uncertainty Estimation” by Phillip Si et al., 2022, hereinafter referred to as Reference 1, describes normalization flows trained using a check score defined according to a tilted absolute error loss function.
[0005] DE 10 2021 204 961 A1 discloses a method for controlling a robot device, comprising providing demonstrations for a robot skill, each demonstration demonstrating a trajectory comprising a sequence of robot configurations, each described by an element of a predetermined configuration space having the structure of a Riemannian manifold, determining a representation of each trajectory as a vector of weights of predetermined basic movements of the robot device by searching a vector of weights that minimizes a distance measure between the combination of the basic movements according to the vector of weights and the demonstrated trajectory, wherein the combination is mapped to the manifold, determining a probability distribution of the vector of weights by fitting a probability distribution to the vector of weights,determined for the demonstrated trajectories and controlling the robot device by performing basic movements according to the determined probability distribution of vectors of weights.,
[0006] DE 10 2020 209 685 A1 discloses a method for controlling a robot device, comprising obtaining demonstrations for controlling the robot device, performing an initial training of a neural actuator network by imitation learning of the demonstrations, controlling the robot device by the initially trained neural actuator network to generate a plurality of trajectories of the robot device, each trajectory comprising a sequence of actions selected by the initially trained neural actuator network in a sequence of states, and observing the return for each of the actions selected by the initially trained neural actuator network, performing an initial training of a neural critic network by supervised learning, wherein the neural critic network is trained to determine the observed returns of the actions,which are selected by the initially trained neural actuator network, training the neural actuator network and the neural critic network by reinforcement learning, starting from the initially trained neural actuator network and the initially trained neural critic network, and controlling the robot device by the trained neural actuator network and the trained neural critic network.
[0007] DE 11 2020 005 156 T5 reveals reinforcement learning of tactile grasping strategies.
[0008] DE 10 2019 220 574 A1 discloses a computer-implemented method and apparatus for testing a machine, the method comprising: providing a set of exemplary trajectories, selecting a subset of the set, determining an output of the machine for a movement of the machine according to an exemplary trajectory of a subset, determining the result of the testing depending on the output.
[0009] DE 11 2019 007 601 T5 discloses a learning device and a learning method.
[0010] US 2022 / 0 395 975 A1 discloses a demonstration-conditioned reinforcement learning for imitation with few recordings. Disclosure of the invention
[0011] According to various embodiments, a method for evaluating a control of a robot device is provided, comprising determining demonstrations for the behavior of at least one object in control situations that include the at least one object and the robot device, training a machine learning model to map information about control situations to information about the behavior of the at least one object by means of a quantile regression loss for one or more undershoot proportions, performing a simulation, wherein the behavior of the robot device is simulated according to the control to be evaluated and the behavior of the at least one object in at least one control situation occurring in the simulation is simulated according to an output that the trained machine learning model outputs in response to information about the occurring control situation,and evaluating the control depending on events in the simulation.,
[0012] By training the machine learning model according to which the at least one object in the simulation behaves based on a quantile regression loss, it is possible to use behavior according to borderline cases (or within certain borderline cases), in particular a worst-case behavior, in the simulation for the object and thus to test or validate the control system to be evaluated (e.g., a control procedure to be tested) against such borderline cases. This allows for a thorough test within a realistic range of possible behaviors (e.g., other drivers) and ensures that a control procedure that is used is safe even for atypical behaviors that fall within such a spectrum.
[0013] A worst-case behavior that is covered does not have to be a theoretically possible worst-case behavior, but a realistic worst-case behavior, such as a particularly daring driver, as can occur in road traffic, ie a case that is most likely the worst possible case or a behavior that is the worst possible except for a residual risk (with a small probability, e.g. 0.001%).
[0014] An event on the basis of which the control is evaluated is, for example, the occurrence or non-occurrence of an accident such as a collision, etc. For example, the control is evaluated as unsuitable if a collision occurs in the simulation.
[0015] By training based on a quantile regression loss, the focus during training can be placed on learning the long tails of a learned probability distribution (which specifies or reflects the output of the trained machine learning model) with high accuracy, so that the output of the machine learning model includes or specifies extreme but realistic behavior.
[0016] Various examples of implementation are given below.
[0017] Embodiment 1 is a method for evaluating a control of a robot device as described above.
[0018] Embodiment 2 is a method according to embodiment 1, wherein the machine learning model is trained to output a quantile (in particular a high quantile) for an action of the at least one object in the control situation for a predetermined undershoot proportion in response to information about a control situation for the at least one object, and wherein the behavior of the at least one object in the at least one control situation occurring in the simulation is simulated according to the quantile that the machine learning model outputs for the occurring control situation.
[0019] This allows realistic boundary cases of behavior to be captured. The machine learning model (or multiple machine learning models) can also be used to determine multiple quantiles that capture multiple boundary cases (e.g., particularly harsh braking and particularly harsh acceleration).
[0020] Embodiment 3 is a method according to embodiment 1, wherein the machine learning model is trained to output, in response to information about a control situation for the at least one object, a plurality of quantiles for an action of the at least one object in the control situation for predetermined undershoot proportions and a plurality of simulations are carried out, wherein the behavior of the at least one object in the occurring control situation is simulated in each of the plurality of simulations according to a respective (e.g. sampled) value from a quantile set that is defined by the quantiles that the machine learning model outputs for the occurring control situation.
[0021] For example, a range between a lower bound for a behavior (e.g., an acceleration) and an upper bound for a behavior can be determined and sampled from this in several simulations to ensure that the control to be evaluated is robust to different behaviors of the object that are possible according to the quantile set.
[0022] Embodiment 4 is a method according to embodiment 1, wherein the machine learning model is trained to output, in response to information about a control situation for the at least one object, a plurality of quantiles for an action of the at least one object in the control situation for predetermined undershoot proportions, wherein the behavior of the at least one object in the occurring control situation is simulated according to a worst-case behavior from a quantile set that is defined by the quantiles that the machine learning model outputs for the occurring control situation.
[0023] In other words, it can be ensured that the control to be evaluated is robust against different behaviors of the object that are possible according to the quantile set by determining a worst-case behavior from the quantile set (e.g., particularly close collision) and using this as the basis for the simulation.
[0024] Embodiment 5 is a method according to embodiment 1, wherein the machine learning model specifies a normalization flow (or “normalizing flow”, see also reference 1) and the behavior of the at least one object in the at least one occurring control situation is simulated according to a sample from a probability distribution that the normalization flow outputs in response to the input of information about the occurring control situation.
[0025] By specifically training a normalization flow using a quantile regression loss, it is possible to ensure that the output of the normalization flow accurately reflects quantiles. For example, the output distribution of the normalization flow is a uniform distribution, and the training loss is a quantile regression loss for a specific undershoot fraction or the sum of quantile regression losses for multiple undershoot fractions.
[0026] As explained above, high quantiles can be used as a worst-case or borderline behavior in testing. A probabilistic approach can also be used, focusing on ensuring that certain quantiles are learned well (through the quantile loss function). For example, the following possibilities exist: • “Worst case”: o use the worst-case behavior in the simulation o sample uniformly from the quantile set in the simulation • “Probabilistic”: Sampling from the normalizing flow (where training targets the correct relevant quantiles)
[0027] The probabilistic approach ensures that even rare (or extreme) behavior is correctly reflected, so that any behavior that can realistically occur is covered.
[0028] Embodiment 6 is a method according to any one of embodiments 1 to 5, wherein the quantile regression loss is calculated using a tilted absolute error loss function.
[0029] This allows for an effective and simple calculation of the quantile regression loss.
[0030] Embodiment 7 is a test device configured to carry out the method according to one of embodiments 1 to 6.
[0031] Embodiment 8 is a computer program including instructions that, when executed by a processor, cause the processor to perform a method according to any one of embodiments 1 to 6.
[0032] Embodiment 9 is a computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method according to any one of embodiments 1 to 6.
[0033] In the drawings, like reference characters generally refer to the same parts throughout the several views. The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various aspects are described with reference to the following drawings. Fig. 1 shows a vehicle. Fig. Figure 2 illustrates an example of a simulation for testing a vehicle control system. Fig. 3 shows a flowchart illustrating a method for evaluating a controller of a robotic device according to an embodiment.
[0034] The following detailed description refers to the accompanying drawings, which, by way of illustration, show specific details and aspects of this disclosure in which the invention may be practiced. Other aspects may be utilized, and structural, logical, and electrical changes may be made without departing from the scope of the invention. The various aspects of this disclosure are not necessarily mutually exclusive, as some aspects of this disclosure may be combined with one or more other aspects of this disclosure to form new aspects.
[0035] Various examples are described in more detail below.
[0036] Fig. 1 shows a vehicle 101.
[0037] In the example of Fig. 1, a vehicle 101, for example a car or truck, is provided with a vehicle control device 102.
[0038] The vehicle control device 102 has data processing components, e.g., a processor (e.g., a CPU (central processing unit)) 103 and a memory 104 for storing control software according to which the vehicle control device 102 operates and data processed by the processor 103.
[0039] For example, the stored control software (computer program) includes instructions that, when executed by the processor, cause the processor 103 to implement one or more neural networks 107.
[0040] The data stored in memory 104 may include, for example, image data captured by one or more cameras 105. The one or more cameras 105 may, for example, capture one or more grayscale or color photos of the surroundings of the vehicle 101.
[0041] The vehicle control device 102 can detect objects in the surroundings of the vehicle 101, in particular other vehicles, using the image data (or data from other information sources, such as other types of sensors or vehicle-to-vehicle communication).
[0042] The vehicle control device 102 can examine the sensor data and control the vehicle 101 according to the results, ie, determine control actions for the vehicle and signal them to respective actuators of the vehicle. For example, the vehicle control device 102 can control an actuator 106 (e.g., a brake) to control the speed of the vehicle, e.g., to decelerate the vehicle.
[0043] The control strategy used by vehicle control device 102 must be extensively tested before its use in real road traffic. This is typically done with simulations that simulate other vehicles. For this purpose, driver models can be learned (i.e., trained) according to which the other vehicles in the simulation behave.
[0044] Learning driver models for automated driving simulation promises better scaling to the enormous number of relevant traffic scenarios than the manual development of heuristic driver models. The dominant learning approach for learning driver models is imitation learning from human demonstrations. Two approaches for this are GAIL (Generative Adversarial Imitation Learning), which is based on GANs (Generative Adversarial Networks), and the "TrafficSim" approach, which is based on a VAE (Variational Autoencoder).
[0045] The basic idea of imitation learning is that recordings (demonstrations) of a (demonstrating) agent's behavior are determined (i.e., acquired), and then a behavioral model for an agent (e.g., for another vehicle) is trained using this data so that it behaves similarly to the demonstrated behavior. In behavior cloning, imitation learning is treated as a simple, superimposed learning problem in which the objective is to learn the mapping from observation to action, i.e., the control strategy, where the observation-action pairs of the demonstrating agent are treated as independent, uniformly distributed samples.
[0046] When using such approaches to test a control system (e.g., an autonomous vehicle), the behavior of another vehicle is typically generated by a correspondingly trained model or sampled from a probabilistic behavior provided by the trained model. The other vehicle thus usually exhibits a "typical" behavior.
[0047] If other vehicles behave typically and the control system being tested copes well with it, i.e., interacts correctly with the other vehicles, for example, in such a way that no accident occurs, this does not mean that this will also be the case if another vehicle behaves particularly poorly in a particular traffic situation. In a real traffic situation, for example, another driver might drive particularly dangerously. Such worst-case behavior could therefore be within the realm of possibility, even if it is not "typical."
[0048] In order to specifically test such borderline cases of possible behavior, according to various embodiments, a type of imitation learning is used to generate a behavior of another vehicle (generally an agent), in which the respective model (e.g. a neural network) is trained with the aim of explicitly and correctly predicting (according to the respective training data) one or more quantiles of the action distribution of the agent whose behavior is to be modeled (or ultimately simulated) for a given situation (e.g. traffic situation).
[0049] Behavioral cloning can be viewed as a regression problem. The above training to predict a quantile can accordingly be viewed as a specific regression method, called quantile regression.
[0050] It should be noted that models trained using generative imitation learning methods such as GAIL or TrafficSim predict the distribution across actions (according to the respective training data). Knowledge of a distribution, in turn, implies knowledge of the quantiles. However, in quantile imitation learning according to various embodiments, a model is trained to directly and explicitly predict a quantile.
[0051] Furthermore, methods like GAIL are very complex, heuristic and difficult to understand: • It is difficult to understand whether they estimate the quantiles correctly or not and they can be very biased in their assumptions about the action distribution, e.g., assuming that it is Gaussian. • Furthermore, there is the problem of calculating the quantiles from a model according to GAIL or TrafficSim, even if the quantiles are uniquely determined by knowledge of the respective generative distribution (ie by knowledge of the GAN or the VAE).
[0052] Fig. 2 illustrates an example of a simulation for testing a control of a vehicle 201 (referred to as an ego vehicle or generally as an ego agent).
[0053] The ego agent 201 is driving with other vehicles (i.e., agents) 202, 203 from left to right on a three-lane highway. For the ego agent 201 and the other agents 202, the solid line shows the current route (i.e., the previous trajectory), and a dashed line shows the planned trajectory for the ego agent 201 and possible routes (trajectories) for the other agents 202, 203.
[0054] These possible travel paths for the other agents 202, 203 are modeled according to various embodiments using quantile imitation learning as mentioned above.
[0055] For the other agent 202, which is located below the ego-agent 201 in the current traffic situation, two possible upwardly curved future trajectories are shown. The lower of these two could, for example, be given by the 99.999% quantile of a distribution for the vehicle's behavior and does not collide with the future trajectory of the ego-vehicle 201. The uppermost future trajectory of this other agent 201 is given, for example, by the 99.9999% quantile and would lead to a collision with the ego-vehicle 201, which is moving along the shown future trajectory. Thus, if the collision probability is to be kept below 0.001%, the shown future path of the ego-agent 201 can be taken (for example, by the vehicle control device 102).However, if there is to be 99.9999% certainty that there will be no collision, then the vehicle control device 102 must not take this route. The quantile can refer to the quantile of the one-dimensional lateral acceleration action of the other agent 202.
[0056] The q-quantile (q is in [0, 1] and is called the undershoot fraction) for a one-dimensional, i.e., real-valued, random variable X is the point c on the real axis at which X lies to the left of c (i.e., below c) with probability q. There are slightly different definitions of quantiles, but this one will be used below. The median, for example, is the 0.5-quantile. In the example of Fig. 2 the value of the random variable X is the lateral acceleration, but it could also be, for example, the distance a driver maintains or the reaction time a driver needs (e.g. to brake).
[0057] Quantile regression is similar to conventional machine learning regression, but instead of the mean, the q-quantile is predicted for a given q. To train a model accordingly, the so-called "tilted absolute error loss function" can be used. This is a version of the absolute error function that has a general q-quantile as the optimum (instead of the median, i.e., the 0.5-quantile, as results from the non-tilted absolute error function).
[0058] The tilted absolute error loss function for training a machine learning model to predict the q-quantile is given by Lq(x,p)={q(x−p)if x≥p(1−q)(x−p)otherwise where x is an observation from a demonstration (e.g., an acceleration) and p is the prediction (prediction for X) of the machine learning model (for a particular traffic situation as training input data element). This loss can be summed over many observations, thus calculating a total loss, which the machine learning model is trained to reduce (e.g., adjusting the weights of a neural network so that the total loss is reduced). For example, a total loss can be determined for each batch of multiple batches of training data, and the machine learning model can be trained over the batches.
[0059] According to various embodiments, for testing a controller (e.g., a control algorithm or control software) in a simulation, one or more other agents are controlled according to one or more models trained using a behavior clone approach. However, such a model is trained to predict the q-quantile of the action distribution (i.e., distribution of actions) of an agent based on a current observation of the agent (instead of, e.g., the most probable action or the averaged action). This means that, essentially, a quantile regression is performed to model the behavior (or control strategy) of the other agents.
[0060] The other agents can be vehicles themselves or robotic devices in general, but they can also be humans and animals, for example. For example, the behavior of factory workers could be modeled (and simulated) to test a control system for a factory robot.
[0061] According to various embodiments, the following components or process steps are provided (given the undershoot proportion q as a parameter): - A control strategy for another agent ("imitator"), represented, for example, by a parameterized deep neural network. This control strategy has: ◯ The agent's current observation ("o") is used as input. This can be, for example, a simple state variable such as position or velocity, or a more complex image-based representation that includes other agents, such as vehicles, in the respective scene (e.g., traffic situation), which is then fed into a convolutional neural network (CNN). ◯ As output: a prediction for the q-quantile of the agent's action ("a") under the current observation. In the simple case where the agent has only a one-dimensional action space (e.g., lateral acceleration), two quantiles are predicted: the q-quantile c oben , so that the action with probability q under c oben lie g t, and analogously c unten , so that the action with probability q over c untenMore complicated approaches can also be used to predict a subset of the action space such that the action lies in this subset with high probability (given by q). In the following, this set is also referred to as the "quantile set." ▪ For the driver modeling use case, the action could be a two-dimensional action consisting of acceleration and steering angle. ▪ It should be noted that the definition of quantile given above only considers the one-dimensional case. For n-dimensional actions for n > 1, the dimensions can be treated separately, thus reducing this case to n one-dimensional quantile problems, or an n-dimensional quantile set can be defined in a suitable manner. ▪ Optionally, a quantile set can be defined based on conditional quantiles (analogous to Reference 1). For example, a q-quantile can be defined for a random variable X in the case that the random variable Y has a certain value. This can, for example, be applied to the individual dimensions of the action a (in the case that the action a has more than one dimension, ie, n > 1): For example, one can determine (predict) the quantiles (upper and lower) for the second dimension for each possible value of the first action dimension (ie condition on the first action dimension in addition to the condition on the observation o). - A machine learning model (or the control strategy it represents) that is trained to predict the q-quantile using the absolute error loss function tilted (corresponding to the undershoot fraction q), as explained above. As explained above, the parameters of the control strategy (e.g., weights of the neural network) are adjusted for training such that it behaves according to the quantiles with respect to the given demonstration data (i.e., similar to behavior cloning, only for the quantiles of the distribution in the training data). The demonstrator agent that provides the demonstrations (i.e., the demonstrations) is, for example, a human driver. ◯ As training data, i.e. demonstration data, i.e. recordings of the behavior of the demonstrator agent (e.g. human driver), the highD dataset (or other data, e.g. data recorded by human-controlled vehicles equipped with environmental sensors) can be used. ◯ The training data contains, for example, a set of trajectories (time series) of state-action pairs or, more generally, observation-action pairs; i.e., each such trajectory has the form (o 1 , a 1 ), (o 2 , a 2 ), ..., (o T , a T ), where each observation o t can also contain information about the environment, other agents in the scene, etc. - Application of the trained control strategy in a simulation to control an agent (from the perspective of an ego agent, a "different" agent) in the environment of the ego agent whose control is being tested: There are several ways to determine the current agent action from the current state o t in time step t for the (“other”) agent: ◯ Sampling methods: q-quantiles actually only specify a set (in the one-dimensional case, intervals) in which an action lies with a certain probability. In the simple one-dimensional case, the action lies with probability q in the interval (-infinity, c obere); these sets are not distributions. However, several ways are possible to convert such sets into distributions: the simplest would be to use a uniform distribution over the set in question. A more sophisticated approach is to learn quantiles and a distribution simultaneously, as in Reference 1, and to sample according to the distribution (with a focus on correct quantiles, especially on high (i.e., extreme or adversarial) quantiles) and potentially a focus on the support of the distribution). ◯ Worst-case cases or borderline cases: As mentioned above, a model trained for quantile regression can also be used to generate a form of worst-case behavior for the agent that lies within the quantile set. This allows the ego agent's control to be tested for particularly unfavorable behavior. This is used, for example, for testing the ego agent's control for at least some of the other agents for at least some of the time steps in the simulation. ▪ This worst-case behavior can be “a priori adversarial”, ie it is determined a priori which would be the most extreme or most dangerous behavior within the quantile set, e.g. an extreme braking or starting behavior. ▪ Alternatively, it can be the worst-case behavior with respect to a given control (in particular, the one being evaluated): In this case, the action of the other agent within the quantile set can be chosen so that it comes as close as possible to a collision with the ego (and, if possible, actually causes a collision). If no action from the quantile set results in a collision, then under certain assumptions it is guaranteed that there will be no collision with probability q (up to the uncertainty of the estimate).
[0062] Beyond the approach described above, it is also conceivable to predict not only quantiles, but also other relevant properties of the action distribution. In particular, it might be useful to predict the support of the action distribution (i.e., where it has any mass / probability at all; in a sense, this can be viewed as the 100% quantile). If the most unfavorable actions within the support for a (different) agent are taken and these do not lead to an accident, then it is 100% certain that no accident can occur with the control system being evaluated in the respective traffic situation (except for the uncertainty of the estimate).
[0063] Instead of a simple imitation learning approach similar to behavioral cloning (which, as described above, simply considers 1-time-step state-action pairs as independent uniformly distributed samples), more sophisticated imitation learning approaches such as GAIL can be used as a basis for imitation learning to train a quantile regression model.
[0064] Not only can the extreme corner cases (quantiles) of individual actions be predicted and used in the simulation, but the behavior used for a particular agent in the simulation can also be a higher-level extreme case (such as an extreme maneuver trajectory consisting of a sequence of actions). For example, the most extreme merging maneuver that would be expected from a human driver in the respective traffic situation can be used in the simulation for an agent. This can be derived from individual action quantiles, or a model can be trained to predict such higher-level behavior (e.g., an entire trajectory or a set of such trajectories).
[0065] One or more models can be trained to predict quantiles for multiple undershoot proportions and used for prediction (e.g., a model is trained to simultaneously output q-quantiles for several different q). In this context, quantiles can be predicted simultaneously with the full action distribution (i.e., one or more models are trained accordingly). For this purpose, for example, a normalization flow can be trained, as described in Reference 1, which is then used for predictions and action selection for one or more (different) agendas in a simulation.
[0066] In the above examples, a driver model was always trained to test a vehicle controller. However, a similar approach can also be used to test the controllers of other robotic devices. For example, the movement of a robot arm interacting with other robot arms can be demonstrated, a model for the robot arm can be trained, and then a robot controller can be tested in a simulation to determine whether it interacts correctly (or without accidents or damage to processed objects) with one or more robot arms controlled according to the model. Accordingly, the driver model is generally a behavior model for a robotic device, i.e., a robotic device behavior model. The term “robot device” can be understood as referring to any technical system (with a mechanical part whose movement is controlled), such as:a computer-controlled machine, a vehicle, a household appliance, a power tool, a manufacturing machine, a personal assistant, or an access control system. According to one embodiment, a control rule for the technical system is tested using the trained (robot device) behavior model, and the technical system is then controlled (or operated) accordingly (if the test is successful).
[0067] In summary, according to various embodiments, a method is provided as described in Fig. 3 shown.
[0068] Fig. 3 shows a flowchart 300 illustrating a method for evaluating a controller of a robotic device according to an embodiment.
[0069] In 301, demonstrations of the behavior of at least one object in control situations including the at least one object and the robot device are determined.
[0070] In 302, a machine learning model for mapping information about control situations to information about the behavior of the at least one object (i.e., actions of the at least one object or actions according to which the at least one object is controlled) is trained by means of a quantile regression loss (which may in particular be a tilted absolute error loss function) for one or more undershoot proportions.
[0071] In 303, a simulation is performed, wherein the behavior of the robot device is simulated according to the control to be evaluated and the behavior of the at least one object in at least one control situation occurring in the simulation is simulated according to an output that the trained machine learning model outputs in response to information about the occurring control situation.
[0072] In 304, the control is evaluated depending on events in the simulation.
[0073] The control system being evaluated can be a complete control method, i.e., a control strategy. Accordingly, the simulation can be a comprehensive simulation over a longer period of time (e.g., several seconds), during which multiple control actions may be selected according to the control strategy. This can be done as part of testing a control method, with the result of the testing then being used, for example, to decide whether the control method should be implemented.
[0074] Alternatively, the evaluation can also take place during operation of a control method, i.e. a control method has, for example, several possible controls available and for each one it is evaluated whether it would lead to a collision with the at least one other object (or another undesirable event). In this case, the simulation is therefore the determination of whether a collision (or other undesirable behavior) occurs when controlling the robot device according to the control to be evaluated and the behavior of the at least one object according to the output of the trained machine learning model. The output of the trained machine learning model can therefore be used here as a prediction for a possible behavior of the at least one object. Several control steps (i.e., for example, each selection of a control action) can be called by a control device, and the control action can thus trigger one or more control actions (e.g.B: a trajectory) in such a way that a collision is unlikely to occur. Since the sequence of events is extrapolated from the respective behavior, this can also be considered a simulation.
[0075] The at least one object is, for example, another robot device (possibly of the same type as the robot device whose control is to be tested, e.g., as in the above examples, one or more other vehicles), but it can also be, for example, a human (e.g., a pedestrian or factory worker) or an animal (e.g., a dog in traffic).
[0076] The procedure of Fig.3 may be performed by one or more computers having one or more data processing units. The term “data processing unit” may be understood as any type of entity that enables the processing of data or signals. The data or signals may, for example, be handled according to at least one (i.e., one or more than one) specific function performed by the data processing unit. A data processing unit may comprise or be formed from an analog circuit, a digital circuit, a logic circuit, a microprocessor, a microcontroller, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable gate array (FPGA) integrated circuit, or any combination thereof.Any other way of implementing the respective functions described in more detail herein may also be understood as a data processing unit or logic circuit arrangement. One or more of the method steps described in detail herein may be performed (e.g., implemented) by a data processing unit through one or more specific functions performed by the data processing unit.
[0077] According to various embodiments, the method is therefore particularly computer-implemented.
[0078] Various embodiments may receive and use sensor signals from various sensors such as video, radar, LiDAR, ultrasound, motion, thermal imaging, etc., for example, to obtain sensor data regarding trajectories, particularly demonstrations.
[0079] The trained model can be used to generate a control signal for a vehicle, particularly in a simulation for testing a vehicle control system, i.e., for controlling a technical system. The driver model can be trained using reinforcement learning and then one or more (simulated) vehicles can be controlled accordingly.
Claims
[1] A method (300) for evaluating a control of a robot device, comprising: Determining (301) demonstrations of the behavior of at least one object in control situations including the at least one object and the robot device; Training (302) a machine learning model for mapping information about control situations to information about the behavior of the at least one object by means of a quantile regression loss for one or more undershoot proportions; Carrying out (303) a simulation, wherein the behavior of the robot device is simulated according to the control to be evaluated and the behavior of the at least one object in at least one control situation occurring in the simulation is simulated according to an output that the trained machine learning model outputs in response to information about the occurring control situation; and Evaluate (304) the control depending on events in the simulation. [2] The method (300) of claim 1, wherein the machine learning model is trained to output a quantile for an action of the at least one object in the control situation for a predetermined undershoot proportion in response to information about a control situation for the at least one object, and wherein the behavior of the at least one object in the at least one control situation occurring in the simulation is simulated according to the quantile that the machine learning model outputs for the occurring control situation. [3] The method (300) of claim 1, wherein the machine learning model is trained to output, in response to information about a control situation for the at least one object, a plurality of quantiles for an action of the at least one object in the control situation for predetermined undershoot proportions, and a plurality of simulations are carried out, wherein the behavior of the at least one object in the occurring control situation is simulated in each of the plurality of simulations according to a respective value from a quantile set defined by the quantiles that the machine learning model outputs for the occurring control situation. [4] The method (300) of claim 1, wherein the machine learning model is trained to output, in response to information about a control situation for the at least one object, a plurality of quantiles for an action of the at least one object in the control situation for predetermined undershoot proportions, wherein the behavior of the at least one object in the occurring control situation is simulated according to a worst-case behavior from a quantile set defined by the quantiles that the machine learning model outputs for the occurring control situation. [5] The method (300) of claim 1, wherein the machine learning model specifies a normalization flow and the behavior of the at least one object in the at least one occurring control situation is simulated according to at least one sample from a probability distribution that the normalization flow outputs in response to the input of information about the occurring control situation. [6] Method (300) according to one of claims 1 to 5, wherein the quantile regression loss is calculated using a tilted absolute error loss function. [7] Test device arranged to carry out the method according to one of claims 1 to 6. [8] A computer program comprising instructions which, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 6. [9] A computer-readable medium storing instructions which, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and apparatus for testing a machine
DE102019220574A1
SYSTEMS AND METHODS FOR FEATURE ENGINEERING
DE102020126568A1
METHOD FOR CONTROLLING A ROBOT DEVICE AND ROBOT DEVICE CONTROL
DE102020209685A1
Method for controlling a robot device
DE102021204961A1
LEARNING DEVICE, LEARNING PROCEDURE, LEARNING DATA GENERATION DEVICE, LEARNING DATA GENERATION PROCEDURE, INFERENCE DEVICE AND INFERENCE PROCEDURE
DE112019007601T5