Method and device for the optimal parameterization of a vehicle dynamics control system
A model-based approach with a trainable neuronal network optimizes driving dynamic controllers, enhancing tire control and reducing braking distance through automated parameterization, addressing inefficiencies in existing methods.
Patent Information
- Application Number
- DE102021206880
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-30
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2041-06-30
AI Technical Summary
Existing driving dynamic controllers require complex manual parameterization that does not always lead to optimal settings, and existing automated methods lack efficiency and scalability.
A model-based approach using a trainable model, such as a neuronal network, to optimize driving dynamic controller parameters by predicting vehicle states and actions, combined with a correction model to refine predictions, allowing for automated and efficient parameterization.
The solution achieves significantly better performance than manual settings, enabling more optimal tire control, reducing braking distance, and ensuring robustness across various driving conditions without requiring extensive real-world maneuvers.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for optimally parameterizing a vehicle dynamics controller, a training device, a computer program and a machine-readable storage medium. State of the art
[0002] Vehicle dynamics controllers are generally known from the state of the art. The term vehicle dynamics controller (ESC, English: Electronic Stability Control) refers to a controller for any vehicle which, in certain situations, e.g., when tire grip on the road surface is no longer optimal, intervenes in the vehicle's driving operation to, for example, restore optimal tire grip. These specific situations can be identified, for example, by an anomaly occurring during continuous monitoring of tire grip or other vehicle conditions.
[0003] For example, the vehicle dynamics controller can counteract a vehicle skidding by selectively braking individual wheels. This can be used, for instance, to prevent the vehicle from spinning out at the limit in corners, both in cases of oversteer and understeer, thus ensuring the driver maintains control. Another application of vehicle dynamics controllers is, for example, providing optimal brake pressure during emergency braking to prevent wheel lock-up and minimize the braking distance.
[0004] When adapting the braking system to a specific vehicle type, a large number of parameters must be set according to the vehicle type during the application process. This is therefore complex and does not always lead to optimal settings.
[0005] One object of the present invention is to provide an efficient and automated method for the optimal parameterization of a vehicle dynamics controller.
[0006] DE 100 03 739 A1 discloses a method for identifying system parameters in vehicles by measuring measured values representing vehicle state variables during the vehicle's operation and evaluating them in a computing unit according to a calculation rule for determining the system parameters, taking into account equations of motion of a vehicle calculation model.
[0007] DE 10 2006 054 425 A1 discloses a method for determining a value of a model parameter of a vehicle reference model, with which a reference value of a first driving state variable can be determined.
[0008] DE 10 2016 214 064 A1 discloses a method for determining driving condition variables of a motor vehicle.
[0009] DE 10 2019 127 906 A1 discloses a method for determining an estimated value of at least one vehicle parameter. Advantages of the invention
[0010] The invention with the features of independent claim 1 has the advantage that parameterizations can be found which achieve significantly more optimized control, particularly of tire grip, compared to previous parameterizations. For example, it has been shown that this can further reduce the braking distance during emergency braking compared to current emergency braking systems. This increases the safety of the vehicle occupants.
[0011] Furthermore, the invention has the advantage that the optimal parameterization can be found automatically, thus eliminating the need for time-consuming manual testing and evaluation.
[0012] Furthermore, the invention has the advantage that the parameterization can be learned exclusively through vehicle simulations. This is particularly advantageous because simulations allow the limits of vehicle dynamics to be pushed further, ultimately enabling more parameterizations to be tested more cost-effectively. It can therefore be said that this results in a particularly efficient and effective application.
[0013] Furthermore, the invention has the advantage that minimal intervention by the application engineer can be carried out through domain knowledge (such as high-level decisions: more comfort vs. more performance), and thus the parameterization can be specifically tailored to customer needs.
[0014] Furthermore, the invention has the advantage that higher robustness, i.e., good performance in all operating points and thus not only optimized for individual test scenarios, can be achieved. Disclosure of the invention
[0015] In a first aspect, the invention relates to a method, particularly a computer-implemented one, for the optimal parameterization of a vehicle's driving dynamics controller. The driving dynamics controller can intervene in or control the vehicle's driving dynamics, whereby the driving dynamics controller depends on a, in particular current, vehicle state (s t ) an action (a t ) determined in order to, for example, positively intervene in the driving dynamics. Intervention in the driving dynamics is therefore made to keep the vehicle stable on its original driving trajectory.
[0016] Vehicle dynamics encompasses the movements of a vehicle, including its path, speed, acceleration, and the forces and moments acting on it in and around the three directions of motion. Examples of vehicle movements include straight-line and cornering, vertical, pitch, and roll movements, as well as driving at a constant speed, braking, and acceleration. Furthermore, the vibrations of the vehicle that arise during these movements can also be considered part of vehicle dynamics.
[0017] Under vehicle condition (s tThe term "vehicle state" can be understood as a quantity that characterizes a vehicle's state with respect to its driving dynamics and / or the state of a vehicle component. Preferably, the vehicle state characterizes a part of the current vehicle dynamics, in particular a spatial motion of the vehicle. Spatial motion can encompass the three translational movements along the principal axes: longitudinal movement along a longitudinal axis (the actual change in position), lateral movement along a transverse axis, and vertical movement along a vertical axis, usually combined with longitudinal movement when driving downhill or uphill. Spatial movements can also be expressed as accelerations. Rotational movements around the three principal axes are also conceivable, e.g., yaw around the vertical axis, pitch around the transverse axis, and roll around the longitudinal axis.These movements can be expressed as angles or angular velocities. Furthermore, translational and rotational vibrations can also influence the vehicle's condition. Furthermore, the vehicle condition preferably characterizes a steering wheel angle or steering wheel torque. The vehicle condition also particularly preferably characterizes tire grip and / or tire behavior.
[0018] Under the action (a t ) can be understood as a quantity that characterizes a movement of the vehicle, i.e., when the action is performed by the vehicle, the vehicle performs this movement. Preferably, the action (a t ) a control variable, such as a braking force or even a braking pressure.
[0019] The procedure comprises the following steps: The process begins with the provision of a model P to predict a vehicle state (s t+1The Model P is set up for this, depending on the vehicle's condition (see below). t ) and the action (a t ) a subsequent vehicle condition (s t+1 ) to predict. Under the following vehicle condition (s t+1 ) can be understood as the vehicle state that occurs immediately when the action (a t ) from the current vehicle state (s t ) was carried out by the vehicle.
[0020] This is followed by the determination of at least one data tuple (s0, ..., s). t , ... s T ; a0, ..., a t , ... a T ) comprising a sequence of vehicle states (s0, ..., s t , ... s T ) and their respective assigned actions (a0, ..., a t , ... a T ), wherein the vehicle states are determined by the model (P) depending on a determined action by the vehicle dynamics controller.
[0021] This is followed by an adjustment of the parameters (θ) of the vehicle dynamics controller such that a cost function (c), which determines the cost of the recorded trajectory depending on the vehicle states of the data tuple and the actions determined for each associated vehicle state, is minimized. This cost function is dependent on the parameters of the vehicle dynamics controller. The adjustment of the parameters (θ) of the vehicle dynamics controller can be performed for each vehicle state of the data tuple or for an entire sequence from the data tuple.
[0022] The adjustment of the parameters (θ) can be carried out by an optimization algorithm, preferably by means of a gradient descent method, particularly preferably by means of back-prop-through time.
[0023] It was recognized that a so-called model-based approach to optimizing the vehicle dynamics controller can determine the most informative parameter adjustments θ due to the optimization using the relatively accurate model. Alternatively, there are so-called model-free approaches, but these are less effective because they do not provide sufficient information for parameter optimization to obtain a vehicle dynamics controller that optimally manages vehicle dynamics. Therefore, it can be said that only the proposed method can provide vehicle dynamics controllers that exhibit significantly better performance than manually configured controllers. Compared to other learning paradigms, this approach has the advantage of being scalable, meaning it can handle arbitrarily complex datasets, and it can also optimize high-dimensional vehicle dynamics controllers.
[0024] It is proposed that the model (P) is a trainable model whose parameterization has been trained based on recorded driving maneuvers of the vehicle or another vehicle, or that the model (P) is a physical model describing the vehicle's driving dynamics, particularly along a longitudinal and a lateral axis. The trainable model can be, for example, a machine learning system, preferably a neural network. In general, the trainable model can be a black-box model, such as linear or feature-based regression, Gaussian process models, recurrent neural networks (RNNs, LSTMs), or (deep) neural networks, or a white-box model, such as (simplified) physical models with parameters or combinations thereof (grey-box model).
[0025] The physical model has the advantage of enabling analytical gradients, which allows for more precise adjustments, resulting in vehicle dynamics controller parameterization that is closer to an ideal parameterization. Furthermore, no actual measurements are required, which is why the method can advantageously be performed purely through simulation.
[0026] Furthermore, it is proposed that the trajectory of a real vehicle maneuver be additionally captured. Depending on the captured trajectory and the model (P), a correction model (g) is created such that the correction model (g) corrects the outputs of model (P) so that they substantially correspond to the captured trajectory. The trajectory can describe a sequence of vehicle states and the action chosen in each vehicle state for the real driving maneuver. "Substantially" here means that the accuracy achieved by this correction lies within the measurement tolerances for the vehicle states being corrected, or within the accuracy achievable with the respective optimization methods for creating the correction model, or within the maximum achievable accuracy of the correction model through a given power of the correction model.
[0027] It is conceivable that this additional step of capturing the actual driving maneuver could be performed again after the step of adjusting the parameters of the vehicle dynamics controller, with the actual driving maneuver now being carried out using the adjusted vehicle dynamics controller. The trajectory thus captured can then be used again to fine-tune the correction model and also to readjust the subsequent steps of determining at least one data tuple using model P, and subsequently to adjust the parameters of the vehicle dynamics controller.
[0028] In particular, the correction model is created by minimizing the difference between the output of the correction model and the difference between the recorded state of the training data and the state predicted by the model.
[0029] If the model (P) is a learned model, its parameterization can be learned depending on the recorded trajectory. For this purpose, for example, optimization of the model parameters by calculating the maximum likelihood or maximum a posteriori solution for single- or multi-step model predictions in an open (feedforward) or closed (feedback) control loop is suitable, e.g., using (stochastic) gradient descent.
[0030] Surprisingly, the use of the correction model offers the advantage that very few real-world maneuvers of the vehicle are needed, both to optimize the vehicle dynamics controller and to create the learned model.
[0031] Furthermore, it is proposed that the model (P) is deterministic and the correction model is time-dependent. In other words, the correction model depends on the state or on time, where time characterizes the interval of time elapsed since the beginning of the recorded trajectory. That is, the correction model determines the correction value for model P as a function of time. Surprisingly, this type of correction model has been found to produce the best parameterizations. Time can also be a discrete value that characterizes the number of actions performed since a predefined starting point (e.g., the time at which the vehicle dynamics controller intervenes in the vehicle dynamics).
[0032] The combination of model P and the correction model, which is configured to correct the model's outputs, can be viewed as a global model for predicting state changes. In other words, the global state change model is a superposition of these two models.
[0033] The correction model is configured to correct errors in the first model regarding the true state of the environment after an action has been performed. For example, the model predicts a state based on a current state and an action. It should be noted that the action can be determined by the vehicle dynamics controller or, for example, by a driver. The correction model then adjusts the predicted state so that it is as close as possible to the actual state of the environment after the agent has performed this action for the current state. In other words, the correction model modifies the output of the first model to obtain a predicted state that is as close as possible to the state the environment would actually assume, or the state captured during the maneuvers.Therefore, the correction model corrects the first model to obtain a more accurate state regarding the environment, especially the environmental dynamics.
[0034] Preferably, the correction model depends either on a time step and / or the current state. Alternatively, the correction model is a correction term, which is an extracted correction value determined by the difference between the recorded vehicle states of the recorded real-world maneuver and the states predicted by the model. The correction model can output discrete corrections that can be directly added to the model's prediction. A special case of the correction model exists in which the correction model outputs time-discrete correction values.
[0035] Furthermore, it is proposed that the correction model be selected by minimizing a measure of the difference between the output of the correction model and the difference between the observed vehicle state along the trajectory and the predicted vehicle states of the model. This minimization can also be performed using the well-known gradient descent method.
[0036] Furthermore, it is proposed that a plurality of different models (P) be provided, with the data tuple being randomly selected for one of the plurality of different models.
[0037] The advantage here is that, firstly, model uncertainties can be modeled, ultimately leading to robust controller behavior and thus enabling the controller to handle temporal changes more effectively. It is also conceivable that the model(s) could output an uncertainty value for their predictions, characterizing the uncertainty of those predictions. This uncertainty would then be taken into account when adjusting the parameters of the vehicle dynamics controller. The uncertainty can be determined as follows: 1) Statistical statements about the performance, robustness and safety of the controller, and / or 2) Avoidance of unsafe / unknown vehicle behavior during / after optimization, and / or 3) Acceleration of learning through additional exploration in previously unknown regions.
[0038] Propagating the (model) uncertainty over several time steps yields the uncertainty in the long-term predictions (a distribution over possible future system behavior). (Methods used here include, for example, closed-form analysis (for simple models, e.g., linear Gaussian), sampling, numerical integration, moment matching, and linearization (for more complex models).
[0039] Furthermore, it is proposed that the different models differ in that they each consider or characterize different dynamics of external variables or different dynamics of vehicle variables.
[0040] An example of a dynamic quantity is a changing road surface or different scenarios, such as a changing course of the road with respect to all three possible spatial axes of the road.
[0041] Furthermore, it is proposed that a separate data tuple be recorded for each model, with parameter changes depending on all data tuples. This approach has proven to result in particularly optimal and robust controller behavior.
[0042] Furthermore, it is proposed that the recorded vehicle states be filtered using a Kalman filter, whereby a parameterization of the Kalman filter is determined depending on a predicted trajectory of the vehicle, and the Kalman filter is applied to the recorded states.
[0043] Furthermore, it is proposed that the vehicle dynamics controller has a modular controller structure, whereby when parameters are adjusted, they are adapted in such a way that the changed parameters lie within predefined value ranges. This means the controller is divided into modules, and each module is responsible for a sub-function, e.g., a PID controller that adjusts the current slip to the target slip. Another module would then be a gain scheduler that adjusts the PID gains depending on the driving situation, road surface, and speed. Yet another module can estimate the driving situation based on wheel speeds, brake pressure profile, and vehicle speed, etc.
[0044] The advantage, besides optimizing parameters only within confidence ranges, is that the vehicle dynamics controller will not exhibit any safety-critical behavior and, furthermore, that exploration of vehicle dynamics is limited to meaningful vehicle states.
[0045] Furthermore, it is proposed that the vehicle dynamics controller is a neural network, in particular a radial basis function network.
[0046] A key advantage is that neural networks are very good at learning complex relationships and exhibit exceptional flexibility in learning a wide variety of controller behaviors. RBF networks are particularly favored because their compact design makes them especially suitable for implementation on a vehicle's electronic control unit (ECU).
[0047] Furthermore, it is proposed that after adjusting the parameters, a vehicle state is recorded during vehicle operation, whereby an actuator of the vehicle is controlled depending on the action that is initiated by the vehicle dynamics controller based on this recorded vehicle state.
[0048] Furthermore, it is proposed that the cost function is a weighted superposition of a plurality of functions, where the functions characterize a difference between a current slip of the vehicle's tires and a target slip, a distance traveled since the intervention of the vehicle dynamics controller, and time derivatives of the distance traveled.
[0049] Furthermore, it is proposed that the vehicle dynamics controller is an ABS controller and outputs an action that characterizes a braking force, with the physical model comprising a plurality of sub-models that are a physical model of a component of the vehicle.
[0050] The action can, for example, be determined separately for each of the vehicle's wheels or axles, so that the wheels / axles can be controlled individually depending on the respective brake pressure.
[0051] It should be noted that the vehicle dynamics controller can be, for example, an ABS, TCS or ESP controller, etc., or a combination of these controllers.
[0052] It should also be noted that the procedure can also be used to readjust an already optimized vehicle dynamics controller for a first vehicle type / instance for a second vehicle type / instance, or even to readjust it for the first vehicle type / instance, e.g. if the vehicle has been fitted with new tires.
[0053] In further aspects, the invention relates to a device and a computer program, each configured to perform the above methods, and a machine-readable storage medium on which this computer program is stored.
[0054] Embodiments of the invention are explained in more detail below with reference to the accompanying drawings. The drawings show: Fig. 1 schematically an embodiment of the control of a vehicle with a vehicle dynamics controller; Fig. 2 schematically a flowchart of a procedure for parameterizing the vehicle dynamics controller; Fig. 3. A possible design for a training device. Description of the exemplary implementations
[0055] Fig. Figure 1 shows an example of a vehicle 100 with a control system 40.
[0056] Vehicle 100 can generally be a motor vehicle controlled by a driver, or it can be a semi-autonomous or even a fully autonomous vehicle. In other embodiments, the motor vehicle can be a wheeled, tracked, or rail vehicle. It is also conceivable that the motor vehicle is a two-wheeler, such as a bicycle, motorcycle, etc.
[0057] The vehicle's condition is detected at preferably regular intervals using at least one sensor 30, which may also be determined by a plurality of sensors. The condition can also be determined based on the detected sensor values. The sensor 30 is preferably an acceleration sensor (in the longitudinal direction of the vehicle, but could also be a 3D sensor in all axes), a wheel speed sensor (on all wheels), or a yaw rate sensor about a vertical axis (but could also be about all other axes).
[0058] The control system 40 receives a sequence of sensor signals S from sensor 30 in an optional receiving unit, which converts the sequence of sensor signals S into a sequence of pre-processed sensor signals.
[0059] The sequence of sensor signals S or pre-processed sensor signals is fed to a vehicle dynamics controller 60 of the control system 40. The vehicle dynamics controller 60 is preferably parameterized by parameters θ, which are stored in and provided by a parameter memory P.
[0060] The vehicle dynamics controller 60 determines an action, hereinafter also referred to as control signal A, based on the sensor signals S and its parameters θ, which is transmitted to an actuator 10 of the vehicle. The actuator 10 receives the control signals A, is controlled accordingly, and then executes the corresponding action. It is also conceivable that the actuator 10 is configured to convert the control signal A into a direct control signal. For example, if the actuator 10 receives a braking force as a control signal A, the actuator can convert this into a corresponding brake pressure, which is then used to directly control the brakes. The actuator 10 can be a braking system, comprising the brakes of the vehicle 100. Additionally or alternatively, the actuator 10 can be a drive or steering system of the vehicle 100.
[0061] In further preferred embodiments, the control system 40 comprises one or more processors 45 and at least one machine-readable storage medium 46 on which instructions are stored which, when executed on the processors 45, cause the control system 40 to execute the method according to the invention.
[0062] In further embodiments, a display unit 10a is provided in addition to the actuator 10. The display unit 10a is intended, for example, to indicate an intervention of the vehicle dynamics controller 60 and / or to issue a warning that the vehicle dynamics controller 60 is about to intervene.
[0063] The vehicle dynamics controller 60 is defined by a parameterized function a = f (s, θ) which outputs the control signal A depending on the state s and / or the sensor signals S of the sensor 30. If the vehicle dynamics controller 60 outputs a control signal A for the actuator 10, where the actuator 10 has a plurality of actuators, the control signal A can represent a control signal for each of the actuators. The individual actuators can be the individual brakes of the vehicle 100.
[0064] In a preferred embodiment, the vehicle dynamics controller 60 is an ABS controller, wherein this controller outputs a braking force or a braking pressure as a control signal. Preferably, the vehicle dynamics controller 60 outputs a braking pressure or braking force for each of the brakes of the wheels or for each of the axles of the vehicle 100 in order to be able to control the wheels individually.
[0065] Preferably, the vehicle dynamics controller 60 has an interpretable controller structure. This can be achieved, for example, by defining valid parameter limits within the controller. This has the advantage that the behavior of the vehicle dynamics controller 60 is comprehensible in every situation.
[0066] Examples of the parameterized function f of the vehicle dynamics controller 60 are as follows: A vehicle dynamics controller 60 that has an interpretable controller structure can, for example, be a structured vehicle dynamics controller 60 which is structured like a decision tree.
[0067] To determine an action using the decision tree, the process starts at a root node and proceeds along the tree. At each node, an attribute is queried (e.g., a vehicle status), and a decision is made based on this information regarding the selection of the next node. This procedure continues until a leaf node of the decision tree is reached. The leaf node characterizes one action from a plurality of possible actions. For example, the leaf node might characterize the application / decrease of brake pressure. In this example, the parameters θ are decision thresholds or similar.
[0068] The vehicle dynamics controller 60 can alternatively be provided by an RBF network (radial basis functions) or by a Deep RNN policy. It should be noted that the parameterized function f can also be any other mathematical function that maps the state of the vehicle to a control signal depending on the parameters.
[0069] Fig. Figure 2 shows a schematic representation of a flowchart 20 for parameterizing the vehicle dynamics controller 60 and optionally for subsequently operating the vehicle dynamics controller 60 in the vehicle 100.
[0070] The procedure begins with step S21. In this step, driving data from vehicle 100 is collected. The driving data is, for example, a data series s0, s1, ..., s t , ...,s T , which describe a state s of vehicle 100 during a driving maneuver. This driving data is preferably a data tuple comprising the state data s t as well as action data a t at any point in time t of the maneuver.
[0071] In the event that the vehicle dynamics controller 60 is an ABS controller, a braking process can be recorded, e.g., using a known ABS controller or by having a driver operate the vehicle 100, whereby the state data (s t) The following sensor data may be included as examples: vehicle speed v veh , acceleration a veh , preferably the following sensor data per wheel of the vehicle 100: Wheel speed v wheel , acceleration a wheel , jerking (engl. jerk) j wheel The action data (a t ) are the braking forces selected in the respective state, preferably also a quantity that characterizes a road surface.
[0072] Alternatively, the driving data can be generated via simulation, in which a fictitious vehicle performs one or more (braking) maneuvers in a simulated environment.
[0073] Step S22 can then follow. Here, the recorded state data s are partially reconstructed. This is because not all the necessary information about the state of a system (car) is typically measured by internal sensors (e.g., vehicle tilt, suspension behavior, wheel acceleration). For learning and modeling, this latent information must be recovered to enable predictions and optimizations. This area is typically referred to as latent state inference (e.g., hidden Markov models) and is solved using filtering / smoothing algorithms (e.g., Kalman filters). Preferably, the reconstruction of the vehicle states in step S22 is performed using the Kalman filter.
[0074] After completion of step S21 or step S22, step S23 follows. In this step, a model P is provided. This provision can be done either by creating the model P based on the records from step S21 or by providing a physical model.
[0075] The model P(s t+1 |s t , a t ) is a model which depends on a vehicle's condition s t at a time t and a control signal selected depending on it, a subsequent vehicle state s t+1 predicts the immediately following time t + 1.
[0076] The preferred model is P(s) t+1 |s t , a t ) a first-order physical model. That is, the physical model comprises equations that describe physical relationships and depend on the current vehicle state s. t and the action a tthe following vehicle condition s t+1 In particular, predict deterministically. For example, the physical model for the vehicle dynamics controller 60 for ABS can be composed of one or more sub-models from the following list: a first sub-model, which is a physical model of a wheel of the vehicle 100; a second sub-model, which describes the center of gravity of the vehicle; a third sub-model, which is a physical model of the damper; a fourth sub-model, which is a physical model of the tire; and a fifth sub-model, which is a multi-dimensional model of a hydraulic system. It should be noted that the list is not exhaustive and further physical characteristics such as tire / brake temperature, etc., can be taken into account.
[0077] It should be noted that, in addition to model P, other approaches are conceivable for optimizing the parameterization. As an alternative to the model, a so-called model-free reinforcement learning approach or a value-based reinforcement learning approach can also be chosen. Accordingly, in step S23, for example, the Q-function for value-based reinforcement learning is created based on the recordings from step S21.
[0078] Step S24 can follow step S23. Step S24 can be described as an "on-policy correction." Here, a correction model g is created, which makes predictions of the model P(s). t+1 |s t , a t ) about vehicle states are corrected in such a way that the corrected predictions essentially match the predictions recorded in step S21.
[0079] Preferably, the corrected vehicle condition is corrected as follows: st+1′=P(st+1|st,at)+g(st,at)
[0080] The correction model g is created in such a way that it is optimized such that, given s t and a t outputs a single value that represents the error of the model P (s t+1 |s t , a t ) corresponds to the recorded vehicle conditions according to S21.
[0081] Furthermore, the correction model g has the advantage that it corrects a lack of agreement between model P and the actual behavior of the vehicle.
[0082] To correct the lack of agreement between model P and the actual vehicle behavior, the following measures can be taken, either alternatively or additionally. It is conceivable that a so-called transfer learning technique could be used, which incorporates previously determined vehicle states and thus allows for faster learning of the model for the specific vehicle instance. It is also conceivable that a multiple of different models could be used, enabling a more robust controller behavior to be learned through this ensemble.
[0083] Step S25 follows step S23 or step S24. In this step, a multiple rollouts are executed. That is, using the current parameterization θ. kThe vehicle dynamics controller 60 and using the model P, in particular additionally with the correction model g, is applied to a maneuver and the resulting trajectory, in particular determined sequences of vehicle states, is recorded.
[0084] It should be noted that, in addition to model P, other approaches are conceivable for optimizing the parameterization (model-free reinforcement learning approach or value-based reinforcement learning approach). Accordingly, the trajectory capture must be adapted in this rollout step.
[0085] After step S25 has been executed, or after step S25 has been executed several times, step S26 follows. In this step, an evaluation of the costs for the recorded trajectory(ies) from step S25 is performed.
[0086] The costs for the trajectory can be determined as follows. Preferably, costs are determined for each proposed action of the vehicle dynamics controller 60. For this purpose, a cost function c(s, a) can be used to calculate the costs depending on the previous trajectory or the current vehicle state s. t as well as the currently selected action a t Determine the total costs for a trajectory. These costs can then be accumulated over the entire maneuver, i.e., over all time points t: J(θ)=∑t=0Tc(st,f(st,θ))
[0087] The cost function c(s, a) can be composed as follows: c(s,a)=α1∗mean deceleration+α2∗steerability+α3∗ where α n Predefinable coefficients are those that are specified, for example, by an application engineer or set to initial values. These coefficients can take a value between 0 and 1.
[0088] Steerability can be understood as the controllability of the vehicle. This can be determined depending on a force acting laterally on the vehicle (F_lat) (e.g., F_lat, max - F_lat, current), possibly also depending on a normalized lateral force: (F_lat,max−F_lat,current) / F_lat,max,non_braking.
[0089] Controllability can also be defined negatively if the cost function is to be minimized. Additionally or alternatively, controllability can also be determined as a function of longitudinal forces, so that the longitudinal forces are not fully utilized in order to leave "room" for lateral forces. For this purpose, a target slip range can be defined (e.g., slip ∈ [slip_min, slip_max]) and then mapped to corresponding costs, for example, using a sigmoid function.
[0090] Mean deceleration can be understood as an average over all accelerations of the trajectory, e.g. 1n∑i ai ∀ i in ABSactive.
[0091] Further components of the cost function can be determined by any behavior of the vehicle that is to be penalized or rewarded. Examples include: comfort / jerk (i.e., how jerky the braking is), hardware requirements (how stressful the braking is on the braking system, vehicle, hydraulics, tires), performance (e.g., braking distance, acceleration), and straight-line stability (i.e., behavior around the vertical axis).
[0092] All these components can be evaluated based on different signals (sensor signals or estimates) and different cost functions, e.g., Mean Absolute Error, Mean Squared Error, Root Mean Squared Error between the current and target state (e.g., in slip, coefficient of friction, jerk), or the standard deviation of a signal.
[0093] Another component of the cost function can be the total braking distance. This can be a single value that is only available in the time step in which the braking is complete (e.g., v). veh < v treshold -> c, where c is the total braking distance, or a sum of v * dt for each time step in which the braking is active.
[0094] Another component of the cost function can be a deviation via a slip from a target slip: ||slip-lip_target||^2 and / or a mean acceleration: mean std(acceleration).
[0095] After the total costs for the trajectory or the majority of trajectories have been determined in step S26, step S27 follows. In this step, the parameters θ of the vehicle dynamics controller 60 are iteratively adjusted to reduce the total costs. An optimization can be defined as follows: θ*=argminθJ(θ)
[0096] This optimization via the parameters θ can be carried out using a gradient descent method over the total costs J.
[0097] For each iteration k of the optimization, the current parameters θ are then used. k as follows: θk+1=θk+λdJdθ where λ is a coefficient which takes on a value less than 1.
[0098] The iteration can continue until a termination criterion is met. Termination criteria could be, for example: a maximum number of iterations, a minimum change in J < J treshold, a minimal change in the parameters θ < θ threshold .
[0099] If multiple total costs, especially for different maneuvers, have been determined, the parameters can be adjusted batchwise across most of these total costs. This batch approach can be performed as a batch across model parameters, a batch across scenarios / maneuvers, or a batch across sub-trajectories.
[0100] After step S27 has been completed, steps S25 to S27 can be repeated, or alternatively, steps S21 to S27 can be repeated.
[0101] In the optional step S28, the control system 40 of the vehicle 100 is initialized with the adapted vehicle dynamics controller 60 from step S27.
[0102] In the subsequent optional step S29, the vehicle 100 is operated with the adapted vehicle dynamics controller 60. The vehicle 100 can be controlled by this vehicle dynamics controller 60 when it is activated in a corresponding situation, for example, when emergency braking is performed.
[0103] In a further embodiment of the method according to Fig. 2. After step S27 has been completed, steps S21 to S27 can also be repeated; however, the adapted vehicle dynamics controller 30 is transferred to another vehicle of a different vehicle model or type after step S27, and new measurements are taken after step S21. This makes it possible to optimize an optimized vehicle dynamics controller 60 for another vehicle with minimal effort.
[0104] In a further embodiment of the method according to Fig. In step 2, a plurality of different models P are provided or generated, and trajectories are created for each of the models. These models differ in that they describe different scenarios and / or consider different dynamics of external variables, such as road surface or other driving characteristics of different vehicle model types, or wear of tires / other vehicle components. This has the advantage that the vehicle dynamics controller learns to handle changes over time.
[0105] Fig.Figure 3 schematically shows a training device 300 for parameterizing the vehicle dynamics controller 60. The training device 300 comprises a provider 31, which either provides the recorded driving data from step S21 or is a simulated environment that generates a state according to the actions performed by the vehicle dynamics controller 60. The data from the provider 31 is forwarded to the vehicle dynamics controller 60, which determines the respective action from this data. The data from the provider 31 and the actions are fed to an evaluator 33, which determines parameters adapted according to step S27. These parameters are transmitted to the parameter memory P and replace the current parameters there.
[0106] The steps performed by the training device 300 can be implemented as a computer program, stored on a machine-readable storage medium 34, and executed by a processor 35.
[0107] The term "computer" encompasses any device for executing predefined computational instructions. These instructions can be in the form of software, hardware, or a hybrid of both.
Claims
[1] Method (20) for parameterizing a driving dynamics controller (60) of a vehicle 100, which can intervene in a controlling manner in a driving dynamics of the vehicle (100), wherein the driving dynamics controller (60) is configured to control a driving dynamics controller (60) depending on a vehicle state (s t ) an action (a t ), comprising the following steps: Providing (S23) a model P for predicting a vehicle state (s t+1 ), which is set up depending on the vehicle condition (s t ) and the action (a t ) a subsequent vehicle condition (s t+1 ) predict; Determine at least one data tuple (s0, ..., s t , ... s T ; a0, ..., a t , ... a T ) comprising a sequence of vehicle states (s0, ...,s t , ... s T ) and the associated actions (a0, ..., a t , ... a T), wherein the vehicle states are determined by means of the model (P) depending on a determined action by the vehicle dynamics controller; Adapting (S27) the parameters (θ) of the vehicle dynamics controller (60) such that a cost function (c) which determines costs of the recorded data tuple depending on the vehicle states of the data tuple and the determined actions of the respectively assigned vehicle states and wherein the cost function (c) is dependent on the parameters of the vehicle dynamics controller (60) is minimized. [2] Method according to claim 1, wherein the model (P) is a machine learning system whose parameterization has been learned as a function of detected driving maneuvers of the vehicle (100) or of another vehicle, or the model (P) is a physical model which describes driving dynamics of the vehicle, in particular along a longitudinal, lateral and horizontal axis of vehicles. [3] Method according to claim 1 or 2, wherein additionally a detection (S21) of a trajectory of a real driving maneuver of the vehicle (100) takes place, wherein depending on the detected trajectory and the model (P), a correction model g is created, so that the correction model corrects outputs of the model (P) such that they substantially correspond to the detected trajectory. [4] Method according to claim 3, wherein the model (P) is deterministic and the correction model is time-dependent. [5] Method according to one of the preceding claims, wherein a plurality of different models (P) are provided, wherein the data tuple is randomly acquired for one of the plurality of different models. [6] Method according to claim 5, wherein the different models differ in that they each describe different dynamics of external variables or different dynamics of variables of the vehicle. [7] Method according to claim 5 or 6, wherein a data tuple is acquired for each of the models, wherein the change of the parameters occurs depending on all data tuples. [8] Method according to one of the preceding claims, wherein the vehicle states are filtered by means of a Kalman filter (S22). [9] Method according to one of the preceding claims, wherein the vehicle dynamics controller has a modular controller structure, wherein when adjusting the parameters, these are adjusted in such a way that the changed parameters lie within predetermined value ranges. [10] Method according to one of the preceding claims 1 to 8, wherein the vehicle dynamics controller (60) is a neural network, in particular a network with radial basis functions (radial basis function network). [11] Method according to one of the preceding claims, wherein after the adaptation (S27) of the parameters, a vehicle state is detected during operation of the vehicle (100), wherein an actuator (10) of the vehicle is controlled depending on the action which is carried out by means of the vehicle dynamics controller (60) depending on this detected vehicle state. [12] Method according to one of the preceding claims 2 to 11, wherein the vehicle dynamics controller (60) is an ABS controller and outputs an action which characterizes a braking force, wherein the physical model comprises a plurality of sub-models, wherein the sub-models are each a physical model of a component of the vehicle. [13] Method according to one of the preceding claims, wherein the cost function is a weighted superposition of a plurality of functions, wherein the functions characterize a difference between a current slip of tires of the vehicle and a target slip, a distance traveled since intervention of the vehicle dynamics controller, and time derivatives of the distance traveled. [14] Vehicle dynamics controller (60) parameterized according to the method of claim 9. [15] Device which is arranged to carry out the method according to one of the preceding claims 1 to 13. [16] A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 13. [17] A machine-readable storage medium on which the computer program according to claim 16 is stored.
Citation Information
Patent Citations
System parameters identification in vehicle, involves obtaining movement equations with respect to measured speed and acceleration based on vehicle substitute model
DE10003739A1
Method for determination of value of model parameter of reference vehicle model, involves determination of statistical value of model parameter whereby artificial neural network is adapted with learning procedure
DE102006054425A1
determination of driving status variables
DE102016214064A1
Method and device for determining the value of a vehicle parameter
DE102019127906A1