Method and device for optimally parameterizing vehicle dynamic control system of vehicle
Patent Information
- Application Number
- JP2022098820
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-30
- Filing Date
- 2022-06-20
- Publication Date
- 2025-06-12
AI Technical Summary
Existing vehicle dynamic control systems require costly and inefficient manual parameter adjustments to optimize tire grip and braking performance, lacking an automated and efficient method for optimal parameterization.
A method utilizing a trainable model, preferably a neural network, to predict vehicle states and actions, adjusting parameters through gradient descent and a modified model to minimize errors, allowing for automated and optimized vehicle dynamic control.
Significantly reduces braking distance and enhances safety by providing a more optimized control of tire grip, eliminating the need for manual tests and ensuring robust performance across various operating conditions.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method, training apparatus, computer program, and machine-readable storage medium for optimally parameterizing a vehicle dynamics control system. [Background technology]
[0002] Vehicle dynamic control systems are generally known through the prior art. The concept of a vehicle dynamic control system (ESC, Electronic Stability Control) represents a controller for any vehicle that intervenes to adjust the vehicle's driving conditions under certain circumstances, for example, when the tire grip on the road surface is no longer optimal, in order to restore optimal tire grip. These predetermined circumstances can be identified, for example, by the occurrence of an anomaly when constantly monitoring tire grip or other vehicle conditions.
[0003] For example, a vehicle dynamic control system can prevent vehicle slippage by appropriately applying the brakes to individual tires, thereby preventing the vehicle from skidding within the limits of the curve, whether in the case of oversteer or understeer, during cornering, for instance. Another application of vehicle dynamic control is to provide optimal brake pressure during emergency braking to prevent wheel lock-up and minimize braking distance.
[0004] When adapting a brake system to a specific vehicle model, it is necessary to adjust multiple parameters to match the vehicle model during use. Therefore, this is costly, and optimal adjustments are not always achieved. [Overview of the project] [Problems that the invention aims to solve]
[0005] The object of the present invention is to provide an effective automated method for the optimal parameterization of a vehicle dynamic control device.
Advantages of the Invention
[0006] The present invention having the features of independent claim 1 has the advantage that a parameterization that is clearly more optimized compared to conventional parameterizations, especially for the control of tire grip, can be found. For example, it has been proven that by this means, the braking distance during an emergency brake can be made even more clearly shorter compared to the current emergency brake system. This enhances the safety of the vehicle occupants.
[0007] Furthermore, the present invention has the advantage that an optimal parameterization can be automatically found, thereby eliminating costly manual testing and evaluation.
[0008] Furthermore, the present invention has the advantage that the parameterization can be learned solely by simulating the vehicle. This is particularly advantageous because in simulation, the limits of driving dynamics can be further increased to the maximum, so that ultimately, on the one hand, more parameterizations can be attempted, and on the other hand, this can be done at a low cost. Therefore, it can be said that a particularly effective and efficient application is obtained thereby.
[0009] Furthermore, the present invention has the advantage that the minimum intervention of an application engineer is carried out by domain knowledge (e.g., by a high-level decision: higher comfort vs. higher performance), thereby enabling the parameterization to be accurately adapted to the customer's requirements.
[0010] Furthermore, the present invention has the advantage that a higher robustness, i.e., good performance, is obtained by optimizing not only for individual test scenarios but also at all operating points.
Means for Solving the Problem
[0011] In a first aspect, the present invention relates to a method, particularly computer-based, for optimally parameterizing a vehicle dynamics control system of a vehicle. The vehicle dynamics control system can adjust to or control the driving dynamics of a vehicle, in which case the vehicle dynamics control system can, for example, positively intervene in the driving dynamics, particularly in the actual vehicle conditions (s t ) and operation (a t It is calculated depending on the following: In other words, the vehicle dynamics control system intervenes in the driving dynamics to keep the vehicle stable in its initial trajectory.
[0012] Driving dynamics can be interpreted as the motion of a vehicle in and around three vehicle motion directions, namely the path, velocity, acceleration, and the forces and moments acting on the vehicle. Vehicle motion includes, for example, straight-line driving, cornering, vertical motion, pitching and rolling motion, as well as driving at a constant speed, braking and acceleration processes. Furthermore, in this case, the resulting vehicle vibrations can also be interpreted as driving dynamics.
[0013] Vehicle condition (s t) may be interpreted as a value characterizing the state of a vehicle with respect to the driving dynamics of the vehicle and / or the state of components of the vehicle. In a preferred form, the vehicle state characterizes a part of the actual vehicle dynamics, in particular the three-dimensional movement of the vehicle. As the three-dimensional movement, three translational movements in the main axis direction, that is, the longitudinal movement along the longitudinal axis, the self-location movement, the lateral movement along the lateral axis, and the stroke movement along the vertical axis can be considered, which is generally combined with the longitudinal movement when driving downhill or uphill. The three-dimensional movement may be generated as an acceleration. Rotational movements about the three main axes, for example, yawing about the vertical axis, pitching about the lateral axis, and rolling about the longitudinal axis are also conceivable. These movements can be generated as an angle or an angular velocity. Furthermore, translational vibrations and rotational vibrations may additionally be included in the vehicle state.
[0014] Furthermore, the vehicle state characterizes the steering angle or the steering torque in a preferred form. Particularly preferably, the vehicle state also characterizes the tire grip and / or the behavior of the tires.
[0015] Operation (a t ) may be interpreted as a value characterizing the movement of the vehicle, that is, when this operation is executed by the vehicle, it is the value by which the vehicle executes this movement. In a preferred form, operation (a t ) is a control value such as braking force or brake pressure, for example.
[0016] This method has the steps described below: This method starts with the provision of a model P for predicting the vehicle state (s t+1 ). This model P is designed to predict the subsequent vehicle state (s t ) depending on the vehicle state (s t ) and the operation (a t+1 ). The subsequent vehicle state (s t+1 ) means that when the operation (a t ) is executed by the vehicle, the actual vehicle state (s tIt may be interpreted that this is a vehicle state that occurs directly when executed from ).
[0017] Following that, multiple vehicle states (s0, ..., s t ,…s T ) and multiple actions assigned to each (a0, ..., a t ,…a T At least one data tuple (s0, ..., s) having a sequence of ) t ,…s T a0,…,a t ,…a T The calculation of the vehicle state is performed, and at this time, multiple vehicle states are calculated by the vehicle dynamic control device using model (P) depending on the calculated operation.
[0018] Subsequently, the parameters (θ) of the vehicle dynamic control system are adjusted so as to minimize a cost function (c) that depends on the parameters of the vehicle dynamic control system, which calculates the cost of the recorded trajectory depending on the multiple vehicle states of the data tuple and the calculated multiple actions of each assigned vehicle state. The adjustment of the parameters (θ) of the vehicle dynamic control system can be done for each vehicle state of the data tuple or over the entire sequence of data tuples.
[0019] The parameter (θ) can be adjusted using an optimization algorithm, preferably using gradient descent, and more preferably using back-prop-through time.
[0020] So-called model-based formulas for optimizing vehicle dynamic control systems, based on optimization using relatively accurate models, are known to be able to determine the adjustment of the parameter θ, which has the most information. Alternatively, there are so-called model-free formulas, but these are less effective. This is because model-free formulas do not provide sufficient information when optimizing parameters to obtain a vehicle dynamic control system that controls driving dynamics as best as possible. Therefore, it can be said that the proposed method provides a vehicle dynamic control system with significantly better performance than a manually tuned controller. Another advantage of this formula is that it is scalable, meaning it can handle arbitrarily complex amounts of data and can optimize high-dimensional vehicle dynamic control systems.
[0021] It is proposed that model (P) is a trainable model whose parameterization is trained depending on detected driving maneuvers of the vehicle or another vehicle, or that model (P) is a physical model that describes the driving dynamics of the vehicle, particularly along the longitudinal and lateral axes. The trainable model may be, for example, a mechanical learning system, preferably a neural network. Generally, this trainable model may be a black-box model, e.g., linear or feature-based regression, Gaussian process models, recurrent neural networks (RNNs, LSTMs), (deep) neural networks, or a white-box model, e.g., a (simplified) physical model (Grey-Box-Model) with parameters or a combination thereof.
[0022] Physical models have the advantage of enabling analytical gradients, thereby allowing for more precise adjustments, and thus bringing the parameterization of vehicle dynamic control systems closer to ideal parameterization. Furthermore, since actual measurements are not required, this method can advantageously be performed purely through simulation.
[0023] Furthermore, it is proposed that the trajectory of the vehicle's actual driving maneuvers is detected, and in this case, a modified model (g) is created depending on the detected trajectory and model (P), thereby modifying the output of model (P) so that it roughly matches the detected trajectory. The trajectory can describe a sequence of vehicle states and the operation of the selected actual driving maneuver in each vehicle state. "Roughly" here can be interpreted as meaning that the accuracy obtained by this modification is within the measurement tolerance for the vehicle state to be modified, or within the accuracy obtainable by the size of the modified model using the respective optimization method for creating the modified model, or by using the maximum achievable accuracy of the modified model.
[0024] This additional step of detecting actual driving maneuvers may be performed again after the step of adjusting the parameters of the vehicle dynamic control system, in which case the actual driving maneuvers may be executed by the adjusted vehicle dynamic control system. The trajectories thus detected may be used again to adjust the modified model and another step of calculating at least one data tuple by Model P, and then to adjust the parameters of the vehicle dynamic control system.
[0025] In particular, the modified model is created so as to minimize the difference between the outputs of the modified model and the difference between the received state of the training data and the state predicted by the model.
[0026] If the model (P) is a pre-trained model, its parameterization can be learned in a way that depends on the detected trajectory. For this purpose, optimization of model parameters is suitable, for example, by calculating the maximum likelihood in an open (feedforward) or closed (feedback) control circuit using (stochastic) gradient descent, or by using a maximum posterior stochastic solution for single-step or multi-step model prediction.
[0027] Unexpectedly, using a modified model offers the advantage of requiring only minimal actual vehicle operation, both to optimize the vehicle dynamics control system and to create a trained model.
[0028] Furthermore, it is proposed that model (P) is deterministic and the modified model is time-dependent. In other words, the modified model is state-dependent or time-dependent, in which case time characterizes the time interval elapsed from the start of the recorded trajectory. That is, the modified model calculates the modified values for model P in a time-dependent manner. Unexpectedly, it was found that the best parameterization was obtained from this form of modified model. This time may be a discrete value characterizing the number of actions performed from a predetermined starting point (e.g., the point in time when the vehicle dynamics control system intervenes in the driving dynamics).
[0029] The combination of model P and a modification model configured to correct the output of the model can be understood as a global model for predicting state changes. In other words, the global model of state changes is the superposition of these two models.
[0030] Therefore, the modified model is configured to correct errors in the first model regarding the actual state of the surroundings after the operation has been performed. For example, this model predicts a state that depends on the actual state and the operation. It should be noted that the operation can be determined by the vehicle dynamics control system or, for example, by the driver. The modified model then corrects the predicted state of the model so that the predicted state is as similar as possible to the actual state of the surroundings after the operator has performed the operation for this actual state. In other words, the modified model modifies the output of the first model to obtain a predicted state that is as close as possible to the actual state considered to be the surroundings or the state detected during the operation. Therefore, the modified model modifies the first model to obtain a more accurate state with respect to the surroundings, especially surrounding dynamics.
[0031] In a preferred form, the modification model is time-step dependent and / or dependent on the actual state. Selectively, the modification model is a modification term, which is an extracted modification value determined by the difference between the detected vehicle state of the detected actual operation and the state predicted by the model. The modification model can output discrete modifications that can be directly added to the model's predictions. A special case of the modification model may exist where the modification model outputs time-discrete modification values.
[0032] Furthermore, it is proposed that the modified model is selected by minimizing the magnitude of the difference between the outputs of the modified models, and the magnitude of the difference between the detected vehicle state along the trajectory and the vehicle state predicted by the model. Minimization may be performed via known gradient descent methods.
[0033] Furthermore, it is proposed that multiple different models (P) are provided, in which case the data tuple is randomly selected for one of these different models.
[0034] In this case, preferably, the model risk is modeled, ultimately resulting in stable controller behavior, and consequently, the controller can effectively utilize, for example, time variations. It is also conceivable that one or more models may output risk for predictions of these models, characterizing the risk of those predictions, in which case the risk is taken into consideration in tuning the parameters of the vehicle dynamic control system. The risk may be calculated as follows: 1) Statistical judgments regarding the performance, robustness and safety of the controller, and / or 2) Avoid dangerous / unknown vehicle behaviors during / after optimization, and / or 3) Accelerating learning by conducting additional investigations into previously unknown areas.
[0035] By advertising (model) risk over multiple time steps, the risk of long-term predictions (the distribution of possible future system behaviors) can be obtained. [Methods here include, for example, analytical closure (for simple models, e.g., linear Gaussian models), sampling, numerical integration, moment matching, and linearization (for complex models).]
[0036] Furthermore, it is suggested that various models differ in that they take into account or characterize the various dynamics of each external value or the various dynamics of the vehicle's values.
[0037] Examples of dynamic values include changing road surfaces or various scenarios, such as the changing shape of a road relative to all three possible spatial axis coordinates of the road.
[0038] Furthermore, it is proposed that one data tuple be detected for each model, and in this case, parameter changes are made dependent on all data tuples, because it has been shown that this results in particularly optimal and stable controller behavior.
[0039] Furthermore, it is proposed that the detected vehicle state is filtered using a Kalman filter, in which case the parameterization of the Kalman filter is calculated depending on the vehicle's described trajectory, and this Kalman filter is applied to the detected state.
[0040] Furthermore, the vehicle dynamic control system has a modular controller structure, and it is proposed that when adjusting parameters, these parameters be adjusted so that the changed parameters fall within a preset range. In other words, the controller is divided into modules, each module responsible for a sub-function, such as a PID controller that adjusts actual slip to target slip. Another module is a gain scheduler that adapts PID gains according to driving conditions, ground conditions, and speed. Yet another module can estimate driving conditions based on wheel rotation speed, brake pressure changes, and vehicle speed, etc.
[0041] In addition to parameter optimization being achieved only within the confidence interval, the advantages include the vehicle dynamics control system explaining behavior that does not pose a safety risk, and the fact that the investigation of driving dynamics is limited to meaningful driving conditions.
[0042] Furthermore, it is proposed that the vehicle dynamic control system is a neural network, specifically a radial basis function network.
[0043] Preferably, neural networks are highly flexible in that they can learn complex relationships very well and acquire a wide variety of controller behaviors. RBF networks are particularly preferred because, based on their compact structure, they are especially well-suited for implementation in vehicle controllers.
[0044] Furthermore, it is proposed that after parameter adjustment, the vehicle state is detected while the vehicle is in operation, and the vehicle's actuators are driven and controlled using a vehicle dynamic control device, depending on the detected vehicle state.
[0045] Furthermore, it is proposed that the cost function is a weighted superposition of multiple functions, which characterize the difference between the actual tire slip of the vehicle and the target slip, the distance traveled since the intervention of the vehicle dynamic control system, and the time derivative of the distance traveled.
[0046] Furthermore, it is proposed that the vehicle dynamic control device is an ABS controller that outputs an action characterizing the braking force, and that the physical model includes multiple submodels, each of which is a physical model of a vehicle component.
[0047] The operation can be determined separately for each wheel, for example, or individually for the vehicle's axles, so that the wheels / axles can be controlled individually depending on their respective brake pressures.
[0048] It should be noted that the vehicle dynamic control system may be, for example, an ABS, TCS, or ESP controller, or a combination of these controllers.
[0049] It should be further noted that this method may be used to retrospectively adjust an already optimized vehicle dynamic control system for a first vehicle type / example, a second vehicle type / example, or, for example, when new tires are fitted to the vehicle, for the first vehicle type / example.
[0050] In another aspect, the present invention relates to apparatus and computer programs designed to carry out the above methods, and to machine-readable storage media in which the computer programs are stored. [Brief explanation of the drawing]
[0051] [Figure 1] This is a schematic diagram of one embodiment for controlling a vehicle equipped with a vehicle dynamic control system. [Figure 2] This is a schematic flowchart of a method for parameterizing a vehicle dynamic control system. [Figure 3] This figure shows possible configurations of the training device. [Modes for carrying out the invention]
[0052] Embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0053] Figure 1 specifically shows a vehicle 100 equipped with a control system 40.
[0054] Vehicle 100 may generally be an automobile that is controlled by a driver, or it may be a semi-autonomous or fully autonomous vehicle. In another embodiment, the automobile may be a wheeled vehicle, a tracked vehicle, or a rail vehicle. The automobile may also be a two-wheeled vehicle, such as a bicycle or a motorcycle.
[0055] Preferably, the state of the vehicle is detected at regular time intervals by at least one sensor 30, which may be detected by multiple sensors. The state may also be calculated depending on the detected sensor values. The sensors 30 are preferably acceleration sensors (in the longitudinal direction of the vehicle, but may be 3D sensors on all axles), wheel speed sensors (on all wheels), and rotation rate sensors (around the vertical axis, but may be around all other axes).
[0056] The control system 40 receives a sequence of sensor signals S from the sensor 30 in an optional receiving unit, and the optional receiving unit converts the sequence of sensor signals S into a sequence of pre-processed sensor signals.
[0057] A sensor signal S or a sequence of pre-processed sensor signals is supplied to the vehicle dynamic control device 60 of the control system 40. The vehicle dynamic control device 60 is parameterized by parameters θ provided from a parameter storage device P, which is stored in a preferred form.
[0058] The vehicle dynamic control device 60 calculates an operation, also referred to hereafter as a control signal A, depending on the sensor signal S and its parameter θ, and this control is transmitted to the vehicle's actuator 10. The actuator 10 receives the control signal A, is controlled accordingly, and then performs the corresponding operation. It is also possible that the actuator 10 is designed to convert the control signal A into a direct control signal. For example, when the actuator 10 receives a braking force as a control signal A, the actuator can convert the braking force into a corresponding brake pressure, and the brakes are directly controlled by this brake pressure. In this case, the actuator 10 may be a brake system including the brakes of the vehicle 100. Additionally or selectively, the actuator 10 may be a drive unit or steering unit of the vehicle 100.
[0059] In another preferred embodiment, the control system 40 has one or more processors 45 and at least one machine-readable storage medium 46 in which commands are stored, which, when executed by the processors 45, instruct the control system 40 to perform the method according to the present invention.
[0060] In another embodiment, a display unit 10a is provided in addition to the actuator 10. The display unit 10a is pre-configured to, for example, display the operation of the vehicle dynamic control device 60 and / or to provide a warning that the vehicle dynamic control device 60 will soon be activated.
[0061] The vehicle dynamic control device 60 is generated by a parameterized function a=f(s, θ) that outputs a control signal A depending on the state s of the sensor 30 and / or the sensor signal S. If the vehicle dynamic control device 60 outputs a control signal A for the actuator 10 and the actuator 10 has multiple actuators, then the control signal A may have one control signal for each actuator. The individual actuators may be individual brakes of the vehicle 100.
[0062] In a preferred embodiment, the vehicle dynamic control device 60 is an ABS controller, in which case the control device outputs a braking force or brake pressure as a control signal. In a preferred form, in order to allow individual control of the wheels, the vehicle dynamic control device 60 here outputs brake pressure or braking force for each brake of the wheels or for each axle of the vehicle 100.
[0063] In a preferred form, the vehicle dynamic control device 60 has an interpretable controller structure. This is achieved, for example, by allowing the correct parameter limits to be defined within the control device. This has the advantage that the behavior of the vehicle dynamic control device 60 in each situation can be retrofitted.
[0064] An example of a parameterized function f for the vehicle dynamic control device 60 is as follows: The vehicle dynamic control device 60 having an interpretable controller structure may be generated by a structured vehicle dynamic control device 60 that is structured, for example, as a decision tree.
[0065] To calculate an action using a decision tree, the process proceeds from the root node along the tree. At each node, an attribute is queried (e.g., vehicle state), and a decision is made regarding the selection of the next node based on this attribute. This procedure continues for a long time until a leaf of the decision tree is obtained. Each leaf characterizes one of many possible actions. For example, a leaf might characterize brake pressure increase / decrease. In this example, the parameter θ is the decision threshold, etc.
[0066] The vehicle dynamic control device 60 may be selectively generated by an RBF network (radial basis functions) or a Deep RNN policy.
[0067] It should be noted that the parameterized function f may be any other mathematical function that expresses the state of the vehicle in terms of a control signal, depending on the parameters.
[0068] Figure 2 shows a schematic diagram of a flowchart 20 for parameterizing the vehicle dynamic control device 60 and for selectively parameterizing the subsequent operation of the vehicle dynamic control device 60 within the vehicle 100.
[0069] This method begins in step S21. In this step, driving data of vehicle 100 is collected. The driving data is, for example, a data sequence s0, s1, ..., s that describes the state s of vehicle 100 according to vehicle operation. t ,…s T These driving data preferably include state data s at each point t of the operation. t and operation data a t It is a data tuple containing [the specified data].
[0070] If the vehicle dynamic control device 60 is an ABS controller, the braking process is recorded, for example, by using a known ABS controller or by the driver operating the vehicle 100, and in this case, state data (s t) For example, the following sensor data: vehicle speed v veh , acceleration a veh In a preferred format, the following sensor data for each wheel of vehicle 100: wheel speed v wheel , wheel acceleration a wheel , wheel jerk (English: jerk) wheel It has operation data (a t ) is the braking force selected for each condition, and preferably a value that characterizes the road surface.
[0071] Selectively, driving data is generated for each simulation in which a fictional vehicle performs one or more (braking) operations in a simulated environment.
[0072] Next comes step S22, in which the recorded state data s is partially reconstructed. This is because not all the necessary information about the system's (automatic) state is measured by internal sensors (e.g., vehicle tilt, suspension characteristics, wheel acceleration). For learning and modeling, this latent information needs to be re-collected to enable prediction and optimization. This area is essentially called latent state inference (e.g., hidden Markov models) and is solved by filtering / smoothing algorithms (e.g., Kalman filters).
[0073] In a preferred format, the reconstruction of the vehicle state in step S22 is performed by a Kalman filter.
[0074] Step S23 follows the completion of either step S21 or step S22. In this step, model P is provided. The provision is made either by creating model P after step S21 based on records or by providing a physical model.
[0075] Model P(s t+1 |s t ,a t ) is the vehicle state s at time t. tDepending on and the control signals selected in relation to this, the subsequent vehicle state s at the immediately following time point t+1 t+1 This is a model that predicts [something].
[0076] Preferably, Model P(s t+1 |s t ,a t ) is a first-order physical model. In other words, the physical model describes the physical relationships and the actual vehicle state s t and operation a t The subsequent vehicle status depends on t+1 It includes equations that predict particularly deterministically. Typically, for a vehicle dynamic control device 60 for ABS, the physical model may consist of one or more submodels from the following list of submodels: a first submodel which is the physical model of one wheel of the vehicle 100; a second submodel which describes the center of gravity of the vehicle; a third submodel which is the physical model of the damper; a fourth submodel which is the physical model of the tire; and a fifth submodel which is a multidimensional model of the hydraulic model. It should be noted that this list is not final and other physical characteristics such as tire / brake temperature may be considered.
[0077] It should be noted that, in addition to model P, other formulas can be considered to optimize parameterization. Instead of this model, a so-called model-free reinforcement learning formula or a value-based reinforcement learning formula can also be selected. Then, accordingly, in step S23, a function for value-based reinforcement learning, such as Q, is created based on the record from step S21.
[0078] Step S24 follows step S23. Step 24 may be called "on-policy modification". In this case, a modified model g is created, and this modified model g is a model P(s) relating to the vehicle state. t+1 |s t ,a t The prediction is modified so that the revised prediction is in general agreement with the prediction detected in step S21.
[0079] Vehicle conditions corrected in the preferred format are modified as follows:
number
[0080] The modified model g is obtained from s t and a t In S21, the model P(s) for the vehicle state detected was calculated. t+1 |s t ,a t It is designed to be optimized to output a value equivalent to the error of ).
[0081] Furthermore, the modified model g has the advantage of correcting the shortcomings of model P by comparing it with the actual vehicle behavior.
[0082] To correct the shortcomings of Model P by comparing it with the actual vehicle behavior, the following measures may be taken selectively or additionally. For this purpose, so-called transfer learning may be used, which incorporates previously calculated vehicle states, thereby enabling rapid learning of the model for specific vehicle examples. Multiple different models may also be used, thereby enabling the learning of stable controller behavior through this harmonization.
[0083] Step S23 or step S24 is followed by step S25. Here, multiple rollouts are performed, that is, the actual parameterization θ of the vehicle dynamic control device 60. k Using and, in particular, using model P with additional modified model g, a vehicle dynamic control device is used for operation, and the resulting trajectory, in particular the calculated sequence of vehicle states, is detected.
[0084] It should be noted that, in order to optimize parameterization (using either a model-free or value-based reinforcement learning formula), other formulas besides Model P can be considered. Accordingly, trajectory detection must be adjusted in this rollout step.
[0085] Step S26 follows after step S25 has been performed, or after step S25 has been performed multiple times. In this step, the cost for the trajectory detected in step S25 is evaluated.
[0086] The cost for trajectory can be calculated as follows. In a preferred form, the cost for each proposed operation of the vehicle dynamic control device 60 is calculated. For this purpose, the cost function c(s,a) is calculated for conventional trajectory or actual vehicle state s t Depending on, and the actually selected behavior a t The cost can be calculated depending on the following: The total cost for the trajectory can be accumulated over all operations, that is, over all time points t:
number
[0087] The cost function c(s,a) can be constructed as follows: c(s,a)=α1*mean deceleration+α2*steerability+α3*… In this formula, α n This is a pre-configurable coefficient, which can be set by, for example, an application engineer or incorporated into the initial value. This coefficient is a value between 0 and 1.
[0088] "Steerability" can be interpreted as the controllability of a vehicle. Controllability is calculated based on the lateral force acting on the vehicle (F_lat) (e.g., F_lat, max-F_lat, current), and in some cases, it may also depend on the normalized lateral force: (F_lat,max-F_lat,current) / F_lat,max,non_braking.
[0089] Controllability may be defined as negative if the cost function should be minimized. Additionally or selectively, controllability may be calculated depending on longitudinal forces, so that longitudinal forces are not fully utilized to allow for "play" for lateral forces. For this purpose, a target slip range is defined (e.g., slip ∈ [slip_min, slip_max]), and then the corresponding cost is shown, for example, by a sigmoid function.
[0090] Mean deceleration can be interpreted as the average value of all accelerations in the trajectory.
number
[0091] Other elements of the cost function may be given by the vehicle's respective behaviors, which should be penalized or rewarded. Specifically, these may be: comfort / jerk, i.e., the degree of braking jerk; hardware requirements (how much load braking places on the braking system, vehicle, hydraulics, and tires); performance (e.g., braking distance, acceleration); and straight-line drivability, i.e., behavior around the vertical axis.
[0092] All of these elements can be evaluated using various signals (sensor signals or evaluations) and various cost functions, such as the mean absolute error, mean square error, mean square error, or standard deviation of the signal between the actual state and the target state (e.g., slip, friction value, jerk).
[0093] Another element of the cost function may be the total braking distance. This is the distance from when braking is finished (for example, v veh <v threshold ->c, where c is either the total braking distance or the sum over v*dt for the time step while braking is applied, which may be individual values available first in the time step.
[0094] Other elements of the cost function may be the slip deviation relative to the target slip:||slip-lip_target|| ∧ 2 and / or average acceleration: average std(acceleration).
[0095] Step S27 follows, after the total cost has been calculated in step S26 for a trajectory or multiple trajectories. In this step, the parameters θ of the vehicle dynamic control device 60 are repeatedly adjusted to reduce the total cost. In this case, the optimization may be defined as follows:
number
[0096] This optimization with respect to the parameter θ may be performed via gradient descent over the total cost J.
[0097] Next, for each iteration k of the optimization, the actual parameter θ k The following adjustments are made:
number
[0098] The iteration may continue until a termination criterion is met. The termination criterion is, for example: the maximum number of iterations, J <J threshold Minimum change of parameter θ < θ threshold It is sufficient if it is the minimum change.
[0099] In particular, if multiple total costs are calculated for various operations, the parameters may be adjusted in a batch manner for the multiple total costs. Batch processing may be performed once for model parameters, once for scenarios / operations, or once for sub-trajectories.
[0100] After step S27 is completed, steps S25 to S27 may be executed again, and steps S21 to S27 may be selectively executed again.
[0101] In an optional step S28, the control system 40 of the vehicle 100 is initialized with the vehicle dynamic control device 60 adjusted in step 27.
[0102] In the next optional step S29, the vehicle 100 is driven by a tuned vehicle dynamic control device 60. In this case, when the vehicle dynamic control device 60 is activated in appropriate circumstances, for example when the emergency brake is applied, the vehicle 100 can be controlled by this vehicle dynamic control device 60.
[0103] In another embodiment of the method shown in Figure 2, steps S21 to S27 may be performed again after the completion of step S27, but the vehicle dynamic control device 60 adjusted in step S27 is used in a different vehicle of a different vehicle model or type, and new measurements are performed according to step S21. This allows the optimized vehicle dynamic control device 60 to be retrospectively optimized for another vehicle at a low cost.
[0104] In another embodiment of the method shown in Figure 2, several different models P are provided or generated, and trajectories are generated for each model. This model differs in that it describes various scenarios and / or takes into account external values, such as various road surface dynamics or different driving characteristics of a different vehicle model type or wear of the vehicle's tires / other components. This has the advantage that the vehicle dynamics control system can learn to use temporal changes.
[0105] Figure 3 schematically shows a training device 300 for parameterizing the vehicle dynamic control device 60. The training device 300 has a providing device 31, which provides driving data recorded in step S21 or a simulated environment that produces a state according to the actions performed by the vehicle dynamic control device 60. The data from the providing device 31 is transferred to the vehicle dynamic control device 60, which calculates each action from this data. The data and actions from the providing device 31 are supplied to an evaluation device 33, which calculates adjusted parameters according to step S27, and these parameters are transmitted to a parameter storage device P, where they are replaced with the latest parameters.
[0106] The steps performed by the training device 300 may be implemented as a computer program, stored in a machine-readable storage medium 34, and executed by the processor 35.
[0107] The concept of "computer" includes any device for processing pre-configurable mathematical formulas. These formulas may be provided in software form, in hardware form, or in a hybrid form of software and hardware. [Explanation of Symbols]
[0108] 10 Actuators 10a Display Unit 20 flowcharts, methods 30 sensors 31 Providing device 33 Evaluation device 34 Storage medium 35 processors 40 Control Systems 45 processors 46 Storage medium 60. Vehicle Dynamic Control System 100 vehicles 300 training devices S21~S29 Step
Claims
1. A method (20) for parameterizing a vehicle dynamics control device (60) of a vehicle (100) capable of adjusting the driving dynamics of said vehicle (100), said vehicle dynamics control device (60) being adapted to a vehicle state (s t ) depending on the operation (a t In the method (20) for calculating The vehicle state (s t ) and the operation (a t ) to predict a subsequent vehicle state (s t+1 ), a step (S23) of providing a model P for predicting a vehicle state (s t+1 ) designed to be: A plurality of vehicle states (s 0 , …, s t , … s T ) and a plurality of respectively assigned operations (a 0 , …, a t , … a T ) having at least one data tuple (s 0 , …, s t , … s T ; a 0 , …, a t , … a T ) is calculated, and calculating the plurality of vehicle states by the vehicle dynamic control device using the model (P) depending on the calculated operations, Calculating the cost of the recorded data tuples depending on a plurality of vehicle states of the data tuples and depending on a plurality of calculated operations of the respectively assigned vehicle states, and adjusting the parameter (θ) of the vehicle dynamic control device (60) such that a cost function (c) depending on the parameters of the vehicle dynamic control device (60) is minimized (step S27); A method (20) for parameterizing a vehicle dynamic control device (60) of a vehicle (100) having the above.
2. The method according to claim 1, wherein the model (P) is a machine learning system, and the parameterization of this learning system is trained depending on the detected driving maneuvers of the vehicle (100) or another vehicle, or the model (P) is a physical model that describes the vehicle dynamics, particularly along the longitudinal, lateral and vertical axes of the vehicle.
3. Additionally performing detection (S21) of the trajectory of the actual driving maneuvers of the vehicle (100), and creating a correction model g depending on the detected trajectory and the model (P), whereby the correction model corrects the output of the model (P) such that the output substantially coincides with the detected trajectory. The method according to claim 1 or 2.
4. The method according to claim 3, wherein the model (P) is deterministic and the correction model is time-dependent.
5. Providing a plurality of various models (P), and randomly detecting the data tuples for one of the plurality of various models. The method according to claim 1 or 2.
6. The method according to claim 5, wherein the various models differ in that these models describe the respective various dynamics of external values or the various dynamics of vehicle values.
7. Detecting one data tuple for each of the models, and performing the change of the parameters depending on all the data tuples. The method according to claim 5.
8. Filtering the vehicle state by a Kalman filter (S22). The method according to claim 1 or 2.
9. The method according to claim 1 or 2, wherein the vehicle dynamic control device has a modular controller structure, and when adjusting a parameter, the parameter is adjusted so that the changed parameter is within a preset value range.
10. The method according to claim 1 or 2, wherein the vehicle dynamic control device (60) is a neural network, particularly a radial basis function network.
11. After adjusting the parameter (S27), the vehicle state is detected during operation of the vehicle (100), and the actuator (10) of the vehicle is driven and controlled depending on an operation depending on the detected vehicle state using the vehicle dynamic control device (60). The method according to claim 1 or 2.
12. The method according to claim 2, wherein the vehicle dynamic control device (60) is an ABS controller that outputs an operation characterizing a braking force, and the physical model includes a plurality of partial models, and these partial models are each one physical model of a vehicle component.
13. The method according to claim 1 or 2, wherein the cost function is a weighted overlap of a plurality of functions, and the functions characterize the difference between the actual slip of the vehicle's tire with respect to the target slip, the distance traveled since the vehicle dynamic control device intervened, and the time derivative of the traveled distance.
14. A vehicle dynamic control device (60) obtainable according to claim 9.
15. An apparatus designed to execute the method according to claim 1 or 2.
16. A computer program including instructions for instructing the computer to execute the method according to claim 1 or 2 when the program is executed by the computer.
17. A machine-readable storage medium storing the computer program according to claim 16.