Method for determining a control parameter of a control system

By deriving reward functions and clustering from driving trajectories using machine learning methods, the problem of existing adjustment systems being unable to adapt to individual driving behaviors is solved, enabling drivers to set personalized adjustment parameters and improving the efficiency and comfort of the driving assistance system.

CN112977461BActive Publication Date: 2026-03-31ROBERT BOSCH GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-11
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing adjustment systems cannot be individually adapted to the driving behavior of individual drivers.

Method used

Using machine learning methods, a reward function is derived from driving trajectories through inverse reinforcement learning, forming clusters specific to driver types. Based on these clusters, adjustment parameters are determined to achieve adaptation to the behavior of individual drivers.

Benefits of technology

It achieves personalized adaptation of the adjustment system to individual driver behavior, improving the efficiency and comfort of the driving assistance system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112977461B_ABST
    Figure CN112977461B_ABST
Patent Text Reader

Abstract

The invention relates to a method (200) for determining adjustment parameters (θj) of an adjustment system (100), in particular of a motor vehicle (110), in particular for adjusting a driving operation of a motor vehicle (110), using machine learning, wherein the method (200) comprises: providing (210) a set of driving trajectories (D); deriving (220) a reward function (Rj) from the driving trajectories (D) using an inverse reinforcement learning method; deriving (230) driver type-specific clusters (Cj) based on the reward function (Rj); determining (240) adjustment parameters (θj) for the respective driver type-specific clusters (Cj).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method according to claim 1 for using machine learning to determine the adjustment parameters of an adjustment system, particularly an adjustment system for a motor vehicle, and especially an adjustment system for adjusting the driving operation of a motor vehicle.

[0002] This disclosure also relates to a method for adjusting a motor vehicle using an adjustment system, as described in claim 6.

[0003] This disclosure also relates to an adjustment system according to claim 10. Background Technology

[0004] Adjustment systems are used in motor vehicles, for example, as driver assistance systems, to assist the driver or reduce the driver's workload in certain driving situations.

[0005] To achieve this assistance function, the driver assistance system includes ambient sensors, such as radar sensors, lidar sensors, laser scanners, video sensors, and ultrasonic sensors. If the vehicle is equipped with a navigation system, the driver assistance system can also utilize the data from that system. Furthermore, the driver assistance system, preferably connected to the vehicle's onboard electrical network via at least one bus, preferably a CAN bus, can also actively intervene in onboard systems, such as, in particular, the steering system, braking system, powertrain system, and alarm system.

[0006] Typically, when a control system is available within the fleet, a unified data compilation (Bedatung) for that system is used. If necessary, the control system can also be adapted to Sport or Comfort modes. Individual adaptation to the driving behavior of individual drivers is not currently known.

[0007] Therefore, it is desirable to provide an adjustment system that can achieve this individual adaptation to the driving behavior of an individual driver. Summary of the Invention

[0008] This is achieved by the adjustment system and computer-implemented method as described in the independent claims.

[0009] A preferred embodiment relates to a computer-implemented method for determining regulation parameters of a regulation system, particularly a regulation system for a motor vehicle, and especially a regulation system for regulating the driving operation of a motor vehicle, using machine learning, wherein the method includes:

[0010] Provides a set D of driving trajectories;

[0011] The reward function is derived from the driving trajectory using an inverse reinforcement learning method.

[0012] These reward functions are used to derive driver-type specific clusters;

[0013] Adjustment parameters are determined for the corresponding driver-type-specific clusters.

[0014] During the learning phase, different driver types are clustered based on the set of driving trajectories. The characteristic of clustering is that objects within the same cluster possess similar, especially identical, characteristics, thus distinguishing them from objects in different clusters. Then, in the application phase of the system, the system can individually adapt to the driving behavior of corresponding drivers by selecting specific driver-type-specific clusters. Advantageously, the driving trajectories are based on driving demonstrations by different drivers or driver types.

[0015] A reward function is a function that assigns a reward value to the value of an adjustment amount. Advantageously, the reward function is chosen such that the smaller the deviation between the adjustment amount and the nominal amount, the larger the reward function becomes. According to the present invention, a corresponding reward function is determined for a given driving trajectory, and this reward function is optimized for that driving trajectory.

[0016] The reward function is derived using inverse reinforcement learning methods, such as in the case of using inverse reinforcement learning algorithms. This method and exemplary algorithms are disclosed, for example, at https: / / arxiv.org / pdf / 1712.05514.pdf: Inverse Reinforce Learning with Nonparametric Behavior Clustering, Siddharthan Rajasekaran, Jinwei Zhang, and Jie Fu.

[0017] Next, driver type clustering is derived based on these reward functions.

[0018] The reward function specifically describes the desired state and actions of the corresponding driver. Therefore, the reward function can particularly correspond to the goals and desires of an individual driver, such as maintaining a specific distance from third-party vehicles, acceleration, and speed. Thus, the reward function represents the driver's rational actions and can better generalize the situation as a direct imitation of driving behavior. Generalizations can be advantageously obtained by clustering the reward functions derived from these driving trajectories, and especially not by clustering the driving trajectories themselves.

[0019] In another preferred embodiment, the driving trajectory includes the vehicle's operating data and / or the vehicle's reference data about its surrounding environment, and the reward function takes into account this operating data and / or reference data.

[0020] For example, from the published literature Kuderer, Markus, Shilpa Gulati, and Wolfram Burgard, “Learning driving styles for autonomous vehicles from demonstration.” 2015 IEEE, International Conference on Robotics and Automation (ICRA). IEEE, 2015, it is exemplarily known that features that can affect the reward function, such as, in particular, acceleration, speed, and distance from the lane centerline. Advantageously, in particular, other features, such as distance from third-party vehicles, especially the vehicle in front and / or other vehicles, and the relative speed between the vehicle and third-party vehicles, may also have an effect.

[0021] In another preferred embodiment, it is specified that: for driver type-specific clusters, driving strategies are calculated, especially driver type-specific driving strategies.

[0022] In another preferred embodiment, the adjustment parameters for driver type-specific clusters are specified to be optimized based on the reward function of the corresponding cluster and / or based on vehicle operating data and / or reference data of the vehicle regarding its surrounding environment. Advantageously, these adjustment parameters can be optimized using an optimization function.

[0023]

[0024] In the case of optimization, r is optimized. In the exemplary optimization function shown, r j The reward function θ is described for cluster j. j The controller π describes cluster j θj The adjustment parameters, and It describes the distribution of future states, which include states constituted by the forward model of this vehicle and the behavior of reference objects, especially third-party vehicles, where state x t This includes the state of the vehicle at time point t, as well as the states of reference objects, especially third-party vehicles. The solution to the optimization function identifies the following parameter θ. j Under the given parameters, the reward function is maximized and therefore optimal in relation to the driver's goals and desires extracted in the first step.

[0025] In another preferred embodiment, these adjustment parameters are specified to be optimized for at least one adjustment case. Adjustment cases include controller application cases, such as distance adjustment (Adaptive Cruise Control, ACC) or parking assist or lane keeping assist (LKS).

[0026] Other preferred embodiments relate to a method for regulating a motor vehicle using a regulation system, wherein the method includes: providing a set of driver-type-specific clusters, each driver-type-specific cluster including a reward function and regulation parameters, wherein the driver-type-specific clusters and / or regulation parameters are determined according to a method according to at least one of these embodiments; observing the driver's driving behavior while the motor vehicle is in motion; and identifying driver-type-specific clusters from the set of driver-type-specific clusters based on the observed driving behavior.

[0027] Furthermore, the regulation parameters of the driver-type-specific clusters are used to parameterize the regulation system, and in particular, the model of the regulation system.

[0028] In another preferred embodiment, the identification of clusters includes evaluating driving behavior based on a reward function specific to driver type clusters. Advantageously, a derived reward function is used to identify clusters. Driver behavior, particularly behavior over a specific time period, is evaluated based on a reward function specific to driver type clusters, and specific driver type clusters are selected based on average rewards. Advantageously, the selected driver type-specific clusters are paired with the reward function.

[0029]

[0030] Optimize, where D D This includes the observed common states of both the current vehicle and the vehicle ahead. Correspondingly, driver-type-specific clusters are selected based on driver types that have the most similar goals and demands.

[0031] In another preferred embodiment, the identification of driver-type-specific clusters includes evaluating driving behavior based on the driver's driving strategy. Advantageously, for cluster identification, a driver-type-specific driving strategy, learned in conjunction with an inverse reinforcement learning method, is used. The driver's behavior, particularly within a specific time period, is based on selected driving actions, such as acceleration, braking, steering, etc., and the learned driving strategy π from the inverse reinforcement steps. jThe driving actions are compared and driving type-specific clusters are selected. The most similar, especially identical, driving actions are selected by applying this driving type-specific cluster. Advantageously, the selected driver type-specific clusters are used to evaluate the function.

[0032]

[0033] Optimization was carried out, including It is the driving strategy π formed by observing state-action tuples and driver type-specific clustering. j The distance measure for comparison. Correspondingly, the following driving type-specific clusters are selected, with the most similar, especially identical, driving actions being selected by applying this driving type-specific cluster.

[0034] In another preferred embodiment, the identification of driver-type-specific clusters includes: evaluating driving behavior based on conditioning parameters according to conditioning conditions. Advantageously, conditioning parameters learned by applying an inverse reinforcement learning method are used for cluster identification. Driver behavior, especially behavior within a specific time period, is based on selected driving actions, such as acceleration, braking, steering, etc., and the learned controller policy π for the selected conditioning conditions. θj The driving actions are compared and driving type-specific clusters are selected. The most similar, especially identical, driving actions are selected by applying this driving type-specific cluster. Advantageously, the selected driver type-specific clusters are used to evaluate the function.

[0035]

[0036] Optimization was carried out, including It is the observed state-action tuple and the driving strategy π θj The distance measure for comparison. Correspondingly, the following driving type-specific clusters are selected, with the most similar, especially identical, driving actions being selected by applying this driving type-specific cluster.

[0037] In another preferred embodiment, the adaptive driver model of the adjustment system is parameterized using adjustment parameters of the selected driver-type-specific clusters, and the adjustment system is used to adjust the motor vehicle, especially its driving operation. After one time step, the vehicle and reference objects, especially third-party vehicles, especially the preceding vehicle, have moved, especially relative to each other, and / or continue to move, and the steps of the method for adjusting the motor vehicle using the adjustment system, especially the step of identifying driver-type-specific clusters, and / or parameterizing the adjustment system using the adjustment parameters of the identified driver-type-specific clusters, are re-implemented. Advantageously, this method can be repeatedly implemented during the driving operation of the motor vehicle.

[0038] Other preferred embodiments relate to a regulation system for motor vehicles, particularly for regulating the operation of motor vehicles, the regulation system comprising: an identification module for identifying driver-type-specific clusters according to a method of at least one of these embodiments; and a controller configured to output at least one control quantity when using a model, wherein the model can be parameterized based on the driver-type-specific clusters identified according to the method of at least one of these embodiments.

[0039] In another preferred embodiment, the model is specified to depict the behavior of the motor vehicle and the behavior of the motor vehicle's surrounding environment.

[0040] Other features, applications, and advantages of the invention will become apparent from the subsequent description of embodiments of the invention, illustrated in the accompanying drawings. Hereinafter, all features described or shown, either alone or in any combination, form the subject matter of the invention, regardless of their generalization in the claims or their references thereto, and regardless of their expression or presentation in the specification or drawings. Attached Figure Description

[0041] In the attached diagram:

[0042] Figure 1 A schematic diagram showing the adjustment function of the motor vehicle's adjustment system is provided.

[0043] Figure 2 A schematic diagram illustrates the steps of a method for using machine learning to determine the regulation parameters of a regulation system, particularly a regulation system for a motor vehicle;

[0044] Figure 3 A schematic diagram illustrating the steps of a method for using a regulating system is shown; and

[0045] Figure 4 It shows Figure 3A schematic overview of the methods used. Detailed Implementation

[0046] Figure 1 Taking distance adjustment as an example, a schematic diagram of the adjustment system 100 of a motor vehicle 110, particularly the adjustment function of a driver assistance system, is shown. To achieve the assistance function, the driver assistance system includes ambient environmental sensors, such as radar sensors, lidar sensors, laser scanners, video sensors, and ultrasonic sensors. As long as the vehicle is equipped with a navigation system, the driver assistance system can also utilize the data from that system. Furthermore, the driver assistance system, preferably connected to the vehicle's onboard electrical network via at least one bus, preferably a CAN bus, can also actively intervene in onboard systems, such as, in particular, the steering system, braking system, powertrain system, and alarm system.

[0047] In one design, the control system 100 is configured as a control device for a motor vehicle 110. The control device may include a computer, particularly a microprocessor or calculator. The control device may include a command memory, which the computer can execute.

[0048] The adjustment system 100 of the motor vehicle 110 is configured to output a control quantity u. Based on the control quantity u, the adjustment quantity y of the motor vehicle can be adjusted through an appropriate control process so that the adjustment quantity y is adapted to the command variable w of the adjustment system.

[0049] Typically, in the case of distance adjustment, the vehicle's position is compared to the distance to the vehicle 120 traveling in front, and this distance is adjusted to a pre-defined rated value through targeted acceleration and / or braking intervention. This rated value should be selected such that it is not less than the legally required distance between the two vehicles. On the other hand, this distance should not become so large that the vehicle in front can no longer be reliably detected. For this purpose, a camera- or radar-based sensor system is typically used, and a real-valued value is then generated from it, which is used as input to the adjustment algorithm for distance adjustment. For example, distance adjustment begins when the driver 130 activates the function. The function is either deactivated by the driver 130 or automatically deactivated when, for example, a sudden braking intervention occurs.

[0050] See below for reference. Figures 2 to 4 This will explain how the adjustment system 100 can be adapted to the driving behavior of an individual driver.

[0051] Figure 2 The steps of a method 200 for determining regulation parameters of a regulation system 100 using machine learning are shown. Method 200 includes the following steps:

[0052] Step 210 is used to provide the set D of driving trajectories;

[0053] This is used to derive the reward function R from the driving trajectory D using inverse reinforcement learning. j Step 220;

[0054] Used for reward function R j To derive driver-type specific clustering C j Step 230; and

[0055] Used to determine the corresponding driver-type-specific cluster C j Adjustment parameter θ j Step 240.

[0056] Method 200 illustrates the steps of a learning phase for adjusting system 100. In the learning phase, different driver types C are adjusted based on a set D of driving trajectories. j Clusters are formed. The characteristic of clustering is that objects in the same cluster share similar, especially identical, characteristics, and are thus distinguished from objects in different clusters. Next, in the application phase of the control system 100, the control system 100 can select a specific cluster C specific to the driver type. j This allows for individual adaptation to the driving behavior of the corresponding driver. Advantageously, the driving trajectory D is based on driving demonstrations by different drivers or driver types.

[0057] Reward function R j The reward function, in English, is a function that assigns a reward value to the value of the adjustment amount. Advantageously, the reward function is chosen such that the smaller the deviation between the adjustment amount and the nominal amount, the larger the value of the reward function. According to the present invention, a corresponding reward function r is determined for the corresponding driving trajectory d. j The reward function is optimized for the driving trajectory d.

[0058] The reward function R is derived by using inverse reinforcement learning methods, such as in the case of using an inverse reinforcement learning algorithm. j The method and exemplary algorithms are disclosed, for example, at https: / / arxiv.org / pdf / 1712.05514.pdf: Inverse Reinforce Learning with Nonparametric Behavior Clustering, SiddharthanRajasekaran, Jinwei Zhang, and Jie Fu.

[0059] Next, based on these reward functions R j To derive driver type clustering C j .

[0060] Reward function R j Specifically, it describes the desired state and actions of the corresponding driver. Therefore, the reward function can particularly correspond to the goals and demands of an individual driver, such as maintaining a specific distance of 120 from a third-party vehicle, acceleration, and speed. Thus, the reward function R... j This represents the driver's rational actions and can be better summarized as a direct imitation of driving behavior. The reward function R is derived from these driving trajectories D. j Clustering of the trajectories themselves, rather than clustering of the trajectories themselves, can advantageously yield generalized results.

[0061] In another preferred embodiment, the driving trajectory includes the vehicle's operating data and / or reference data about the vehicle's surrounding environment, and the reward function takes into account this operating data and / or reference data. The vehicle's operating data includes, for example, speed, acceleration, steering angle, and tilt. The vehicle's surrounding environment data includes, for example, information about road conditions, weather, lane gradient, road direction, etc.

[0062] For example, features that can influence the reward function, such as, in particular, acceleration, speed, and distance from the lane centerline, are exemplarily known from the published literature Kuderer, Markus, Shilpa Gulati, and Wolfram Burgard, “Learning driving styles for autonomous vehicles from demonstration.”, 2015 IEEE, International Conference on Robotics and Automation (ICRA), IEEE, 2015. Advantageously, other features, such as distance from third-party vehicle 120, especially the distance to the preceding vehicle and / or other vehicles, and the relative speed between motor vehicle 110 and third-party vehicle 120, may also have an effect.

[0063] In another preferred embodiment, it is specified that: cluster C is specific to driver type. j Calculate driving strategies, especially driver-type-specific driving strategies.

[0064] In another preferred embodiment, a driver-type-specific cluster C is specified. j Adjustment parameter θ j According to the corresponding clustering C j Reward function R jAnd / or based on the operating data of the motor vehicle 110 and / or reference data about the surrounding environment of the motor vehicle 110. Advantageously, these adjustment parameters can be optimized using an optimization function.

[0065]

[0066] In the case of optimization, r is optimized. In the exemplary optimization function shown, r j The reward function θ is described for cluster j. j The controller π describes cluster j θj The adjustment parameters, and It describes the distribution of future states, which include states constituted by the forward model of this vehicle and the behavior of reference objects, especially third-party vehicles, where state x t This includes the state of the vehicle at time point t, as well as the states of reference objects, especially third-party vehicles. The solution to the optimization function identifies the following parameter θ. j Under the given parameters, the reward function r j It is optimal in terms of the driver's goals and demands extracted in the first step.

[0067] In another preferred embodiment, these adjustment parameters θ are specified. j Optimization is provided for at least one adjustment scenario. Adjustment scenarios include controller application scenarios, such as distance adjustment (Adaptive Cruise Control, ACC) or parking assist or lane keeping assist (LKS).

[0068] Figure 3 The steps of a method 300 for adjusting a motor vehicle 110 using an adjustment system 100 are shown.

[0069] Method 300 includes the following steps:

[0070] Used to provide driver type-specific clustering C j Step 310 of the set, the corresponding driver-type-specific cluster c j Including the reward function r j and adjusting parameter θ j Among them, driver-type specific cluster C j and / or adjust parameter θ j It is determined according to method 200 of the above-described embodiment;

[0071] Step 320 for observing the driving behavior of driver 130 during the operation of motor vehicle 110;

[0072] Used to select driver-type specific clusters C based on observed driving behavior j The set identifies clusters c that are specific to driver type. j Step 330;

[0073] and for utilizing the identified driver-type-specific clusters c j Adjustment parameter θ j Step 340 is to parameterize the control system 100, and especially the model of the control system 100.

[0074] In another preferred embodiment, it is specified that: for cluster c j The identifier 330 includes: evaluating driving behavior based on a reward function for driver type-specific clusters. Advantageously, a derived reward function is used to identify the clusters. Driver behavior, especially behavior over a specific time period, is evaluated based on the reward function for driver type-specific clusters, and specific clusters for driver type are selected based on average rewards. Advantageously, the selected driver type-specific clusters are paired with the function.

[0075]

[0076] Optimize, where D D This includes the observed common states of both the current vehicle and the vehicle ahead. Correspondingly, driver-type-specific clusters are selected based on driver types that have the most similar goals and demands.

[0077] In another preferred embodiment, it is specified that: clusters c specific to driver type are defined. j The labeling 330 includes: evaluating driving behavior based on the driver's driving strategy. Advantageously, for labeling clusters, a driving strategy, particularly driver-type specific, learned with the application of an inverse reinforcement learning method is used. The driver's behavior, especially behavior within a specific time period, is based on selected driving actions, such as acceleration, braking, steering, etc., and the learned driving strategy π after the inverse reinforcement steps. j The driving actions are compared and driving type-specific clusters are selected. The most similar, especially identical, driving actions are selected by applying this driving type-specific cluster. Advantageously, the selected driver type-specific clusters are used to evaluate the function.

[0078]

[0079] Optimization was carried out, including It is the driving strategy π formed by observing state-action tuples and driver type-specific clustering. jThe distance measure for comparison. Correspondingly, the following driving type-specific clusters are selected, with the most similar, especially identical, driving actions being selected by applying this driving type-specific cluster.

[0080] In another preferred embodiment, it is specified that: clusters c specific to driver type are defined. j The identifier 330 includes: based on the adjustment parameter θ, depending on the adjustment situation. j This is used to evaluate driving behavior. Advantageously, to label clusters, a modulating parameter θ learned in the context of applying an inverse reinforcement learning method is used. j The driver's behavior, especially within a specific time period, is based on the learned controller strategy π, which relates to the selected driving actions, such as acceleration, braking, steering, etc., and the chosen adjustment conditions. θj The driving actions are compared and driving type-specific clusters are selected. The most similar, especially identical, driving actions are selected by applying this driving type-specific cluster. Advantageously, the selected driver type-specific clusters are used to evaluate the function.

[0081]

[0082] Optimization was carried out, including It is the observed state-action tuple and the driving strategy π θj The distance measure for comparison. Correspondingly, the following driving type-specific clusters are selected, with the most similar, especially identical, driving actions being selected by applying this driving type-specific cluster.

[0083] Figure 4 It shows Figure 3 A schematic overview of the methods used.

[0084] The regulation system 100 includes a model M, specifically an adaptive driver model. It utilizes a selected driver-type-specific clustering c. j Adjustment parameter θ j The model M is parameterized, and the adjustment system is used to adjust the driving operation of motor vehicles, especially motor vehicle 110.

[0085] x t 110 This describes the state of vehicle 110 at time point t, and x t 120 The state of the reference object, especially the third-party vehicle 120, is described at time t.

[0086] The adjustment system 100 also includes an identification module 140 for identifying driver type-specific clusters cj. This identification is performed at time point t according to the method 300 described above, as shown in the embodiment. The driver type-specific cluster cj selected at time point t is then used. t Adjustment parameter θ j The model M is parameterized, and the adjustment system is used to adjust the driving operation of motor vehicles, especially motor vehicle 110.

[0087] After one time step, for example at time point t+1, vehicle 110 and reference object 120, especially third-party vehicles, especially the preceding vehicle, have moved, especially moved relative to each other, and / or continue to move forward. At time point t+1, the steps of the method for regulating the motor vehicle using the regulation system are re-implemented, especially the steps for identifying driver-type-specific clusters and / or parameterizing the regulation system using the regulation parameters θj of the identified driver-type-specific clusters. Advantageously, method 300 can be repeatedly implemented during the operation of motor vehicle 110.

Claims

1. A method (200) for determining adjustment parameters (θj) of an adjustment system (100) of a motor vehicle (110) using machine learning, wherein the method (200) comprises: providing (210) a set of driving trajectories (D) of the motor vehicle (110); deriving (220) a reward function (Rj) from the driving trajectories (D) using an inverse reinforcement learning method, wherein the reward function describes a state and an action desired by a respective driver; deriving (230) driver type specific clusters (Cj) based on the reward function (Rj); determining (240) adjustment parameters (θj) for a respective driver type specific cluster (Cj).

2. The method (200) according to claim 1, wherein the adjustment system (100) is an adjustment system (100) for adjusting a driving operation of the motor vehicle (110).

3. The method (200) according to claim 1, wherein the driving trajectories (D) comprise operational data of the motor vehicle (110) and / or reference data of the motor vehicle (110) with respect to a surrounding of the motor vehicle, and the reward function (Rj) takes the operational data and / or reference data into account.

4. The method (200) according to any one of the preceding claims, wherein a driving strategy is calculated for a driver type specific cluster (Cj).

5. The method (200) according to claim 4, wherein the driving strategy is a driver type specific driving strategy.

6. The method (200) according to any one of the preceding claims 1 to 3, wherein the adjustment parameters (θj) of a driver type specific cluster (Cj) are optimized depending on the reward function (Rj) of the respective driver type specific cluster (Cj) and / or depending on operational data of the motor vehicle (110) and / or reference data of the motor vehicle (110) with respect to a surrounding of the motor vehicle (110).

7. The method (200) according to any one of the preceding claims 1 to 3, wherein the adjustment parameters (θj) are optimized for at least one adjustment situation.

8. A method (300) for adjusting a motor vehicle (110) with an adjustment system (100), wherein the method (300) comprises: providing (310) a set of driver type specific clusters (Cj), a respective driver type specific cluster (Cj) comprising a reward function (Rj) and adjustment parameters (θj), wherein the driver type specific clusters (Cj) and / or the adjustment parameters (θj) are determined according to the method (200) of any one of claims 1 to 7; observing (320) a driving behavior of a driver (130) during a driving operation of the motor vehicle (110); identifying (330) a driver type specific cluster (Cj) from the set of driver type specific clusters (Cj) based on the observed driving behavior; and parameterizing (340) the regulation system (100) using the identified driver type specific cluster (Cj) specific regulation parameter (θj).

9. The method (300) of claim 8, wherein parameterizing (340) the regulation system (100) comprises parameterizing (340) a model (M) of the regulation system (100).

10. The method (300) of claim 8, wherein the identifying (330) of a cluster (cj) specific to a driver type comprises: evaluating the driving behavior based on a reward function (Rj) of the driver type specific cluster (Cj).

11. The method (300) according to any one of claims 8 to 10, wherein the identifying (330) of a cluster (cj) specific to a driver type comprises: evaluating the driving behavior based on a driving strategy of the driver (130).

12. The method (300) of claim 11, wherein the driving strategy is a driver type specific driving strategy.

13. The method (300) according to any one of claims 8 to 10, wherein the identifying (330) of a cluster (cj) specific to a driver type comprises: evaluating the driving behavior based on the regulation parameter (θj) depending on a regulation situation.

14. A conditioning system (100) for a motor vehicle (110), the conditioning system comprising: an identification module (140) for identifying a driver type specific cluster (Cj) according to any one of claims 8 to 13; and a controller configured to output at least one control quantity (u) using a model (M), wherein the model (M) is parameterizable according to a driver type specific cluster identified according to any one of claims 8 to 13.

15. The regulation system of claim 14, wherein the regulation system (100) is set up for regulating a driving operation of the motor vehicle (110).

16. The regulation system of claim 14 or 15, wherein the model (M) depicts a behavior of the motor vehicle (110) and a behavior of a surrounding of the motor vehicle (110).

Citation Information

Patent Citations

  • Driver behavior modeling method based on reverse reinforcement learning

    CN108819948A